Most guides stop at “install Telegraf, import a dashboard, done”. That works on a lab with twelve VMs and falls apart at four hundred. This build walks the full path from a vSphere performance counter to a rendered pixel, explains the statistics model underneath it, and calls out every decision point where you are trading accuracy against load on vCenter.
Level: intermediate to advanced · Stack: vCenter 7 or 8, Telegraf, InfluxDB, Grafana · Build time: about 90 minutes
Contents
- What we are building
- Architecture and data flow
- How vCenter actually exposes metrics
- The design decisions that matter
- Prerequisites and sizing
- Build, step by step
- From counter to pixel
- Panels worth building
- Use cases
- Alerting without the noise
- Scaling, and what breaks first
01. What we are building
The goal is a self hosted metrics pipeline that pulls performance counters out of vCenter on a schedule, stores them as time series, and renders them in Grafana with alerting attached. Nothing is installed on ESXi hosts and nothing is installed on the vCenter Server Appliance. The collector talks to the vSphere Web Services API over HTTPS on port 443 with a read only service account, exactly the same interface the vSphere Client uses.
| Component | Job | Why this one |
|---|---|---|
| vCenter Server | Source of truth for inventory and performance counters | It already aggregates every host and VM. Polling hosts directly means N connections and no cluster level rollups. |
Telegraf with the vsphere input plugin | Discovers inventory, batches QueryPerf calls, converts to line protocol | Written against govmomi, understands real time versus historical intervals, and respects the vCenter query limit automatically. |
| InfluxDB | Time series storage, retention and downsampling | Tag based model fits vSphere inventory cleanly. Prometheus is a valid alternative, covered in decision 03. |
| Grafana | Query, visualise, alert | Panel level transformations let you do the counter arithmetic without pre computing it at ingest. |
Scope. This pipeline covers performance data. Configuration and state data (snapshot age, VM tags, orphaned VMDKs, licence expiry, certificate expiry) does not live in performance counters and needs a second path, covered in step 09.
02. Architecture and data flow
SOURCE
ESXi hosts vCenter Server Storage array
20s samples ----> /sdk Web Services API REST API
kept 1 hour rollups 5m / 30m / 2h / 1d array side latency
| |
+--------------------------+-------------+ |
| | | |
COLLECT v v v v
Telegraf realtime Telegraf historical Telegraf http input
interval 60s interval 300s interval 60s
hosts + virtual VMs datastores, DC array metrics
| | |
+--------------------------+--------------------------+
|
STORE v
InfluxDB bucket: vsphere
vsphere_vm_cpu, vsphere_host_datastore, ...
tags: vcenter, dcname, clustername, vmname, moid
|
PRESENT v
Grafana
dashboards, template variables, transformations
unified alerting, contact points
The pipeline splits deliberately into two collector instances because vCenter serves real time and historical counters from different places at different granularities. Merging them into one polling loop is the single most common cause of a pipeline that half works.
The flow in words
- ESXi hosts sample their own counters every 20 seconds and keep roughly one hour of that detail.
- vCenter pulls those samples and writes rolled up copies into its database at 5 minute, 30 minute, 2 hour and 1 day granularity, governed by the statistics level configured per interval.
- Telegraf logs into
https://vcenter/sdk, walks the inventory tree to discover managed object references, then issues batchedQueryPerfcalls for the counters you asked for. - The plugin turns each returned sample into line protocol, one measurement per resource type and counter group, and writes it to InfluxDB with inventory tags attached.
- Grafana queries InfluxDB. Any unit conversion happens in the query or a panel transformation, not at ingest.
- Alert rules evaluate the same queries on a schedule and fire into contact points.
03. How vCenter actually exposes metrics
Skip this section and you will spend a weekend wondering why datastore latency has gaps and cluster CPU sometimes returns nothing. Three properties of the vSphere statistics model drive almost every design decision later on.
Real time versus historical
Real time counters exist only for ESXi hosts and virtual machines, at 20 second granularity, retained about an hour. Datastores, clusters, resource pools and datacenters have no real time series at all. Their smallest available granularity is the 5 minute historical rollup, served out of the vCenter database.
| Interval | Granularity | Typical retention | Applies to |
|---|---|---|---|
| Real time | 20 seconds | About 1 hour, on the host | Hosts, virtual machines |
| Past day | 5 minutes | 1 day | All object types |
| Past week | 30 minutes | 1 week | All object types |
| Past month | 2 hours | 1 month | All object types |
| Past year | 1 day | 1 year | All object types |
Consequence. Polling datastore or cluster metrics on a 20 or 60 second loop does not give you faster data. It gives you the same 5 minute value repeated, plus real query load on the vCenter database. Poll historical resource types at 300 seconds and nothing faster.
Statistics levels
Each historical interval has a statistics level from 1 to 4. Level 1 gives you the basic average counters. Device level and per instance detail (per vCPU, per vmnic, per virtual disk) only appears at level 3 and above. If a counter you configured never shows up in InfluxDB, the usual cause is that the level for that interval is too low, not that the plugin is broken.
Raising levels globally is expensive. The vCenter database grows quickly and the rollup jobs get heavier. Raise level 1 to level 2 or 3 only if you genuinely need per device breakdowns in the longer intervals, and leave levels 2 through 4 alone.
The query limit
vCenter caps how many metrics a single QueryPerf call may request, through the advanced setting config.vpxd.stats.maxQueryMetrics. Per the Telegraf plugin documentation, the value defaults to 64 on vSphere 5.5 and earlier and to 256 on more recent versions. The plugin always checks this setting and automatically reduces its own batch size when the vCenter limit is lower than the configured value, logging a warning when it does.
Cluster metrics make this worse. Cluster counters are aggregated from host and virtual machine metrics, and when the most recent values are not already available vCenter performs that aggregation on the fly. Every internal subquery needed for that aggregation counts towards the same limit, so even a small query can fail with a fault about vpxd.stats.maxQueryMetrics.
The documented remedies are to raise config.vpxd.stats.maxQueryMetrics above the total number of virtual machines managed by that vCenter, or to exclude cluster metrics entirely and derive cluster level figures by aggregating host data in the query layer instead. The second option is the one I prefer, and decision 05 explains why.
04. The design decisions that matter
These are the choices that determine whether the pipeline survives contact with a production estate. Each is written as the option set, the choice, and the cost of that choice, because there is no universally right answer and you should be able to argue yours.
Decision 01. Pull from vCenter, or push from hosts
- Options. Poll the vCenter API, poll each ESXi host directly, or run an agent inside every guest.
- Choice. Poll vCenter.
- Why. One credential, one endpoint, inventory relationships already resolved, and cluster and datastore objects that simply do not exist at the host level. Host direct polling scales linearly in connections and loses every rollup.
- Cost. vCenter becomes a single point of failure for monitoring, and your collection load lands on the appliance that also runs DRS, HA and the client. Budget for it explicitly rather than discovering it.
Decision 02. Split the collector by interval class
- Options. One Telegraf input block for everything, or two input blocks with different intervals.
- Choice. Two separate input blocks, ideally two separate config files.
- Why. Real time objects want a 60 second loop. Historical objects want 300 seconds. A single block forces one interval on both, so you either hammer the vCenter database or you blur your VM data.
- Cost. Two config files to keep in sync, two sets of credentials, two failure surfaces to monitor. Worth it every time.
Decision 03. InfluxDB or Prometheus
- Options. Telegraf into InfluxDB, or a govmomi based exporter scraped by Prometheus.
- Choice. InfluxDB for vSphere specifically.
- Why. vSphere counters arrive with timestamps that are not the scrape time, especially for historical rollups. Prometheus assigns the scrape timestamp to a sample, which silently shifts 5 minute rollup data. Telegraf preserves the sample timestamp from vCenter. The push model also survives a collector that takes 90 seconds to walk a large inventory, where a scrape would time out.
- Cost. Friction if your organisation is standardised on Prometheus. The workable compromise is Telegraf with a Prometheus remote write output, keeping the timestamp handling but landing in the platform you already run.
Decision 04. Which counters, and therefore how much cardinality
- Options. Wildcard everything, or an explicit include list per resource type.
- Choice. Explicit include list. Never wildcard virtual machine counters.
- Why. Series count is roughly objects multiplied by counters multiplied by instances. Two thousand VMs with 40 wildcard counters and per device instances runs into millions of series, and the pain lands on query latency months later when the dashboard is already business critical.
- Cost. You will occasionally need a counter you did not collect, and it is not retroactive. Keep the include list in version control with a comment per counter explaining which panel uses it.
Decision 05. Collect cluster metrics, or aggregate host metrics at query time
- Options. Enable
cluster_metric_include, or exclude clusters and sum host series in Grafana. - Choice. Exclude cluster metrics, aggregate in the query layer.
- Why. Cluster counters are the main trigger for the maxQueryMetrics fault. Summing host CPU and memory in Flux or InfluxQL gives the same numbers without loading vCenter, and it stays correct when hosts move between clusters.
- Cost. Slightly more complex queries, and cluster level counters that have no host equivalent are lost. If you need those, raise the vCenter limit above your VM count and accept the load.
Decision 06. Credential scope
- Options. Reuse an administrator account, or create a dedicated read only service account.
- Choice. Dedicated account with the built in Read only role, assigned at the vCenter root object with propagation enabled.
- Why. Performance queries need nothing beyond read. Root with propagation is required because the plugin discovers inventory from the top, and a role scoped to one datacenter produces confusing partial discovery rather than a clean permission error.
- Cost. One more identity to rotate. Put the password in an environment file the collector reads, never inline in the config, and set the account not to expire.
Decision 07. TLS verification
- Options. Set
insecure_skip_verify = true, or trust the vCenter certificate authority properly. - Choice. Trust the CA. Export the vCenter root certificate and install it in the collector trust store.
- Why. Skipping verification is the default in every tutorial and it means your monitoring credentials will be handed to anything that answers on that address. It is a read only account, but it is still an authenticated path into vCenter.
- Cost. Five extra minutes at build time, and a certificate to rotate. Every guide skips this and every audit flags it.
Decision 08. Retention and downsampling
- Options. Keep raw samples forever, or keep raw briefly and downsample for the long view.
- Choice. Raw for 30 days, hourly rollups for 13 months.
- Why. Troubleshooting needs full resolution and happens within days. Capacity planning needs trend, not detail, and 13 months lets you compare against the same month last year.
- Cost. A downsampling task to write and monitor, and dashboards that must pick the right bucket based on the selected time range.
Decision 09. Where the unit maths lives
- Options. Convert counters at ingest with a Telegraf processor, or convert in the Grafana query.
- Choice. Convert in Grafana.
- Why. The raw counter stays raw and auditable, so when someone disputes a CPU ready number you can show the summation value and the arithmetic. Converting at ingest bakes an assumption about the sample interval into stored data, and that assumption breaks the day you change the collection interval.
- Cost. The same expression repeated across panels. Solve it with a library panel, not by pre computing at ingest.
Decision 10. Correlating with the storage array
- Options. vSphere latency only, or vSphere latency alongside array side latency on the same time axis.
- Choice. Pull array metrics into the same bucket, tagged with the datastore name.
- Why. This is the single highest value thing you can add. Guest latency minus array latency tells you whether the delay is in the array, the fabric or the host queue, and that answer normally takes three teams and a bridge call to reach.
- Cost. You have to get the tag values to match exactly across two systems that name things differently. Do the normalisation in a Telegraf rename or override processor, once, at ingest.
05. Prerequisites and sizing
- vCenter Server 7.0 or 8.0, reachable on TCP 443 from the collector.
- A Linux VM for the monitoring stack, kept off the vCenter appliance itself. For up to 500 VMs, 4 vCPU, 8 GB RAM and 100 GB of fast disk is comfortable. Above 2000 VMs, separate InfluxDB onto its own host with 16 GB and SSD backed storage.
- Docker and the compose plugin, or native packages if you prefer.
- A vCenter service account, created in step 01.
- NTP working on the collector. Time series with skewed clocks look like data loss.
Placement. Put the monitoring VM on a different cluster and a different datastore from the workloads you are watching, if you have one. A monitoring stack that goes down with the thing it monitors tells you nothing during the outage you built it for.
06. Build, step by step
Step 01. Create the read only vCenter service account
Create the user in vCenter Single Sign On, then assign the built in Read only role at the root vCenter Server object with Propagate to children ticked. Assigning at a datacenter or cluster instead produces partial discovery that is hard to diagnose.
# PowerCLI
Connect-VIServer -Server vcenter.lab.local
New-VIPermission -Entity (Get-Folder -NoRecursion) `
-Principal 'LAB\svc-grafana' `
-Role 'ReadOnly' `
-Propagate $true
# verify it landed at the root
Get-VIPermission | Where-Object { $_.Principal -like '*svc-grafana*' } |
Format-Table Entity, Role, Propagate
Step 02. Raise the query limit, or plan to live without cluster metrics
In the vSphere Client go to the vCenter object, Configure, Advanced Settings, and check config.vpxd.stats.maxQueryMetrics. If the setting is absent it is running at the default. Following decision 05 I leave it alone and exclude cluster metrics. If you do need cluster counters, set it above your total VM count.
# check the current value
Get-AdvancedSetting -Entity $global:DefaultVIServer `
-Name 'config.vpxd.stats.maxQueryMetrics'
# raise it only if you are collecting cluster metrics
Get-AdvancedSetting -Entity $global:DefaultVIServer `
-Name 'config.vpxd.stats.maxQueryMetrics' |
Set-AdvancedSetting -Value 4096 -Confirm:$false
Do not set this to -1. Removing the cap entirely lets a single malformed query ask vCenter for everything at once. Pick a number tied to your inventory size and revisit it when the estate grows.
Step 03. Trust the vCenter certificate
Download the root certificate bundle from the vCenter certs endpoint, extract the .0 file for Linux, and install it. This is what lets you leave insecure_skip_verify at false.
curl -sk -o /tmp/vc-certs.zip https://vcenter.lab.local/certs/download.zip
unzip -o /tmp/vc-certs.zip -d /tmp/vc-certs
# the lin folder holds the OpenSSL hashed certificate
cp /tmp/vc-certs/certs/lin/*.0 /usr/local/share/ca-certificates/vcenter-root.crt
update-ca-certificates
# confirm the chain now validates
openssl s_client -connect vcenter.lab.local:443 -brief < /dev/null
Step 04. Stand up InfluxDB and Grafana
One compose file, three services. The CA bundle is mounted into the Telegraf container so it can validate the vCenter certificate.
# docker-compose.yml
services:
influxdb:
image: influxdb:2.7
container_name: influxdb
restart: unless-stopped
ports:
- "8086:8086"
volumes:
- influx-data:/var/lib/influxdb2
- influx-conf:/etc/influxdb2
environment:
DOCKER_INFLUXDB_INIT_MODE: setup
DOCKER_INFLUXDB_INIT_USERNAME: admin
DOCKER_INFLUXDB_INIT_PASSWORD: ${INFLUX_PASSWORD}
DOCKER_INFLUXDB_INIT_ORG: lab
DOCKER_INFLUXDB_INIT_BUCKET: vsphere
DOCKER_INFLUXDB_INIT_RETENTION: 30d
DOCKER_INFLUXDB_INIT_ADMIN_TOKEN: ${INFLUX_TOKEN}
telegraf:
image: telegraf:1.34
container_name: telegraf
restart: unless-stopped
depends_on:
- influxdb
env_file:
- ./vsphere.env
volumes:
- ./telegraf.d:/etc/telegraf/telegraf.d:ro
- ./telegraf.conf:/etc/telegraf/telegraf.conf:ro
- /etc/ssl/certs:/etc/ssl/certs:ro
grafana:
image: grafana/grafana:11.6.0
container_name: grafana
restart: unless-stopped
depends_on:
- influxdb
ports:
- "3000:3000"
volumes:
- grafana-data:/var/lib/grafana
- ./provisioning:/etc/grafana/provisioning:ro
environment:
GF_SECURITY_ADMIN_PASSWORD: ${GRAFANA_PASSWORD}
GF_USERS_ALLOW_SIGN_UP: "false"
volumes:
influx-data:
influx-conf:
grafana-data:
# vsphere.env, kept out of version control
VSPHERE_URL=https://vcenter.lab.local/sdk
VSPHERE_USER=svc-grafana@lab.local
VSPHERE_PASSWORD=change-me
INFLUX_TOKEN=paste-the-admin-token
INFLUX_ORG=lab
INFLUX_BUCKET=vsphere
# then
chmod 600 vsphere.env
docker compose up -d
Step 05. Configure the real time collector
Hosts and virtual machines, 60 second loop, explicit counter list. Note that datastore, cluster and datacenter are all excluded here because they belong to the historical block.
# telegraf.d/vsphere-realtime.conf
[[inputs.vsphere]]
interval = "60s"
vcenters = ["${VSPHERE_URL}"]
username = "${VSPHERE_USER}"
password = "${VSPHERE_PASSWORD}"
## certificate is trusted at the OS level, see step 03
insecure_skip_verify = false
## ---- virtual machines ----
vm_metric_include = [
"cpu.demand.average", # what the VM wants
"cpu.usagemhz.average",
"cpu.usage.average",
"cpu.ready.summation", # contention, the headline counter
"cpu.costop.summation", # oversized vCPU tell
"cpu.latency.average",
"mem.active.average", # true working set
"mem.consumed.average",
"mem.granted.average",
"mem.vmmemctl.average", # ballooning
"mem.swapinRate.average", # host is out of memory
"mem.swapoutRate.average",
"mem.usage.average",
"net.bytesRx.average",
"net.bytesTx.average",
"net.droppedRx.summation",
"net.droppedTx.summation",
"virtualDisk.read.average",
"virtualDisk.write.average",
"virtualDisk.numberReadAveraged.average",
"virtualDisk.numberWriteAveraged.average",
"virtualDisk.totalReadLatency.average",
"virtualDisk.totalWriteLatency.average",
"sys.uptime.latest",
]
## ---- esxi hosts ----
host_metric_include = [
"cpu.usage.average",
"cpu.usagemhz.average",
"cpu.ready.summation",
"cpu.utilization.average",
"mem.usage.average",
"mem.active.average",
"mem.totalCapacity.average",
"mem.vmmemctl.average",
"mem.swapused.average",
"net.bytesRx.average",
"net.bytesTx.average",
"net.droppedRx.summation",
"net.droppedTx.summation",
"disk.maxTotalLatency.latest", # the number that matters
"disk.numberReadAveraged.average",
"disk.numberWriteAveraged.average",
"power.power.average",
"sys.uptime.latest",
]
## historical objects belong to the other input block
datastore_metric_exclude = ["*"]
cluster_metric_exclude = ["*"]
datacenter_metric_exclude = ["*"]
resource_pool_metric_exclude = ["*"]
vsan_metric_exclude = ["*"]
## batching and concurrency, see decision 04
max_query_objects = 256
max_query_metrics = 256
collect_concurrency = 3
discover_concurrency = 3
object_discovery_interval = "300s"
timeout = "60s"
Tuning. The plugin documentation suggests a rule of thumb of setting the concurrency parameters to the number of virtual machines divided by 1500, rounded up. Start at 3, watch the gather duration, and raise it only if collection is not finishing inside the interval.
Step 06. Configure the historical collector
Datastores only, 300 second loop, clusters excluded per decision 05. A generous timeout matters here because the vCenter database is the bottleneck, not the network.
# telegraf.d/vsphere-historical.conf
[[inputs.vsphere]]
interval = "300s"
vcenters = ["${VSPHERE_URL}"]
username = "${VSPHERE_USER}"
password = "${VSPHERE_PASSWORD}"
insecure_skip_verify = false
vm_metric_exclude = ["*"]
host_metric_exclude = ["*"]
cluster_metric_exclude = ["*"] # aggregate hosts in Grafana instead
vsan_metric_exclude = ["*"]
datastore_metric_include = [
"disk.capacity.latest",
"disk.used.latest",
"disk.provisioned.latest",
"disk.capacity.provisioned.average",
"disk.capacity.usage.average",
"datastore.numberReadAveraged.average",
"datastore.numberWriteAveraged.average",
"datastore.totalReadLatency.average",
"datastore.totalWriteLatency.average",
]
datacenter_metric_include = ["*"]
max_query_objects = 256
collect_concurrency = 2
timeout = "240s"
force_discover_on_init = true
Validate before restarting the service. The test flag runs one collection cycle and prints line protocol to stdout without writing anything. If you see line protocol here, vCenter connectivity, permissions and TLS are all correct.
docker compose exec telegraf telegraf \
--config /etc/telegraf/telegraf.conf \
--config-directory /etc/telegraf/telegraf.d \
--input-filter vsphere --test | head -40
Step 07. Provision the Grafana datasource
Provision it as a file rather than clicking through the UI, so the whole stack rebuilds from the repository.
# provisioning/datasources/influx.yml
apiVersion: 1
datasources:
- name: InfluxDB-vSphere
type: influxdb
access: proxy
url: http://influxdb:8086
jsonData:
version: Flux
organization: lab
defaultBucket: vsphere
tlsSkipVerify: false
timeInterval: "60s"
secureJsonData:
token: ${INFLUX_TOKEN}
isDefault: true
The timeInterval value tells Grafana the minimum meaningful resolution of the data, which stops it drawing misleading interpolation between sparse points.
Step 08. Import a baseline dashboard, then replace it
Community vSphere dashboards for the Telegraf plugin are published on the Grafana dashboard library and on the InfluxData integration page. Import one to confirm data is flowing end to end, and pick a dashboard whose panels reference the vsphere_* measurements.
Treat that import as a smoke test, not a destination. Imported dashboards assume a counter list you did not configure, and half the panels will be empty. Once you have confirmed data is arriving, build your own with the panels in section 08.
Step 09. Add the configuration state path
Snapshot age, VM tags, orphaned disks and licence expiry are not performance counters and will never appear through this plugin. Two ways to get them: the Grafana Infinity datasource querying the vCenter REST API directly, good for small always current tables; or a scheduled script using govc or pyVmomi writing line protocol into the same bucket, better when you want to trend the value.
#!/usr/bin/env bash
# snapshot inventory into InfluxDB via govc, run hourly from cron
set -euo pipefail
export GOVC_URL="https://vcenter.lab.local"
export GOVC_USERNAME="svc-grafana@lab.local"
export GOVC_PASSWORD="${VSPHERE_PASSWORD}"
now=$(date +%s%N)
govc find / -type m | while read -r vm; do
name=$(basename "$vm")
count=$(govc snapshot.tree -vm "$vm" -f 2>/dev/null | wc -l)
[ "$count" -eq 0 ] && continue
oldest=$(govc snapshot.tree -vm "$vm" -D -f 2>/dev/null | head -1 | awk '{print $NF}')
age=$(( ( now/1000000000 - $(date -d "$oldest" +%s) ) / 86400 ))
echo "vsphere_vm_snapshot,vmname=${name} count=${count}i,age_days=${age}i ${now}"
done | curl -s --data-binary @- \
-H "Authorization: Token ${INFLUX_TOKEN}" \
"http://localhost:8086/api/v2/write?org=lab&bucket=vsphere&precision=ns"
Step 10. Set retention and downsampling
Raw data stays 30 days in the vsphere bucket. A task rolls hourly means into vsphere_long, which keeps 13 months.
// Flux task, created in the InfluxDB UI under Tasks
option task = {name: "vsphere-hourly-rollup", every: 1h, offset: 5m}
from(bucket: "vsphere")
|> range(start: -1h)
|> filter(fn: (r) => r._measurement =~ /^vsphere_/)
|> aggregateWindow(every: 1h, fn: mean, createEmpty: false)
|> to(bucket: "vsphere_long", org: "lab")
In Grafana, point capacity and trend dashboards at vsphere_long and troubleshooting dashboards at vsphere. Make the bucket a template variable so one dashboard can serve both.
07. From counter to pixel
This is the part that gets skipped, and it is the part that makes the whole thing debuggable. Here is a single metric traced from the hypervisor to the rendered line, using CPU ready because it is the counter most often misread.
| Stage | What exists there | What to check when it breaks |
|---|---|---|
| ESXi | The VMkernel scheduler records milliseconds a vCPU spent ready to run but waiting for a physical core, per 20 second sample, as cpu.ready.summation. The unit is milliseconds, not a percentage. | Real time chart in the vSphere Client for that VM. If it is empty there, it is empty everywhere. |
| vCenter | Exposes the same counter through QueryPerf against the VM managed object reference, at real time granularity or as a 5 minute rollup. | Statistics level for the interval you are querying. Per vCPU instance detail needs level 3. |
| Telegraf | Writes measurement vsphere_vm_cpu, field ready_summation, tags vcenter, dcname, clustername, esxhostname, vmname, moid. Field naming is the counter name with the rollup type appended. | The --test output, and the Telegraf log for maxQueryMetrics warnings. |
| InfluxDB | One series per unique tag combination, timestamped with the vCenter sample time rather than the write time. | Query the bucket directly with a 5 minute range. If Telegraf shows data and Influx does not, it is the output plugin or the token. |
| Grafana | Flux query selects the field, then arithmetic converts milliseconds to a percentage of the sample window. | The interval assumption in your divisor. This is where almost every wrong CPU ready number comes from. |
The arithmetic, stated plainly
ready_percent = ready_summation_ms / (interval_seconds * 1000) * 100
20 second real time samples : ready_ms / 20000 * 100 = ready_ms / 200
5 minute historical rollups : ready_ms / 300000 * 100 = ready_ms / 3000
For a multi vCPU VM the aggregate figure sums across all vCPUs, so divide by the vCPU count for a per core view. A VM showing 40 percent aggregate ready with 8 vCPUs is at 5 percent per core, which is normal. Reading that 40 as per core is how a healthy VM gets escalated.
// Flux, per VM CPU ready percentage from real time data
from(bucket: v.bucket)
|> range(start: v.timeRangeStart, stop: v.timeRangeStop)
|> filter(fn: (r) => r._measurement == "vsphere_vm_cpu")
|> filter(fn: (r) => r._field == "ready_summation")
|> filter(fn: (r) => r.clustername == "${cluster}")
|> aggregateWindow(every: v.windowPeriod, fn: mean, createEmpty: false)
|> map(fn: (r) => ({r with _value: r._value / 200.0}))
|> group(columns: ["vmname"])
|> yield(name: "cpu_ready_pct")
-- the equivalent in InfluxQL, if you are on InfluxDB 1.x
SELECT mean("ready_summation") / 200 AS "cpu_ready_pct"
FROM "vsphere_vm_cpu"
WHERE $timeFilter AND "clustername" =~ /^$cluster$/
GROUP BY time($__interval), "vmname" fill(none)
Template variables that make one dashboard serve everything
// variable: cluster
import "influxdata/influxdb/schema"
schema.tagValues(bucket: "vsphere",
tag: "clustername",
predicate: (r) => r._measurement == "vsphere_host_cpu")
// variable: vm, chained to cluster
import "influxdata/influxdb/schema"
schema.tagValues(bucket: "vsphere",
tag: "vmname",
predicate: (r) => r._measurement == "vsphere_vm_cpu"
and r.clustername == "${cluster}")
Chaining the VM list to the cluster selection is what keeps a two thousand VM dropdown usable. Add a bucket variable with the values vsphere and vsphere_long so the same dashboard covers both live troubleshooting and the annual view.
08. Panels worth building
Ordered by how often they end an argument.
| Panel | Metric | What it answers |
|---|---|---|
| CPU ready leaderboard | vsphere_vm_cpu.ready_summation, converted and normalised per vCPU | Which VMs are actually waiting for CPU, as opposed to which VMs are busy. Sort descending, top 20, table panel. |
| Co stop | vsphere_vm_cpu.costop_summation | Whether a wide VM is being penalised for its vCPU count. High co stop with low utilisation is the textbook right sizing signal. |
| Memory pressure stack | vmmemctl_average, swapinRate_average, active_average | Whether the host is reclaiming. Ballooning is a warning, swap in is already hurting. |
| Latency split | Guest virtualDisk.totalWriteLatency, host disk.maxTotalLatency, array side write latency | Where the milliseconds are being spent. Three series, one axis, one glance. See use case 2. |
| Datastore headroom | vsphere_datastore_disk used against capacity, projected | Days until full, using a linear fit over the last 30 days from the long bucket. |
| Allocated against consumed | Configured vCPU and memory against demand_average and active_average | The reclaim list. Usually 20 to 30 percent of a mature estate. |
| Dropped packets | net.droppedRx.summation, net.droppedTx.summation | Ring buffer exhaustion and uplink saturation, both invisible in throughput graphs. |
| Snapshot age table | vsphere_vm_snapshot.age_days from step 09 | The snapshots nobody deleted. Threshold at 3 days, colour at 7. |
09. Use cases
Right sizing, with evidence
Compare configured vCPU against cpu.demand.average, and configured memory against mem.active.average, over 30 days at the 95th percentile. Application owners resist right sizing because it usually arrives as an assertion. A 30 day graph showing an 8 vCPU VM peaking at 1.2 vCPU of demand, with co stop climbing because of the width, is a different conversation. Track reclaimed vCPU and GB as a running total and it becomes a reportable outcome.
Splitting storage latency three ways
This is the highest value use case in the whole build, and it needs decision 10 in place. Put three series on one time axis for the same workload:
- Guest latency from
virtualDisk.totalWriteLatency.average, what the application experiences. - Host latency from
disk.maxTotalLatency.latest, what the kernel sees after queuing. - Array latency from the array REST API, what the storage system reports for the same volume.
Read it as a subtraction. Guest high and array low means the delay is in the host queue or the fabric, not the array. All three high and moving together means the array is genuinely saturated. Guest high, host normal, array normal points at the guest itself, often a filesystem or queue depth setting. That determination normally takes a bridge call with three teams, and here it is one panel.
Capacity forecasting that survives review
Query the 13 month bucket, fit a linear trend to datastore used capacity and cluster memory consumed, and project forward. Two numbers matter to a budget conversation: days until a resource is exhausted, and the growth rate per month. Both come out of the same query. Add a marker for the last hardware purchase so the trend line has context.
Change validation
Snapshot a dashboard before a firmware upgrade, a driver change or a datastore migration, then compare after. Grafana’s snapshot feature freezes the data with the panel, so the before state is preserved even after retention expires. This turns “it feels slower since the upgrade” into a measurable claim within an hour.
Noisy neighbour identification
Group CPU ready by esxhostname and overlay total host CPU. When one host shows elevated ready across many VMs, it is a placement problem for DRS. When one VM shows ready while its host is idle, it is a limit or a shares setting. Both are configuration rather than capacity.
Showback
Tag VMs in vCenter with a cost centre, expose the tag as a Telegraf custom attribute, and aggregate consumed CPU, memory and storage by tag. Even without billing attached, a monthly consumption table per business unit changes request behaviour quickly.
10. Alerting without the noise
Grafana unified alerting evaluates the same queries the panels use. Four rules that earn their keep.
| Rule | Condition | Why this threshold |
|---|---|---|
| Sustained CPU contention | Per vCPU ready above 10 percent for 15 minutes | Below 5 percent is normal on any consolidated host. Brief spikes during boot storms are not incidents. The 15 minute window is what removes the noise. |
| Memory reclamation active | Swap in rate above zero for 5 minutes on any host | Ballooning is the hypervisor working as designed. Swap in means it has run out of gentler options and guest performance is already affected. |
| Datastore projected full | Linear projection crosses 90 percent inside 14 days | Alerting at a static 85 percent fires on datastores stable for two years. Alerting on trajectory gives you time to act and stays quiet otherwise. |
| Collector health | No new points in the bucket for 10 minutes | The rule that catches every other rule failing silently. Build it first. |
Do not duplicate vCenter alarms. vCenter already alarms on host connectivity, datastore full and hardware health, and it does it better because it has state you do not. Use Grafana for trend, contention and cross system correlation, which are exactly the things vCenter alarms cannot express. Two systems alerting on the same condition means both get muted.
11. Scaling, and what breaks first
| Symptom | Cause | Fix |
|---|---|---|
Fault mentioning vpxd.stats.maxQueryMetrics | Cluster aggregation subqueries exceeding the cap, per section 03 | Exclude cluster metrics, or raise the setting above your VM count. Do not set it to -1. |
| Collection takes longer than the interval | Too many objects for the configured concurrency, or a slow vCenter database | Raise collect_concurrency in steps of 2 and measure. If vCenter CPU rises, the answer is a second collector sharded by datacenter instead. |
| Sawtooth gaps in graphs | Historical objects polled faster than their 5 minute source granularity | Move them to the 300 second input block. This is decision 02 asserting itself. |
| Dashboards slow after several months | Series cardinality, usually from wildcard VM counters or per instance detail | Measure cardinality per measurement, prune the include list, and move long range panels to the downsampled bucket. |
| Series for deleted VMs persist | Time series databases do not delete, they expire | Let retention handle it, and filter panels on a recent time window rather than the full history. |
| Everything stops after a vCenter upgrade | Certificate replaced, or the service account password expired | The collector health alert from section 10 catches this in ten minutes instead of on Monday. |
Above roughly 3000 VMs
- Shard collectors by vCenter, then by datacenter within a vCenter. Each Telegraf instance handles one shard.
- Separate InfluxDB onto dedicated hardware with SSD backed storage, and give it memory proportional to your series count.
- Drop VM level real time collection for non production clusters and keep only the 5 minute rollups there. Half your inventory usually does not need 60 second data.
- Consider VictoriaMetrics or Mimir if a single InfluxDB instance becomes the bottleneck. Telegraf can write to either without changing the input configuration, which is a large part of why the collector and the store are separate components in the first place.
Closing thought
The install is the easy part and it is what every guide covers. What determines whether this pipeline is still running and trusted a year from now is the set of choices in section 04: splitting the collectors by interval class, refusing to wildcard VM counters, keeping the unit arithmetic in the query layer where it can be audited, and getting array side latency onto the same time axis as guest latency. Get those right and the dashboards mostly build themselves.
Tested against vCenter 8.0, Telegraf 1.34, InfluxDB 2.7 and Grafana 11.6. Counter names and plugin options are stable across recent versions, but check the Telegraf vSphere plugin documentation before copying configuration verbatim. Configuration here assumes a lab or non production first pass. Validate collection load against your own vCenter before pointing it at production.
Leave a Reply