Most guides stop at “install Telegraf, import a dashboard, done”. That works on a lab with twelve VMs and falls apart at four hundred. This build walks the full path from a vSphere performance counter to a rendered pixel, explains the statistics model underneath it, and calls out every decision point where you are trading accuracy against load on vCenter.

Level: intermediate to advanced  ·  Stack: vCenter 7 or 8, Telegraf, InfluxDB, Grafana  ·  Build time: about 90 minutes


Contents

  1. What we are building
  2. Architecture and data flow
  3. How vCenter actually exposes metrics
  4. The design decisions that matter
  5. Prerequisites and sizing
  6. Build, step by step
  7. From counter to pixel
  8. Panels worth building
  9. Use cases
  10. Alerting without the noise
  11. Scaling, and what breaks first

01. What we are building

The goal is a self hosted metrics pipeline that pulls performance counters out of vCenter on a schedule, stores them as time series, and renders them in Grafana with alerting attached. Nothing is installed on ESXi hosts and nothing is installed on the vCenter Server Appliance. The collector talks to the vSphere Web Services API over HTTPS on port 443 with a read only service account, exactly the same interface the vSphere Client uses.

ComponentJobWhy this one
vCenter ServerSource of truth for inventory and performance countersIt already aggregates every host and VM. Polling hosts directly means N connections and no cluster level rollups.
Telegraf with the vsphere input pluginDiscovers inventory, batches QueryPerf calls, converts to line protocolWritten against govmomi, understands real time versus historical intervals, and respects the vCenter query limit automatically.
InfluxDBTime series storage, retention and downsamplingTag based model fits vSphere inventory cleanly. Prometheus is a valid alternative, covered in decision 03.
GrafanaQuery, visualise, alertPanel level transformations let you do the counter arithmetic without pre computing it at ingest.

Scope. This pipeline covers performance data. Configuration and state data (snapshot age, VM tags, orphaned VMDKs, licence expiry, certificate expiry) does not live in performance counters and needs a second path, covered in step 09.

02. Architecture and data flow

SOURCE
    ESXi hosts            vCenter Server              Storage array
    20s samples   ---->   /sdk Web Services API       REST API
    kept 1 hour           rollups 5m / 30m / 2h / 1d  array side latency
                                    |                          |
         +--------------------------+-------------+            |
         |                          |             |            |
COLLECT  v                          v             v            v
    Telegraf realtime      Telegraf historical    Telegraf http input
    interval 60s           interval 300s          interval 60s
    hosts + virtual VMs    datastores, DC         array metrics
         |                          |                          |
         +--------------------------+--------------------------+
                                    |
STORE                               v
                     InfluxDB   bucket: vsphere
                     vsphere_vm_cpu, vsphere_host_datastore, ...
                     tags: vcenter, dcname, clustername, vmname, moid
                                    |
PRESENT                             v
                     Grafana
                     dashboards, template variables, transformations
                     unified alerting, contact points

The pipeline splits deliberately into two collector instances because vCenter serves real time and historical counters from different places at different granularities. Merging them into one polling loop is the single most common cause of a pipeline that half works.

The flow in words

  1. ESXi hosts sample their own counters every 20 seconds and keep roughly one hour of that detail.
  2. vCenter pulls those samples and writes rolled up copies into its database at 5 minute, 30 minute, 2 hour and 1 day granularity, governed by the statistics level configured per interval.
  3. Telegraf logs into https://vcenter/sdk, walks the inventory tree to discover managed object references, then issues batched QueryPerf calls for the counters you asked for.
  4. The plugin turns each returned sample into line protocol, one measurement per resource type and counter group, and writes it to InfluxDB with inventory tags attached.
  5. Grafana queries InfluxDB. Any unit conversion happens in the query or a panel transformation, not at ingest.
  6. Alert rules evaluate the same queries on a schedule and fire into contact points.

03. How vCenter actually exposes metrics

Skip this section and you will spend a weekend wondering why datastore latency has gaps and cluster CPU sometimes returns nothing. Three properties of the vSphere statistics model drive almost every design decision later on.

Real time versus historical

Real time counters exist only for ESXi hosts and virtual machines, at 20 second granularity, retained about an hour. Datastores, clusters, resource pools and datacenters have no real time series at all. Their smallest available granularity is the 5 minute historical rollup, served out of the vCenter database.

IntervalGranularityTypical retentionApplies to
Real time20 secondsAbout 1 hour, on the hostHosts, virtual machines
Past day5 minutes1 dayAll object types
Past week30 minutes1 weekAll object types
Past month2 hours1 monthAll object types
Past year1 day1 yearAll object types

Consequence. Polling datastore or cluster metrics on a 20 or 60 second loop does not give you faster data. It gives you the same 5 minute value repeated, plus real query load on the vCenter database. Poll historical resource types at 300 seconds and nothing faster.

Statistics levels

Each historical interval has a statistics level from 1 to 4. Level 1 gives you the basic average counters. Device level and per instance detail (per vCPU, per vmnic, per virtual disk) only appears at level 3 and above. If a counter you configured never shows up in InfluxDB, the usual cause is that the level for that interval is too low, not that the plugin is broken.

Raising levels globally is expensive. The vCenter database grows quickly and the rollup jobs get heavier. Raise level 1 to level 2 or 3 only if you genuinely need per device breakdowns in the longer intervals, and leave levels 2 through 4 alone.

The query limit

vCenter caps how many metrics a single QueryPerf call may request, through the advanced setting config.vpxd.stats.maxQueryMetrics. Per the Telegraf plugin documentation, the value defaults to 64 on vSphere 5.5 and earlier and to 256 on more recent versions. The plugin always checks this setting and automatically reduces its own batch size when the vCenter limit is lower than the configured value, logging a warning when it does.

Cluster metrics make this worse. Cluster counters are aggregated from host and virtual machine metrics, and when the most recent values are not already available vCenter performs that aggregation on the fly. Every internal subquery needed for that aggregation counts towards the same limit, so even a small query can fail with a fault about vpxd.stats.maxQueryMetrics.

The documented remedies are to raise config.vpxd.stats.maxQueryMetrics above the total number of virtual machines managed by that vCenter, or to exclude cluster metrics entirely and derive cluster level figures by aggregating host data in the query layer instead. The second option is the one I prefer, and decision 05 explains why.

04. The design decisions that matter

These are the choices that determine whether the pipeline survives contact with a production estate. Each is written as the option set, the choice, and the cost of that choice, because there is no universally right answer and you should be able to argue yours.

Decision 01. Pull from vCenter, or push from hosts

  • Options. Poll the vCenter API, poll each ESXi host directly, or run an agent inside every guest.
  • Choice. Poll vCenter.
  • Why. One credential, one endpoint, inventory relationships already resolved, and cluster and datastore objects that simply do not exist at the host level. Host direct polling scales linearly in connections and loses every rollup.
  • Cost. vCenter becomes a single point of failure for monitoring, and your collection load lands on the appliance that also runs DRS, HA and the client. Budget for it explicitly rather than discovering it.

Decision 02. Split the collector by interval class

  • Options. One Telegraf input block for everything, or two input blocks with different intervals.
  • Choice. Two separate input blocks, ideally two separate config files.
  • Why. Real time objects want a 60 second loop. Historical objects want 300 seconds. A single block forces one interval on both, so you either hammer the vCenter database or you blur your VM data.
  • Cost. Two config files to keep in sync, two sets of credentials, two failure surfaces to monitor. Worth it every time.

Decision 03. InfluxDB or Prometheus

  • Options. Telegraf into InfluxDB, or a govmomi based exporter scraped by Prometheus.
  • Choice. InfluxDB for vSphere specifically.
  • Why. vSphere counters arrive with timestamps that are not the scrape time, especially for historical rollups. Prometheus assigns the scrape timestamp to a sample, which silently shifts 5 minute rollup data. Telegraf preserves the sample timestamp from vCenter. The push model also survives a collector that takes 90 seconds to walk a large inventory, where a scrape would time out.
  • Cost. Friction if your organisation is standardised on Prometheus. The workable compromise is Telegraf with a Prometheus remote write output, keeping the timestamp handling but landing in the platform you already run.

Decision 04. Which counters, and therefore how much cardinality

  • Options. Wildcard everything, or an explicit include list per resource type.
  • Choice. Explicit include list. Never wildcard virtual machine counters.
  • Why. Series count is roughly objects multiplied by counters multiplied by instances. Two thousand VMs with 40 wildcard counters and per device instances runs into millions of series, and the pain lands on query latency months later when the dashboard is already business critical.
  • Cost. You will occasionally need a counter you did not collect, and it is not retroactive. Keep the include list in version control with a comment per counter explaining which panel uses it.

Decision 05. Collect cluster metrics, or aggregate host metrics at query time

  • Options. Enable cluster_metric_include, or exclude clusters and sum host series in Grafana.
  • Choice. Exclude cluster metrics, aggregate in the query layer.
  • Why. Cluster counters are the main trigger for the maxQueryMetrics fault. Summing host CPU and memory in Flux or InfluxQL gives the same numbers without loading vCenter, and it stays correct when hosts move between clusters.
  • Cost. Slightly more complex queries, and cluster level counters that have no host equivalent are lost. If you need those, raise the vCenter limit above your VM count and accept the load.

Decision 06. Credential scope

  • Options. Reuse an administrator account, or create a dedicated read only service account.
  • Choice. Dedicated account with the built in Read only role, assigned at the vCenter root object with propagation enabled.
  • Why. Performance queries need nothing beyond read. Root with propagation is required because the plugin discovers inventory from the top, and a role scoped to one datacenter produces confusing partial discovery rather than a clean permission error.
  • Cost. One more identity to rotate. Put the password in an environment file the collector reads, never inline in the config, and set the account not to expire.

Decision 07. TLS verification

  • Options. Set insecure_skip_verify = true, or trust the vCenter certificate authority properly.
  • Choice. Trust the CA. Export the vCenter root certificate and install it in the collector trust store.
  • Why. Skipping verification is the default in every tutorial and it means your monitoring credentials will be handed to anything that answers on that address. It is a read only account, but it is still an authenticated path into vCenter.
  • Cost. Five extra minutes at build time, and a certificate to rotate. Every guide skips this and every audit flags it.

Decision 08. Retention and downsampling

  • Options. Keep raw samples forever, or keep raw briefly and downsample for the long view.
  • Choice. Raw for 30 days, hourly rollups for 13 months.
  • Why. Troubleshooting needs full resolution and happens within days. Capacity planning needs trend, not detail, and 13 months lets you compare against the same month last year.
  • Cost. A downsampling task to write and monitor, and dashboards that must pick the right bucket based on the selected time range.

Decision 09. Where the unit maths lives

  • Options. Convert counters at ingest with a Telegraf processor, or convert in the Grafana query.
  • Choice. Convert in Grafana.
  • Why. The raw counter stays raw and auditable, so when someone disputes a CPU ready number you can show the summation value and the arithmetic. Converting at ingest bakes an assumption about the sample interval into stored data, and that assumption breaks the day you change the collection interval.
  • Cost. The same expression repeated across panels. Solve it with a library panel, not by pre computing at ingest.

Decision 10. Correlating with the storage array

  • Options. vSphere latency only, or vSphere latency alongside array side latency on the same time axis.
  • Choice. Pull array metrics into the same bucket, tagged with the datastore name.
  • Why. This is the single highest value thing you can add. Guest latency minus array latency tells you whether the delay is in the array, the fabric or the host queue, and that answer normally takes three teams and a bridge call to reach.
  • Cost. You have to get the tag values to match exactly across two systems that name things differently. Do the normalisation in a Telegraf rename or override processor, once, at ingest.

05. Prerequisites and sizing

  • vCenter Server 7.0 or 8.0, reachable on TCP 443 from the collector.
  • A Linux VM for the monitoring stack, kept off the vCenter appliance itself. For up to 500 VMs, 4 vCPU, 8 GB RAM and 100 GB of fast disk is comfortable. Above 2000 VMs, separate InfluxDB onto its own host with 16 GB and SSD backed storage.
  • Docker and the compose plugin, or native packages if you prefer.
  • A vCenter service account, created in step 01.
  • NTP working on the collector. Time series with skewed clocks look like data loss.

Placement. Put the monitoring VM on a different cluster and a different datastore from the workloads you are watching, if you have one. A monitoring stack that goes down with the thing it monitors tells you nothing during the outage you built it for.

06. Build, step by step

Step 01. Create the read only vCenter service account

Create the user in vCenter Single Sign On, then assign the built in Read only role at the root vCenter Server object with Propagate to children ticked. Assigning at a datacenter or cluster instead produces partial discovery that is hard to diagnose.

# PowerCLI
Connect-VIServer -Server vcenter.lab.local

New-VIPermission -Entity (Get-Folder -NoRecursion) `
                 -Principal 'LAB\svc-grafana' `
                 -Role 'ReadOnly' `
                 -Propagate $true

# verify it landed at the root
Get-VIPermission | Where-Object { $_.Principal -like '*svc-grafana*' } |
  Format-Table Entity, Role, Propagate

Step 02. Raise the query limit, or plan to live without cluster metrics

In the vSphere Client go to the vCenter object, Configure, Advanced Settings, and check config.vpxd.stats.maxQueryMetrics. If the setting is absent it is running at the default. Following decision 05 I leave it alone and exclude cluster metrics. If you do need cluster counters, set it above your total VM count.

# check the current value
Get-AdvancedSetting -Entity $global:DefaultVIServer `
  -Name 'config.vpxd.stats.maxQueryMetrics'

# raise it only if you are collecting cluster metrics
Get-AdvancedSetting -Entity $global:DefaultVIServer `
  -Name 'config.vpxd.stats.maxQueryMetrics' |
  Set-AdvancedSetting -Value 4096 -Confirm:$false

Do not set this to -1. Removing the cap entirely lets a single malformed query ask vCenter for everything at once. Pick a number tied to your inventory size and revisit it when the estate grows.

Step 03. Trust the vCenter certificate

Download the root certificate bundle from the vCenter certs endpoint, extract the .0 file for Linux, and install it. This is what lets you leave insecure_skip_verify at false.

curl -sk -o /tmp/vc-certs.zip https://vcenter.lab.local/certs/download.zip
unzip -o /tmp/vc-certs.zip -d /tmp/vc-certs

# the lin folder holds the OpenSSL hashed certificate
cp /tmp/vc-certs/certs/lin/*.0 /usr/local/share/ca-certificates/vcenter-root.crt
update-ca-certificates

# confirm the chain now validates
openssl s_client -connect vcenter.lab.local:443 -brief < /dev/null

Step 04. Stand up InfluxDB and Grafana

One compose file, three services. The CA bundle is mounted into the Telegraf container so it can validate the vCenter certificate.

# docker-compose.yml
services:
  influxdb:
    image: influxdb:2.7
    container_name: influxdb
    restart: unless-stopped
    ports:
      - "8086:8086"
    volumes:
      - influx-data:/var/lib/influxdb2
      - influx-conf:/etc/influxdb2
    environment:
      DOCKER_INFLUXDB_INIT_MODE: setup
      DOCKER_INFLUXDB_INIT_USERNAME: admin
      DOCKER_INFLUXDB_INIT_PASSWORD: ${INFLUX_PASSWORD}
      DOCKER_INFLUXDB_INIT_ORG: lab
      DOCKER_INFLUXDB_INIT_BUCKET: vsphere
      DOCKER_INFLUXDB_INIT_RETENTION: 30d
      DOCKER_INFLUXDB_INIT_ADMIN_TOKEN: ${INFLUX_TOKEN}

  telegraf:
    image: telegraf:1.34
    container_name: telegraf
    restart: unless-stopped
    depends_on:
      - influxdb
    env_file:
      - ./vsphere.env
    volumes:
      - ./telegraf.d:/etc/telegraf/telegraf.d:ro
      - ./telegraf.conf:/etc/telegraf/telegraf.conf:ro
      - /etc/ssl/certs:/etc/ssl/certs:ro

  grafana:
    image: grafana/grafana:11.6.0
    container_name: grafana
    restart: unless-stopped
    depends_on:
      - influxdb
    ports:
      - "3000:3000"
    volumes:
      - grafana-data:/var/lib/grafana
      - ./provisioning:/etc/grafana/provisioning:ro
    environment:
      GF_SECURITY_ADMIN_PASSWORD: ${GRAFANA_PASSWORD}
      GF_USERS_ALLOW_SIGN_UP: "false"

volumes:
  influx-data:
  influx-conf:
  grafana-data:
# vsphere.env, kept out of version control
VSPHERE_URL=https://vcenter.lab.local/sdk
VSPHERE_USER=svc-grafana@lab.local
VSPHERE_PASSWORD=change-me
INFLUX_TOKEN=paste-the-admin-token
INFLUX_ORG=lab
INFLUX_BUCKET=vsphere

# then
chmod 600 vsphere.env
docker compose up -d

Step 05. Configure the real time collector

Hosts and virtual machines, 60 second loop, explicit counter list. Note that datastore, cluster and datacenter are all excluded here because they belong to the historical block.

# telegraf.d/vsphere-realtime.conf
[[inputs.vsphere]]
  interval = "60s"
  vcenters = ["${VSPHERE_URL}"]
  username = "${VSPHERE_USER}"
  password = "${VSPHERE_PASSWORD}"

  ## certificate is trusted at the OS level, see step 03
  insecure_skip_verify = false

  ## ---- virtual machines ----
  vm_metric_include = [
    "cpu.demand.average",          # what the VM wants
    "cpu.usagemhz.average",
    "cpu.usage.average",
    "cpu.ready.summation",         # contention, the headline counter
    "cpu.costop.summation",        # oversized vCPU tell
    "cpu.latency.average",
    "mem.active.average",          # true working set
    "mem.consumed.average",
    "mem.granted.average",
    "mem.vmmemctl.average",        # ballooning
    "mem.swapinRate.average",      # host is out of memory
    "mem.swapoutRate.average",
    "mem.usage.average",
    "net.bytesRx.average",
    "net.bytesTx.average",
    "net.droppedRx.summation",
    "net.droppedTx.summation",
    "virtualDisk.read.average",
    "virtualDisk.write.average",
    "virtualDisk.numberReadAveraged.average",
    "virtualDisk.numberWriteAveraged.average",
    "virtualDisk.totalReadLatency.average",
    "virtualDisk.totalWriteLatency.average",
    "sys.uptime.latest",
  ]

  ## ---- esxi hosts ----
  host_metric_include = [
    "cpu.usage.average",
    "cpu.usagemhz.average",
    "cpu.ready.summation",
    "cpu.utilization.average",
    "mem.usage.average",
    "mem.active.average",
    "mem.totalCapacity.average",
    "mem.vmmemctl.average",
    "mem.swapused.average",
    "net.bytesRx.average",
    "net.bytesTx.average",
    "net.droppedRx.summation",
    "net.droppedTx.summation",
    "disk.maxTotalLatency.latest",   # the number that matters
    "disk.numberReadAveraged.average",
    "disk.numberWriteAveraged.average",
    "power.power.average",
    "sys.uptime.latest",
  ]

  ## historical objects belong to the other input block
  datastore_metric_exclude     = ["*"]
  cluster_metric_exclude       = ["*"]
  datacenter_metric_exclude    = ["*"]
  resource_pool_metric_exclude = ["*"]
  vsan_metric_exclude          = ["*"]

  ## batching and concurrency, see decision 04
  max_query_objects  = 256
  max_query_metrics  = 256
  collect_concurrency = 3
  discover_concurrency = 3
  object_discovery_interval = "300s"
  timeout = "60s"

Tuning. The plugin documentation suggests a rule of thumb of setting the concurrency parameters to the number of virtual machines divided by 1500, rounded up. Start at 3, watch the gather duration, and raise it only if collection is not finishing inside the interval.

Step 06. Configure the historical collector

Datastores only, 300 second loop, clusters excluded per decision 05. A generous timeout matters here because the vCenter database is the bottleneck, not the network.

# telegraf.d/vsphere-historical.conf
[[inputs.vsphere]]
  interval = "300s"
  vcenters = ["${VSPHERE_URL}"]
  username = "${VSPHERE_USER}"
  password = "${VSPHERE_PASSWORD}"
  insecure_skip_verify = false

  vm_metric_exclude      = ["*"]
  host_metric_exclude    = ["*"]
  cluster_metric_exclude = ["*"]     # aggregate hosts in Grafana instead
  vsan_metric_exclude    = ["*"]

  datastore_metric_include = [
    "disk.capacity.latest",
    "disk.used.latest",
    "disk.provisioned.latest",
    "disk.capacity.provisioned.average",
    "disk.capacity.usage.average",
    "datastore.numberReadAveraged.average",
    "datastore.numberWriteAveraged.average",
    "datastore.totalReadLatency.average",
    "datastore.totalWriteLatency.average",
  ]

  datacenter_metric_include = ["*"]

  max_query_objects = 256
  collect_concurrency = 2
  timeout = "240s"
  force_discover_on_init = true

Validate before restarting the service. The test flag runs one collection cycle and prints line protocol to stdout without writing anything. If you see line protocol here, vCenter connectivity, permissions and TLS are all correct.

docker compose exec telegraf telegraf \
  --config /etc/telegraf/telegraf.conf \
  --config-directory /etc/telegraf/telegraf.d \
  --input-filter vsphere --test | head -40

Step 07. Provision the Grafana datasource

Provision it as a file rather than clicking through the UI, so the whole stack rebuilds from the repository.

# provisioning/datasources/influx.yml
apiVersion: 1
datasources:
  - name: InfluxDB-vSphere
    type: influxdb
    access: proxy
    url: http://influxdb:8086
    jsonData:
      version: Flux
      organization: lab
      defaultBucket: vsphere
      tlsSkipVerify: false
      timeInterval: "60s"
    secureJsonData:
      token: ${INFLUX_TOKEN}
    isDefault: true

The timeInterval value tells Grafana the minimum meaningful resolution of the data, which stops it drawing misleading interpolation between sparse points.

Step 08. Import a baseline dashboard, then replace it

Community vSphere dashboards for the Telegraf plugin are published on the Grafana dashboard library and on the InfluxData integration page. Import one to confirm data is flowing end to end, and pick a dashboard whose panels reference the vsphere_* measurements.

Treat that import as a smoke test, not a destination. Imported dashboards assume a counter list you did not configure, and half the panels will be empty. Once you have confirmed data is arriving, build your own with the panels in section 08.

Step 09. Add the configuration state path

Snapshot age, VM tags, orphaned disks and licence expiry are not performance counters and will never appear through this plugin. Two ways to get them: the Grafana Infinity datasource querying the vCenter REST API directly, good for small always current tables; or a scheduled script using govc or pyVmomi writing line protocol into the same bucket, better when you want to trend the value.

#!/usr/bin/env bash
# snapshot inventory into InfluxDB via govc, run hourly from cron
set -euo pipefail

export GOVC_URL="https://vcenter.lab.local"
export GOVC_USERNAME="svc-grafana@lab.local"
export GOVC_PASSWORD="${VSPHERE_PASSWORD}"

now=$(date +%s%N)

govc find / -type m | while read -r vm; do
  name=$(basename "$vm")
  count=$(govc snapshot.tree -vm "$vm" -f 2>/dev/null | wc -l)
  [ "$count" -eq 0 ] && continue
  oldest=$(govc snapshot.tree -vm "$vm" -D -f 2>/dev/null | head -1 | awk '{print $NF}')
  age=$(( ( now/1000000000 - $(date -d "$oldest" +%s) ) / 86400 ))
  echo "vsphere_vm_snapshot,vmname=${name} count=${count}i,age_days=${age}i ${now}"
done | curl -s --data-binary @- \
  -H "Authorization: Token ${INFLUX_TOKEN}" \
  "http://localhost:8086/api/v2/write?org=lab&bucket=vsphere&precision=ns"

Step 10. Set retention and downsampling

Raw data stays 30 days in the vsphere bucket. A task rolls hourly means into vsphere_long, which keeps 13 months.

// Flux task, created in the InfluxDB UI under Tasks
option task = {name: "vsphere-hourly-rollup", every: 1h, offset: 5m}

from(bucket: "vsphere")
  |> range(start: -1h)
  |> filter(fn: (r) => r._measurement =~ /^vsphere_/)
  |> aggregateWindow(every: 1h, fn: mean, createEmpty: false)
  |> to(bucket: "vsphere_long", org: "lab")

In Grafana, point capacity and trend dashboards at vsphere_long and troubleshooting dashboards at vsphere. Make the bucket a template variable so one dashboard can serve both.

07. From counter to pixel

This is the part that gets skipped, and it is the part that makes the whole thing debuggable. Here is a single metric traced from the hypervisor to the rendered line, using CPU ready because it is the counter most often misread.

StageWhat exists thereWhat to check when it breaks
ESXiThe VMkernel scheduler records milliseconds a vCPU spent ready to run but waiting for a physical core, per 20 second sample, as cpu.ready.summation. The unit is milliseconds, not a percentage.Real time chart in the vSphere Client for that VM. If it is empty there, it is empty everywhere.
vCenterExposes the same counter through QueryPerf against the VM managed object reference, at real time granularity or as a 5 minute rollup.Statistics level for the interval you are querying. Per vCPU instance detail needs level 3.
TelegrafWrites measurement vsphere_vm_cpu, field ready_summation, tags vcenter, dcname, clustername, esxhostname, vmname, moid. Field naming is the counter name with the rollup type appended.The --test output, and the Telegraf log for maxQueryMetrics warnings.
InfluxDBOne series per unique tag combination, timestamped with the vCenter sample time rather than the write time.Query the bucket directly with a 5 minute range. If Telegraf shows data and Influx does not, it is the output plugin or the token.
GrafanaFlux query selects the field, then arithmetic converts milliseconds to a percentage of the sample window.The interval assumption in your divisor. This is where almost every wrong CPU ready number comes from.

The arithmetic, stated plainly

ready_percent = ready_summation_ms / (interval_seconds * 1000) * 100

  20 second real time samples : ready_ms / 20000  * 100  = ready_ms / 200
  5 minute historical rollups : ready_ms / 300000 * 100  = ready_ms / 3000

For a multi vCPU VM the aggregate figure sums across all vCPUs, so divide by the vCPU count for a per core view. A VM showing 40 percent aggregate ready with 8 vCPUs is at 5 percent per core, which is normal. Reading that 40 as per core is how a healthy VM gets escalated.

// Flux, per VM CPU ready percentage from real time data
from(bucket: v.bucket)
  |> range(start: v.timeRangeStart, stop: v.timeRangeStop)
  |> filter(fn: (r) => r._measurement == "vsphere_vm_cpu")
  |> filter(fn: (r) => r._field == "ready_summation")
  |> filter(fn: (r) => r.clustername == "${cluster}")
  |> aggregateWindow(every: v.windowPeriod, fn: mean, createEmpty: false)
  |> map(fn: (r) => ({r with _value: r._value / 200.0}))
  |> group(columns: ["vmname"])
  |> yield(name: "cpu_ready_pct")
-- the equivalent in InfluxQL, if you are on InfluxDB 1.x
SELECT mean("ready_summation") / 200 AS "cpu_ready_pct"
FROM "vsphere_vm_cpu"
WHERE $timeFilter AND "clustername" =~ /^$cluster$/
GROUP BY time($__interval), "vmname" fill(none)

Template variables that make one dashboard serve everything

// variable: cluster
import "influxdata/influxdb/schema"
schema.tagValues(bucket: "vsphere",
                 tag: "clustername",
                 predicate: (r) => r._measurement == "vsphere_host_cpu")

// variable: vm, chained to cluster
import "influxdata/influxdb/schema"
schema.tagValues(bucket: "vsphere",
                 tag: "vmname",
                 predicate: (r) => r._measurement == "vsphere_vm_cpu"
                                   and r.clustername == "${cluster}")

Chaining the VM list to the cluster selection is what keeps a two thousand VM dropdown usable. Add a bucket variable with the values vsphere and vsphere_long so the same dashboard covers both live troubleshooting and the annual view.

08. Panels worth building

Ordered by how often they end an argument.

PanelMetricWhat it answers
CPU ready leaderboardvsphere_vm_cpu.ready_summation, converted and normalised per vCPUWhich VMs are actually waiting for CPU, as opposed to which VMs are busy. Sort descending, top 20, table panel.
Co stopvsphere_vm_cpu.costop_summationWhether a wide VM is being penalised for its vCPU count. High co stop with low utilisation is the textbook right sizing signal.
Memory pressure stackvmmemctl_average, swapinRate_average, active_averageWhether the host is reclaiming. Ballooning is a warning, swap in is already hurting.
Latency splitGuest virtualDisk.totalWriteLatency, host disk.maxTotalLatency, array side write latencyWhere the milliseconds are being spent. Three series, one axis, one glance. See use case 2.
Datastore headroomvsphere_datastore_disk used against capacity, projectedDays until full, using a linear fit over the last 30 days from the long bucket.
Allocated against consumedConfigured vCPU and memory against demand_average and active_averageThe reclaim list. Usually 20 to 30 percent of a mature estate.
Dropped packetsnet.droppedRx.summation, net.droppedTx.summationRing buffer exhaustion and uplink saturation, both invisible in throughput graphs.
Snapshot age tablevsphere_vm_snapshot.age_days from step 09The snapshots nobody deleted. Threshold at 3 days, colour at 7.

09. Use cases

Right sizing, with evidence

Compare configured vCPU against cpu.demand.average, and configured memory against mem.active.average, over 30 days at the 95th percentile. Application owners resist right sizing because it usually arrives as an assertion. A 30 day graph showing an 8 vCPU VM peaking at 1.2 vCPU of demand, with co stop climbing because of the width, is a different conversation. Track reclaimed vCPU and GB as a running total and it becomes a reportable outcome.

Splitting storage latency three ways

This is the highest value use case in the whole build, and it needs decision 10 in place. Put three series on one time axis for the same workload:

  • Guest latency from virtualDisk.totalWriteLatency.average, what the application experiences.
  • Host latency from disk.maxTotalLatency.latest, what the kernel sees after queuing.
  • Array latency from the array REST API, what the storage system reports for the same volume.

Read it as a subtraction. Guest high and array low means the delay is in the host queue or the fabric, not the array. All three high and moving together means the array is genuinely saturated. Guest high, host normal, array normal points at the guest itself, often a filesystem or queue depth setting. That determination normally takes a bridge call with three teams, and here it is one panel.

Capacity forecasting that survives review

Query the 13 month bucket, fit a linear trend to datastore used capacity and cluster memory consumed, and project forward. Two numbers matter to a budget conversation: days until a resource is exhausted, and the growth rate per month. Both come out of the same query. Add a marker for the last hardware purchase so the trend line has context.

Change validation

Snapshot a dashboard before a firmware upgrade, a driver change or a datastore migration, then compare after. Grafana’s snapshot feature freezes the data with the panel, so the before state is preserved even after retention expires. This turns “it feels slower since the upgrade” into a measurable claim within an hour.

Noisy neighbour identification

Group CPU ready by esxhostname and overlay total host CPU. When one host shows elevated ready across many VMs, it is a placement problem for DRS. When one VM shows ready while its host is idle, it is a limit or a shares setting. Both are configuration rather than capacity.

Showback

Tag VMs in vCenter with a cost centre, expose the tag as a Telegraf custom attribute, and aggregate consumed CPU, memory and storage by tag. Even without billing attached, a monthly consumption table per business unit changes request behaviour quickly.

10. Alerting without the noise

Grafana unified alerting evaluates the same queries the panels use. Four rules that earn their keep.

RuleConditionWhy this threshold
Sustained CPU contentionPer vCPU ready above 10 percent for 15 minutesBelow 5 percent is normal on any consolidated host. Brief spikes during boot storms are not incidents. The 15 minute window is what removes the noise.
Memory reclamation activeSwap in rate above zero for 5 minutes on any hostBallooning is the hypervisor working as designed. Swap in means it has run out of gentler options and guest performance is already affected.
Datastore projected fullLinear projection crosses 90 percent inside 14 daysAlerting at a static 85 percent fires on datastores stable for two years. Alerting on trajectory gives you time to act and stays quiet otherwise.
Collector healthNo new points in the bucket for 10 minutesThe rule that catches every other rule failing silently. Build it first.

Do not duplicate vCenter alarms. vCenter already alarms on host connectivity, datastore full and hardware health, and it does it better because it has state you do not. Use Grafana for trend, contention and cross system correlation, which are exactly the things vCenter alarms cannot express. Two systems alerting on the same condition means both get muted.

11. Scaling, and what breaks first

SymptomCauseFix
Fault mentioning vpxd.stats.maxQueryMetricsCluster aggregation subqueries exceeding the cap, per section 03Exclude cluster metrics, or raise the setting above your VM count. Do not set it to -1.
Collection takes longer than the intervalToo many objects for the configured concurrency, or a slow vCenter databaseRaise collect_concurrency in steps of 2 and measure. If vCenter CPU rises, the answer is a second collector sharded by datacenter instead.
Sawtooth gaps in graphsHistorical objects polled faster than their 5 minute source granularityMove them to the 300 second input block. This is decision 02 asserting itself.
Dashboards slow after several monthsSeries cardinality, usually from wildcard VM counters or per instance detailMeasure cardinality per measurement, prune the include list, and move long range panels to the downsampled bucket.
Series for deleted VMs persistTime series databases do not delete, they expireLet retention handle it, and filter panels on a recent time window rather than the full history.
Everything stops after a vCenter upgradeCertificate replaced, or the service account password expiredThe collector health alert from section 10 catches this in ten minutes instead of on Monday.

Above roughly 3000 VMs

  • Shard collectors by vCenter, then by datacenter within a vCenter. Each Telegraf instance handles one shard.
  • Separate InfluxDB onto dedicated hardware with SSD backed storage, and give it memory proportional to your series count.
  • Drop VM level real time collection for non production clusters and keep only the 5 minute rollups there. Half your inventory usually does not need 60 second data.
  • Consider VictoriaMetrics or Mimir if a single InfluxDB instance becomes the bottleneck. Telegraf can write to either without changing the input configuration, which is a large part of why the collector and the store are separate components in the first place.

Closing thought

The install is the easy part and it is what every guide covers. What determines whether this pipeline is still running and trusted a year from now is the set of choices in section 04: splitting the collectors by interval class, refusing to wildcard VM counters, keeping the unit arithmetic in the query layer where it can be audited, and getting array side latency onto the same time axis as guest latency. Get those right and the dashboards mostly build themselves.

Tested against vCenter 8.0, Telegraf 1.34, InfluxDB 2.7 and Grafana 11.6. Counter names and plugin options are stable across recent versions, but check the Telegraf vSphere plugin documentation before copying configuration verbatim. Configuration here assumes a lab or non production first pass. Validate collection load against your own vCenter before pointing it at production.

Leave a Reply

Discover more from VMwareBlogs

Subscribe now to keep reading and get access to the full archive.

Continue reading