Upgrades are the operation where VCF earns its licence fee, and also where an unprepared estate discovers every shortcut it has taken. The good news is that the prechecks are genuinely thorough. The bad news is that they will fail on things nobody connected to upgrading, and reading them correctly is a skill.
Part 5, and the last of this series.
01. The dependency chain
1. RELEASE which target release applies to this domain
/v1/releases, /v1/target-upgrade-version
|
2. BUNDLES download what that release needs
/v1/bundles, depot settings must be configured first
|
3. UPGRADABLES what can move, and in what order
/v1/upgradables, /v1/sddc-manager-upgradable
|
4. PRECHECKS is the estate ready
/v1/system-prechecks, /v1/check-sets
|
5. UPGRADE execute, component by component
/v1/upgrades
|
6. TASKS watch, and read failed subtasks
/v1/tasks
SDDC Manager upgrades FIRST, and separately. It is the thing
that performs every other upgrade, so it cannot upgrade itself
as part of the same run.
The ordering is not advisory. The
upgradablesAPI tells you what is eligible now, and that set changes as you complete each step. A component may be ineligible simply because something it depends on has not moved yet. If an expected component is missing from the list, that is usually the answer rather than a fault.
02. Why prechecks fail on things that look unrelated
A precheck is not testing the upgrade. It is testing whether the estate is in the state VCF believes it is in. Anything that has drifted since the last operation shows up here, which is why prechecks are the place stale problems surface.
| Precheck failure | What it is really telling you | Where it came from |
|---|---|---|
| Certificate or trust validation | SDDC Manager holds a thumbprint the component no longer presents | A certificate replaced outside VCF, possibly months ago |
| Password expired or expiring | A managed credential is close to or past expiry | A password changed directly on the component |
| Inventory mismatch | A host or cluster differs between SDDC Manager and vCenter | A change made in the vSphere Client instead of through VCF |
| Version mismatch | A component is not at the version VCF recorded | Out-of-band patching |
| Resource lock held | A previous workflow never completed | A failed task nobody cleared |
| NTP or DNS | Same checks as bring-up, still enforced | Infrastructure change since deployment |
| Disk space | Not enough room for the bundle or its extraction | Log growth on an appliance |
| Licensing state | Domain or vCenter in subscription | A licensing transition. Blocks all day-N operations |
Six of those eight are drift, not readiness. That is the insight worth carrying: a precheck failure usually means somebody made a change outside VCF at some point, and the upgrade is simply the first operation strict enough to notice. This is exactly the drift the reconciliation script from the SDDC Manager and vCenter post detects, which is why running that monthly turns upgrade day from an investigation into a formality.
03. A precheck driver that groups by cause
#!/usr/bin/env python3
"""vcf_precheck.py - run prechecks, group failures, and gate the upgrade.
Groups results by likely cause rather than listing them flat, because
fifteen failures with one root cause is a very different situation
from fifteen independent problems.
Read only unless --start-upgrade is given, and that requires a clean
precheck in the same run.
Uses SddcManager from vcf_client.py (part 1).
"""
import argparse
import logging
import re
import sys
from collections import defaultdict
from vcf_client import SddcManager, VcfApiError, ResourceLocked
logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")
log = logging.getLogger("precheck")
# map text seen in a failed check to a cause and the action that clears it
CAUSES = [
(r"certificat|thumbprint|trust|ssl", "CERTIFICATE TRUST",
"A component presents a certificate SDDC Manager did not record. "
"Re-run discovery for that resource, or fix the trust store."),
(r"password|credential|expir", "CREDENTIALS",
"A managed password is expired or expiring. Rotate it through "
"the credentials API, not on the component directly."),
(r"lock|already locked|in progress", "RESOURCE LOCK",
"A previous workflow holds a lock. Resolve or fail that task first."),
(r"ntp|time sync|clock", "TIME",
"NTP unreachable or out of sync. Fix before anything else; "
"time problems cause certificate errors downstream."),
(r"dns|resolve|fqdn|reverse", "NAME RESOLUTION",
"Forward or reverse DNS is missing for a component."),
(r"disk|space|storage full|capacity", "DISK SPACE",
"An appliance lacks room for the bundle or its extraction."),
(r"version|build|mismatch|inventor", "INVENTORY DRIFT",
"SDDC Manager and reality disagree. Usually an out-of-band change."),
(r"licen|subscription", "LICENSING",
"Licence state blocks the operation. A domain in subscription "
"supports no day-N operations at all."),
(r"connect|unreachable|timeout|not responding", "CONNECTIVITY",
"A component is unreachable from SDDC Manager."),
]
def classify(text):
low = (text or "").lower()
for pattern, label, advice in CAUSES:
if re.search(pattern, low):
return label, advice
return "OTHER", "No pattern matched. Read the full message and remedy."
def walk(node, out):
"""Precheck results nest arbitrarily deep. Flatten every leaf check."""
if isinstance(node, dict):
if node.get("resultStatus") or node.get("status"):
out.append(node)
for key in ("checks", "resources", "presultDetails", "resultDetails",
"nestedChecks", "subTasks", "elements"):
for child in (node.get(key) or []):
walk(child, out)
elif isinstance(node, list):
for child in node:
walk(child, out)
def main():
p = argparse.ArgumentParser()
p.add_argument("--sddc", required=True)
p.add_argument("--user", required=True)
p.add_argument("--password", required=True)
p.add_argument("--domain", help="domain name; omit for the whole system")
p.add_argument("--start-upgrade", action="store_true",
help="start the upgrade if prechecks are clean")
a = p.parse_args()
vcf = SddcManager(a.sddc, a.user, a.password)
# ---- context first: what release, what bundles, what is upgradable
domains = {d["name"]: d for d in vcf.paginate("v1", "/domains")}
target = None
if a.domain:
d = domains.get(a.domain)
if not d:
sys.exit(f"domain '{a.domain}' not found. "
f"Known: {', '.join(domains)}")
target = d["id"]
print(f"domain {a.domain} status={d.get('status')}")
try:
bundles = list(vcf.paginate("v1", "/bundles"))
pending = [b for b in bundles
if str(b.get("downloadStatus", "")).upper()
not in ("SUCCESSFUL", "SUCCESS", "COMPLETED")]
print(f"bundles: {len(bundles)} known, {len(pending)} not downloaded")
for b in pending[:5]:
print(f" pending: {b.get('type')} {b.get('version')} "
f"[{b.get('downloadStatus')}]")
except VcfApiError as e:
log.warning("bundle query failed: %s", e)
try:
upgradables = list(vcf.paginate("v1", "/upgradables"))
print(f"\nupgradable components: {len(upgradables)}")
for u in upgradables:
print(f" {str(u.get('bundleType')):<18} "
f"{str(u.get('status')):<14} {u.get('bundleId','')[:18]}")
if not upgradables:
print(" nothing eligible right now. Either you are current, "
"or a prerequisite component has not been upgraded yet.")
except VcfApiError as e:
log.warning("upgradables query failed: %s", e)
# ---- prechecks
print("\nrunning prechecks, this takes several minutes")
spec = {"resources": [{"resourceId": target, "resourceType": "DOMAIN"}]} \
if target else {}
try:
run = vcf.post("v1", "/system-prechecks", json=spec)
except VcfApiError as e:
log.error("could not start prechecks:\n%s", e)
return 2
task_id = run.get("id") or run.get("taskId")
result = vcf.wait_for_task(task_id, poll=20, timeout=3600) \
if task_id else run
checks = []
walk(result, checks)
grouped = defaultdict(list)
passed = warned = 0
for c in checks:
status = str(c.get("resultStatus") or c.get("status") or "").upper()
name = c.get("name") or c.get("description") or "(unnamed)"
msg = ""
err = c.get("errorResponse") or {}
if isinstance(err, dict):
msg = err.get("message") or ""
remedy = err.get("remediationMessage") or ""
else:
remedy = ""
if status in ("FAILED", "ERROR", "RED"):
label, advice = classify(f"{name} {msg}")
grouped[label].append((name, msg, remedy, advice))
elif status in ("WARNING", "YELLOW"):
warned += 1
elif status:
passed += 1
total_failed = sum(len(v) for v in grouped.values())
print(f"\npassed {passed} warnings {warned} FAILED {total_failed}")
for label in sorted(grouped, key=lambda k: -len(grouped[k])):
items = grouped[label]
print(f"\n{'=' * 66}\n{label} ({len(items)} check(s))")
print(f" {items[0][3]}")
for name, msg, remedy, _ in items[:6]:
print(f"\n - {name}")
if msg:
print(f" {msg[:200]}")
if remedy:
print(f" remedy: {remedy[:200]}")
if len(items) > 6:
print(f"\n ... and {len(items) - 6} more in this group")
if total_failed:
print(f"\n{'=' * 66}")
print(f"{len(grouped)} distinct cause(s) behind {total_failed} "
"failed check(s).")
print("Fix by cause, not check by check. One certificate problem "
"often produces a dozen failures.")
return 2
print("\nprechecks clean")
if not a.start_upgrade:
print("Re-run with --start-upgrade to proceed.")
return 0
if not upgradables:
print("nothing eligible to upgrade")
return 0
print("\nUpgrade order reminder: SDDC Manager first and separately, "
"then NSX, then vCenter, then ESXi.")
if input("Type UPGRADE to continue: ").strip() != "UPGRADE":
return 1
for u in upgradables:
if str(u.get("status", "")).upper() not in ("AVAILABLE", "ELIGIBLE"):
continue
body = {"bundleId": u.get("bundleId"),
"resourceType": u.get("resourceType"),
"resourceUpgradeSpecs": u.get("resourceUpgradeSpecs", [])}
try:
t = vcf.post("v1", "/upgrades", json=body)
log.info("upgrade started for %s", u.get("bundleType"))
vcf.wait_for_task(t.get("id"), poll=60, timeout=6 * 3600)
except ResourceLocked as e:
log.error("blocked by a lock:\n%s", e)
return 2
except VcfApiError as e:
log.error("upgrade failed for %s:\n%s", u.get("bundleType"), e)
return 2
return 0
if __name__ == "__main__":
sys.exit(main())
Grouping by cause is the point of that script. A flat list of fifteen failures looks like fifteen problems and reads as a disaster. Grouped, it is often one expired credential producing twelve failures and one genuine disk space issue. The response to those two situations is completely different, and the flat list actively hides which one you are in.
04. Order of operations on the day
| Step | Component | Why here | Rough duration |
|---|---|---|---|
| 1 | SDDC Manager | It performs every other upgrade. It cannot be part of its own run | 30 to 60 minutes |
| 2 | Re-run prechecks | The new SDDC Manager may check things the old one did not | 10 minutes |
| 3 | NSX | Managers then edges then host components. Longest single step | Several hours |
| 4 | vCenter | After NSX, before hosts | 1 to 2 hours |
| 5 | ESXi | Rolling, one host at a time. Bounded by evacuation time | See the evacuation arithmetic |
| 6 | Drift audit | Confirm SDDC Manager and vCenter still agree afterwards | Minutes |
Step 5 is where the vSphere cluster design arithmetic comes back. ESXi upgrade time is evacuation time plus reboot time, multiplied by host count, and that is usually the dominant term in the whole window. If you sized the cluster for M = 1, you patch one host at a time and the window is what it is.
05. Habits that make upgrades boring
- Run prechecks monthly, not before upgrades. They are read-only and they surface drift while it is still cheap. A precheck run in a quiet week is a report; the same run on upgrade day is an incident.
- Download bundles days ahead. Depot connectivity and disk space are avoidable reasons to lose a window.
- Never change anything through the component’s own UI. Every certificate, password and host change made outside VCF becomes a precheck failure later, with a message that names the symptom rather than your change.
- Clear failed tasks when they happen. A failed task from six weeks ago holding a resource lock will refuse your upgrade with an error about locking, not about upgrading.
- Check the compatibility matrix before planning. The
compatibility-matrixandv-san-hclAPIs answer whether the target combination is supported before you commit to a date.
That is the series: the two control planes and the token model, bring-up, network pools and commissioning, workload domain design, and lifecycle. The through-line is that VCF is strict because it maintains its own model of the estate, and almost every difficult failure is that model disagreeing with reality. Keep them in agreement and the platform is genuinely low-drama.
API categories, the versioning model and error semantics are taken from the published VMware Cloud Foundation API reference. Endpoint paths, request body shapes and precheck result structures vary between releases, so confirm against the operation index for your version before scheduling any of this. Durations given are indicative only. Nothing here is official guidance from VMware or Broadcom.
Leave a Reply