Bring-up is the one operation in VCF that you cannot partially undo. If it fails at 70 percent you are usually rebuilding hosts and starting again. That is why Cloud Builder separates validation from execution, and why the single most valuable habit you can form is to run validation repeatedly until it is completely clean before you ever start a build.
Part 2 of the series. Part 1 covered the two control planes and the client. This part uses the Cloud Builder half of it.
01. The shape of the operation
1. Deploy the Cloud Builder OVA manual, once
2. Build the SDDC spec JSON, this is the whole design
3. POST the spec for VALIDATION repeat until clean
4. Read every failed check, fix, re-run the part that matters
5. POST the spec for EXECUTION now it should just work
6. Poll until complete 1 to 4 hours typically
7. Log in to SDDC Manager Cloud Builder's job is done
Steps 3 and 4 are cheap and repeatable.
Step 5 is expensive and effectively one-way.
Spend your time in 3 and 4.
The sddc API category on Cloud Builder covers this: submitting a spec for validation, retrieving validation results, starting the bring-up, and querying its progress. Check the operation index for the exact paths on your release, since resource versions move independently.
The spec is your design document in machine-readable form. Every hostname, VLAN, IP pool, licence key and password ends up in it. Treat it as a versioned artefact in source control, not as something you generate once from a spreadsheet. When you build the second environment, the diff between the two specs is the difference between the two sites.
02. What the spec has to describe
Field names vary by release, so build yours from the schema for your version rather than copying anyone’s example, including mine. What does not change is the set of things you must have decided before you start typing.
| Section | What you are declaring | Decide this beforehand |
|---|---|---|
| Infrastructure | NTP servers, DNS servers, DNS zone, subdomain | Both DNS directions must already resolve for every FQDN in the spec |
| Management network | Management VLAN, subnet, gateway, MTU | Static addressing only. DHCP is rejected |
| vMotion and vSAN networks | VLANs, subnets, IP ranges | Ranges must be large enough for growth, not just for today |
| ESXi hosts | FQDN, credentials, per-host addressing | All hosts at the same build, same disk layout |
| vCenter | FQDN, IP, root and SSO passwords, SSO domain name | The SSO domain name is essentially permanent |
| NSX | Manager FQDNs, VIP, transport VLAN, licence | The VIP needs its own DNS record, and it is the one people forget |
| SDDC Manager | FQDN, IP, credentials | |
| vSphere Distributed Switch | Name, uplinks, MTU | MTU must be consistent from vmk to fabric |
Two entries in that table are effectively irreversible. The SSO domain name is baked into the management domain and changing it later is a rebuild. The NSX VIP FQDN must exist in DNS before you start, and its absence produces a failure late in bring-up rather than at validation time on some releases.
03. The validation loop
#!/usr/bin/env python3
"""vcf_bringup.py - validate an SDDC spec, then optionally execute it.
Validation is cheap and repeatable. Execution is not. This defaults to
validate-only and refuses to execute unless validation passed in the
same run and --execute was given explicitly.
Uses CloudBuilder from vcf_client.py (part 1).
"""
import argparse
import json
import logging
import sys
import time
from vcf_client import CloudBuilder, VcfApiError
logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")
log = logging.getLogger("bringup")
TERMINAL = {"COMPLETED_WITH_SUCCESS", "COMPLETED_WITH_FAILURE",
"FAILED", "SUCCEEDED", "COMPLETED"}
def poll(cb, path, poll_secs=20, timeout=6 * 3600):
"""Poll a Cloud Builder resource until it reaches a terminal state."""
deadline = time.time() + timeout
last = None
while time.time() < deadline:
body = cb.request("GET", "v1", path)
status = (body.get("executionStatus")
or body.get("status")
or body.get("sddcStatus"))
if status != last:
log.info("status: %s", status)
last = status
if status and status.upper() in TERMINAL:
return body
time.sleep(poll_secs)
raise TimeoutError(f"{path} did not finish within {timeout}s")
def report(result):
"""Print every check, failures first, with the description and remedy."""
checks = []
for section in (result.get("validationChecks")
or result.get("resultDetails")
or []):
checks.append(section)
for nested in (section.get("nestedValidationChecks") or []):
checks.append(nested)
def rank(c):
s = str(c.get("resultStatus") or c.get("status") or "").upper()
return {"FAILED": 0, "ERROR": 0, "WARNING": 1}.get(s, 2)
checks.sort(key=rank)
failed = warned = passed = 0
for c in checks:
status = str(c.get("resultStatus") or c.get("status") or "?").upper()
name = c.get("description") or c.get("name") or "(unnamed check)"
if status in ("FAILED", "ERROR"):
failed += 1
print(f"\nFAILED {name}")
for e in (c.get("errorResponse", {}) or {}).get("nestedErrors", []) \
or [c.get("errorResponse") or {}]:
if not e:
continue
if e.get("message"):
print(f" {e['message']}")
if e.get("remediationMessage"):
print(f" remedy: {e['remediationMessage']}")
elif status == "WARNING":
warned += 1
print(f"WARNING {name}")
else:
passed += 1
print(f"\npassed {passed} warnings {warned} FAILED {failed}")
return failed
def main():
p = argparse.ArgumentParser()
p.add_argument("--cloud-builder", required=True)
p.add_argument("--password", required=True, help="admin account password")
p.add_argument("--spec", required=True, help="path to the SDDC spec JSON")
p.add_argument("--execute", action="store_true",
help="run bring-up if validation passes. One way.")
a = p.parse_args()
spec = json.load(open(a.spec))
cb = CloudBuilder(a.cloud_builder, a.password)
# ---- validate
log.info("submitting spec for validation")
v = cb.request("POST", "v1", "/sddcs/validations", json=spec)
vid = v.get("id")
result = poll(cb, f"/sddcs/validations/{vid}") if vid else v
failures = report(result)
if failures:
log.error("%d check(s) failed. Fix these and re-run validation.",
failures)
return 2
log.info("validation clean")
if not a.execute:
log.info("validate-only. Re-run with --execute to build.")
return 0
# ---- execute
print("\nThis starts bring-up. It cannot be cleanly undone.")
if input("Type the management domain name to confirm: ").strip() \
!= spec.get("sddcId", ""):
log.info("aborted")
return 1
run = cb.request("POST", "v1", "/sddcs", json=spec)
rid = run.get("id")
log.info("bring-up started, id %s", rid)
final = poll(cb, f"/sddcs/{rid}", poll_secs=60, timeout=8 * 3600)
status = str(final.get("status") or final.get("sddcStatus") or "").upper()
if "SUCCESS" in status or status == "COMPLETED":
log.info("bring-up complete. SDDC Manager is now the control plane.")
return 0
log.error("bring-up failed: %s", status)
report(final)
return 2
if __name__ == "__main__":
sys.exit(main())
Two deliberate design choices. The script cannot execute unless validation passed in the same run, so there is no path where a stale clean result authorises a build against a changed spec. And it requires you to type the management domain name to confirm, because an accidental bring-up against the wrong Cloud Builder is a genuinely expensive mistake.
04. What the validations are testing
Validation failures cluster into five families. Recognising the family tells you which team to talk to.
| Family | What is really being tested | Whose problem |
|---|---|---|
| Name resolution | Every FQDN resolves forward, and every IP resolves back to the same name | DNS team. Reverse records are the usual gap |
| Time | NTP reachable from hosts and appliances, and actually in sync, not merely configured | Network or platform. Reachable is not the same as synchronised |
| Host reachability and credentials | SSH login works with the supplied account, and the account is not locked | You. Test one by hand first |
| Host configuration | Static management addressing, expected ESXi build, disks eligible for the chosen storage | Build process. Usually a host imaged differently from the rest |
| Network | VLANs present on the uplinks, MTU consistent end to end, gateways reachable | Network team. MTU is the one that passes small tests and fails large ones |
Run these by hand before you even build the spec. They take ten minutes and they catch most of what validation would tell you an hour later.
# from a machine on the management network, for every host and appliance FQDN
for n in esxi-01 esxi-02 esxi-03 esxi-04 vcenter-mgmt nsx-a nsx-b nsx-vip sddc-mgr; do
fqdn="$n.mgmt.lab.local"
ip=$(dig +short "$fqdn")
rev=$(dig +short -x "$ip" 2>/dev/null)
printf '%-28s %-16s %s\n' "$fqdn" "${ip:-NO-A-RECORD}" "${rev:-NO-PTR}"
done
# every row must have both. NO-PTR is the single most common bring-up blocker.
# on each host: static addressing, NTP synchronised, correct build
for h in esxi-01 esxi-02 esxi-03 esxi-04; do
echo "== $h"
ssh root@$h.mgmt.lab.local '
esxcli network ip interface ipv4 get -i vmk0 | tail -1
esxcli system ntp get | grep -E "Enabled|Server"
vmware -vl
esxcli network nic list | head -4
'
done
# prove the MTU on the path you intend to use for storage or vSAN
ssh root@esxi-01.mgmt.lab.local 'vmkping -I vmk0 -s 8972 -d <gateway>'
Reverse DNS is the number one cause of a failed first attempt. Forward records get created because somebody has to reach the host. Reverse records get forgotten because nothing else needs them. VCF checks both.
05. If bring-up fails partway
- Read the failed subtask, not the top-level status. The parent says the workflow failed; the child says which host and which step.
- Retry is sometimes available and sometimes not. If the failure was transient, such as an NTP blip, retry is reasonable. If it was configuration, fix the configuration first or you will fail at the same point.
- Do not hand-fix things in vCenter and then resume. You will produce exactly the inventory drift described in the earlier post on SDDC Manager and vCenter disagreeing, except at the worst possible moment.
- Assume a rebuild is possible. Re-imaging four hosts and re-running a validated spec is often faster and always cleaner than nursing a half-built management domain.
Next: network pools and commissioning hosts at scale, within the documented payload limits.
Endpoint paths and spec field names for the sddc category vary between VCF releases; build your spec from the schema for your version and confirm paths against the operation index. The validation and bring-up flow, and the separation between them, is as described in the VMware Cloud Foundation API reference. Nothing here is official guidance from VMware or Broadcom.
Leave a Reply