Bring-up is the one operation in VCF that you cannot partially undo. If it fails at 70 percent you are usually rebuilding hosts and starting again. That is why Cloud Builder separates validation from execution, and why the single most valuable habit you can form is to run validation repeatedly until it is completely clean before you ever start a build.

Part 2 of the series. Part 1 covered the two control planes and the client. This part uses the Cloud Builder half of it.

01. The shape of the operation

1. Deploy the Cloud Builder OVA          manual, once
2. Build the SDDC spec                   JSON, this is the whole design
3. POST the spec for VALIDATION          repeat until clean
4. Read every failed check, fix, re-run  the part that matters
5. POST the spec for EXECUTION           now it should just work
6. Poll until complete                   1 to 4 hours typically
7. Log in to SDDC Manager                Cloud Builder's job is done

Steps 3 and 4 are cheap and repeatable.
Step 5 is expensive and effectively one-way.
Spend your time in 3 and 4.

The sddc API category on Cloud Builder covers this: submitting a spec for validation, retrieving validation results, starting the bring-up, and querying its progress. Check the operation index for the exact paths on your release, since resource versions move independently.

The spec is your design document in machine-readable form. Every hostname, VLAN, IP pool, licence key and password ends up in it. Treat it as a versioned artefact in source control, not as something you generate once from a spreadsheet. When you build the second environment, the diff between the two specs is the difference between the two sites.

02. What the spec has to describe

Field names vary by release, so build yours from the schema for your version rather than copying anyone’s example, including mine. What does not change is the set of things you must have decided before you start typing.

SectionWhat you are declaringDecide this beforehand
InfrastructureNTP servers, DNS servers, DNS zone, subdomainBoth DNS directions must already resolve for every FQDN in the spec
Management networkManagement VLAN, subnet, gateway, MTUStatic addressing only. DHCP is rejected
vMotion and vSAN networksVLANs, subnets, IP rangesRanges must be large enough for growth, not just for today
ESXi hostsFQDN, credentials, per-host addressingAll hosts at the same build, same disk layout
vCenterFQDN, IP, root and SSO passwords, SSO domain nameThe SSO domain name is essentially permanent
NSXManager FQDNs, VIP, transport VLAN, licenceThe VIP needs its own DNS record, and it is the one people forget
SDDC ManagerFQDN, IP, credentials
vSphere Distributed SwitchName, uplinks, MTUMTU must be consistent from vmk to fabric

Two entries in that table are effectively irreversible. The SSO domain name is baked into the management domain and changing it later is a rebuild. The NSX VIP FQDN must exist in DNS before you start, and its absence produces a failure late in bring-up rather than at validation time on some releases.

03. The validation loop

#!/usr/bin/env python3
"""vcf_bringup.py - validate an SDDC spec, then optionally execute it.

Validation is cheap and repeatable. Execution is not. This defaults to
validate-only and refuses to execute unless validation passed in the
same run and --execute was given explicitly.

Uses CloudBuilder from vcf_client.py (part 1).
"""

import argparse
import json
import logging
import sys
import time

from vcf_client import CloudBuilder, VcfApiError

logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")
log = logging.getLogger("bringup")

TERMINAL = {"COMPLETED_WITH_SUCCESS", "COMPLETED_WITH_FAILURE",
            "FAILED", "SUCCEEDED", "COMPLETED"}


def poll(cb, path, poll_secs=20, timeout=6 * 3600):
    """Poll a Cloud Builder resource until it reaches a terminal state."""
    deadline = time.time() + timeout
    last = None
    while time.time() < deadline:
        body = cb.request("GET", "v1", path)
        status = (body.get("executionStatus")
                  or body.get("status")
                  or body.get("sddcStatus"))
        if status != last:
            log.info("status: %s", status)
            last = status
        if status and status.upper() in TERMINAL:
            return body
        time.sleep(poll_secs)
    raise TimeoutError(f"{path} did not finish within {timeout}s")


def report(result):
    """Print every check, failures first, with the description and remedy."""
    checks = []
    for section in (result.get("validationChecks")
                    or result.get("resultDetails")
                    or []):
        checks.append(section)
        for nested in (section.get("nestedValidationChecks") or []):
            checks.append(nested)

    def rank(c):
        s = str(c.get("resultStatus") or c.get("status") or "").upper()
        return {"FAILED": 0, "ERROR": 0, "WARNING": 1}.get(s, 2)

    checks.sort(key=rank)

    failed = warned = passed = 0
    for c in checks:
        status = str(c.get("resultStatus") or c.get("status") or "?").upper()
        name = c.get("description") or c.get("name") or "(unnamed check)"
        if status in ("FAILED", "ERROR"):
            failed += 1
            print(f"\nFAILED  {name}")
            for e in (c.get("errorResponse", {}) or {}).get("nestedErrors", []) \
                     or [c.get("errorResponse") or {}]:
                if not e:
                    continue
                if e.get("message"):
                    print(f"        {e['message']}")
                if e.get("remediationMessage"):
                    print(f"        remedy: {e['remediationMessage']}")
        elif status == "WARNING":
            warned += 1
            print(f"WARNING {name}")
        else:
            passed += 1

    print(f"\npassed {passed}   warnings {warned}   FAILED {failed}")
    return failed


def main():
    p = argparse.ArgumentParser()
    p.add_argument("--cloud-builder", required=True)
    p.add_argument("--password", required=True, help="admin account password")
    p.add_argument("--spec", required=True, help="path to the SDDC spec JSON")
    p.add_argument("--execute", action="store_true",
                   help="run bring-up if validation passes. One way.")
    a = p.parse_args()

    spec = json.load(open(a.spec))
    cb = CloudBuilder(a.cloud_builder, a.password)

    # ---- validate
    log.info("submitting spec for validation")
    v = cb.request("POST", "v1", "/sddcs/validations", json=spec)
    vid = v.get("id")
    result = poll(cb, f"/sddcs/validations/{vid}") if vid else v

    failures = report(result)
    if failures:
        log.error("%d check(s) failed. Fix these and re-run validation.",
                  failures)
        return 2

    log.info("validation clean")
    if not a.execute:
        log.info("validate-only. Re-run with --execute to build.")
        return 0

    # ---- execute
    print("\nThis starts bring-up. It cannot be cleanly undone.")
    if input("Type the management domain name to confirm: ").strip() \
            != spec.get("sddcId", ""):
        log.info("aborted")
        return 1

    run = cb.request("POST", "v1", "/sddcs", json=spec)
    rid = run.get("id")
    log.info("bring-up started, id %s", rid)
    final = poll(cb, f"/sddcs/{rid}", poll_secs=60, timeout=8 * 3600)

    status = str(final.get("status") or final.get("sddcStatus") or "").upper()
    if "SUCCESS" in status or status == "COMPLETED":
        log.info("bring-up complete. SDDC Manager is now the control plane.")
        return 0

    log.error("bring-up failed: %s", status)
    report(final)
    return 2


if __name__ == "__main__":
    sys.exit(main())

Two deliberate design choices. The script cannot execute unless validation passed in the same run, so there is no path where a stale clean result authorises a build against a changed spec. And it requires you to type the management domain name to confirm, because an accidental bring-up against the wrong Cloud Builder is a genuinely expensive mistake.

04. What the validations are testing

Validation failures cluster into five families. Recognising the family tells you which team to talk to.

FamilyWhat is really being testedWhose problem
Name resolutionEvery FQDN resolves forward, and every IP resolves back to the same nameDNS team. Reverse records are the usual gap
TimeNTP reachable from hosts and appliances, and actually in sync, not merely configuredNetwork or platform. Reachable is not the same as synchronised
Host reachability and credentialsSSH login works with the supplied account, and the account is not lockedYou. Test one by hand first
Host configurationStatic management addressing, expected ESXi build, disks eligible for the chosen storageBuild process. Usually a host imaged differently from the rest
NetworkVLANs present on the uplinks, MTU consistent end to end, gateways reachableNetwork team. MTU is the one that passes small tests and fails large ones

Run these by hand before you even build the spec. They take ten minutes and they catch most of what validation would tell you an hour later.

# from a machine on the management network, for every host and appliance FQDN
for n in esxi-01 esxi-02 esxi-03 esxi-04 vcenter-mgmt nsx-a nsx-b nsx-vip sddc-mgr; do
  fqdn="$n.mgmt.lab.local"
  ip=$(dig +short "$fqdn")
  rev=$(dig +short -x "$ip" 2>/dev/null)
  printf '%-28s %-16s %s\n' "$fqdn" "${ip:-NO-A-RECORD}" "${rev:-NO-PTR}"
done
# every row must have both. NO-PTR is the single most common bring-up blocker.

# on each host: static addressing, NTP synchronised, correct build
for h in esxi-01 esxi-02 esxi-03 esxi-04; do
  echo "== $h"
  ssh root@$h.mgmt.lab.local '
    esxcli network ip interface ipv4 get -i vmk0 | tail -1
    esxcli system ntp get | grep -E "Enabled|Server"
    vmware -vl
    esxcli network nic list | head -4
  '
done

# prove the MTU on the path you intend to use for storage or vSAN
ssh root@esxi-01.mgmt.lab.local 'vmkping -I vmk0 -s 8972 -d <gateway>'

Reverse DNS is the number one cause of a failed first attempt. Forward records get created because somebody has to reach the host. Reverse records get forgotten because nothing else needs them. VCF checks both.

05. If bring-up fails partway

  • Read the failed subtask, not the top-level status. The parent says the workflow failed; the child says which host and which step.
  • Retry is sometimes available and sometimes not. If the failure was transient, such as an NTP blip, retry is reasonable. If it was configuration, fix the configuration first or you will fail at the same point.
  • Do not hand-fix things in vCenter and then resume. You will produce exactly the inventory drift described in the earlier post on SDDC Manager and vCenter disagreeing, except at the worst possible moment.
  • Assume a rebuild is possible. Re-imaging four hosts and re-running a validated spec is often faster and always cleaner than nursing a half-built management domain.

Next: network pools and commissioning hosts at scale, within the documented payload limits.

Endpoint paths and spec field names for the sddc category vary between VCF releases; build your spec from the schema for your version and confirm paths against the operation index. The validation and bring-up flow, and the separation between them, is as described in the VMware Cloud Foundation API reference. Nothing here is official guidance from VMware or Broadcom.

Leave a Reply

Discover more from VMwareBlogs

Subscribe now to keep reading and get access to the full archive.

Continue reading