Creating a workload domain takes an afternoon. Living with the choices made in that afternoon takes years. This part is about the handful of decisions inside domain creation that are effectively permanent, and about the allocation ceilings that will refuse a domain you have already promised somebody.
01. What a workload domain actually bundles
A WORKLOAD DOMAIN is a bundle of four decisions:
1. its own vCenter always dedicated to the domain
2. an NSX cluster dedicated, or SHARED with other domains
3. an SSO domain join the management SSO, or stand alone
4. one or more clusters each with a storage type and a host count
Only item 4 is easy to change afterwards.
Items 2 and 3 are the ones to think hardest about.
02. Decision one: share an NSX cluster, or not
| Dedicated NSX per domain | Shared NSX across domains | |
|---|---|---|
| Blast radius | An NSX problem affects one domain | An NSX problem affects every domain sharing it |
| Resource cost | Three manager appliances per domain | One set serves several domains |
| Upgrade coupling | Independent per domain | Coupled. You upgrade all of them together |
| Ceiling | None from sharing | A documented maximum number of domains may reuse one NSX cluster |
| Operational effect | More to patch, more IPs, more certificates | Removing a host is refused while NSX is unhealthy, for every sharing domain |
Exceed the sharing limit and creation fails with NSXT_CLUSTER_RESOURCE_ALLOCATION_NOT_AVAILABLE_DUE_TO_MAXIMUM_NUMBER_OF_DOMAINS_REUSING_NSXT_REACHED. Note what that error is: an allocation failure at creation time. Nothing warns you as you approach the ceiling. The design has to account for it.
The question that settles it: would you accept these domains going down together? If two domains serve the same business function and share a maintenance window, sharing NSX is reasonable. If one is production and one is a regulated environment with separate change control, dedicate. Sharing infrastructure between things that must fail independently is how a design fails its own requirements.
03. Decision two: join the management SSO domain, or stand alone
Joining gives you a single identity source, shared roles and tags, and cross-vCenter vMotion between domains. Standing alone gives you isolation. There is also a hard ceiling.
MANAGEMENT_SSO_RESOURCE_ALLOCATION_NOT_AVAILABLE_DUE_TO_
MAXIMUM_NUMBER_OF_JOINED_DOMAINS_REACHED
"Allocation of Management SSO {0} is not available as the
maximum supported number ({1}) of joined domains is reached"
The limit is reported in the error, not before you hit it.
Count your planned domains at design time and check the
configuration maximums for your release.
This pairs with the SSO topology reasoning from the vSphere scale post: the deciding question is whether you migrate VMs between these domains as a normal operation. If you never do, you are buying a shared identity fault domain and consuming a scarce allocation slot for a feature you do not use.
04. Decision three: storage type, per cluster, permanently
Covered in part 3 from the host side; here is the consequence at domain level. A cluster’s storage type cannot be changed in place. Changing it means building a new cluster with the correct type, migrating every workload, and removing the old one. Whatever you choose on day one is what that cluster is.
The second-order effect is on your spare capacity. Hosts are commissioned with a storage type, so a pool of unassigned vSAN hosts cannot rescue an NFS cluster that needs capacity urgently. If you run mixed storage types, you need spare hosts of each type, and that is a real cost that belongs in the design rather than being discovered during an incident.
05. What blocks removing a domain later
Worth knowing on the way in, because these are the conditions that make a domain sticky.
| Blocker | Error code | What you must do first |
|---|---|---|
| A cluster in the domain is stretched | PUBLIC_REMOVE_DOMAIN_NOT_SUPPORTED_WHEN_STRETCHED_CLUSTERS_ARE_PRESENT | Unstretch every stretched cluster |
| Cross vCenter datastore mounts exist | PUBLIC_REMOVE_DOMAIN_NOT_SUPPORTED_WHEN_XVC_DATASTORES_ARE_PRESENT | Unmount them and remove the datastore sources |
| A cluster is stretched | FAILED_TO_REMOVE_STRETCHED_CLUSTER | Unstretch before removing the cluster |
| Domain or vCenter is in subscription licensing | VALIDATION_OF_DOMAIN_LICENSING_STATE_FAILED | Transition out of subscription. No day-N operations at all until you do |
That last row is worth reading twice. The message states that a domain in subscription does not support any day-N operations. That is not a removal-specific block; it stops routine work. If you are planning a licensing change, understand its operational effect before the change window, not during it.
Also relevant to stretching: PUBLIC_UNSTRETCH_NOT_SUPPORTED exists. Stretching a cluster is not universally reversible, so treat the decision to stretch as close to permanent and confirm the unstretch path for your release before you rely on it.
06. A design validator to run before you build
#!/usr/bin/env python3
"""vcf_domain_plan.py - check a planned domain against current state.
Answers the questions that are cheap now and expensive later:
* are there enough unassigned hosts, of the right storage type?
* how many domains already share the NSX cluster you intend to reuse?
* how many domains have already joined the management SSO domain?
* is anything holding a resource lock that will block creation?
* do the network pools have addresses left?
Read only. Uses SddcManager from vcf_client.py (part 1).
"""
import argparse
import logging
import sys
from collections import Counter, defaultdict
from vcf_client import SddcManager, VcfApiError
logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")
log = logging.getLogger("plan")
def main():
p = argparse.ArgumentParser()
p.add_argument("--sddc", required=True)
p.add_argument("--user", required=True)
p.add_argument("--password", required=True)
p.add_argument("--hosts-needed", type=int, required=True)
p.add_argument("--storage-type", required=True,
help="VSAN, VSAN_ESA, NFS, VMFS_FC or VVOL")
p.add_argument("--reuse-nsx", help="NSX cluster ID you intend to share")
p.add_argument("--join-sso", action="store_true",
help="plan to join the management SSO domain")
a = p.parse_args()
vcf = SddcManager(a.sddc, a.user, a.password)
findings = []
# ---------- capacity, by storage type
hosts = list(vcf.paginate("v1", "/hosts"))
free = [h for h in hosts
if str(h.get("status", "")).startswith("UNASSIGNED")]
by_type = Counter(str(h.get("storageType") or "UNKNOWN").upper()
for h in free)
print("UNASSIGNED HOSTS BY STORAGE TYPE")
for t, n in sorted(by_type.items()):
marker = " <-- your type" if t == a.storage_type.upper() else ""
print(f" {t:<12} {n:>4}{marker}")
available = by_type.get(a.storage_type.upper(), 0)
if available < a.hosts_needed:
findings.append(
f"need {a.hosts_needed} {a.storage_type} hosts, "
f"{available} unassigned. Commission {a.hosts_needed - available} "
"more, and remember storage type is fixed at commission time")
# ---------- NSX sharing
domains = list(vcf.paginate("v1", "/domains"))
nsx_use = defaultdict(list)
for d in domains:
nsx = d.get("nsxtCluster") or d.get("nsxCluster") or {}
nid = nsx.get("id") if isinstance(nsx, dict) else None
if nid:
nsx_use[nid].append(d.get("name"))
print(f"\nDOMAINS: {len(domains)}")
for d in domains:
print(f" {d.get('name'):<28} type={d.get('type'):<12} "
f"status={d.get('status')}")
if nsx_use:
print("\nNSX CLUSTER SHARING")
for nid, names in nsx_use.items():
print(f" {nid[:20]:<22} used by {len(names)}: {', '.join(names)}")
if a.reuse_nsx:
current = len(nsx_use.get(a.reuse_nsx, []))
print(f"\nplanned NSX cluster already serves {current} domain(s); "
f"yours would be number {current + 1}")
findings.append(
f"confirm the maximum domains per NSX cluster for your release. "
f"Exceeding it fails at creation with an allocation error, "
f"not a warning")
# ---------- SSO joins
if a.join_sso:
joined = [d for d in domains if d.get("ssoName") or d.get("ssoId")]
print(f"\ndomains carrying an SSO reference: {len(joined)}")
findings.append(
"confirm the maximum joined domains for the management SSO "
"domain. The ceiling is reported only when you hit it")
# ---------- network pool headroom
print("\nNETWORK POOLS")
for pool in vcf.paginate("v1", "/network-pools"):
nets = pool.get("networks") or []
desc = ", ".join(f"{n.get('type')}/{n.get('vlanId')}" for n in nets)
print(f" {pool.get('name'):<28} {desc}")
findings.append(
"verify each pool has at least one free address per new host per "
"network before you build. Exhaustion fails mid-operation")
# ---------- locks and open work
open_tasks = [t for t in vcf.paginate("v1", "/tasks")
if t.get("status") in ("IN_PROGRESS", "Pending", "Failed")]
if open_tasks:
print(f"\nOPEN OR FAILED TASKS: {len(open_tasks)}")
for t in open_tasks[:10]:
print(f" {t.get('status'):<12} {t.get('name')} [{t.get('id')}]")
findings.append(
f"{len(open_tasks)} task(s) open or failed. These may hold "
"resource locks that will refuse your domain creation")
# ---------- verdict
print("\n" + "=" * 66)
if findings:
print(f"{len(findings)} thing(s) to resolve or confirm:\n")
for i, f in enumerate(findings, 1):
print(f" {i}. {f}\n")
else:
print("no blockers found from the API's point of view")
print("Decisions this tool cannot make for you:")
print(" * dedicated or shared NSX blast radius versus cost")
print(" * join or stand alone on SSO xVC-vMotion versus isolation")
print(" * storage type per cluster cannot be changed in place")
print(" * stretch or not may not be reversible")
return 2 if findings else 0
if __name__ == "__main__":
sys.exit(main())
07. The decision register for a domain
| Decision | Reversible? | Cost of getting it wrong |
|---|---|---|
| Domain name and vCenter FQDN | No | Rebuild |
| SSO domain join | Effectively no | Rebuild the domain |
| NSX dedicated or shared | No | Rebuild, or live with the coupling |
| Storage type per cluster | No, not in place | New cluster plus full workload migration |
| Stretched or standard | Sometimes not | Confirm the unstretch path for your release before relying on it |
| Host count per cluster | Yes | Expansion, within pool and licence limits |
| Cluster count per domain | Yes | Add clusters as needed |
Five of the seven rows are one-way. That ratio is the argument for spending real time on a domain design before anyone opens the API, and for recording the rejected alternatives alongside the choices.
Next, and last in this series: prechecks and upgrades, and how to read a failed precheck properly.
Allocation ceilings, removal preconditions and error codes are taken from the published VMware Cloud Foundation API reference. The specific maximum values are release dependent and are reported in the error message rather than published inline, so check the configuration maximums for your version at design time. Nothing here is official guidance from VMware or Broadcom.
Leave a Reply