Creating a workload domain takes an afternoon. Living with the choices made in that afternoon takes years. This part is about the handful of decisions inside domain creation that are effectively permanent, and about the allocation ceilings that will refuse a domain you have already promised somebody.

01. What a workload domain actually bundles

A WORKLOAD DOMAIN is a bundle of four decisions:

  1. its own vCenter               always dedicated to the domain
  2. an NSX cluster                dedicated, or SHARED with other domains
  3. an SSO domain                 join the management SSO, or stand alone
  4. one or more clusters          each with a storage type and a host count

Only item 4 is easy to change afterwards.
Items 2 and 3 are the ones to think hardest about.

02. Decision one: share an NSX cluster, or not

Dedicated NSX per domainShared NSX across domains
Blast radiusAn NSX problem affects one domainAn NSX problem affects every domain sharing it
Resource costThree manager appliances per domainOne set serves several domains
Upgrade couplingIndependent per domainCoupled. You upgrade all of them together
CeilingNone from sharingA documented maximum number of domains may reuse one NSX cluster
Operational effectMore to patch, more IPs, more certificatesRemoving a host is refused while NSX is unhealthy, for every sharing domain

Exceed the sharing limit and creation fails with NSXT_CLUSTER_RESOURCE_ALLOCATION_NOT_AVAILABLE_DUE_TO_MAXIMUM_NUMBER_OF_DOMAINS_REUSING_NSXT_REACHED. Note what that error is: an allocation failure at creation time. Nothing warns you as you approach the ceiling. The design has to account for it.

The question that settles it: would you accept these domains going down together? If two domains serve the same business function and share a maintenance window, sharing NSX is reasonable. If one is production and one is a regulated environment with separate change control, dedicate. Sharing infrastructure between things that must fail independently is how a design fails its own requirements.

03. Decision two: join the management SSO domain, or stand alone

Joining gives you a single identity source, shared roles and tags, and cross-vCenter vMotion between domains. Standing alone gives you isolation. There is also a hard ceiling.

MANAGEMENT_SSO_RESOURCE_ALLOCATION_NOT_AVAILABLE_DUE_TO_
  MAXIMUM_NUMBER_OF_JOINED_DOMAINS_REACHED

  "Allocation of Management SSO {0} is not available as the
   maximum supported number ({1}) of joined domains is reached"

The limit is reported in the error, not before you hit it.
Count your planned domains at design time and check the
configuration maximums for your release.

This pairs with the SSO topology reasoning from the vSphere scale post: the deciding question is whether you migrate VMs between these domains as a normal operation. If you never do, you are buying a shared identity fault domain and consuming a scarce allocation slot for a feature you do not use.

04. Decision three: storage type, per cluster, permanently

Covered in part 3 from the host side; here is the consequence at domain level. A cluster’s storage type cannot be changed in place. Changing it means building a new cluster with the correct type, migrating every workload, and removing the old one. Whatever you choose on day one is what that cluster is.

The second-order effect is on your spare capacity. Hosts are commissioned with a storage type, so a pool of unassigned vSAN hosts cannot rescue an NFS cluster that needs capacity urgently. If you run mixed storage types, you need spare hosts of each type, and that is a real cost that belongs in the design rather than being discovered during an incident.

05. What blocks removing a domain later

Worth knowing on the way in, because these are the conditions that make a domain sticky.

BlockerError codeWhat you must do first
A cluster in the domain is stretchedPUBLIC_REMOVE_DOMAIN_NOT_SUPPORTED_WHEN_STRETCHED_CLUSTERS_ARE_PRESENTUnstretch every stretched cluster
Cross vCenter datastore mounts existPUBLIC_REMOVE_DOMAIN_NOT_SUPPORTED_WHEN_XVC_DATASTORES_ARE_PRESENTUnmount them and remove the datastore sources
A cluster is stretchedFAILED_TO_REMOVE_STRETCHED_CLUSTERUnstretch before removing the cluster
Domain or vCenter is in subscription licensingVALIDATION_OF_DOMAIN_LICENSING_STATE_FAILEDTransition out of subscription. No day-N operations at all until you do

That last row is worth reading twice. The message states that a domain in subscription does not support any day-N operations. That is not a removal-specific block; it stops routine work. If you are planning a licensing change, understand its operational effect before the change window, not during it.

Also relevant to stretching: PUBLIC_UNSTRETCH_NOT_SUPPORTED exists. Stretching a cluster is not universally reversible, so treat the decision to stretch as close to permanent and confirm the unstretch path for your release before you rely on it.

06. A design validator to run before you build

#!/usr/bin/env python3
"""vcf_domain_plan.py - check a planned domain against current state.

Answers the questions that are cheap now and expensive later:
  * are there enough unassigned hosts, of the right storage type?
  * how many domains already share the NSX cluster you intend to reuse?
  * how many domains have already joined the management SSO domain?
  * is anything holding a resource lock that will block creation?
  * do the network pools have addresses left?

Read only. Uses SddcManager from vcf_client.py (part 1).
"""

import argparse
import logging
import sys
from collections import Counter, defaultdict

from vcf_client import SddcManager, VcfApiError

logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")
log = logging.getLogger("plan")


def main():
    p = argparse.ArgumentParser()
    p.add_argument("--sddc", required=True)
    p.add_argument("--user", required=True)
    p.add_argument("--password", required=True)
    p.add_argument("--hosts-needed", type=int, required=True)
    p.add_argument("--storage-type", required=True,
                   help="VSAN, VSAN_ESA, NFS, VMFS_FC or VVOL")
    p.add_argument("--reuse-nsx", help="NSX cluster ID you intend to share")
    p.add_argument("--join-sso", action="store_true",
                   help="plan to join the management SSO domain")
    a = p.parse_args()

    vcf = SddcManager(a.sddc, a.user, a.password)
    findings = []

    # ---------- capacity, by storage type
    hosts = list(vcf.paginate("v1", "/hosts"))
    free = [h for h in hosts
            if str(h.get("status", "")).startswith("UNASSIGNED")]
    by_type = Counter(str(h.get("storageType") or "UNKNOWN").upper()
                      for h in free)

    print("UNASSIGNED HOSTS BY STORAGE TYPE")
    for t, n in sorted(by_type.items()):
        marker = "  <-- your type" if t == a.storage_type.upper() else ""
        print(f"  {t:<12} {n:>4}{marker}")

    available = by_type.get(a.storage_type.upper(), 0)
    if available < a.hosts_needed:
        findings.append(
            f"need {a.hosts_needed} {a.storage_type} hosts, "
            f"{available} unassigned. Commission {a.hosts_needed - available} "
            "more, and remember storage type is fixed at commission time")

    # ---------- NSX sharing
    domains = list(vcf.paginate("v1", "/domains"))
    nsx_use = defaultdict(list)
    for d in domains:
        nsx = d.get("nsxtCluster") or d.get("nsxCluster") or {}
        nid = nsx.get("id") if isinstance(nsx, dict) else None
        if nid:
            nsx_use[nid].append(d.get("name"))

    print(f"\nDOMAINS: {len(domains)}")
    for d in domains:
        print(f"  {d.get('name'):<28} type={d.get('type'):<12} "
              f"status={d.get('status')}")

    if nsx_use:
        print("\nNSX CLUSTER SHARING")
        for nid, names in nsx_use.items():
            print(f"  {nid[:20]:<22} used by {len(names)}: {', '.join(names)}")

    if a.reuse_nsx:
        current = len(nsx_use.get(a.reuse_nsx, []))
        print(f"\nplanned NSX cluster already serves {current} domain(s); "
              f"yours would be number {current + 1}")
        findings.append(
            f"confirm the maximum domains per NSX cluster for your release. "
            f"Exceeding it fails at creation with an allocation error, "
            f"not a warning")

    # ---------- SSO joins
    if a.join_sso:
        joined = [d for d in domains if d.get("ssoName") or d.get("ssoId")]
        print(f"\ndomains carrying an SSO reference: {len(joined)}")
        findings.append(
            "confirm the maximum joined domains for the management SSO "
            "domain. The ceiling is reported only when you hit it")

    # ---------- network pool headroom
    print("\nNETWORK POOLS")
    for pool in vcf.paginate("v1", "/network-pools"):
        nets = pool.get("networks") or []
        desc = ", ".join(f"{n.get('type')}/{n.get('vlanId')}" for n in nets)
        print(f"  {pool.get('name'):<28} {desc}")
    findings.append(
        "verify each pool has at least one free address per new host per "
        "network before you build. Exhaustion fails mid-operation")

    # ---------- locks and open work
    open_tasks = [t for t in vcf.paginate("v1", "/tasks")
                  if t.get("status") in ("IN_PROGRESS", "Pending", "Failed")]
    if open_tasks:
        print(f"\nOPEN OR FAILED TASKS: {len(open_tasks)}")
        for t in open_tasks[:10]:
            print(f"  {t.get('status'):<12} {t.get('name')} [{t.get('id')}]")
        findings.append(
            f"{len(open_tasks)} task(s) open or failed. These may hold "
            "resource locks that will refuse your domain creation")

    # ---------- verdict
    print("\n" + "=" * 66)
    if findings:
        print(f"{len(findings)} thing(s) to resolve or confirm:\n")
        for i, f in enumerate(findings, 1):
            print(f"  {i}. {f}\n")
    else:
        print("no blockers found from the API's point of view")

    print("Decisions this tool cannot make for you:")
    print("  * dedicated or shared NSX      blast radius versus cost")
    print("  * join or stand alone on SSO   xVC-vMotion versus isolation")
    print("  * storage type per cluster     cannot be changed in place")
    print("  * stretch or not               may not be reversible")

    return 2 if findings else 0


if __name__ == "__main__":
    sys.exit(main())

07. The decision register for a domain

DecisionReversible?Cost of getting it wrong
Domain name and vCenter FQDNNoRebuild
SSO domain joinEffectively noRebuild the domain
NSX dedicated or sharedNoRebuild, or live with the coupling
Storage type per clusterNo, not in placeNew cluster plus full workload migration
Stretched or standardSometimes notConfirm the unstretch path for your release before relying on it
Host count per clusterYesExpansion, within pool and licence limits
Cluster count per domainYesAdd clusters as needed

Five of the seven rows are one-way. That ratio is the argument for spending real time on a domain design before anyone opens the API, and for recording the rejected alternatives alongside the choices.

Next, and last in this series: prechecks and upgrades, and how to read a failed precheck properly.

Allocation ceilings, removal preconditions and error codes are taken from the published VMware Cloud Foundation API reference. The specific maximum values are release dependent and are reported in the error message rather than published inline, so check the configuration maximums for your version at design time. Nothing here is official guidance from VMware or Broadcom.

Leave a Reply

Discover more from VMwareBlogs

Subscribe now to keep reading and get access to the full archive.

Continue reading