, ,

How to Decommission a Cluster or Workload Domain in VCF 9.1 (VCF 9.1 Day-2 Operations Step by Step Guide, Part 14)

Retire capacity in VCF 9.1 the safe way. Delete a vSphere cluster in SDDC Manager or a whole workload domain in VCF Operations, then reclaim the hosts.

VCF 9.1 Day-2 Operations · Part 14 of 20
Task type: Scenario. Run this only when you retire capacity, either removing a single vSphere cluster from a domain or deleting a whole workload domain. It depends on a healthy fleet, workloads already migrated off the target, and a fresh backup of SDDC Manager, vCenter, and NSX Manager.
Before you begin. Confirm a healthy, fully backed up fleet before you remove any capacity, because these steps destroy vSAN datastores and cannot be reversed. Work through the VCF 9.x pre-installation checklist first, then take a fresh SDDC Manager, vCenter, and NSX Manager backup so you can recover if a workflow stalls.
Quick summary
  • Console: delete a single cluster in SDDC Manager, delete a whole domain in VCF Operations.
  • You cannot delete the last cluster in a domain, so remove the domain instead.
  • Clear NSX Edge clusters and unmount remote vSAN datastores before either delete.
  • vSAN datastores are destroyed on delete, and other storage types are unmounted.
  • Released hosts return to the free pool in a need cleanup state and must be reimaged before reuse.
  • Network pools are not removed automatically and must be deleted separately.

Decommissioning removes capacity you no longer need, either a single vSphere cluster inside a domain or an entire workload domain with its vCenter and NSX Manager. In VCF 9.1 the console you use depends on the scope. A single cluster is removed from SDDC Manager, on the Clusters tab of the domain. A full domain is removed from VCF Operations, which now owns fleet and SDDC lifecycle. Both actions are destructive: vSAN datastores on the affected hosts are destroyed, other datastores are unmounted, and released hosts return to the free pool for cleanup. This part walks both paths end to end and shows how to reclaim the hardware afterward.

Run this scenario only when a cluster or domain is genuinely retiring, for example when you consolidate onto newer hardware, collapse a test domain, or move a workload tier elsewhere. To grow capacity instead, see the reverse task, adding a cluster to a workload domain. For host level context and how commissioning works, see adding ESX hosts to a cluster. For where each console fits in day to day operations, see the Day-2 operations overview, and for how these domains were first built, see the VCF 9.1 Deployment guide.

Prerequisites

Work through each row before you start. Most failed removals trace back to a skipped prerequisite, usually an Edge node still on the cluster, a mounted remote datastore, or a workload left running on the target.

RequirementDetailConsole
Fleet health and backupFleet healthy, with fresh SDDC Manager, vCenter, and NSX Manager backupsVCF Operations, SDDC Manager
Workloads migratedVirtual machines you keep moved off the target with cross vCenter vMotionvCenter
NSX Edge clearedEdge clusters on the target deleted, or shrunk to at least two nodes eachNSX Manager
Remote datastores unmountedNo remote vSAN datastore mounted on the target clustervCenter
Scope confirmedTarget is not the last cluster in a domain, or plan a domain delete insteadSDDC Manager
AccessAccount holds the Administrator role in SDDC Manager and VCF OperationsVCF Operations

Step 1, Decide the decommission scope

Scope decides the console. Choose whether you are removing one cluster or a whole domain, and check for the constraints that block each path.

  1. Open VCF Operations and confirm the fleet shows a healthy state before any change.
  2. Identify whether you are removing a single vSphere cluster or an entire workload domain.
  3. If the target is the last cluster in a domain, plan a full domain deletion, because you cannot delete the last cluster on its own.
  4. If the cluster is the domain default, assign the default role to another cluster first, so Delete Cluster becomes available.
  5. Record the domain name, cluster name, host FQDNs, and network pool so you can reclaim them later.
AspectDelete a clusterDelete a domain
ConsoleSDDC Manager, Clusters tabVCF Operations, Detailed View
Scope removedOne vSphere cluster in a domainWhole domain, all clusters, vCenter, NSX Manager
Not allowed onLast cluster in a domainManagement domain
Domain vCenterStays in placeRemoved with the domain
Hosts afterwardReturn to free pool as need cleanupReturn to free pool as need cleanup
Typical timeMinutes, varies by sizeUp to 20 minutes
What are you removingOne vSphere clusterWhole workload domainSDDC Manager, Delete ClusterVCF Operations, Delete DomainLast cluster in a domain forces the domain path
Scope decides the console, and the last cluster in a domain forces a domain deletion.

Step 2, Migrate or back up the workloads you keep

Anything left on the target is lost. Move or capture every workload and dataset you want to keep before you start either delete.

  1. Open the domain vCenter in the vSphere Client.
  2. Use Migrate to move any virtual machines you keep to another cluster or domain with cross vCenter vMotion.
  3. Back up any data on the target datastore, because vSAN datastores are destroyed and other datastores are unmounted on delete.
  4. Confirm no running workloads remain on the cluster or domain you are about to remove.

Step 3, Clear NSX Edge clusters and remote datastores

Edge nodes and mounted remote datastores both block a delete. Clear them from the target before you run either workflow.

  1. Open NSX Manager for the domain.
  2. Delete any NSX Edge cluster hosted on the target, or shrink it by removing NSX Edge nodes, keeping at least two nodes on each remaining Edge cluster.
  3. In the domain vCenter, unmount any remote vSAN datastore mounted on the target cluster, because a mounted remote datastore blocks deletion.
  4. If the domain shares an NSX Manager and VCF Operations for Networks latency collection is enabled, disable latency collection so referenced NSX objects do not block the delete.
Note. You cannot remove NSX Edge nodes if the removal would leave an Edge cluster with fewer than two nodes. Relocate the Edge nodes or shrink the hosted workloads first, then delete the Edge cluster, as described in Broadcom KB 78635.

Step 4, Delete a vSphere cluster from a domain

Use this path when the domain keeps at least one other cluster. Cluster deletion runs from SDDC Manager, on the Clusters tab of the domain. vSAN datastores on the cluster hosts are destroyed during the workflow, while NFS and Fibre Channel datastores are only unmounted. You cannot run other domain tasks until the delete completes.

  1. Log in to SDDC Manager at https://sddc-manager.vcf.example.com with an account that holds the Administrator role.
  2. In the navigation pane, click Inventory, then Workload Domains.
  3. Click the name of the workload domain that holds the cluster.
  4. Click the Clusters tab.
  5. Click the vertical ellipsis next to the cluster name, then click Delete Cluster.
  6. Click Delete Cluster again to confirm, then wait for the workflow to finish before you start another domain task.

Step 5, Delete an entire workload domain

Use this path when the whole domain retires, including its last cluster, vCenter, and any unshared NSX Manager. Domain deletion runs from VCF Operations. A domain delete also removes the domain vCenter and any NSX Manager that no other domain shares. A shared NSX Manager and the network pools stay in place, so plan to clean those up separately.

  1. In VCF Operations, select Inventory, then Detailed View.
  2. Expand VCF Instances and browse to the VCF instance that holds the domain.
  3. Click the workload domain name, open the Actions menu, then select Delete Domain.
  4. Click Delete Workload Domain to start the workflow, which can take up to 20 minutes.
  5. Wait for the task to complete, and avoid other workload domain operations while it runs.

You can capture the domain identifier from the SDDC Manager API first, which helps you confirm you are acting on the right domain before you delete in the console.

GET https://sddc-manager.vcf.example.com/v1/domains
Authorization: Bearer eyJhbGciOiJSUzI1NiIsInR5cCI6IkpXVCJ9
Accept: application/json
Migrate or back up workloadsRemove NSX Edge and remote datastoresDelete cluster or domain in the consoleHosts return to free pool as need cleanupDecommission and reimage the hostsDelete the leftover network pool
Decommission order, from workload migration through host and network pool reclamation.

Step 6, Reclaim the hosts and network pools

A delete releases the hardware, but does not clean it up for you. Reclaim the hosts and remove the leftover network pool so the capacity is ready to reuse.

  1. Confirm the released hosts appear in the free pool with a need cleanup state.
  2. Decommission those hosts, reimage them with ESX, then commission them again before reuse, as covered in the host commissioning steps.
  3. Delete the network pool that the removed domain used, because domain deletion leaves it in place.
  4. Verify the fleet inventory no longer lists the removed cluster or domain.

Step 7, Recover a stalled deletion

Occasionally a delete workflow fails partway and leaves a cluster or domain in an inconsistent state. Use the supported recovery path rather than editing the inventory database by hand.

  1. Open the Tasks view in the console you used, then read the failed subtask detail to find the blocking object.
  2. Clear the specific blocker, for example an Edge node, a mounted remote datastore, or an enabled latency collection, then use Restart Task to run the workflow again.
  3. If the workflow cannot restart, follow the Broadcom KB for manual workload domain removal, which cleans up inventory safely after a failed deletion.
  4. Reconcile the fleet inventory in VCF Operations, then confirm the removed object is gone before you reuse the hosts.

After a domain delete finishes, the fleet reflects several changes at once. VCF Operations drops the domain from inventory, the management domain loses the domain vCenter and any unshared NSX Manager, released hosts move to the free pool in a need cleanup state, and the network pools remain until you delete them. Reconcile each of these before you commission the hosts into another domain, so the fleet stays consistent.

How to confirm the removal completed

Confirm the workflow finished cleanly in the console you used, then check that inventory, storage, and the host pool all reflect the change.

  • In SDDC Manager or VCF Operations, confirm the removed cluster or domain no longer appears under Inventory.
  • Confirm the delete task shows a Successful state in the Tasks view, with no failed subtasks.
  • In the domain vCenter, confirm the deleted cluster is gone and its datastores are no longer listed.
  • In the free pool, confirm the released hosts show a need cleanup state, which means they are ready to decommission and reimage.
  • Confirm the network pool that the domain used is deleted, and that no orphaned NSX objects remain for the domain.

Common errors and fixes

Most removal failures come from a leftover dependency on the target, and each clears with a specific action. Match the symptom below, apply the fix, then restart the workflow from the Tasks view.

SymptomCauseFix
Delete Cluster is greyed outCluster is the domain default, or it is the last cluster in the domainAssign the default role to another cluster, or delete the whole domain instead
Delete blocked by NSX Edge nodesEdge cluster still hosts nodes on the target, or removal would drop below two nodesDelete the Edge cluster or relocate its nodes, keeping at least two nodes per remaining Edge cluster, then retry
Domain deletion stalls on NSX objectsShared NSX Manager with VCF Operations for Networks latency collection enabledDisable latency collection for that NSX Manager, then restart the delete
Remote datastore blocks removalA remote vSAN datastore is still mounted on the clusterMigrate VMs to local storage, unmount the remote datastore in vCenter, then retry

Common questions

Can I undo a cluster or domain deletion?

No. Deletion is irreversible. vSAN datastores are destroyed, and a domain delete also removes the domain vCenter and any unshared NSX Manager. Restore from backup only if you captured one first.

What happens to the hosts after a delete?

They return to the free pool in a need cleanup state. Decommission them, reimage with ESX, and commission them again before reuse.

Does deleting a domain remove its network pool?

No. Network pools are left in place and must be deleted separately once the domain is gone.

How long does a domain deletion take?

Up to 20 minutes. During that window you cannot run other workload domain operations.

Is a shared NSX Manager deleted with the domain?

No. If the NSX Manager cluster is shared with another domain, it stays in place. Only the domain vCenter and any unshared NSX Manager are removed.

References

VCF 9.1 Day-2 Operations · Part 14 of 20
« Previous: Part 13  |  Complete Guide  |  Next: Part 15 »

About The Author


Discover more from Journal of Intelligent Infrastructure

Subscribe to get the latest posts sent to your email.

Leave a Reply

Your email address will not be published. Required fields are marked *

Architect’s Toolkit

About the Author

Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.

Discover more from Journal of Intelligent Infrastructure

Subscribe now to keep reading and get access to the full archive.

Continue reading