, ,

How to Monitor Fleet Health with VCF Operations in VCF 9.1 (VCF 9.1 Day-2 Operations Step by Step Guide, Part 18)

A routine, step-by-step guide to monitoring VCF 9.1 fleet health in VCF Operations: read VCF Health, review findings, tune alerts, and route notifications.

VCF 9.1 Day-2 Operations · Part 18 of 20
Task type: ROUTINE. Every fleet runs this, and it never stops. Monitoring is the background task that tells you when any other Day-2 task is due.
What it does: Uses VCF Operations as one console to watch the health of every workload domain, vCenter, NSX instance, vSAN cluster, ESX host and VCF management service, then routes alerts to the people who act on them.
Depends on: a deployed VCF 9.1 fleet where VCF Operations already collects from vCenter, NSX and vSAN, which the deployment work sets up.
Before you begin: Confirm a healthy, fully backed up fleet before you rely on its health data. Data sources should be connected and collecting, and recent backups of SDDC Manager, vCenter and NSX Manager should already exist. Work through the VCF 9.x pre-installation checklist so you start from a known good state.
TL;DR
  • Fleet health monitoring is routine and continuous. It runs in VCF Operations, the console that owns fleet lifecycle and diagnostics in VCF 9.1.
  • VCF Health gives two entry points: Component View, a single summary of every object, and VCF View, a drill down from a VCF instance to its domains, vCenter, NSX, vSAN and ESX hosts.
  • VCF Operations 9.1 tracks health for certificates, NTP and DNS, vCenter, ESX, vSAN, NSX, vSphere High Availability, VM operations, vMotion, snapshots and VCF management services.
  • Findings collect data from every component to surface existing or potential issues, each with a recommended fix.
  • Alerts combine symptoms and recommendations. Route them outside the console with notification rules and an outbound plug-in such as Standard Email, Slack or Webhook.
  • This part watches health and raises alerts. Fixing what an alert points to, such as an expiring certificate, belongs to the task that owns it, for example Part 5 for certificate rotation.

This part shows you how to monitor VCF 9.1 fleet health from VCF Operations, and how to make sure the right alerts reach the right people. You will open VCF Health, read the fleet summary, drill into component tiles, review Findings, check the health of VCF Operations itself, tune a few alert definitions, and configure outbound notifications. Each step names the exact console so you always know where you are working.

Treat monitoring as a routine that runs every day rather than a one time setup. VCF Operations owns fleet health and diagnostics in VCF 9.1, so this is where you watch the whole stack in one place instead of logging into each vCenter and NSX Manager. For the wider operating model behind these tasks, see the Day-2 operations overview, and for a primer on the console itself, see what VCF Operations is in VCF 9.

1. Detect VCF Health 2. Investigate Findings 3. Notify Outbound 4. Remediate Owning task 5. Verify VCF Health
Figure 1. The monitoring loop you run continuously, and the surface that owns each stage.

Prerequisites

Requirement Detail
VCF Operations accessAn administrator account, or a role with view and notification rights, reached from the fleet single sign on.
Data collectionvCenter, NSX and vSAN adapters already collecting, so health and metrics are populated. Set up during deployment.
Delivery endpointA reachable SMTP server, Slack webhook, or other endpoint if you plan to send notifications outside the console.
Baseline healthA fleet that currently reports green, so new alerts stand out against a known good state.
Named recipientsThe teams or distribution lists that should receive alerts, agreed before you build notification rules.

Step 1, Open VCF Health and read the fleet summary

Start where the whole stack rolls up. VCF Health gives two entry points, and both live in VCF Operations.

  1. Sign in to VCF Operations at https://vcf-operations-fqdn as an administrator.
  2. In the left navigation pane, open VCF Health.
  3. Click Component View to see a summary of every object of each type across the fleet.
  4. Scan the tiles for any red or yellow status that marks a component needing attention.
  5. Click VCF View to switch to the inventory tree, then expand a VCF instance to its management and workload domains.
  6. Select a domain to see its vCenter, NSX, vSAN and ESX hosts with their reported health.

VCF Health reports on active inventory only. If you delete a resource, its events are removed at once, so a tile that reads green means the objects it still tracks are healthy, not that nothing was ever removed.

Step 2, Drill into component health tiles

Each tile summarizes one operational area and links to the detail behind it. Use the table to see what each area tracks and which console you move to when you need to act.

Health area What it tracks Where you act
CertificatesValidity and days to expiry across componentsRotate in Part 5
NTP and DNSTime sync drift and forward and reverse lookupsDNS and NTP servers
vCenterConnectivity, utilization and services per vCentervSphere Client
ESXHost connection, hardware and performance across all hostsvSphere Client
vSANDatastore capacity, resync and object healthvSphere Client, Skyline Health
NSXEdge and host networking capacity and manager healthNSX Manager
VCF management servicesHealth of services on the VCF services runtimeVCF Operations and SDDC Manager
  1. In VCF Health, click a component tile such as vCenter to open its detail view.
  2. Click the View Details link on a tile to open the panel with per object metrics.
  3. Click the View Dashboard link to open the full dashboard for that component in VCF Operations.
  4. For an ESX or vSAN concern, note the affected host or cluster, then confirm it in the vSphere Client.
  5. For a certificate nearing expiry, record the component and expiry date so you can hand rotation to the task that owns it.

Reading a health tile and fixing what it reports are two different jobs. This part surfaces the condition. Certificate rotation is a separate task, covered in Part 5, so link out to it rather than rotating from here.

VCF Health Component View summary VCF View inventory tree Per component tiles View Details and Dashboard Findings and Diagnostics Known and potential issues Recommended fix per finding Log Assist for support Public REST API Alerts and Notifications Symptom definitions Alert definitions Outbound plug ins Notification rules
Figure 2. Three monitoring surfaces in VCF Operations and what each one gives you.

Step 3, Review Findings and act on them

Findings collect data from every VCF component to flag existing or potential issues, and each one carries a recommended fix. Work them in severity order.

  1. In VCF Operations, open the Findings view to list issues across the fleet.
  2. Sort by severity so critical findings sit at the top of the list.
  3. Click a finding to read its description, affected objects and recommended fix.
  4. Apply the recommended fix, or route it to the Day-2 task that owns it, such as a backup or an upgrade.
  5. Click Log Assist when a finding needs a support case, to generate a log bundle with the finding data attached.

For deeper log analysis behind a finding, VCF Operations for Logs sits alongside diagnostics. See the logs troubleshooting workflow for how to pivot from a finding to the underlying events.

Step 4, Check the health of VCF Operations itself

If the monitoring platform is unhealthy, every other tile lies by omission. Confirm the analytics cluster is sound before you trust its data.

  1. In VCF Operations, open the VCF Operations Health dashboard to see the state of the analytics cluster.
  2. Confirm every node reports online and that no node sits at a capacity limit.
  3. Confirm collection is current, so no adapter has stopped gathering data.
  4. Review the self monitoring alerts for the cluster and clear any that point to a node, disk or collection problem.

A stalled adapter is a common blind spot. A tile can read green simply because no fresh data arrived to change it, so a current collection time matters as much as the color.

Step 5, Tune alert definitions and symptoms

An alert definition combines symptoms with recommendations. Tune the noisy ones so responders trust the alerts they receive.

  1. In VCF Operations, open Alert Definitions to see the alert library.
  2. Filter the list to the definition you want to change, then open it.
  3. Review the Symptom Definitions that trigger the alert, and adjust a threshold that fires too often or too late.
  4. Add a Recommendation so responders see the fix at the moment the alert fires.
  5. Click Save to keep the change.
  6. To build a new alert, click Add and combine a symptom, an object type and a recommendation.

Apply threshold changes through a policy where you can, so a single edit covers a group of objects rather than one at a time. That keeps the alert library consistent as the fleet grows.

Step 6, Configure outbound notifications

An outbound plug-in is how an alert leaves VCF Operations and reaches a person or another system. Add and activate one before you build rules.

  1. In the VCF Operations navigation bar, click Operate.
  2. In the left navigation pane, click Administration, then Configurations.
  3. Click the Outbound Settings tile.
  4. Click Add to open the Outbound Plug-In dialog box.
  5. From the Plug-In Type drop-down menu, select a delivery method such as Standard Email, Slack or Webhook Notification.
  6. Enter the destination details, for email the SMTP host, port, sender address and recipients.
  7. Click Save, select the new instance, then click Activate so it starts sending.

Available plug-in types include Standard Email, SNMP Trap, Webhook Notification, Slack, Service Now and Log File. Deactivating an instance stops delivery without deleting the configuration, which is useful during a maintenance window.

Step 7, Route alerts with notification rules

A notification rule decides which alerts flow to which plug-in, so a team receives only what it owns. Build one rule per audience.

  1. In VCF Operations, open Notifications under the alerts configuration.
  2. Click Add to create a rule.
  3. Select the outbound plug-in instance you activated in Step 6 as the target.
  4. Set the filter by object type and alert criticality, so only relevant alerts reach that team.
  5. Click Save to activate the rule.

A workable split is to route critical vSAN and ESX alerts to the storage and platform team, and certificate or password warnings to whoever owns rotation. Fewer, sharper rules beat one rule that copies every alert to everyone.

How to verify fleet health monitoring

Confirm the loop works from detection through to delivery before you call it done.

  1. In VCF Health, confirm the fleet summary reads green once you have cleared open findings.
  2. Send a test message from the outbound plug-in and confirm the target receives it.
  3. Confirm a known alert appears in the Alerts view and in the routed channel.
  4. Confirm the VCF Operations Health dashboard shows all nodes online and collection current.

You can also pull open alerts from the VCF Operations API to feed a dashboard or a ticketing system.

# acquire a token first, then reuse it in the header below
curl -k 'https://vcf-operations-fqdn/suite-api/api/alerts?activeOnly=true' \
  -H 'Authorization: vRealizeOpsToken YOUR_TOKEN_VALUE' \
  -H 'Accept: application/json'
# an empty alert list means no active alerts on the queried objects

Common errors and fixes

Symptom Likely cause Fix
No alerts arrive by emailThe outbound plug-in is added but not activated, or SMTP is blocked.Activate the Standard Email instance and confirm the SMTP host and port are reachable from VCF Operations.
A component tile shows no dataThe adapter for that component stopped collecting.Reconnect the data source, confirm its credentials, then wait for the next collection cycle.
Too many alerts, real ones get lostSymptom thresholds are too sensitive, or every alert routes to one channel.Raise the noisy thresholds through a policy and split notification rules by object type and criticality.
Health is green but a host is downHealth covers active inventory only, so a removed object drops its events.Confirm the host is still in inventory, recommission it if it was removed, then recheck the tile.

Common questions

Where do I watch fleet health in VCF 9.1? In VCF Operations, under VCF Health. It brings the whole stack into one console instead of a separate tool per component.

Do I fix a failing certificate from here? No. VCF Health shows you a certificate nearing expiry, but rotation is a separate task, covered in Part 5.

What is the difference between Component View and VCF View? Component View summarizes every object of a type across the fleet. VCF View lets you drill down one VCF instance at a time through its domains and components.

Can alerts reach Slack or a ticketing system? Yes. Add a Slack, Webhook or Service Now outbound plug-in, activate it, then point a notification rule at it.

References

VCF 9.1 Day-2 Operations · Part 18 of 20
« Previous: Part 17  |  Complete Guide  |  Next: Part 19 »

About The Author


Discover more from Journal of Intelligent Infrastructure

Subscribe to get the latest posts sent to your email.

Leave a Reply

Your email address will not be published. Required fields are marked *

Architect’s Toolkit

About the Author

Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.

Discover more from Journal of Intelligent Infrastructure

Subscribe now to keep reading and get access to the full archive.

Continue reading