Tag: Observability
-
Vertex AI Observability and Tracing, from Dashboard to Span (Google Cloud Gen AI Series, Part 25)
A green status code on an eight second request tells you nothing. Here is how Cloud Monitoring, Cloud Trace, and Cloud Logging on Vertex AI tell you which call was slow, what it cost, and what to never log.
-
Azure Monitor Observability for GenAI, from Metrics to Traces (Azure Gen AI Series, Part 25)
Azure GenAI observability comes in three layers: free platform metrics, per request diagnostic logs in Log Analytics, and OpenTelemetry traces in Application Insights. Here is what each one sees, what it costs, and the order I turn them on.
-
Azure AI Foundry Evaluation and Observability, from CI Gate to Live Traffic (Azure Gen AI Series, Part 23)
Evaluation scores catch a bad agent; tracing tells you why it went bad. Here is how Microsoft Foundry runs the same evaluators at dev time, in your CI gate, and against live traffic, and what continuous evaluation actually costs.
-
Amazon Bedrock Observability with CloudWatch and Invocation Logging (AWS Gen AI Series, Part 25)
Bedrock ships almost no history by default. Here is how I turn on model invocation logging, pick the CloudWatch metrics worth an alarm, and pull token cost per model straight from the logs.
-
Monitoring GPU Resources in a Private AI Platform: Metrics, Dashboards, and Tools
Which metrics tell you the truth about GPU health, and which tools to use to see them, with real dashboard patterns, alert thresholds, and practical habits for a private AI estate.
-

Observability for VKS: Metrics, Logs and VCF Operations (VKS Series, Part 11)
Kubernetes-only tooling is blind below the node. Here is how metrics, logs and VCF Operations fit together so you can tell the app from the cluster from the infrastructure.
-

VCF Operations in VCF 9: Monitoring and Observability Explained (VCF 9 Series, Part 19)
A field guide to VCF Operations in VCF 9: what replaced the Aria suite, the architecture you actually deploy, and the observability gotchas that bite during the move.
Architect’s Toolkit
PJ’s Tools
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed

DrJha