Category: Tech Notes
-
Vertex AI Observability and Tracing, from Dashboard to Span (Google Cloud Gen AI Series, Part 25)
A green status code on an eight second request tells you nothing. Here is how Cloud Monitoring, Cloud Trace, and Cloud Logging on Vertex AI tell you which call was slow, what it cost, and what to never log.
-
Multimodal on Vertex AI, from Nano Banana to Veo and Lyria (Google Cloud Gen AI Series, Part 24)
A field guide to generative media on Vertex AI: when to use Gemini native image generation, Imagen, Veo, Lyria, and Chirp, with real model IDs and a per-image cost model.
-
Vertex AI Gen AI Evaluation Service, Pointwise to Trajectory (Google Cloud Gen AI Series, Part 23)
The Vertex AI Gen AI evaluation service turns a two-example spot check into a number you can gate a release on. How pointwise, pairwise, rubric, and trajectory metrics work, and how to validate the autorater judge before you trust it.
-
Gemini Enterprise and Agentspace, Enterprise Agents for Every Employee (Google Cloud Gen AI Series, Part 22)
Google Agentspace is now Gemini Enterprise, the per seat front door that puts search and a gallery of agents in front of every employee. What it includes, how a request flows, what a seat costs, and when to buy it instead of building your own.
-
Multi-Agent Systems on Vertex AI with ADK and Agent Engine (Google Cloud Gen AI Series, Part 21)
One big agent with twenty tools rots fast. Here is how to split it into a coordinator and typed sub-agents with the ADK, choose deterministic versus LLM-driven flows, connect across boundaries with A2A, and deploy to Vertex AI Agent Engine.
-
TPU Pods and Multislice Distributed Training on GKE (Google Cloud Gen AI Series, Part 19)
Where a single TPU slice stops fitting your model, Multislice takes over. How v6e Pods, ICI and the DCN, and GKE JobSets scale training from 16 chips to thousands.
-
Gemma Open Models on Vertex AI, from Model Garden to Endpoint (Google Cloud Gen AI Series, Part 18)
When a managed Gemini bill stops making sense, Gemma is the open model you host yourself on Vertex AI. Here is the 2026 lineup, the deploy path through Model Garden, and the volume where self-hosting actually pays.
-
Distilling Gemini on Vertex AI, from Teacher to Student Model (Google Cloud Gen AI Series, Part 17)
Distillation trains a small Gemini student to copy a large teacher, so the teacher writes the labels and you serve the result at Flash prices. When it pays, what it costs, and where it fails.
-
Fine-Tuning Gemini with Supervised Tuning on Vertex AI (Google Cloud Gen AI Series, Part 16)
Supervised fine-tuning on Vertex AI adjusts Gemini to your task with a few hundred labelled examples, and because it uses LoRA the tuned model costs the same to run as the base. Here is when to tune, how to build the dataset, the knobs that matter, and what it costs.
-
Grounding Gemini with Google Search and Your Own Data (Google Cloud Gen AI Series, Part 15)
How to ground Gemini on Vertex AI against the live web and your own Vertex AI Search data store, read the grounding metadata, render citations correctly, and size what it costs.
-
Vertex AI Safety Filters and Model Armor (Google Cloud Gen AI Series, Part 14)
Gemini safety filters and Model Armor are two separate layers on Vertex AI. Here is what each one catches, why the built-in filters default to off, and the configuration I would actually run in production.
Architect’s Toolkit
PJ’s Tools
- Infra 360 Hub – All Series
- VCF 9 Interactive Walkthroughs
- VCF Design Cheatsheet
- VCF Upgrade Planner
- VCF 9 Series Hub
- VCF Deployment Hub
- AI Stack Hub
- AI Infra Sizing & Cost Calculator
- LLM & RAG Cost Calculator
- DrJhaGPT – Ask Pranay
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed


DrJha