Category: AI/ML
-
Vertex AI Pipelines for LLMOps, from Notebook to Nightly Retrain (Google Cloud Gen AI Series, Part 28)
How I turn a Gemini tuning notebook into a Vertex AI pipeline that reruns nightly, gates on evals, registers versions, and does not quietly ship a worse model.
-
Vertex AI Responsible AI, Governance, and Audit Logging (Google Cloud Gen AI Series, Part 27)
Data Access audit logs, SynthID provenance, Model Registry, and Access Transparency, wired into a governance baseline for a Gemini workload on Vertex AI. What is on by default, what you have to switch on, and what it costs.
-
Vertex AI Cost Governance and FinOps (Google Cloud Gen AI Series, Part 26)
Read a Vertex AI bill line by line, then cut it with model tiering, context caching, Batch mode, and provisioned throughput judged against a real break-even. With budgets, billing export, and label-based attribution.
-
Vertex AI Observability and Tracing, from Dashboard to Span (Google Cloud Gen AI Series, Part 25)
A green status code on an eight second request tells you nothing. Here is how Cloud Monitoring, Cloud Trace, and Cloud Logging on Vertex AI tell you which call was slow, what it cost, and what to never log.
-
Multimodal on Vertex AI, from Nano Banana to Veo and Lyria (Google Cloud Gen AI Series, Part 24)
A field guide to generative media on Vertex AI: when to use Gemini native image generation, Imagen, Veo, Lyria, and Chirp, with real model IDs and a per-image cost model.
-
Vertex AI Gen AI Evaluation Service, Pointwise to Trajectory (Google Cloud Gen AI Series, Part 23)
The Vertex AI Gen AI evaluation service turns a two-example spot check into a number you can gate a release on. How pointwise, pairwise, rubric, and trajectory metrics work, and how to validate the autorater judge before you trust it.
-
Gemini Enterprise and Agentspace, Enterprise Agents for Every Employee (Google Cloud Gen AI Series, Part 22)
Google Agentspace is now Gemini Enterprise, the per seat front door that puts search and a gallery of agents in front of every employee. What it includes, how a request flows, what a seat costs, and when to buy it instead of building your own.
-
Multi-Agent Systems on Vertex AI with ADK and Agent Engine (Google Cloud Gen AI Series, Part 21)
One big agent with twenty tools rots fast. Here is how to split it into a coordinator and typed sub-agents with the ADK, choose deterministic versus LLM-driven flows, connect across boundaries with A2A, and deploy to Vertex AI Agent Engine.
-
TPU Pods and Multislice Distributed Training on GKE (Google Cloud Gen AI Series, Part 19)
Where a single TPU slice stops fitting your model, Multislice takes over. How v6e Pods, ICI and the DCN, and GKE JobSets scale training from 16 chips to thousands.
-
Gemma Open Models on Vertex AI, from Model Garden to Endpoint (Google Cloud Gen AI Series, Part 18)
When a managed Gemini bill stops making sense, Gemma is the open model you host yourself on Vertex AI. Here is the 2026 lineup, the deploy path through Model Garden, and the volume where self-hosting actually pays.
-
Distilling Gemini on Vertex AI, from Teacher to Student Model (Google Cloud Gen AI Series, Part 17)
Distillation trains a small Gemini student to copy a large teacher, so the teacher writes the labels and you serve the result at Flash prices. When it pays, what it costs, and where it fails.
Architect’s Toolkit
PJ’s Tools
- Infra 360 Hub – All Series
- VCF 9 Interactive Walkthroughs
- VCF Design Cheatsheet
- VCF Upgrade Planner
- VCF 9 Series Hub
- VCF Deployment Hub
- AI Stack Hub
- AI Infra Sizing & Cost Calculator
- LLM & RAG Cost Calculator
- DrJhaGPT – Ask Pranay
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed




DrJha