Category: AI/ML
-
Third-Party Models in Vertex AI Model Garden, from Claude to Self-Deploy (Google Cloud Gen AI Series, Part 4)
Vertex AI Model Garden runs Claude, Mistral, Grok and more as managed APIs, or as proprietary models you license into your own VPC. Here is how each mode works and when to use it.
-
Gemini Flash vs Pro, and When to Pay for Reasoning (Google Cloud Gen AI Series, Part 3)
Flash or Pro? On Vertex AI the two Gemini tiers differ by about 4x on output tokens. Here is how to pick per request and route only the hard ones to Pro.
-
Vertex AI and Model Garden, from Catalog to Endpoint (Google Cloud Gen AI Series, Part 2)
Vertex AI is Google Cloud’s managed platform and Model Garden is its 200-plus model catalog. Here is how the three access paths, managed API, MaaS, and self-deploy, decide your cost, latency, and data isolation, and which one to start on.
-
Google Cloud Generative AI Stack, End to End (Google Cloud Gen AI Series, Part 1)
A map of the Google Cloud generative AI stack in 2026, from Gemini Enterprise Agent Platform (the service that used to be Vertex AI) down to Ironwood TPUs, and where a real project plugs in.
-
Azure Generative AI vs the Field, the Verdict (Azure Gen AI Series, Part 30)
After twenty-nine parts on the Azure GenAI stack, here is the plain verdict: where Azure wins, where it loses to AWS and Google Cloud, what the bill really looks like, and who should build on it.
-
Reference Architectures for Azure GenAI, from Chatbot to Batch (Azure Gen AI Series, Part 29)
The four architectures almost every Azure GenAI project actually needs, chatbot, retrieval, agentic, and batch, with the baseline Foundry components, real sizing numbers, and when to add an API Management gateway.
-
LLMOps and CI/CD on Azure for GenAI Apps (Azure Gen AI Series, Part 28)
An Azure LLMOps pipeline looks like MLOps until the gate. Here is how to block a merge on an evaluation score, register the flow, and roll it out blue to green on a managed online endpoint, plus why new pipelines should skip Prompt Flow.
-
Azure Responsible AI, the RAI Dashboard and Scorecard (Azure Gen AI Series, Part 27)
Responsible AI on Azure is two toolchains: the Responsible AI dashboard for tabular models and Foundry safety evaluators for generative apps. Here is which one your workload needs, how to run each, and how to keep a scorecard for the EU AI Act.
-
Azure OpenAI Cost Governance and FinOps (Azure Gen AI Series, Part 26)
When a Provisioned Throughput reservation actually beats pay as you go on Azure OpenAI, how to read a bill hidden under Cognitive Services, and the Batch and caching discounts you get for free.
-
Azure Monitor Observability for GenAI, from Metrics to Traces (Azure Gen AI Series, Part 25)
Azure GenAI observability comes in three layers: free platform metrics, per request diagnostic logs in Log Analytics, and OpenTelemetry traces in Application Insights. Here is what each one sees, what it costs, and the order I turn them on.
-
Vision, Audio, Image, and Document Models on Azure OpenAI (Azure Gen AI Series, Part 24)
Multimodal on Azure is not one switch, it is four services. Where vision, voice, image generation, and Content Understanding each live, what they cost, and which one owns each input.
-
Azure AI Foundry Evaluation and Observability, from CI Gate to Live Traffic (Azure Gen AI Series, Part 23)
Evaluation scores catch a bad agent; tracing tells you why it went bad. Here is how Microsoft Foundry runs the same evaluators at dev time, in your CI gate, and against live traffic, and what continuous evaluation actually costs.
Architect’s Toolkit
PJ’s Tools
- Infra 360 Hub – All Series
- VCF 9 Interactive Walkthroughs
- VCF Design Cheatsheet
- VCF Upgrade Planner
- VCF 9 Series Hub
- VCF Deployment Hub
- AI Stack Hub
- AI Infra Sizing & Cost Calculator
- LLM & RAG Cost Calculator
- DrJhaGPT – Ask Pranay
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed




DrJha