Tag: RAG
-
Reference Architectures on Google Cloud, from Chatbot to Batch (Google Cloud Gen AI Series, Part 29)
Most generative AI on Google Cloud reduces to four shapes: a direct call, RAG, an agent, and batch. Here is how each maps to managed services, what it costs, and which one to reach for first.
-
Grounding Gemini with Google Search and Your Own Data (Google Cloud Gen AI Series, Part 15)
How to ground Gemini on Vertex AI against the live web and your own Vertex AI Search data store, read the grounding metadata, render citations correctly, and size what it costs.
-
Vertex AI Search and RAG Engine, from Data Store to Grounded Answer (Google Cloud Gen AI Series, Part 12)
Vertex AI Search gives you managed retrieval with almost no plumbing, while RAG Engine hands you the chunking, embeddings, and vector backend. Here is when each one wins.
-
Azure AI Search and RAG, from Index to Answer (Azure Gen AI Series, Part 12)
How Azure AI Search grounds a model in your own documents: the ingestion pipeline, keyword and vector retrieval, RRF hybrid merging, and the semantic ranker that decides what the model actually reads.
-
Bedrock Reference Architectures for Chatbot, RAG, Agentic, and Batch (AWS Gen AI Series, Part 29)
Most AWS generative AI features are one of four shapes: chatbot, RAG, agentic, or batch. Here is how each maps to Amazon Bedrock services, what it costs, and which one to reach for first.
-
Amazon Bedrock Knowledge Bases and Managed RAG, End to End (AWS Gen AI Series, Part 12)
How Amazon Bedrock Knowledge Bases turn your documents into retrieval augmented generation, which vector store and chunking to pick, and why the bill is dominated by a vector-store floor, not tokens.
-
Data Sources in Private AI: Connectors and Supported File Formats
The four data source connectors in Private AI (Google Drive, Confluence, Amazon S3, SharePoint) and the file formats the platform can index for retrieval.
-
What Is in the Private AI Catalog: A Guide to the Blueprints
A plain-language field guide to every blueprint in the Private AI catalog, grouped by job: compute, model serving, RAG retrieval, OCR and speech, and the access layer.
Architect’s Toolkit
PJ’s Tools
- Infra 360 Hub – All Series
- VCF 9 Interactive Walkthroughs
- VCF Design Cheatsheet
- VCF Upgrade Planner
- VCF 9 Series Hub
- VCF Deployment Hub
- AI Stack Hub
- AI Infra Sizing & Cost Calculator
- LLM & RAG Cost Calculator
- DrJhaGPT – Ask Pranay
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed

DrJha