Tag: RAG
-
Path to AI Engineer, and What to Learn Next (AI Engineering Series, Part 30)
Thirty parts on, here is the honest version of the AI engineering career path: what the market pays, which routes into the role actually work, what a hiring loop tests, and a twelve month plan to close the gaps this series left open.
-
Cost Control and Model Routing for LLM Applications (AI Engineering Series, Part 27)
Model routing saved our documentation assistant 22 percent. Trimming retrieval saved more, in one line of config. Here is the cost model, the routing decision table, and the measured numbers behind both.
-
Caching, Batching and Latency Engineering for LLM Applications (AI Engineering Series, Part 26)
Prompt caching, batch processing and streaming, measured on a production docs assistant. Where the breakpoint really goes, why the obvious block to cache is the wrong one, and the cost arithmetic that follows.
-
Retrieval Evaluation for RAG, Separated From Answer Evaluation (AI Engineering Series, Part 13)
Most RAG teams report one number and tune the wrong component. Here is how I split retrieval evaluation from answer evaluation on a 120 query set, with ranx, real metric output and the failure attribution table that came out of it.
-
Hybrid Search and Reranking for RAG Retrieval (AI Engineering Series, Part 12)
Dense retrieval cannot find an error code. This part adds BM25, fuses the two ranked lists with reciprocal rank fusion, and reranks the survivors with a cross encoder, with measured nDCG and latency for each step.
-
Vector Stores in Practice: pgvector, Chroma and Qdrant Compared (AI Engineering Series, Part 11)
pgvector, Chroma and Qdrant measured on the same 50,000 chunk corpus: build time, p95 latency, resident memory and the filtered recall collapse that default settings will hand you. With the configuration fixes that actually work.
-
Choosing an Embedding Model: Benchmarking, Dimensions and Cost (AI Engineering Series, Part 10)
Embedding model choice sets your recall ceiling and, far more expensively, the memory your vector index needs forever. Here is the sizing arithmetic, a recall benchmark run on a real corpus, and why 3072 dimensions is rarely the right default.
-
Chunking Strategies for RAG and How to Choose One (AI Engineering Series, Part 9)
Chunk size, overlap, header aware splitting and atomic tables, measured on a real documentation page with langchain-text-splitters 1.1.2. Includes the orphaned table row that cost me three weeks.
-
Document Ingestion for RAG: Parsing, Cleaning and Formats That Break (AI Engineering Series, Part 8)
A rate limit table came out of our handbook PDF as a vertical column of orphaned numbers, and retrieval answered 18 of 40 questions wrong for three weeks before anyone checked. Ingestion decides what your assistant can ever know.
-
Why Retrieval Beats Fine Tuning for Most Business Problems (AI Engineering Series, Part 7)
Retrieval is not cheaper per request; it made our documentation assistant 2.1x more expensive per call. What it buys is a cheap unit of change, and that is the trade that decides most business projects.
-
watsonx Reference Architectures for RAG, Agentic, and Regulated Workloads (IBM Gen AI Series, Part 23)
The three watsonx reference architectures that recur in real builds: enterprise RAG on watsonx.data, agentic systems on watsonx Orchestrate, and a watsonx.governance overlay for regulated work. Which to build first, what each costs, and where they break.
-
RAG on watsonx.data with Milvus Vector Search (IBM Gen AI Series, Part 9)
How to build governed RAG on watsonx.data with the embedded Milvus vector database, from choosing a slate or Granite embedding model to sizing the service and picking an index.
Architect’s Toolkit
PJ’s Tools
- Infra 360 Hub – All Series
- VCF 9 Interactive Walkthroughs
- VCF Design Cheatsheet
- VCF Upgrade Planner
- VCF 9 Series Hub
- VCF Deployment Hub
- AI Stack Hub
- AI Infra Sizing & Cost Calculator
- LLM & RAG Cost Calculator
- DrJhaGPT – Ask Pranay
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed

DrJha