Category: Generative AI
-
AI & GenAI — Complete Guides & Series
Deep-dive series across the AI and generative-AI stack — self-hosted, private cloud and public cloud, explained part by part. Pick one below to read the full series. Red Hat Gen AI SeriesSelf hosting generative AI on Red Hat: RHEL AI, InstructLab, Granite, OpenShift AI, KServe, vLLM, llm-d, guardrails and GPU cost, in 30 parts.Read →Infra…
-
How to Reduce LLM API Costs, From Prompt Hygiene to Self Hosting
A practical, beginner to expert guide to cutting LLM API costs: token math, prompt and context hygiene, caching and batching, model routing, and when self hosting actually pays off.
-
The Economics and Future of Generative AI: An Honest Take (GenAI Series, Part 30)
An honest take to close the series: why GPU utilization is the real cost lever, a blunt verdict on the hype, what is actually coming, and a recap with reading paths.
-
Mixture-of-Experts and Where AI Architecture Is Heading (GenAI Series, Part 29)
Mixture-of-experts models hold enormous capacity but activate only a few experts per token, so they run cheaply. How MoE works, its memory catch, and the trends to watch.
-
What It Takes to Train a Model Across Thousands of GPUs (GenAI Series, Part 28)
Training a frontier model coordinates thousands of GPUs for months. How data, tensor, pipeline and expert parallelism, the memory math, and checkpointing make it possible.
-
On-Prem vs Cloud vs Hybrid for GenAI: An Honest Verdict (GenAI Series, Part 27)
Where should generative AI run? An honest framework weighing data sovereignty, the cost crossover, and control, and why most large organisations end up hybrid.
-
The Network and Storage Behind Large-Scale AI (GenAI Series, Part 26)
At scale, the network between GPUs is often the real bottleneck. How NVLink, InfiniBand and RoCE, collective operations like all-reduce, and high-throughput storage keep GPUs fed.
-
Scaling Inference: The Latency vs Throughput Trade-Off (GenAI Series, Part 25)
Scaling AI inference means choosing a point on the latency-versus-throughput curve. How batching, tensor and pipeline parallelism, and autoscaling on the right signal work.
-
vLLM vs TensorRT-LLM vs SGLang: Which Inference Engine, and When (GenAI Series, Part 24)
The inference engine decides whether a GPU serves five users or fifty. How continuous batching and paged attention work, and when to choose vLLM, TensorRT-LLM, SGLang or NIM.
-
Why GenAI Runs on GPUs, and the Memory Wall That Limits It (GenAI Series, Part 23)
Models run on GPUs for parallel matrix math, but generating text is limited by memory, not compute. Why bandwidth caps speed, VRAM caps what runs, and the KV cache fills the gap.
-
Where the Money Actually Goes in Generative AI (GenAI Series, Part 22)
Almost every dollar in generative AI is GPU time, metered as tokens. The real cost drivers, why output tokens cost more than input, and the build-versus-buy decision.
-
Guardrails and Responsible AI: What They Catch, and What They Miss (GenAI Series, Part 21)
Guardrails screen what goes into and out of an AI model. What they catch, harmful content, jailbreaks, prompt injection, data leaks, and why safety must be layered, not a single filter.
Architect’s Toolkit
PJ’s Tools
- Infra 360 Hub – All Series
- VCF 9 Interactive Walkthroughs
- VCF Design Cheatsheet
- VCF Upgrade Planner
- VCF 9 Series Hub
- VCF Deployment Hub
- AI Stack Hub
- AI Infra Sizing & Cost Calculator
- LLM & RAG Cost Calculator
- DrJhaGPT – Ask Pranay
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed

DrJha