Category: AI Stack
-
Choosing a First Model and Accelerator for Red Hat AI (Red Hat Gen AI Series, Part 6)
Sizing a first Granite model and GPU for Red Hat AI is a memory problem, not a benchmark one. Here is how weights, KV cache and the InstructLab training floor decide what you actually buy or rent.
-
InstructLab and the LAB Method for Taxonomy Driven Alignment (Red Hat Gen AI Series, Part 5)
InstructLab turns a handful of hand written questions into thousands of training examples. Here is how the LAB method and a taxonomy tree tune a Granite model on your own documents.
-
Red Hat AI Hybrid Cloud Deployment, and Where It Actually Runs (Red Hat Gen AI Series, Part 4)
Red Hat AI runs on bare metal, in your private cloud, on public cloud GPU instances and in air gapped sites. Here is how to place a self hosted GenAI project when its data cannot leave the building.
-
RHEL AI vs RHEL vs OpenShift AI, and Where a Project Belongs (Red Hat Gen AI Series, Part 3)
RHEL, RHEL AI and OpenShift AI get confused constantly. One is an operating system, one runs a single model on one server, one runs many across a cluster. Here is how to pick the right one for a project, with the trade offs named.
-
Granite Model Family and Choosing a Size Under Apache 2.0 (Red Hat Gen AI Series, Part 2)
Granite 4.0 comes in four practical sizes under Apache 2.0. Here is how total versus active parameters decide GPU memory, and why H-Tiny, not H-Small, is the right first model for a self hosted support assistant.
-
Red Hat AI Explained, and How RHEL AI, OpenShift AI and the Inference Server Fit (Red Hat Gen AI Series, Part 1)
Red Hat AI is not one product but three: RHEL AI, OpenShift AI and the AI Inference Server. Here is what each does, the open source thesis behind them, and where to start when you have used a hosted model API but never run your own inference.
-
Path to AI Engineer, and What to Learn Next (AI Engineering Series, Part 30)
Thirty parts on, here is the honest version of the AI engineering career path: what the market pays, which routes into the role actually work, what a hiring loop tests, and a twelve month plan to close the gaps this series left open.
-
Choosing and Switching Providers Without a Rewrite (AI Engineering Series, Part 29)
Provider lock in for an LLM application does not live in the API call. I compare six portability strategies with measured latency overhead, show the compatibility endpoint failure that cost us three days, and name the one I would ship.
-
Deployment, Versioning and Rollback for Prompts and Models (AI Engineering Series, Part 28)
A model string in three files is not a deployment story. Here is how I pin model snapshots, version prompts in a registry, canary by tenant hash, and get rollback down from 14 minutes to 8 seconds.
-
Cost Control and Model Routing for LLM Applications (AI Engineering Series, Part 27)
Model routing saved our documentation assistant 22 percent. Trimming retrieval saved more, in one line of config. Here is the cost model, the routing decision table, and the measured numbers behind both.
-
Caching, Batching and Latency Engineering for LLM Applications (AI Engineering Series, Part 26)
Prompt caching, batch processing and streaming, measured on a production docs assistant. Where the breakpoint really goes, why the obvious block to cache is the wrong one, and the cost arithmetic that follows.
-
Observability for LLM Applications: Tracing and Debugging a Non Deterministic System (AI Engineering Series, Part 25)
A request level trace is the only artifact that explains why an LLM answer was wrong, slow or expensive. Here is how I instrument one with OpenTelemetry GenAI conventions and Langfuse, which attributes actually earn their storage, and what tracing costs once traffic is real.
Architect’s Toolkit
PJ’s Tools
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed
