Red Hat Gen AI, Complete Guide

Self hosting generative AI on Red Hat, in 30 parts, for engineers who have used a hosted model API but never stood up their own inference. The series starts with what Red Hat AI actually is, then builds in order: standing a Granite model up on one RHEL AI server, tuning it with InstructLab, moving it onto OpenShift AI for scale, serving it fast through the Red Hat AI Inference Server and llm-d, then guarding, grounding and cost controlling it. One self hosted internal support assistant runs the whole way through. It assumes the GenAI Series for concepts, the AI Engineering Series for LLM engineering practice, and the Data Science Series for Python, rather than repeating them.

Complete · all 30 parts published
Phase 1 · Foundations
  1. 01What Red Hat AI Actually Is, and How RHEL AI, OpenShift AI and the Inference Server Fit
  2. 02Granite Model Family, Apache 2.0 Licensing and Choosing a Size
  3. 03RHEL AI vs RHEL vs OpenShift AI, and Where a Project Belongs
  4. 04Hybrid Cloud Deployment Model and Where Red Hat AI Runs
  5. 05InstructLab and the LAB Method for Taxonomy Driven Alignment
  6. 06Choosing a First Model and Accelerator
Phase 2 · RHEL AI on one server
  1. 07Installing RHEL AI, the Bootable Image and First Boot
  2. 08Serving Granite Locally With ilab and vLLM
  3. 09Building a Taxonomy and Generating Synthetic Data With InstructLab
  4. 10Fine Tuning and Multi Phase Alignment With InstructLab
  5. 11Evaluating a Tuned Model, MMLU, MT-Bench and Honest Scoring
  6. 12Hardware and Accelerator Sizing for RHEL AI
Phase 3 · OpenShift AI at scale
  1. 13OpenShift AI Platform Architecture
  2. 14Data Science Pipelines on OpenShift AI
  3. 15Distributed Training and the Training Operator
  4. 16Model Registry and Versioning on OpenShift AI
  5. 17Model Serving With KServe and Serving Runtimes
  6. 18GPU Scheduling and Sharing, Time Slicing, MIG and Node Management
  7. 19Multi Tenancy, Projects and Resource Quota
Phase 4 · Inference and serving
  1. 20Red Hat AI Inference Server, a Hardened vLLM Distribution
  2. 21vLLM Internals, PagedAttention and Continuous Batching
  3. 22Model Compression and Quantization With LLM Compressor
  4. 23llm-d, Distributed Inference on Kubernetes
  5. 24Token Economics, Throughput and Latency Tuning
  6. 25Benchmarking an Inference Deployment
Phase 5 · Production and governance
  1. 26AI Guardrails, Input and Output Safety on OpenShift AI
  2. 27RAG on OpenShift AI With a Self Hosted Vector Store
  3. 28Cost and FinOps for Self Hosted GenAI, GPU Utilisation and Unit Cost
  4. 29Security, Air Gapped and Disconnected Deployments
  5. 30Red Hat AI vs the Managed Clouds, the Verdict and What to Learn Next

Architect’s Toolkit

About the Author

Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.