Self hosting generative AI on Red Hat, in 30 parts, for engineers who have used a hosted model API but never stood up their own inference. The series starts with what Red Hat AI actually is, then builds in order: standing a Granite model up on one RHEL AI server, tuning it with InstructLab, moving it onto OpenShift AI for scale, serving it fast through the Red Hat AI Inference Server and llm-d, then guarding, grounding and cost controlling it. One self hosted internal support assistant runs the whole way through. It assumes the GenAI Series for concepts, the AI Engineering Series for LLM engineering practice, and the Data Science Series for Python, rather than repeating them.
- 01What Red Hat AI Actually Is, and How RHEL AI, OpenShift AI and the Inference Server Fit
- 02Granite Model Family, Apache 2.0 Licensing and Choosing a Size
- 03RHEL AI vs RHEL vs OpenShift AI, and Where a Project Belongs
- 04Hybrid Cloud Deployment Model and Where Red Hat AI Runs
- 05InstructLab and the LAB Method for Taxonomy Driven Alignment
- 06Choosing a First Model and Accelerator
- 07Installing RHEL AI, the Bootable Image and First Boot
- 08Serving Granite Locally With ilab and vLLM
- 09Building a Taxonomy and Generating Synthetic Data With InstructLab
- 10Fine Tuning and Multi Phase Alignment With InstructLab
- 11Evaluating a Tuned Model, MMLU, MT-Bench and Honest Scoring
- 12Hardware and Accelerator Sizing for RHEL AI · coming soon
- 13OpenShift AI Platform Architecture · coming soon
- 14Data Science Pipelines on OpenShift AI · coming soon
- 15Distributed Training and the Training Operator · coming soon
- 16Model Registry and Versioning on OpenShift AI · coming soon
- 17Model Serving With KServe and Serving Runtimes · coming soon
- 18GPU Scheduling and Sharing, Time Slicing, MIG and Node Management · coming soon
- 19Multi Tenancy, Projects and Resource Quota · coming soon
- 20Red Hat AI Inference Server, a Hardened vLLM Distribution · coming soon
- 21vLLM Internals, PagedAttention and Continuous Batching · coming soon
- 22Model Compression and Quantization With LLM Compressor · coming soon
- 23llm-d, Distributed Inference on Kubernetes · coming soon
- 24Token Economics, Throughput and Latency Tuning · coming soon
- 25Benchmarking an Inference Deployment · coming soon
- 26AI Guardrails, Input and Output Safety on OpenShift AI · coming soon
- 27RAG on OpenShift AI With a Self Hosted Vector Store · coming soon
- 28Cost and FinOps for Self Hosted GenAI, GPU Utilisation and Unit Cost · coming soon
- 29Security, Air Gapped and Disconnected Deployments · coming soon
- 30Red Hat AI vs the Managed Clouds, the Verdict and What to Learn Next · coming soon
