Tag: OpenShift AI
-
Red Hat AI vs the Managed Clouds, the Verdict and What to Learn Next (Red Hat Gen AI Series, Part 30)
Capstone of the series: where Red Hat AI beats AWS, Google, Azure and IBM watsonx, where it loses on cost, and an honest verdict on self hosting Granite versus renting a frontier API.
-
Air Gapped and Disconnected Red Hat AI Deployments (Red Hat Gen AI Series, Part 29)
Stand up a self hosted Granite assistant in a disconnected data center: mirror images with oc-mirror v2, carry model weights across the gap, trust your registry, and lock down egress.
-
Cost and FinOps for Self Hosted GenAI on OpenShift AI (Red Hat Gen AI Series, Part 28)
Self hosting Granite rarely wins on unit cost until you reach billions of tokens a month. Here is how to price a self hosted token, measure real GPU utilisation on OpenShift AI, and find where a managed API stops being cheaper.
-
RAG on OpenShift AI With a Self Hosted Vector Store (Red Hat Gen AI Series, Part 27)
Build self hosted RAG on OpenShift AI: Llama Stack, a Milvus vector store and your served Granite model, with the vector-only gotcha that quietly costs recall.
-
OpenShift AI Guardrails for a Self Hosted Granite Assistant (Red Hat Gen AI Series, Part 26)
Input and output guardrails for a self hosted Granite assistant on OpenShift AI: deploy the FMS Guardrails Orchestrator, block PII and prompt injection, and keep the added latency inside the tail budget.
-
llm-d Distributed Inference on Kubernetes for Granite at Scale (Red Hat Gen AI Series, Part 23)
llm-d spreads vLLM inference across a Kubernetes cluster with prefill decode disaggregation and KV cache aware routing. When it pays off, and how to deploy it behind an inference gateway.
-
Multi Tenancy, Projects and Resource Quota on OpenShift AI (Red Hat Gen AI Series, Part 19)
How to give shared GPU nodes accountable owners on OpenShift AI: data science projects, ResourceQuota, LimitRange and Kueue fair-share queues, with the failures each one hides.
-
GPU Sharing on OpenShift AI With Time Slicing and MIG (Red Hat Gen AI Series, Part 18)
One 8B model on a whole A100 is a card billed at full price and used at a fraction. Here is how to split a GPU on OpenShift AI with time slicing and MIG, and which one to pick.
-
Serving Granite With KServe and Serving Runtimes on OpenShift AI (Red Hat Gen AI Series, Part 17)
Take the production Granite version from the registry and turn it into a live OpenAI compatible endpoint with KServe on OpenShift AI, then learn where serving runtimes, deployment modes and autoscaling actually break under real traffic.
-
Model Registry and Versioning on OpenShift AI (Red Hat Gen AI Series, Part 16)
A tuned model with no version is a rollback you cannot make. Stand up the OpenShift AI model registry, register every Granite version from Python, and turn an hours long re-tune into a seconds long repoint.
-
Distributed Training on OpenShift AI With the Training Operator (Red Hat Gen AI Series, Part 15)
Spread the Granite retrain across GPUs with the OpenShift AI Training Operator, PyTorchJob and Kueue, and see the measurements that show when multi node is the wrong call.
-
Data Science Pipelines on OpenShift AI, From Notebook to Scheduled Retrain (Red Hat Gen AI Series, Part 14)
Build the support assistant’s nightly retrain as a Kubeflow pipeline on OpenShift AI: wire object storage, move data as artifacts, and stop a green run from shipping a worse model.
Architect’s Toolkit
PJ’s Tools
- Infra 360 Hub – All Series
- VCF 9 Interactive Walkthroughs
- VCF Design Cheatsheet
- VCF Upgrade Planner
- VCF 9 Series Hub
- VCF Deployment Hub
- AI Stack Hub
- AI Infra Sizing & Cost Calculator
- LLM & RAG Cost Calculator
- DrJhaGPT – Ask Pranay
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed

DrJha