Tag: Red Hat AI
-
Serving Granite With KServe and Serving Runtimes on OpenShift AI (Red Hat Gen AI Series, Part 17)
Take the production Granite version from the registry and turn it into a live OpenAI compatible endpoint with KServe on OpenShift AI, then learn where serving runtimes, deployment modes and autoscaling actually break under real traffic.
-
Model Registry and Versioning on OpenShift AI (Red Hat Gen AI Series, Part 16)
A tuned model with no version is a rollback you cannot make. Stand up the OpenShift AI model registry, register every Granite version from Python, and turn an hours long re-tune into a seconds long repoint.
-
Distributed Training on OpenShift AI With the Training Operator (Red Hat Gen AI Series, Part 15)
Spread the Granite retrain across GPUs with the OpenShift AI Training Operator, PyTorchJob and Kueue, and see the measurements that show when multi node is the wrong call.
-
Data Science Pipelines on OpenShift AI, From Notebook to Scheduled Retrain (Red Hat Gen AI Series, Part 14)
Build the support assistant’s nightly retrain as a Kubeflow pipeline on OpenShift AI: wire object storage, move data as artifacts, and stop a green run from shipping a worse model.
-
OpenShift AI Platform Architecture, Component by Component (Red Hat Gen AI Series, Part 13)
A working map of Red Hat OpenShift AI, the meta operator, the DataScienceCluster, and the components you should actually turn on, with the serving change most tutorials miss.
-
Red Hat AI Hybrid Cloud Deployment, and Where It Actually Runs (Red Hat Gen AI Series, Part 4)
Red Hat AI runs on bare metal, in your private cloud, on public cloud GPU instances and in air gapped sites. Here is how to place a self hosted GenAI project when its data cannot leave the building.
-
RHEL AI vs RHEL vs OpenShift AI, and Where a Project Belongs (Red Hat Gen AI Series, Part 3)
RHEL, RHEL AI and OpenShift AI get confused constantly. One is an operating system, one runs a single model on one server, one runs many across a cluster. Here is how to pick the right one for a project, with the trade offs named.
-
Granite Model Family and Choosing a Size Under Apache 2.0 (Red Hat Gen AI Series, Part 2)
Granite 4.0 comes in four practical sizes under Apache 2.0. Here is how total versus active parameters decide GPU memory, and why H-Tiny, not H-Small, is the right first model for a self hosted support assistant.
-
Red Hat AI Explained, and How RHEL AI, OpenShift AI and the Inference Server Fit (Red Hat Gen AI Series, Part 1)
Red Hat AI is not one product but three: RHEL AI, OpenShift AI and the AI Inference Server. Here is what each does, the open source thesis behind them, and where to start when you have used a hosted model API but never run your own inference.
Architect’s Toolkit
PJ’s Tools
- Infra 360 Hub – All Series
- VCF 9 Interactive Walkthroughs
- VCF Design Cheatsheet
- VCF Upgrade Planner
- VCF 9 Series Hub
- VCF Deployment Hub
- AI Stack Hub
- AI Infra Sizing & Cost Calculator
- LLM & RAG Cost Calculator
- DrJhaGPT – Ask Pranay
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed




DrJha