Tag: Granite
-
RAG on OpenShift AI With a Self Hosted Vector Store (Red Hat Gen AI Series, Part 27)
Build self hosted RAG on OpenShift AI: Llama Stack, a Milvus vector store and your served Granite model, with the vector-only gotcha that quietly costs recall.
-
OpenShift AI Guardrails for a Self Hosted Granite Assistant (Red Hat Gen AI Series, Part 26)
Input and output guardrails for a self hosted Granite assistant on OpenShift AI: deploy the FMS Guardrails Orchestrator, block PII and prompt injection, and keep the added latency inside the tail budget.
-
Token Economics and Latency Tuning for Self Hosted Granite (Red Hat Gen AI Series, Part 24)
What one answer from a self hosted Granite model actually costs, and the three vLLM flags that decide it. A latency aware guide to throughput, TTFT and cost per token on the Red Hat AI Inference Server.
-
Model Compression and Quantization With LLM Compressor (Red Hat Gen AI Series, Part 22)
Quantizing Granite with LLM Compressor cuts the support assistant from 16 GB to under 5 GB of weights. A format by format comparison of FP8, INT8, INT4 and NVFP4, with the memory, accuracy recovery and hardware trade offs named.
-
Red Hat AI Inference Server, a Hardened vLLM Distribution (Red Hat Gen AI Series, Part 20)
Red Hat AI Inference Server is upstream vLLM packaged as a supported, hardened container. Here is what it adds, how to serve Granite behind it, and how to benchmark it honestly.
-
Model Registry and Versioning on OpenShift AI (Red Hat Gen AI Series, Part 16)
A tuned model with no version is a rollback you cannot make. Stand up the OpenShift AI model registry, register every Granite version from Python, and turn an hours long re-tune into a seconds long repoint.
-
Distributed Training on OpenShift AI With the Training Operator (Red Hat Gen AI Series, Part 15)
Spread the Granite retrain across GPUs with the OpenShift AI Training Operator, PyTorchJob and Kueue, and see the measurements that show when multi node is the wrong call.
-
Data Science Pipelines on OpenShift AI, From Notebook to Scheduled Retrain (Red Hat Gen AI Series, Part 14)
Build the support assistant’s nightly retrain as a Kubeflow pipeline on OpenShift AI: wire object storage, move data as artifacts, and stop a green run from shipping a worse model.
-
Hardware and Accelerator Sizing for RHEL AI (Red Hat Gen AI Series, Part 12)
Sizing GPUs for RHEL AI is a KV cache problem, not a weights problem. A worked memory calculator, a per accelerator concurrency table for Granite 3.1 8B, and the one flag that wakes a stalled server.
-
Evaluating a Tuned Granite Model With MMLU and MT-Bench (Red Hat Gen AI Series, Part 11)
How to score a tuned Granite model honestly on RHEL AI with MMLU, MT-Bench, MMLU Branch and DK-Bench, and why the branch scores, not MMLU, decide whether to ship.
-
Fine Tuning Granite With InstructLab Multi Phase Alignment (Red Hat Gen AI Series, Part 10)
Run lab-multiphase training on RHEL AI to tune Granite on your own docs, read the checkpoints MT-Bench actually picks, and avoid the restart prompt that wipes hours of work.
-
InstructLab Taxonomy and Synthetic Data Generation on RHEL AI (Red Hat Gen AI Series, Part 9)
Building an InstructLab knowledge taxonomy and running synthetic data generation on RHEL AI, from qna.yaml seed examples to the training JSONL that Part 10 tunes on.
Architect’s Toolkit
PJ’s Tools
- Infra 360 Hub – All Series
- VCF 9 Interactive Walkthroughs
- VCF Design Cheatsheet
- VCF Upgrade Planner
- VCF 9 Series Hub
- VCF Deployment Hub
- AI Stack Hub
- AI Infra Sizing & Cost Calculator
- LLM & RAG Cost Calculator
- DrJhaGPT – Ask Pranay
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed




DrJha