Tag: vLLM
-
Text Generation Inference (TGI) in Production: A Real Serving Example (Hugging Face Series, Part 12)
TGI turns a Hugging Face model into an OpenAI-compatible endpoint with one docker run. Here are the flags that decide whether it fits your VRAM, how to consume it, and an honest verdict now that TGI is in maintenance mode and Hugging Face points new builds at vLLM.
-
What Is in the Private AI Catalog: A Guide to the Blueprints
A plain-language field guide to every blueprint in the Private AI catalog, grouped by job: compute, model serving, RAG retrieval, OCR and speech, and the access layer.
-
vLLM vs TensorRT-LLM vs SGLang: Which Inference Engine, and When (GenAI Series, Part 24)
The inference engine decides whether a GPU serves five users or fifty. How continuous batching and paged attention work, and when to choose vLLM, TensorRT-LLM, SGLang or NIM.
-
VMware Private AI Services: Deploying Models with the Model Store and Model Runtime (Private AI Series, Part 12)
A hands-on runbook for Private AI Services 2.1: stand up a Harbor model gallery, validate and push models with the vcf pais CLI, then serve them as endpoints through Model Runtime and the ML API Gateway.
Architect’s Toolkit
PJ’s Tools
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed

DrJha