Tag: quantization
-
optimum and Quantization: ONNX, GPTQ and AWQ for the GPUs You Have (Hugging Face Series, Part 13)
Quantization is how you fit a model on the GPUs you already own. A field guide to optimum, ONNX Runtime, and 4-bit GPTQ vs AWQ, written for the infra engineer who has to make the capacity math work.
-
TensorRT and TensorRT-LLM: Optimization, Quantization, and Engine Building (NVIDIA AI Series, Part 18)
What TensorRT does at build time versus what TensorRT-LLM adds at runtime — kernel fusion, paged KV cache, in-flight batching, and quantization choices from FP8 to NVFP4 — and when to hand-build engines instead of relying on a NIM.
-
GPU Memory and Precision: HBM3e, HBM4 and What Actually Fits (NVIDIA AI Series, Part 4)
A 70B model in FP16 needs 140 GB of weights before a single token of context. Here is the GPU memory and precision math that decides what fits, why HBM (not FLOPS) is the real ceiling, and where FP8 and NVFP4 buy you headroom.
-
Quantization: Running Big Models on Smaller GPUs (GenAI Series, Part 20)
Quantization stores a model at lower precision so it needs far less memory. How FP16, INT8 and INT4 trade a little quality for big savings, plus distillation and pruning.
Architect’s Toolkit
PJ’s Tools
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed

DrJha