Tag: GPU
-
GPUDirect Storage: DMA From NVMe Straight to GPU Memory (NVIDIA AI Series, Part 9)
GPUDirect Storage (GDS) creates a direct DMA path from NVMe or networked storage straight into GPU HBM, bypassing the CPU bounce buffer entirely. Here is when it helps, what the cuFile API requires, and the filesystem and NIC prerequisites to validate before enabling in production.
-
GPU Power, Cooling and Density: Why Blackwell Forces Liquid (NVIDIA AI Series, Part 10)
The GB200 NVL72 draws ~120 kW per rack and ships liquid-cooled by design. Learn why Blackwell-class systems make direct-to-chip cooling mandatory, how CDUs and facility water loops work, and what to validate before ordering.
-
NVLink and NVSwitch: How NVIDIA Builds the Scale-Up Fabric (NVIDIA AI Series, Part 7)
Fifth-generation NVLink delivers 1.8 TB/s per GPU, and NVSwitch builds a non-blocking 130 TB/s all-to-all fabric across 72 GPUs in the GB200 NVL72. Here is how the domain forms, why it determines your tensor and expert parallelism strategy, and where the boundary falls.
-
GPU Partitioning on NVIDIA Data-Center GPUs: MIG vs vGPU vs Time-Slicing vs Passthrough (NVIDIA AI Series, Part 6)
Four ways to partition an NVIDIA H100, H200, or B200 GPU: MIG, vGPU, CUDA time-slicing, and full passthrough. This post covers the isolation guarantees, profile geometry, Kubernetes GPU Operator configuration, and a sizing worked example to help you pick the right mode for your cluster.
-
NVIDIA AI Factory Systems: DGX, HGX, MGX and NVL72 (NVIDIA AI Series, Part 5)
DGX, HGX and MGX are not performance tiers, they are three ways to integrate the same NVIDIA GPUs. Here is how they differ and where the GB200 and GB300 NVL72 rack actually earns its 120 kW.
-
GPU Memory and Precision: HBM3e, HBM4 and What Actually Fits (NVIDIA AI Series, Part 4)
A 70B model in FP16 needs 140 GB of weights before a single token of context. Here is the GPU memory and precision math that decides what fits, why HBM (not FLOPS) is the real ceiling, and where FP8 and NVFP4 buy you headroom.
-
NVIDIA Data-Center GPU Lineup: Hopper vs Blackwell vs Rubin (NVIDIA AI Series, Part 3)
The NVIDIA data-center GPU lineup from Hopper to Blackwell to Rubin, compared for training and inference: memory, bandwidth, FP4 and rack-scale NVL72, with a clear way to choose.
-
Governing Private AI Consumption in VCF Automation: GPU Quotas, Policies and Roles (VCF Automation 9 Series, Part 35)
GPU is the most expensive resource you will ever put behind self-service. Here is how to govern Private AI consumption in VCF Automation 9.1 with namespace GPU quotas, consumption policies, roles and per-namespace activation, without touching the AI stack itself.
-
Self-Service Private AI and GPU Catalog Items in VCF Automation (VCF Automation 9 Series, Part 27)
VCF Automation can turn scarce GPUs into self-service catalog items: deep learning VMs, RAG workstations and GPU Kubernetes clusters. The Quickstart builds the blueprints in minutes. The real work is editing them and rationing the GPUs with policy.
-
Why GenAI Runs on GPUs, and the Memory Wall That Limits It (GenAI Series, Part 23)
Models run on GPUs for parallel matrix math, but generating text is limited by memory, not compute. Why bandwidth caps speed, VRAM caps what runs, and the KV cache fills the gap.
-
Where the Money Actually Goes in Generative AI (GenAI Series, Part 22)
Almost every dollar in generative AI is GPU time, metered as tokens. The real cost drivers, why output tokens cost more than input, and the build-versus-buy decision.
-
Quantization: Running Big Models on Smaller GPUs (GenAI Series, Part 20)
Quantization stores a model at lower precision so it needs far less memory. How FP16, INT8 and INT4 trade a little quality for big savings, plus distillation and pruning.
Architect’s Toolkit
PJ’s Tools
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed
