Category: AI Stack
-
NVIDIA AI Factory Systems: DGX, HGX, MGX and NVL72 (NVIDIA AI Series, Part 5)
DGX, HGX and MGX are not performance tiers, they are three ways to integrate the same NVIDIA GPUs. Here is how they differ and where the GB200 and GB300 NVL72 rack actually earns its 120 kW.
-
GPU Memory and Precision: HBM3e, HBM4 and What Actually Fits (NVIDIA AI Series, Part 4)
A 70B model in FP16 needs 140 GB of weights before a single token of context. Here is the GPU memory and precision math that decides what fits, why HBM (not FLOPS) is the real ceiling, and where FP8 and NVFP4 buy you headroom.
-
NVIDIA Data-Center GPU Lineup: Hopper vs Blackwell vs Rubin (NVIDIA AI Series, Part 3)
The NVIDIA data-center GPU lineup from Hopper to Blackwell to Rubin, compared for training and inference: memory, bandwidth, FP4 and rack-scale NVL72, with a clear way to choose.
-
NVIDIA AI Enterprise: What the Subscription Includes and What It Costs (NVIDIA AI Series, Part 2)
NVIDIA AI Enterprise is the supported, secured wrapper around the open-source NVIDIA stack, licensed per GPU. Part 2 covers what is in the box (NIM, NeMo, Run:ai, the operators), how it is licensed (subscription, consumption, perpetual), what it costs, and when it is worth it.
-
What the NVIDIA AI Stack Actually Is, End to End (NVIDIA AI Series, Part 1)
NVIDIA AI is not one product, it is a stack roughly nine layers deep from silicon to agents. Part 1 maps the whole thing: GPUs, CUDA, the operators, TensorRT-LLM, Triton, Dynamo, NIM, NeMo, Nemotron, Blueprints and the AI Enterprise wrapper that supports it all.
-
Governing Private AI Consumption in VCF Automation: GPU Quotas, Policies and Roles (VCF Automation 9 Series, Part 35)
GPU is the most expensive resource you will ever put behind self-service. Here is how to govern Private AI consumption in VCF Automation 9.1 with namespace GPU quotas, consumption policies, roles and per-namespace activation, without touching the AI stack itself.
-
Self-Service Private AI and GPU Catalog Items in VCF Automation (VCF Automation 9 Series, Part 27)
VCF Automation can turn scarce GPUs into self-service catalog items: deep learning VMs, RAG workstations and GPU Kubernetes clusters. The Quickstart builds the blueprints in minutes. The real work is editing them and rationing the GPUs with policy.
-
The Economics and Future of Generative AI: An Honest Take (GenAI Series, Part 30)
An honest take to close the series: why GPU utilization is the real cost lever, a blunt verdict on the hype, what is actually coming, and a recap with reading paths.
-
Mixture-of-Experts and Where AI Architecture Is Heading (GenAI Series, Part 29)
Mixture-of-experts models hold enormous capacity but activate only a few experts per token, so they run cheaply. How MoE works, its memory catch, and the trends to watch.
-
What It Takes to Train a Model Across Thousands of GPUs (GenAI Series, Part 28)
Training a frontier model coordinates thousands of GPUs for months. How data, tensor, pipeline and expert parallelism, the memory math, and checkpointing make it possible.
-
On-Prem vs Cloud vs Hybrid for GenAI: An Honest Verdict (GenAI Series, Part 27)
Where should generative AI run? An honest framework weighing data sovereignty, the cost crossover, and control, and why most large organisations end up hybrid.
-
The Network and Storage Behind Large-Scale AI (GenAI Series, Part 26)
At scale, the network between GPUs is often the real bottleneck. How NVLink, InfiniBand and RoCE, collective operations like all-reduce, and high-throughput storage keep GPUs fed.
Architect’s Toolkit
PJ’s Tools
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed
