Tag: GPU
-
Running NVIDIA AI On-Prem and on VCF: Cost, Trade-offs and the Verdict (NVIDIA AI Series, Part 30)
The finale: running the NVIDIA AI stack on bare metal, on VMware Cloud Foundation, or in the cloud; the real total cost of an AI factory; and the verdict on when to build versus rent.
-
GPU Observability and Multi-Tenancy: DCGM, Honest Utilization, and Sharing (NVIDIA AI Series, Part 29)
Why GPU utilization lies, the DCGM profiling fields that tell the truth (SM and Tensor activity), dcgm-exporter into Prometheus, and choosing MIG vs time-slicing for multi-tenancy.
-
NVIDIA NeMo Framework: Training and Fine-Tuning at Scale (NVIDIA AI Series, Part 22)
What the NVIDIA NeMo framework is: Megatron-Core parallelism, NeMo 2.0 Python recipes and NeMo-Run, Megatron Bridge for Hugging Face interop, and when to fine-tune instead of pretrain.
-
Multi-Node LLM Training: Scheduling, Checkpointing and Fault Tolerance (NVIDIA AI Series, Part 25)
At thousands of GPUs, failures are routine. This part covers gang scheduling (Slurm vs Kubernetes vs NVIDIA Run:ai), async distributed checkpointing with NeMo, and the NVIDIA Resiliency Extension stack for fault tolerance, straggler detection, and elastic restart.
-
Inference Economics: Throughput, Latency, Batching and Cost Per Token (NVIDIA AI Series, Part 21)
TTFT, ITL, continuous batching, KV cache pressure, FP8 quantization — this is how you compute and actually drive down $/1M tokens on NVIDIA H100, H200, and B200 GPUs without breaking your latency SLOs.
-
NeMo Customization: LoRA, SFT, and RLHF on NVIDIA NeMo (NVIDIA AI Series, Part 23)
A practical decision guide for AI infrastructure architects on the full NeMo customization spectrum: when to use LoRA, full SFT, DPO, or GRPO, what data and GPU budget each method needs, and how the NeMo Customizer microservice ties it all together.
-
Data Preparation at Scale with NeMo Curator (NVIDIA AI Series, Part 24)
NeMo Curator is NVIDIA’s GPU-accelerated data curation toolkit that runs exact dedup, fuzzy dedup, semantic dedup, heuristic filtering, classifier-based quality filters, and PII redaction at trillion-token scale using RAPIDS cuDF and Dask. Learn why investing in data curation beats buying more GPUs.
-
NVIDIA Network Operator on Kubernetes: RDMA, SR-IOV, and the Accelerated Fabric (NVIDIA AI Series, Part 13)
The NVIDIA Network Operator provisions MOFED drivers, RDMA shared device plugin, SR-IOV VFs, and Multus secondary networks to Kubernetes pods. This is how GPUDirect RDMA actually works at scale on ConnectX-7 and NDR InfiniBand clusters.
-
NVIDIA Drivers, CUDA, and the Container Toolkit: Building a Clean GPU Host Baseline (NVIDIA AI Series, Part 11)
The GPU host stack has three distinct layers: the data-center driver (open kernel module now required for Hopper and Blackwell), the CUDA Toolkit, and the NVIDIA Container Toolkit. Get the install order or versions wrong and containers fail silently. Here is the right sequence, the compatibility matrix, and the failure modes.
-
NGC Catalog: Containers, Models, Helm Charts and How to Consume Them (NVIDIA AI Series, Part 14)
The NGC catalog is your upstream source for NVIDIA GPU-optimized containers, pretrained models, and Helm charts. Here is how the nvcr.io registry, org/team/API-key model, and NVAIE entitlement actually work, with a full operational pull-and-deploy walkthrough.
-
NVIDIA GPU Operator on Kubernetes: ClusterPolicy, Components, and Day-2 Ops (NVIDIA AI Series, Part 12)
The NVIDIA GPU Operator automates every software layer a GPU node needs in Kubernetes, from kernel driver to DCGM metrics, via a single ClusterPolicy CRD. Here is what it installs, how the reconciliation loop works, when to disable the driver component, and the failure modes that will catch you on first install.
-
InfiniBand vs Spectrum-X Ethernet: Choosing Your AI Cluster Scale-Out Fabric (NVIDIA AI Series, Part 8)
InfiniBand Quantum-X800 and Spectrum-X Ethernet both run at 800 Gb/s — but they are not the same choice. A direct comparison of SHARPv4 in-network reduction, lossless fabric mechanisms, rail-optimized topology, multi-tenant isolation, and operational trade-offs, with a clear verdict on which fabric wins for dedicated AI training versus shared enterprise GPU platforms.
Architect’s Toolkit
PJ’s Tools
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed
