Tag: AI cost
-
Token Budgets, Context Limits and Cost per LLM Request (AI Engineering Series, Part 5)
One tool schema adds 389 tokens to every request. A newer tokenizer can add 30 percent to your bill with no price change. Here is how to measure token cost before a request leaves your process, and what our documentation assistant actually costs per 1,000 calls.
-
GPU Cost, Scale and Sizing Decisions for Machine Learning Workloads (Data Science Series, Part 28)
Nine percent average GPU utilisation on a cluster I inherited turned out to be the whole story. Here is how I size, buy and scale machine learning compute, with current instance rates and the arithmetic that decides each call.
-
Where the Money Actually Goes in Generative AI (GenAI Series, Part 22)
Almost every dollar in generative AI is GPU time, metered as tokens. The real cost drivers, why output tokens cost more than input, and the build-versus-buy decision.
-
Training vs Inference: Why Using AI Is the Real Cost (GenAI Series, Part 9)
Training builds a model once in three stages; inference runs it on every request, forever. Why the recurring inference bill, not the headline training cost, decides AI economics.
Architect’s Toolkit
PJ’s Tools
- Infra 360 Hub – All Series
- VCF 9 Interactive Walkthroughs
- VCF Design Cheatsheet
- VCF Upgrade Planner
- VCF 9 Series Hub
- VCF Deployment Hub
- AI Stack Hub
- AI Infra Sizing & Cost Calculator
- LLM & RAG Cost Calculator
- DrJhaGPT – Ask Pranay
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed

DrJha