Category: AI Stack
-
Amazon Bedrock Pricing Across On-Demand, Provisioned, and Batch (AWS Gen AI Series, Part 6)
The five ways Amazon Bedrock charges for the same model, from on-demand tokens to reserved model units, and the break-even math that tells you which mode your workload actually belongs on.
-
Amazon Bedrock vs SageMaker AI, and When to Use Each (AWS Gen AI Series, Part 5)
Bedrock gives you models behind an API; SageMaker AI gives you the whole ML platform. Here is how I decide between them, with the cost math that usually settles it.
-
Amazon Nova Models and Where Each One Fits (AWS Gen AI Series, Part 4)
A working tour of Amazon Nova on Bedrock: Micro, Lite, Pro and Premier, the creative and speech models, and what Nova 2 changes. With model IDs, context sizes, real cost math and the inference-profile trap that breaks first calls.
-
Amazon Bedrock Model Catalog and Choosing a Model (AWS Gen AI Series, Part 3)
Bedrock ships more than a hundred models across fifteen providers, and the price gap between the cheapest and the priciest is over 400x. Here is how I read the catalog and pick a model without overpaying.
-
Amazon Bedrock and the Shared Responsibility Model (AWS Gen AI Series, Part 2)
Bedrock is a managed service, but security is still split between AWS and you. Here is exactly which half is yours, the defaults that catch teams out, and the baseline I deploy before any prompt goes live.
-

AWS Generative AI Stack, End to End (AWS Gen AI Series, Part 1)
AWS generative AI is really three layers: Amazon Bedrock for managed models, SageMaker AI to build your own, and Trainium and Inferentia underneath. Here is the whole map, with a real cost example and a first call that works.
-
Which Model for What: a Hugging Face Model Map for Text, Vision, Audio and Video (Hugging Face Series, Part 17)
A task-to-model map for Hugging Face: which model family to use for chat, search, transcription, speech, captioning, image and video, with sizes, licenses, the right library, and the GPU it needs.
-
Hugging Face Air-Gapped: Enterprise Hub, Offline Mirroring, and the On-Prem Build (Hugging Face Series, Part 16)
How to run Hugging Face models on a segmented or air-gapped network: mirror the artifacts to local storage, force offline mode at runtime, and use the Enterprise Hub for identity, governance, and the rate limits a proxy needs.
-
Hugging Face Security and Governance: Gated Models, Malicious Weights, and Scanning Before Anything Enters Your Registry (Hugging Face Series, Part 15)
A model download is an unvetted artifact, and a pickle checkpoint can run code the moment you load it. Here is how to gate the Hugging Face Hub the way you already gate a container registry: safetensors over pickle, no blind trust_remote_code, scan before promote, pin for provenance.
-
Hugging Face Spaces and Gradio: Great for Demos, Wrong for Production (Hugging Face Series, Part 14)
A Hugging Face Space turns a model into a shareable Gradio demo in three files. Here is how a Space runs, a worked example, and the hard line where a demo must graduate to real serving.
-
optimum and Quantization: ONNX, GPTQ and AWQ for the GPUs You Have (Hugging Face Series, Part 13)
Quantization is how you fit a model on the GPUs you already own. A field guide to optimum, ONNX Runtime, and 4-bit GPTQ vs AWQ, written for the infra engineer who has to make the capacity math work.
-
Text Generation Inference (TGI) in Production: A Real Serving Example (Hugging Face Series, Part 12)
TGI turns a Hugging Face model into an OpenAI-compatible endpoint with one docker run. Here are the flags that decide whether it fits your VRAM, how to consume it, and an honest verdict now that TGI is in maintenance mode and Hugging Face points new builds at vLLM.
Architect’s Toolkit
PJ’s Tools
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed
