Tag: Azure OpenAI
-
Azure Generative AI vs the Field, the Verdict (Azure Gen AI Series, Part 30)
After twenty-nine parts on the Azure GenAI stack, here is the plain verdict: where Azure wins, where it loses to AWS and Google Cloud, what the bill really looks like, and who should build on it.
-
Reference Architectures for Azure GenAI, from Chatbot to Batch (Azure Gen AI Series, Part 29)
The four architectures almost every Azure GenAI project actually needs, chatbot, retrieval, agentic, and batch, with the baseline Foundry components, real sizing numbers, and when to add an API Management gateway.
-
Azure OpenAI Cost Governance and FinOps (Azure Gen AI Series, Part 26)
When a Provisioned Throughput reservation actually beats pay as you go on Azure OpenAI, how to read a bill hidden under Cognitive Services, and the Batch and caching discounts you get for free.
-
Azure Monitor Observability for GenAI, from Metrics to Traces (Azure Gen AI Series, Part 25)
Azure GenAI observability comes in three layers: free platform metrics, per request diagnostic logs in Log Analytics, and OpenTelemetry traces in Application Insights. Here is what each one sees, what it costs, and the order I turn them on.
-
Vision, Audio, Image, and Document Models on Azure OpenAI (Azure Gen AI Series, Part 24)
Multimodal on Azure is not one switch, it is four services. Where vision, voice, image generation, and Content Understanding each live, what they cost, and which one owns each input.
-
Azure OpenAI Distillation and Stored Completions (Azure Gen AI Series, Part 17)
Capture production traffic with store=True, then distill a small Azure OpenAI model that answers like a flagship. The workflow, the real costs, and the traffic volume where it pays off.
-
Fine-Tuning Azure OpenAI, from SFT to DPO and RFT (Azure Gen AI Series, Part 16)
SFT, DPO, and RFT on Azure OpenAI: which models take which method, what the training and hosting actually cost, and how to read the loss curve before you deploy.
-
Calling Azure OpenAI Models with REST, the SDKs, and the Responses API (Azure Gen AI Series, Part 11)
Azure gives you three ways to call a model, and they are not interchangeable. How Chat Completions, the Responses API, and raw REST differ, and when to pick each on the v1 path.
-
Azure OpenAI Entra ID Authentication, Managed Identities, and Encryption (Azure Gen AI Series, Part 10)
Move Azure OpenAI off static API keys to Microsoft Entra ID and managed identities, disable local auth without locking yourself out, and know when customer-managed keys are worth the cost.
-
Azure OpenAI Private Link and VNet Isolation (Azure Gen AI Series, Part 9)
Private networking for Azure OpenAI, done in the right order: private endpoints, the DNS zones that make them work, when a managed VNet earns its keep, and the On Your Data trap that closes every call.
-
Azure OpenAI Regions, Quotas, and Data Zones (Azure Gen AI Series, Part 8)
Azure OpenAI quota is a grid: per region, per subscription, per model, per deployment type. Here are the real 2026 numbers, the new quota tiers, how data zones change residency, and the two commands I run before promising any capacity.
-
Azure OpenAI Deployment Types, Standard vs PTU vs Batch (Azure Gen AI Series, Part 6)
Standard, provisioned throughput units, and Batch are the three ways to bill an Azure OpenAI deployment. How to pick with a utilization break-even, size PTUs, and use spillover so you are not throttled or overpaying.
Architect’s Toolkit
PJ’s Tools
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed

DrJha