Tag: llm
-
Prompt Engineering That Actually Works (GenAI Series, Part 12)
Prompt engineering is not secret incantations, it is clear communication. The four moves that do most of the work, system vs user prompts, and the anti-patterns that waste tokens.
-
Why AI Models Make Things Up (and What Temperature Does) (GenAI Series, Part 11)
AI models generate by sampling likely words from a probability distribution. Why that produces confident hallucinations, what the temperature setting really does, and how to reduce it.
-
The Context Window, and Why Models Forget (GenAI Series, Part 10)
The context window is everything an AI can see at once. Why models have no memory between turns, why longer prompts cost more, and why details get lost in the middle.
-
Attention, the Idea That Made Modern AI Work (GenAI Series, Part 8)
How attention lets every word in a sentence weigh every other word, why it replaced slow left-to-right models, and why running in parallel is what let AI scale.
-
What a Model Really Is (GenAI Series, Part 4)
A model is not a database of answers. It is one large function that predicts the next token, built from billions of parameters. What model sizes and open vs closed weights really mean.
-
The GenAI Words Everyone Uses, and What They Actually Mean (GenAI Series, Part 2)
Model, tokens, parameters, inference, embeddings, hallucination: the words everyone uses about generative AI, sorted into build time and use time and explained in plain English.
-
What Generative AI Actually Is, Explained Without the Hype (GenAI Series, Part 1)
Generative AI does not look answers up. It builds new text, images and code by predicting one piece at a time. Here is what that really means, in plain English, with GPT-4o and Llama as the examples.
-
Building Enterprise AI with NVIDIA NeMo Microservices: From Data to Guardrails
The GenAI wave is no longer about just calling an LLM API. It’s about building reliable, scalable, secure, and continuously improving AI systems. While many teams are still experimenting with prompts, enterprises are moving toward something bigger: 👉 AI factories powered by microservices And that’s exactly where NVIDIA NeMo comes in. The Big Picture: Enterprise…
-
What is NVIDIA NeMo — and Why It Matters for Agentic AI
When people talk about AI systems, they often focus on models or APIs. But once you move beyond simple use cases, a bigger challenge appears: How do you control, guide, and manage AI behavior in real-world systems? This is where NVIDIA NeMo becomes critical. If NIM is the layer that runs AI models, then NeMo…
-
What is NVIDIA NIM — and Why It Matters for Modern AI Systems
When most people start learning AI, they focus on models—LLMs, vision models, embeddings, and so on. But in real-world systems, models alone are not enough. The real challenge is how to run these models reliably, at scale, and in a way that applications can actually use them. This is exactly where NVIDIA NIM comes into…
-
Preparing for NVIDIA Certified Professional – AI Infrastructure? Here’s What Actually Happened to Me
“I’m not from AI… can I even do this?” When I first decided to prepare for the NVIDIA Certified Professional – AI Infrastructure, I had a very honest thought: “I’m not from a pure AI background… how am I going to understand all this?” Yes, I did have some exposure to AI/ML during my PhD.…
-
What is Ollama? A Simple explanation to understand!
Artificial Intelligence might sound complicated, but tools like Ollama are making it simple for everyone. Imagine having a smart personal assistant that can answer questions, help you write, or even solve problems, right on your own computer. That’s what Ollama does! Usually, AI tools like ChatGPT live in the cloud. That means your questions and…
Architect’s Toolkit
PJ’s Tools
- Infra 360 Hub – All Series
- VCF 9 Interactive Walkthroughs
- VCF Design Cheatsheet
- VCF Upgrade Planner
- VCF 9 Series Hub
- VCF Deployment Hub
- AI Stack Hub
- AI Infra Sizing & Cost Calculator
- LLM & RAG Cost Calculator
- DrJhaGPT – Ask Pranay
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed

DrJha