Tag: Statistics
-
Model Evaluation Without Fooling Yourself (Infra to Data Science Series, Part 14)
A single new feature lifted this model AUC from 0.889 to 0.975 on infra telemetry, and none of it was real. How to catch leakage in features, preprocessing and folds, and read a score you can defend.
-
Machine Learning Fundamentals for Infra Engineers (Infra to Data Science Series, Part 12)
Machine learning fundamentals for infra engineers, using your own incident labels: why accuracy lies on rare events, and how a baseline, recall and precision decide a real model.
-
Probability and Distributions for Infra Telemetry (Infra to Data Science Series, Part 11)
Fit named distributions to your own latency and arrival data in Python, and see exactly where a clean fitted curve understates the tail that breaches your SLO.
-
Statistics for Infra Engineers, Percentiles Over Averages (Infra to Data Science Series, Part 10)
The percentiles you already trust from SLOs, made the core of statistics for infra data. Why the mean sat at the 68th percentile, why mean plus two standard deviations gave an impossible negative latency, and how a change significant at p 5.57e-05 moved the median by 0.13 ms.
-
Statistical Inference for Machine Learning: Sampling, Confidence Intervals and What a p Value Is Not (Data Science Series, Part 8)
Every metric you report is one draw from a distribution. Here is how to put a confidence interval on it, when to bootstrap, and the three readings of a p value that quietly wreck model selection.
-
Probability and Distributions a Modeller Actually Needs (Data Science Series, Part 7)
Probability is what separates a model that ranks customers from a model you can attach money to. Here is the working subset a modeller needs, with a churn worked example showing why a well ranked model can still lose cash.
-
Correlation vs Causation: The Traps Between Two Columns (Data Analyst Series, Part 12)
Two columns move together on a chart and someone declares one drives the other. Here is how to tell correlation from causation, spot the confounder, and avoid Simpson’s paradox before it reaches your slide deck.
-
Statistics a Data Analyst Actually Uses (Data Analyst Series, Part 11)
The working statistics an analyst uses daily: mean versus median, the standard deviation, the normal curve and its 68, 95, 99.7 rule, percentiles, and the sampling ideas behind margins of error and significance.
Architect’s Toolkit
PJ’s Tools
- Infra 360 Hub – All Series
- VCF 9 Interactive Walkthroughs
- VCF Design Cheatsheet
- VCF Upgrade Planner
- VCF 9 Series Hub
- VCF Deployment Hub
- AI Stack Hub
- AI Infra Sizing & Cost Calculator
- LLM & RAG Cost Calculator
- DrJhaGPT – Ask Pranay
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed




DrJha