Tag: multimodal AI
-
Multimodal on Vertex AI, from Nano Banana to Veo and Lyria (Google Cloud Gen AI Series, Part 24)
A field guide to generative media on Vertex AI: when to use Gemini native image generation, Imagen, Veo, Lyria, and Chirp, with real model IDs and a per-image cost model.
-
Vision, Audio, Image, and Document Models on Azure OpenAI (Azure Gen AI Series, Part 24)
Multimodal on Azure is not one switch, it is four services. Where vision, voice, image generation, and Content Understanding each live, what they cost, and which one owns each input.
-
Amazon Bedrock Data Automation for Multimodal Content (AWS Gen AI Series, Part 24)
A practical walk through Amazon Bedrock Data Automation: standard output versus custom blueprints, the async API, real per-page and per-minute pricing, and when to wire it into a Knowledge Base.
Architect’s Toolkit
PJ’s Tools
VMware Cloud Foundation
- VCF Documentation
- VCF 9 Planning & Preparation Workbook
- VCF Bill of Materials (BoM)
- VMware Compatibility Guide
- VMware Interoperability Matrix
- VMware Configuration Maximums
- VMware Ports & Protocols
- VMware Hands-on Labs
- RVTools Download
Nutanix
AI & Cloud-Native Platform
- NVIDIA Build (Model Catalog)
- NVIDIA AI Enterprise Reference Architecture
- NVIDIA NIM Performance Benchmarking
- NVIDIA NGC Catalog
- NeMo Microservices Helm Chart
- Helm Charts Repository
- Hugging Face Models
Architecture & Design
About the Author

Dr Pranay Jha
Dr. Pranay Jha is a Cloud and AI Consultant with 18+ years of experience in hybrid cloud, virtualization, and enterprise infrastructure transformation. He specializes in VMware technologies, multi-cloud strategy, and Generative AI solutions. He holds a PhD in Computer Applications with research focused on Cloud and AI, has published multiple research papers, and has been a VMware vExpert since 2016 and a VMUG Community Leader.
You May Have Missed

DrJha