The Google Cloud generative AI stack, end to end, for engineers and architects: Gemini Enterprise Agent Platform, the service formerly called Vertex AI, for managed models, Model Garden for everything else, Gemini and Gemma as the models, and TPUs underneath, plus the retrieval, agent, safety, cost and governance layers that turn a model into a product. A 30 part series that reads from first principles to production. Where it meets vendor neutral ground it links to the Generative AI guide and the NVIDIA AI guide rather than repeating them.
- 01What the Google Cloud GenAI Stack Is, End to End
- 02Vertex AI and Model Garden
- 03The Gemini Family, Flash vs Pro
- 04Model Garden Third-Party Models
- 05Vertex AI vs the Gemini API in AI Studio
- 06Pricing, Provisioned Throughput and Context Caching
- 07Cloud TPUs vs GPUs
- 08Regions, Quotas and Global Endpoints
- 09Private Service Connect and VPC-SC
- 10IAM, CMEK and Data Governance
- 11Calling Models: Gemini API, Streaming, Function Calling
- 12Vertex AI Search and RAG Engine
- 13Vertex AI Agent Builder and the ADK
- 14Safety Filters and Model Armor
- 15Grounding with Google Search and Your Data
- 16Tuning Gemini with Supervised Fine-Tuning
- 17Distillation on Vertex AI
- 18Gemma Open Models
- 19Distributed Training on TPU Pods and GKE
- 20Data Prep and Embeddings
- 21Multi-Agent with ADK and Agent Engine
- 22Gemini Enterprise and Agentspace Agents
- 23The Gen AI Evaluation Service
- 24Multimodal with Veo, Imagen and Audio
- 25Observability and Tracing on Vertex
- 26Cost Governance and FinOps
- 27Responsible AI, Governance and Audit
- 28LLMOps with Vertex AI Pipelines

DrJha