Intrigued by the art of cloud architecture? Discover how to design, develop, and manage robust, secure, scalable, and dynamic solutions on Google Cloud as you prepare for the Professional Cloud Architect exam!
Prepare and test your skills
Prepare and test your skills
Worked example. The correct answer is already marked and every option is explained below, so there is nothing to select here. To answer questions yourself, start the free trial.
Keep the momentum going with these hand-picked practice scenarios
Want more questions like this?
Get a free certification question every week.
Last updated
An enterprise is architecting an AI-powered customer support ecosystem on Google Cloud with two distinct workload requirements:
You need to analyze Model Garden offerings and select the appropriate foundation models and deployment architectures to optimize both performance and cost.
Which model selection and architecture strategy should you recommend?
Deploy ShieldGemma 2 to handle agent multi-step planning and tool calling, and use Gemini Pro with maximum thinking budget for the high-volume classification tasks.
Use PaliGemma 2 via the managed Gemini API for the core orchestrator, and deploy Gemini Pro on dedicated GKE clusters for high-volume triage.
Perform full fine-tuning of Gemini Pro for the triage workload and deploy it as a self-hosted container, while using MedGemma for general agent orchestration.
Use Gemini Pro via Vertex AI managed APIs for the core orchestrator, and use a smaller language model such as an instruction-tuned Gemma open model deployed to a self-managed endpoint or Gemini Flash for the high-volume triage workload.
Deploy ShieldGemma 2 to handle agent multi-step planning and tool calling, and use Gemini Pro with maximum thinking budget for the high-volume classification tasks.
Use PaliGemma 2 via the managed Gemini API for the core orchestrator, and deploy Gemini Pro on dedicated GKE clusters for high-volume triage.
Perform full fine-tuning of Gemini Pro for the triage workload and deploy it as a self-hosted container, while using MedGemma for general agent orchestration.
Use Gemini Pro via Vertex AI managed APIs for the core orchestrator, and use a smaller language model such as an instruction-tuned Gemma open model deployed to a self-managed endpoint or Gemini Flash for the high-volume triage workload.
This architecture leverages Model Garden to implement a tiered model selection strategy (model routing) across Google Cloud AI services. It pairs the high-reasoning capabilities of Gemini Pro as a managed service for complex orchestration with a lightweight, cost-effective model like Gemma (open-weights) or Gemini Flash for high-volume, lower-complexity classification tasks.
gemma-2b-it) or a Gemini Flash model handles high-throughput categorization and entity extraction with minimal inference latency and substantially lower cost per token.Using specialized models aligned to task complexity avoids over-provisioning expensive compute for simple categorization while ensuring adequate reasoning power for multi-step agent orchestration.