Intrigued by the art of cloud architecture? Discover how to design, develop, and manage robust, secure, scalable, and dynamic solutions on Google Cloud as you prepare for the Professional Cloud Architect exam!
Prepare and test your skills
Prepare and test your skills
Worked example. The correct answer is already marked and every option is explained below, so there is nothing to select here. To answer questions yourself, start the free trial.
Keep the momentum going with these hand-picked practice scenarios
Want more questions like this?
Get a free certification question every week.
Last updated
An e-commerce platform running on Google Cloud uses Regional Managed Instance Groups (MIGs) behind an external Application Load Balancer, with Cloud Monitoring tracking Service Level Indicators (SLIs) against a 99.99% availability Service Level Objective (SLO). The Site Reliability Engineering (SRE) team wants to conduct a chaos engineering experiment in a staging environment that mirrors production to validate that automated recovery mechanisms function properly during infrastructure disruptions without exceeding the allocated error budget.
Which strategy should the team implement to design and validate this failure injection experiment?
Simulate a live migration event on all instances and verify that Compute Engine host maintenance policies restart the virtual machines sequentially without using load balancing.
Establish a steady-state baseline using SLI metrics, inject automated VM terminations and network latency within a single zone under synthetic load, and evaluate whether Regional MIG autohealing and load balancing preserve the SLO within error budget thresholds.
Disable health checks on the Application Load Balancer backend service, terminate all virtual machines across all zones simultaneously, and measure the time required for engineers to execute manual disaster recovery runbooks.
Increase the target CPU utilization threshold on the Managed Instance Group autoscaler to 100%, disable scale-out actions, and monitor system degradation under unthrottled load.
Simulate a live migration event on all instances and verify that Compute Engine host maintenance policies restart the virtual machines sequentially without using load balancing.
Establish a steady-state baseline using SLI metrics, inject automated VM terminations and network latency within a single zone under synthetic load, and evaluate whether Regional MIG autohealing and load balancing preserve the SLO within error budget thresholds.
Chaos engineering is the discipline of experimenting on a system to build confidence in its capability to withstand turbulent conditions in production. Combining chaos engineering with Site Reliability Engineering (SRE) principles involves defining a steady-state baseline using Service Level Indicators (SLIs) and verifying that automated resilience mechanisms operate within the boundaries of the defined Service Level Objectives (SLOs) and error budgets.
This method aligns with Google Cloud reliability engineering best practices. It tests the resilience of regional compositions against zonal failure domains, verifies automated platform recovery primitives (MIG autohealing and load balancer failover), and uses empirical SRE metrics to measure overall system health.
Disable health checks on the Application Load Balancer backend service, terminate all virtual machines across all zones simultaneously, and measure the time required for engineers to execute manual disaster recovery runbooks.
Increase the target CPU utilization threshold on the Managed Instance Group autoscaler to 100%, disable scale-out actions, and monitor system degradation under unthrottled load.