Professional Cloud DevOps Engineer
Prepare and test your skills
Prepare and test your skills
Worked example. The correct answer is already marked and every option is explained below, so there is nothing to select here. To answer questions yourself, start the free trial.
Keep the momentum going with these hand-picked practice scenarios
Want more questions like this?
Get a free certification question every week.
Last updated
A business-critical microservices application deployed across Google Kubernetes Engine (GKE) and managed with Cloud Service Mesh recently experienced significant latency degradation and an elevated rate of HTTP 5xx responses following a new release rollout. The automated deployment pipeline executed a rollback to restore the previous stable application revision.
To verify service functionality and performance metrics post-rollback, determine the incident's root cause, and establish preventive measures for future deployment strategies, what should the SRE team do?
Enable Binary Authorization breakglass overrides across all deployments, remove automated canary analysis gates from the deployment pipeline, and patch container images directly inside running pods using manual command-line overrides.
Request a quota increase for Global internal traffic director backend services in the fleet project, set all ServiceEntry resolution fields to DNS_ROUND_ROBIN, and restart all GKE cluster worker nodes simultaneously.
Purge all historical telemetry in Cloud Logging sinks to minimize storage overhead, execute a synthetic load test at 200% peak capacity against the rolled-back pods, and redeploy the failed revision immediately with disabled timeout thresholds.
Verify baseline recovery across golden signals (error rates, latency, and inter-service round-trip times) in Cloud Monitoring, analyze Cloud Service Mesh access logs and distributed traces to isolate whether errors stemmed from backend application logic or proxy routing, and conduct a blameless post-incident review to document preventive action items.
Enable Binary Authorization breakglass overrides across all deployments, remove automated canary analysis gates from the deployment pipeline, and patch container images directly inside running pods using manual command-line overrides.
Request a quota increase for Global internal traffic director backend services in the fleet project, set all ServiceEntry resolution fields to DNS_ROUND_ROBIN, and restart all GKE cluster worker nodes simultaneously.
Purge all historical telemetry in Cloud Logging sinks to minimize storage overhead, execute a synthetic load test at 200% peak capacity against the rolled-back pods, and redeploy the failed revision immediately with disabled timeout thresholds.
Verify baseline recovery across golden signals (error rates, latency, and inter-service round-trip times) in Cloud Monitoring, analyze Cloud Service Mesh access logs and distributed traces to isolate whether errors stemmed from backend application logic or proxy routing, and conduct a blameless post-incident review to document preventive action items.
This approach encompasses comprehensive post-rollback system state verification, observability-driven root cause analysis (RCA), and structured post-incident learning consistent with Site Reliability Engineering (SRE) best practices in Google Cloud.
RESPONSE_FLAGS, RESPONSE_CODE, and UPSTREAM_LOCAL_ADDRESS) and distributed traces via Cloud Trace allows engineers to differentiate between application-level failures (e.g., unhandled exceptions or database connection exhaustion) and data-plane routing issues (e.g., Envoy proxy misconfigurations or circuit breaking).Validating post-rollback stability using quantitative metrics and analyzing distributed telemetry provides the definitive evidence needed to understand failures and prevent regressions, making it the most robust operational strategy.