professional-cloud-data-engineer
Prepare and test your skills
Prepare and test your skills
Worked example. The correct answer is already marked and every option is explained below, so there is nothing to select here. To answer questions yourself, start the free trial.
Keep the momentum going with these hand-picked practice scenarios
Want more questions like this?
Get a free certification question every week.
Last updated
An enterprise data analytics platform ingests high-frequency transaction records through a distributed data processing pipeline before storing them in Google Cloud databases. During recent production runs, upstream schema mutations and data source drift caused unexpected null distributions and value outliers, degrading downstream analytical fidelity.
You need to design an automated validation and monitoring architecture that accomplishes the following:
Which architecture should you implement?
Enable Spark dynamic allocation on Dataproc with dynamic resource scaling up to 100% scaleUpFactor to automatically drop pending task containers whenever data validation thresholds fail in YARN.
Deploy the AlloyDB Omni Kubernetes Operator to poll DBCluster status YAML files for critical incident error codes, and execute heap dumps to automatically isolate corrupted data rows.
Configure Cloud SQL to automatically scale up compute size and memory when lock contention spikes, and use the database error log to trigger near-zero downtime scale-down events once outlier records are committed.
Instrument the pipeline to log structured validation errors to Cloud Logging, create log-based custom metrics in Cloud Monitoring with anomaly detection alerting policies against baseline observation thresholds, and route alert notifications via Pub/Sub to a Cloud Run remediation service that isolates drifting records to a dead-letter quarantine table.
Enable Spark dynamic allocation on Dataproc with dynamic resource scaling up to 100% scaleUpFactor to automatically drop pending task containers whenever data validation thresholds fail in YARN.
Deploy the AlloyDB Omni Kubernetes Operator to poll DBCluster status YAML files for critical incident error codes, and execute heap dumps to automatically isolate corrupted data rows.
Configure Cloud SQL to automatically scale up compute size and memory when lock contention spikes, and use the database error log to trigger near-zero downtime scale-down events once outlier records are committed.
Instrument the pipeline to log structured validation errors to Cloud Logging, create log-based custom metrics in Cloud Monitoring with anomaly detection alerting policies against baseline observation thresholds, and route alert notifications via Pub/Sub to a Cloud Run remediation service that isolates drifting records to a dead-letter quarantine table.
This architecture establishes an end-to-end automated observability and data fidelity framework using Cloud Logging, Cloud Monitoring, Pub/Sub, and Cloud Run to detect, alert on, and remediate data drift and anomalies.