professional-cloud-data-engineer
Prepare and test your skills
Prepare and test your skills
Worked example. The correct answer is already marked and every option is explained below, so there is nothing to select here. To answer questions yourself, start the free trial.
Keep the momentum going with these hand-picked practice scenarios
Want more questions like this?
Get a free certification question every week.
Last updated
A data platform team is managing an enterprise data lake on Google Cloud where petabyte-scale datasets in BigQuery and Cloud Storage receive continuous updates. The team needs to implement an automated data quality and monitoring framework to meet the following requirements:
Which architecture should the data platform team deploy to satisfy these requirements?
Execute BigQuery scheduled queries with full table joins to validate row counts, trigger on-demand DLP inspection templates using Cloud Functions upon object finalization in Cloud Storage, and track Dataflow worker CPU utilization in Cloud Monitoring.
Configure Dataplex auto data quality scans with an incremental scope tracking the ingestion timestamp column, enable Sensitive Data Protection discovery to export data profiles to Dataplex Universal Catalog, and configure Cloud Monitoring alert policies on Dataflow data freshness metrics.
Create Cloud Composer DAGs that execute Dataflow batch jobs with Join transforms to verify table consistency, configure Security Command Center built-in detectors for database encryption posture, and evaluate pipeline execution duration via Cloud Logging metrics.
Deploy Dataproc ephemeral Spark jobs executing custom assertion scripts against full BigQuery tables, run synchronous DLP content.deidentify API calls on every batch output, and monitor pipeline health using Cloud Logging log sinks with BigQuery destinations.
Execute BigQuery scheduled queries with full table joins to validate row counts, trigger on-demand DLP inspection templates using Cloud Functions upon object finalization in Cloud Storage, and track Dataflow worker CPU utilization in Cloud Monitoring.
Configure Dataplex auto data quality scans with an incremental scope tracking the ingestion timestamp column, enable Sensitive Data Protection discovery to export data profiles to Dataplex Universal Catalog, and configure Cloud Monitoring alert policies on Dataflow data freshness metrics.
This architecture combines Dataplex auto data quality scans, Sensitive Data Protection (Cloud DLP) discovery, and Cloud Monitoring with Dataflow metrics to establish an automated, scalable data lake monitoring framework.
DATE or TIMESTAMP column to evaluate only newly appended data records. This avoids scanning entire multi-terabyte or petabyte tables, drastically reducing compute query costs and processing time.data_freshness metric (measuring the age of the oldest unprocessed element) directly into Cloud Monitoring, enabling threshold-based alerting against defined freshness SLOs.This solution relies entirely on native, fully managed Google Cloud services designed for data governance, automated scanning, and streaming observability, avoiding custom script maintenance while ensuring minimal compute overhead.
Create Cloud Composer DAGs that execute Dataflow batch jobs with Join transforms to verify table consistency, configure Security Command Center built-in detectors for database encryption posture, and evaluate pipeline execution duration via Cloud Logging metrics.
Deploy Dataproc ephemeral Spark jobs executing custom assertion scripts against full BigQuery tables, run synchronous DLP content.deidentify API calls on every batch output, and monitor pipeline health using Cloud Logging log sinks with BigQuery destinations.