professional-cloud-data-engineer
Prepare and test your skills
Prepare and test your skills
Worked example. The correct answer is already marked and every option is explained below, so there is nothing to select here. To answer questions yourself, start the free trial.
Keep the momentum going with these hand-picked practice scenarios
Want more questions like this?
Get a free certification question every week.
Last updated
An enterprise data platform spans multiple Google Cloud projects containing BigQuery tables and multi-terabyte Cloud Storage buckets containing semi-structured JSON and Parquet datasets.
Your data governance team requires a centralized metadata architecture that fulfills the following criteria:
Which architecture should you implement?
Build a Cloud Data Fusion batch pipeline on an hourly schedule to ingest Cloud Storage metadata, load technical schemas into Dataproc Metastore, and apply Google Cloud resource labels to underlying storage buckets.
Organize Cloud Storage assets into Dataplex zones with Discovery enabled to register metadata in BigQuery publishing datasets. Create schematized Tag Templates in Dataplex/Data Catalog, using private tag templates for sensitive compliance fields and public tag templates for general metadata, and control discovery through IAM roles.
Store metadata in JSON Lines sidecar files alongside raw data in Cloud Storage, ingest these files into a Vertex AI Search data store, and query the data store using natural language document search.
Deploy Cloud Functions triggered by Cloud Storage object finalization to parse file headers, execute BigQuery DDL to generate external tables, and write metadata records into a centralized BigQuery governance dataset secured with authorized views.
Build a Cloud Data Fusion batch pipeline on an hourly schedule to ingest Cloud Storage metadata, load technical schemas into Dataproc Metastore, and apply Google Cloud resource labels to underlying storage buckets.
Organize Cloud Storage assets into Dataplex zones with Discovery enabled to register metadata in BigQuery publishing datasets. Create schematized Tag Templates in Dataplex/Data Catalog, using private tag templates for sensitive compliance fields and public tag templates for general metadata, and control discovery through IAM roles.
Dataplex Universal Catalog and Data Catalog provide an automated, policy-driven metadata management and discovery service. Dataplex manages distributed data lakes and data zones, automatically extracting technical metadata through Discovery jobs, while Data Catalog provides schematized Tag Templates to capture business and operational metadata across resources like BigQuery, Cloud Storage, and Pub/Sub.
bigquery.tables.get or roles/bigquery.metadataViewer). A user cannot view search results for assets they lack permissions to access. Furthermore, configuring private tag templates restricts the visibility and searchability of sensitive compliance metadata exclusively to principals granted explicit IAM permissions on the template.This architecture leverages fully managed Google Cloud governance services without requiring custom ETL pipelines, third-party metastores, or bespoke metadata synchronization scripts.
Store metadata in JSON Lines sidecar files alongside raw data in Cloud Storage, ingest these files into a Vertex AI Search data store, and query the data store using natural language document search.
Deploy Cloud Functions triggered by Cloud Storage object finalization to parse file headers, execute BigQuery DDL to generate external tables, and write metadata records into a centralized BigQuery governance dataset secured with authorized views.