professional-cloud-data-engineer
Prepare and test your skills
Prepare and test your skills
Worked example. The correct answer is already marked and every option is explained below, so there is nothing to select here. To answer questions yourself, start the free trial.
Keep the momentum going with these hand-picked practice scenarios
Want more questions like this?
Get a free certification question every week.
Last updated
A financial analytics company is migrating its daily batch data pipeline to Google Cloud. The workload has the following characteristics and constraints:
Which execution framework and architecture should the company select?
Load raw transaction files directly into BigQuery staging tables and rewrite the transformation logic as BigQuery SQL stored procedures
Execute the existing Spark jobs on ephemeral Cloud Dataproc clusters orchestrated by Cloud Composer, creating clusters on demand and deleting them immediately after job completion
Rewrite the data pipelines using the Apache Beam SDK and execute them as batch jobs on Cloud Dataflow
Provision a persistent 24/7 Cloud Dataproc cluster configured with HDFS and autoscaling to handle the nightly Spark batch jobs
Load raw transaction files directly into BigQuery staging tables and rewrite the transformation logic as BigQuery SQL stored procedures
Execute the existing Spark jobs on ephemeral Cloud Dataproc clusters orchestrated by Cloud Composer, creating clusters on demand and deleting them immediately after job completion
Cloud Dataproc is a fully managed service for executing open-source data processing engines, such as Apache Spark, Apache Hadoop, and Hive. An ephemeral cluster pattern creates dedicated compute resources specifically for the duration of a discrete workflow and terminates the cluster immediately upon job completion, persisting data externally in Cloud Storage and BigQuery.
Compared to refactoring pipelines into Apache Beam or BigQuery SQL, ephemeral Dataproc allows the team to migrate legacy Spark applications with minimal operational friction while completely eliminating the continuous cost overhead of persistent servers.
Rewrite the data pipelines using the Apache Beam SDK and execute them as batch jobs on Cloud Dataflow
Provision a persistent 24/7 Cloud Dataproc cluster configured with HDFS and autoscaling to handle the nightly Spark batch jobs