professional-cloud-data-engineer
Prepare and test your skills
Prepare and test your skills
Worked example. The correct answer is already marked and every option is explained below, so there is nothing to select here. To answer questions yourself, start the free trial.
Keep the momentum going with these hand-picked practice scenarios
Want more questions like this?
Get a free certification question every week.
Last updated
An enterprise analytics team manages batch Spark workloads for multiple internal business units. The workloads have the following operational requirements and constraints:
Which architecture and orchestration strategy should the team implement?
Deploy a single, long-running persistent Dataproc cluster shared across all business units, enabling YARN FairScheduler and autoscaling to handle concurrent tenant jobs.
Execute all tenant Spark data transformations directly inside Cloud Composer worker pods using the PythonOperator and local PySpark libraries.
Author Cloud Composer DAGs that create an ephemeral Dataproc cluster configured with tenant-specific dependencies, execute the batch job, stage and output data to Cloud Storage, and delete the cluster upon job completion.
Maintain a single persistent Dataproc cluster and use Cloud Composer DAGs to execute dynamic initialization action scripts against the master node before every job submission.
Deploy a single, long-running persistent Dataproc cluster shared across all business units, enabling YARN FairScheduler and autoscaling to handle concurrent tenant jobs.
Execute all tenant Spark data transformations directly inside Cloud Composer worker pods using the PythonOperator and local PySpark libraries.
Author Cloud Composer DAGs that create an ephemeral Dataproc cluster configured with tenant-specific dependencies, execute the batch job, stage and output data to Cloud Storage, and delete the cluster upon job completion.
Cloud Composer acts as an orchestration engine using Apache Airflow Directed Acyclic Graphs (DAGs) to programmatically manage the lifecycle of ephemeral (job-scoped) Dataproc clusters. In this pattern, the cluster is dynamically provisioned just before job execution and deleted immediately following completion, while all persistent input and output data resides externally in Cloud Storage.
--image-version) do not conflict across different teams.DataprocCreateClusterOperator, DataprocSubmitJobOperator, DataprocDeleteClusterOperator) fully automate workflow execution.For intermittent, scheduled batch workloads with heterogeneous dependencies and multi-tenant billing requirements, ephemeral Dataproc clusters orchestrated via Cloud Composer offer the best balance of cost efficiency, fault tolerance, and security boundary isolation.
Maintain a single persistent Dataproc cluster and use Cloud Composer DAGs to execute dynamic initialization action scripts against the master node before every job submission.