Intrigued by the art of cloud architecture? Discover how to design, develop, and manage robust, secure, scalable, and dynamic solutions on Google Cloud as you prepare for the Professional Cloud Architect exam!
Prepare and test your skills
Prepare and test your skills
Worked example. The correct answer is already marked and every option is explained below, so there is nothing to select here. To answer questions yourself, start the free trial.
Keep the momentum going with these hand-picked practice scenarios
Want more questions like this?
Get a free certification question every week.
Last updated
An enterprise is planning the migration and modernization of its data processing pipelines to Google Cloud. The solution must address two distinct workload requirements:
Which combination of data processing services should the cloud architect recommend?
Migrate Workload 1 to Cloud Dataflow using Apache Beam connectors; build Workload 2 on Cloud Dataproc persistent clusters with Spark Streaming.
Migrate Workload 1 to BigQuery scheduled queries; build Workload 2 using Datastream with BigQuery change data capture (CDC).
Deploy both Workload 1 and Workload 2 onto a single large persistent Compute Engine cluster configured with Apache Flink and Apache Hadoop.
Migrate Workload 1 to Cloud Dataproc using ephemeral clusters; build Workload 2 on Cloud Dataflow using the Apache Beam SDK.
Migrate Workload 1 to Cloud Dataflow using Apache Beam connectors; build Workload 2 on Cloud Dataproc persistent clusters with Spark Streaming.
Migrate Workload 1 to BigQuery scheduled queries; build Workload 2 using Datastream with BigQuery change data capture (CDC).
Deploy both Workload 1 and Workload 2 onto a single large persistent Compute Engine cluster configured with Apache Flink and Apache Hadoop.
Migrate Workload 1 to Cloud Dataproc using ephemeral clusters; build Workload 2 on Cloud Dataflow using the Apache Beam SDK.
Cloud Dataproc is a managed service for deploying and running open-source data processing engines like Apache Spark, Apache Hadoop, and Apache Flink. Cloud Dataflow is a fully managed, serverless data processing service that executes batch and streaming pipelines written with the Apache Beam SDK.
This architecture pairs the ideal execution environment to each workload's lifecycle: Dataproc eliminates refactoring costs for existing Spark/Hadoop jobs, while Dataflow delivers a modern, serverless, unified batch and stream runtime for the new analytics system.