Intrigued by the art of cloud architecture? Discover how to design, develop, and manage robust, secure, scalable, and dynamic solutions on Google Cloud as you prepare for the Professional Cloud Architect exam!
Prepare and test your skills
Prepare and test your skills
Worked example. The correct answer is already marked and every option is explained below, so there is nothing to select here. To answer questions yourself, start the free trial.
Keep the momentum going with these hand-picked practice scenarios
Want more questions like this?
Get a free certification question every week.
Last updated
An enterprise operates several on-premises data processing pipelines that process both historical batch logs and real-time streaming transactional data. The operations team spends significant time managing cluster nodes, provisioning compute capacity for peak hours, and handling cascading job failures.
Leadership wants to modernize this architecture to minimize operational maintenance, achieve automatic horizontal scaling for variable data volumes, and unify their batch and streaming processing models.
Which modernization and re-platforming strategy should the enterprise implement?
Rehost the existing data processing jobs directly on Compute Engine virtual machines using custom autoscaling and OS-level cron scheduling
Containerize the data processing workloads and manage them manually on self-managed Apache Spark pods inside Google Kubernetes Engine (GKE)
Offload batch records into Cloud Storage and trigger event-driven Cloud Functions to process all batch transformations and continuous streaming analytics
Refactor the data pipelines using the Apache Beam SDK and execute them as serverless jobs on Cloud Dataflow
Rehost the existing data processing jobs directly on Compute Engine virtual machines using custom autoscaling and OS-level cron scheduling
Containerize the data processing workloads and manage them manually on self-managed Apache Spark pods inside Google Kubernetes Engine (GKE)
Offload batch records into Cloud Storage and trigger event-driven Cloud Functions to process all batch transformations and continuous streaming analytics
Refactor the data pipelines using the Apache Beam SDK and execute them as serverless jobs on Cloud Dataflow
Cloud Dataflow is a fully managed, serverless execution engine for parallel data processing pipelines built on the open-source Apache Beam SDK. It enables unified stream and batch data processing with a single programming model.
Re-platforming pipelines to Dataflow replaces high-maintenance cluster management with a serverless architecture. This enables developers to focus purely on business transformation logic rather than infrastructure operations, directly satisfying the enterprise's scalability and maintenance reduction goals.