professional-cloud-data-engineer
Prepare and test your skills
Prepare and test your skills
Worked example. The correct answer is already marked and every option is explained below, so there is nothing to select here. To answer questions yourself, start the free trial.
An enterprise data analytics team is redesigning their batch processing architecture. Every night, multi-terabyte structured datasets from operational databases and Cloud Storage are ingested, requiring complex multi-table joins and analytical aggregations before being delivered to reporting dashboards in BigQuery.
The existing pipeline runs programmatic ETL transformations entirely on Apache Spark within Cloud Dataproc clusters. However, resource-intensive shuffle operations during large-scale joins frequently cause cluster memory bottlenecks and lead to inflated compute costs.
Which architectural approach should the team implement to optimize performance and reduce compute costs?
Keep the momentum going with these hand-picked practice scenarios
Want more questions like this?
Get a free certification question every week.