professional-cloud-data-engineer
Prepare and test your skills
Prepare and test your skills
Worked example. The correct answer is already marked and every option is explained below, so there is nothing to select here. To answer questions yourself, start the free trial.
An enterprise data engineering team maintains an ETL pipeline built with Cloud Data Fusion running on a managed Cloud Dataproc Spark cluster. The pipeline ingests multi-terabyte datasets, performs several complex relational joins and aggregations across multiple tables, and then executes custom feature-engineering transformations in Spark before writing outputs to downstream analytics sinks.
During peak runs, the team observes severe performance degradation and high compute costs driven by massive data shuffle operations during the distributed Spark joins.
Which strategy should the data engineering team implement to optimize pipeline performance and reduce compute costs?
Keep the momentum going with these hand-picked practice scenarios
Want more questions like this?
Get a free certification question every week.