professional-cloud-data-engineer
Prepare and test your skills
Prepare and test your skills
Worked example. The correct answer is already marked and every option is explained below, so there is nothing to select here. To answer questions yourself, start the free trial.
A data engineering team runs a PySpark batch transformation pipeline on an existing Google Cloud Dataproc cluster. The job reads several terabytes of raw logs stored in Google Cloud Storage and joins this massive dataset with a 15 MB dimension table containing account metadata.
During job execution, the team observes two major performance bottlenecks:
How should the team optimize this pipeline for a specific job submission without altering cluster-wide settings?
Keep the momentum going with these hand-picked practice scenarios
Want more questions like this?
Get a free certification question every week.