professional-cloud-data-engineer
Prepare and test your skills
Prepare and test your skills
Worked example. The correct answer is already marked and every option is explained below, so there is nothing to select here. To answer questions yourself, start the free trial.
A data engineering team maintains a nightly batch Apache Beam pipeline running on Cloud Dataflow that aggregates multi-terabyte transactional logs. The pipeline groups data by merchant ID using a GroupByKey transform and writes aggregated metrics into BigQuery.
Monitoring shows the following bottlenecks during job execution:
GroupByKey transform suffers from severe stragglers due to extreme key skew caused by a few high-volume merchants, leaving most worker vCPUs idle while a single worker runs out of memory.Which combination of optimization strategies should the team implement to resolve the performance bottlenecks and minimize execution costs?
Keep the momentum going with these hand-picked practice scenarios
Want more questions like this?
Get a free certification question every week.