professional-cloud-data-engineer
Prepare and test your skills
Prepare and test your skills
Worked example. The correct answer is already marked and every option is explained below, so there is nothing to select here. To answer questions yourself, start the free trial.
Keep the momentum going with these hand-picked practice scenarios
Want more questions like this?
Get a free certification question every week.
Last updated
An enterprise is architecting a multi-tiered data storage and analytics platform on Google Cloud to handle hundreds of terabytes of transactional log data. The solution must satisfy the following architectural requirements:
customer_id, store_id, and product_category.Which multi-tiered storage and processing architecture should you implement?
Ingest raw logs into Cloud Storage Standard and apply Object Lifecycle Management to transition objects to Nearline, Coldline, and Archive storage over time. Ingest analytical data into a BigQuery table partitioned by day on the event timestamp and clustered by customer_id, store_id, and product_category. Create BigQuery Materialized Views for executive summary metrics.
Ingest raw logs into Cloud Storage Standard without lifecycle policies. Query the raw files directly in BigQuery using external tables, partitioning the external tables by customer_id and clustering by event timestamp. Create BigQuery standard (logical) views to serve executive dashboard aggregations.
Ingest raw logs into BigQuery directly using logical storage billing. Export older data annually to Cloud Storage Standard using manual BigQuery extract jobs. Create a non-partitioned BigQuery table clustered by customer_id, and build BigQuery BI Engine SQL views for real-time aggregation.
Ingest raw logs into Cloud Storage Archive storage directly. Query the archive files using BigLake external tables. In BigQuery, create a table partitioned on customer_id and clustered by event timestamp, and use scheduled queries to refresh a summary aggregation table every 5 minutes.
Ingest raw logs into Cloud Storage Standard and apply Object Lifecycle Management to transition objects to Nearline, Coldline, and Archive storage over time. Ingest analytical data into a BigQuery table partitioned by day on the event timestamp and clustered by customer_id, store_id, and product_category. Create BigQuery Materialized Views for executive summary metrics.
This architecture establishes an end-to-end multi-tiered storage strategy by integrating Cloud Storage for scalable, tiered data lifecycle management with BigQuery native storage optimized through date partitioning, multi-column clustering, and materialized views.
customer_id, store_id, and product_category (up to four columns) collocates related data into organized storage blocks, dramatically improving filter and aggregation performance.This architecture pairs the ideal services for each access pattern: Cloud Storage provides unbounded, tiered, cost-effective object storage, while BigQuery native tables with partitioning, clustering, and materialized views deliver maximum query performance and minimal query execution overhead.
Ingest raw logs into Cloud Storage Standard without lifecycle policies. Query the raw files directly in BigQuery using external tables, partitioning the external tables by customer_id and clustering by event timestamp. Create BigQuery standard (logical) views to serve executive dashboard aggregations.
Ingest raw logs into BigQuery directly using logical storage billing. Export older data annually to Cloud Storage Standard using manual BigQuery extract jobs. Create a non-partitioned BigQuery table clustered by customer_id, and build BigQuery BI Engine SQL views for real-time aggregation.
Ingest raw logs into Cloud Storage Archive storage directly. Query the archive files using BigLake external tables. In BigQuery, create a table partitioned on customer_id and clustered by event timestamp, and use scheduled queries to refresh a summary aggregation table every 5 minutes.