Free GCP-PDE sample questions
Real questions from the Google Cloud Professional Data Engineer practice bank, with the correct answer and an explanation for each one. No junk, no filler.
Try them in the simulator Same questions, with study, timed and flashcard modes.
Showing 6 of 12 free sample questions.
A financial institution requires a new data warehouse solution. The security team mandates that all data stored must be encrypted using keys that the institution manages and rotates annually. Additionally, specific columns containing PII (Personally Identifiable Information) like Social Security Numbers must be restricted so that only the 'HR-Auditors' group can view the plaintext values, while data analysts see masked data. You plan to use BigQuery. Which combination of features should you configure?
Your team is migrating a legacy Hadoop workload to Google Cloud. The workload currently runs on an on-premises cluster and processes 50 TB of log data daily using Spark jobs. The processing demand is highly variable: it spikes significantly between 2 AM and 6 AM and is idle for long periods. You need to minimize costs while ensuring the jobs complete within the 4-hour window. What architecture should you design?
You are designing a database schema for Cloud Bigtable to store time-series data from environmental sensors. The query patterns involve retrieving data for a specific sensor over a specific time range. You also need to perform periodic batch analytics on all sensors in a specific geographic region. The current row key design is `[SensorID]#[Timestamp]`. What potential issue does this design introduce, and how should you fix it?
A retail company uses BigQuery for its data warehouse. They have a table named `Sales` that is 500 TB in size. Analysts frequently run queries filtering by `TransactionDate` and aggregating by `StoreId`. These queries are becoming slower and more expensive. How should you optimize the table schema to improve performance and reduce costs?
Your organization is building a machine learning pipeline to predict customer churn. The data resides in BigQuery, and your data science team wants to use SQL to build and train the model because they are less proficient in Python/TensorFlow. The model needs to be retrained weekly with fresh data. Which Google Cloud service should you use?
6 more free samples are waiting
Create a free account to unlock the whole GCP-PDE sample bank, or get full access to all 145 practice questions in the simulator.