WorksheetsData 50 preg (260-211) v2
Total questions: 50
Worksheet time: 2hrs 40mins
You need to ensure that both projects have the appropriate compute resources available. What should you do?
You need to recommend a solution that they can use. What should you do?
Use Dataprep to clean the data, and write the results to BigQuery. Analyze the data by using Connected Sheets.
Use Dataprep to clean the data, and write the results to BigQuery. Analyze the data by using Looker Studio.
Use Dataflow to clean the data, and write the results to BigQuery. Analyze the data by using Connected Sheets.
Use Dataflow to clean the data, and write the results to BigQuery. Analyze the data by using Looker Studio.
You have several different file type data sources, such as Apache Parquet and CSV. You want to store the data in Cloud Storage. You need to set up an object sink for your data that allows you to use your own encryption keys. You want to use a GUI-based solution. What should you do?
Use Storage Transfer Service to move files into Cloud Storage.
Use Cloud Data Fusion to move files into Cloud Storage.
Use Dataflow to move files into Cloud Storage.
Use BigQuery Data Transfer Service to move files into BigQuery.
What should you do?
Create a Cloud Storage bucket with Autoclass enabled.
Create a Cloud Storage bucket with an Object Lifecycle Management policy to transition objects from Standard to Coldline storage class if an object age reaches 30 days.
Create a Cloud Storage bucket with an Object Lifecycle Management policy to transition objects from Standard to Coldline storage class if an object is not live.
Create two Cloud Storage buckets. Use the Standard storage class for the first bucket, and use the Coldline storage class for the second bucket. Migrate objects from the first bucket to the second bucket after 30 days.
What should you do?
You have an Oracle database deployed in a VM as part of a Virtual Private Cloud (VPC) network. You want to replicate and continuously synchronize 50 tables to BigQuery. You want to minimize the need to manage infrastructure. What should you do?
Deploy Apache Kafka in the same VPC network, use Kafka Connect Oracle Change Data Capture (CDC), and Dataflow to stream the Kafka topic to BigQuery.
Create a Pub/Sub subscription to write to BigQuery directly. Deploy the Debezium Oracle connector to capture changes in the Oracle database, and sink to the Pub/Sub topic.
Deploy Apache Kafka in the same VPC network, use Kafka Connect Oracle change data capture (CDC), and the Kafka Connect Google BigQuery Sink Connector.
Create a Datastream service from Oracle to BigQuery, use a private connectivity configuration to the same VPC network, and a connection profile to BigQuery.
What should you do to improve performance?
Enable Vertical Autoscaling to let the pipeline use larger workers.
Change the pipeline code, and introduce a Reshuffle step to prevent fusion.
Update the job to increase the maximum number of workers.
Use Dataflow Prime, and enable Right Fitting to increase the worker resources.
What should you do?
Ensure that your workers have network tags to access Cloud Storage and BigQuery. Use Dataflow with only internal IP addresses.
Ensure that the firewall rules allow access to Cloud Storage and BigQuery. Use Dataflow with only internal IPs.
Create a VPC Service Controls perimeter that contains the VPC network and add Dataflow, Cloud Storage, and BigQuery as allowed services in the perimeter. Use Dataflow with only internal IP addresses.
Ensure that Private Google Access is enabled in the subnetwork. Use Dataflow with only internal IP addresses.
What should you do?
Create a normalized model with tables for each entity. Use snapshots before updates to track historical data.
Create a normalized model with tables for each entity. Keep all input files in a Cloud Storage bucket to track historical data.
Create a denormalized model with nested and repeated fields. Update the table and use snapshots to track historical data.
Create a denormalized, append-only model with nested and repeated fields. Use the ingestion timestamp to track historical data.
You have important legal hold documents in a Cloud Storage bucket. You need to ensure that these documents are not deleted or modified. What should you do?
Set a retention policy. Lock the retention policy.
Set a retention policy. Set the default storage class to Archive for long-term digital preservation.
Enable the Object Versioning feature. Add a lifecycle rule.
Enable the Object Versioning feature. Create a copy in a bucket in a different region.
What should you do?
What should you do to optimize for cost?
Create the bucket with the Autoclass storage class feature.
Create an Object Lifecycle Management policy to modify the storage class for data older than 30 days to nearline, 90 days to coldline, and 365 days to archive storage class. Delete old data as needed.
Create an Object Lifecycle Management policy to modify the storage class for data older than 30 days to coldline, 90 days to nearline, and 365 days to archive storage class. Delete old data as needed.
Create an Object Lifecycle Management policy to modify the storage class for data older than 30 days to nearline, 45 days to coldline, and 60 days to archive storage class. Delete old data as needed.
You have an inventory of VM data stored in the BigQuery table. You want to prepare the data for regular reporting in the most cost-effective way. You need to exclude VM rows with fewer than 8 vCPU in your report. What should you do?
Create a view with a filter to drop rows with fewer than 8 vCPU, and use the UNNEST operator.
Create a materialized view with a filter to drop rows with fewer than 8 vCPU, and use the WITH common table expression.
Create a view with a filter to drop rows with fewer than 8 vCPU, and use the WITH common table expression.
Use Dataflow to batch process and write the result to another BigQuery table.
What should you do?
You have one BigQuery dataset which includes customers’ street addresses. You want to retrieve all occurrences of street addresses from the dataset. What should you do?
Write a SQL query in BigQuery by using REGEXP_CONTAINS on all tables in your dataset to find rows where the word “street” appears.
Create a deep inspection job on each table in your dataset with Cloud Data Loss Prevention and create an inspection template that includes the STREET_ADDRESS infoType.
Create a discovery scan configuration on your organization with Cloud Data Loss Prevention and create an inspection template that includes the STREET_ADDRESS infoType.
Create a de-identification job in Cloud Data Loss Prevention and use the masking transformation.
You are developing a model to identify the factors that lead to sales conversions for your customers. You have completed processing your data. You want to continue through the model development lifecycle. What should you do next?
Use your model to run predictions on fresh customer input data.
Monitor your model performance, and make any adjustments needed.
Delineate what data will be used for testing and what will be used for training the model.
Test and evaluate your model on your curated data to determine how well the model performs.
What should you do?
Ask each team to create authorized views of their data. Grant the biquery.jobUser role to each team.
Create a BigQuery scheduled query to replicate all customer data into team projects.
Ask each team to publish their data in Analytics Hub. Direct the other teams to subscribe to them.
Enable each team to create materialized views of the data they need to access in their projects.
Which query should you use?
What should you do?
What should you do?
Adopt multi-regional Cloud Storage buckets in your architecture.
Adopt two regional Cloud Storage buckets, and update your application to write the output on both buckets.
Adopt a dual-region Cloud Storage bucket, and enable turbo replication in your architecture.
Adopt two regional Cloud Storage buckets, and create a daily task to copy from one bucket to the other.
You need to assign access rights to these two groups. What should you do?
What should you do?
Increase the slot capacity of the project with baseline as 0 and maximum reservation size as 3000.
Update SQL pipelines to run as a batch query, and run ad-hoc queries as interactive query jobs.
Increase the slot capacity of the project with baseline as 2000 and maximum reservation size as 3000.
Update SQL pipelines and ad-hoc queries to run as interactive query jobs.
You want to encrypt the customer data stored in BigQuery. You need to implement per-user crypto-deletion on data stored in your tables. You want to adopt native features in Google Cloud to avoid custom solutions. What should you do?
Implement Authenticated Encryption with Associated Data (AEAD) BigQuery functions while storing your data in BigQuery.
Create a customer-managed encryption key (CMEK) in Cloud KMS. Associate the key to the table while creating the table.
Create a customer-managed encryption key (CMEK) in Cloud KMS. Use the key to encrypt data before storing in BigQuery.
Encrypt your data during ingestion by using a cryptographic library supported by your ETL pipeline.
What should you do?
Use Cloud Data Fusion to design your pipeline, use the Cloud DLP plug-in to de-identify data within your pipeline, and then move the data into BigQuery.
Use the BigQuery Data Transfer Service to schedule your migration. After the data is populated in BigQuery, use the connection to the Cloud Data Loss Prevention (Cloud DLP) API to de-identify the necessary data.
Create your pipeline with Dataflow through the Apache Beam SDK for Python, customizing separate options within your code for streaming, batch processing, and Cloud DLP. Select BigQuery as your data sink.
Set up Datastream to replicate your on-premise data on BigQuery
What should you do?
What should you do?
What should you do?
What should you do?
Use slot reservations for your project to ensure that you have enough query processing capacity and are able to allocate available slots to the slower queries.
Use Cloud Monitoring to view BigQuery metrics and set up alerts that let you know when a certain percentage of slots were used.
Use available administrative resource charts to determine how slots are being used and how jobs are performing over time. Run a query on the INFORMATION_SCHEMA to review query performance.
Use Cloud Logging to determine if any users or downstream consumers are changing or deleting access grants on tagged resources.
You are on the data governance team and are implementing security requirements to deploy resources. You need to ensure that resources are limited to only the europe-west3 region. You want to follow Google-recommended practices.
Set the constraints/gcp.resourceLocations organization policy constraint to in:europe-west3- locations.
Deploy resources with Terraform and implement a variable validation rule to ensure that the region is set to the europe-west3 region for all resources.
Set the constraints/gcp.resourceLocations organization policy constraint to in:eu-locations.
Create a Cloud Function to monitor all resources created and automatically destroy the ones created outside the europe-west3 region.
What should you do? (Choose two.)
Increase the directed acyclic graph (DAG) file parsing interval.
Increase the Cloud Composer 2 environment size from medium to large.
Increase the maximum number of workers and reduce worker concurrency.
Increase the memory available to the Airflow workers.
Increase the memory available to the Airflow triggerer.
What should you do?
Use Bigtable for your large workloads, with connections to Cloud Storage to handle any HDFS use cases. Orchestrate your pipelines with Cloud Composer.
Use Dataproc to migrate Hadoop clusters to Google Cloud, and Cloud Storage to handle any HDFS use cases. Orchestrate your pipelines with Cloud Composer.
Use Dataproc to migrate Hadoop clusters to Google Cloud, and Cloud Storage to handle any HDFS use cases. Convert your ETL pipelines to Dataflow.
Use Dataproc to migrate your Hadoop clusters to Google Cloud, and Cloud Storage to handle any HDFS use cases. Use Cloud Data Fusion to visually design and deploy your ETL pipelines.
What should you do?
You have a streaming pipeline that ingests data from Pub/Sub in production. You need to update this streaming pipeline with improved business logic. You need to ensure that the updated pipeline reprocesses the previous two days of delivered Pub/Sub messages. What should you do? (Choose two.)
Use the Pub/Sub subscription clear-retry-policy flag.
Use Pub/Sub Snapshot capture two days before the deployment.
Create a new Pub/Sub subscription two days before the deployment.
Use the Pub/Sub subscription retain-acked-messages flag.
Use Pub/Sub Seek with a timestamp.
What should you do?
Create a new Memorystore for Redis instance with Standard Tier. Set capacity to 4 GB and read replica to No read replicas (high availability only). Delete the old instance.
Create a new Memorystore for Redis instance with Standard Tier. Set capacity to 5 GB and create multiple read replicas. Delete the old instance.
Create a new Memorystore for Memcached instance. Set a minimum of three nodes, and memory per node to 4 GB. Modify the Dataflow pipeline and all clients to use the Memcached instance. Delete the old instance.
Create multiple new Memorystore for Redis instances with Basic Tier (4 GB capacity). Modify the Dataflow pipeline and new clients to use all instances.
What should you do?
Add firewall rules in project A so only traffic from the VPC in project A is permitted.
Configure VPC Service Controls in the organization with a perimeter around project A.
Use Identity and Access Management conditions to ensure that only users and service accounts in project A. can access resources in project A.
Configure VPC Service Controls in the organization with a perimeter around the VPC of project A.
What should you do?
Migrate your data to Cloud Storage and migrate the metadata to Dataproc Metastore (DPMS). Refactor Spark pipelines to write and read data on Cloud Storage, and run them on Dataproc Serverless.
Migrate your data to Cloud Storage and register the bucket as a Dataplex asset. Refactor Spark pipelines to write and read data on Cloud Storage, and run them on Dataproc Serverless.
Migrate your data to BigQuery. Refactor Spark pipelines to write and read data on BigQuery, and run them on Dataproc Serverless.
Migrate your data to BigLake. Refactor Spark pipelines to write and read data on Cloud Storage, and run them on Dataproc on Compute Engine.
What is the problem and what should you do?
The advertising department is causing delays when consuming the messages. Work with the advertising department to fix this.
Messages in your Dataflow job are taking more than 30 seconds to process. Optimize your job or increase the number of workers to fix this.
Messages in your Dataflow job are processed in less than 30 seconds, but your job cannot keep up with the backlog in the Pub/Sub subscription. Optimize your job or increase the number of workers to fix this.
The web server is not pushing messages fast enough to Pub/Sub. Work with the web server team to fix this.
You are building an ELT solution in BigQuery by using Dataform. You need to perform uniqueness and null value checks on your final tables. What should you do to efficiently integrate these checks into your pipeline?
Build BigQuery user-defined functions (UDFs)
Create Dataplex data quality tasks.
Build Dataform assertions into your code.
Write a Spark-based stored procedure.
What should you do?
Provide the data science team access to Dataflow to create a pipeline to prepare and validate the raw data and load data into BigQuery for data exploration.
Create an external table in BigQuery and use SQL to transform the data as necessary. Provide the data science team access to the external tables to explore the raw data.
Load the data into BigQuery and use SQL to transform the data as necessary. Provide the data science team access to staging tables to explore the raw data.
Provide the data science team access to Dataprep to prepare, validate, and explore the data within Cloud Storage.
What should you do?
Use BigQuery Data Transfer Service to load files from Azure and AWS into BigQuery.
Create a Dataflow pipeline to ingest files from Azure and AWS to BigQuery.
Load files from AWS and Azure to Cloud Storage with Cloud Shell gsutil rsync arguments.
Use the BigQuery Omni functionality and BigLake tables to query files in Azure and AWS.
What should you do?
You orchestrate ETL pipelines by using Cloud Composer. One of the tasks in the Apache Airflow directed acyclic graph (DAG) relies on a third-party service. You want to be notified when the task does not succeed. What should you do?
Assign a function with notification logic to the on_retry_callback parameter for the operator responsible for the task at risk.
Configure a Cloud Monitoring alert on the sla_missed metric associated with the task at risk to trigger a notification.
Assign a function with notification logic to the on_failure_callback parameter tor the operator responsible for the task at risk.
Assign a function with notification logic to the sla_miss_callback parameter for the operator responsible for the task at risk.
What should you do?
Enable zonal high availability on the primary instance. Create a new read replica in a new region.
Create a cascading read replica from the existing read replica in Region3.
Create two new read replicas from the new primary instance, one in Region3 and one in a new region.
Create a new read replica in Region1, promote the new read replica to be the primary instance, and enable zonal high availability.
What should you do? (Choose two.)
Create two separate authorized datasets; one for the data analytics team and another for the consumer support team.
Ensure that the data analytics team members do not have the Data Catalog Fine-Grained Reader role for the policy tags.
Replace the authorized dataset with an authorized view. Use row-level security and apply filter_expression to limit data access.
Remove the bigquery.dataViewer role from the data analytics team on the authorized datasets.
Enforce access control in the policy tag taxonomy.
What should you do?
Set up VPC Network Peering between Project A and Project B. Add a firewall rule to allow the peered subnet range to access all instances on the network.
Turn off the external IP addresses on the Dataflow worker. Enable Cloud NAT in Project A.
Add the external IP addresses of the Dataflow worker as authorized networks in the Cloud SQL instance.
Set up VPC Network Peering between Project A and Project B. Create a Compute Engine instance without external IP address in Project B on the peered subnet to serve as a proxy server to the Cloud SQL database.
You are administering a BigQuery dataset that uses a customer-managed encryption key (CMEK). You need to share the dataset with a partner organization that does not have access to your CMEK. What should you do?
Provide the partner organization a copy of your CMEKs to decrypt the data.
Export the tables to parquet files to a Cloud Storage bucket and grant the storageinsights.viewer role on the bucket to the partner organization.
Copy the tables you need to share to a dataset without CMEKs. Create an Analytics Hub listing for this dataset.
Create an authorized view that contains the CMEK to decrypt the data when accessed.
You have a Standard Tier Memorystore for Redis instance deployed in a production environment. You need to simulate a Redis instance failover in the most accurate disaster recovery situation, and ensure that the failover has no impact on production data. What should you do?
Create a Standard Tier Memorystore for Redis instance in the development environment. Initiate a manual failover by using the limited-data-loss data protection mode.
Create a Standard Tier Memorystore for Redis instance in a development environment. Initiate a manual failover by using the force-data-loss data protection mode.
Increase one replica to Redis instance in production environment. Initiate a manual failover by using the force-data-loss data protection mode.
Initiate a manual failover by using the limited-data-loss data protection mode to the Memorystore for Redis instance in the production environment.
How should you redesign the BigQuery table to support faster access?
Cluster the table by country and username fields.
Cluster the table by country field, and partition by username field.
Partition the table by country and username fields.
Partition the table by _PARTITIONTIME.
What should you do?
Determine whether your Dataflow pipeline has a custom network tag set.
Determine whether there is a firewall rule set to allow traffic on TCP ports 12345 and 12346 for the Dataflow network tag.
Determine whether there is a firewall rule set to allow traffic on TCP ports 12345 and 12346 on the subnet used by Dataflow workers.
Determine whether your Dataflow pipeline is deployed with the external IP address option enabled.
You are using BigQuery with a multi-region dataset that includes a table with the daily sales volumes. This table is updated multiple times per day. You need to protect your sales table in case of regional failures with a recovery point objective (RPO) of less than 24 hours, while keeping costs to a minimum. What should you do?
Schedule a daily export of the table to a Cloud Storage dual or multi-region bucket.
Schedule a daily copy of the dataset to a backup region.
Schedule a daily BigQuery snapshot of the table.
Modify ETL job to load the data into both the current and another backup region.
