WorksheetsMCQsPDE Personal 1st 50
Total questions: 50
Worksheet time: 25mins
Your company built a TensorFlow neutral-network model with a large number of neurons and layers. The model fits well for the training data. However, when tested against new data, it performs poorly. What method can you employ to address this?
Threading
Serialization
Dropout Methods
Dimensionality Reduction
You are building a model to make clothing recommendations. You know a user's fashion preference is likely to change over time, so you build a data pipeline to stream new data back to the model as it becomes available. How should you use this data to train the model?
Continuously retrain the model on just the new data.
Continuously retrain the model on a combination of existing data and the new data.
Train on the existing data while using the new data as your test set.
Train on the new data while using the existing data as your test set.
You designed a database for patient records as a pilot project to cover a few hundred patients in three clinics. Your design used a single database table to represent all patients and their visits, and you used self-joins to generate reports. The server resource utilization was at 50%. Since then, the scope of the project has expanded. The database must now store 100 times more patient records. You can no longer run the reports, because they either take too long or they encounter errors with insufficient compute resources. How should you adjust the database design?
Add capacity (memory and disk space) to the database server by the order of 200.
Shard the tables into smaller ones based on date ranges, and only generate reports with prespecified date ranges.
Normalize the master patient-record table into the patient table and the visits table, and create other necessary tables to avoid self-join.
Partition the table into smaller tables, with one for each clinic. Run queries against the smaller table pairs, and use unions for consolidated reports.
You create an important report for your large team in Google Data Studio 360. The report uses Google BigQuery as its data source. You notice that visualizations are not showing data that is less than 1 hour old. What should you do?
Disable caching by editing the report settings.
Disable caching in BigQuery by editing table details.
Refresh your browser tab showing the visualizations.
Clear your browser history for the past hour then reload the tab showing the virtualizations.
An external customer provides you with a daily dump of data from their database. The data flows into Google Cloud Storage GCS as comma-separated values (CSV) files. You want to analyze this data in Google BigQuery, but the data could have rows that are formatted incorrectly or corrupted. How should you build this pipeline?
Use federated data sources, and check data in the SQL query.
Enable BigQuery monitoring in Google Stackdriver and create an alert.
Import the data into BigQuery using the gcloud CLI and set max_bad_records to 0.
Run a Google Cloud Dataflow batch pipeline to import the data into BigQuery, and push errors to another dead-letter table for analysis.
Your weather app queries a database every 15 minutes to get the current temperature. The frontend is powered by Google App Engine and serve millions of users. How should you design the frontend to respond to a database failure?
Issue a command to restart the database servers.
Retry the query with exponential backoff, up to a cap of 15 minutes.
Retry the query every second until it comes back online to minimize staleness of data.
Reduce the query frequency to once every hour until the database comes back online.
You are creating a model to predict housing prices. Due to budget constraints, you must run it on a single resource-constrained virtual machine. Which learning algorithm should you use?
Linear regression
Logistic classification
Recurrent neural network
Feedforward neural network
You are building new real-time data warehouse for your company and will use Google BigQuery streaming inserts. There is no guarantee that data will only be sent in once but you do have a unique ID for each row of data and an event timestamp. You want to ensure that duplicates are not included while interactively querying data. Which query type should you use?
Include ORDER BY DESK on timestamp column and LIMIT to 1.
Use GROUP BY on the unique ID column and timestamp column and SUM on the values.
Use the LAG window function with PARTITION by unique ID along with WHERE LAG IS NOT NULL.
Use the ROW_NUMBER window function with PARTITION by unique ID along with WHERE row equals 1.
Your company is using WILDCARD tables to query data across multiple tables with similar names. The SQL statement is currently failing with the following error:
#Syntax error: Expected end of statement but got "-" at [4:11]
SELECT age
FROM bigquery-public-data.noaa_gsod.gsod
WHERE age != 99
AND_TABLE_SUFFIX = '1929'
ORDER BY age DESC
Which table name will make the SQL statement work correctly?
'bigquery-public-data.noaa_gsod.gsod'
bigquery-public-data.noaa_gsod.gsod*
'bigquery-public-data.noaa_gsod.gsod'*
'bigquery-public-data.noaa_gsod.gsod*`
Your company is in a highly regulated industry. One of your requirements is to ensure individual users have access only to the minimum amount of information required to do their jobs. You want to enforce this requirement with Google BigQuery. Which three approaches can you take? (Choose three.)
Use Google Stackdriver Audit Logging to determine policy violations.
Restrict access to tables by role.
Ensure that the data is encrypted at all times.
Restrict BigQuery API access to approved users.
Segregate data across multiple tables or databases.
You are working on optimizing BigQuery for a query that is run repeatedly on a single table. The data queried is about 1 GB, and some rows are expected to change about 10 times every hour. You have optimized the SQL statements as much as possible. You want to further optimize the query's performance. What should you do?
Create a materialized view based on the table, and query that view.
Enable caching of the queried data so that subsequent queries are faster.
Create a scheduled query, and run it a few minutes before the report has to be created.
Reserve a larger number of slots in advance so that you have maximum compute power to execute the query.
Several years ago, you built a machine learning model for an ecommerce company. Your model made good predictions. Then a global pandemic occurred, lockdowns were imposed, and many people started working from home. Now the quality of your model has degraded. You want to improve the quality of your model and prevent future performance degradation. What should you do?
Retrain the model with data from the first 30 days of the lockdown.
Monitor data until usage patterns normalize, and then retrain the model.
Retrain the model with data from the last 30 days. After one year, return to the older model.
Retrain the model with data from the last 30 days. Add a step to continuously monitor model input data for changes, and retrain the model.
A new member of your development team works remotely. The developer will write code locally on their laptop, which will connect to a MySQL instance on Cloud SQL. The instance has an external (public) IP address. You want to follow Google-recommended practices when you give access to Cloud SQL to the new team member. What should you do?
Ask the developer for their laptop's IP address, and add it to the authorized networks list.
Remove the external IP address, and replace it with an internal IP address. Add only the IP address for the remote developer's laptop to the authorized list.
Give instance access permissions in Identity and Access Management (IAM), and have the developer run Cloud SQL Auth proxy to connect to a MySQL instance.
Give instance access permissions in Identity and Access Management (IAM), change the access to "private service access" for security, and allow the developer to access Cloud SQL from their laptop.
Your Cloud Spanner database stores customer address information that is frequently accessed by the marketing team. When a customer enters the country and the state where they live, this information is stored in different tables connected by a foreign key. The current architecture has performance issues. You want to follow Google-recommended practices to improve performance. What should you do?
Create interleaved tables, and store states under the countries.
Denormalize the data, and have a row for each state with its corresponding country.
Retain the existing architecture, but use short, two-letter codes for the countries and states.
Combine the countries in a single cell's text, for example "country:state1,state2, …" and when required, split the data.
Your company runs its business-critical system on PostgreSQL. The system is accessed simultaneously from many locations around the world and supports millions of customers. Your database administration team manages the redundancy and scaling manually. You want to migrate the database to Google Cloud. You need a solution that will provide global scale and availability and require minimal maintenance. What should you do?
Migrate to BigQuery.
Migrate to Cloud Spanner.
Migrate to a Cloud SQL for PostgreSQL instance.
Migrate to bare metal machines with PostgreSQL installed.
Your company collects data about customers to regularly check their health vitals. You have millions of customers around the world. Data is ingested at an average rate of two events per 10 seconds per user. You need to be able to visualize data in Bigtable on a per user basis. You need to construct the Bigtable key so that the operations are performant. What should you do?
Construct the key as user-id#device-id#activity-id#timestamp.
Construct the key as timestamp#user-id#device-id#activity-id.
Construct the key as timestamp#device-id#activity-id#user-id.
Construct the key as user-id#timestamp#device-id#activity-id.
Your company is hiring several business analysts who are new to BigQuery. The analysts will use BigQuery to analyze large quantities of data. You need to control costs in BigQuery and ensure that there is no budget overrun while you maintain the quality of query results. What should you do?
Set a customized project-level or user-level daily quota to acceptable values.
Reduce the data in the BigQuery table so that the analysts query less data, and then archive the remaining data.
Train the analysts to use the query validator or --dry_run to estimate costs so that the analysts can self-regulate usage.
Export the BigQuery daily costs, and visualize the data on Looker on a per-analyst basis so that the analysts can self-regulate usage.
Your Bigtable database was recently deployed into production. The scale of data ingested and analyzed has increased significantly, but the performance has degraded. You want to identify the performance issue. What should you do?
Use Key Visualizer to analyze performance
Use Cloud Trace to identify the performance issue.
Add logging statements into the code to see which inserts cause the delay.
Add more nodes to the cluster to see if that resolves the performance issue.
Your company is moving your data analytics to BigQuery. Your other operations will remain on-premises. You need to transfer 800 TB of historic data. You also need to plan for 30 Gbps of daily data transfers that must be appended for analysis the next day. You want to follow Google-recommended practices to transfer your data. What should you do?
As early as possible every day, use Cloud VPN to transfer the existing data over the internet.
Use a Transfer Appliance to move the existing data to Google Cloud. Use Cloud VPN to transfer data daily.
Use a Transfer Appliance to move the existing data to Google Cloud.. Use VPC Network Peering to transfer data daily.
Use a Transfer Appliance to move the existing data to Google Cloud. Set up a Dedicated or Partner Interconnect for daily transfers.
Your team runs Dataproc workloads where the worker node takes about 45 minutes to process. You have been exploring various options to optimize the system for cost, including shutting down worker nodes aggressively. However, in your metrics you see that the entire job takes even longer. You want to optimize the system for cost without increasing job completion time. What should you do?
Set a graceful decommissioning timeout greater than 45 minutes.
Rewrite the processing in Cloud Data Fusion, and run the job automatically.
Rewrite the processing in Dataflow, and use stream processing of the same data.
Increase the number of vCPUs on each worker node so that the processing finishes sooner.
Your customer has a SQL Server database that contains about 5 TB of data in another public cloud. You expect the data to grow to a maximum of 25 TB. The database is the backend of an internal reporting application that is used once a week. You want to migrate the application to Google Cloud to reduce administrative effort while keeping costs the same or reducing them. What should you do?
Migrate the database to Bigtable.
Migrate the database to Cloud Spanner.
Install SQL Server on a Compute Engine VM.
Migrate the database to SQL Server in Cloud SQL.
Your IT team uses BigQuery for storing structured data. Your finance team recently moved to Google Workspace Enterprise edition from a standalone, desktop-based spreadsheet processor. When the finance team needs data insights, the IT team runs a query on BigQuery, exports the data to a CSV file, and sends the file as an email attachment to the finance team members. You want to improve the process while you retain familiar methods of data analysis for the finance team. What should you do?
Run the query in BigQuery, and give the finance team access to the results view, which can be analyzed.
Run the query in BigQuery, and give the finance team access to the data visualizations in Google Data Studio.
Run the query in BigQuery, export the data to CSV, upload the file to a Cloud Storage bucket, and share the file with the finance team.
Run the query in BigQuery, and save the results to a Google Sheets shared spreadsheet that can be accessed and analyzed by the finance team.
Your scooter-sharing company collects information about their scooters, such as location, battery level, and speed. The company visualizes this data in real time. To guard against intermittent connectivity, each scooter sends repeats of certain messages within a short interval. Occasional data errors have been noticed. The messages are received in Pub/Sub and stored in BigQuery. You need to ensure that the data does not contain duplicates and that erroneous data with empty fields is rejected. What should you do?
Store the data in BigQuery, and run delete queries on erroneous and duplicate data.
Use Dataflow to subscribe to Pub/Sub, process the data, and store the data in BigQuery.
Use Kubernetes to create a microservices application that can remove duplicates and erroneous data. Then insert the data into BigQuery.
Create an application on Compute Engine with Managed Instance Groups that can remove duplicates and erroneous data. Then insert the data into BigQuery.
Your cryptocurrency trading company visualizes prices to help your customers make trading decisions. Because different trades happen in real time, the price data is fed to a data pipeline that uses Dataflow for processing. You want to compute moving averages. What should you do?
Use hopping windows in Dataflow.
Use session windows in Dataflow.
Use tumbling windows in Dataflow.
Use Dataflow SQL, and compute averages grouped by time.
You are building the trading platform for a stock exchange with millions of traders. Trading data is written rapidly. You need to retrieve data quickly to show visualizations to the traders, such as the changing price of a particular stock over time. You need to choose a storage solution in Google Cloud. What should you do?
Use Bigtable.
Use Firestore.
Use Cloud SQL.
Use Memorystore.
Your customer uses Hadoop and Spark to run data analytics on-premises. The main data is stored in hard disks that are centrally accessed. Your customer needs to migrate their workloads to Google Cloud efficiently while considering scalability. You want to select an architecture that requires minimal effort. What should you do?
Use Dataproc to run Hadoop and Spark jobs. Move the data to Cloud Storage.
Use Dataflow to recreate the jobs in a serverless approach. Move the data to Cloud Storage.
Use Dataproc to run Hadoop and Spark jobs. Retain the data on a Compute Engine VM with an attached persistent disk.
Use Dataflow to recreate the jobs in a serverless approach. Retain the data on a Compute Engine VM with an attached persistent disk.
You used a small amount of data to build a machine learning model that gives you good inferences during testing. However, the results show more errors when real-world data is used to run the model. No additional data can be collected for testing. You want to get a more accurate view of the model's capability. What should you do?
Reduce the amount of data to improve the model.
Cross-validate the data, and re-run the model building process.
Create feature crosses that will add new columns to increase the data.
Duplicate the data twice to increase the data, and re-run the model building process.
Your organization has been collecting information for many years about your customers, including their address and credit card details. You plan to use this customer data to build machine learning models on Google Cloud. You are concerned about private data leaking into the machine learning model. Your management is also concerned that direct leaks of personal data could damage the company's reputation. You need to address these concerns about data security. What should you do?
Remove all the tables that contain sensitive data.
Use libraries like SciPy to build the ML models on your local computer.
Remove the sensitive data by using the Cloud Data Loss Prevention (DLP) API.
Identify the rows that contain sensitive data, and apply SQL queries to remove only those rows.
Your healthcare application has a backend system that accepts event data directly from IoT devices. Recent increases of the application's users and devices are causing a sudden influx of data that overwhelms the system. You need to redesign the data pipeline to ensure that all data is processed and that no events are lost. You want to follow Google-recommended practices. What should you do?
Use Kafka with pull mode.
Use Pub/Sub with pull mode.
Use Pub/Sub with push mode.
Run Cloud Scheduler at fixed intervals.
You have 250,000 devices which produce a JSON device status event every 10 seconds. You want to capture this event data for outlier time series analysis. What should you do?
Ship the data into BigQuery. Develop a custom application that uses the BigQuery API to query the dataset and displays device outlier data based on your business requirements.
Ship the data into BigQuery. Use the BigQuery console to query the dataset and display device outlier data based on your business requirements.
Ship the data into Cloud Bigtable. Use the Cloud Bigtable cbt tool to display device outlier data based on your business requirements.
Ship the data into Cloud Bigtable. Install and use the HBase shell for Cloud Bigtable to query the table for device outlier data based on your business requirements.
You are designing storage for CSV files and using an I/O-intensive custom Apache Spark transform as part of deploying a data pipeline on Google Cloud. You intend to use ANSI SQL to run queries for your analysts.
How should you transform the input data?
Use BigQuery for storage. Use Dataflow to run the transformations.
Use BigQuery for storage. Use Dataproc to run the transformations.
Use Cloud Storage for storage. Use Dataflow to run the transformations.
Use Cloud Storage for storage. Use Dataproc to run the transformations.
You are designing storage for CSV files and using an I/O-intensive custom Apache Spark transform as part of deploying a data pipeline on Google Cloud. You intend to use ANSI SQL to run queries for your analysts.
How should you transform the input data?
Use BigQuery for storage. Use Dataflow to run the transformations.
Use BigQuery for storage. Use Dataproc to run the transformations.
Use Cloud Storage for storage. Use Dataflow to run the transformations.
Use Cloud Storage for storage. Use Dataproc to run the transformations.
Your company is loading comma-separated values (CSV) files into BigQuery. The data is fully imported successfully; however, the imported data is not matching byte-to-byte to the source file.
What is the most likely cause of this problem?
The CSV data loaded in BigQuery is not flagged as CSV.
The CSV data had invalid rows that were skipped on import.
The CSV data has not gone through an ETL phase before loading into BigQuery.
The CSV data loaded in BigQuery is not using BigQuery’s default encoding.
You are using Pub/Sub to stream inventory updates from many point-of-sale (POS) terminals into BigQuery.
Each update event has the following information: product identifier "prodSku", change increment "quantityDelta", POS identification "termId", and "messageId" which is created for each push attempt from the terminal.
During a network outage, you discovered that duplicated messages were sent, causing the inventory system to over-count the changes. You determine that the terminal application has design problems and may send the same event more than once during push retries.
You want to ensure that the inventory update is accurate. What should you do?
Add another attribute orderId to the message payload to mark the unique check-out order across all terminals. Make sure that messages whose "orderId" and "prodSku" values match corresponding rows in the BigQuery table are discarded.
Inspect the "messageId" of each message. Make sure that any messages whose "messageId" values match corresponding rows in the BigQuery table are discarded.
Instead of specifying a change increment for "quantityDelta", always use the derived inventory value after the increment has been applied. Name the new attribute "adjustedQuantity".
Inspect the "publishTime" of each message. Make sure that messages whose "publishTime" values match rows in the BigQuery table are discarded.
You are building storage for files for a data pipeline on Google Cloud. You want to support JSON files. The schema of these files will occasionally change.
Your analyst teams will use running aggregate ANSI SQL queries on this data. What should you do?
Use BigQuery for storage. Provide format files for data load. Update the format files as needed.
Use BigQuery for storage. Select "Automatically detect" in the Schema section.
Use Cloud Storage for storage. Link data as temporary tables in BigQuery and turn on the "Automatically detect" option in the Schema section of BigQuery.
Use Cloud Storage for storage. Link data as permanent tables in BigQuery and turn on the "Automatically detect" option in the Schema section of BigQuery.
You need to stream time-series data in Avro format, and then write this to both BigQuery and Cloud Bigtable simultaneously using Dataflow. You want to achieve minimal end-to-end latency.
Your business requirements state this needs to be completed as quickly as possible. What should you do?
Create a pipeline and use ParDo transform.
Create a pipeline that groups the data into a PCollection and uses the Combine transform.
Create a pipeline that groups data using a PCollection, and then use Avro I/O transform to write to Cloud Storage. After the data is written, load the data from Cloud Storage into BigQuery and Bigtable.
Create a pipeline that groups data using a PCollection and then uses Bigtable and BigQueryIO transforms.
You are working on a project with two compliance requirements. The first requirement states that your developers should be able to see the Google Cloud billing charges for only their own projects.
The second requirement states that your finance team members can set budgets and view the current charges for all projects in the organization.
The finance team should not be able to view the project contents. You want to set permissions. What should you do?
Add the finance team members to the Billing Administrator role for each of the billing accounts that they need to manage. Add the developers to the Viewer role for the Project.
Add the finance team members to the default IAM Owner role. Add the developers to a custom role that allows them to see their own spend only.
Add the developers and finance managers to the Viewer role for the Project.
Add the finance team to the Viewer role for the Project. Add the developers to the Security Reviewer role for each of the billing accounts.
You want to publish system metrics to Google Cloud from a large number of on-prem hypervisors and VMs for analysis and creation of dashboards.
You have an existing custom monitoring agent deployed to all the hypervisors and your on-prem metrics system is unable to handle the load. You want to design a system that can collect and store metrics at scale. You don't want to manage your own time series database.
Metrics from all agents should be written to the same table but agents must not have permission to modify or read data written by other agents. What should you do?
Modify the monitoring agent to write protobuf messages directly to BigTable.
Modify the monitoring agent to publish protobuf messages to Pub/Sub. Use a Dataproc cluster or Dataflow job to consume messages from Pub/Sub and write to BigTable.
Modify the monitoring agent to write protobuf messages to HBase deployed on Compute Engine VM Instances
Modify the monitoring agent to write protobuf messages to Pub/Sub. Use a Dataproc cluster or Dataflow job to consume messages from Pub/Sub and write to Cassandra deployed on Compute Engine VM Instances.
Your company is streaming real-time sensor data from their factory floor into Bigtable and they have noticed extremely poor performance.
How should the row key be redesigned to improve Bigtable performance on queries that populate real-time dashboards?
Use a row key of the form <timestamp>.
Use a row key of the form <sensorid>.
Use a row key of the form <timestamp>#<sensorid>.
Use a row key of the form <sensorid>#<timestamp>.
You are designing a relational data repository on Google Cloud to grow as needed. The data will be transactionally consistent and added from any location in the world.
You want to monitor and adjust node count for input traffic, which can spike unpredictably. What should you do?
Use Cloud Spanner for storage. Monitor CPU utilization and increase node count if more than 70% utilized for your time span.
Use Cloud Spanner for storage. Monitor storage usage and increase node count if more than 70% utilized.
Use Cloud Bigtable for storage. Monitor data stored and increase node count if more than 70% utilized.
Use Cloud Bigtable for storage. Monitor CPU utilization and increase node count if more than 70% utilized for your time span.
A company is migrating its current infrastructure from on-premise to Google cloud. It stores over 280TB of data on its on-premise HDFS servers. You were tasked to move data from HDFS to Google Storage in a secure and efficient manner. Which of the following approaches are best to fulfill this task?
Install Google Storage gsutil tool on servers and copy the data from HDFS to Google Storage.
Use Cloud Data Transfer Service to migrate the data to Google Storage.
Import the data from HDFS to BigQuery. Then, export the data to Google Storage in AVRO format.
Use Transfer Appliance Service to migrate the data to Google Storage.
You have a Dataflow pipeline to run and process a set of data files received from a client, for transformation and loading into a data warehouse. This pipeline should run each morning so that metrics can be ready when stakeholders need the latest stats based on data sent the day before. Which tool should you use?
Cloud Functions
Compute Engine
Kubernetes Engine
Cloud Scheduler
Your company is migrating their 30-node Apache Hadoop cluster to the cloud. They want to re-use Hadoop jobs they have already created and minimize the management of the cluster as much as possible. They also want to be able to persist data beyond the life of the cluster. What should you do?
Create a Google Cloud Dataflow job to process the data.
Create a Google Cloud Dataproc cluster that uses persistent disks for HDFS.
Create a Hadoop cluster on Google Compute Engine that uses persistent disks.
Create a Cloud Dataproc cluster that uses the Google Cloud Storage connector.
You work for a bank. You have a labelled dataset that contains information on already granted loan application and whether these applications have been defaulted. You have been asked to train a model to predict default rates for credit applicants.
What should you do?
Increase the size of the dataset by collecting additional data.
Train a linear regression to predict a credit default risk score
Remove the bias from the data and collect applications that have been declined loans.
Match loan applicants with their social profiles to enable feature engineering.
You have an Apache Kafka cluster on-prem with topics containing web application logs. You need to replicate the data to Google Cloud for analysis in BigQuery and Cloud Storage. The preferred replication method is mirroring to avoid deployment of Kafka Connect plugins.
What should you do?
Deploy a Kafka cluster on GCE VM Instances. Configure your on-prem cluster to mirror your topics to the cluster running in GCE. Use a Dataproc cluster or Dataflow job to read from Kafka and write to GCS.
Deploy a Kafka cluster on GCE VM Instances with the Pub/Sub Kafka connector configured as a Sink connector. Use a Dataproc cluster or Dataflow job to read from Kafka and write to GCS.
Deploy the Pub/Sub Kafka connector to your on-prem Kafka cluster and configure Pub/Sub as a Source connector. Use a Dataflow job to read from Pub/Sub and write to GCS.
Deploy the Pub/Sub Kafka connector to your on-prem Kafka cluster and configure Pub/Sub as a Sink connector. Use a Dataflow job to read from Pub/Sub and write to GCS.
Your company maintains a hybrid deployment with GCP, where analytics are performed on your anonymized customer data. The data are imported to Cloud Storage from your data center through parallel uploads to a data transfer server running on GCP.
Management informs you that the daily transfers take too long and have asked you to fix the problem. You want to maximize transfer speeds. Which action should you take?
Increase the CPU size on your server.
Increase the size of the Google Persistent Disk on your server.
Increase your network bandwidth from your datacenter to GCP.
Increase your network bandwidth from Compute Engine to Cloud Storage.
You are responsible for writing your company's ETL pipelines to run on an Apache Hadoop cluster. The pipeline will require some checkpointing and splitting pipelines. Which method should you use to write the pipelines?
PigLatin using Pig
HiveQL using Hive
Java using MapReduce
Python using MapReduce
You are designing storage for two relational tables that are part of a 10-TB database on Google Cloud. You want to support transactions that scale horizontally.
You also want to optimize data for range queries on non-key columns. What should you do?
Use Cloud SQL for storage. Add secondary indexes to support query patterns.
Use Cloud SQL for storage. Use Cloud Dataflow to transform data to support query patterns.
Use Cloud Spanner for storage. Add secondary indexes to support query patterns.
Use Cloud Spanner for storage. Use Cloud Dataflow to transform data to support query patterns.
Your company is selecting a system to centralize data ingestion and delivery. You are considering messaging and data integration systems to address the requirements. The key requirements are:
✑ The ability to seek to a particular offset in a topic, possibly back to the start of all data ever captured
✑ Support for publish/subscribe semantics on hundreds of topics
✑ Retain per-key ordering
Which system should you choose?
Apache Kafka
Cloud Storage
Cloud Pub/Sub
Firebase Cloud Messaging
You plan to deploy Cloud SQL using MySQL. You need to ensure high availability in the event of a zone failure. What should you do?
Create a Cloud SQL instance in one zone, and create a failover replica in another zone within the same region.
Create a Cloud SQL instance in one zone, and create a read replica in another zone within the same region.
Create a Cloud SQL instance in one zone, and configure an external read replica in a zone in a different region.
Create a Cloud SQL instance in a region, and configure automatic backup to a Cloud Storage bucket in the same region.
