NEW
Font size
WorksheetsThe Recap Quiz !
Total questions: 25
Worksheet time: 19mins
During a class discussion, Jerin and Chrisha are talking about where data comes from in their project. Which statement best describes a data source?
Only internal databases
Any system or location where data originates
Only systems generating structured data
Only APIs
John and Chrisha are working on a data integration project. What is the primary difference between ETL and ELT?
ELT requires no transformations
ETL loads data only into cloud systems
ETL transforms before load; ELT transforms after load
ELT is only for batch processing
Jeevan and Rashima are working on a project to move company data from various sources to a central database. Which of the following is NOT a purpose of data ingestion?
Extracting data from sources
Profiling business processes
Delivering data to a target system
Scheduling and orchestrating loads
During a data profiling exercise in Jeevan and Jeshna's project, a high null percentage is detected in a mandatory column. What does this indicate?
Valid scenario
Data quality issue
Metadata enrichment
Business rule compliance
While working on a data quality project, Rashima and Jerlin are using data profiling to assess their dataset. Data profiling primarily supports which DQ dimension?
Accuracy
Completeness
Timeliness
All of the above
At a financial firm where Jerin and Jeshna work, what is a common consequence of poor data quality?
Increased automation
Regulatory fines
Reduced storage costs
Faster processing
Adil and Jerin are working on a project where they need to store and exchange data between different systems. They are considering different file formats. Which is considered a semi-structured source?
SQL table
XML file
PDF report
MP4 file
While working on a data analysis project, Jeshna and Jeevan use column profiling to help identify:
Cross-system lineage
Data precision issues
User access violations
Backup failures
While working on a data project, Rashima and Jeevan noticed that ELT became more popular due to:
Mainframes
Cloud warehouses capable of large-scale transformation
Lack of ETL tools
Less need for governance
In a school database, if Adil and Agnes are both assigned the same student ID as a primary key, which DQ dimension is violated?
Usability
Integrity
Accuracy
Lineage
Jeevan and Chrisha are starting a new data analytics project. Which is a key driver for profiling early in their project?
To speed up data load
To understand actual data behavior before writing rules
To avoid mapping
To reduce metadata
When Adil and Agnes work with data from external vendors for their project, which additional challenge might they face?
You can control their schema fully
Zero latency
No ownership
Volatile formats or unexpected changes
In a school database, Jerin and Jeshna's gender values are recorded as 'M' and 'F'. A transformation is applied to standardize all gender values to 'M' or 'F'. This improves:
Auditability
Completeness
Consistency
Security
Jerin and Agnes are building an AI system for their school project. Impact of poor data in their AI system commonly includes:
Faster model training
Model bias and poor predictions
More accurate outcomes
Unlimited scaling
Rashima and John are managing a source system with rapidly changing schemas. What does their system require?
No monitoring
Strong metadata tracking
Manual patching
No documentation
John and Chrisha are working on an ETL project for their class. Which step in ETL is responsible for handling schema alignment?
Extract
Transform
Load
Archive
During a group analytics project, Jerlin and Agnes detected data quality defects late in their analysis. This situation most likely leads to:
Faster dashboards
Expensive rework
Safer decisions
Better latency
Adil and Agnes are working on a data validation project and need to use profiling repeated patterns (regex-style). Which task are they most likely supporting?
Security classification
Pattern-based validation rule creation
Pipeline scheduling
Log backup
Rashima and Chrisha are discussing their company's database systems. Which is true about source systems?
Capture data for operational needs
Designed for analytical consumption only
Always have perfect data
Cannot be modified
While working on an ETL project, Agnes and John noticed several failures. These failures often occur because:
Too much metadata
Poor understanding of source data
Lack of cloud systems
No SQL usage
In a university database, Jerin is entering student enrollment records. A missing foreign key link in the enrollment table reflects a failure in:
Completeness
Integrity
Timeliness
Security
John and Chrisha are working on a new business analytics project. The first step they should take before designing business rules is:
Load data into dashboards
Profile data to understand behavior
Create reports
Define KPIs
Chrisha and Jeevan are working on a project where they regularly analyze data to ensure its quality. Why is profiling a continuous activity?
Data never changes
Systems remain static
Data evolves and new anomalies appear
It replaces monitoring
During a data migration project, Jeevan and Chrisha noticed that a poorly designed extract step may cause:
Loss of lineage
Transform errors
Latency increases
All of the above
In a project managed by Rashima and Jeshna, which is the biggest hidden cost of poor data?
Storage
Manual workarounds
ETL licensing
Cloud compute
