Font size
WorksheetsData Science and Big Data Worksheet
Total questions: 53
Worksheet time: 28mins
Big data refers to datasets that are
Small and simple
Easily handled by RDBMS
Massive and complex
Only numerical
Data generated from social media posts is an example of
Structured data
Unstructured data
Graph data
Machine data
Which language is widely used in data science due to its rich libraries?
C
Java
Python
Assembly
Hadoop Distributed File System (HDFS) is an example of
NoSQL database
Distributed file system
Programming language
Visualization tool
Data cleansing mainly deals with
Model evaluation
Removing errors and inconsistencies
Visualization
Deployment
Which type of data arrives continuously in real time?
Structured data
Batch data
Streaming data
Graph data
NLP stands for
Neural Learning Program
Natural Language Processing
Network Learning Process
Numerical Logic Processing
Which database type is best suited for relationship-based data?
Key-value store
Document store
Graph database
Column store
MapReduce is associated with
Data visualization
Distributed processing
Data cleaning
Data security
Veracity in big data refers to
Speed of data
Trustworthiness of data
Data size
Data format
Why are traditional RDBMS systems insufficient for big data?
They are expensive
They cannot handle scale and variety
They lack security
They are open source
Volume and velocity together mainly affect
Data storage and processing
Visualization only
Programming syntax
Data security
The relationship between big data and data science is best described as
Cause and effect
Oil and refinery
Server and client
Input and output
Which activity best represents data exploration?
Training a model
Creating dashboards
Studying patterns using visualization
Automating workflows
Why is data preparation the most time-consuming step?
Data is always clean
Real-world data is messy
Models are complex
Visualization is difficult
Which example best represents machine-generated data?
Blog post
Sensor readings
PDF document
Email message
Graph-based data is useful for
Text analysis
Image processing
Relationship analysis
Numerical computation
Why is Python preferred for data science?
Low-level programming
Rich ecosystem of libraries
Faster hardware execution
Strong typing
Feature engineering is part of
Data exploration
Data transformation
Data retrieval
Presentation
Streaming systems like Kafka are used to
Store structured tables
Process real-time data
Create dashboards
Secure data
What is the main goal of data modeling?
Cleaning data
Extracting patterns and predictions
Visualizing trends
Data storage
Why are NoSQL databases used in big data?
They enforce strict schemas
They support horizontal scaling
They replace SQL entirely
They reduce data size
Which phase ensures alignment with business objectives?
Data modeling
Data retrieval
Setting research goal
Data exploration
Automation in data science ensures
Manual reporting
One-time execution
Continuous use of models
Data deletion
Visualization helps primarily in
Model training
Understanding results
Data storage
Data encryption
The primary goal of following a structured data science process is to
Increase programming speed
Maximize success and minimize cost
Avoid teamwork
Eliminate data cleaning
How many steps are there in the typical data science process?
Four
Five
Six
Seven
Which document formally defines the research goal and deliverables?
Business report
Project charter
Data dictionary
Which step focuses on understanding the what, why, and how of the project?
Data Retrieval
Data Preparation
Setting a Research Goal
Model Building
What is the immediate output of the data retrieval step?
Clean data
Modeled data
Raw data
Visualized data
Which repository stores data in its raw or natural format?
Data warehouse
Data mart
Database
Data lake
The phrase "Garbage in equals garbage out" emphasizes the importance of
Data visualization
Data modeling
Data preparation
Automation
Which step primarily uses descriptive and visual techniques?
Data Preparation
Data Exploration
Model Building
Presentation
Which plot shows median, minimum, and maximum values?
Histogram
Line graph
Boxplot
Scatter plot
Which measure indicates how well a regression model fits the data?
Accuracy
R-squared
Precision
Recall
A p-value less than 0.05 indicates that a predictor is
Insignificant
Random
Significant
Redundant
Which matrix compares predicted and actual classification results?
Correlation matrix
Confusion matrix
Distance matrix
Transition matrix
What portion of data is kept aside for testing model performance?
Training set
Validation set
Holdout sample
Raw data
The final step of the data science process focuses on
Data cleansing
Model selection
Presentation and automation
Data storage
Why is iteration important in the data science process?
To avoid documentation
To reduce computing power
To refine results based on findings
To skip steps
Why are data warehouses preferred for analysis?
They store raw data
They allow real-time updates
They contain preprocessed data
They eliminate errors
Why are joins used in data integration?
To stack observations
To enrich records
To remove duplicates
To reduce variables
Why do models using relative measures often perform better?
They reduce computation
They add business context
They remove noise
They eliminate scaling
Why is visualization emphasized in EDA?
It replaces modeling
It simplifies coding
It improves data storage
It makes patterns easier to understand
Why are simple models often preferred over complex ones?
They need more data
They are harder to explain
They often generalize better
They reduce accuracy
Why is adjusted R-squared used instead of R-squared?
It increases model error
It penalizes model complexity
It removes outliers
It improves visualization
Why is model explainability important?
For faster computation
For stakeholder trust
For data storage
For automation
Why are soft skills crucial in the final stage?
To write code
To clean data
To influence decision-making
To collect data
Why is automation required after model success?
To redesign the process
To repeat analysis efficiently
To remove errors
To store raw data
Why is the data science process not always linear?
Due to hardware issues
Because steps may need revisiting
Due to lack of tools
Because of data loss
Data Cleansing means
(a)
Why Data Science used in Modern World
(a)
Whether strong Programming knowledge is required to become data analyst? Why?
(a)
