wayground logo

Free Printable Worksheets

NEW

Font size

S
M
L
XL
Worksheets

Chapter 1 – Data Science Worksheet (Grade 12)

Total questions: 79

Worksheet time: 40mins

Name
Class
Date
1.

Who first proposed the term Data Science?

a)

William S. Cleveland

b)

Chien-Fu Jeff Wu

c)

DJ Patil

d)

John Tukey

2.

What was William S. Cleveland’s key contribution to Data Science?

a)

Popularized the concept of Big Data

b)

Proposed Data Science as an independent discipline

c)

Developed the R programming language

d)

Founded Kaggle

3.

The growth of Data Science in the 21st century is mainly due to which factor?

a)

Decline of the Internet

b)

Slow progress in software

c)

Big data and high computational power

d)

Decreased demand for analytics

4.

Who helped popularize the term Data Scientist?

a)

Google

b)

LinkedIn and Facebook

c)

IBM

d)

Microsoft

5.

How does Data Science mainly differ from Data Analysis?

a)

It focuses only on statistics

b)

It covers the entire process from data collection to deployment

c)

It does not require programming

d)

It ignores understanding data

6.

Which factor makes Data Science important for businesses?

a)

Reduced productivity

b)

Increased ability to make decisions based on data

c)

Lower reliability

d)

Technology limitations

7.

In healthcare, Data Science can help with which task?

a)

Analyzing medical images and predicting diseases

b)

Reducing processing speed

c)

Increasing treatment costs

d)

Removing data

8.

An application of Data Science in finance is:

a)

Predicting market trends and credit risk

b)

Designing website interfaces

c)

Recording transactions

d)

Deleting financial history

9.

In education, Data Science can help with:

a)

Randomly ranking students

b)

Analyzing learning outcomes and optimizing curricula

c)

Reducing the amount of data

d)

Limiting the number of students

10.

In agriculture, Data Science helps mainly by:

a)

Reducing productivity

b)

Predicting weather and optimizing crop yields

c)

Deleting sensitive data

d)

Lowering product quality

11.

A notable characteristic of Data Science in the modern era is:

a)

Lack of data

b)

Information abundance and the need for smart processing tools

c)

Decreased demand for analysis

d)

No need for statistics

12.

Big Data is often described by the 3Vs. Which of the following is not one of the original 3Vs?

a)

Volume

b)

Velocity

c)

Variety

d)

Value

13.

Which is an example of Data Science in e-commerce?

a)

Personalized product recommendations

b)

Storing website pages

c)

Encoding HTML data

d)

Random advertising

14.

What is the primary role of a Data Scientist?

a)

Manually collect data

b)

Build models and derive insights

c)

Write technical manuals

d)

Draw network diagrams

15.

Which factor distinguishes modern Data Science from traditional approaches?

a)

Relying entirely on intuition

b)

Using advanced algorithms and statistical models

c)

Not needing software

d)

Using paper and pen

16.

What does Data-driven Decision Making mean?

a)

Making decisions based on intuition

b)

Basing decisions on evidence from data

c)

Randomly choosing outcomes

d)

Avoiding analysis

17.

What is a major challenge in Data Science today?

a)

Lack of data

b)

Ensuring privacy and security

c)

Absence of machine learning

d)

Inability to store data

18.

In Data Science, Predictive Analytics refers to:

a)

Analyzing historical data only

b)

Predicting future trends

c)

Describing current data

d)

Classifying data

19.

Which is an example of Data Science in logistics?

a)

Predicting delivery time

b)

Manually sorting packages

c)

Deleting orders

d)

Randomly changing routes

20.

What is the ultimate goal of Data Science?

a)

Collect as much data as possible

b)

Turn data into value and knowledge to support decisions

c)

Store data indefinitely

d)

Reduce processing costs

21.

Structured data is commonly stored in:

a)

.txt files

b)

Relational databases (SQL)

c)

Images and videos

d)

PDF files

22.

Unstructured data typically includes:

a)

CSV files

b)

Excel spreadsheets

c)

Images, videos, and free text

d)

Numeric data

23.

Semi-structured data is:

a)

Data in JSON or XML format

b)

Data in SQL tables

c)

Audio data

d)

Unreadable data

24.

The Data Cleaning process involves:

a)

Collecting data

b)

Processing to remove errors and missing values

c)

Deploying models

d)

Evaluating models

25.

What is an outlier in a dataset?

a)

Data close to the mean

b)

Missing data

c)

Abnormal data that is far from the rest

d)

Duplicate data

26.

Which is an example of discrete data?

a)

Height

b)

Number of students

c)

Temperature

d)

Weight

27.

Which is an example of continuous data?

a)

Eye color

b)

Number of products

c)

Temperature or height

d)

Student ID

28.

Which is a categorical variable?

a)

Age

b)

Satisfaction level: Low, Medium, High

c)

Weight

d)

Revenue

29.

Which is a numerical variable?

a)

Gender

b)

Occupation

c)

Income

d)

Product name

30.

Feature Engineering is the process of:

a)

Collecting new data

b)

Creating new features from raw data

c)

Evaluating models

d)

Deleting data

31.

In Machine Learning, the label is:

a)

An attribute that describes the data

b)

The target variable to be predicted

c)

The test dataset

d)

A random variable

32.

Data normalization aims to:

a)

Increase data size

b)

Standardize data to the same scale

c)

Delete all data

d)

Split data into many tables

33.

One-Hot Encoding is used to:

a)

Convert variables to strings

b)

Encode categorical variables into binary vectors

c)

Remove unnecessary variables

d)

Shuffle the data

34.

Exploratory Data Analysis helps you:

a)

Predict the future

b)

Understand features and relationships among variables

c)

Increase processing speed

d)

Reduce the amount of data

35.

Missing data is a common issue and is often handled by:

a)

Removing all data

b)

Imputing with mean, median, or predictive models

c)

Assigning entirely different values

d)

Deleting the entire dataset

36.

Time-series data is characterized by:

a)

No inherent order

b)

A time factor associated with each record

c)

No classification

d)

Only categorical values

37.

Data Integration is the process of:

a)

Combining data from multiple sources

b)

Splitting data into smaller parts

38.

Data Pipeline is used for:

a)

Manual storage

b)

Automating the collection, processing, and storage of data

c)

Copying data

d)

Visualization

39.

Big Data is characterized by:

a)

Small data volume

b)

Easy processing with Excel

c)

Large volume, velocity, and variety of data

d)

Impossible to store

40.

The ultimate goal of data processing is:

a)

Having as much data as possible

b)

Accurate, usable data for analysis and modeling

c)

Reducing the amount of data

d)

Deleting erroneous data

41.

The first step in the Data Science process is:

a)

Data collection

b)

Defining business goals and problems

c)

Model training

d)

Model deployment

42.

Data Cleaning helps to:

a)

Make data clean by removing missing and duplicate records

b)

Reduce file size

c)

Change the data format

d)

Insert data into a report

43.

Exploratory Data Analysis (EDA) is:

a)

Surveying and visualizing to understand the data

b)

Predicting future values

c)

Storing data

d)

Training models

44.

Feature Selection is the step that:

a)

Removes irrelevant or redundant variables

b)

Adds new data

c)

Predicts labels

d)

Standardizes data

45.

Train/Test Split is used to:

a)

Separate data into training and testing sets

b)

Copy data

c)

Increase dataset size

d)

Randomly shuffle data

46.

Model Training is:

a)

The process of teaching a model to learn from data

b)

Deleting erroneous data

c)

Saving the model to a file

d)

Storing data

47.

Supervised Learning requires:

a)

Labeled data

b)

Unlabeled data

c)

Random data

d)

Noisy data

48.

Unsupervised Learning is used when:

a)

Labeled data is available

b)

Labels are not available and hidden structure is sought

c)

Data is erroneous

d)

Data is small

49.

An example of an Unsupervised Learning algorithm is:

a)

Linear Regression

b)

Decision Tree

c)

K-Means Clustering

d)

Logistic Regression

50.

Model Evaluation aims to:

a)

Assess model performance on unseen data

b)

Reduce the size of data

c)

Increase data

d)

Delete variables

51.

Cross Validation helps to:

a)

Stably evaluate a model by splitting data into multiple folds

b)

Train faster

c)

Increase bias

d)

Reduce the number of features

52.

Overfitting occurs when:

a)

A model learns the training data too well and performs poorly on the test set

b)

The model has not learned enough

c)

The data is faulty

d)

The algorithm is wrong

53.

Underfitting occurs when:

a)

The model is too complex

b)

The model is too simple and fails to learn the patterns

c)

The data is faulty

d)

There is excessive data

54.

Hyperparameter Tuning is:

a)

Optimizing the model’s controlling parameters

b)

Collecting data

c)

Reducing the number of samples

d)

Increasing model size

55.

After a model performs well, the next step is:

a)

Delete the model

b)

Deployment

c)

Stop using it

d)

Cease analysis

56.

Model Monitoring aims to:

a)

Track model performance after deployment

b)

Delete faulty models

c)

Speed up processing

d)

Shut down the entire system

57.

A Feedback Loop in Data Science is:

a)

A cycle of updating the model from real-world feedback

b)

Unlimited data replication

c)

Deleting old data

d)

Stopping the model

58.

Data Visualization is typically performed during:

a)

EDA and result reporting

b)

Data collection

c)

Model training

d)

Deployment

59.

The ultimate result of the Data Science process is:

a)

An effective model that supports real-world decision-making

b)

Simple charts only

c)

Raw data

d)

An Excel file

60.

The most popular programming language in Data Science is:

a)

Java

b)

Python

c)

C++

d)

PHP

61.

Besides Python, a language widely used in statistical analysis is:

a)

R

b)

C#

c)

Kotlin

d)

Swift

62.

SQL is used to:

a)

Train models

b)

Query and manage relational databases

c)

Draw charts

d)

Create PowerPoint reports

63.

The Python library specialized for tabular data processing is:

a)

NumPy

b)

Matplotlib

c)

Pandas

d)

TensorFlow

64.

The Python library for basic Machine Learning is:

a)

Seaborn

b)

Scikit-learn

c)

Plotly

d)

PySpark

65.

TensorFlow is a framework widely used for:

a)

Database administration

b)

Deep Learning

c)

Statistical analysis

d)

Data storage

66.

PyTorch is developed by:

a)

Microsoft

b)

Amazon

c)

Meta (Facebook)

d)

IBM

67.

NumPy is primarily used for:

a)

Charting

b)

Numerical array processing and matrix computations

c)

Building web applications

d)

Data storage

68.

Matplotlib is used for:

a)

Model training

b)

Charting and data visualization

c)

Dashboard creation

d)

Database creation

69.

Seaborn is:

a)

An extension of Matplotlib for more aesthetic statistical plots

b)

A data cleaning tool

c)

A web scraping tool

d)

A data browser

70.

Power BI and Tableau are tools for:

a)

Database administration

b)

Data visualization and reporting

c)

Writing Python code

d)

Data storage

71.

Jupyter Notebook is commonly used to:

a)

Write code and present interactive data analysis

b)

Run office software

c)

Store databases

d)

Host a web server

72.

Google Colab is:

a)

An offline programming tool

b)

Google’s free online notebook environment

c)

A database management system

d)

Visualization software

73.

Apache Spark is a tool for:

a)

Storing static data

b)

Distributed processing of Big Data

c)

Creating dashboards

d)

Compressing data

74.

Hadoop is used for:

a)

Storing and processing big data using the MapReduce model

b)

Image analysis

c)

Designing neural networks

d)

Building web applications

75.

In Data Science, Docker is used to:

a)

Write Python code

b)

Package and deploy model environments

c)

Save Excel files

d)

Create reports

76.

Apache Airflow is used to:

a)

Schedule and automate data workflows

b)

Manage databases

c)

Create charts

d)

Train models

77.

MLflow is used to:

a)

Track experiments and manage machine learning models

b)

Visualize data

c)

Write SQL code

d)

Store JSON files

78.

Which of the following cloud platforms is not a primary tool for Data Science?

a)

Google Cloud

b)

AWS

c)

Microsoft Azure

d)

Canva

79.

In Data Science, the purpose of a dashboard is:

a)

Store data

b)

Visualize and monitor performance metrics

c)

Train models

d)

Clean data