wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Data Science - Chapter 1 Long Test

Total questions: 40

Worksheet time: 20mins

Name
Class
Date
1.

How does Emma show curiosity as a data scientist?

a)

By creating reports without feedback

b)

By asking clients “why” and “how” questions

c)

By applying the same model each time

d)

By avoiding meetings with stakeholders

2.

What is the first step Emma takes when working on a project?

a)

Collecting survey responses

b)

Creating a dashboard

c)

Clarifying the client’s business problem

d)

Running a machine learning model

3.

Which tools might Emma use for complex data transformations?

a)

Canva and Photoshop

b)

Talend and Informatica

c)

Tableau and Power BI

d)

Zoom and MS Teams

4.

Which programming language does Emma prefer for modeling?

a)

Python

b)

PHP

c)

Java

d)

C++

5.

Which tools does Emma use to create reports and dashboards?

a)

Photoshop, Canva, and Figma

b)

SQL, Hadoop, and Spark

c)

Tableau, Power BI, and QlikView

d)

Word, PowerPoint, and Publisher

6.

In Emma’s project, data transformation and exploratory data analysis (EDA) serve the same purpose, so performing one makes the other unnecessary.

a)

True

b)

False

7.

Emma’s ability to ask “why” and “how” questions is considered a technical skill

a)

True

b)

False

8.

In Emma’s workflow, data acquisition can only be performed manually through surveys or interviews with clients.

a)

True

b)

False

9.

XML files use predefined tags like <p> and <h1> from HTML, making them less flexible.

a)

True

b)

False

10.

PDF files are frequently used in data science for analytics because they are highly structured and easy to process.

a)

True

b)

False

11.

Data science is transforming how organizations make decisions, but it does not change how they understand the world.

a)

True

b)

False

12.

The first step of the data science process is to collect as much data as possible before clarifying the problem.

a)

True

b)

False

13.

Even if Emma selects the best-performing machine learning model, poor communication of results could make her work ineffective for the client.

a)

True

b)

False

14.

The role of the data scientist can be compared to that of a detective, uncovering hidden patterns and guiding decisions.

a)

True

b)

False

15.

High achievers in the student studies preferred practice tests as a review method, showing how pattern recognition can reveal unexpected findings.

a)

True

b)

False

16.

XLSX files, unlike CSVs, can support multiple worksheets, formulas, and pivot tables.

a)

True

b)

False

17.

A strength of XML is that it is platform- and language-independent, which makes it excellent for sharing data between systems.

a)

True

b)

False

18.

JSON stores data as key–value pairs, making it both lightweight and easy to parse by most programming languages.

a)

True

b)

False

19.

Data cleaning may involve correcting misspelled attributes, removing duplicates, and filling in missing values.

a)

True

b)

False

20.

A data scientist who fails to communicate insights effectively leaves their findings essentially useless.

a)

True

b)

False

21.

Which example shows data transformation in Emma’s project?

a)

Asking stakeholders to provide more survey responses

b)

Deleting entire datasets that seem unnecessary

c)

Copying results into PowerPoint slides for presentation

d)

Converting dates into a standard format across all records

22.

Which of the following is the main goal of Exploratory Data Analysis (EDA) in Emma’s project?

a)

To identify patterns and refine variables before modeling

b)

To test whether the client agrees with her assumptions

c)

To reduce the size of the dataset for faster storage

d)

To replace missing values with zeros by default

23.

Which scenario shows Emma using logistic regression appropriately?

a)

Identifying which products often appear together in shopping carts

b)

Predicting whether students will pass or fail an exam

c)

Grouping customers into clusters based on purchase history

d)

Recommending movies based on viewing history

24.

Which real-world example illustrates model deployment?

a)

A student calculating GPA manually with a calculator

b)

A teacher posting grades on the bulletin board

c)

A bank using a fraud detection system to flag suspicious transactions

d)

A developer editing images using graphic design tools

25.

Which sequence correctly represents Emma’s data science workflow?

a)

Clarify → Collect → Clean/Transform → EDA → Model → Communicate → Deploy → Monitor

b)

Clean/Transform → Communicate → EDA → Collect → Model → Deploy → Monitor → Clarify

c)

Collect → Model → Clean/Transform → Deploy → Clarify → Communicate → EDA → Monitor

d)

Monitor → Deploy → Model → Clarify → Collect → Communicate → Clean/Transform → EDA

26.

Which punctuation mark is most commonly used as the delimiter in a CSV file?

a)

Tab

b)

Semicolon

c)

Comma

d)

Space

27.

Which symbol is used as the delimiter in a TSV file?

a)

Colon

b)

Tab

c)

Vertical bar

d)

Comma

28.

Which best describes the main function of CSV and TSV files?

a)

To represent data in plain text rows, separated by delimiters

b)

To transmit hierarchical data using user-defined tags

c)

To store complex data relationships with formulas and macros

d)

To preserve document formatting across multiple devices

29.

Why has JSON become the standard for web services and APIs?

a)

It uses predefined HTML tags for storing data

b)

It is used mainly to design user interfaces

c)

It can only be read by JavaScript and not by other languages

d)

It is more concise, lightweight, and easier to parse than XML

30.

Why are PDFs not commonly used for direct data analysis?

a)

They preserve formatting but are difficult to extract structured data from

b)

They require a specific programming language to open

c)

They are incompatible with legal and financial industries

d)

They are too large to be shared

31.

Which example reflects unstructured data that a data scientist might analyze?

a)

Exam scores stored in a gradebook

b)

Customer complaints in social media posts

c)

Employee payroll records in spreadsheets

d)

Sales transactions recorded in a database

32.

Which sequence correctly follows the data science process?

a)

Clarify → Collect → Analyze → Communicate

b)

Collect → Clarify → Model → Communicate

c)

Analyze → Clarify → Collect → Monitor

d)

Collect → Model → Deploy → Clarify

33.

Which tool would be most suitable for helping stakeholders see and act on insights?

a)

Notepad

b)

Power BI

c)

Spreadsheet

d)

Calculator

34.

According to Professor Murtaza Haider, which quality ensures a data scientist continues asking meaningful questions?

a)

Argumentation

b)

Flexibility

c)

Curiosity

d)

Judgment

35.

Which quality enables a data scientist to refine assumptions as new data emerges?

a)

Argumentation

b)

Inflexibility

c)

Judgment

d)

Rigidity

36.

Which dataset is considered structured?

a)

A transcript of an interview saved as a Word document

b)

A sales database with product IDs, prices, and quantities

c)

A folder of digital photos from a school event

d)

A collection of YouTube video uploads

37.

Which of the following is NOT an example of structured data?

a)

Library catalog entries stored in a relational database

b)

Payroll records with employee IDs and monthly salaries

c)

Online survey results exported as CSV files

d)

Recorded classroom lectures stored as MP4 files

38.

Which dataset would a data scientist classify as structured?

a)

Voice notes recorded by doctors describing patient symptoms

b)

A healthcare database listing patient IDs, age, and blood pressure readings

c)

CT scan images stored in a hospital’s image archive

d)

Patient complaints expressed as handwritten forms scanned into PDFs

39.

Which of the following is structured?

a)

Classroom debates recorded as MP3 files

b)

Essays submitted by students in Word documents

c)

Research posters photographed at a school event

d)

Student names, IDs, and grades stored in Excel

40.

Which dataset is structured rather than unstructured or semi-structured?

a)

SQL database tables with student exam scores

b)

JSON output listing user activity logs

c)

Blog posts describing student experiences

d)

XML files containing nested tags for course descriptions