Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

FAM_MODULE_456

Total questions: 30

Worksheet time: 15mins

Name
Class
Date
1.

Which of the following best describes data collection?

a)

 Transforming data into numerical format

b)

Gathering information from various sources for analysis

c)

 Eliminating duplicates in a dataset

d)

Visualizing data in graphs

2.

A retail store uses an online survey to gather customer feedback. Which method is being used?

a)

 Observational study

b)

Survey/Questionnaire

c)

 Web scraping

d)

 Interviews

3.

Which method involves observing people in their natural setting?

a)

 Survey

b)

Focus group

c)

Observational study

d)

 Web scraping

4.

Using past company sales records is an example of:

a)

 Web scraping

b)

Existing data sources

c)

 Interview

d)

Survey

5.

Automated tools that extract competitor website data is called:

a)

Survey

b)

 Observational study

c)

Web scraping

d)

Data validation

6.

Which method is best for qualitative insights?

a)

 Surveys

b)

 Observations

c)

Web scraping

d)

Interviews and Focus Groups

7.

Duplicate entries must be removed because they:

a)

Can skew analysis results

b)

 Increase accuracy

c)

Reduce sample size

d)

 Enhance performance

8.

Standardizing all dates into MM/DD/YYYY is:

a)

Feature engineering

b)

 Data validation

c)

Standardizing formats

d)

 Imputation

9.

Which is NOT a data cleaning technique?

a)

 Removing duplicates

b)

 Correcting errors

c)

 Outlier detection

d)

Regression modeling

10.

A dataset shows “Manila” and “MNL.” This requires:

a)

Outlier detection

b)

Correcting errors

c)

 Deletion

d)

Encoding

11.

Which technique is best for predicting a continuous outcome?

a)

 Decision Tree

b)

Regression

c)

Clustering

d)

Time-Series

12.

Logistic regression is commonly used for:

a)

Predicting sales

b)

Forecasting electricity demand

c)

 Binary outcomes (Yes/No)

d)

Customer segmentation

13.

Which regression method assumes a straight-line relationship?

a)

 Linear Regression

b)

 Logistic Regression

c)

Hierarchical Clustering

d)

Decision Tree

14.

Predicting house prices using size, location, and age is an example of:

a)

 Logistic Regression

b)

Multiple Regression

c)

Clustering

d)

ARIMA

15.

Which model visually represents decision-making with branches?

a)

Regression

b)

Decision Tree

c)

Clustering

d)

 Exponential Smoothing

16.

What is the biggest disadvantage of decision trees?

a)

High interpretability

b)

 Handles both numerical and categorical data

c)

Overfitting

d)

Uses probabilities

17.

A bank wants to classify loan applicants into approve/deny categories. The best model is:

a)

 Clustering

b)

Regression

c)

Decision Tree

d)

ARIMA

18.

Which model groups customers with similar behaviors?

a)

Regression

b)

Decision Trees

c)

Clustering

d)

Time-Series

19.

K-Means clustering works by:

a)

Assigning points to the nearest cluster mean

b)

Building a tree of nodes

c)

Drawing a regression line

d)

 Using exponential smoothing

20.

Hierarchical clustering is best for:

a)

 Forecasting demand

b)

Predicting binary outcomes

c)

Understanding relationships between clusters

d)

Pruning decision trees

21.

Which metric is most misleading in imbalanced datasets?

a)

Precision

b)

Recall

c)

Accuracy

d)

F1 Score

22.

A hospital needs to ensure that patients with a disease are detected. Which metric is most important?

a)

Precision

b)

Recall

c)

 Accuracy

d)

Silhouette Score

23.

The F1 Score is the:

a)

 Arithmetic mean of precision and recall

b)

Weighted sum of precision and recall

c)

Harmonic mean of precision and recall

d)

 Ratio of true positives to total samples

24.

ROC-AUC measures:

a)

Regression variance explained

b)

 Trade-off between true positive rate and false positive rate

c)

 Model interpretability

d)

Cluster separation

25.

K-fold cross-validation divides the dataset into:

a)

 k subsets used for training and testing in rotation

b)

Equal halves

c)

Training only

d)

 Testing only

26.

Which is most interpretable?

a)

 Linear regression

b)

Random forests

c)

Neural networks

d)

 Gradient boosting

27.

Fairness in models means:

a)

 Highest accuracy possible

b)

Equal treatment of different groups

c)

 Minimum computation cost

d)

No retraining needed

28.

A/B testing helps in:

a)

 Balancing datasets

b)

 Feature selection

c)

 Comparing old and new models

d)

Removing outliers

29.

Model performance drift means:

a)

 Accuracy suddenly increases

b)

 Performance degrades over time due to changing data

c)

 Dataset duplication

d)

 Model stops training

30.

Retraining a model is done to:

a)

 Save storage space

b)

Adapt to new data patterns

c)

 Increase interpretability

d)

 Avoid testing