Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

LIBA | Data Preprocessing

Total questions: 10

Worksheet time: 3mins

Name
Class
Date
1.

You have a dataset where a column has 95% missing values. The best preprocessing step is:

a)

Replace missing values with mean

b)

Replace missing values with zero

c)

Drop the column

d)

Use KNN imputation

2.

Which method is most robust to outliers when imputing missing numeric values?

a)

Mean

b)

Median

c)

Mode

d)

Min

3.

When performing one-hot encoding, the number of new columns created is:

a)

Equal to the number of unique values

b)

Equal to number of rows

c)

Always 1

d)

Equal to number of unique values minus 1

4.

If two features are highly correlated (say r > 0.9), what is the main concern?

a)

Missing values

b)

Multicollinearity

c)

Outliers

d)

Overfitting

5.

In Min-Max normalization, the transformed value of x is given by:

a)

(x - mean) / std

b)

(x - min) / (max - min)

c)

log(x)

d)

sqrt(x)

6.

Which of the following is a method to detect outliers in numeric data?

a)

Boxplot and IQR

b)

Scatter plot

c)

Z-score method

d)

All of the above

7.

When using Z-score standardization, the transformed data has:

a)

Min = 0, Max = 1

b)

Mean = 0, Std = 1

c)

Median = 0, Std = 1

d)

Mean = 1, Std = 0

8.

In handling missing categorical data, the most common approach is:

a)

Replace with mean

b)

Replace with mode

c)

Replace with median

d)

Drop column

9.

If two features are strongly negatively correlated (r ≈ -0.95), which is true?

a)

One can be removed to reduce redundancy

b)

They are independent

c)

Both are essential

d)

Correlation does not matter in preprocessing

10.

After applying standardization (Z-score), a value of -2 indicates:

a)

Two standard deviations below the mean

b)

Two times the mean

c)

Two times the standard deviation above the mean

d)

Cannot be interpreted