WorksheetsLIBA | Data Preprocessing
Total questions: 10
Worksheet time: 3mins
You have a dataset where a column has 95% missing values. The best preprocessing step is:
Replace missing values with mean
Replace missing values with zero
Drop the column
Use KNN imputation
Which method is most robust to outliers when imputing missing numeric values?
Mean
Median
Mode
Min
When performing one-hot encoding, the number of new columns created is:
Equal to the number of unique values
Equal to number of rows
Always 1
Equal to number of unique values minus 1
If two features are highly correlated (say r > 0.9), what is the main concern?
Missing values
Multicollinearity
Outliers
Overfitting
In Min-Max normalization, the transformed value of x is given by:
(x - mean) / std
(x - min) / (max - min)
log(x)
sqrt(x)
Which of the following is a method to detect outliers in numeric data?
Boxplot and IQR
Scatter plot
Z-score method
All of the above
When using Z-score standardization, the transformed data has:
Min = 0, Max = 1
Mean = 0, Std = 1
Median = 0, Std = 1
Mean = 1, Std = 0
In handling missing categorical data, the most common approach is:
Replace with mean
Replace with mode
Replace with median
Drop column
If two features are strongly negatively correlated (r ≈ -0.95), which is true?
One can be removed to reduce redundancy
They are independent
Both are essential
Correlation does not matter in preprocessing
After applying standardization (Z-score), a value of -2 indicates:
Two standard deviations below the mean
Two times the mean
Two times the standard deviation above the mean
Cannot be interpreted
