Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

CSS330

Total questions: 100

Worksheet time: 39mins

Name
Class
Date
1.

What is the most common simple technique to handle missing numerical data without deleting rows?

a)

 One-Hot Encoding

b)

Feature Scaling

c)

 Dropping the column

d)

 Imputation

2.

 If a dataset has a large number of missing values (e.g., 90%) in a single column, what is usually the best approach?

a)

 Drop the column

b)

Fill with the median

c)

 Fill with the mean

d)

 Leave it as is

3.

. Which term describes missing data where the probability of missingness is completely random and unrelated to any other variable?

a)

MNAR (Missing Not At Random)

b)

MAR (Missing At Random)

c)

MCAR (Missing Completely At Random)

d)

OAR (Observed At Random)

4.

Removing rows that contain missing data is technically known as:

a)

 Imputation

b)

Listwise deletion

c)

 Binning

d)

Normalization

5.

When using "Mean Imputation" to fill missing values, what happens to the overall variance of the dataset?

a)

It decreases

b)

 It increases

c)

 It becomes infinite

d)

It remains exactly the same

6.

 Which value is generally best to use for imputation if the numerical feature contains significant outliers?

a)

Maximum value

b)

Mean

c)

Standard Deviation

d)

Median

7.

Which Python library is most commonly used for high-level DataFrame manipulation and detecting missing values (e.g., .isnull())?

a)

NumPy

b)

Matplotlib

c)

Scikit-Learn

d)

 Pandas

8.

Forward fill (ffill) is a technique most often effective in what type of data?

a)

Time-series data

b)

Categorical data

c)

Image data

d)

 Unstructured text data

9.

Imputing missing values with a constant placeholder (e.g., -1 or 0) is specifically useful when:

a)

The data follows a Gaussian distribution

b)

The missingness itself might be meaningful information

c)

You want to increase the variance of the model

d)

You are using Linear Regression

10.

 Min-Max Normalization transforms data into which specific range?

a)

infinity to +infinity

b)

0 to 1

c)

-1 to 1

d)

0 to 100

11.

Standardization (Z-score normalization) transforms data to have a mean of:

a)

1

b)

0

c)

0.5

d)

10

12.

 Which scaling technique is most sensitive to outliers, often causing the majority of valid data to be compressed into a tiny interval?

a)

Robust Scaler

b)

Standardization

c)

Log Transformation

d)

Min-Max Normalization

13.

What is the standard deviation of a variable after it has been Standardized (Z-score)?

a)

1

b)

0

c)

100

d)

 The same as the original data

14.

When using algorithms based on distance (like k-NN or K-Means), why is feature scaling crucial?

a)

 It converts categorical text data into numbers

b)

It prevents features with large magnitudes from dominating the distance calculation

c)

It removes missing values from the dataset

d)

It guarantees that the model will not overfit

15.

Which formula represents Min-Max Scaling?

a)

 x^2

b)

(x − mean) / std_dev

c)

 (x − min) / (max − min)

d)

log(x)

16.

If your data does not follow a Gaussian (Normal) distribution, which scaling method is generally recommended as the default?

a)

Standardization

b)

Mean Removal

c)

 Log Transformation

d)

Min-Max Normalization

17.

Which scaling method is specifically designed to be robust against outliers by using the Interquartile Range (IQR)?

a)

Robust Scaler

b)

Min-Max Scaler

c)

Standard Scaler

d)

Max Abs Scaler

18.

Standardization is generally preferred (and often required for faster convergence) for which type of algorithm?

a)

Decision Trees

b)

Gradient Descent-based algorithms

c)

Random Forests

d)

Rule-based systems

19.

Which statistical measure is the primary component used to define the "whiskers" and detect outliers in a Boxplot?

a)

Mean

b)

 Standard Deviation

c)

 Interquartile Range (IQR)

d)

Mode

20.

According to the standard method (used in boxplots), a data point is considered an outlier if it falls outside the range of ___ times the IQR.

a)

0.5

b)

1.5

c)

5.0

d)

1.0

21.

In the context of outlier treatment, "Trimming" specifically refers to:

a)

Capping values at a specific percentile

b)

Replacing outliers with the mean

c)

Transforming the feature using a Log function

d)

Removing the rows containing outliers from the dataset

22.

 When using the Z-score method, which threshold is most commonly used to flag a data point as an outlier?

a)

> 1

b)

 > 10

c)

> 3

d)

> 0.5

23.

 "Winsorization" is a technique that involves:

a)

Transforming the distribution using a Log scale

b)

Deleting all rows with outliers

c)

Replacing extreme values with specified boundary values (e.g., the 5th and 95th percentiles)

d)

Normalizing the data to a 0-1 range

24.

Which visualization is considered the standard tool for spotting outliers in a single numerical variable?

a)

Heatmap

b)

Boxplot

c)

Pie Chart

d)

Bar Chart

25.

Outliers in a dataset can be caused by:

a)

Data entry or processing errors

b)

Measurement instrument errors

c)

Natural variation in the population

d)

 All of the above

26.

Which statistical measure is most sensitive to outliers (meaning its value changes drastically if an extreme value is added)?

a)

Median

b)

 Interquartile Range (IQR)

c)

Mean

d)

Mode

27.

If you identify an outlier that is a valid, natural data point (not an error), what is generally the best course of action?

a)

 Delete it immediately to improve model accuracy

b)

Change its value to 0 to neutralize it

c)

Investigate the context to decide whether to keep, treat, or model it separately

d)

Hide it from the training set but keep it in the test set

28.

 Label Encoding converts categorical text values into:

a)

 Floating point numbers (0.1, 0.2...)

b)

Vectors of 0s and 1s

c)

 Integers (0, 1, 2...)

d)

Text strings

29.

What is the main disadvantage of using Label Encoding for nominal data (categories with no inherent order)?

a)

It creates too many new columns, increasing dimensionality

b)

It cannot handle missing values

c)

It takes up too much memory compared to One-Hot Encoding

d)

The model might misinterpret the integer numbers as a mathematical ranking (e.g., 2 > 1)

30.

If a categorical feature has k unique values, how many new columns does standard One-Hot Encoding create?

a)

1

b)

k (or k-1 if avoiding collinearity)

c)

k^2

d)

 log(k)

31.

Which encoding technique is most appropriate for Ordinal data (e.g., "Low", "Medium", "High") where the order matters?

a)

Binary Encoding

b)

Frequency Encoding

c)

One-Hot Encoding

d)

 Label/Ordinal Encoding

32.

The "Dummy Variable Trap" generally refers to:

a)

Encoding numerical variables as text strings

b)

Perfect multicollinearity (high correlation) between one-hot encoded columns

c)

Using too many categories in a single feature

d)

Having missing data in categorical columns

33.

 If a categorical feature has 1,000 unique values (high cardinality), which method is problematic because it drastically increases the dataset's dimensionality?

a)

Frequency Encoding

b)

Target Encoding

c)

Label Encoding

d)

One-Hot Encoding

34.

Frequency Encoding replaces categorical labels with:

a)

A random integer

b)

The count or percentage of their occurrence in the dataset

c)

The mean of the target variable

d)

 Their alphabetical rank

35.

The function pd.get_dummies() in the Pandas library is specifically used for:

a)

Normalization

b)

One-Hot Encoding

c)

Missing Value Imputation

d)

Label Encoding

36.

Which of the following is an example of a Nominal variable (no intrinsic order)?

a)

Temperature (0°C, 10°C, 20°C)

b)

T-shirt size (S, M, L)

c)

City (Paris, London, Tokyo)

d)

Salary ($50k, $60k, $70k)

37.

 Feature Engineering is the process of:

a)

Cleaning missing data

b)

 Selecting the best model

c)

Creating new features from existing data to improve model performance

d)

Visualizing data

38.

 "Binning" allows you to convert:

a)

Images into Vectors

b)

Text into Numbers

c)

Categorical variables into Numerical

d)

Numerical variables into Categorical or Intervals

39.

Extracting "Day of Week" from a "Date" column is an example of:

a)

Data Imputation

b)

Feature Extraction/Engineering

c)

 Feature Scaling

d)

Dimensionality Reduction

40.

Which of the following is an "Interaction Feature"?

a)

log(x1)

b)

x1^2

c)

x1 + x2

d)

x1 * x2

41.

Polynomial features are created to:

a)

Capture non-linear relationships in the data

b)

Reduce the number of features

c)

Remove outliers

d)

Normalize the data

42.

Calculating the length of a text string to use as a predictor is:

a)

Clustering

b)

Data Cleaning

c)

Feature Engineering

d)

Data Leakage

43.

"Domain Knowledge" is useful in feature engineering because:

a)

 It is not useful

b)

It helps identify irrelevant features and construct meaningful ones

c)

It allows you to skip the testing phase

d)

 It is the only way to handle missing values

44.

Creating a "Total Income" feature by adding "Applicant Income" and "Co-applicant Income" is:

a)

Feature Selection

b)

 One-Hot Encoding

c)

Feature Scaling

d)

Feature Creation

45.

The primary goal of feature engineering is to:

a)

Make the data compatible with the algorithm and improve predictive power

b)

Hide sensitive data

c)

Make the dataset larger

d)

Remove all categorical variables

46.

The main goal of PCA (Principal Component Analysis) is to:

a)

Impute missing values

b)

Classify data into groups

c)

Reduce the dimensionality of data while preserving variance

d)

Increase the number of features

47.

PCA is what type of learning algorithm?

a)

Supervised

b)

Semi-supervised

c)

Reinforcement

d)

Unsupervised

48.

The new features created by PCA are called:

a)

Support Vectors

b)

Principal Components

c)

Centroids

d)

Decision Boundaries

49.

In PCA, the first Principal component captures:

a)

The mean of the data

b)

The most variance

c)

The noise only

d)

The least variance

50.

Principle components are always:

a)

Larger than the original features

b)

Categorical variables

c)

Correlated with each other

d)

Orthogonal( uncorrelated ) to each other

51.

Before applying PCA, it is crucial to:

a)

Convert all data to text

b)

Scale/Standardize the data

c)

Remove all negative numbers

d)

Sort the data

52.

If you reduce 100 features to 2 using PCA, you can easily:

a)

E;iminate all bias

b)

Visualize the data in a 2D scatter plot

c)

Calculate the accuracy

d)

Ignore the target variable

53.

The "Explained Variance Ratio" tells you:

a)

The accuracy of the model

b)

How many clusters to pick

c)

How much information (variance) each component holds

d)

The number of missing values

54.

PCA works best on variables that are:

a)

Completely independent

b)

Constant

c)

Highly correlated

d)

Categorical

55.

The Pearson correlation coefficient ranges from:

a)

0 to 100

b)

-1 to 1

c)

-10 to 10

d)

0 to 1

56.

A correlation of -0.9 indicates:

a)

A strong positive relationship

b)

 No relationship

c)

A weak negative relationship

d)

A strong negative relationship

57.

Why is high multicollinearity a problem for linear models?

a)

 It makes the data unreadable

b)

It reduces predictive accuracy heavily

c)

It makes the coefficient estimates unstable and hard to interpret

d)

 It causes the model to crash

58.

Multicollinearity generally refers to:

a)

Low variance in dataLow variance in data

b)

High correlation between features and the target

c)

Missing data in multiple columns

d)

 High correlation between two or more independent variables

59.

Which metric is used to detect multicollinearity?

a)

VIF

b)

RMSE

c)

AUC

d)

 R-squared

60.

If two features have a correlation of 1.0, you should:

a)

Square them

b)

Remove one of them

c)

Keep both

d)

Multiply them together

61.

Which correlation method works best for non-linear or ordinal relationships?

a)

 Euclidean Distance

b)

Spearman (Rank Correlation)

c)

 Mean Correlation

d)

 Pearson

62.

 A correlation of 0 implies:

a)

Perfect negative relationship

b)

No linear relationship

c)

An error in calculation

d)

Strong linear relationship

63.

Which tool is commonly used to visualize correlation matrices?

a)

Pie Chart

b)

Histogram

c)

Heatmap

d)

Boxplot

64.

The primary purpose of EDA is to:

a)

Deploy the model

b)

Write the documentation

c)

Understand the data structure, patterns, and anomalies

d)

Build the final model immediately

65.

df.describe() in Pandas provides:

a)

 Summary statistics (mean, std, min, max)

b)

The first 5 rows of data

c)

A correlation matrix

d)

A histogram plot

66.

Which plot is best for seeing the distribution of a single numerical variable?

a)

Scatter Plot

b)

Histogram

c)

Scatter Matrix

d)

Heatmap

67.

A Scatter Plot is primarily used to show the relationship between:

a)

Two categorical variables

b)

 Time and frequency

c)

One categorical and one numerical variable

d)

Two numerical variables

68.

 To visualize the count of items in different categories, use a:

a)

Scatter Plot

b)

Line Chart

c)

Bar Chart

d)

Boxplot

69.

df.head() displays:

a)

The column names only

b)

The missing values

c)

The correlation matrix

d)

 The first N rows of the dataset

70.

A "Pair Plot" allows you to visualize:

a)

Only missing data

b)

Relationships between all pairs of variables in a dataset

c)

Only the target variable

d)

The time taken to train

71.

Univariate analysis involves analyzing:

a)

One variable at a time

b)

Two variables at a time

c)

Three variables at a time

d)

 The entire database

72.

Skewness measures:

a)

 The height of the distribution

b)

The flatness of the distribution

c)

The asymmetry of the probability distribution

d)

The number of outliers

73.

Data Leakage occurs when:

a)

Data is lost during transfer

b)

The model performs poorly on training data

c)

The database is hacked

d)

Information from outside the training dataset (like the test set) is used to create the model

74.

Which of the following is a common cause of data leakage?

a)

Imputing using only the training set mean

b)

Using Cross-Validation

c)

Removing outliers

d)

Imputing missing values using the mean of the entire dataset (train + test)

75.

 If your model has 99.9% accuracy on both training and test set immediately, you should suspect:

a)

Your model is perfect

b)

Underfitting

c)

Data Leakage

d)

Bad data

76.

In time-series forecasting, using future data to predict the past is:

a)

 Feature Engineering

b)

Look-ahead bias (Leakage)

c)

Backpropagation

d)

Lagging

77.

To prevent leakage during scaling/normalization, you should:

a)

fit_transform on the whole dataset

b)

fit on the training set, then transform both training and test sets

c)

fit on the test set only

d)

Not scale the data

78.

 Including the target variable (or a proxy for it) as a feature:

a)

Is a severe form of Data Leakage

b)

Reduces overfitting

c)

Is necessary for supervised learning

d)

 Is good feature engineering

79.

ID columns (e.g., PatientID, TransactionID) should usually be:

a)

Used as a primary feature

b)

Squared

c)

Used for normalization

d)

Dropped, as they can cause leakage if they correlate with the target

80.

When using Cross-Validation, feature selection should be done:

a)

Inside the cross-validation loop (on the training folds only)

b)

After the model is built

c)

Before the cross-validation loop on the whole dataset

d)

Never

81.

Which is the “Golden Rule” to avoid leakage?

a)

The test set should never be seen or touched during training/preprocessing

b)

The test set should be added to the training set

c)

Always use Deep Learning

d)

Preprocessing is not important

82.

Increasing model complexity usually causes:

a)

Decrease bias and increase variance

b)

Increase bias and decrease variance

c)

Decrease both bias and variance

d)

 Increase bias and increase variance

83.

A model with high bias typically:

a)

Performs similarly on train and test, but both poorly

b)

Fits training data extremely well

c)

 Has a large gap between train and test performance

d)

Needs more data to generalize

84.

High variance is most strongly associated with:

a)

Class imbalance

b)

Overfitting

c)

 Data leakage

d)

Underfitting

85.

If both training and validation errors are high, the model likely has:

a)

 Low bias

b)

 High variance

c)

High bias

d)

Low variance

86.

 A common way to reduce variance is to:

a)

Use regularization or a simpler model

b)

Add more features without regularization

c)

 Increase learning rate drastically

d)

Reduce training data size

87.

Which pattern indicates high variance?

a)

Train error low, test error low

b)

Train error high, test error high

c)

Train error high, test error low

d)

Train error low, test error high

88.

Bagging (Bootstrap Aggregation) mainly reduces:

a)

Neither

b)

Bias

c)

Variance

d)

Both bias and variance equally

89.

 Boosting often helps reduce:

a)

Noise only

b)

Bias

c)

Variance only

d)

Feature scaling need

90.

If model complexity increases and training error rises slightly while validation error falls, the best interpretation is:

a)

Generalization is improving even if training fit changes

b)

Data leakage is happening

c)

Model is getting worse overall

d)

Regularization is too weak

91.

Underfitting means:. Underfitting means:

a)

 Test data is too small

b)

Model is too simple to capture patterns

c)

Features are redundant

d)

Model captures noise

92.

Overfitting occurs when:

a)

Train accuracy is much higher than test accuracy

b)

Both accuracies are low

c)

Test accuracy is higher than train accuracy

d)

Model is linear

93.

Which action most directly combats underfitting?

a)

Remove informative features

b)

Add regularization

c)

Reduce model complexity

d)

Increase model complexity

94.

 A very deep decision tree is prone to:

a)

Normalization errors

b)

Underfitting

c)

Class imbalance only

d)

Overfitting

95.

If validation loss starts rising while training loss keeps dropping, you should:

a)

Increase learning rate

b)

Train longer

c)

Stop training or add regularization

d)

Remove cross-validation

96.

 Early stopping helps mainly by reducing:

a)

Variance

b)

Dimensionality

c)

Bias

d)

 Missing values

97.

 A degree-1 polynomial regression underfits curved data because:

a)

It assumes independent errors

b)

 It requires scaling

c)

It can’t model non-linearity

d)

It has too many parameters

98.

 Dropout in neural networks mostly reduces:

a)

Target leakage

b)

Bias

c)

Feature dependence

d)

Variance

99.

After increasing training data size, train error rises while test error drops. Best conclusion:

a)

Data became noisier

b)

Bias increased

c)

Variance decreased

d)

Model became worse

100.

 The main goal of a train–test split is to:

a)

Increase variance

b)

Improve training accuracy

c)

Reduce bias

d)

Estimate performance on unseen data