WorksheetsCSS330
Total questions: 100
Worksheet time: 39mins
What is the most common simple technique to handle missing numerical data without deleting rows?
One-Hot Encoding
Feature Scaling
Dropping the column
Imputation
If a dataset has a large number of missing values (e.g., 90%) in a single column, what is usually the best approach?
Drop the column
Fill with the median
Fill with the mean
Leave it as is
. Which term describes missing data where the probability of missingness is completely random and unrelated to any other variable?
MNAR (Missing Not At Random)
MAR (Missing At Random)
MCAR (Missing Completely At Random)
OAR (Observed At Random)
Removing rows that contain missing data is technically known as:
Imputation
Listwise deletion
Binning
Normalization
When using "Mean Imputation" to fill missing values, what happens to the overall variance of the dataset?
It decreases
It increases
It becomes infinite
It remains exactly the same
Which value is generally best to use for imputation if the numerical feature contains significant outliers?
Maximum value
Mean
Standard Deviation
Median
Which Python library is most commonly used for high-level DataFrame manipulation and detecting missing values (e.g., .isnull())?
NumPy
Matplotlib
Scikit-Learn
Pandas
Forward fill (ffill) is a technique most often effective in what type of data?
Time-series data
Categorical data
Image data
Unstructured text data
Imputing missing values with a constant placeholder (e.g., -1 or 0) is specifically useful when:
The data follows a Gaussian distribution
The missingness itself might be meaningful information
You want to increase the variance of the model
You are using Linear Regression
Min-Max Normalization transforms data into which specific range?
infinity to +infinity
0 to 1
-1 to 1
0 to 100
Standardization (Z-score normalization) transforms data to have a mean of:
1
0
0.5
10
Which scaling technique is most sensitive to outliers, often causing the majority of valid data to be compressed into a tiny interval?
Robust Scaler
Standardization
Log Transformation
Min-Max Normalization
What is the standard deviation of a variable after it has been Standardized (Z-score)?
1
0
100
The same as the original data
When using algorithms based on distance (like k-NN or K-Means), why is feature scaling crucial?
It converts categorical text data into numbers
It prevents features with large magnitudes from dominating the distance calculation
It removes missing values from the dataset
It guarantees that the model will not overfit
Which formula represents Min-Max Scaling?
x^2
(x − mean) / std_dev
(x − min) / (max − min)
log(x)
If your data does not follow a Gaussian (Normal) distribution, which scaling method is generally recommended as the default?
Standardization
Mean Removal
Log Transformation
Min-Max Normalization
Which scaling method is specifically designed to be robust against outliers by using the Interquartile Range (IQR)?
Robust Scaler
Min-Max Scaler
Standard Scaler
Max Abs Scaler
Standardization is generally preferred (and often required for faster convergence) for which type of algorithm?
Decision Trees
Gradient Descent-based algorithms
Random Forests
Rule-based systems
Which statistical measure is the primary component used to define the "whiskers" and detect outliers in a Boxplot?
Mean
Standard Deviation
Interquartile Range (IQR)
Mode
According to the standard method (used in boxplots), a data point is considered an outlier if it falls outside the range of ___ times the IQR.
0.5
1.5
5.0
1.0
In the context of outlier treatment, "Trimming" specifically refers to:
Capping values at a specific percentile
Replacing outliers with the mean
Transforming the feature using a Log function
Removing the rows containing outliers from the dataset
When using the Z-score method, which threshold is most commonly used to flag a data point as an outlier?
> 1
> 10
> 3
> 0.5
"Winsorization" is a technique that involves:
Transforming the distribution using a Log scale
Deleting all rows with outliers
Replacing extreme values with specified boundary values (e.g., the 5th and 95th percentiles)
Normalizing the data to a 0-1 range
Which visualization is considered the standard tool for spotting outliers in a single numerical variable?
Heatmap
Boxplot
Pie Chart
Bar Chart
Outliers in a dataset can be caused by:
Data entry or processing errors
Measurement instrument errors
Natural variation in the population
All of the above
Which statistical measure is most sensitive to outliers (meaning its value changes drastically if an extreme value is added)?
Median
Interquartile Range (IQR)
Mean
Mode
If you identify an outlier that is a valid, natural data point (not an error), what is generally the best course of action?
Delete it immediately to improve model accuracy
Change its value to 0 to neutralize it
Investigate the context to decide whether to keep, treat, or model it separately
Hide it from the training set but keep it in the test set
Label Encoding converts categorical text values into:
Floating point numbers (0.1, 0.2...)
Vectors of 0s and 1s
Integers (0, 1, 2...)
Text strings
What is the main disadvantage of using Label Encoding for nominal data (categories with no inherent order)?
It creates too many new columns, increasing dimensionality
It cannot handle missing values
It takes up too much memory compared to One-Hot Encoding
The model might misinterpret the integer numbers as a mathematical ranking (e.g., 2 > 1)
If a categorical feature has k unique values, how many new columns does standard One-Hot Encoding create?
1
k (or k-1 if avoiding collinearity)
k^2
log(k)
Which encoding technique is most appropriate for Ordinal data (e.g., "Low", "Medium", "High") where the order matters?
Binary Encoding
Frequency Encoding
One-Hot Encoding
Label/Ordinal Encoding
The "Dummy Variable Trap" generally refers to:
Encoding numerical variables as text strings
Perfect multicollinearity (high correlation) between one-hot encoded columns
Using too many categories in a single feature
Having missing data in categorical columns
If a categorical feature has 1,000 unique values (high cardinality), which method is problematic because it drastically increases the dataset's dimensionality?
Frequency Encoding
Target Encoding
Label Encoding
One-Hot Encoding
Frequency Encoding replaces categorical labels with:
A random integer
The count or percentage of their occurrence in the dataset
The mean of the target variable
Their alphabetical rank
The function pd.get_dummies() in the Pandas library is specifically used for:
Normalization
One-Hot Encoding
Missing Value Imputation
Label Encoding
Which of the following is an example of a Nominal variable (no intrinsic order)?
Temperature (0°C, 10°C, 20°C)
T-shirt size (S, M, L)
City (Paris, London, Tokyo)
Salary ($50k, $60k, $70k)
Feature Engineering is the process of:
Cleaning missing data
Selecting the best model
Creating new features from existing data to improve model performance
Visualizing data
"Binning" allows you to convert:
Images into Vectors
Text into Numbers
Categorical variables into Numerical
Numerical variables into Categorical or Intervals
Extracting "Day of Week" from a "Date" column is an example of:
Data Imputation
Feature Extraction/Engineering
Feature Scaling
Dimensionality Reduction
Which of the following is an "Interaction Feature"?
log(x1)
x1^2
x1 + x2
x1 * x2
Polynomial features are created to:
Capture non-linear relationships in the data
Reduce the number of features
Remove outliers
Normalize the data
Calculating the length of a text string to use as a predictor is:
Clustering
Data Cleaning
Feature Engineering
Data Leakage
"Domain Knowledge" is useful in feature engineering because:
It is not useful
It helps identify irrelevant features and construct meaningful ones
It allows you to skip the testing phase
It is the only way to handle missing values
Creating a "Total Income" feature by adding "Applicant Income" and "Co-applicant Income" is:
Feature Selection
One-Hot Encoding
Feature Scaling
Feature Creation
The primary goal of feature engineering is to:
Make the data compatible with the algorithm and improve predictive power
Hide sensitive data
Make the dataset larger
Remove all categorical variables
The main goal of PCA (Principal Component Analysis) is to:
Impute missing values
Classify data into groups
Reduce the dimensionality of data while preserving variance
Increase the number of features
PCA is what type of learning algorithm?
Supervised
Semi-supervised
Reinforcement
Unsupervised
The new features created by PCA are called:
Support Vectors
Principal Components
Centroids
Decision Boundaries
In PCA, the first Principal component captures:
The mean of the data
The most variance
The noise only
The least variance
Principle components are always:
Larger than the original features
Categorical variables
Correlated with each other
Orthogonal( uncorrelated ) to each other
Before applying PCA, it is crucial to:
Convert all data to text
Scale/Standardize the data
Remove all negative numbers
Sort the data
If you reduce 100 features to 2 using PCA, you can easily:
E;iminate all bias
Visualize the data in a 2D scatter plot
Calculate the accuracy
Ignore the target variable
The "Explained Variance Ratio" tells you:
The accuracy of the model
How many clusters to pick
How much information (variance) each component holds
The number of missing values
PCA works best on variables that are:
Completely independent
Constant
Highly correlated
Categorical
The Pearson correlation coefficient ranges from:
0 to 100
-1 to 1
-10 to 10
0 to 1
A correlation of -0.9 indicates:
A strong positive relationship
No relationship
A weak negative relationship
A strong negative relationship
Why is high multicollinearity a problem for linear models?
It makes the data unreadable
It reduces predictive accuracy heavily
It makes the coefficient estimates unstable and hard to interpret
It causes the model to crash
Multicollinearity generally refers to:
Low variance in dataLow variance in data
High correlation between features and the target
Missing data in multiple columns
High correlation between two or more independent variables
Which metric is used to detect multicollinearity?
VIF
RMSE
AUC
R-squared
If two features have a correlation of 1.0, you should:
Square them
Remove one of them
Keep both
Multiply them together
Which correlation method works best for non-linear or ordinal relationships?
Euclidean Distance
Spearman (Rank Correlation)
Mean Correlation
Pearson
A correlation of 0 implies:
Perfect negative relationship
No linear relationship
An error in calculation
Strong linear relationship
Which tool is commonly used to visualize correlation matrices?
Pie Chart
Histogram
Heatmap
Boxplot
The primary purpose of EDA is to:
Deploy the model
Write the documentation
Understand the data structure, patterns, and anomalies
Build the final model immediately
df.describe() in Pandas provides:
Summary statistics (mean, std, min, max)
The first 5 rows of data
A correlation matrix
A histogram plot
Which plot is best for seeing the distribution of a single numerical variable?
Scatter Plot
Histogram
Scatter Matrix
Heatmap
A Scatter Plot is primarily used to show the relationship between:
Two categorical variables
Time and frequency
One categorical and one numerical variable
Two numerical variables
To visualize the count of items in different categories, use a:
Scatter Plot
Line Chart
Bar Chart
Boxplot
df.head() displays:
The column names only
The missing values
The correlation matrix
The first N rows of the dataset
A "Pair Plot" allows you to visualize:
Only missing data
Relationships between all pairs of variables in a dataset
Only the target variable
The time taken to train
Univariate analysis involves analyzing:
One variable at a time
Two variables at a time
Three variables at a time
The entire database
Skewness measures:
The height of the distribution
The flatness of the distribution
The asymmetry of the probability distribution
The number of outliers
Data Leakage occurs when:
Data is lost during transfer
The model performs poorly on training data
The database is hacked
Information from outside the training dataset (like the test set) is used to create the model
Which of the following is a common cause of data leakage?
Imputing using only the training set mean
Using Cross-Validation
Removing outliers
Imputing missing values using the mean of the entire dataset (train + test)
If your model has 99.9% accuracy on both training and test set immediately, you should suspect:
Your model is perfect
Underfitting
Data Leakage
Bad data
In time-series forecasting, using future data to predict the past is:
Feature Engineering
Look-ahead bias (Leakage)
Backpropagation
Lagging
To prevent leakage during scaling/normalization, you should:
fit_transform on the whole dataset
fit on the training set, then transform both training and test sets
fit on the test set only
Not scale the data
Including the target variable (or a proxy for it) as a feature:
Is a severe form of Data Leakage
Reduces overfitting
Is necessary for supervised learning
Is good feature engineering
ID columns (e.g., PatientID, TransactionID) should usually be:
Used as a primary feature
Squared
Used for normalization
Dropped, as they can cause leakage if they correlate with the target
When using Cross-Validation, feature selection should be done:
Inside the cross-validation loop (on the training folds only)
After the model is built
Before the cross-validation loop on the whole dataset
Never
Which is the “Golden Rule” to avoid leakage?
The test set should never be seen or touched during training/preprocessing
The test set should be added to the training set
Always use Deep Learning
Preprocessing is not important
Increasing model complexity usually causes:
Decrease bias and increase variance
Increase bias and decrease variance
Decrease both bias and variance
Increase bias and increase variance
A model with high bias typically:
Performs similarly on train and test, but both poorly
Fits training data extremely well
Has a large gap between train and test performance
Needs more data to generalize
High variance is most strongly associated with:
Class imbalance
Overfitting
Data leakage
Underfitting
If both training and validation errors are high, the model likely has:
Low bias
High variance
High bias
Low variance
A common way to reduce variance is to:
Use regularization or a simpler model
Add more features without regularization
Increase learning rate drastically
Reduce training data size
Which pattern indicates high variance?
Train error low, test error low
Train error high, test error high
Train error high, test error low
Train error low, test error high
Bagging (Bootstrap Aggregation) mainly reduces:
Neither
Bias
Variance
Both bias and variance equally
Boosting often helps reduce:
Noise only
Bias
Variance only
Feature scaling need
If model complexity increases and training error rises slightly while validation error falls, the best interpretation is:
Generalization is improving even if training fit changes
Data leakage is happening
Model is getting worse overall
Regularization is too weak
Underfitting means:. Underfitting means:
Test data is too small
Model is too simple to capture patterns
Features are redundant
Model captures noise
Overfitting occurs when:
Train accuracy is much higher than test accuracy
Both accuracies are low
Test accuracy is higher than train accuracy
Model is linear
Which action most directly combats underfitting?
Remove informative features
Add regularization
Reduce model complexity
Increase model complexity
A very deep decision tree is prone to:
Normalization errors
Underfitting
Class imbalance only
Overfitting
If validation loss starts rising while training loss keeps dropping, you should:
Increase learning rate
Train longer
Stop training or add regularization
Remove cross-validation
Early stopping helps mainly by reducing:
Variance
Dimensionality
Bias
Missing values
A degree-1 polynomial regression underfits curved data because:
It assumes independent errors
It requires scaling
It can’t model non-linearity
It has too many parameters
Dropout in neural networks mostly reduces:
Target leakage
Bias
Feature dependence
Variance
After increasing training data size, train error rises while test error drops. Best conclusion:
Data became noisier
Bias increased
Variance decreased
Model became worse
The main goal of a train–test split is to:
Increase variance
Improve training accuracy
Reduce bias
Estimate performance on unseen data
