WorksheetsExtra Quiz for MIS 447 ML – Fall 2025-2026
Total questions: 21
Worksheet time: 32mins
You are given an imbalanced dataset for binary classification. Which metric provides a balanced view of performance by combining precision and recall?
Specificity directly replaces precision
Accuracy is sufficient for imbalanced datasets
F1-Score balances precision and recall
ROC AUC equals precision times recall
You must choose k for K-Means on high-dimensional data. What is a reasonable plan to select k and prepare features?
Pick k as number of classes in labels
Use elbow plot and apply PCA before K-Means
Set k to sqrt of samples always
Tune k by maximizing training accuracy
In the context of clustering, what is a common method to evaluate the quality of clusters formed by K-Means?
Silhouette Score
Mean Squared Error
Cross-Validation Score
Log-Likelihood
Which method is commonly used to reduce the dimensionality of data before applying clustering algorithms?
Decision Trees
Support Vector Machines
Linear Regression
Principal Component Analysis (PCA)
In the context of evaluating classification models, what does the ROC curve represent?
True Positive Rate vs False Positive Rate
F1-Score vs Specificity
Precision vs Recall
Accuracy vs Error Rate
In the context of model evaluation, what does the term 'confusion matrix' refer to?
A technique for dimensionality reduction
A method for visualizing data distributions
A graph showing model training progress
A table used to describe the performance of a classification model
What is the primary purpose of using cross-validation in model evaluation?
To visualize the performance of the model
To assess how the results of a statistical analysis will generalize to an independent dataset
To reduce the computational cost of training
To increase the size of the training dataset
Which metric is most suitable for evaluating the performance of a regression model?
Recall
Precision
Mean Squared Error
F1-Score
Which algorithm is typically used for supervised learning tasks?
Hierarchical Clustering
Principal Component Analysis
Linear Regression
K-Means Clustering
What is the purpose of using a validation set during model training?
To evaluate the model on unseen data only
To visualize the model's predictions
To increase the size of the training dataset
To tune hyperparameters and prevent overfitting
What is the main advantage of using Standardization for feature scaling in machine learning?
It ensures all features contribute equally to the distance calculations
It reduces the dimensionality of the dataset
It increases the complexity of the model
It eliminates the need for feature selection
What does the 'K' parameter represent in K-Means clustering?
The number of clusters to form
The maximum number of iterations
The number of features in the dataset
The distance metric used for clustering
What is the main benefit of using ElasticNet Regression over Lasso Regression?
It only uses L1 regularization
It combines L1 and L2 regularization to improve model performance
It is faster to compute than Lasso
It requires fewer hyperparameters to tune
What is the primary purpose of using the Elbow Method in K-Means clustering?
To reduce the dimensionality of the data
To visualize the clusters formed
To determine the optimal number of clusters
To assess the quality of the clustering
Which algorithm is particularly effective for handling high-dimensional data due to its ability to find a hyperplane that maximizes the margin between classes?
K-Nearest Neighbors (KNN)
Naive Bayes
Decision Trees
Support Vector Machine (SVM)
What is a key assumption made by the Naive Bayes classifier regarding the features used for classification?
Features are correlated
Features are independent given the class label
Features have a normal distribution
Features are equally important
What is the main purpose of the Yeo-Johnson transformation in data preprocessing?
To normalize data that may not follow a Gaussian distribution
To eliminate outliers from the dataset
To reduce the dimensionality of the dataset
To increase the variance of the dataset
In the context of ensemble learning, Bagging and Boosting use different mathematical strategies to reduce the total error. Which of the following correctly identifies the error type targeted by each and their training nature?
Bagging reduces Bias (parallel); Boosting reduces Variance (sequential).
Bagging reduces Variance (parallel); Boosting reduces Bias (sequential).
Bagging reduces Variance (sequential); Boosting reduces Bias (parallel).
Both reduce Bias, but Bagging uses weighted voting while Boosting uses simple averaging.
Unlike traditional Gradient Boosting Machines that grow trees level-by-level (Level-wise), LightGBM uses a Leaf-wise (Best-first) strategy. What is the primary operational difference of this strategy?
It splits all nodes at the same depth before moving to the next level to keep the tree balanced.
It only splits the root node and then stops to prevent complexity.
It chooses the leaf node that provides the maximum reduction in loss to split, regardless of its depth in the tree.
It uses a random search to decide which leaf should be expanded next.
XGBoost and Overfitting (Regularization) Why is XGBoost often better at "Generalization" (performing well on new data) compared to the original Gradient Boosting Machine (GBM)?
Because XGBoost is much slower, which gives it more time to learn.
Because XGBoost includes built-in L1 and L2 Regularization to punish complex models.
Because XGBoost removes all outliers from the dataset automatically.
Because XGBoost does not use Decision Trees.
Please, write your Student Id and Name.
