WorksheetsMachine Learning Quiz
Total questions: 20
Worksheet time: 10mins
What is precision in the context of machine learning metrics?
The ratio of true positives to the total number of predictions.
The ratio of true positives to the sum of true positives and false negatives.
The ratio of true positives to the sum of true positives and false positives.
The ratio of true negatives to the total number of predictions.
Which of the following is an advantage of ensemble methods?
They reduce the complexity of the models.
They improve the robustness and accuracy of predictions.
They require less data for training.
They eliminate the need for cross-validation.
In the voting ensemble method, the final prediction is made based on:
The majority vote of the base models.
The average of the predictions of the base models.
The weighted sum of the predictions of the base models.
All of the above.
Bagging helps in improving model performance by:
Reducing bias.
Reducing variance.
Increasing bias.
Increasing variance.
Which of the following is a boosting algorithm?
Random Forest
Gradient Boosting
K-Nearest Neighbors
Support Vector Machine
In stacking, the model that combines the predictions of base models is called:
Primary model
Secondary model
Meta model
Base model
A decision tree splits data at each node based on:
Random selection
The feature that gives the highest reduction in impurity
The feature with the most missing values
The feature with the least variance
Which measure is commonly used to evaluate the splits in a decision tree for classification tasks?
Variance
Mean Absolute Error
Gini Impurity
Euclidean Distance
In regression trees, the best split is chosen based on the reduction in:
Gini Impurity
Entropy
Variance
Accuracy
Random Forests improve the accuracy of predictions by:
Using a single decision tree
Combining multiple decision trees trained on random subsets of the data
Using linear regression
Ignoring certain features
Which metric is most appropriate to use when the dataset has a significant class imbalance?
Accuracy
Precision
Recall
F1-Score
Which ensemble method involves training multiple models on different subsets of the training data with replacement?
Voting
Bagging
Boosting
Stacking
In a weighted voting ensemble method, the model with the ________ accuracy typically has the highest weight.
Lowest
Highest
Average
Most recent
Which of the following is a common algorithm used in bagging?
Random Forest
AdaBoost
Gradient Boosting
Support Vector Machines
Boosting algorithms sequentially train models by giving more weight to ________.
Correctly classified samples
Misclassified samples
Random samples
Large samples
The meta-model in stacking is trained on:
The original dataset
The predictions of the base models
A subset of the features
The residual errors of the base models
In a decision tree, what is a leaf node?
A node that splits into further sub-nodes
A node that does not split further and represents a class label or value
The topmost node in the tree
A node with the highest impurity
Which splitting criterion is used by default in the Sklearn implementation of decision trees for classification?
Entropy
Information Gain
Gini Impurity
Variance Reduction
In the context of regression trees, which of the following measures is minimized to find the best split?
Gini Impurity
Mean Squared Error (MSE)
Cross-Entropy
Information Gain
Random Forests create diversity among individual trees by:
Using the entire dataset for each tree
Using different subsets of features and data for each tree
Training each tree sequentially
Using the same features for all trees
