NEW
Font size
WorksheetsMachine learning quiz
Total questions: 17
Worksheet time: 9mins
Q.1) Which of the following best describes the goal of linear regression?
A. Classify data into categories
B. Predict a continuous numerical value
C. Cluster data points into groups
D. Reduce the dimensionality of data
Q2 Decision trees split nodes based on which criterion to maximize separation of classes?
A. Euclidean distance
B. Gini impurity / Information gain
C. Gradient descent
D. K-distance
Q3 Which of the following is not an assumption of classical linear regression?
A. Linearity between predictors and response
B. Independence of errors
C. Homoscedasticity (constant variance of errors)
D. The predictors must all be integers
Q4 Which metric is not typically used to evaluate regression models?
A. Mean Absolute Error (MAE)
B. R-squared
C. Adjusted R-squared
D. F1-score
Q5 In a decision tree for classification, common impurity measures for splitting are:
A. Gini impurity and entropy
B. Euclidean distance and Manhattan distance
C. Mean squared error and R-squared
D. Gradient descent and backpropagation
Q6 Random Forest improves accuracy mainly by:
A. Using one deep tree on the whole dataset
B. Combining many trees trained on bootstrapped samples
C. Choosing medoids instead of centroids
D. Applying L1 regularization to each tree
Q7 Gradient Boosting differs from Bagging because it:
A. Builds trees in paralle
B. Builds trees sequentially, each correcting the previous one
C. Uses k-nearest neighbors for splits
D. Requires distance metrics
Q7 Which statement about decision trees is true?
A. They can only handle numerical data
B. They require feature scaling
C. They can naturally handle both categorical and numerical features
D. They cannot be used for regression
Q. Pruning a decision tree is primarily done to:
A. Increase training accuracy
B. Reduce overfitting and improve generalization
C. Make the tree deeper
D. Convert it into a random forest
Q. Which of the following is not a type of machine learning?
A. Supervised
B. Unsupervised
C. Reinforcement
D. Compilation
Q. K-Nearest Neighbors (KNN) is a:
A. Parametric model
B. Non-parametric model
C. Ensemble model
D. Dimensionality-reduction method
Q. Which of these algorithms is not distance-based?
A. K-Nearest Neighbors (KNN)
B. K-Means
C. Naïve Bayes
D. DBSCAN
Q. The most commonly used distance metric in Euclidean space is:
A. Cosine similarity
B. Manhattan distance
C. Euclidean distance
D. Hamming distance
Q. In KNN, the parameter k refers to:
A. Number of clusters
B. Number of neighbors considered when predicting a label
C. Number of features in the dataset
D. Number of decision trees
Q. Feature scaling (e.g., standardization) is critical for KNN because:
Q. Which voting scheme is common in KNN classification?
Q. K-Means clustering minimizes:
