WorksheetsMachine Learning and Linear Algebra Quiz
Total questions: 20
Worksheet time: 10mins
Which of the following best describes the experience E in supervised learning?
A model trained without any label
A dataset containing only input features
A dataset of (input, label) pairs used for learning
The performance metric used to evaluate models
In the context of ML model evaluation, which metric is not derived from a confusion matrix?
F1 Score
Precision
Recall
Log-likelihood
Empirical risk is minimized during training. What does it represent?
The model's confidence in predictions
The average loss on the training data
The regularization term added to prevent overfitting
The variance of the prediction errors
Which of the following is not an axiom of probability?
Non-negativity
Independence
Total probability equals 1
Additivity
Bayes' theorem can be used to:
Estimate model weights
Convert prior probabilities to posterior probabilities
Normalize feature vectors
Find eigenvectors of the covariance matrix
The matrix A is invertible if:
Its rank is less than n
Its determinant is zero
Its null space is non-trivial
There exists a matrix A⁻¹ such that AA⁻¹=I
Which of the following is true about the dot product u⋅v?
It's the sum of the element-wise division
It equals the L2 norm of the vector
It is zero if u and v are orthogonal
It can only be computed for square matrices
What does the rank of a matrix correspond to?
The number of zero rows
The dimension of its column space
The number of elements in the null space
The number of pivots in its transpose
The condition number of a matrix gives insight into:
The number of non-zero eigenvalues
How sparse the matrix is
The matrix’s numerical stability when solving Ax=b
The number of dimensions in its eigenspace
If a matrix A is symmetric and positive definite, then:
All its eigenvalues are negative
It has complex eigenvalues
It is guaranteed to be invertible
Its rank is always 1
Which supervised learning algorithm uses information gain to build its structure?
Logistic Regression
K-Nearest Neighbors
Decision Tree
Support Vector Machine
What does the entropy of a node in a decision tree measure?
The number of features selected
The variance in feature values
The uncertainty or impurity in the class distribution
The maximum likelihood of prediction
In K-Nearest Neighbors, increasing the value of k usually leads to:
Higher variance
Lower bias
Better performance on all datasets
Smoother decision boundaries
Which of the following statements about logistic regression is true?
It minimizes squared error
It produces multi-class outputs by default
It models the probability using a sigmoid function
It uses k-nearest neighbors for estimation
Support Vector Machines aim to:
Minimize classification error directly
Maximize the distance between support vectors
Maximize the geometric margin
Minimize the entropy of the decision boundary
In K-means clustering, the objective is to minimize:
The silhouette score
The variance between clusters
The within-cluster sum of squared distances
The information gain
What does the elbow method help determine?
The cluster initialization points
The number of relevant features
The optimal number of clusters K
The performance of agglomerative clustering
In PCA, the principal components are:
Randomly chosen from the dataset
Orthogonal vectors capturing maximum variance
Cluster centers from K-means
Linear regression coefficients
Which of the following is not a typical use case for PCA?
Dimensionality reduction
Feature decorrelation
Supervised classification
Data visualization
The eigenvectors of the covariance matrix in PCA:
Define the directions of least variance
Form a dependent basis for the data
Are ranked by the size of their corresponding eigenvalues
Are always complex-valued
