WorksheetsMachine Learning Multiple Choice Questions
Total questions: 80
Worksheet time: 40mins
Which of the following is a type of Machine Learning?
Supervised
Unsupervised
Reinforcement
All of the above
Which of the following is not a Machine Learning algorithm?
Linear Regression
K-Means
Naive Bayes
Bubble Sort
Machine Learning is a subset of:
Deep Learning
Artificial Intelligence
Neural Networks
Robotics
Which of the following uses labeled data?
A. Unsupervised Learning
B. Supervised Learning
C. Reinforcement Learning
D. None
Which is an example of classification problem?
Predicting house prices
Diagnosing a disease as positive/negative
Finding groups in data
Reducing dimensions
Which term refers to the difference between the predicted and actual values?
Loss
Error
Cost
All of the above
Which of the following is used for dimensionality reduction?
PCA
KNN
SVM
Naive Bayes
The curse of dimensionality is related to:
Overfitting
Too many features
Overfitting occurs when:
Model performs well on test data
Model learns noise in training data
Model generalizes well
None of the above
Which of these is a linear model?
Decision Tree
SVM with RBF Kernel
Linear Regression
K-Means
Which of the following is a regression algorithm?
Logistic Regression
Linear Regression
Naive Bayes
KNN
Which algorithm can be used for both classification and regression?
KNN
Linear Regression
Naive Bayes
None
What is the output of a classification algorithm?
Continuous value
Discrete label
Which of the following is not used for classification?
SVM
Decision Tree
KNN
Linear Regression
What does the sigmoid function output range between?
-1 to 1
0 to 1
-∞ to ∞
1 to ∞
Which of these is sensitive to outliers?
K-Means
Decision Tree
Linear Regression
All of the above
In Logistic Regression, the output is:
Probability
Label
Class name
Integer
A confusion matrix is used in:
Clustering
Regression
Classification
All of the above
Which is not a classification metric?
Accuracy
Precision
R-Squared
Recall
Which ensemble method averages predictions?
Random Forest
Bagging
Boosting
Stacking
Which of these is an unsupervised learning task?
Classification
Regression
Clustering
Linear Regression
K-Means algorithm is sensitive to:
Initial centroids
Data scaling
Outliers
All of the above
Which technique is commonly used for market segmentation?
PCA
Clustering
Regression
Classification
What is the value of K in K-Means?
Maximum iterations
Number of clusters
Number of features
None
Which is a clustering algorithm?
Naive Bayes
DBSCAN
Linear Regression
Logistic Regression
What does ROC curve represent?
Recall vs. Precision
Accuracy vs. Time
TPR vs. FPR
Loss vs. Epoch
Which score is best for imbalanced datasets?
Accuracy
F1-Score
Precision
Specificity
Which evaluation metric is used for regression tasks?
Precision
MSE
Recall
AUC
R-Squared value indicates:
Feature importance
Variance explained
Model complexity
Dataset size
Which is not a loss function?
Cross-Entropy
Mean Squared Error
Gini Index
Hinge Loss
Neural networks are inspired by:
Heart
Lungs
Brain
Bones
Which activation function is used in hidden layers?
Sigmoid
ReLU
Softmax
Which optimizer is widely used in deep learning?
Gradient Descent
AdaGrad
Adam
Newton-Raphson
What is dropout used for?
Adding layers
Regularization
Removing bias
Decreasing learning rate
Convolutional Neural Networks are used for:
Text data
Image data
Tabular data
Audio only
What is Transfer Learning?
Transferring models
Using pre-trained models
Transferring data
Sharing parameters
Which technique is used in NLP?
CNN
LSTM
GAN
DBSCAN
Reinforcement Learning involves:
Rewards and penalties
Labeled data
Unlabeled data
Clusters
Which is not a feature selection technique?
A. Chi-Square
B. PCA
Which of the following is used in anomaly detection?
KNN
Isolation Forest
Random Forest
Logistic Regression
What is the purpose of feature scaling?
Reduce number of features
Increase training data
Standardize feature range
Decrease model accuracy
Which of these is a common scaling technique?
One-hot encoding
Label encoding
Min-Max Scaling
Binning
Which technique is used for handling missing data?
Data augmentation
Imputation
Feature reduction
Scaling
One-hot encoding is used for:
Numerical features
Text features
Categorical features
Continuous labels
Which of the following can cause data leakage?
Train-test split
Using future data during training
Feature selection
Normalization
What does label encoding do?
Normalizes numeric values
Converts categories to integers
Drops irrelevant features
Which of these techniques reduces overfitting?
Adding features
Increasing epochs
Regularization
Using noisy data
Which regularization technique uses L1 norm?
Ridge
Lasso
Elastic Net
Batch Norm
Which term describes input variables in a dataset?
Labels
Targets
Features
Errors
Which of the following splits data into training and test sets?
KMeans
train_test_split
OneHotEncoder
StandardScaler
Which algorithm is best suited for non-linear decision boundaries?
Logistic Regression
KNN
SVM with RBF Kernel
Linear Regression
What does ‘K’ in KNN represent?
A. Number of classes
B. Number of features
C. Number of neighbors
D. Kernel used
Naive Bayes assumes:
Feature independence
Linear relationships
Non-linear decision boundaries
Clustering nature
Decision Trees split data based on:
Mean
Variance
Gini or Entropy
Mode
Which algorithm is prone to overfitting?
Linear Regression
Decision Tree
Ridge Regression
Naive Bayes
What is an ensemble model?
Single model
Model trained on time series
Combination of multiple models
Reinforcement model
Which of these is a boosting algorithm?
Bagging
Random Forest
AdaBoost
KMeans
Gradient Boosting works by:
Reducing bias iteratively
Voting
Clustering data
Penalizing error
In Random Forest, each tree is trained on:
Same data
Different subset
Random noise
Which algorithm is good for high-dimensional data?
KNN
Decision Tree
SVM
Linear Regression
Which ML task is used in spam filtering?
Clustering
Classification
Regression
Reinforcement
Recommendation systems use:
Regression
Classification
Collaborative filtering
Reinforcement learning only
What is the primary goal of unsupervised learning?
Predict output
Classify labels
Find structure in data
Calculate loss
Which of the following best describes underfitting?
Model fits noise
Model is too complex
Model misses patterns
Model generalizes well
A model performs well on training data but poorly on test data. It is:
Underfitted
Overfitted
Regularized
Accurate
AUC stands for:
Area Under Curve
Average Under Class
Which ML type is used in robotics for learning from feedback?
Supervised
Unsupervised
Reinforcement
Semi-Supervised
What is the full form of SVM?
Support Vector Machine
Sample Vector Model
Supervised Variance Model
Statistical Vector Map
Which of the following is used to avoid overfitting in neural networks?
Increasing layers
Batch normalization
Dropout
Both B and C
What is the role of learning rate in training?
Measures accuracy
Sets data split
Controls weight updates
Defines model size
Which Python library is commonly used for ML?
NumPy
Pandas
Scikit-learn
Flask
Which function is used to train a model in scikit-learn?
model.run()
model.train()
model.fit()
model.predict()
Which language is most used in Machine Learning?
(a)
TensorFlow is developed by:
Microsoft
OpenAI
Which of the following is not a framework?
PyTorch
TensorFlow
NumPy
Keras
The train-test split ratio commonly used is:
50:50
60:40
80:20
95:5
Which file format is commonly used for datasets?
.doc
.csv
.exe
Which is used for text classification?
CNN
LSTM
RNN
All of the above
Which method detects overfitting during training?
Validation loss
Training accuracy
Feature scaling
Epochs
Which of the following models is most interpretable?
Neural Networks
Random Forest
Decision Tree
XGBoost
