WorksheetsChapter 5_Introduction to Machine Learning
Total questions: 55
Worksheet time: 28mins
Which statement best captures Arthur Samuel’s 1959 definition of machine learning?
Programs that follow fixed rules to solve tasks
Hardware that accelerates numerical computations for AI
A study enabling computers to learn without explicit programming
Systems that only store and retrieve predefined answers
Which contribution is most closely associated with Alan Turing in the history of machine learning?
Coining the term machine learning in academic literature
Inventing the backpropagation algorithm for neural nets
Introducing a test for machine intelligence called the Turing Test
Creating the first commercial expert system for medicine
Which key idea describes how machine learning systems improve over time?
They become better as they see more data
They reset performance after each prediction
They stabilize when rules are hard-coded
They degrade as they use more memory
Which option best differentiates supervised, unsupervised, and reinforcement learning at a high level?
All use labeled data for training
Reinforcement learns from static datasets without interaction
Supervised uses labels, unsupervised finds structure, reinforcement optimizes actions via rewards
Unsupervised uses labels, supervised discovers clusters, reinforcement ignores feedback
Which outcome aligns with interpreting a confusion matrix for model evaluation?
Determining hardware requirements for training
Calculating precision, recall, and F1-score
Selecting the best optimizer for gradient descent
Designing new features using domain knowledge
A checker-playing program that improves by learning from experience exemplifies which principle?
Rule-based symbolic reasoning
Data-driven performance improvement over time
Pattern identification without updates
Manual tuning of fixed strategies
Which statement best defines machine learning in practical terms?
Systems that identify patterns to make predictions or decisions
Databases that store data for later retrieval only
Tools that execute prewritten instructions without adaptation
Compilers that translate code to machine instructions
Which statement best contrasts traditional programming with machine learning?
Rules and data produce coded outputs
Data and rules yield fixed procedures
Data and answers create learned rules
Rules and answers generate training labels
In supervised learning, what is the primary characteristic of the training data?
Contains unlabeled examples for clustering
Includes labeled inputs with known outputs
Offers rewards but no explicit labels
Provides raw features without any structure
Which task is most appropriate for unsupervised learning?
Grouping customers by similar behaviors
Estimating student grades from history
Classifying emails as spam or not spam
Predicting housing prices from features
Reinforcement learning primarily trains an agent by which mechanism?
Supervising with labeled target values
Discovering hidden clusters in data
Receiving feedback through rewards
Fitting curves to continuous variables
Which scenario best matches reinforcement learning rather than supervised learning?
Predicting stock prices from features
Detecting product review sentiment
Learning optimal moves to win a game
Diagnosing diseases from labeled scans
For linear regression with one feature x, which equation is used to model predictions?
y=ax0+ax1x+ax2x2+b
y=mx+b with slope and intercept
y=m1x1+m2x2+m3x3+b
y=a0+a1x+a2x2+...+anxn
Which term in linear regression corresponds to the parameter controlling the line’s steepness?
Feature X representing input value
Intercept b determining vertical shift
Slope m adjusting gradient angle
Loss MSE measuring average error
Why might polynomial regression be preferred over linear regression for some problems?
It requires fewer model parameters
It eliminates the need for targets
It guarantees zero training error
It captures curved relationships
Which metric reports the average magnitude of errors in the same units as the regression output?
R2 Score reports variance proportion only
Mean Squared Error reports squared error units
Mean Absolute Error reports average absolute differences
Accuracy reports proportion of correct predictions
In regression, which metric penalizes larger errors more heavily due to squaring?
Mean Squared Error penalizes large errors more
Accuracy penalizes both false positives equally
R2 Score penalizes only small residuals
Mean Absolute Error penalizes squared residuals
What is Root Mean Squared Error primarily valued for when comparing models?
It is undefined for continuous targets
It measures proportion of variance explained
It has the same units as the output
It gives percentages of correct classifications
Which statement best describes the R2 (coefficient of determination) in regression?
Average of absolute residual differences
Average of squared residual differences
Proportion of variance explained by the model
Square root of mean squared error values
A classifier’s accuracy is defined as which proportion?
Total correct predictions over all predictions
True positives over actual positives
Harmonic mean of precision and recall
True positives over predicted positives
Precision answers which question about a classifier’s predictions?
How many predictions were correct overall
How many actual positives were correctly found
How many predicted positives are truly positive
How strongly the model explains variance
Recall (sensitivity) focuses on which aspect of classification performance?
Average magnitude of errors in regression
Proportion of predicted positives that are correct
Proportion of actual positives correctly detected
Balance between precision and recall performance
The F1 score is defined as which type of mean of precision and recall?
Quadratic mean of the two metrics
Geometric mean of the two metrics
Harmonic mean of the two metrics
Arithmetic mean of the two metrics
In a binary email classifier, the confusion matrix shows TP=70, FP=10, FN=30, TN=90. Which term describes emails correctly identified as spam?
True Negatives, correctly labeled trusted
False Negatives, missed spam emails
False Positives, incorrectly flagged spam
True Positives, correctly labeled spam
Using TP=70, FP=10, FN=30, TN=90 for the email classifier, what is the overall accuracy?
0.50 using (TP+FN)/total
0.70 using (TP+FP)/total
0.80 using (TP+TN)/total
0.65 using (TP+TN)/predicted
In the fruit multiclass confusion matrix: Predicted A against Actual A,B,C is 3,2,1; Predicted B against Actual A,B,C is 2,3,2; Predicted C against Actual A,B,C is 1,1,2. What is the number of correct predictions across all classes?
8 from diagonal 3+3+2
7 from diagonal 3+2+2
9 from diagonal 4+3+2
6 from diagonal 2+2+2
For the same fruit matrix, which entry represents a banana predicted as cherry?
Predicted A, Actual B = 2
Predicted C, Actual A = 1
Predicted C, Actual B = 1
Predicted B, Actual C = 2
With totals in the fruit matrix equal to 18 samples, what is the accuracy computed from the matrix?
0.333 from 6 correct over 18
0.600 from 12 correct over 20
0.444 from 8 correct over 18
0.500 from 9 correct over 18
A model designer wants to reduce false positives in the email spam example. Which change targets this specifically without changing class definitions?
Ignore true negatives entirely
Raise decision threshold for spam
Lower decision threshold for spam
Increase dataset class imbalance
Which stage primarily tunes hyperparameters and helps select the best version of a model?
Training builds internal logic
Deployment serves predictions
Testing measures performance
Validation tunes hyperparameters
A model shows high training accuracy but low validation accuracy. What issue is most likely present?
Class imbalance in labels
Data leakage in test set
Overfitting to training data
Underfitting due to simple model
Which validation approach provides a more reliable estimate of performance by rotating the validation fold across k splits?
Stratified holdout split
Time-series split only
Bootstrapped resampling
K-Fold cross-validation
In classification tasks, what type of outputs are predicted?
Confidence intervals only
Ranked similarity scores
Discrete class labels
Continuous numeric values
You observe low training and validation accuracy. Which adjustment is the most appropriate first step?
Reduce features and depth
Increase model complexity
Collect more validation data
Lower regularization strength
Which pair correctly matches model family with both classification and regression variants?
KNeighborsClassifier and KNeighborsRegressor
RandomForestClassifier and DecisionTreeRegressor
DecisionTreeClassifier and KNeighborsRegressor
DecisionTreeRegressor and RandomForestRegressor
For a quick baseline, you split data into 80% training and 20% validation. What is the main limitation of this method?
Performance depends on the split
Prevents model generalization
Requires complex hyperparameters
Ignores training set labels
Which stage evaluates a final model on unseen data to measure real‑world performance?
Testing stage only
Training stage only
Validation stage only
Deployment stage only
Which model family lists SVC for classification and SVR for regression?
Support Vector Machines family
Neural Networks family
Gradient Boosting family
Naïve Bayes family
Which statement best describes neural networks variants for classification and regression?
Depend on architecture and loss
Always use SVC and SVR
Use GradientBoostingClassifier only
Assume discrete class probabilities
Which task-specific model is correctly matched to task type?
Logistic Regression — Regression
Poisson Regression — Classification
Linear Regression — Classification
Naïve Bayes — Classification
What is the correct note for logistic regression in this context?
Despite its name, for categorical output only
Predict continuous values only for outputs
For count data only in practice
Assumes discrete class probabilities always
What is the primary goal of clustering in unsupervised learning?
Identify rare or unusual data points
Discover relationships in large datasets
Reduce number of input variables
Group similar data points together
Which unsupervised task focuses on reducing input variables while preserving important information?
Dimensionality Reduction task
Association Rule Learning task
Manifold Learning task
Anomaly Detection task
Which models are commonly used for dimensionality reduction?
Isolation Forest, One-Class SVM
Apriori, Eclat algorithms
K-Means, DBSCAN, Hierarchical
PCA, t-SNE, Autoencoders
Which use case aligns with association rule learning?
Noise reduction pipelines
Image compression tasks
Fraud detection workflows
Market basket analysis use
Which Python import correctly brings in the linear regression estimator used in scikit-learn?
import linear_model.LinearRegression from sklearn
from sklearn.model import LinearRegressionClass
import sklearn.linear_model as LinearRegression
from sklearn.linear_model import LinearRegression
Given house sizes X = [[500],[750],[1000],[1250],[1500]] and prices y = [150000, 200000, 250000, 300000, 350000], which statement correctly trains the model?
model = LinearRegression(); model.fit(predictions, y)
model = LinearRegression(); model.fit(X, y)
model = LinearRegression(); model.fit(y, X)
model = LinearRegression(); model.fit(X, predictions)
After training a LinearRegression model named model on X and y, which code generates predicted prices for the same inputs?
predictions = model.predict(y)
predictions = model.transform(y)
predictions = model.predict(X)
predictions = model.score(X)
You want a plot with actual points in blue and the regression line in red. Which pair of matplotlib calls best achieves this?
plt.plot(y, X, color='blue'); plt.scatter(predictions, X, color='red')
plt.scatter(y, X, color='blue'); plt.plot(predictions, y, color='red')
plt.scatter(X, y, color='blue'); plt.plot(X, predictions, color='red')
plt.bar(X, y, color='blue'); plt.plot(y, predictions, color='red')
Which function in scikit-learn expands input features to include polynomial terms for regression?
LinearRegression from linear_model module
confusion_matrix from metrics module
PolynomialFeatures from preprocessing module
mean_squared_error from metrics module
In the provided workflow, what is the correct order to create a polynomial regression prediction pipeline?
Predict with LinearRegression, then transform with PolynomialFeatures
Transform with PolynomialFeatures, fit LinearRegression, predict
Visualize with matplotlib, then fit LinearRegression, transform
Fit LinearRegression, transform with PolynomialFeatures, predict
Given y_true = [1, 0, 1, 1, 0, 1, 0] and y_pred = [1, 0, 1, 0, 0, 1, 1], which classification metric counts true positives divided by predicted positives?
Precision score for positive predictive value
Recall score for sensitivity measure
F1 score for harmonic mean of rates
Accuracy score for overall correctness
Which plot combination best visualizes actual data points and the polynomial fit curve?
plt.scatter for points and plt.bar for curve
plt.bar for points and plt.plot for curve
plt.plot for points and plt.scatter for curve
plt.scatter for points and plt.plot for curve
You transform X with PolynomialFeatures(degree=2) and fit LinearRegression. To predict on a new X_range, what must you do first?
Normalize X_range using accuracy_score
Transform X_range using the same PolynomialFeatures
Compute confusion_matrix before prediction
Fit PolynomialFeatures again on y values
For regression metrics with y_true = [150000, 200000, 250000] and y_pred = [140000, 210000, 260000], which metric is computed as the square root of MSE?
MAE the average absolute error magnitude
RMSE the root mean squared error value
MSE the mean squared error measure
R2 the proportion of variance explained
