Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Chapter 5_Introduction to Machine Learning

Total questions: 55

Worksheet time: 28mins

Name
Class
Date
1.

Which statement best captures Arthur Samuel’s 1959 definition of machine learning?

a)

Programs that follow fixed rules to solve tasks

b)

Hardware that accelerates numerical computations for AI

c)

A study enabling computers to learn without explicit programming

d)

Systems that only store and retrieve predefined answers

2.

Which contribution is most closely associated with Alan Turing in the history of machine learning?

a)

Coining the term machine learning in academic literature

b)

Inventing the backpropagation algorithm for neural nets

c)

Introducing a test for machine intelligence called the Turing Test

d)

Creating the first commercial expert system for medicine

3.

Which key idea describes how machine learning systems improve over time?

a)

They become better as they see more data

b)

They reset performance after each prediction

c)

They stabilize when rules are hard-coded

d)

They degrade as they use more memory

4.

Which option best differentiates supervised, unsupervised, and reinforcement learning at a high level?

a)

All use labeled data for training

b)

Reinforcement learns from static datasets without interaction

c)

Supervised uses labels, unsupervised finds structure, reinforcement optimizes actions via rewards

d)

Unsupervised uses labels, supervised discovers clusters, reinforcement ignores feedback

5.

Which outcome aligns with interpreting a confusion matrix for model evaluation?

a)

Determining hardware requirements for training

b)

Calculating precision, recall, and F1-score

c)

Selecting the best optimizer for gradient descent

d)

Designing new features using domain knowledge

6.

A checker-playing program that improves by learning from experience exemplifies which principle?

a)

Rule-based symbolic reasoning

b)

Data-driven performance improvement over time

c)

Pattern identification without updates

d)

Manual tuning of fixed strategies

7.

Which statement best defines machine learning in practical terms?

a)

Systems that identify patterns to make predictions or decisions

b)

Databases that store data for later retrieval only

c)

Tools that execute prewritten instructions without adaptation

d)

Compilers that translate code to machine instructions

8.

Which statement best contrasts traditional programming with machine learning?

a)

Rules and data produce coded outputs

b)

Data and rules yield fixed procedures

c)

Data and answers create learned rules

d)

Rules and answers generate training labels

9.

In supervised learning, what is the primary characteristic of the training data?

a)

Contains unlabeled examples for clustering

b)

Includes labeled inputs with known outputs

c)

Offers rewards but no explicit labels

d)

Provides raw features without any structure

10.

Which task is most appropriate for unsupervised learning?

a)

Grouping customers by similar behaviors

b)

Estimating student grades from history

c)

Classifying emails as spam or not spam

d)

Predicting housing prices from features

11.

Reinforcement learning primarily trains an agent by which mechanism?

a)

Supervising with labeled target values

b)

Discovering hidden clusters in data

c)

Receiving feedback through rewards

d)

Fitting curves to continuous variables

12.

Which scenario best matches reinforcement learning rather than supervised learning?

a)

Predicting stock prices from features

b)

Detecting product review sentiment

c)

Learning optimal moves to win a game

d)

Diagnosing diseases from labeled scans

13.

For linear regression with one feature x, which equation is used to model predictions?

a)

y=ax0+ax1x+ax2x2+b

b)

y=mx+b with slope and intercept

c)

y=m1x1+m2x2+m3x3+b

d)

y=a0+a1x+a2x2+...+anxn

14.

Which term in linear regression corresponds to the parameter controlling the line’s steepness?

a)

Feature X representing input value

b)

Intercept b determining vertical shift

c)

Slope m adjusting gradient angle

d)

Loss MSE measuring average error

15.

Why might polynomial regression be preferred over linear regression for some problems?

a)

It requires fewer model parameters

b)

It eliminates the need for targets

c)

It guarantees zero training error

d)

It captures curved relationships

16.

Which metric reports the average magnitude of errors in the same units as the regression output?

a)

R2 Score reports variance proportion only

b)

Mean Squared Error reports squared error units

c)

Mean Absolute Error reports average absolute differences

d)

Accuracy reports proportion of correct predictions

17.

In regression, which metric penalizes larger errors more heavily due to squaring?

a)

Mean Squared Error penalizes large errors more

b)

Accuracy penalizes both false positives equally

c)

R2 Score penalizes only small residuals

d)

Mean Absolute Error penalizes squared residuals

18.

What is Root Mean Squared Error primarily valued for when comparing models?

a)

It is undefined for continuous targets

b)

It measures proportion of variance explained

c)

It has the same units as the output

d)

It gives percentages of correct classifications

19.

Which statement best describes the R2 (coefficient of determination) in regression?

a)

Average of absolute residual differences

b)

Average of squared residual differences

c)

Proportion of variance explained by the model

d)

Square root of mean squared error values

20.

A classifier’s accuracy is defined as which proportion?

a)

Total correct predictions over all predictions

b)

True positives over actual positives

c)

Harmonic mean of precision and recall

d)

True positives over predicted positives

21.

Precision answers which question about a classifier’s predictions?

a)

How many predictions were correct overall

b)

How many actual positives were correctly found

c)

How many predicted positives are truly positive

d)

How strongly the model explains variance

22.

Recall (sensitivity) focuses on which aspect of classification performance?

a)

Average magnitude of errors in regression

b)

Proportion of predicted positives that are correct

c)

Proportion of actual positives correctly detected

d)

Balance between precision and recall performance

23.

The F1 score is defined as which type of mean of precision and recall?

a)

Quadratic mean of the two metrics

b)

Geometric mean of the two metrics

c)

Harmonic mean of the two metrics

d)

Arithmetic mean of the two metrics

24.

In a binary email classifier, the confusion matrix shows TP=70, FP=10, FN=30, TN=90. Which term describes emails correctly identified as spam?

a)

True Negatives, correctly labeled trusted

b)

False Negatives, missed spam emails

c)

False Positives, incorrectly flagged spam

d)

True Positives, correctly labeled spam

25.

Using TP=70, FP=10, FN=30, TN=90 for the email classifier, what is the overall accuracy?

a)

0.50 using (TP+FN)/total

b)

0.70 using (TP+FP)/total

c)

0.80 using (TP+TN)/total

d)

0.65 using (TP+TN)/predicted

26.

In the fruit multiclass confusion matrix: Predicted A against Actual A,B,C is 3,2,1; Predicted B against Actual A,B,C is 2,3,2; Predicted C against Actual A,B,C is 1,1,2. What is the number of correct predictions across all classes?

a)

8 from diagonal 3+3+2

b)

7 from diagonal 3+2+2

c)

9 from diagonal 4+3+2

d)

6 from diagonal 2+2+2

27.

For the same fruit matrix, which entry represents a banana predicted as cherry?

a)

Predicted A, Actual B = 2

b)

Predicted C, Actual A = 1

c)

Predicted C, Actual B = 1

d)

Predicted B, Actual C = 2

28.

With totals in the fruit matrix equal to 18 samples, what is the accuracy computed from the matrix?

a)

0.333 from 6 correct over 18

b)

0.600 from 12 correct over 20

c)

0.444 from 8 correct over 18

d)

0.500 from 9 correct over 18

29.

A model designer wants to reduce false positives in the email spam example. Which change targets this specifically without changing class definitions?

a)

Ignore true negatives entirely

b)

Raise decision threshold for spam

c)

Lower decision threshold for spam

d)

Increase dataset class imbalance

30.

Which stage primarily tunes hyperparameters and helps select the best version of a model?

a)

Training builds internal logic

b)

Deployment serves predictions

c)

Testing measures performance

d)

Validation tunes hyperparameters

31.

A model shows high training accuracy but low validation accuracy. What issue is most likely present?

a)

Class imbalance in labels

b)

Data leakage in test set

c)

Overfitting to training data

d)

Underfitting due to simple model

32.

Which validation approach provides a more reliable estimate of performance by rotating the validation fold across k splits?

a)

Stratified holdout split

b)

Time-series split only

c)

Bootstrapped resampling

d)

K-Fold cross-validation

33.

In classification tasks, what type of outputs are predicted?

a)

Confidence intervals only

b)

Ranked similarity scores

c)

Discrete class labels

d)

Continuous numeric values

34.

You observe low training and validation accuracy. Which adjustment is the most appropriate first step?

a)

Reduce features and depth

b)

Increase model complexity

c)

Collect more validation data

d)

Lower regularization strength

35.

Which pair correctly matches model family with both classification and regression variants?

a)

KNeighborsClassifier and KNeighborsRegressor

b)

RandomForestClassifier and DecisionTreeRegressor

c)

DecisionTreeClassifier and KNeighborsRegressor

d)

DecisionTreeRegressor and RandomForestRegressor

36.

For a quick baseline, you split data into 80% training and 20% validation. What is the main limitation of this method?

a)

Performance depends on the split

b)

Prevents model generalization

c)

Requires complex hyperparameters

d)

Ignores training set labels

37.

Which stage evaluates a final model on unseen data to measure real‑world performance?

a)

Testing stage only

b)

Training stage only

c)

Validation stage only

d)

Deployment stage only

38.

Which model family lists SVC for classification and SVR for regression?

a)

Support Vector Machines family

b)

Neural Networks family

c)

Gradient Boosting family

d)

Naïve Bayes family

39.

Which statement best describes neural networks variants for classification and regression?

a)

Depend on architecture and loss

b)

Always use SVC and SVR

c)

Use GradientBoostingClassifier only

d)

Assume discrete class probabilities

40.

Which task-specific model is correctly matched to task type?

a)

Logistic Regression — Regression

b)

Poisson Regression — Classification

c)

Linear Regression — Classification

d)

Naïve Bayes — Classification

41.

What is the correct note for logistic regression in this context?

a)

Despite its name, for categorical output only

b)

Predict continuous values only for outputs

c)

For count data only in practice

d)

Assumes discrete class probabilities always

42.

What is the primary goal of clustering in unsupervised learning?

a)

Identify rare or unusual data points

b)

Discover relationships in large datasets

c)

Reduce number of input variables

d)

Group similar data points together

43.

Which unsupervised task focuses on reducing input variables while preserving important information?

a)

Dimensionality Reduction task

b)

Association Rule Learning task

c)

Manifold Learning task

d)

Anomaly Detection task

44.

Which models are commonly used for dimensionality reduction?

a)

Isolation Forest, One-Class SVM

b)

Apriori, Eclat algorithms

c)

K-Means, DBSCAN, Hierarchical

d)

PCA, t-SNE, Autoencoders

45.

Which use case aligns with association rule learning?

a)

Noise reduction pipelines

b)

Image compression tasks

c)

Fraud detection workflows

d)

Market basket analysis use

46.

Which Python import correctly brings in the linear regression estimator used in scikit-learn?

a)

import linear_model.LinearRegression from sklearn

b)

from sklearn.model import LinearRegressionClass

c)

import sklearn.linear_model as LinearRegression

d)

from sklearn.linear_model import LinearRegression

47.

Given house sizes X = [[500],[750],[1000],[1250],[1500]] and prices y = [150000, 200000, 250000, 300000, 350000], which statement correctly trains the model?

a)

model = LinearRegression(); model.fit(predictions, y)

b)

model = LinearRegression(); model.fit(X, y)

c)

model = LinearRegression(); model.fit(y, X)

d)

model = LinearRegression(); model.fit(X, predictions)

48.

After training a LinearRegression model named model on X and y, which code generates predicted prices for the same inputs?

a)

predictions = model.predict(y)

b)

predictions = model.transform(y)

c)

predictions = model.predict(X)

d)

predictions = model.score(X)

49.

You want a plot with actual points in blue and the regression line in red. Which pair of matplotlib calls best achieves this?

a)

plt.plot(y, X, color='blue'); plt.scatter(predictions, X, color='red')

b)

plt.scatter(y, X, color='blue'); plt.plot(predictions, y, color='red')

c)

plt.scatter(X, y, color='blue'); plt.plot(X, predictions, color='red')

d)

plt.bar(X, y, color='blue'); plt.plot(y, predictions, color='red')

50.

Which function in scikit-learn expands input features to include polynomial terms for regression?

a)

LinearRegression from linear_model module

b)

confusion_matrix from metrics module

c)

PolynomialFeatures from preprocessing module

d)

mean_squared_error from metrics module

51.

In the provided workflow, what is the correct order to create a polynomial regression prediction pipeline?

a)

Predict with LinearRegression, then transform with PolynomialFeatures

b)

Transform with PolynomialFeatures, fit LinearRegression, predict

c)

Visualize with matplotlib, then fit LinearRegression, transform

d)

Fit LinearRegression, transform with PolynomialFeatures, predict

52.

Given y_true = [1, 0, 1, 1, 0, 1, 0] and y_pred = [1, 0, 1, 0, 0, 1, 1], which classification metric counts true positives divided by predicted positives?

a)

Precision score for positive predictive value

b)

Recall score for sensitivity measure

c)

F1 score for harmonic mean of rates

d)

Accuracy score for overall correctness

53.

Which plot combination best visualizes actual data points and the polynomial fit curve?

a)

plt.scatter for points and plt.bar for curve

b)

plt.bar for points and plt.plot for curve

c)

plt.plot for points and plt.scatter for curve

d)

plt.scatter for points and plt.plot for curve

54.

You transform X with PolynomialFeatures(degree=2) and fit LinearRegression. To predict on a new X_range, what must you do first?

a)

Normalize X_range using accuracy_score

b)

Transform X_range using the same PolynomialFeatures

c)

Compute confusion_matrix before prediction

d)

Fit PolynomialFeatures again on y values

55.

For regression metrics with y_true = [150000, 200000, 250000] and y_pred = [140000, 210000, 260000], which metric is computed as the square root of MSE?

a)

MAE the average absolute error magnitude

b)

RMSE the root mean squared error value

c)

MSE the mean squared error measure

d)

R2 the proportion of variance explained