Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Linear Regression Quiz

Total questions: 30

Worksheet time: 15mins

Name
Class
Date
1.

Who is credited with the quote, "All models are wrong, but some are useful," as mentioned in the learning material?

a)

Isaac Newton

b)

George Box

c)

Albert Einstein

d)

Carl Gauss

2.

What is the primary purpose of linear regression as described in the material?

a)

To classify data into categories

b)

To model the relationship between input and output variables using a linear function

c)

To cluster data points into groups

d)

To encrypt data for security

3.

According to the material, what must the output variables in linear regression be?

a)

Only binary values

b)

Only integer values

c)

Real or float values

d)

Only categorical values

4.

Given the equations y₁ = w₁₁x₁ + w₁₂x₂ + w₁₃x₃ + b₁ and y₂ = w₂₁x₁ + w₂₂x₂ + w₂₃x₃ + b₂, what does the term w₁₁ represent?

a)

The bias for the first output

b)

The weight connecting the first input to the first output

c)

The output variable

d)

The input variable

5.

What does Figure 4.1 in the material illustrate?

a)

The process of data encryption

b)

The relationship between input and output variables in linear regression

c)

The steps of clustering algorithms

d)

The architecture of a neural network for image recognition

6.

Why is house price prediction given as an example in the material?

a)

It is an example of predicting a categorical variable

b)

It is an example of predicting a float variable using input features

c)

It is an example of clustering data points

d)

It is an example of classifying images

7.

Which equation represents a linear regression single-layer network as described in the material?

a)

y = Wx + b

b)

y = Wx - b

c)

y = W + bx

d)

y = xW + b

8.

Which of the following best describes the objective function E as defined in the material?

a)

It measures how close the predicted value y is to the target value t over all training data.

b)

It calculates the sum of input variables.

c)

It determines the number of layers in the network.

d)

It finds the maximum value of the output variable.

9.

What does the Mean Square Error (MSE) loss function measure in linear regression?

a)

The number of input variables

b)

How well the linear model fits the data

c)

The number of output nodes

d)

The number of training samples

10.

Which formula represents the Mean Square Error (MSE) loss function for a set of training samples?

a)

E(W, b) = 1N∑n=1NEn\frac{1}{N} \sum_{n=1}^{N} E^n

b)

E(W, b) = ∑n=1NWnbn\sum_{n=1}^{N} W_n b_n

c)

E(W,b)=N∑n=1NEnE(W, b) = N \sum_{n=1}^{N} E^n

d)

E(W, b) = 1N∑n=1NWnbn\frac{1}{N} \sum_{n=1}^{N} W_n b_n

11.

What is the result of simplifying the cost function E(\theta) as shown in the document?

a)

E(θ)=θTXTXθ−2tTXθ+tTtE(\theta) = \theta^T X^T X \theta - 2 t^T X \theta + t^T t

b)

E(θ)=XTXθ+tTtE(\theta) = X^T X \theta + t^T t

c)

E(θ)=θTXTXθ+2tTXθE(\theta) = \theta^T X^T X \theta + 2 t^T X \theta

d)

E(θ)=θTXTXθ−tTXθE(\theta) = \theta^T X^T X \theta - t^T X \theta

12.

What is the first step to find the optimal value of θ in the context of the learning problem described?

a)

Take the first derivative of the cost function with respect to the parameters and set it to zero.

b)

Multiply the cost function by the parameters.

c)

Integrate the cost function with respect to the parameters.

d)

Set the cost function equal to one.

13.

Why is the analytical solution for finding the values of parameters W and b not always preferred for large training datasets?

a)

It does not scale well with the amount of training data.

b)

It is less accurate than other methods.

c)

It requires more memory than gradient descent.

d)

It is only applicable to non-linear models.

14.

What does the gradient of the MSE loss function with respect to the parameters W and b help to compute in the context of gradient descent?

a)

The direction and magnitude to update the parameters.

b)

The final value of the cost function.

c)

The number of training samples.

d)

The learning rate.

15.

In the gradient descent update rule for the weights and bias, what does the symbol η represent?

a)

The learning rate.

b)

The number of epochs.

c)

The cost function.

d)

The bias term.

16.

How does the mini-batch stochastic gradient variant help when scaling to large amounts of training data?

a)

It estimates the parameters of the linear regression model using subsets of the data.

b)

It increases the learning rate.

c)

It reduces the number of parameters.

d)

It eliminates the need for a cost function.

17.

Given the update rule wrst+1=wrst−η∂E(W,b)/∂wrsw_{rs}^{t+1} = w_{rs}^t - η ∂E(W, b)/∂w_{rs} , what is the purpose of subtracting the gradient term?

a)

To move the parameters in the direction that reduces the loss.

b)

To increase the value of the weights.

c)

To keep the weights constant.

d)

To maximize the loss function.

18.

What is one advantage of mini-batch gradient descent?

a)

It can handle large datasets that do not fit in memory

b)

It always converges faster than any other method

c)

It does not require tuning any hyperparameters

d)

It cannot exploit parallelism

19.

Why is tuning the batch size and learning rate important in mini-batch gradient descent?

a)

Because they affect the convergence and performance of the algorithm

b)

Because they determine the number of input features

c)

Because they are not used in the algorithm

d)

Because they only affect the output format

20.

Suppose you are given a dataset too large to fit in memory. Which gradient descent method would be most suitable?

a)

Mini-batch stochastic gradient descent

b)

Batch gradient descent

c)

Newton's method

d)

Random search

21.

In the context of the provided example, what do the first three columns of the training data represent?

a)

Input features (x1, x2, x3)

b)

Output targets (t1, t2)

c)

Batch sizes

d)

Learning rates

22.

Which of the following best describes the input and output variables in the synthesized data for linear regression modeling shown in Table 4.1?

a)

The input variables are 3 and the output variables are 2, all are float variables.

b)

The input variables are 2 and the output variables are 3, all are integer variables.

c)

The input variables are 3 and the output variables are 2, all are integer variables.

d)

The input variables are 2 and the output variables are 3, all are float variables.

23.

In the provided Python code, what is the purpose of the function analytical_solution(X, T)?

a)

To generate random data for regression modeling.

b)

To calculate the analytical solution for linear regression parameters.

c)

To plot the regression results.

d)

To normalize the input data.

24.

Which Python library function is used in the code to compute the inverse of a matrix?

a)

np.dot()

b)

np.ones()

c)

np.linalg.inv()

d)

np.array()

25.

Given the code snippet, what is the role of X_b in the analytical_solution function?

a)

It stores the output variables.

b)

It appends a column of ones to X for the bias term in linear regression.

c)

It normalizes the input data.

d)

It stores the inverse of X.

26.

Suppose you want to use the analytical_solution function to solve a linear regression problem with a new dataset. What must be true about the shapes of X and T?

a)

X and T must have the same number of columns.

b)

X must have the same number of rows as T.

c)

X must be a square matrix.

d)

T must be a scalar.

27.

Which Python method is suggested for loading the Boston housing dataset in the assignment?

a)

tf.keras.datasets.mnist.load_data

b)

tf.keras.datasets.boston_housing.load_data

c)

tf.keras.datasets.cifar10.load_data

d)

tf.keras.datasets.imdb.load_data

28.

According to the text, what is the expected outcome when using both analytical and gradient descent methods for linear regression?

a)

The results will be different

b)

Only gradient descent will work

c)

Both methods will have the same result

d)

Only the analytical method will work

29.

Why does the text state that both analytical and gradient descent methods yield the same result for linear regression?

a)

Because the dataset is small

b)

Because linear regression has a unique solution

c)

Because gradient descent is faster

d)

Because analytical methods are outdated

30.

Suppose you are asked to predict house prices using the Keras deep learning library and Google Colab. What is the first step you should take according to the assignment instructions?

a)

Visualize the data

b)

Load the dataset using tf.keras.datasets.boston_housing.load_data

c)

Train a neural network

d)

Normalize the features