WorksheetsLinear Regression Quiz
Total questions: 30
Worksheet time: 15mins
Who is credited with the quote, "All models are wrong, but some are useful," as mentioned in the learning material?
Isaac Newton
George Box
Albert Einstein
Carl Gauss
What is the primary purpose of linear regression as described in the material?
To classify data into categories
To model the relationship between input and output variables using a linear function
To cluster data points into groups
To encrypt data for security
According to the material, what must the output variables in linear regression be?
Only binary values
Only integer values
Real or float values
Only categorical values
Given the equations y₁ = w₁₁x₁ + w₁₂x₂ + w₁₃x₃ + b₁ and y₂ = w₂₁x₁ + w₂₂x₂ + w₂₃x₃ + b₂, what does the term w₁₁ represent?
The bias for the first output
The weight connecting the first input to the first output
The output variable
The input variable
What does Figure 4.1 in the material illustrate?
The process of data encryption
The relationship between input and output variables in linear regression
The steps of clustering algorithms
The architecture of a neural network for image recognition
Why is house price prediction given as an example in the material?
It is an example of predicting a categorical variable
It is an example of predicting a float variable using input features
It is an example of clustering data points
It is an example of classifying images
Which equation represents a linear regression single-layer network as described in the material?
y = Wx + b
y = Wx - b
y = W + bx
y = xW + b
Which of the following best describes the objective function E as defined in the material?
It measures how close the predicted value y is to the target value t over all training data.
It calculates the sum of input variables.
It determines the number of layers in the network.
It finds the maximum value of the output variable.
What does the Mean Square Error (MSE) loss function measure in linear regression?
The number of input variables
How well the linear model fits the data
The number of output nodes
The number of training samples
Which formula represents the Mean Square Error (MSE) loss function for a set of training samples?
E(W, b) = N1n=1∑NEn
E(W, b) = n=1∑NWnbn
E(W,b)=Nn=1∑NEn
E(W, b) = N1n=1∑NWnbn
What is the result of simplifying the cost function E(\theta) as shown in the document?
E(θ)=θTXTXθ−2tTXθ+tTt
E(θ)=XTXθ+tTt
E(θ)=θTXTXθ+2tTXθ
E(θ)=θTXTXθ−tTXθ
What is the first step to find the optimal value of θ in the context of the learning problem described?
Take the first derivative of the cost function with respect to the parameters and set it to zero.
Multiply the cost function by the parameters.
Integrate the cost function with respect to the parameters.
Set the cost function equal to one.
Why is the analytical solution for finding the values of parameters W and b not always preferred for large training datasets?
It does not scale well with the amount of training data.
It is less accurate than other methods.
It requires more memory than gradient descent.
It is only applicable to non-linear models.
What does the gradient of the MSE loss function with respect to the parameters W and b help to compute in the context of gradient descent?
The direction and magnitude to update the parameters.
The final value of the cost function.
The number of training samples.
The learning rate.
In the gradient descent update rule for the weights and bias, what does the symbol η represent?
The learning rate.
The number of epochs.
The cost function.
The bias term.
How does the mini-batch stochastic gradient variant help when scaling to large amounts of training data?
It estimates the parameters of the linear regression model using subsets of the data.
It increases the learning rate.
It reduces the number of parameters.
It eliminates the need for a cost function.
Given the update rule wrst+1=wrst−η∂E(W,b)/∂wrs , what is the purpose of subtracting the gradient term?
To move the parameters in the direction that reduces the loss.
To increase the value of the weights.
To keep the weights constant.
To maximize the loss function.
What is one advantage of mini-batch gradient descent?
It can handle large datasets that do not fit in memory
It always converges faster than any other method
It does not require tuning any hyperparameters
It cannot exploit parallelism
Why is tuning the batch size and learning rate important in mini-batch gradient descent?
Because they affect the convergence and performance of the algorithm
Because they determine the number of input features
Because they are not used in the algorithm
Because they only affect the output format
Suppose you are given a dataset too large to fit in memory. Which gradient descent method would be most suitable?
Mini-batch stochastic gradient descent
Batch gradient descent
Newton's method
Random search
In the context of the provided example, what do the first three columns of the training data represent?
Input features (x1, x2, x3)
Output targets (t1, t2)
Batch sizes
Learning rates
Which of the following best describes the input and output variables in the synthesized data for linear regression modeling shown in Table 4.1?
The input variables are 3 and the output variables are 2, all are float variables.
The input variables are 2 and the output variables are 3, all are integer variables.
The input variables are 3 and the output variables are 2, all are integer variables.
The input variables are 2 and the output variables are 3, all are float variables.
In the provided Python code, what is the purpose of the function analytical_solution(X, T)?
To generate random data for regression modeling.
To calculate the analytical solution for linear regression parameters.
To plot the regression results.
To normalize the input data.
Which Python library function is used in the code to compute the inverse of a matrix?
np.dot()
np.ones()
np.linalg.inv()
np.array()
Given the code snippet, what is the role of X_b in the analytical_solution function?
It stores the output variables.
It appends a column of ones to X for the bias term in linear regression.
It normalizes the input data.
It stores the inverse of X.
Suppose you want to use the analytical_solution function to solve a linear regression problem with a new dataset. What must be true about the shapes of X and T?
X and T must have the same number of columns.
X must have the same number of rows as T.
X must be a square matrix.
T must be a scalar.
Which Python method is suggested for loading the Boston housing dataset in the assignment?
tf.keras.datasets.mnist.load_data
tf.keras.datasets.boston_housing.load_data
tf.keras.datasets.cifar10.load_data
tf.keras.datasets.imdb.load_data
According to the text, what is the expected outcome when using both analytical and gradient descent methods for linear regression?
The results will be different
Only gradient descent will work
Both methods will have the same result
Only the analytical method will work
Why does the text state that both analytical and gradient descent methods yield the same result for linear regression?
Because the dataset is small
Because linear regression has a unique solution
Because gradient descent is faster
Because analytical methods are outdated
Suppose you are asked to predict house prices using the Keras deep learning library and Google Colab. What is the first step you should take according to the assignment instructions?
Visualize the data
Load the dataset using tf.keras.datasets.boston_housing.load_data
Train a neural network
Normalize the features
