Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Week2_S2

Total questions: 10

Worksheet time: 8mins

Name
Class
Date
1.

Why is generalization more important than training accuracy in neural networks?

a)

 Generalization proves convergence.

b)

Training accuracy ensures bias reduction.

c)

 It prevents vanishing gradients.

d)

It reflects the ability to predict unseen data, the true goal.

2.

A multilayer network without nonlinearities collapses into a (a)   model.

3.

When using ReLU in hidden layers instead of sigmoid, which benefit typically emerges?

a)
  1. Guarantees linear separability in the input space.

b)

Prevents exploding gradients.

c)

Ensures all neurons remain active.

d)

Reduces vanishing gradient problems

4.

Which of the following statements about gradient descent is correct?

a)

It always finds the global minimum.

b)

 It updates weights by moving against the gradient of the loss function.

c)

It requires linear separability.

d)

It is identical to the perceptron update rule.

5.

Why is nonlinearity essential in deep neural networks?

a)

It guarantees zero error.

b)

 It reduces training time.

c)

It allows the composition of layers to model complex, non-linear boundaries.

d)

It simplifies the optimization problem.

6.

Backpropagation uses the (a)   rule to propagate gradients backward through layers

7.

In a neural network, forward propagation refers to:

a)

Feeding inputs through the network to generate predictions

b)

Updating weights using gradient descent

c)

Reversing gradients to find errors

d)

Adjusting biases to prevent saturation

8.

Which activation function is most commonly used in modern deep networks for hidden layers?

a)

Sigmoid

b)

Tanh

c)

ReLU

d)

Sign function

9.

A perceptron without a bias term may struggle because:

a)

It cannot represent decision boundaries that don't pass through the origin

b)

It cannot compute loss correctly

c)

It requires too many hidden layers

d)

It cannot apply nonlinear activations

10.

Which statement about cross-entropy loss is correct?

a)

It measures the squared difference between predicted and actual values

b)

It is used to calculate the probability of linear separability

c)

It penalizes wrong predictions by increasing loss logarithmically

d)

It works only with ReLU activations