WorksheetsNeural Networks Quiz
Total questions: 25
Worksheet time: 13mins
A simple Multi-Layer Perceptron (MLP) without any non-linear activation functions is mathematically equivalent to:
A universal function approximator.
A single, wider linear layer.
A model incapable of learning.
A support vector machine with a linear kernel.
The primary motivation for using the ReLU activation function over Sigmoid in deep hidden layers is to:
Ensure the output is always positive.
Confine the activation values between 0 and 1.
Mitigate the vanishing gradient problem.
Make the network more computationally expensive but more accurate.
The "dying ReLU" problem is characterized by:
A neuron's weights being updated such that its pre-activation input is consistently negative, causing its output and gradient to be zero.
The network becoming too deep, causing all ReLU activations to eventually become zero.
The learning rate being too low, preventing weights from being updated.
A neuron's weights exploding to infinity.
Why is breaking symmetry by initializing weights randomly (e.g., using Xavier or He initialization) crucial for training?
It guarantees faster convergence to the global minimum.
It prevents all neurons in a layer from learning the same features, as they would with zero initialization.
It acts as a form of L2 regularization.
It ensures the initial loss of the network is exactly 1.0.
What is the primary purpose of the bias term in a neuron?
To scale the output of the activation function.
To act as a learnable offset, allowing the activation function to be shifted left or right.
To prevent the weights from becoming zero during training.
To control the learning rate for that specific neuron.
The Universal Approximation Theorem suggests that:
Any neural network can solve any problem.
A single-layer perceptron can approximate any linear function.
A feed-forward network with one hidden layer and a non-linear activation can approximate any continuous function to arbitrary precision.
Deeper networks are always better than wider networks.
A model is exhibiting high variance and low bias. This is a classic case of:
Underfitting, where the model is too simple for the data.
Overfitting, where the model has learned the training data too well, including its noise.
A well-generalized model.
A model trained with an incorrect loss function.
If you observe that both your training loss and validation loss are high and have plateaued, your model is likely:
Overfitting.
Converging perfectly.
Underfitting.
In need of a higher dropout rate.
What is the fundamental mechanism of Backpropagation?
It uses the chain rule of calculus to efficiently compute the gradient of the loss function with respect to every parameter in the network.
It randomly perturbs weights to see if the loss decreases.
It is a specific type of optimizer, similar to Adam.
It calculates the loss by propagating a signal forward through the network.
In a multi-class classification problem with 10 classes, what is the most appropriate combination for the final layer and loss function?
10 neurons with Sigmoid activation and Binary Cross-Entropy loss.
1 neuron with a linear activation and Mean Squared Error loss.
10 neurons with Softmax activation and Categorical Cross-Entropy loss.
10 neurons with ReLU activation and Hinge loss.
What is the key advantage of "parameter sharing" in a convolutional layer?
It allows each neuron to learn a unique feature.
It dramatically reduces the total number of learnable parameters, making the network translationally equivariant.
It forces the network to only learn features that are centered in the image.
It increases the spatial dimensions of the feature maps.
What does a single filter (or kernel) in a convolutional layer learn to detect?
The overall class of the image.
A specific, local pattern such as an edge, a color blob, a texture, or a corner.
The optimal padding value for the layer.
The degree of overfitting.
In a CNN, what is the primary purpose of a Max Pooling layer?
To introduce non-linearity into the network.
To reduce the spatial dimensions of the feature maps and provide a degree of translational invariance.
To increase the number of channels in the feature map.
To perform data augmentation.
If you use a convolutional layer with a stride of 2, what is the primary effect on the output feature map?
The number of channels is doubled.
The spatial dimensions (height and width) are roughly halved.
The receptive field of the neurons decreases.
The number of parameters in the filter is reduced.
What does 'VALID' padding signify in a convolutional operation?
The input is padded with zeros so the output has the same dimensions as the input.
No padding is applied, and the filter is only applied where it fully overlaps with the input.
The padding is learned during training.
The input is padded with ones instead of zeros.
As you move to deeper layers in a typical CNN, the features learned generally become:
More simple and generic (e.g., edges and colors).
More complex and abstract (e.g., object parts and entire objects).
Smaller in their receptive field.
Less useful for classification.
The "receptive field" of a neuron in a deep CNN layer refers to:
The set of all possible activation values for that neuron.
The region in the original input image that affects the activation of that neuron.
The number of parameters associated with that neuron's filter.
The total number of layers in the network.
What is a primary, non-obvious use case for a 1x1 convolution?
To act as a simple linear transformation without changing spatial dimensions.
To blur the image slightly.
To reduce or increase the number of channels (feature map depth), often used in "bottleneck" layers for computational efficiency.
To detect single-pixel features only.
The VGG architecture is famous for demonstrating the power of:
Using very large (e.g., 11x11) filters.
A very simple and uniform architecture, stacking 3x3 convolutions to achieve great depth.
Inception modules with parallel convolutions.
Skip connections to train extremely deep networks.
How does a Residual Network (ResNet) effectively combat the degradation and vanishing gradient problems in very deep networks?
By using a better optimizer like Adam.
By using "skip connections" that allow the gradient to flow directly to earlier layers.
By using only linear activation functions.
By applying batch normalization after every layer.
In an Inception (GoogLeNet) module, what is the core architectural idea?
To stack 3x3 convolutions sequentially to create a very deep block.
To perform several different-sized convolutions (1x1, 3x3, 5x5) and pooling in parallel and concatenate their outputs.
To replace all convolutions with fully connected layers.
To use a single, very large filter to capture global context.
What is Global Average Pooling (GAP) and how is it often used?
It's another name for a Max Pooling layer with a very large window.
It takes each feature map and reduces it to a single number by averaging all its values, often used to replace the final Flatten and Dense layers in a classifier.
It averages all pixel values in the input image before it enters the network.
It's a type of data augmentation.
A Transposed Convolution (sometimes called deconvolution) is an operation primarily used for:
Down-sampling the feature maps to reduce computational cost.
Performing an inverse convolution to perfectly reconstruct the input.
Up-sampling the feature maps, effectively increasing their spatial resolution, common in segmentation and generative models.
Rotating the feature maps by 90 degrees.
Which of these is a key characteristic of a fully convolutional network (FCN)?
It contains only convolutional layers and no other layer types.
It replaces the final dense (fully connected) layers with 1x1 convolutions, allowing it to process inputs of any spatial size.
It is only used for image regression tasks.
It does not use any activation functions.
Batch Normalization in a CNN is typically applied:
Only to the raw input image.
After the activation function (e.g., after ReLU).
Between the convolutional operation and the non-linear activation function.
As a replacement for the pooling layer.
