WorksheetsHow might LLMs store facts | Deep Learning Chapter 7
Total questions: 12
Worksheet time: 6mins
How do Large Language Models (LLMs) primarily store factual knowledge?
In external databases linked via API calls.
Through explicit rule-based programming.
Within their hundreds of billions of parameters.
By dynamically searching the internet for information.
What are the two primary types of computational blocks that constitute a Transformer architecture?
Convolutional Layers and Recurrent Layers.
Encoder and Decoder Stacks.
Attention and Multilayer Perceptron (MLP) blocks.
Gated Recurrent Units (GRUs) and Long Short-Term Memory (LSTMs).
What is the fundamental computational structure of a Multilayer Perceptron (MLP) block within a Transformer?
A sequence of convolutional, pooling, and fully connected layers.
A pair of matrix multiplications separated by a ReLU activation function.
A recurrent neural network with feedback loops.
A self-attention mechanism followed by a feed-forward network.
In the context of Transformer models, how can specific "directions" in a high-dimensional vector space encode different types of meaning or features?
By assigning a unique numerical ID to each meaning.
Through the magnitude of the vector, where larger magnitudes indicate more important meanings.
By the dot product of a token's embedding vector with a direction vector, indicating alignment with that feature.
By the Euclidean distance between token embeddings, where closer embeddings share similar meanings.
What is the primary role of the matrix in the initial linear transformation of an input vector in a Multilayer Perceptron (MLP)?
To perform element-wise multiplication with the input vector.
To store the learned model parameters that define the transformation.
To normalize the input vector's values before processing.
To directly output the final prediction of the model.
What is the primary function of the Rectified Linear Unit (ReLU) activation function in a neural network?
To scale all input values to a range between 0 and 1.
To introduce non-linearity by setting negative inputs to zero and passing positive inputs unchanged.
To compute the dot product between the input and weight vectors.
To reduce the dimensionality of the input vector.
In the context of the Multilayer Perceptron (MLP) architecture described, what is the relationship between the "up projection" and "down projection" matrices?
The up projection matrix reduces dimensionality, while the down projection matrix expands it.
The up projection matrix maps to a higher-dimensional space, and the down projection matrix maps back to the original embedding space.
Both matrices perform identical transformations on the input vector.
The up projection matrix applies a non-linear activation, while the down projection matrix applies a linear one.
How does a Multilayer Perceptron (MLP) combine information from an input vector to produce an output vector, particularly regarding feature representation?
It exclusively uses non-linear transformations to directly map input features to output features.
It applies a series of linear transformations and non-linear activations, where active intermediate neurons contribute specific feature directions to the output.
It performs a single, complex mathematical operation that directly calculates the output without intermediate steps.
It only considers the most dominant feature in the input vector and ignores others.
How many distinct Multilayer Perceptron (MLP) layers are present in GPT-3's architecture?
12
24
48
96
In the context of large language models and the concept of superposition, how are complex features like "Michael Jordan" typically represented?
By a single, dedicated neuron that activates when the feature is present.
By a specific combination of activations across multiple neurons.
By a unique, perfectly orthogonal vector in the embedding space.
By a sparse autoencoder that directly maps to a single output.
What is the primary advantage of "superposition" in high-dimensional embedding spaces for large language models?
It simplifies model interpretation by assigning one feature per neuron.
It allows for the storage of exponentially more features than the number of dimensions.
It reduces the total number of parameters required for the model.
It ensures all feature vectors are perfectly orthogonal.
According to the Johnson-Lindenstrauss Lemma, what happens to the maximum number of nearly perpendicular vectors that can be crammed into an N-dimensional space as N increases?
It remains constant, regardless of N.
It grows linearly with N.
It grows exponentially with N.
It decreases logarithmically with N.
