WorksheetsIntroduction and Background
Total questions: 62
Worksheet time: 31mins
Which limitation of recurrent models primarily motivates replacing recurrence with self-attention in sequence transduction tasks?
Difficulty modeling long-range dependencies efficiently
Inability to learn local patterns in short sequences
Requirement for convolutional filters at every layer
Lack of an encoder–decoder structure by design
What does sequence transduction most commonly refer to in machine translation?
Mapping an input symbol sequence to an output sequence
Clustering input tokens into semantic groups
Ranking candidate sentences by fluency score
Detecting anomalies across parallel corpora
Why do recurrent models face parallelization constraints during training?
They require GPU-specific fused kernels
They depend on fixed-length input padding
They need convolutional strides for efficiency
They compute states sequentially across positions
Self-attention relates positions within a single sequence to compute what?
A representation of the sequence itself
A probability distribution over vocabularies
A hierarchical parse of syntax trees
A compressed index of token embeddings
Compared to RNNs and CNNs, what computational advantage does self-attention provide in the Transformer?
Higher-resolution features at every layer
More parameters to capture complex patterns
Fewer operations to connect distant positions
Guaranteed linear-time decoding for outputs
Which architecture component does the Transformer remove entirely in favor of attention mechanisms?
Positional encodings and embeddings
Softmax output and loss functions
Recurrence and convolution layers
Residual connections and layer norms
In the described model, what enables improved parallelization across input and output positions?
Using teacher forcing during training only
Computing representations for all positions in parallel
Tying encoder and decoder parameters
Applying fixed sinusoidal positional encodings
What is one empirical outcome reported for the Transformer regarding training time and quality?
Better translation quality with significantly less training time
Similar quality but longer training due to attention
Lower quality offset by extreme parameter reduction
Unchanged quality with massive compute requirements
End-to-end memory networks attempt to replace what with attention-aligned mechanisms?
Token-level dropout and noise scheduling
Convolutional pooling across feature maps
Sequence-aligned recurrence across positions
Beam search during decoding stages
In an encoder–decoder scheme, what does the decoder condition on at each generation step?
Aligned source positions through max pooling
External language models for rescoring
Randomly sampled tokens for exploration
Previously generated symbols as additional input
Which claim distinguishes the Transformer among prior transduction models using attention?
It uses attention only in the decoder while keeping recurrent encoders
It introduces attention solely for cross-lingual alignment
It relies entirely on self-attention without recurrence or convolution
It combines CNN encoders with RNN decoders via attention
What fundamental constraint of sequential computation remains in recurrent architectures despite efficiency tricks?
Requirement for large batch sizes to converge
Inability to represent continuous sequences
Dependence on fixed-width convolutions per layer
Dependence on previous hidden states across positions
In the encoder stack, which sub-layer is applied after Multi-Head Attention within each block?
Feed Forward with Add & Norm
Positional Encoding addition
Masked Multi-Head Attention
Output Softmax projection
What is the primary role of the masked self-attention in the decoder?
Prevent attending to future positions
Amplify gradients from earlier tokens
Replace cross-attention to encoder
Normalize embeddings to unit length
Which component connects the decoder to the encoder outputs?
Linear plus Softmax head
Positional Encoding on outputs
Feed Forward with residuals
Multi-Head Attention over encoder
In Scaled Dot-Product Attention, which operation immediately follows the dot product of Q and K^T?
Scaling by 1/√d_k
Concatenation across heads
Residual addition and norm
Linear projection of values
Why is the scaling factor 1/√d_k applied before softmax in dot-product attention?
Increase variance of key vectors
Reduce large logits to stabilize softmax
Enable concatenation across heads
Ensure outputs match d_v dimension
In Multi-Head Attention, queries, keys, and values are first transformed by which operation per head?
Concatenation across all heads
Softmax normalization over values
Elementwise scaling with 1/√d_v
Separate linear projections
After computing attention separately in each head, how are the head outputs combined?
Concatenation then linear projection
Averaging across heads directly
Max-pooling across head channels
Residual subtraction across heads
Which vectors define compatibility in the attention mechanism?
Embedding with Positional codes
Value with Output logits
Query with Key pairs
Decoder with Softmax weights
What added structure enables the decoder to produce outputs one position ahead of the input embeddings?
Shifted-right output embeddings
Extra feed-forward block
Global positional normalization
Encoder residual connections
Within each encoder or decoder sub-layer, what normalization is used along residual connections?
Layer normalization
Batch normalization
Instance normalization
Group normalization
Which mathematical expression gives attention weights over values?
softmax(QKT/dk)
concat(Q,V)TK
linear(V) + residual
norm(Q) · norm(K)
What additional sub-layer does the decoder have that the encoder does not?
Double feed-forward
Extra Softmax head
Masked self-attention
Positional embedding layer
In the attention pipeline, where can an optional mask be applied?
After final linear projection
After concatenation of heads
Before the initial MatMul
Between scaling and softmax
What is the dimension of outputs from each attention head before concatenation?
d_v per head
d_k per head
h·d_v per head
d_model per head
Which sequence of operations best describes an encoder block?
Masked attention then cross-attention
Concatenation then positional encoding
Feed-forward then softmax head
Self-attention then feed-forward
Why use multiple heads instead of a single attention head of d_model?
Capture information from different subspaces
Reduce need for positional encodings
Guarantee constant gradient magnitude
Avoid residual connections entirely
Which statement best describes a position-wise feed-forward network in a Transformer layer?
Uses convolutional kernels that slide over tokens
Shares different weights across positions dynamically
Applies same two linear layers to each position
Computes attention scores between all token pairs
In the standard Transformer, what activation is used between the two linear layers of the position-wise feed-forward network?
Tanh activation followed by dropout
ReLU activation between linear layers
Sigmoid activation in residual branch
Softmax activation over positions
What are the typical dimensions in the described FFN for inner-layer and model size?
dff = 2048, dmodel = 512
dff = 4096, dmodel = 256
dff = 512, dmodel = 2048
dff = 1024, dmodel = 1024
Why are positional encodings necessary in a Transformer without recurrence or convolution?
To inject relative or absolute token positions
To reduce parameter count of attention heads
To increase embedding dimensionality
To normalize gradients during training
Which functions are used for sinusoidal positional encoding components at even and odd indices?
sin for even, cos for odd indices
softmax for all indices
tanh for even, sigmoid for odd
cos for even, sin for odd indices
What property of sinusoidal positional encoding allows attending to relative positions easily?
Encodings are invariant to token content
Any fixed offset can be represented linearly
Higher dimensions capture syntax trees
Values remain constant across sequences
In the embedding layers, why are weights multiplied by sqrt(dmodel)?
To match attention bias terms
To reduce overfitting in training
To enforce unit-norm embeddings
To scale embeddings to model dimension
Which layer type has per-layer complexity O(n^2 · d) and constant sequential operations O(1)?
Restricted self-attention type
Self-Attention layer type
Recurrent layer type
Convolutional layer type
For recurrent layers, what is the minimum number of sequential operations per layer?
O(n) sequential operations
O(1) sequential operations
O(log n) sequential operations
O(n2) sequential operations
Which layer type can have maximum path length O(1) between any two positions?
Self-Attention layer type
Recurrent layer type
Convolutional layer type
Restricted self-attention type
What is a key advantage of self-attention over recurrent layers when sequence length n is large relative to representation dimension d?
Faster due to lower sequential operations
Better at local pattern detection
Requires fewer parameters overall
Avoids the need for embeddings
In encoder-decoder attention, where do queries originate and what do they attend to?
Queries from FFN attend to previous heads
Queries from encoder attend to decoder inputs
Queries from embeddings attend to positions
Queries from decoder attend to encoder outputs
What masking is applied in decoder self-attention to preserve autoregressive generation?
Set future positions to negative infinity
Drop all past positions randomly
Zero out diagonal attention entries
Normalize logits by sequence length
Which statement about learned versus sinusoidal positional embeddings was observed?
Both produced nearly similar results
Learned required fewer parameters
Learned embeddings always outperformed
Sinusoidal always failed on long sequences
Which dataset was used for English-to-German training in the described experiments?
Europarl v7 corpus
Multi30k captions dataset
IWSLT 2017 TED talks
WMT 2014 full training corpus
WMT 2014 newstest2014 set
What hardware configuration was used to train the models?
Single machine with 8 P100 GPUs
Cluster of 32 K80 GPUs
TPU v2 pod with 64 cores
Single RTX 4090 GPU
CPU-only 64-core server
For base models, what was the total number of training steps?
100,000 steps
300,000 steps
200,000 steps
75,000 steps
50,000 steps
Which optimizer was employed during training?
Adam optimizer
LAMB optimizer
SGD with momentum
AdaGrad optimizer
RMSProp optimizer
Which hyperparameters were used for Adam in these experiments?
β1=0.9 β2=0.98 ε=1e−9
β1=0.9 β2=0.9 ε=1e−10
β1=0.99 β2=0.999 ε=1e−8
β1=0.9 β2=0.999 ε=1e−8
β1=0.95 β2=0.97 ε=1e−9
How was the learning rate scheduled over training?
Step decay at fixed intervals
Cosine annealing with restarts
Cyclical triangular schedule
Exponential decay from the start
Linear warmup then inverse square root decay
Which regularization method was applied to the output of each sub-layer?
Weight decay L2 penalty
Stochastic depth on blocks
Residual dropout before addition
Spatial dropout on embeddings
Zoneout in recurrent units
What residual dropout rate was used for the base model?
p=0.3
p=0.2
p=0.1
p=0.4
p=0.05
During training, what label smoothing value was employed?
εls=0.3
εls=0.4
εls=0.05
εls=0.1
εls=0.2
On WMT14 English-to-German, what BLEU did the big Transformer achieve on newstest2014?
28.4 BLEU
30.5 BLEU
32.1 BLEU
26.0 BLEU
27.3 BLEU
Compared to previous state-of-the-art, how did the big model perform on English-to-French?
Higher BLEU at fraction of cost
Equal BLEU with same cost
Similar BLEU but higher cost
Lower BLEU and higher cost
Unchanged BLEU across models
What inference beam search settings were used for big models’ checkpoint averaging?
Beam size 6 length penalty 1.0
Beam size 4 length penalty 0.6
Beam size 8 length penalty 0.8
Beam size 2 length penalty 0.4
Beam size 1 length penalty 0.0
Which change to the attention key size dk most negatively impacted model quality in experiments?
Reducing dk notably
Keeping dk unchanged
Randomly varying dk
Increasing dk moderately
What training strategy helped avoid overfitting in larger Transformer variants?
Smaller batch size
Lower learning rate
Higher dropout rate
More epochs
Replacing sinusoidal positional encoding with learned positional embeddings produced what effect on results?
Major improvement overall
Slight improvement only
Significantly worse scores
Nearly identical results
Compared to recurrent or convolutional architectures, the Transformer can be trained how for translation tasks?
Significantly faster overall
About the same speed
Significantly slower overall
Only faster on small data
On WMT14 English-to-German, what performance claim is made for the Transformer?
Matches prior baselines
Sets a new state of the art
Underperforms prior ensembles
Improves only BLEU, not PPL
Which observation suggests exploring a compatibility function more sophisticated than dot product?
Increasing model depth helped
Reducing dk hurt quality
Learned positions improved
Dropout decreased BLEU
What future direction aims to handle large inputs and outputs across modalities like images and audio?
Restricted attention mechanisms
Wider embedding vectors
Longer training schedules
Deeper feedforward layers
Where is the code used to train and evaluate the models available?
pytorch/examples repo
github.com/tensorflow/tensor2tensor
tensorflow/models repo
apache/incubator/mxnet
