Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Introduction and Background

Total questions: 62

Worksheet time: 31mins

Name
Class
Date
1.

Which limitation of recurrent models primarily motivates replacing recurrence with self-attention in sequence transduction tasks?

a)

Difficulty modeling long-range dependencies efficiently

b)

Inability to learn local patterns in short sequences

c)

Requirement for convolutional filters at every layer

d)

Lack of an encoder–decoder structure by design

2.

What does sequence transduction most commonly refer to in machine translation?

a)

Mapping an input symbol sequence to an output sequence

b)

Clustering input tokens into semantic groups

c)

Ranking candidate sentences by fluency score

d)

Detecting anomalies across parallel corpora

3.

Why do recurrent models face parallelization constraints during training?

a)

They require GPU-specific fused kernels

b)

They depend on fixed-length input padding

c)

They need convolutional strides for efficiency

d)

They compute states sequentially across positions

4.

Self-attention relates positions within a single sequence to compute what?

a)

A representation of the sequence itself

b)

A probability distribution over vocabularies

c)

A hierarchical parse of syntax trees

d)

A compressed index of token embeddings

5.

Compared to RNNs and CNNs, what computational advantage does self-attention provide in the Transformer?

a)

Higher-resolution features at every layer

b)

More parameters to capture complex patterns

c)

Fewer operations to connect distant positions

d)

Guaranteed linear-time decoding for outputs

6.

Which architecture component does the Transformer remove entirely in favor of attention mechanisms?

a)

Positional encodings and embeddings

b)

Softmax output and loss functions

c)

Recurrence and convolution layers

d)

Residual connections and layer norms

7.

In the described model, what enables improved parallelization across input and output positions?

a)

Using teacher forcing during training only

b)

Computing representations for all positions in parallel

c)

Tying encoder and decoder parameters

d)

Applying fixed sinusoidal positional encodings

8.

What is one empirical outcome reported for the Transformer regarding training time and quality?

a)

Better translation quality with significantly less training time

b)

Similar quality but longer training due to attention

c)

Lower quality offset by extreme parameter reduction

d)

Unchanged quality with massive compute requirements

9.

End-to-end memory networks attempt to replace what with attention-aligned mechanisms?

a)

Token-level dropout and noise scheduling

b)

Convolutional pooling across feature maps

c)

Sequence-aligned recurrence across positions

d)

Beam search during decoding stages

10.

In an encoder–decoder scheme, what does the decoder condition on at each generation step?

a)

Aligned source positions through max pooling

b)

External language models for rescoring

c)

Randomly sampled tokens for exploration

d)

Previously generated symbols as additional input

11.

Which claim distinguishes the Transformer among prior transduction models using attention?

a)

It uses attention only in the decoder while keeping recurrent encoders

b)

It introduces attention solely for cross-lingual alignment

c)

It relies entirely on self-attention without recurrence or convolution

d)

It combines CNN encoders with RNN decoders via attention

12.

What fundamental constraint of sequential computation remains in recurrent architectures despite efficiency tricks?

a)

Requirement for large batch sizes to converge

b)

Inability to represent continuous sequences

c)

Dependence on fixed-width convolutions per layer

d)

Dependence on previous hidden states across positions

13.

In the encoder stack, which sub-layer is applied after Multi-Head Attention within each block?

a)

Feed Forward with Add & Norm

b)

Positional Encoding addition

c)

Masked Multi-Head Attention

d)

Output Softmax projection

14.

What is the primary role of the masked self-attention in the decoder?

a)

Prevent attending to future positions

b)

Amplify gradients from earlier tokens

c)

Replace cross-attention to encoder

d)

Normalize embeddings to unit length

15.

Which component connects the decoder to the encoder outputs?

a)

Linear plus Softmax head

b)

Positional Encoding on outputs

c)

Feed Forward with residuals

d)

Multi-Head Attention over encoder

16.

In Scaled Dot-Product Attention, which operation immediately follows the dot product of Q and K^T?

a)

Scaling by 1/√d_k

b)

Concatenation across heads

c)

Residual addition and norm

d)

Linear projection of values

17.

Why is the scaling factor 1/√d_k applied before softmax in dot-product attention?

a)

Increase variance of key vectors

b)

Reduce large logits to stabilize softmax

c)

Enable concatenation across heads

d)

Ensure outputs match d_v dimension

18.

In Multi-Head Attention, queries, keys, and values are first transformed by which operation per head?

a)

Concatenation across all heads

b)

Softmax normalization over values

c)

Elementwise scaling with 1/√d_v

d)

Separate linear projections

19.

After computing attention separately in each head, how are the head outputs combined?

a)

Concatenation then linear projection

b)

Averaging across heads directly

c)

Max-pooling across head channels

d)

Residual subtraction across heads

20.

Which vectors define compatibility in the attention mechanism?

a)

Embedding with Positional codes

b)

Value with Output logits

c)

Query with Key pairs

d)

Decoder with Softmax weights

21.

What added structure enables the decoder to produce outputs one position ahead of the input embeddings?

a)

Shifted-right output embeddings

b)

Extra feed-forward block

c)

Global positional normalization

d)

Encoder residual connections

22.

Within each encoder or decoder sub-layer, what normalization is used along residual connections?

a)

Layer normalization

b)

Batch normalization

c)

Instance normalization

d)

Group normalization

23.

Which mathematical expression gives attention weights over values?

a)

softmax(QKT/dk)softmax(QK^T/\sqrt{d_k})

b)

concat(Q,V)TKconcat(Q,V)^T K

c)

linear(V) + residual

d)

norm(Q) · norm(K)

24.

What additional sub-layer does the decoder have that the encoder does not?

a)

Double feed-forward

b)

Extra Softmax head

c)

Masked self-attention

d)

Positional embedding layer

25.

In the attention pipeline, where can an optional mask be applied?

a)

After final linear projection

b)

After concatenation of heads

c)

Before the initial MatMul

d)

Between scaling and softmax

26.

What is the dimension of outputs from each attention head before concatenation?

a)

d_v per head

b)

d_k per head

c)

h·d_v per head

d)

d_model per head

27.

Which sequence of operations best describes an encoder block?

a)

Masked attention then cross-attention

b)

Concatenation then positional encoding

c)

Feed-forward then softmax head

d)

Self-attention then feed-forward

28.

Why use multiple heads instead of a single attention head of d_model?

a)

Capture information from different subspaces

b)

Reduce need for positional encodings

c)

Guarantee constant gradient magnitude

d)

Avoid residual connections entirely

29.

Which statement best describes a position-wise feed-forward network in a Transformer layer?

a)

Uses convolutional kernels that slide over tokens

b)

Shares different weights across positions dynamically

c)

Applies same two linear layers to each position

d)

Computes attention scores between all token pairs

30.

In the standard Transformer, what activation is used between the two linear layers of the position-wise feed-forward network?

a)

Tanh activation followed by dropout

b)

ReLU activation between linear layers

c)

Sigmoid activation in residual branch

d)

Softmax activation over positions

31.

What are the typical dimensions in the described FFN for inner-layer and model size?

a)

dff = 2048, dmodel = 512

b)

dff = 4096, dmodel = 256

c)

dff = 512, dmodel = 2048

d)

dff = 1024, dmodel = 1024

32.

Why are positional encodings necessary in a Transformer without recurrence or convolution?

a)

To inject relative or absolute token positions

b)

To reduce parameter count of attention heads

c)

To increase embedding dimensionality

d)

To normalize gradients during training

33.

Which functions are used for sinusoidal positional encoding components at even and odd indices?

a)

sin for even, cos for odd indices

b)

softmax for all indices

c)

tanh for even, sigmoid for odd

d)

cos for even, sin for odd indices

34.

What property of sinusoidal positional encoding allows attending to relative positions easily?

a)

Encodings are invariant to token content

b)

Any fixed offset can be represented linearly

c)

Higher dimensions capture syntax trees

d)

Values remain constant across sequences

35.

In the embedding layers, why are weights multiplied by sqrt(dmodel)?

a)

To match attention bias terms

b)

To reduce overfitting in training

c)

To enforce unit-norm embeddings

d)

To scale embeddings to model dimension

36.

Which layer type has per-layer complexity O(n^2 · d) and constant sequential operations O(1)?

a)

Restricted self-attention type

b)

Self-Attention layer type

c)

Recurrent layer type

d)

Convolutional layer type

37.

For recurrent layers, what is the minimum number of sequential operations per layer?

a)

O(n) sequential operations

b)

O(1) sequential operations

c)

O(log n) sequential operations

d)

O(n2)O(n^2) sequential operations

38.

Which layer type can have maximum path length O(1) between any two positions?

a)

Self-Attention layer type

b)

Recurrent layer type

c)

Convolutional layer type

d)

Restricted self-attention type

39.

What is a key advantage of self-attention over recurrent layers when sequence length n is large relative to representation dimension d?

a)

Faster due to lower sequential operations

b)

Better at local pattern detection

c)

Requires fewer parameters overall

d)

Avoids the need for embeddings

40.

In encoder-decoder attention, where do queries originate and what do they attend to?

a)

Queries from FFN attend to previous heads

b)

Queries from encoder attend to decoder inputs

c)

Queries from embeddings attend to positions

d)

Queries from decoder attend to encoder outputs

41.

What masking is applied in decoder self-attention to preserve autoregressive generation?

a)

Set future positions to negative infinity

b)

Drop all past positions randomly

c)

Zero out diagonal attention entries

d)

Normalize logits by sequence length

42.

Which statement about learned versus sinusoidal positional embeddings was observed?

a)

Both produced nearly similar results

b)

Learned required fewer parameters

c)

Learned embeddings always outperformed

d)

Sinusoidal always failed on long sequences

43.

Which dataset was used for English-to-German training in the described experiments?

a)

Europarl v7 corpus

b)

Multi30k captions dataset

c)

IWSLT 2017 TED talks

d)

WMT 2014 full training corpus

e)

WMT 2014 newstest2014 set

44.

What hardware configuration was used to train the models?

a)

Single machine with 8 P100 GPUs

b)

Cluster of 32 K80 GPUs

c)

TPU v2 pod with 64 cores

d)

Single RTX 4090 GPU

e)

CPU-only 64-core server

45.

For base models, what was the total number of training steps?

a)

100,000 steps

b)

300,000 steps

c)

200,000 steps

d)

75,000 steps

e)

50,000 steps

46.

Which optimizer was employed during training?

a)

Adam optimizer

b)

LAMB optimizer

c)

SGD with momentum

d)

AdaGrad optimizer

e)

RMSProp optimizer

47.

Which hyperparameters were used for Adam in these experiments?

a)

β1=0.9 β2=0.98 ε=1e−9

b)

β1=0.9 β2=0.9 ε=1e−10

c)

β1=0.99 β2=0.999 ε=1e−8

d)

β1=0.9 β2=0.999 ε=1e−8

e)

β1=0.95 β2=0.97 ε=1e−9

48.

How was the learning rate scheduled over training?

a)

Step decay at fixed intervals

b)

Cosine annealing with restarts

c)

Cyclical triangular schedule

d)

Exponential decay from the start

e)

Linear warmup then inverse square root decay

49.

Which regularization method was applied to the output of each sub-layer?

a)

Weight decay L2 penalty

b)

Stochastic depth on blocks

c)

Residual dropout before addition

d)

Spatial dropout on embeddings

e)

Zoneout in recurrent units

50.

What residual dropout rate was used for the base model?

a)

p=0.3

b)

p=0.2

c)

p=0.1

d)

p=0.4

e)

p=0.05

51.

During training, what label smoothing value was employed?

a)

εls=0.3

b)

εls=0.4

c)

εls=0.05

d)

εls=0.1

e)

εls=0.2

52.

On WMT14 English-to-German, what BLEU did the big Transformer achieve on newstest2014?

a)

28.4 BLEU

b)

30.5 BLEU

c)

32.1 BLEU

d)

26.0 BLEU

e)

27.3 BLEU

53.

Compared to previous state-of-the-art, how did the big model perform on English-to-French?

a)

Higher BLEU at fraction of cost

b)

Equal BLEU with same cost

c)

Similar BLEU but higher cost

d)

Lower BLEU and higher cost

e)

Unchanged BLEU across models

54.

What inference beam search settings were used for big models’ checkpoint averaging?

a)

Beam size 6 length penalty 1.0

b)

Beam size 4 length penalty 0.6

c)

Beam size 8 length penalty 0.8

d)

Beam size 2 length penalty 0.4

e)

Beam size 1 length penalty 0.0

55.

Which change to the attention key size dk most negatively impacted model quality in experiments?

a)

Reducing dk notably

b)

Keeping dk unchanged

c)

Randomly varying dk

d)

Increasing dk moderately

56.

What training strategy helped avoid overfitting in larger Transformer variants?

a)

Smaller batch size

b)

Lower learning rate

c)

Higher dropout rate

d)

More epochs

57.

Replacing sinusoidal positional encoding with learned positional embeddings produced what effect on results?

a)

Major improvement overall

b)

Slight improvement only

c)

Significantly worse scores

d)

Nearly identical results

58.

Compared to recurrent or convolutional architectures, the Transformer can be trained how for translation tasks?

a)

Significantly faster overall

b)

About the same speed

c)

Significantly slower overall

d)

Only faster on small data

59.

On WMT14 English-to-German, what performance claim is made for the Transformer?

a)

Matches prior baselines

b)

Sets a new state of the art

c)

Underperforms prior ensembles

d)

Improves only BLEU, not PPL

60.

Which observation suggests exploring a compatibility function more sophisticated than dot product?

a)

Increasing model depth helped

b)

Reducing dk hurt quality

c)

Learned positions improved

d)

Dropout decreased BLEU

61.

What future direction aims to handle large inputs and outputs across modalities like images and audio?

a)

Restricted attention mechanisms

b)

Wider embedding vectors

c)

Longer training schedules

d)

Deeper feedforward layers

62.

Where is the code used to train and evaluate the models available?

a)

pytorch/examples repo

b)

github.com/tensorflow/tensor2tensor

c)

tensorflow/models repo

d)

apache/incubator/mxnet