Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Transformer Architecture and LLMs Worksheet

Total questions: 15

Worksheet time: 8mins

Name
Class
Date
1.

What key advantage made Transformers superior to RNNs?

a)

They use fewer parameters

b)

They process tokens in parallel

c)

They ignore long dependencies

2.

Which layer allows Transformers to identify relationships between all token pairs?

a)

Feed Forward Network

b)

Self-Attention

c)

Residual Block

3.

What do Encoders mainly focus on in Transformer architecture?

a)

Text generation

b)

Understanding and encoding inputs

c)

Adding positional values

4.

What is the main role of the Decoder?

a)

Convert text to embeddings

b)

Generate the next output token

c)

Normalize inputs

5.

Multi-Head Attention helps the model:

a)

Learn only syntax

b)

Learn multiple relationships in parallel

c)

Reduce parameters

6.

Positional Encoding is needed because:

a)

Transformers have no sense of word order

b)

It increases model speed

c)

It reduces overfitting

7.

BERT is best described as:

a)

Decoder-only model

b)

Encoder-only model

c)

Encoder-Decoder model

8.

GPT models are trained using which objective?

a)

Masked Language Modeling

b)

Next Word Prediction

c)

Text-to-Text Mapping

9.

T5 reformulates tasks into:

a)

Image-to-text

b)

Text-to-text

c)

Token-to-embedding

10.

Granite models are known for:

a)

Being closed-source

b)

Being enterprise-grade and transparent

c)

Only supporting small datasets

11.

Pre-training in LLMs involves:

a)

Learning domain-specific data

b)

Learning general language patterns

c)

Generating final responses

12.

Fine-tuning is used to:

a)

Make the model multilingual

b)

Adapt the model to a specific task

c)

Reduce training cost

13.

Layer Normalization helps with:

a)

Stabilizing training

b)

Increasing dataset size

c)

Improving token length

14.

Which component prevents vanishing gradients?

a)

Attention Masking

b)

Residual Connections

c)

Positional Encoding

15.

LLMs like GPT, BERT, and T5 are built on:

a)

RNN architecture

b)

Transformer architecture

c)

Rule-based NLP