WorksheetsTransformer Architecture and LLMs Worksheet
Total questions: 15
Worksheet time: 8mins
What key advantage made Transformers superior to RNNs?
They use fewer parameters
They process tokens in parallel
They ignore long dependencies
Which layer allows Transformers to identify relationships between all token pairs?
Feed Forward Network
Self-Attention
Residual Block
What do Encoders mainly focus on in Transformer architecture?
Text generation
Understanding and encoding inputs
Adding positional values
What is the main role of the Decoder?
Convert text to embeddings
Generate the next output token
Normalize inputs
Multi-Head Attention helps the model:
Learn only syntax
Learn multiple relationships in parallel
Reduce parameters
Positional Encoding is needed because:
Transformers have no sense of word order
It increases model speed
It reduces overfitting
BERT is best described as:
Decoder-only model
Encoder-only model
Encoder-Decoder model
GPT models are trained using which objective?
Masked Language Modeling
Next Word Prediction
Text-to-Text Mapping
T5 reformulates tasks into:
Image-to-text
Text-to-text
Token-to-embedding
Granite models are known for:
Being closed-source
Being enterprise-grade and transparent
Only supporting small datasets
Pre-training in LLMs involves:
Learning domain-specific data
Learning general language patterns
Generating final responses
Fine-tuning is used to:
Make the model multilingual
Adapt the model to a specific task
Reduce training cost
Layer Normalization helps with:
Stabilizing training
Increasing dataset size
Improving token length
Which component prevents vanishing gradients?
Attention Masking
Residual Connections
Positional Encoding
LLMs like GPT, BERT, and T5 are built on:
RNN architecture
Transformer architecture
Rule-based NLP
