WorksheetsNLP Worksheet Questions (Grade 13)
Total questions: 30
Worksheet time: 15mins
Which of the following best defines Natural Language Processing (NLP)?
Processing numeric data using algorithms
Enabling computers to understand, interpret, and generate human language
Converting speech to signals
Encrypting textual data
Which task is NOT a core NLP task?
Machine Translation
Text Summarization
Image Segmentation
Sentiment Analysis
Which step is usually performed first in text preprocessing?
Lemmatization
Tokenization
POS Tagging
Parsing
Removing punctuation and converting text to lowercase is part of:
Parsing
Text normalization
Syntax analysis
Semantic analysis
Word tokenization refers to:
Assigning grammatical tags
Splitting text into meaningful units
Reducing words to root forms
Identifying named entities
Which sentence is most challenging for tokenization?
“Cats are cute.”
“I love NLP.”
“Don’t stop believing.”
“Dogs bark loudly.”
Word normalization aims to:
Increase vocabulary size
Convert words into a standard form
Identify sentence boundaries
Count word frequencies
Which is an example of normalization?
“running” → “run”
“U.S.A.” → “usa”
Sentence splitting
POS tagging
Key difference between stemming and lemmatization is that lemmatization:
Uses only suffix removal
Is faster but less accurate
Uses linguistic knowledge
Produces non-dictionary words
Output of Porter Stemmer for “studies” is:
study
studi
stud
studies
Sentence Segmentation: Sentence segmentation is the task of:
Dividing text into words
Dividing text into sentences
Removing stopwords
Assigning POS tags
Sentence Segmentation: Which symbol commonly causes ambiguity in sentence segmentation?
Comma
Question mark
Period
Exclamation mark
Minimum Edit Distance: Minimum Edit Distance measures:
Semantic similarity
Syntactic correctness
String similarity based on edit operations
Frequency similarity
Minimum Edit Distance: Which operations are used in edit distance calculation?
Merge, split, swap
Insert, delete, substitute
Add, multiply, divide
Encode, decode, parse
N-Gram Language Models: A bigram language model estimates:
P(w)
P(wn | wn-1)
P(wn | wn-2)
P(sentence)
N-Gram Language Models: Increasing the value of n in an n-gram model generally:
Reduces sparsity
Increases context and sparsity
Eliminates unseen words
Removes smoothing
Evaluation of Language Models: Which metric is commonly used to evaluate language models?
Accuracy
Precision
Perplexity
Recall
Evaluation of Language Models: Lower perplexity indicates:
Poor model performance
Higher uncertainty
Better language model
Overfitting
Generalization: Generalization in NLP refers to:
Memorizing training data
Performing well on unseen data
Increasing training accuracy
Reducing vocabulary size
Generalization: Overfitting negatively affects:
Training performance
Model generalization
Vocabulary
Tokenization
Why is smoothing required in language models?
To speed up training
To remove stopwords
To assign non-zero probability to unseen n-grams
To normalize text
Which is a smoothing technique?
Maximum Likelihood Estimation
Add-One (Laplace) Smoothing
Tokenization
Lemmatization
POS tagging assigns:
Sentence boundaries
Word meanings
Grammatical categories to words
Entity labels
The word “book” is ambiguous because it can be:
Only a noun
Only a verb
Both noun and verb
Only an adjective
Named Entity Recognition identifies:
Verbs and nouns
Sentence boundaries
Proper names like person, location, organization
Stopwords
Which is a named entity?
quickly
happiness
New Delhi
beautiful
In HMM-based POS tagging, hidden states represent:
Words
POS tags
Sentences
Characters
Which algorithm is commonly used to find the best tag sequence in HMM?
Forward algorithm
Backward algorithm
Viterbi algorithm
CYK algorithm
Which assumption is made by Hidden Markov Model–based POS tagging?
Each word depends on all previous words
Each tag depends only on the previous tag
Each word is independent of its tag
Tags are observed directly
In Hidden Markov Model–based POS tagging, the emission probability represents:
Probability of a tag given the previous tag
Probability of a word given a POS tag
Probability of a sentence being grammatical
Probability of transitioning between sentences
