Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

NLP Worksheet Questions (Grade 13)

Total questions: 30

Worksheet time: 15mins

Name
Class
Date
1.

Which of the following best defines Natural Language Processing (NLP)?

a)

Processing numeric data using algorithms

b)

Enabling computers to understand, interpret, and generate human language

c)

Converting speech to signals

d)

Encrypting textual data

2.

Which task is NOT a core NLP task?

a)

Machine Translation

b)

Text Summarization

c)

Image Segmentation

d)

Sentiment Analysis

3.

Which step is usually performed first in text preprocessing?

a)

Lemmatization

b)

Tokenization

c)

POS Tagging

d)

Parsing

4.

Removing punctuation and converting text to lowercase is part of:

a)

Parsing

b)

Text normalization

c)

Syntax analysis

d)

Semantic analysis

5.

Word tokenization refers to:

a)

Assigning grammatical tags

b)

Splitting text into meaningful units

c)

Reducing words to root forms

d)

Identifying named entities

6.

Which sentence is most challenging for tokenization?

a)

“Cats are cute.”

b)

“I love NLP.”

c)

“Don’t stop believing.”

d)

“Dogs bark loudly.”

7.

Word normalization aims to:

a)

Increase vocabulary size

b)

Convert words into a standard form

c)

Identify sentence boundaries

d)

Count word frequencies

8.

Which is an example of normalization?

a)

“running” → “run”

b)

“U.S.A.” → “usa”

c)

Sentence splitting

d)

POS tagging

9.

Key difference between stemming and lemmatization is that lemmatization:

a)

Uses only suffix removal

b)

Is faster but less accurate

c)

Uses linguistic knowledge

d)

Produces non-dictionary words

10.

Output of Porter Stemmer for “studies” is:

a)

study

b)

studi

c)

stud

d)

studies

11.

Sentence Segmentation: Sentence segmentation is the task of:

a)

Dividing text into words

b)

Dividing text into sentences

c)

Removing stopwords

d)

Assigning POS tags

12.

Sentence Segmentation: Which symbol commonly causes ambiguity in sentence segmentation?

a)

Comma

b)

Question mark

c)

Period

d)

Exclamation mark

13.

Minimum Edit Distance: Minimum Edit Distance measures:

a)

Semantic similarity

b)

Syntactic correctness

c)

String similarity based on edit operations

d)

Frequency similarity

14.

Minimum Edit Distance: Which operations are used in edit distance calculation?

a)

Merge, split, swap

b)

Insert, delete, substitute

c)

Add, multiply, divide

d)

Encode, decode, parse

15.

N-Gram Language Models: A bigram language model estimates:

a)

P(w)

b)

P(wn | wn-1)

c)

P(wn | wn-2)

d)

P(sentence)

16.

N-Gram Language Models: Increasing the value of n in an n-gram model generally:

a)

Reduces sparsity

b)

Increases context and sparsity

c)

Eliminates unseen words

d)

Removes smoothing

17.

Evaluation of Language Models: Which metric is commonly used to evaluate language models?

a)

Accuracy

b)

Precision

c)

Perplexity

d)

Recall

18.

Evaluation of Language Models: Lower perplexity indicates:

a)

Poor model performance

b)

Higher uncertainty

c)

Better language model

d)

Overfitting

19.

Generalization: Generalization in NLP refers to:

a)

Memorizing training data

b)

Performing well on unseen data

c)

Increasing training accuracy

d)

Reducing vocabulary size

20.

Generalization: Overfitting negatively affects:

a)

Training performance

b)

Model generalization

c)

Vocabulary

d)

Tokenization

21.

Why is smoothing required in language models?

a)

To speed up training

b)

To remove stopwords

c)

To assign non-zero probability to unseen n-grams

d)

To normalize text

22.

Which is a smoothing technique?

a)

Maximum Likelihood Estimation

b)

Add-One (Laplace) Smoothing

c)

Tokenization

d)

Lemmatization

23.

POS tagging assigns:

a)

Sentence boundaries

b)

Word meanings

c)

Grammatical categories to words

d)

Entity labels

24.

The word “book” is ambiguous because it can be:

a)

Only a noun

b)

Only a verb

c)

Both noun and verb

d)

Only an adjective

25.

Named Entity Recognition identifies:

a)

Verbs and nouns

b)

Sentence boundaries

c)

Proper names like person, location, organization

d)

Stopwords

26.

Which is a named entity?

a)

quickly

b)

happiness

c)

New Delhi

d)

beautiful

27.

In HMM-based POS tagging, hidden states represent:

a)

Words

b)

POS tags

c)

Sentences

d)

Characters

28.

Which algorithm is commonly used to find the best tag sequence in HMM?

a)

Forward algorithm

b)

Backward algorithm

c)

Viterbi algorithm

d)

CYK algorithm

29.

Which assumption is made by Hidden Markov Model–based POS tagging?

a)

Each word depends on all previous words

b)

Each tag depends only on the previous tag

c)

Each word is independent of its tag

d)

Tags are observed directly

30.

In Hidden Markov Model–based POS tagging, the emission probability represents:

a)

Probability of a tag given the previous tag

b)

Probability of a word given a POS tag

c)

Probability of a sentence being grammatical

d)

Probability of transitioning between sentences