Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

ytu-yz-yaz NLP EN

Total questions: 14

Worksheet time: 7mins

Name
Class
Date
1.

What is the primary goal of Natural Language Processing (NLP)?

a)

To create new human languages

b)

To enable computers to understand and interact with human language

c)

To replace human communication entirely

d)

To encrypt text data

2.

Which of the following is NOT a common NLP application?

a)

Text summarization

b)

Sentiment analysis

c)

Image recognition

d)

Named entity recognition

3.

In NLP, what does the term "tokenization" refer to?

a)

Encrypting text data

b)

Breaking down text into smaller units, typically words

c)

Combining words into sentences

d)

Translating text from one language to another

4.

What is the main difference between stemming and lemmatization?

a)

Stemming is faster but less accurate, while lemmatization is slower but more accurate

b)

Stemming works only for nouns, while lemmatization works for all parts of speech

c)

Stemming adds suffixes, while lemmatization removes them

d)

There is no difference; they are two terms for the same process

5.

What does TF-IDF stand for in the context of NLP?

a)

Text Formatting-Indirect Document Filtering

b)

Total Findings-Inferred Data Formation

c)

Text Function-Integrated Document Format

d)

Term Frequency-Inverse Document Frequency

6.

Which NLP approach relies on pre-existing lexical resources like dictionaries and thesauri?

a)

Corpus-based approach

b)

Statistical approach

c)

Dictionary-based approach

d)

Deep learning approach

7.

What is the main limitation of the Bag of Words (BoW) model?

a)

It requires too much computational power

b)

It can only be used for short texts

c)

It disregards word order and context

d)

It only works for English language texts

8.

Which of the following is a characteristic of word embeddings like Word2Vec?

a)

They produce sparse, high-dimensional vectors

b)

They generate different representations for a word based on its context

c)

They map semantically similar words to nearby points in a vector space

d)

They require manually labeled training data

9.

What is a key advantage of contextual embeddings (like BERT or GPT) over traditional word embeddings?

a)

They are computationally less intensive

b)

They can handle out-of-vocabulary words better

c)

They generate different representations for a word based on its context

d)

They require less training data

10.

In the context of NLP, what does POS tagging refer to?

a)

Point of Sale tagging

b)

Probability of Sequence tagging

c)

Part of Speech tagging

d)

Parsing of Sentences tagging

11.

N-gram models capture some local context and word order information, unlike the Bag of Words model

a)

True

b)

False

12.

In the sentiment analysis task using a dictionary-based method on books of Jane Austen, what does the process of calculating "positive minus negative words per page"

a)

It determines the overall genre of the book

b)

It identifies the most frequently used words in the book

c)

It calculates the reading difficulty of each page

d)

It quantifies the emotional tone of different parts of the text

13.

In the context of modern AI development, why is unstructured data often referred to as "gold"?

a)

It's rare and difficult to find

b)

It's always more valuable than structured data

c)

It's crucial for training large language models like GPT

d)

It's only useful for financial applications

14.

Consider the sentence "The quick brown fox". Which of the following correctly represents its bigram (2-gram) tokenization?

a)

["The", "quick", "brown", "fox"]

b)

["The quick", "brown fox"]

c)

["The quick", "quick brown", "brown fox"]

d)

["The quick brown", "quick brown fox"]