Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

NLP_B1_19.09

Total questions: 20

Worksheet time: 10mins

Name
Class
Date
1.

N-grams are defined as the combination of N keywords together. How many bi-grams can be generated from the given sentence: Gandhiji is the father of our nation

a)

6

b)

7

c)

9

d)

8

2.

In the sentence, “They bought a blue house”, the underlined part is an example of _____.

a)

Noun phrase

b)

Verb phrase

c)

Adverbial phrase

d)

Prepositional phrase

3.

4-grams are better than trigrams for part-of-speech tagging

a)

False

b)

True

4.

What will be the perplexity value if you calculate the perplexity of an unsmoothed language model on a test corpus with unseen words

a)

Infinity

b)

None of the above

c)

any non-zero value

d)

0

5.

Which of the following NLP tasks use sequential labeling technique

a)

Speech recognition

b)

All of the above

c)

Named Entity Recognition

d)

POS tagging

6.

What is the number of trigrams in a normalized sentence of length n words

a)

n

b)

n-1

c)

n-2

d)

n-3

7.

Which of the following techniques is commonly used for smoothing in n-gram models?

a)

K-means clustering

b)

Latent Dirichlet Allocation (LDA)

c)

Principal Component Analysis (PCA)

d)

Add-one (Laplace) smoothing

8.

What is the primary purpose of an n-gram model in natural language processing?

a)

To classify text into predefined categories

b)

To generate random sequences of text

c)

To estimate the probability of a word sequence

d)

To identify named entities in text

9.

In a bigram model, which probability is used to estimate the likelihood of the word sequence "w1 w2 w3"?

a)


P(w1)⋅P(w2∣w1)⋅P(w3∣w2)

b)


P(w1∣w2)⋅P(w2∣w3)⋅P(w3)

c)


P(w1∣w2,w3)⋅P(w2∣w1)⋅P(w3)

d)

P(w2∣w1)⋅P(w3∣w2)

10.

How does the smoothing technique help in n-gram models?

a)

By reducing the computational complexity

b)

By increasing the number of n-grams in the model

c)

By handling zero probabilities for unseen n-grams

d)

By increasing the corpus size

11.

Which of the following is NOT a common type of token in tokenization?

a)
  • Words

b)
  • Characters

c)
  • Sentences

d)
  • Phrases

12.

What is tokenization in the context of natural language processing?

a)
  • The process of converting text into a series of tokens

b)
  • The process of summarizing a text document

c)
  • The process of encoding text into binary format

d)
  • The process of translating text from one language to another

13.

What is sentence tokenization

a)
  • The process of breaking text into individual words

b)
  • The process of translating sentences into another language

c)
  • The process of encoding text into tokens

d)
  • The process of dividing a text into meaningful sentences

14.

When sentence tokenizing, why might an NLP system use punctuation marks

a)
  • To identify word boundaries

b)
  • To translate text

c)
  • To extract named entities

d)
  • To determine sentence boundaries

15.

What does the Minimum Edit Distance algorithm compute

a)
  • The minimum number of operations required to transform one string into another

b)
  • The maximum similarity score between two strings

c)
  • The maximum number of characters that can be removed to make two strings equal

d)
  • The longest common subsequence between two strings

16.

Which of the following operations is NOT typically considered in the Minimum Edit Distance algorithm

a)
  • Deletion

b)
  • Insertion

c)
  • Transposition

d)
  • Substitution

17.

In the context of the Minimum Edit Distance algorithm, what is the time complexity of the dynamic programming approach?

a)
  • O(m+n)

b)
  • O(n^2)

c)
  • O(m*n)

d)
  • O(n^3)

18.

If the cost of insertion, deletion, and substitution is all equal to 1, which of the following sequences of operations would be optimal for transforming the string "abc" into "yabd"

a)

Insert 'y', replace 'c' with 'd'

b)
  • Replace 'a' with 'y', insert 'd'

c)
  • Insert 'y', insert 'd', replace 'b' with 'a'

d)
  • Replace 'b' with 'd', insert 'y'

19.

In the case where both input strings are identical, what is the minimum edit distance between them

a)
  • The length of the strings

b)
  • The length of the longer string

c)

0

d)

1

20.

Which NLP technique involves analyzing and generating human-like text or speech by computer programs

a)

Text Generation

b)

Dependency Parsing

c)

Text Classification

d)

NER