WorksheetsNLP_B1_19.09
Total questions: 20
Worksheet time: 10mins
N-grams are defined as the combination of N keywords together. How many bi-grams can be generated from the given sentence: Gandhiji is the father of our nation
6
7
9
8
In the sentence, “They bought a blue house”, the underlined part is an example of _____.
Noun phrase
Verb phrase
Adverbial phrase
Prepositional phrase
4-grams are better than trigrams for part-of-speech tagging
False
True
What will be the perplexity value if you calculate the perplexity of an unsmoothed language model on a test corpus with unseen words
Infinity
None of the above
any non-zero value
0
Which of the following NLP tasks use sequential labeling technique
Speech recognition
All of the above
Named Entity Recognition
POS tagging
What is the number of trigrams in a normalized sentence of length n words
n
n-1
n-2
n-3
Which of the following techniques is commonly used for smoothing in n-gram models?
K-means clustering
Latent Dirichlet Allocation (LDA)
Principal Component Analysis (PCA)
Add-one (Laplace) smoothing
What is the primary purpose of an n-gram model in natural language processing?
To classify text into predefined categories
To generate random sequences of text
To estimate the probability of a word sequence
To identify named entities in text
In a bigram model, which probability is used to estimate the likelihood of the word sequence "w1 w2 w3"?
P(w1)⋅P(w2∣w1)⋅P(w3∣w2)
P(w1∣w2)⋅P(w2∣w3)⋅P(w3)
P(w1∣w2,w3)⋅P(w2∣w1)⋅P(w3)
P(w2∣w1)⋅P(w3∣w2)
How does the smoothing technique help in n-gram models?
By reducing the computational complexity
By increasing the number of n-grams in the model
By handling zero probabilities for unseen n-grams
By increasing the corpus size
Which of the following is NOT a common type of token in tokenization?
Words
Characters
Sentences
Phrases
What is tokenization in the context of natural language processing?
The process of converting text into a series of tokens
The process of summarizing a text document
The process of encoding text into binary format
The process of translating text from one language to another
What is sentence tokenization
The process of breaking text into individual words
The process of translating sentences into another language
The process of encoding text into tokens
The process of dividing a text into meaningful sentences
When sentence tokenizing, why might an NLP system use punctuation marks
To identify word boundaries
To translate text
To extract named entities
To determine sentence boundaries
What does the Minimum Edit Distance algorithm compute
The minimum number of operations required to transform one string into another
The maximum similarity score between two strings
The maximum number of characters that can be removed to make two strings equal
The longest common subsequence between two strings
Which of the following operations is NOT typically considered in the Minimum Edit Distance algorithm
Deletion
Insertion
Transposition
Substitution
In the context of the Minimum Edit Distance algorithm, what is the time complexity of the dynamic programming approach?
O(m+n)
O(n^2)
O(m*n)
O(n^3)
If the cost of insertion, deletion, and substitution is all equal to 1, which of the following sequences of operations would be optimal for transforming the string "abc" into "yabd"
Insert 'y', replace 'c' with 'd'
Replace 'a' with 'y', insert 'd'
Insert 'y', insert 'd', replace 'b' with 'a'
Replace 'b' with 'd', insert 'y'
In the case where both input strings are identical, what is the minimum edit distance between them
The length of the strings
The length of the longer string
0
1
Which NLP technique involves analyzing and generating human-like text or speech by computer programs
Text Generation
Dependency Parsing
Text Classification
NER
