Worksheetsytu-yz-yaz NLP EN
Total questions: 14
Worksheet time: 7mins
What is the primary goal of Natural Language Processing (NLP)?
To create new human languages
To enable computers to understand and interact with human language
To replace human communication entirely
To encrypt text data
Which of the following is NOT a common NLP application?
Text summarization
Sentiment analysis
Image recognition
Named entity recognition
In NLP, what does the term "tokenization" refer to?
Encrypting text data
Breaking down text into smaller units, typically words
Combining words into sentences
Translating text from one language to another
What is the main difference between stemming and lemmatization?
Stemming is faster but less accurate, while lemmatization is slower but more accurate
Stemming works only for nouns, while lemmatization works for all parts of speech
Stemming adds suffixes, while lemmatization removes them
There is no difference; they are two terms for the same process
What does TF-IDF stand for in the context of NLP?
Text Formatting-Indirect Document Filtering
Total Findings-Inferred Data Formation
Text Function-Integrated Document Format
Term Frequency-Inverse Document Frequency
Which NLP approach relies on pre-existing lexical resources like dictionaries and thesauri?
Corpus-based approach
Statistical approach
Dictionary-based approach
Deep learning approach
What is the main limitation of the Bag of Words (BoW) model?
It requires too much computational power
It can only be used for short texts
It disregards word order and context
It only works for English language texts
Which of the following is a characteristic of word embeddings like Word2Vec?
They produce sparse, high-dimensional vectors
They generate different representations for a word based on its context
They map semantically similar words to nearby points in a vector space
They require manually labeled training data
What is a key advantage of contextual embeddings (like BERT or GPT) over traditional word embeddings?
They are computationally less intensive
They can handle out-of-vocabulary words better
They generate different representations for a word based on its context
They require less training data
In the context of NLP, what does POS tagging refer to?
Point of Sale tagging
Probability of Sequence tagging
Part of Speech tagging
Parsing of Sentences tagging
N-gram models capture some local context and word order information, unlike the Bag of Words model
True
False
In the sentiment analysis task using a dictionary-based method on books of Jane Austen, what does the process of calculating "positive minus negative words per page"
It determines the overall genre of the book
It identifies the most frequently used words in the book
It calculates the reading difficulty of each page
It quantifies the emotional tone of different parts of the text
In the context of modern AI development, why is unstructured data often referred to as "gold"?
It's rare and difficult to find
It's always more valuable than structured data
It's crucial for training large language models like GPT
It's only useful for financial applications
Consider the sentence "The quick brown fox". Which of the following correctly represents its bigram (2-gram) tokenization?
["The", "quick", "brown", "fox"]
["The quick", "brown fox"]
["The quick", "quick brown", "brown fox"]
["The quick brown", "quick brown fox"]
