NEW
Font size
WorksheetsNLP CLASS TEST - 2
Total questions: 65
Worksheet time: 34mins
ELIZA mainly worked using:
Deep neural networks
Semantic parsing
World knowledge
Pattern matching
ELIZA gave the illusion of intelligence because:
It understood emotions
It reasoned logically
It had memory
Replies sounded relevant
ELIZA did NOT actually understand:
Grammar
Turn-taking
Dialogue flow
Meaning of words
ELIZA rules mainly used:
Parsing trees
Knowledge graphs
Probabilities
IF–THEN patterns
In NLP, words are treated as:
Emotions
Concepts
Sentences
Data units
Large language models usually treat punctuation as:
Noise
Ignored
Errors
Separate tokens
An utterance is:
Only a written sentence
Grammar rule
Paragraph
Continuous spoken language
“uh” and “um” are examples of:
Prefixes
Tokens
Corpus
Fillers
Speech disfluencies include:
Only pauses
Only repetition
Only grammar errors
False starts and fillers
A corpus is:
One document
Grammar rule set
Vocabulary list
Large collection of text/speech
Word types refer to:
All tokens
Word positions
Sentences
Unique words
Vocabulary size equals:
Tokens
Sentences
Characters
Number of word types
Word instances count:
Unique words
Characters
Lines
Every occurrence
Morphemes are:
Letters
Sounds
Sentences
Minimal meaning units
“un-happy” consists of:
Two characters
Two tokens
Two words
Two morphemes
Free morphemes can:
Attach only
Change tense
Modify grammar
Stand alone
“-ed” is a:
Root
Free morpheme
Word
Bound morpheme
Inflectional morphemes change:
Word class
Meaning completely
Vocabulary
Grammar form
Derivational morphemes:
Only pluralize
Only tense
Never change class
Create new words
Tokenization means:
Parsing grammar
Translating
Tagging
Breaking text into units
Tokenization errors affect:
Only parsing
Storage
Fonts
All NLP tasks
Tokenization depends mainly on:
CPU speed
Dataset size
Randomness
Task and language
Regex stands for:
Rule generator
Grammar engine
Parser
Regular expression
Regex is used to:
Translate
Learn meaning
Classify sentiment
Match patterns
Pattern cat will match:
catalog
cats
scatter
cat
Wildcard . matches:
Nothing
Digits only
Letters only
Any single character
Regex [aeiou] matches:
Any word
Numbers
Punctuation
A vowel
[0-9]+ is used to match:
Names
Symbols
Emails
Numbers
[a-zA-Z]+ matches:
Digits
Symbols
Punctuation
Alphabetic words
Tokenization is:
Universal
Fixed
Random
Design choice
A filler may help speech recognition by:
Ending sentences
Creating noise
Removing words
Predicting restarts
Fragment “main-” is:
Token
Word type
Root
Disfluency
Regex for punctuation [.,!?] matches:
Words
Numbers
Emails
Symbols
Tokenization is foundation of:
Hardware
Storage
Displays
NLP pipelines
Corpus examples include:
RAM
CPU
Token
News articles
Word “cats” contains:
One morpheme
Four tokens
Letters only
Two morphemes
“teacher” shows:
Inflection only
Prefixing
Plurality
Derivation
Regex helps mainly with:
Semantics
Pragmatics
Reasoning
Preprocessing
In ELIZA, sadness was:
Understood
Reasoned
Stored
Pattern-matched
Sentence vs utterance differs in:
Meaning
Tokenization
Grammar
Speech vs writing
Counting punctuation depends on:
Font
Storage
Language only
Task
Morpheme study is called:
Syntax
Phonology
Pragmatics
Morphology
Word instance example “the” twice counts as:
One
Two types
Zero
Two instances
Regex email example contains:
Only letters
Spaces
Slashes
@ symbol
Disfluencies may sometimes be:
Removed always
Ignored fully
Forbidden
Kept for prediction
Tokenization precedes:
Printing
Saving
Display
Sentiment analysis
Pattern-based chatbots rely on:
Knowledge bases
Planning
Learning
Rules
Corpus vocabulary is denoted by:
T
C
N
V
Regex describes:
Exact sentence
Meaning
Tree
String patterns
Bound morphemes cannot:
Change grammar
Modify meaning
Add tense
Stand alone
Tokenization example “don’t → do + n’t” shows:
Lemmatization
Parsing
Stemming
Task-dependent split
Utterances include:
Only paragraphs
Only books
Only essays
Single words
NLP pipeline always begins with:
Translation
Parsing
Learning
Tokens
Regex used for searching text is:
Random
Statistical
Neural
Rule-based
“play → played” uses:
Derivation
Prefix
Root change
Inflection
“happy → unhappy” shows:
Inflection
Tokenization
Plurality
Derivation
Speech systems must handle:
Only silence
Only grammar
Only tokens
Disfluencies
Corpus size affects:
Fonts
Grammar
Punctuation
Vocabulary
Morphemes differ from characters because they:
Are letters
Are sounds
Are punctuation
Carry meaning
Regex [a-z]+@[a-z]+\.com matches:
URL
Number
Date
The Kleene closure of a language L, written as L*, represents:
Zero or more concatenations of strings from L
Only one repetition of strings in L
All strings of length exactly two from L
All strings formed by infinite repetition
The Kleene positive closure of a language L, written as L⁺, represents:
Zero or more repetitions
One or more concatenations of strings from L
Only the empty string
At most two repetitions
If L = {ab}, then L⁺ contains:
b
ε
ab
ab, abab, ababab, …
Tokenization is described as a design choice in NLP mainly because:
It depends on the language, application, and task requirements
All languages follow identical grammar rules
Computers cannot store raw text
Tokens must always be characters
Which of the following tokenizations best reflects a task-dependent NLP design choice as described in the notes?
don’t → don + t
New-York → New-York
don’t → do + n’t
don’t → dont
