WorksheetsNLP Quiz 3
Total questions: 30
Worksheet time: 15mins
The main goal of POS tagging is to:
Assign grammatical category labels to words
Assign syntactic trees to sentences
Identify named entities
Identify sentence boundaries
In the Penn Treebank tagset, the tag ‘NNPS’ corresponds to:
Common noun, singular
Proper noun, plural
Possessive pronoun
Proper noun, singular
The transformation-based learning (TBL) approach for POS tagging was proposed by:
Dan Jurafsky
Christopher Manning
Eric Brill
Noam Chomsky
In TBL tagging, the system:
Learns a sequence of rule transformations from an initial tagging
Iteratively applies random transformations
Computes emission probabilities
Uses neural embeddings directly
In Hidden Markov Model (HMM) POS tagging, transition probabilities represent:
P(word | tag)
P(tag | word)
P(tag | previous tag)
P(tag | sentence)
The Viterbi algorithm used in HMM tagging finds:
The most probable tag sequence for a given word sequence
The maximum emission probability
The lexicon probabilities
The shortest parse path
A limitation of rule-based POS taggers is that they:
Depend heavily on manually crafted linguistic rules
Are probabilistic
Require gradient descent
Ignore morphology
Neural POS tagging models replace discrete word features with:
Hand-engineered rules
Binary tag matrices
Continuous vector representations
Token-based dictionaries
Bidirectional LSTM models outperform HMM taggers because they:
Use global sentence-level dependencies
Are deterministic
Require no corpus
Work only for English
A major advantage of neural POS taggers is that they can:
Learn context-sensitive features automatically
Ignore training data
Replace dependency parsers
Achieve unsupervised tagging
A Context-Free Grammar (CFG) consists of:
Terminals, non-terminals, production rules, and a start symbol
Probability matrices and embeddings
Sentences and lexicons only
Annotated dependency graphs
Top-down parsing starts from:
The non-terminal start symbol
Lexical entries
Terminal symbols
Probability matrices
Bottom-up parsing differs by:
Building parse trees from terminals upward
Matching grammar rules backward
Expanding from start symbol downward
Ignoring non-terminals
The CKY algorithm is applicable only to CFGs in:
Greibach Normal Form
Regular Grammar Form
Chomsky Normal Form
Dependency Normal Form
The time complexity of CKY parsing for a sentence of length n is:
O(n⁴)
O(n log n)
O(n³)
O(n²)
A Treebank such as the Penn Treebank provides:
Annotated syntactic structures for sentences
Vectorized token embeddings
Raw, unparsed corpora
Word frequency distributions
A Probabilistic Context-Free Grammar (PCFG) adds which component to a CFG?
Neural word embeddings
Probability distributions over rule expansions
Context constraints
Semantic annotations
In a PCFG, the sum of probabilities of all productions expanding a non-terminal must equal:
1
log(1/p)
0
n
Probabilistic CKY parsing differs from deterministic CKY by:
Computing the highest-probability parse tree
Ignoring syntactic ambiguity
Removing non-terminals
Using only lexicalized grammars
One advantage of PCFG-based parsing over simple CFG parsing is:
Handling syntactic ambiguity quantitatively
Eliminating recursion
Guaranteeing single parse trees
Producing dependency graphs directly
The distributional hypothesis underlying vector semantics states that:
Words that occur in similar contexts have similar meanings
Words are semantically independent
Syntax alone defines meaning
Semantic vectors are random
In a term-document matrix, each cell typically stores:
Word frequency or tf-idf weight
Word embeddings
POS tag counts
Parse tree depth
The cosine similarity between two word vectors measures:
Directional similarity
Logarithmic distance
Frequency variance
Euclidean magnitude only
Singular Value Decomposition (SVD) in Latent Semantic Analysis helps to:
Reduce noise and dimensionality
Expand corpus vocabulary
Normalize text format
Generate parse structures
In Latent Semantic Analysis, latent dimensions represent:
Underlying semantic concepts
POS categories
Morphological forms
Dependency edges
The Skip-gram model in Word2Vec learns to:
Predict surrounding context words from a target word
Predict a word from its context
Predict sentence length
Predict tag sequence
The CBOW model differs from Skip-gram because it:
Predicts target word given context words
Uses context to predict next sentence
Computes global co-occurrence
Uses hierarchical clustering
Word embeddings capture semantic regularities because:
Similar words are mapped to nearby vectors
Each word has a unique orthogonal vector
Word order is ignored
They are manually constructed
In WordNet, the basic unit of semantic organization is:
Synset (set of cognitive synonyms)
Document frequency table
Lemmatized corpus
Embedding layer
Word sense disambiguation refers to:
Determining which meaning of a word is used in context
Mapping POS tags
Parsing dependency relations
Translating between languages
