Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

NLP Quiz 3

Total questions: 30

Worksheet time: 15mins

Name
Class
Date
1.

The main goal of POS tagging is to:

a)

Assign grammatical category labels to words

b)

Assign syntactic trees to sentences

c)

Identify named entities

d)

Identify sentence boundaries

2.

In the Penn Treebank tagset, the tag ‘NNPS’ corresponds to:

a)

Common noun, singular

b)

Proper noun, plural

c)

Possessive pronoun

d)

Proper noun, singular

3.

The transformation-based learning (TBL) approach for POS tagging was proposed by:

a)

Dan Jurafsky

b)

Christopher Manning

c)

Eric Brill

d)

Noam Chomsky

4.

In TBL tagging, the system:

a)

Learns a sequence of rule transformations from an initial tagging

b)

Iteratively applies random transformations

c)

Computes emission probabilities

d)

Uses neural embeddings directly

5.

In Hidden Markov Model (HMM) POS tagging, transition probabilities represent:

a)

P(word | tag)

b)

P(tag | word)

c)

P(tag | previous tag)

d)

P(tag | sentence)

6.

The Viterbi algorithm used in HMM tagging finds:

a)

The most probable tag sequence for a given word sequence

b)

The maximum emission probability

c)

The lexicon probabilities

d)

The shortest parse path

7.

A limitation of rule-based POS taggers is that they:

a)

Depend heavily on manually crafted linguistic rules

b)

Are probabilistic

c)

Require gradient descent

d)

Ignore morphology

8.

Neural POS tagging models replace discrete word features with:

a)

Hand-engineered rules

b)

Binary tag matrices

c)

Continuous vector representations

d)

Token-based dictionaries

9.

Bidirectional LSTM models outperform HMM taggers because they:

a)

Use global sentence-level dependencies

b)

Are deterministic

c)

Require no corpus

d)

Work only for English

10.

A major advantage of neural POS taggers is that they can:

a)

Learn context-sensitive features automatically

b)

Ignore training data

c)

Replace dependency parsers

d)

Achieve unsupervised tagging

11.

A Context-Free Grammar (CFG) consists of:

a)

Terminals, non-terminals, production rules, and a start symbol

b)

Probability matrices and embeddings

c)

Sentences and lexicons only

d)

Annotated dependency graphs

12.

Top-down parsing starts from:

a)

The non-terminal start symbol

b)

Lexical entries

c)

Terminal symbols

d)

Probability matrices

13.

Bottom-up parsing differs by:

a)

Building parse trees from terminals upward

b)

Matching grammar rules backward

c)

Expanding from start symbol downward

d)

Ignoring non-terminals

14.

The CKY algorithm is applicable only to CFGs in:

a)

Greibach Normal Form

b)

Regular Grammar Form

c)

Chomsky Normal Form

d)

Dependency Normal Form

15.

The time complexity of CKY parsing for a sentence of length n is:

a)

O(n⁴)

b)

O(n log n)

c)

O(n³)

d)

O(n²)

16.

A Treebank such as the Penn Treebank provides:

a)

Annotated syntactic structures for sentences

b)

Vectorized token embeddings

c)

Raw, unparsed corpora

d)

Word frequency distributions

17.

A Probabilistic Context-Free Grammar (PCFG) adds which component to a CFG?

a)

Neural word embeddings

b)

Probability distributions over rule expansions

c)

Context constraints

d)

Semantic annotations

18.

In a PCFG, the sum of probabilities of all productions expanding a non-terminal must equal:

a)

1

b)

log(1/p)

c)

0

d)

n

19.

Probabilistic CKY parsing differs from deterministic CKY by:

a)

Computing the highest-probability parse tree

b)

Ignoring syntactic ambiguity

c)

Removing non-terminals

d)

Using only lexicalized grammars

20.

One advantage of PCFG-based parsing over simple CFG parsing is:

a)

Handling syntactic ambiguity quantitatively

b)

Eliminating recursion

c)

Guaranteeing single parse trees

d)

Producing dependency graphs directly

21.

The distributional hypothesis underlying vector semantics states that:

a)

Words that occur in similar contexts have similar meanings

b)

Words are semantically independent

c)

Syntax alone defines meaning

d)

Semantic vectors are random

22.

In a term-document matrix, each cell typically stores:

a)

Word frequency or tf-idf weight

b)

Word embeddings

c)

POS tag counts

d)

Parse tree depth

23.

The cosine similarity between two word vectors measures:

a)

Directional similarity

b)

Logarithmic distance

c)

Frequency variance

d)

Euclidean magnitude only

24.

Singular Value Decomposition (SVD) in Latent Semantic Analysis helps to:

a)

Reduce noise and dimensionality

b)

Expand corpus vocabulary

c)

Normalize text format

d)

Generate parse structures

25.

In Latent Semantic Analysis, latent dimensions represent:

a)

Underlying semantic concepts

b)

POS categories

c)

Morphological forms

d)

Dependency edges

26.

The Skip-gram model in Word2Vec learns to:

a)

Predict surrounding context words from a target word

b)

Predict a word from its context

c)

Predict sentence length

d)

Predict tag sequence

27.

The CBOW model differs from Skip-gram because it:

a)

Predicts target word given context words

b)

Uses context to predict next sentence

c)

Computes global co-occurrence

d)

Uses hierarchical clustering

28.

Word embeddings capture semantic regularities because:

a)

Similar words are mapped to nearby vectors

b)

Each word has a unique orthogonal vector

c)

Word order is ignored

d)

They are manually constructed

29.

In WordNet, the basic unit of semantic organization is:

a)

Synset (set of cognitive synonyms)

b)

Document frequency table

c)

Lemmatized corpus

d)

Embedding layer

30.

Word sense disambiguation refers to:

a)

Determining which meaning of a word is used in context

b)

Mapping POS tags

c)

Parsing dependency relations

d)

Translating between languages