Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Data Mining and Machine Learning Worksheet (Transcribed MCQs)

Total questions: 41

Worksheet time: 21mins

Name
Class
Date
1.

What is the primary goal of classification in data mining?

a)

Group similar data without labels

b)

Predict a continuous numerical value

c)

Assign data items to pre-defined class labels

d)

Reduce the dimensionality of data

2.

Classification is an example of which type of learning?

a)

Unsupervised learning

b)

Supervised learning

c)

Reinforcement learning

d)

Clustering

3.

In a decision tree what does an internal node represent?

a)

A class label

b)

A test on an attribute

c)

The final outcome

d)

The root of the database

4.

What does a leaf node in a decision tree represent?

a)

A test on a feature

b)

A splitting rule

c)

A class label

d)

An attribute measure

5.

Which measure is used by the ID3 algorithm for splitting?

a)

Gini Index

b)

Information Gain

c)

Euclidean Distance

d)

Support Count

6.

Information gain is based on which concept?

a)

Gini Index

b)

Entropy

c)

Standard deviation

d)

Correlation

7.

What is the main purpose of tree pruning?

a)

To increase tree depth

b)

To add more branches

c)

To reduce overfitting and noise

d)

To increase training time

8.

Which of the following describes post‑pruning?

a)

Halting tree construction early

b)

Removing branches from a grown tree

c)

Selecting the best attribute first

d)

Converting tree to rules

9.

Bayesian classification is based on which mathematical theorem?

a)

Pythagoras theorem

b)

Bayes theorem

c)

Central limit theorem

d)

Taylor theorem

10.

What is the naive assumption in Naive Bayes?

a)

Attributes are dependent

b)

Attributes are independent

c)

Data is normally distributed

d)

Classes are equal

11.

In rule‑based classification, rules are represented in which format?

a)

WHILE‑DO

b)

IF‑THEN

c)

FOR‑LOOP

d)

SWITCH‑CASE

12.

What does SVM stand for?

a)

System Vector Model

b)

Support Vector Machine

c)

Standard Variable Mining

d)

Simple Value Mapping

13.

SVM aims to find a hyperplane with the:

a)

Maximum margin

b)

Minimum error

c)

Maximum depth

d)

Lowest support

14.

Which kernel is commonly used in SVM?

a)

Linear

b)

Polynomial

c)

RBF

d)

Sigmoid

15.

Which metric is the ratio of correct predictions to total instances?

a)

Precision

b)

Recall

c)

Accuracy

d)

F1‑score

16.

A confusion matrix is used to:

a)

Clean the data

b)

Evaluate classifier performance

c)

Store association data

d)

Build a decision tree

17.

What is sensitivity also known as?

a)

Precision

b)

Accuracy

c)

F1‑score

d)

Recall

18.

A false positive error is also known as:

a)

Type I error

b)

Type II error

c)

Standard error

d)

Sampling error

19.

In K‑fold cross validation, data is split into:

a)

Two equal halves

b)

K mutually exclusive subsets

c)

One training set only

d)

Multiple random samples

20.

What characterizes the bootstrap method?

a)

Sampling without replacement

b)

Sampling with replacement

c)

Splitting data in two halves

d)

Using entire dataset twice

21.

What does ROC stand for in performance curves?

a)

Random operating code

b)

Receiver operating characteristic

c)

Root of classification

d)

Range of correlation

22.

What is the main objective of market basket analysis?

a)

Predict stock prices

b)

Find co‑occurring items

c)

Classify customers

d)

Route delivery

23.

Which symbol represents an association rule?

a)

IF X THEN Y

b)

X => Y

c)

X -> Y

d)

All of the above

24.

What does support measure in a rule A -> B?

a)

Percent of transactions with A and B

b)

Probability B given A

c)

Total database size

d)

Weight of item A

25.

Confidence in association rule mining measures:

a)

Data popularity

b)

Certainty or predictability

c)

Algorithm speed

d)

Database size

26.

An itemset satisfying the minimum support threshold is a:

a)

Closed itemset

b)

Maximal itemset

c)

Frequent itemset

d)

Candidate itemset

27.

The Apriori property states all subsets of a frequent itemset must be:

a)

Infrequent

b)

Frequent

c)

Closed

d)

Maximal

28.

If itemset {A B} is infrequent then {A B C} is:

a)

Also frequent

b)

Guaranteed infrequent

c)

Possibly frequent

d)

A candidate

29.

Apriori uses which search approach?

a)

Depth‑first

b)

Level‑wise (Breadth‑first)

c)

Random sampling

d)

Genetic evolution

30.

In Apriori what is the Join Step used for?

a)

Merging databases

b)

Generating Lk from Ck

c)

Generating Ck from Lk‑1

d)

Deleting items

31.

Which algorithm avoids candidate generation using a tree?

a)

Apriori

b)

FP‑Growth

c)

Partitioning

d)

Sampling

32.

In FP‑growth, data is compressed into a:

a)

Hash table

b)

FP‑tree

c)

Linked list

d)

Matrix

33.

A disadvantage of the Apriori algorithm is:

a)

Multiple database scans

b)

Too complex

c)

Small datasets only

d)

No rules found

34.

Formula for confidence (A -> B) is:

a)

Support(A ∪ B) / Support(A)

b)

Support(A)

c)

Support(A ∪ B) / Support(B)

d)

Support(A) / Support(B)

e)

Support(A ∩ B) / T

35.

A rule is strong if it satisfies:

a)

Only minimum support

b)

Only minimum confidence

c)

Both minimum thresholds

d)

Maximum support only

36.

Lift(A B) > 11 means A and B are:

a)

Negatively correlated

b)

Positively correlated

c)

Independent

d)

Not frequent

37.

A frequent itemset with no frequent superset is:

a)

Closed

b)

Maximal

c)

Frequent

d)

Dense

38.

An itemset with no superset having the same support is:

a)

Closed

b)

Maximal

c)

Frequent

d)

Sparse

39.

Eclat algorithm uses which data format?

a)

Horizontal

b)

Vertical (TID‑set)

c)

Circular

d)

Hybrid

40.

Mining patterns at different abstraction levels is:

a)

Quantitative

b)

Multilevel

c)

Constraint‑based

d)

Simple

41.

Correlation analysis helps to:

a)

Speed up mining

b)

Filter boring rules

c)

Find more items

d)

Increase support