WorksheetsData Mining and Machine Learning Worksheet (Transcribed MCQs)
Total questions: 41
Worksheet time: 21mins
What is the primary goal of classification in data mining?
Group similar data without labels
Predict a continuous numerical value
Assign data items to pre-defined class labels
Reduce the dimensionality of data
Classification is an example of which type of learning?
Unsupervised learning
Supervised learning
Reinforcement learning
Clustering
In a decision tree what does an internal node represent?
A class label
A test on an attribute
The final outcome
The root of the database
What does a leaf node in a decision tree represent?
A test on a feature
A splitting rule
A class label
An attribute measure
Which measure is used by the ID3 algorithm for splitting?
Gini Index
Information Gain
Euclidean Distance
Support Count
Information gain is based on which concept?
Gini Index
Entropy
Standard deviation
Correlation
What is the main purpose of tree pruning?
To increase tree depth
To add more branches
To reduce overfitting and noise
To increase training time
Which of the following describes post‑pruning?
Halting tree construction early
Removing branches from a grown tree
Selecting the best attribute first
Converting tree to rules
Bayesian classification is based on which mathematical theorem?
Pythagoras theorem
Bayes theorem
Central limit theorem
Taylor theorem
What is the naive assumption in Naive Bayes?
Attributes are dependent
Attributes are independent
Data is normally distributed
Classes are equal
In rule‑based classification, rules are represented in which format?
WHILE‑DO
IF‑THEN
FOR‑LOOP
SWITCH‑CASE
What does SVM stand for?
System Vector Model
Support Vector Machine
Standard Variable Mining
Simple Value Mapping
SVM aims to find a hyperplane with the:
Maximum margin
Minimum error
Maximum depth
Lowest support
Which kernel is commonly used in SVM?
Linear
Polynomial
RBF
Sigmoid
Which metric is the ratio of correct predictions to total instances?
Precision
Recall
Accuracy
F1‑score
A confusion matrix is used to:
Clean the data
Evaluate classifier performance
Store association data
Build a decision tree
What is sensitivity also known as?
Precision
Accuracy
F1‑score
Recall
A false positive error is also known as:
Type I error
Type II error
Standard error
Sampling error
In K‑fold cross validation, data is split into:
Two equal halves
K mutually exclusive subsets
One training set only
Multiple random samples
What characterizes the bootstrap method?
Sampling without replacement
Sampling with replacement
Splitting data in two halves
Using entire dataset twice
What does ROC stand for in performance curves?
Random operating code
Receiver operating characteristic
Root of classification
Range of correlation
What is the main objective of market basket analysis?
Predict stock prices
Find co‑occurring items
Classify customers
Route delivery
Which symbol represents an association rule?
IF X THEN Y
X => Y
X -> Y
All of the above
What does support measure in a rule A -> B?
Percent of transactions with A and B
Probability B given A
Total database size
Weight of item A
Confidence in association rule mining measures:
Data popularity
Certainty or predictability
Algorithm speed
Database size
An itemset satisfying the minimum support threshold is a:
Closed itemset
Maximal itemset
Frequent itemset
Candidate itemset
The Apriori property states all subsets of a frequent itemset must be:
Infrequent
Frequent
Closed
Maximal
If itemset {A B} is infrequent then {A B C} is:
Also frequent
Guaranteed infrequent
Possibly frequent
A candidate
Apriori uses which search approach?
Depth‑first
Level‑wise (Breadth‑first)
Random sampling
Genetic evolution
In Apriori what is the Join Step used for?
Merging databases
Generating Lk from Ck
Generating Ck from Lk‑1
Deleting items
Which algorithm avoids candidate generation using a tree?
Apriori
FP‑Growth
Partitioning
Sampling
In FP‑growth, data is compressed into a:
Hash table
FP‑tree
Linked list
Matrix
A disadvantage of the Apriori algorithm is:
Multiple database scans
Too complex
Small datasets only
No rules found
Formula for confidence (A -> B) is:
Support(A ∪ B) / Support(A)
Support(A)
Support(A ∪ B) / Support(B)
Support(A) / Support(B)
Support(A ∩ B) / T
A rule is strong if it satisfies:
Only minimum support
Only minimum confidence
Both minimum thresholds
Maximum support only
Lift(A B) > 1 means A and B are:
Negatively correlated
Positively correlated
Independent
Not frequent
A frequent itemset with no frequent superset is:
Closed
Maximal
Frequent
Dense
An itemset with no superset having the same support is:
Closed
Maximal
Frequent
Sparse
Eclat algorithm uses which data format?
Horizontal
Vertical (TID‑set)
Circular
Hybrid
Mining patterns at different abstraction levels is:
Quantitative
Multilevel
Constraint‑based
Simple
Correlation analysis helps to:
Speed up mining
Filter boring rules
Find more items
Increase support
