Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

PEIS006 DM Final Quiz No. 2

Total questions: 15

Worksheet time: 8mins

Name
Class
Date
1.

Which of the following is NOT a characteristic of unsupervised learning algorithms?

a)

They can be used to categorize data into clusters.

b)

They rely on labeled data to train the model.

c)

They can help in anomaly detection

d)

They are used to find hidden patterns in data.

2.

What is the purpose of the silhouette score in clustering algorithms?

a)

To assign data points to the nearest cluster center.

b)

To visualize the clusters.

c)

To find the optimal number of clusters.

d)

To evaluate how well each cluster is separated.

3.

What is the key difference between supervised and unsupervised learning?

a)

Supervised learning uses labeled data, while unsupervised learning does not.

b)

Unsupervised learning requires more data than supervised learning.

c)

Supervised learning can handle both structured and unstructured data, while unsupervised cannot.

d)

Unsupervised learning is slower than supervised learning.

4.

Which of the following best describes unsupervised learning?

a)

The algorithm uses feedback from its predictions to improve.

b)

The algorithm makes predictions based on historical data.

c)

The algorithm tries to find patterns in unlabeled data.

d)

The algorithm is trained using labeled data.

5.

What is a dendrogram?

a)

A tree-like diagram that represents the hierarchical relationships between clusters

b)

A table displaying the distance between clusters

c)

A statistical model that generates predictions

d)

A chart that shows the predicted clusters

6.

What is the main advantage of hierarchical clustering over KMeans?

a)

It works better with large datasets

b)

It is faster

c)

It doesn’t require the number of clusters to be specified beforehand

d)

It handles only numerical data

7.

In Python's scipy.cluster.hierarchy, which function is used to create a dendrogram?

a)

hierarchical_plot()

b)

generate_dendrogram()

c)

plot_clusters()

d)

dendrogram()

8.

What is the main difference between agglomerative and divisive hierarchical clustering?

a)

Agglomerative uses KMeans, while divisive uses KNN

b)

Agglomerative starts with separate points and merges them, while divisive starts with one cluster and splits it

c)

Agglomerative uses fixed clusters, and divisive dynamically adjusts the number of clusters

d)

Agglomerative starts with one cluster, and divisive starts with all points as separate clusters

9.

Which of the following best describes hierarchical clustering?

a)

It works only with numerical data

b)

It uses a fixed centroid for each cluster

c)

It creates a hierarchy of clusters, which can be represented in a dendrogram

d)

It requires the number of clusters to be specified in advance

10.

In Python, what does the fit() function do when applied to KMeans?

a)

It trains the model and predicts cluster labels

b)

It assigns data points to the nearest centroid

c)

It initializes centroids

d)

It calculates the distance between points and centroids

11.

What is the role of the n_init parameter in KMeans clustering in Python?

a)

It specifies the number of times the algorithm will be run with different initial centroids

b)

It determines the number of nearest neighbors to check

c)

It controls the maximum number of clusters

d)

It sets the minimum distance between centroids

12.

How do you determine the optimal number of clusters for KMeans?

a)

By using a predefined value for K

b)

By trial and error

c)

By looking at the data distribution

d)

By using the elbow method

13.

Which Python function is used to apply KMeans clustering from the scikit-learn library?

b)

KMeans.cluster()

c)

KMeans.fit_predict()

d)

KMeans.cluster_centers()

14.

Which of the following is true about the KMeans algorithm?

a)

It can handle categorical data efficiently

b)

It always produces optimal clusters

c)

It requires the number of clusters to be specified in advance

d)

It is a supervised learning algorithm

15.

What is the primary goal of the KMeans algorithm?

a)

To divide data into a set number of clusters

b)

To generate new data points

c)

To reduce the number of features in data

d)

To analyze time-series data