WorksheetsPEIS006 DM Final Quiz No. 2
Total questions: 15
Worksheet time: 8mins
Which of the following is NOT a characteristic of unsupervised learning algorithms?
They can be used to categorize data into clusters.
They rely on labeled data to train the model.
They can help in anomaly detection
They are used to find hidden patterns in data.
What is the purpose of the silhouette score in clustering algorithms?
To assign data points to the nearest cluster center.
To visualize the clusters.
To find the optimal number of clusters.
To evaluate how well each cluster is separated.
What is the key difference between supervised and unsupervised learning?
Supervised learning uses labeled data, while unsupervised learning does not.
Unsupervised learning requires more data than supervised learning.
Supervised learning can handle both structured and unstructured data, while unsupervised cannot.
Unsupervised learning is slower than supervised learning.
Which of the following best describes unsupervised learning?
The algorithm uses feedback from its predictions to improve.
The algorithm makes predictions based on historical data.
The algorithm tries to find patterns in unlabeled data.
The algorithm is trained using labeled data.
What is a dendrogram?
A tree-like diagram that represents the hierarchical relationships between clusters
A table displaying the distance between clusters
A statistical model that generates predictions
A chart that shows the predicted clusters
What is the main advantage of hierarchical clustering over KMeans?
It works better with large datasets
It is faster
It doesn’t require the number of clusters to be specified beforehand
It handles only numerical data
In Python's scipy.cluster.hierarchy, which function is used to create a dendrogram?
hierarchical_plot()
generate_dendrogram()
plot_clusters()
dendrogram()
What is the main difference between agglomerative and divisive hierarchical clustering?
Agglomerative uses KMeans, while divisive uses KNN
Agglomerative starts with separate points and merges them, while divisive starts with one cluster and splits it
Agglomerative uses fixed clusters, and divisive dynamically adjusts the number of clusters
Agglomerative starts with one cluster, and divisive starts with all points as separate clusters
Which of the following best describes hierarchical clustering?
It works only with numerical data
It uses a fixed centroid for each cluster
It creates a hierarchy of clusters, which can be represented in a dendrogram
It requires the number of clusters to be specified in advance
In Python, what does the fit() function do when applied to KMeans?
It trains the model and predicts cluster labels
It assigns data points to the nearest centroid
It initializes centroids
It calculates the distance between points and centroids
What is the role of the n_init parameter in KMeans clustering in Python?
It specifies the number of times the algorithm will be run with different initial centroids
It determines the number of nearest neighbors to check
It controls the maximum number of clusters
It sets the minimum distance between centroids
How do you determine the optimal number of clusters for KMeans?
By using a predefined value for K
By trial and error
By looking at the data distribution
By using the elbow method
Which Python function is used to apply KMeans clustering from the scikit-learn library?
Which of the following is true about the KMeans algorithm?
It can handle categorical data efficiently
It always produces optimal clusters
It requires the number of clusters to be specified in advance
It is a supervised learning algorithm
What is the primary goal of the KMeans algorithm?
To divide data into a set number of clusters
To generate new data points
To reduce the number of features in data
To analyze time-series data
