Font size
WorksheetsData Mining
Total questions: 35
Worksheet time: 35mins
1.) ______ is a generalization of both Euclidean distance and Manhattan distance.
Minkowski distance
city block distance
absolute distance
Tanimoto distance
2.) A ___________is a generalization of the binary variable in that it can take on more than two states.
ordinal variable
ratio-scaled variable
categorical variable
None of these options
3.) A ratio-scaled variable makes a _______ measurement on a nonlinear scale.
Negative
Positive
Both positive and negative
Neither positive nor negative
4.) The agglomerative approach, also called the _______
divisive approach
Top-down approach
Bottom-up approach
None of these options
5.) _______ is a typical example of a grid-based method.
DBSCAN
OPTICS
STING
DENCLUE
6.)______ distance, is frequently used in information retrieval and biology taxonomy.
Euclidean
Manhattan
Minkowski
Tanimoto
7.)A _____ variable has only two states (1,0).
Ratio-Scaled Variables
Binary
Ordinal
None of these options
8.) ____is a clustering approach that performs clustering by incorporation of user-specified or application-oriented constraints.
Density-based
Model-based
Constraint-based
Neither positive nor negative
9.) _____ methods quantize the object space into a finite number of cells that form a grid structure.
Density-based
Model-based
Hierarchical-based
Grid-based
10.) ____is an algorithm that performs expectation-maximization analysis based on statistical modeling.
COBWEB
EM
SOM
SNIG
11.) Clustering is the process of grouping the data into (a)
12.)Clustering can also be used for (a) .
13.)clustering is a form of (a) , rather than learning by examples .
14.) (a) is a another clustering methodology, extracts distinct frequent patterns among subsets of dimensions that occur frequently.
15.) (a) type algorithm called CLARANS.
16.) The distance between two binary variables based on the notion of (a) .
17.) (a) was one of the first k-medoids algorithms.
18.)Clustering is also called (a) .
19.) (a) is frequently used in information retrieval and biology taxonomy.
20.)In machine learning, clustering is an example of (a)
21.) The k-medoids method is more robust than k-means in the presence of noise and outliers.
True
False
22.) There are three methods to handle ratio-scaled variables for computing the dissimilarity between objects.
True
False
23.) Data matrix is often called a one-mode matrix.
True
False
24.) A ratio-scaled variable makes a negative measurement on a nonlinear scale.
True
False
25.) WaveCluster applies wavelet transformation for clustering analysis and is both grid-based and density-based
True
False
26.) The k-means method is suitable for discovering clusters with nonconvex shapes or clusters of very different size.
True
False
27.) The main advantage of grid-based approach is its slow processing time .
True
False
28.) CLARANS also enables the detection of outliers .
True
False
29.) The k-means and the k-modes methods cannot be integrated to cluster data with mixed numeric and categorical values.
True
False
30.) PAM works effectively for large data sets .
True
False
31.)Match the following:
Agglomerative approach --
Centeroid based technique
Representative object-based technique
Bottom- up approach
Top-down approach
Object-by-object structure
32.)Match the following:
Divisive approach --
Centeroid based technique
Top-down approach
Bottom- up approach
Representative object-based technique
Object-by-object structure
33.)Match the following:
K-means method --
Centeroid based technique
Representative object-based technique
Top-down approach
Bottom- up approach
Object-by-object structure
34.)Match the following:
K-medoid method --
Centeroid based technique
Object-by-object structure
Representative object-based technique
Top-down approach
Bottom- up approach
35.)Match the following:
Dissimilarity matrix --
Object-by-object structure
Centeroid based technique
Representative object-based technique
Top-down approach
Bottom- up approach
