wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Data Mining

Total questions: 35

Worksheet time: 35mins

Name
Class
Date
1.

1.) ______ is a generalization of both Euclidean distance and Manhattan distance.

a)

Minkowski distance

b)

city block distance

c)

absolute distance

d)

Tanimoto distance

2.

2.) A ___________is a generalization of the binary variable in that it can take on more than two states.

a)

ordinal variable

b)

ratio-scaled variable

c)

categorical variable

d)

None of these options

3.

3.) A ratio-scaled variable makes a _______ measurement on a nonlinear scale.

a)

Negative

b)

Positive

c)

Both positive and negative

d)

Neither positive nor negative

4.

4.) The agglomerative approach, also called the _______

a)

divisive approach

b)

Top-down approach

c)

Bottom-up approach

d)

None of these options

5.

5.) _______ is a typical example of a grid-based method.

a)

DBSCAN

b)

OPTICS

c)

STING

d)

DENCLUE

6.

6.)______ distance, is frequently used in information retrieval and biology taxonomy.

a)

Euclidean

b)

Manhattan

c)

Minkowski

d)

Tanimoto

7.

7.)A _____ variable has only two states (1,0).

a)

Ratio-Scaled Variables

b)

Binary

c)

Ordinal

d)

None of these options

8.

8.) ____is a clustering approach that performs clustering by incorporation of user-specified or application-oriented constraints.

a)

Density-based

b)

Model-based

c)

Constraint-based

d)

Neither positive nor negative

9.

9.) _____ methods quantize the object space into a finite number of cells that form a grid structure.

a)

Density-based

b)

Model-based

c)

Hierarchical-based

d)

Grid-based

10.

10.) ____is an algorithm that performs expectation-maximization analysis based on statistical modeling.

a)

COBWEB

b)

EM

c)

SOM

d)

SNIG

11.

11.) Clustering is the process of grouping the data into (a)  

12.

12.)Clustering can also be used for (a)   .

13.

13.)clustering is a form of (a)   , rather than learning by examples .

14.

14.) (a)   is a another clustering methodology, extracts distinct frequent patterns among subsets of dimensions that occur frequently.

15.

15.) (a)   type algorithm called CLARANS.

16.

16.) The distance between two binary variables based on the notion of (a)   .

17.

17.) (a)   was one of the first k-medoids algorithms.

18.

18.)Clustering is also called (a)   .

19.

19.) (a)   is frequently used in information retrieval and biology taxonomy.

20.

20.)In machine learning, clustering is an example of (a)  

21.

21.) The k-medoids method is more robust than k-means in the presence of noise and outliers.

a)

True

b)

False

22.

22.) There are three methods to handle ratio-scaled variables for computing the dissimilarity between objects.

a)

True

b)

False

23.

23.) Data matrix is often called a one-mode matrix.

a)

True

b)

False

24.

24.) A ratio-scaled variable makes a negative measurement on a nonlinear scale.

a)

True

b)

False

25.

25.) WaveCluster applies wavelet transformation for clustering analysis and is both grid-based and density-based

a)

True

b)

False

26.

26.) The k-means method is suitable for discovering clusters with nonconvex shapes or clusters of very different size.

a)

True

b)

False

27.

27.) The main advantage of grid-based approach is its slow processing time .

a)

True

b)

False

28.

28.) CLARANS also enables the detection of outliers .

a)

True

b)

False

29.

29.) The k-means and the k-modes methods cannot be integrated to cluster data with mixed numeric and categorical values.

a)

True

b)

False

30.

30.) PAM works effectively for large data sets .

a)

True

b)

False

31.

31.)Match the following:

Agglomerative approach --

a)

Centeroid based technique

b)

Representative object-based technique

c)

Bottom- up approach

d)

Top-down approach

e)

Object-by-object structure

32.

32.)Match the following:

Divisive approach --

a)

Centeroid based technique

b)

Top-down approach

c)

Bottom- up approach

d)

Representative object-based technique

e)

Object-by-object structure

33.

33.)Match the following:

K-means method --

a)

Centeroid based technique

b)

Representative object-based technique

c)

Top-down approach

d)

Bottom- up approach

e)

Object-by-object structure

34.

34.)Match the following:

K-medoid method --

a)

Centeroid based technique

b)

Object-by-object structure

c)

Representative object-based technique

d)

Top-down approach

e)

Bottom- up approach

35.

35.)Match the following:

Dissimilarity matrix --

a)

Object-by-object structure

b)

Centeroid based technique

c)

Representative object-based technique

d)

Top-down approach

e)

Bottom- up approach