wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Data Warehousing and Data Mining LAB

Total questions: 69

Worksheet time: 37mins

Name
Class
Date
1.

Data Mining is the set of methodologies used in analyzing data from various dimensions and perspectives, finding previously unknown hidden patterns, classifying and grouping the data and summarizing the identified relationships

a)

True

b)

False

2.

What should be written in the blue box?

a)

Transformed data

b)

Pattern/model

c)

Preprocessed data

d)

Raw data

3.

What should be written in the blue box?

a)

Transformed Data

b)

Preprocessed Data

c)

Pattern/model

d)

Raw data

4.

Using features to predict unknown or future values of the same or other feature is known as ___________ power of data mining

a)

Clustering

b)

Predictive

c)

Associative

d)

Descriptive

5.

"Use data mining to find interesting, human-interpretable patterns that describe the data". What is the type of data mining to achieve this objective?

a)

Association

b)

Anomaly

c)

Descriptive

d)

Categorisation

6.

"Process of sifting through large data sets to identify and describe patterns, discover and establish relationships with an intent to predict future trends based on those patterns and relationships". What does this statement explain about?

a)

Visualization

b)

Representation

c)

Data mining

d)

Business intelligence

7.

Fraud-detection models and risk mitigation models-these are examples of data mining solution for which discipline?

a)

Health

b)

Insurance

c)

Banking

d)

Retail

8.

Select the skills mainly required as a competent data analyst/scientist/miner

a)

SQL

b)

R

c)

Python

d)

Java

9.

Select the tools that can be used for data mining

a)

KNIME

b)

WEKA

c)

RATTLE

d)

TANAGRA

10.

Data mining process includes reporting and analysis

a)

True

b)

False

11.

Focus on the specific organisation data to detect patterns

a)

True

b)

False

12.

If during data mining, some data is incomplete, the team should seeking out the incomplete data

a)

True

b)

False

13.

Techniques such as Self-Organizing-Maps (SOM’s), help to map missing data based by visualizing the model of multi-dimensional complex data.

a)

True

b)

False

14.

A scoreboard, on a manager or supervisor’s computer, fed with real-time from data as it flows in and out of various databases within the company’s environment. Choose below which explains this best.

a)

Data visualization

b)

Dashboard

c)

Anomaly detection

d)

Data analysis

15.

_______ is helpful to automatically find patterns within the text embedded in hordes of text files, word-processed files, PDFs, and presentation files.

a)

SQL

b)

Text Analysis

c)

Ctrl+F

d)

Visualization

16.

Associations/co-relations between product sales, & prediction based on such association is called ____________

a)

Customer profiling

b)

Target marketing

c)

Customer requirement analysis

d)

Cross-market analysis

17.

Class label is unknown: Group data to form new classes, e.g., cluster houses to find distribution patterns

a)

Predictive Analysis

b)

Anomaly Detection

c)

Association Mining

d)

Cluster analysis

18.

A pattern is interesting if it is easily understood by humans, valid on new or test data with some degree of certainty, potentially useful, novel, or validates some hypothesis that a user seeks to confirm

a)

True

b)

False

19.

Data mining is driven by the following:

-Kinds of data to be mined

-Kinds of knowledge to be discovered

-Kinds of techniques utilized

-Kinds of applications adapted

-Kinds of given mining duration

-Kinds of mining period

a)

True

b)

False

20.

Data mining depends on

-Kinds of data to be mined

-Kinds of knowledge to be discovered

-Kinds of techniques utilized

-Kinds of applications adapted

a)

True

b)

False

21.

Choose which data mining task is suitable for the following scenario: first buy digital camera, then buy large SD memory cards

a)

Classification

b)

Sequential pattern analysis

c)

Association rule

d)

Prediction

22.

Choose which data mining task is the most suitable for the following scenario: Identifying an unexpected/unusual amount of spending

a)

Prediction

b)

Sequential pattern analysis

c)

Association rules

d)

Anomaly detection

23.

Choose which data mining task is the most suitable for the following scenario: diagnosing the level of flood severity

a)

Prediction

b)

Classification

c)

Anomaly detection

d)

Association rules

24.

Choose which data mining task is the most suitable for the following scenario: detecting the dosage of medicine for a certain treatment

a)

Prediction

b)

Classification

c)

Association rules

d)

Sequential pattern analysis

25.

Choose which data mining task is the most suitable for the following scenario: grouping participants in a weight loss campaign

a)

Classification

b)

Clustering

c)

Prediction

d)

Association rules

26.

Choose which data mining task is the most suitable for the following scenario: determining the stock value of a certain company

a)

Prediction

b)

Classification

c)

Clustering

d)

Association rules

27.

Choose which data mining task is the most suitable for the following scenario: determining the best location to be recommended to a tourist club members

a)

Association rules

b)

Clustering

c)

Classification

d)

Prediction

28.

Choose which data mining task is the most suitable for the following scenario: determining the rating when a location is recommended to a tourist club member

a)

Classification

b)

Prediction

c)

Clustering

d)

Association rules

29.

Choose which data mining task is the most suitable for the following scenario: determining which tour group is suitable to a new member based on her past location ratings

a)

Prediction

b)

Classification

c)

Clustering

d)

Anomaly detection

30.

Choose which data mining task is the most suitable for the following scenario: determining the thumbs up/thumbs down of a social media post

a)

Prediction

b)

Association rules

c)

Classification

d)

Clustering

31.

Choose which data mining task is the most suitable for the following scenario:

To identify items that are bought concomitantly by a reasonable fraction of customers so that they can be shelved.

a)

Classification

b)

Association rules

c)

Clustering

d)

Prediction

32.

Choose which data mining task is the most suitable for the following scenario:


To subdivide a market into distinct subset of customers where each subset can be targeted with a distinct marketing mix

a)

Classification

b)

Prediction

c)

Clustering

d)

Association rules

33.

Choose which data mining task is the most suitable for the following scenario:

To find groups of documents that are similar to each other based on important terms appearing in them

a)

Classification

b)

Clustering

c)

Prediction

d)

Association rules

34.

Choose which data mining task is the most suitable for the following scenario:

To reduce cost of mailing by targeting a set of consumers likely to buy a new cell phone product

a)

Classification

b)

Association rules

c)

Clustering

d)

Prediction

35.

Choose which data mining task is the most suitable for the following scenario:

Predict fraudulent cases in credit card transactions

a)

Classification

b)

Association rules

c)

Anomaly detection

d)

Clustering

36.

Choose which data mining task is the most suitable for the following scenario:

To guess wind velocities based on temperature, humidity, air pressure, etc

a)

Classification

b)

Association rules

c)

Prediction

d)

Anomaly detection

37.

Choose which data mining task is the most suitable for the following scenario:

Given a set of n points or objects, and k, the expected number of outliers, find the top k objects that considerably dissimilar, exceptional or inconsistent with the remaining data

a)

Classification

b)

Association rules

c)

Clustering

d)

Anomaly detection

38.

Choose which data mining task is the most suitable for the following scenario:

Based on past usage patterns, develop model for authorized credit card transactions

a)

Classification

b)

Association rules

c)

Clustering

d)

Anomaly detection

39.

Choose which data mining task is the most suitable for the following scenario:

Given is a set of objects, with each object associated with its own time of events, find rules that predict strong sequential dependencies among different events

a)

Classification

b)

Sequential pattern analysis

c)

Clustering

d)

Association rules

40.

Choose which data mining task is the most suitable for the following scenario:

Given the records of books that a group of people read, find relationship of the genre pattern

a)

Classification

b)

Association rules

c)

Clustering

d)

Prediction

41.
KDD describes the _________.  
a)
whole process of extraction of knowledge from data
b)
Extraction of data 
c)
extraction of information
42.
The partition of overall data warehouse is _______. 
a)
database
b)
data cube
c)
data mart
d)
operational data.
43.
Metadata describes __________. 
a)
contents of database
b)
structure of contents of database
c)
structure of database.
44.
OLAP stands for ________. 
a)
Online Analytical Processing
b)
Online Linear Analytical Processing
c)
Online Analytical Problem
45.
 OLAP is used to explore the ___________ knowledge. 
a)
shallow
b)
deep
c)
multidimensional
d)
 hidden.
46.
. ________ is the technique which is used for discovering patterns in dataset at the beginning of data mining process. 
a)
Kohenon map
b)
Visualization
c)
OLAP
47.

Why was data warehousing proposed?

a)

To keep track of transactional data

b)

To keep summarized historical information

c)

To manage data from heterogeneous sources

d)

To produce management reports

48.
What is a data warehouse?
a)
Is a relational database that is designed for query and analysis
b)
Is a non-relational database that is designed for query and analysis
49.

In data warehousing, what is time-variant data?

a)

Data in the warehouse is only accurate and valid at some point in time or over time interval

b)

Data in the warehouse is always accurate and valid

c)

Data in the warehouse is only accurate sometimes

d)

Data in the warehouse is not accurate

50.
What is a Data Mart?
a)
A data mart is a subgroup of the data warehouse
b)
A data mart is another type of data warehouse
c)
A data mart is not actually related to data warehouses
d)
None of these
51.
Metadata is data about data
a)
True
b)
False
52.

Is the data in a data warehouse generally updated in real-time?

a)

Yes

b)

No

53.

. An operational system is which of the following?

a)

A system that is used to run the business in real time and is based on historical data.

b)

A system that is used to run the business in real time and is based on current data.

c)

A system that is used to support decision making and is based on current data.

d)

A system that is used to support decision making and is based on historical data.

54.

Which schema is best for data warehouse development?

a)
b)
c)
55.

The data collected in data warehouse can be used for analyzing purposes.

a)

True

b)

False

56.

Which of the following are the characteristics of a data warehouse?

a)

Subject-oriented.

b)

Integrated.

c)

Non-volatile.

d)

All of the above.

57.

A basic concept of data warehouse is which of the following?

a)

Can be updated by end users.

b)

Contains numerous naming conventions and formats.

c)

Store the data in formats suitable for easy access for decision making.

d)

Contains only current data.

58.
What is a Star Schema?
a)
A star schema consists of a fact table with a single table for each dimension
b)
A star schema is a type of database system
c)
A star schema is used when exporting data from the database
d)
None of these
59.

What does OLTP stand for?

a)

Online transaction processing

b)

Offline transaction processing

c)

Outline trajectory processing

d)

Online traffic processing

60.

A snowflake schema is a normalized star schema

a)

TRUE

b)

FALSE

61.

"A fact table is narrow, but deep" means

a)

Number of columns is high, number of rows is high

b)

Number of columns is high, number of rows is low

c)

Number of columns is low, number of rows is high

d)

Number of columns is low, number of rows is low

62.

What is the mode?

a)

# occurring the most

b)

the average

c)

greatest - least

d)

the middle #

63.
The scores awarded to 25 students for an assignment were as follows:
4  7  5  9  8  6  7  7  8  5  6  9  8    5  8  7  4  7  3  6  8  9  7  6  9
What is  the mode?
a)
6
b)
7
c)
8
d)
9
64.
Find the mean of these numbers:
5,11,2,12,4,2
a)
4.1
b)
6
c)
4.5
d)
4
65.
Find the median of these numbers:
4,2,7,4,3
a)
2
b)
5
c)
7
d)
4
66.
Find the median, mode and range: 
3, 5, 7, 9, 11, 8, 3
a)
median:7  mode:3  range:7
b)
median:6  mode:3  range:7
c)
median:7  mode:3  range:8
67.
Find the mean of these numbers:
2, 57, 38, 42, 6
a)
29
b)
38
c)
50
d)
145
68.
Find the median.
12, 5, 9, 18, 22, 25, 5
a)
9
b)
8
c)
12
d)
18
69.

The number of miles that Kyle biked each week for a 7-week period is shown:

36, 42, 28, 52, 48, 36, 31

What is the median number of miles Kyle biked?

a)

24

b)

36

c)

39

d)

52