wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Data Warehousing

Total questions: 72

Worksheet time: 1hrs 13mins

Name
Class
Date
1.

Which one is correct for data warehousing?

a)

It can be updated by end users

b)

It can solve all business questions

c)

It is designed for focus subject areas

d)

It contains only current data

2.

Which are used in Multi-dimensional modeling?

a)

Foreign key

b)

Fact Table

c)

Dimensional Table

d)

Surrogate key

3.

Which one come first to build data warehouse

a)

Data Cleansing

b)

Business questions

c)

Data Acquisition

d)

Data Profiling

4.

Why do we apply snowflake schema?

a)

Aggregation

b)

Normalization

c)

Generalization

d)

Transformation

5.

What is importance of data profiling to data warehousing? You may answer more than one answer

a)

It helps to aggregate data

b)

It helps to understand data

c)

It helps to detect abnormality in data

d)

It helps to generate sample data

6.

What is metadata in data warehousing? You may answer more than one

a)

Business definition

b)

Data dictionary

c)

BI report

d)

Measures

7.

What is CORRECT about data warehouse? You may answer more than one

a)

It makes use of historical data to answer business questions

b)

It reveals insight from data

c)

It supports roll-up and drill-down to explore data

d)

Its design is the same as database

8.

What is CORRECT about dimensional modeling?

a)

It is designed to support query

b)

It is designed to support storage

c)

It is entity-relationship modeling

d)

It does not need normalization

9.

Data warehouse is subject-oriented

a)

True

b)

False

10.

What is NOT aggregation function?

a)

Median

b)

Count

c)

Distinct

d)

Sum

11.

Which mode is supported in data warehouse

a)

Write only

b)

Read only

c)

Write and Read

d)

Copy

12.

What is data to describe data in data warehouse?

a)

Relational data

b)

Operational data

c)

Informational data

d)

Metadata

13.

What is the relationship type between fact and dimension table in star schema?

a)

Many-to-Many

b)

One-to-One

c)

One-to-Many

d)

Many-to-One

14.

Geographical locations are measures

a)

True

b)

False

15.

The candy was sour.

a)

Qualitative Data

b)

Quantitative Data

16.

The flower is red.

a)

Qualitative Data

b)

Quantitative Data

17.

The mass of the beaker was 122 g.

a)

Qualitative Data

b)

Quantitative Data

18.

There are 13 boys and 14 girls in fifth period.

a)

Qualitative Data

b)

Quantitative Data

19.

Which is not a property of data warehouse?

a)

subject oriented

b)

time variant

c)

collection from heterogeneous sources

d)

volatile

20.

data warehousing used in_______________

a)

transaction system

b)

decision support system

21.

What are the characeristics of OLAP systems?

a)

query driven

b)

integrated

c)

less users

d)

store current data

22.

Which system is application oriented?

a)

Online transaction processing system

b)

Online analytical processing system

23.

data warehouse is based on_____________

a)

two dimensional model

b)

three dimensional model

c)

multidimensional model

d)

unidirectional model

24.

Data mining means______

a)

data fetching

b)

data accesing

c)

knowledge discovery

25.

Which book is suggested for reference?

a)

Pieter Adriaans

b)

Arun k Pujari

c)

Han and Kamber

d)

Margerat

26.

multidimensional model of data warehouse called as__

a)

data structure

b)

table

c)

data cube

d)

tree

27.

a data warehouse can include

a)

flat-files

b)

database table

c)

online data

d)

all

28.
Why was Data warehousing proposed?
a)
To keep track of trasnactional data
b)
To keep summarised historical information
29.
What is an operational system in data warehousing?
a)
A system that is used to process the day-to-day transactions of an organisation
b)
A system that sets up the data warehouse
c)
A system that runs the data warehouse
d)
A system that identifies error in the data warehouse
30.
Orangisations focus on ways to use operational data to support ___________ as a means of gaining competitive advantage
a)
Data warehousing
b)
Decision-Making
c)
Competitive-Analyse
d)
Relationl databases
31.
In data warehousing what is time-variant data?
a)
Data in the warehouse is only accurate and valid at some point in time or over time interval
b)
Data in the warehouse is always accurate and valid 
c)
Data in the warehouse is not accurate
d)
Data in the warehouse is only accurate sometimes
32.
What does OLTP stand for?
a)
Online transaction processing
b)
Offline transaction processing
c)
Outline trajectory processing
d)
Online traffic processing
33.
Metadata is data about data
a)
True
b)
False
34.
What is a Data Mart?
a)
A data mart is a subgroup of the data warehouse
b)
A data mart is another type of data warehouse
c)
A data mart is not actually related to data warehouses
d)
None of these
35.
What is a benefit of a data warehouse?
a)
Data flexibility
b)
Cost to implmenet
c)
Competitive Advantage
d)
Data ownership concerns
36.
What is a Star Schema?
a)
A star schema consists of a fact table with a single table for each dimension
b)
A star schema is a type of database system
c)
A star schema is used when exporting data from the database
d)
None of these
37.
What is a Snowflake Schema?
a)
Each dimension table is normalized, which may create additional tables attached to the dimension tables
b)
A Snowflake schema is a type of database system
c)
A Snowflake schema is used when exporting data from the database
d)
None of these
38.

Finding the hidden information from the data base

a)

Data Structures

b)

Data Mining

c)

Data scrubbing

d)

Data Handling

39.

Data Mining is also referred to as

a)

Exploratory Data Analysis

b)

Knowledge discovery from Databases

c)

None of the above

40.

The output of KDD is

a)

Data

b)

Information

c)

Query

d)

Useful information

41.

Identifying the examples of sequence data

a)

Datamatrix

b)

Weather forecast

c)

Market basket data

d)

Genomic data

42.

Which of the following is not example of ordinal attributes

a)

Zip code

b)

Ordered no

c)

Movie rating

d)

Military rank

43.
What type of graph is this?
a)
Bar Graph
b)
Line Plot
c)
Pictograph
d)
Tally Chart
44.

Using features to predict unknown or future values of the same or other feature is known as ___________ power of data mining

a)

Clustering

b)

Predictive

c)

Associative

d)

Descriptive

45.

Fraud-detection models and risk mitigation models-these are examples of data mining solution for which discipline?

a)

Health

b)

Insurance

c)

Banking

d)

Retail

46.

Select the skills mainly required as a competent data analyst/scientist/miner

a)

SQL

b)

R

c)

Python

d)

Java

47.

Select the tools that can be used for data mining

a)

KNIME

b)

WEKA

c)

RATTLE

d)

TANAGRA

48.

Choose which data mining task is the most suitable for the following scenario: detecting the dosage of medicine for a certain treatment

a)

Prediction

b)

Classification

c)

Association rules

d)

Sequential pattern analysis

49.

Choose which data mining task is the most suitable for the following scenario: grouping participants in a weight loss campaign

a)

Classification

b)

Clustering

c)

Prediction

d)

Association rules

50.

Choose which data mining task is the most suitable for the following scenario: determining the stock value of a certain company

a)

Prediction

b)

Classification

c)

Clustering

d)

Association rules

51.

Choose which data mining task is the most suitable for the following scenario: determining the stock value of a certain company

a)

Prediction

b)

Classification

c)

Clustering

d)

Association rules

52.

Associations/co-relations between product sales, & prediction based on such association is called ____________

a)

Customer profiling

b)

Target marketing

c)

Customer requirement analysis

d)

Cross-market analysis

53.

Classification is

a)

A subdivision of a set of examples into a number of classes

b)

A measure of the accuracy, of the classification of a concept that is given by a certain theory

c)

The task of assigning a classification to a set of examples

d)

None of these

54.

Data selection is

a)

The actual discovery phase of a knowledge discovery process

b)

The stage of selecting the right data for a KDD process

c)

A subject-oriented integrated time variant non-volatile collection of data in support of management

d)

None of these

55.

The problem of finding the hidden structure in unlabeled data is called as

a)

supervised learning

b)

unsupervised learning

c)

reinforcement learning

d)

none of these

56.

Telephone company want to segment their customers into different groups.. this is an example of

a)

supervised learning

b)

data extraction

c)

unsupervised learning

d)

data extraction

57.

The task of inferring a model from as set of labeled data is called as

a)

supervised learning

b)

unsupervised learning

c)

reinforcement learning

d)

semi supervised learning

58.

Discrimination of email is spam or not is an classification task ...

a)

TRUE

b)

FALSE

59.

Which of the following tools can be for DATA MINING

a)

weka

b)

R

c)

both a& b

d)

none

60.

Data warehouse is a .....

a)

collection of data marts

b)

collection of data algorithms

c)

collection of processing techniques

d)

all of above

61.

Training Data means....

a)

data used to train a model

b)

data used for analysis

c)

both a & b

d)

only a

62.

Credit card approval is a .... kind of application

a)

classification

b)

prediction

c)

analysing

d)

none

63.

Choose which data mining task is suitable for the following scenario: first buy digital camera, then buy large SD memory cards

a)

Classification

b)

Sequential pattern analysis

c)

Association rule

d)

Prediction

64.

Choose which data mining task is the most suitable for the following scenario: Identifying an unexpected/unusual amount of spending

a)

Prediction

b)

Sequential pattern analysis

c)

Association rules

d)

Anomaly detection

65.

Choose which data mining task is the most suitable for the following scenario: diagnosing the level of flood severity

a)

Prediction

b)

Classification

c)

Anomaly detection

d)

Association rules

66.

Choose which data mining task is the most suitable for the following scenario: determining the best location to be recommended a tourist club (multiple answers)

a)

Association rules

b)

Clustering

c)

Classification

d)

Prediction

67.

Choose which data mining task is the most suitable for the following scenario: determining the rating when a location is recommended to a tourist club member

a)

Classification

b)

Prediction

c)

Clustering

d)

Association rules

68.

Choose which data mining task is the most suitable for the following scenario: determining which tour group is suitable to a new member based on her past location ratings

a)

Prediction

b)

Classification

c)

Clustering

d)

Anomaly detection

69.

Data mining is the exploration and analyzing large data set to discover patterns and correlations to predict outcomes.

a)

True

b)

False

70.

Which of the following technique is categorised as unsupervised learning?

a)

Decision Tree

b)

Clustering

c)

None of the above

71.

"Process of sifting through large data sets to identify and describe patterns, discover and establish relationships with an intent to predict future trends based on those patterns and relationships". What does this statement explain about?

a)

Visualization

b)

Representation

c)

Data mining

d)

Business intelligence

72.

Fraud-detection models and risk mitigation models-these are examples of data mining solution for which discipline?

a)

Health

b)

Insurance

c)

Banking

d)

Retail