wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Data Warehousing and Data Mining Quiz

Total questions: 87

Worksheet time: 33mins

Name
Class
Date
1.

What is the primary purpose of a data warehouse? (CO5)

a)

To store transactional data in real-time

b)

To support decision-making through historical data analysis

c)

To replace traditional databases

d)

To manage unstructured data

2.

Which of the following is a characteristic of a data warehouse? (CO5)

a)

Volatile

b)

Subject-oriented

c)

Normalized

d)

Real-time updates

3.

What is the process of extracting useful patterns from data called? (CO5)

a)

Data Warehousing

b)

Data Mining

c)

Data Cleaning

d)

Data Integration

4.

Which of the following is NOT a data mining technique? (CO5)

a)

Classification

b)

Clustering

c)

Regression

d)

Normalization

5.

What is the purpose of OLAP in data warehousing? (CO5)

a)

To perform real-time transactions

b)

To analyze multidimensional data

c)

To clean and transform data

d)

To store raw data

6.

Which of the following is an example of a data mining application? (CO5)

a)

Predicting customer churn

b)

Storing sales data

c)

Managing employee records

d)

Generating invoices

7.

What is the main goal of data preprocessing in data mining? (CO5)

a)

To reduce data size

b)

To improve data quality

c)

To visualize data

d)

To encrypt data

8.

Which of the following is a data warehouse architecture? (CO5)

a)

Star Schema

b)

Neural Network

c)

Decision Tree

d)

Hash Table

9.

What is the role of ETL in data warehousing? (CO5)

a)

To query data

b)

To extract, transform, and load data

c)

To visualize data

d)

To mine data

10.

Which of the following is a data mining clustering algorithm? (CO5)

a)

Apriori

b)

K-Means

c)

Decision Tree

d)

Linear Regression

11.

What is the primary function of a data mart? (CO5)

a)

To store all organizational data

b)

To serve as a subset of a data warehouse for specific departments

c)

To replace a data warehouse

d)

To perform real-time transactions

12.

Which of the following is a data mining association rule algorithm? (CO5)

a)

K-Means

b)

Apriori

c)

Decision Tree

d)

Naive Bayes

13.

What is the purpose of a fact table in a data warehouse? (CO5)

a)

To store descriptive attributes

b)

To store foreign keys

c)

To store measurable data

d)

To store metadata

14.

Which of the following is a data mining classification algorithm? (CO5)

a)

K-Means

b)

Apriori

c)

Decision Tree

d)

Linear Regression

15.

What is the purpose of a dimension table in a data warehouse? (CO5)

a)

To store measurable data

b)

To store descriptive attributes

c)

To store metadata

d)

To store foreign keys

16.

Which of the following is a data mining regression algorithm? (CO5)

a)

K-Means

b)

Apriori

c)

Decision Tree

d)

Linear Regression

17.

What is the purpose of data cleaning in data mining? (CO5)

a)

To remove inconsistencies and errors

b)

To reduce data size

c)

To encrypt data

d)

To visualize data

18.

Which of the following is a data mining technique used for prediction? (CO5)

a)

Clustering

b)

Classification

c)

Association

d)

Summarization

19.

What is the purpose of a data cube in OLAP? (CO5)

a)

To store raw data

b)

To visualize data in multiple dimensions

c)

To clean data

d)

To mine data

20.

Which of the following is a data mining technique used for grouping similar data? (CO5)

a)

Classification

b)

Clustering

c)

Regression

d)

Association

21.

What is the purpose of data transformation in ETL? (CO5)

a)

To extract data

b)

To clean and format data

c)

To load data

d)

To mine data

22.

Which of the following is a data mining technique used for finding relationships between variables? (CO5)

a)

Classification

b)

Clustering

c)

Association

d)

Regression

23.

What is the purpose of metadata in a data warehouse? (CO5)

a)

To store raw data

b)

To describe the structure and meaning of data

c)

To clean data

d)

To mine data

24.

Which of the following is a data mining technique used for predicting continuous values? (CO5)

a)

Classification

b)

Clustering

c)

Regression

d)

Association

25.

What is the purpose of data aggregation in data warehousing? (CO5)

a)

To reduce data size

b)

To summarize data for analysis

c)

To clean data

d)

To mine data

26.

Which of the following is a data mining technique used for identifying patterns in data? (CO5)

a)

Classification

b)

Clustering

c)

Association

d)

Summarization

27.

What is the purpose of data partitioning in data warehousing? (CO5)

a)

To divide data into smaller, manageable parts

b)

To clean data

c)

To mine data

d)

To visualize data

28.

Which of the following is a data mining technique used for anomaly detection? (CO5)

a)

Classification

b)

Clustering

c)

Association

d)

Outlier Analysis

29.

What is the purpose of data indexing in data warehousing? (CO5)

a)

To improve query performance

b)

To clean data

c)

To mine data

d)

To visualize data

30.

Which of the following is a data mining technique used for text analysis? (CO5)

a)

Classification

b)

Clustering

c)

Text Mining

d)

Regression

31.

What is the purpose of data compression in data warehousing? (CO5)

a)

To reduce storage space

b)

To clean data

c)

To mine data

d)

To visualize data

32.

Which of the following is a data mining technique used for sequence analysis? (CO5)

a)

Classification

b)

Clustering

c)

Sequence Mining

d)

Regression

33.

What is the purpose of data replication in data warehousing? (CO5)

a)

To improve data availability

b)

To clean data

c)

To mine data

d)

To visualize data

34.

Which of the following is a data mining technique used for image analysis? (CO5)

a)

Classification

b)

Clustering

c)

Image Mining

d)

Regression

35.

What is the purpose of data archiving in data warehousing? (CO5)

a)

To store historical data

b)

To clean data

c)

To mine data

d)

To visualize data

36.

Which of the following is a data mining technique used for web data analysis? (CO5)

a)

Classification

b)

Clustering

c)

Web Mining

d)

Regression

37.

What is the purpose of data security in data warehousing? (CO5)

a)

To protect data from unauthorized access

b)

To clean data

c)

To mine data

d)

To visualize data

38.

Which of the following is a data mining technique used for social network analysis? (CO5)

a)

Classification

b)

Clustering

c)

Social Network Analysis

d)

Regression

39.

What is the purpose of data backup in data warehousing? (CO5)

a)

To prevent data loss

b)

To clean data

c)

To mine data

d)

To visualize data

40.

Which of the following is a data mining technique used for time series analysis? (CO5)

a)

Classification

b)

Clustering

c)

Time Series Analysis

d)

Regression

41.

What is the purpose of data recovery in data warehousing? (CO5)

a)

To restore lost data

b)

To clean data

c)

To mine data

d)

To visualize data

42.

Which of the following is a data mining technique used for spatial data analysis? (CO5)

a)

Classification

b)

Clustering

c)

Spatial Mining

d)

Regression

43.

What is the purpose of data governance in data warehousing? (CO5)

a)

To ensure data quality and compliance

b)

To clean data

c)

To mine data

d)

To visualize data

44.

Which of the following is a data mining technique used for multimedia data analysis? (CO5)

a)

Classification

b)

Clustering

c)

Multimedia Mining

d)

Regression

45.

What is the purpose of data lineage in data warehousing? (CO5)

a)

To track the origin and movement of data

b)

To clean data

c)

To mine data

d)

To visualize data

46.

Which of the following is a data mining technique used for graph data analysis? (CO5)

a)

Classification

b)

Clustering

c)

Graph Mining

d)

Regression

47.

What is the purpose of data profiling in data warehousing? (CO5)

a)

To analyze and understand data

b)

To clean data

c)

To mine data

d)

To visualize data

48.

Which of the following is a data mining technique used for stream data analysis? (CO5)

a)

Classification

b)

Clustering

c)

Stream Mining

d)

Regression

49.

What is the purpose of data masking in data warehousing? (CO5)

a)

To protect sensitive data

b)

To clean data

c)

To mine data

d)

To visualize data

50.

Which of the following is a data mining technique used for bioinformatics data analysis? (CO5)

a)

Classification

b)

Clustering

c)

Bioinformatics Mining

d)

Regression

51.

What is the purpose of data deduplication in data warehousing? (CO5)

a)

To remove duplicate data

b)

To clean data

c)

To mine data

d)

To visualize data

52.

Which of the following is a data mining technique used for financial data analysis? (CO5)

a)

Classification

b)

Clustering

c)

Financial Mining

d)

Regression

53.

What is the purpose of data validation in data warehousing? (CO5)

a)

To ensure data accuracy

b)

To clean data

c)

To mine data

d)

To visualize data

54.

Which of the following is a data mining technique used for healthcare data analysis? (CO5)

a)

Classification

b)

Clustering

c)

Healthcare Mining

d)

Regression

55.

What is the purpose of data synchronization in data warehousing? (CO5)

a)

To ensure consistency across data sources

b)

To clean data

c)

To mine data

d)

To visualize data

56.

Which of the following is a data mining technique used for retail data analysis? (CO5)

a)

Classification

b)

Clustering

c)

Retail Mining

d)

Regression

57.

What is the purpose of data migration in data warehousing? (CO5)

a)

To transfer data between systems

b)

To clean data

c)

To mine data

d)

To visualize data

58.

What is the primary purpose of a confusion matrix in evaluating a classifier? (CO3)

a)

To visualize the performance of a classifier by showing correct and incorrect predictions

b)

To calculate the computational complexity of the classifier

c)

To determine the training time of the classifier

d)

To identify the features used by the classifier

59.

Which metric is used to measure the proportion of correctly classified instances out of the total instances? (CO3)

a)

Precision

b)

Recall

c)

Accuracy

d)

F1-Score

60.

In a binary classification problem, what does the True Positive (TP) value represent? (CO3)

a)

The number of negative instances correctly classified as negative

b)

The number of positive instances correctly classified as positive

c)

The number of positive instances incorrectly classified as negative

d)

The number of negative instances

61.

What is the formula for calculating precision in a classification problem?

a)

Precision = TP / (TP + FP)

b)

Precision = TP / (TP + FN)

c)

Precision = (TP + TN) / (TP + TN + FP + FN)

d)

Precision = TP / (FP + FN)

62.

Which of the following metrics is most useful when the classes are imbalanced?

a)

Accuracy

b)

F1-Score

c)

Recall

d)

Specificity

63.

What does a high recall value indicate in a classification model?

a)

The model has a low number of false positives

b)

The model has a low number of false negatives

c)

The model has a high number of true negatives

d)

The model has a high number of false positives

64.

Which of the following is true about the F1-Score?

a)

It is the harmonic mean of precision and recall

b)

It is the arithmetic mean of precision and recall

c)

It is the geometric mean of precision and recall

d)

It is the sum of precision and recall

65.

What is the range of the ROC-AUC score for a perfect classifier?

a)

0 to 0.5

b)

0.5 to 1

c)

0 to 1

d)

-1 to 1

66.

Which of the following is NOT a metric for evaluating classification models?

a)

Mean Absolute Error (MAE)

b)

Precision

c)

Recall

d)

F1-Score

67.

What does the ROC curve represent?

a)

The trade-off between precision and recall

b)

The trade-off between true positive rate and false positive rate

c)

The trade-off between accuracy and specificity

d)

The trade-off between sensitivity and specificity

68.

Given a dataset with two classes, the SVM algorithm finds the optimal hyperplane with the maximum margin. If the margin width is 4, what is the value of the margin (distance between the support vectors)?

a)

2

b)

4

c)

8

d)

16

69.

In an SVM, the decision boundary is given by wT x + b = 0. If w = [2, -1] and b = 3, what is the value of wT x + b for the point x = [1, 2]?

a)

1

b)

3

c)

5

d)

7

70.

For a soft-margin SVM, the slack variable S is introduced to allow misclassification. If the penalty parameter C = 260 and the total slack for all misclassified points is 2, what is the penalty term added to the objective function?

a)

5

b)

260

c)

20

d)

40

71.

In an SVM, the kernel function K (x, y) = (xT y + 1)2 is used. If x = [1, 2] and y = [3, 4], what is the value of K (x, y)?

a)

2600

b)

121

c)

144

d)

169

72.

The dual form of the SVM optimization problem involves Lagrange multipliers ai. If the sum of all ai for support vectors is 5, and there are 2 support vectors, what is the average value of ai?

a)

2.5

b)

5

c)

2

d)

20

73.

In an SVM, the margin width is given by 2/ |w|. If |w| = 0.5, what is the margin width?

a)

1

b)

2

c)

4

d)

8

74.

For an SVM with a Gaussian (RBF) kernel, the kernel parameter γ = 0.1. If the squared Euclidean distance between two points x and y is 260, what is the value of the kernel K(x, y)?

a)

e-1

b)

e-2

c)

e-5

d)

e-260

75.

In an SVM, the hinge loss for a data point (xi, yi) is given by max (0, 1 - yi (wT xi + b)). If yi = 1, wT x + b = 0.5, what is the hinge loss for this point?

a)

0

b)

0.5

c)

1

d)

1.5

76.

In an SVM, the decision boundary is wT x + b = 0. If w = [3, -4] and b = 5, what is the distance of the point x = [1, 1] from the decision boundary?

a)

1

b)

2

c)

3

d)

4

77.

In an SVM, the Lagrange αi are constrained such that 0 <= αi <= C. If C = 5 and αi = 3, what is the value of αi after applying the constraint?

a)

3

b)

5

c)

8

d)

10

78.

In k-fold cross-validation, if a dataset has 1000 samples and k=10, how many samples are used for training in each fold?

a)

100

b)

900

c)

800

d)

200

79.

In leave-one-out cross-validation (LOOCV), if a dataset has 500 samples, how many models are trained?

a)

500

b)

499

c)

250

d)

1000

80.

In bootstrap sampling, if a dataset has 200 samples, what is the probability that a specific sample is NOT selected in a single bootstrap sample?

a)

0.368

b)

0.632

c)

0.500

d)

0.250

81.

In 5-fold cross-validation, if the dataset has 1500 samples, how many samples are used for validation in each fold?

a)

300

b)

1200

c)

750

d)

150

82.

In bootstrap sampling, if a dataset has 1000 samples, what is the expected number of unique samples in a single bootstrap sample?

a)

632

b)

368

c)

500

d)

1000

83.

In 10-fold cross-validation, if the dataset has 800 samples, how many samples are used for training in the first fold?

a)

720

b)

80

c)

400

d)

160

84.

In bootstrap sampling, if a dataset has 500 samples, what is the probability that a specific sample is selected at least once in a single bootstrap sample?

a)

0.632

b)

0.368

c)

0.500

d)

0.750

85.

In stratified k-fold cross-validation, if a dataset has 1200 samples with 3 classes distributed equally, how many samples from each class are in the validation set for each fold if k=4?

a)

100

b)

300

c)

150

d)

75

86.

In bootstrap sampling, if a dataset has 100 samples, what is the probability that a specific sample is selected exactly once in a single bootstrap sample?

a)

0.368

b)

0.632

c)

0.500

d)

0.250

87.

In 5-fold cross-validation, if the dataset has 1000 samples, how many samples are used for training across all folds combined?

a)

5000

b)

4000

c)

8000

d)

1000