wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

MIDTERM Business Intelligence

Total questions: 50

Worksheet time: 35mins

Name
Class
Date
1.

What does data mining refer to?

a)

Extracting knowledge from large amounts of data

b)

Extracting minerals from the earth

c)

Extracting oil from underground reservoirs

d)

Extracting water from rivers

2.

Who coined the term 'Knowledge Discovery in Databases'?

a)

Emily Johnson

b)

John Smith

c)

Gregory Piatetsky-Shapiro

d)

Ryan Lungcay

3.

What is the first step in the KDD process?

a)

Data Mining

b)

Data Preprocessing

c)

Data Selection

d)

Data Transformation

4.

Which technique is used for market basket or transaction data analysis?

a)

Classification

b)

Clustering

c)

Regression

d)

Association

5.

What is the main advantage of Decision Trees for classification?

a)

Require high computational power

b)

Easy to interpret

c)

Difficult to understand

d)

Not suitable for simple data sets

6.

Which classifier is based on the Bayes theorem?

a)

SVM

b)

K-NN Classifier

c)

Decision Tree

d)

Bayesian Classification

7.

What is the purpose of Rule-Based Classification?

a)

To identify outliers

b)

To perform regression analysis

c)

To represent knowledge in the form of rules

d)

To create decision trees

8.

What is the goal of Frequent-Pattern Based Classification?

a)

To identify outliers

b)

To discover relevant patterns in large datasets

c)

To perform regression analysis

d)

To classify data based on rules

9.

What is the main disadvantage of data mining related to data quality?

a)

Technical Complexity

b)

Ethical Considerations

c)

Data Quality

d)

Data Privacy and Security

10.

Which algorithm simulates the process of natural selection for solving problems?

a)

Regression

b)

Artificial Neural Network

c)

Genetic Algorithm

d)

Outlier Detection

11.

What is the purpose of data smoothing?

a)

Remove noise from the data set

b)

Calculate the mean

c)

Add noise to the data set

d)

Sort the data

12.

What is the primary goal of supervised learning?

a)

To predict outcomes accurately

b)

To group similar data points

c)

To discover hidden patterns

d)

To reduce the number of features

13.

Which type of learning uses labeled datasets?

a)

Semi-Supervised Learning

b)

Reinforcement Learning

c)

Unsupervised Learning

d)

Supervised Learning

14.

What is the task of clustering in unsupervised learning?

a)

Reducing the number of features

b)

Grouping similar data points

c)

Identifying associations among data items

d)

Predicting numerical values

15.

Which algorithm is commonly used for classification problems?

a)

Support vector machines

b)

Logistic regression

c)

Linear regression

d)

Principal component analysis

16.

What is the drawback of unsupervised learning?

a)

Uses predefined output labels

b)

Requires labeled data

c)

Lack of Ground Truth

d)

Predicts outcomes accurately

17.

What does regression in supervised learning help in predicting?

a)

Numerical values

b)

Essential information

c)

Hidden patterns

d)

Similar data points

18.

What is the first step in the data mining process?

a)

Data Preparation

b)

Modeling

c)

Business Understanding

d)

Evaluation

19.

Which step involves building a predictive model using machine learning algorithms?

a)

Data Understanding

b)

Data Preparation

c)

Deployment

d)

Modeling

20.

What is the overall goal of the data mining process?

a)

Analyzing data quality

b)

Creating data repositories

c)

Extracting information from data sets

d)

Extracting raw data

21.

What is another name for Data Mining?

a)

Data Harvesting

b)

Data Analysis

c)

Knowledge Mining

d)

Data Extraction

22.

What is the main task during the 'Estimate model' phase of data mining?

a)

Model Selection and Implementation

b)

Data Preprocessing

c)

Data Collection

d)

Model Interpretation

23.

What is one of the major issues in Data Mining related to different users' interests?

a)

Data Preprocessing

b)

Pattern Evaluation

c)

Data Cleaning

d)

Mining Different Kinds of Knowledge

24.

What is the advantage of Data Mining related to decision making?

a)

Fraud Detection

b)

Improved Decision Making

c)

Better Customer Service

d)

Increased Efficiency

25.

What is one of the disadvantages of Data Mining related to privacy concerns?

a)

Complexity

b)

Data Quality

c)

High Cost

d)

Privacy Concerns

26.

What is the process of interactive mining of knowledge at multiple levels of abstraction?

a)

Data Mining Query Languages

b)

Data Preprocessing

c)

Pattern Evaluation

d)

Incorporation of Background Knowledge

27.

What is the need for efficient and scalable data mining algorithms related to?

a)

Data Preprocessing

b)

Data Normalization

c)

Parallel, Distributed, and Incremental Mining Algorithms

d)

Pattern Evaluation

28.

What is the goal of data preprocessing in data mining?

a)

To introduce errors in the data

b)

To skip the data cleaning process

c)

To increase the size of the dataset

d)

To make the data more suitable for analysis

29.

What is data cleaning in the data preprocessing process?

a)

Adding noise to the data

b)

Identifying and correcting errors or inconsistencies in the data

c)

Transforming the data into a lower-dimensional space

d)

Increasing the size of the dataset

30.

What is data integration in data preprocessing?

a)

Scaling the data to a common range

b)

Combining data from multiple sources to create a unified dataset

c)

Removing data from the dataset

d)

Dividing continuous data into discrete categories

31.

What does data transformation involve in data preprocessing?

a)

Converting the data into a suitable format for analysis

b)

Selecting a subset of relevant features from the dataset

c)

Grouping similar data points together into clusters

d)

Fitting the data to a regression function

32.

What is data reduction in the data preprocessing process?

a)

Dividing the data into discrete categories

b)

Reducing the size of the dataset while preserving the important information

c)

Increasing the complexity of the dataset

d)

Adding noise to the dataset

33.

What is feature selection in data reduction?

a)

Compressing the dataset

b)

Transforming the data into a lower-dimensional space

c)

Selecting a subset of relevant features from the dataset

d)

Grouping similar data points into clusters

34.

What is feature extraction in data reduction?

a)

Grouping similar data points into clusters

b)

Selecting a subset of relevant features from the dataset

c)

Transforming the data into a lower-dimensional space while preserving the important information

d)

Compressing the dataset

35.

What is sampling in data reduction?

a)

Compressing the dataset

b)

Grouping similar data points into clusters

c)

Transforming the data into a lower-dimensional space

d)

Selecting a subset of data points from the dataset

36.

What is clustering in data reduction?

a)

Compressing the dataset

b)

Transforming the data into a lower-dimensional space

c)

Selecting a subset of relevant features from the dataset

d)

Grouping similar data points together into clusters

37.

What is compression in data reduction?

a)

Transforming the data into a lower-dimensional space

b)

Grouping similar data points into clusters

c)

Selecting a subset of relevant features from the dataset

d)

Compressing the dataset while preserving the important information

38.

What significantly impacts the results obtained from data mining?

a)

Data complexity

b)

Scalability

c)

Data quality

d)

Data privacy

39.

Which techniques are essential to improve data quality?

a)

Data anonymization and encryption

b)

Clustering and classification

c)

Data cleaning and preprocessing

d)

Association rule mining

40.

What poses challenges due to the vast amounts of data generated from various sources?

a)

Data complexity

b)

Data quality

c)

Data privacy and security

d)

Scalability

41.

Which regulations impose strict rules on data collection and usage?

a)

GDPR, CCPA, and HIPAA

b)

Data cleaning and preprocessing

c)

Clustering, classification, and association rule mining

d)

Data anonymization and encryption

42.

What becomes critical factors as dataset size increases?

a)

Data quality

b)

Data complexity

c)

Computational resources and processing time

d)

Data privacy and security

43.

What is the MEANS value for 8,9,15,16?

a)

23

b)

30

c)

12

d)

13

44.

What is the BOUNDARIES value for 8,9,15,16?

a)

8, 8, 15, 16

b)

8, 8, 8, 16

c)

8, 8, 16, 16

d)

8, 16, 16, 16

45.

What is the MEANS value for 21,21,24,26?

a)

23

b)

30

c)

12

d)

13

46.

What is the BOUNDARIES value for 21,21,24,26?

a)

21, 21, 26, 26

b)

21, 21, 21, 26

c)

21, 26, 26, 26

d)

21, 21, 24 24

47.

What is the MEDIAN value for 21,21,24,26?

a)

21, 21, 21, 21

b)

24

c)

24, 24, 24, 24

d)

223, 23, 23, 23

48.

What is the MEANS value for 27,30,30,34?

a)

23

b)

30

c)

12

d)

13

49.

What is the BOUNDARIES value for 27,30,30,34?

a)

27, 27, 34, 34

b)

27, 27, 27, 34

c)

27, 34, 34, 34

d)

27, 27, 30, 34

50.

What are some real-world applications of data mining, give one and explain?

4 lines