Font size
WorksheetsMIDTERM Business Intelligence
Total questions: 50
Worksheet time: 35mins
What does data mining refer to?
Extracting knowledge from large amounts of data
Extracting minerals from the earth
Extracting oil from underground reservoirs
Extracting water from rivers
Who coined the term 'Knowledge Discovery in Databases'?
Emily Johnson
John Smith
Gregory Piatetsky-Shapiro
Ryan Lungcay
What is the first step in the KDD process?
Data Mining
Data Preprocessing
Data Selection
Data Transformation
Which technique is used for market basket or transaction data analysis?
Classification
Clustering
Regression
Association
What is the main advantage of Decision Trees for classification?
Require high computational power
Easy to interpret
Difficult to understand
Not suitable for simple data sets
Which classifier is based on the Bayes theorem?
SVM
K-NN Classifier
Decision Tree
Bayesian Classification
What is the purpose of Rule-Based Classification?
To identify outliers
To perform regression analysis
To represent knowledge in the form of rules
To create decision trees
What is the goal of Frequent-Pattern Based Classification?
To identify outliers
To discover relevant patterns in large datasets
To perform regression analysis
To classify data based on rules
What is the main disadvantage of data mining related to data quality?
Technical Complexity
Ethical Considerations
Data Quality
Data Privacy and Security
Which algorithm simulates the process of natural selection for solving problems?
Regression
Artificial Neural Network
Genetic Algorithm
Outlier Detection
What is the purpose of data smoothing?
Remove noise from the data set
Calculate the mean
Add noise to the data set
Sort the data
What is the primary goal of supervised learning?
To predict outcomes accurately
To group similar data points
To discover hidden patterns
To reduce the number of features
Which type of learning uses labeled datasets?
Semi-Supervised Learning
Reinforcement Learning
Unsupervised Learning
Supervised Learning
What is the task of clustering in unsupervised learning?
Reducing the number of features
Grouping similar data points
Identifying associations among data items
Predicting numerical values
Which algorithm is commonly used for classification problems?
Support vector machines
Logistic regression
Linear regression
Principal component analysis
What is the drawback of unsupervised learning?
Uses predefined output labels
Requires labeled data
Lack of Ground Truth
Predicts outcomes accurately
What does regression in supervised learning help in predicting?
Numerical values
Essential information
Hidden patterns
Similar data points
What is the first step in the data mining process?
Data Preparation
Modeling
Business Understanding
Evaluation
Which step involves building a predictive model using machine learning algorithms?
Data Understanding
Data Preparation
Deployment
Modeling
What is the overall goal of the data mining process?
Analyzing data quality
Creating data repositories
Extracting information from data sets
Extracting raw data
What is another name for Data Mining?
Data Harvesting
Data Analysis
Knowledge Mining
Data Extraction
What is the main task during the 'Estimate model' phase of data mining?
Model Selection and Implementation
Data Preprocessing
Data Collection
Model Interpretation
What is one of the major issues in Data Mining related to different users' interests?
Data Preprocessing
Pattern Evaluation
Data Cleaning
Mining Different Kinds of Knowledge
What is the advantage of Data Mining related to decision making?
Fraud Detection
Improved Decision Making
Better Customer Service
Increased Efficiency
What is one of the disadvantages of Data Mining related to privacy concerns?
Complexity
Data Quality
High Cost
Privacy Concerns
What is the process of interactive mining of knowledge at multiple levels of abstraction?
Data Mining Query Languages
Data Preprocessing
Pattern Evaluation
Incorporation of Background Knowledge
What is the need for efficient and scalable data mining algorithms related to?
Data Preprocessing
Data Normalization
Parallel, Distributed, and Incremental Mining Algorithms
Pattern Evaluation
What is the goal of data preprocessing in data mining?
To introduce errors in the data
To skip the data cleaning process
To increase the size of the dataset
To make the data more suitable for analysis
What is data cleaning in the data preprocessing process?
Adding noise to the data
Identifying and correcting errors or inconsistencies in the data
Transforming the data into a lower-dimensional space
Increasing the size of the dataset
What is data integration in data preprocessing?
Scaling the data to a common range
Combining data from multiple sources to create a unified dataset
Removing data from the dataset
Dividing continuous data into discrete categories
What does data transformation involve in data preprocessing?
Converting the data into a suitable format for analysis
Selecting a subset of relevant features from the dataset
Grouping similar data points together into clusters
Fitting the data to a regression function
What is data reduction in the data preprocessing process?
Dividing the data into discrete categories
Reducing the size of the dataset while preserving the important information
Increasing the complexity of the dataset
Adding noise to the dataset
What is feature selection in data reduction?
Compressing the dataset
Transforming the data into a lower-dimensional space
Selecting a subset of relevant features from the dataset
Grouping similar data points into clusters
What is feature extraction in data reduction?
Grouping similar data points into clusters
Selecting a subset of relevant features from the dataset
Transforming the data into a lower-dimensional space while preserving the important information
Compressing the dataset
What is sampling in data reduction?
Compressing the dataset
Grouping similar data points into clusters
Transforming the data into a lower-dimensional space
Selecting a subset of data points from the dataset
What is clustering in data reduction?
Compressing the dataset
Transforming the data into a lower-dimensional space
Selecting a subset of relevant features from the dataset
Grouping similar data points together into clusters
What is compression in data reduction?
Transforming the data into a lower-dimensional space
Grouping similar data points into clusters
Selecting a subset of relevant features from the dataset
Compressing the dataset while preserving the important information
What significantly impacts the results obtained from data mining?
Data complexity
Scalability
Data quality
Data privacy
Which techniques are essential to improve data quality?
Data anonymization and encryption
Clustering and classification
Data cleaning and preprocessing
Association rule mining
What poses challenges due to the vast amounts of data generated from various sources?
Data complexity
Data quality
Data privacy and security
Scalability
Which regulations impose strict rules on data collection and usage?
GDPR, CCPA, and HIPAA
Data cleaning and preprocessing
Clustering, classification, and association rule mining
Data anonymization and encryption
What becomes critical factors as dataset size increases?
Data quality
Data complexity
Computational resources and processing time
Data privacy and security
What is the MEANS value for 8,9,15,16?
23
30
12
13
What is the BOUNDARIES value for 8,9,15,16?
8, 8, 15, 16
8, 8, 8, 16
8, 8, 16, 16
8, 16, 16, 16
What is the MEANS value for 21,21,24,26?
23
30
12
13
What is the BOUNDARIES value for 21,21,24,26?
21, 21, 26, 26
21, 21, 21, 26
21, 26, 26, 26
21, 21, 24 24
What is the MEDIAN value for 21,21,24,26?
21, 21, 21, 21
24
24, 24, 24, 24
223, 23, 23, 23
What is the MEANS value for 27,30,30,34?
23
30
12
13
What is the BOUNDARIES value for 27,30,30,34?
27, 27, 34, 34
27, 27, 27, 34
27, 34, 34, 34
27, 27, 30, 34
What are some real-world applications of data mining, give one and explain?
