WorksheetsFAM_MODULE_456
Total questions: 30
Worksheet time: 15mins
Which of the following best describes data collection?
Transforming data into numerical format
Gathering information from various sources for analysis
Eliminating duplicates in a dataset
Visualizing data in graphs
A retail store uses an online survey to gather customer feedback. Which method is being used?
Observational study
Survey/Questionnaire
Web scraping
Interviews
Which method involves observing people in their natural setting?
Survey
Focus group
Observational study
Web scraping
Using past company sales records is an example of:
Web scraping
Existing data sources
Interview
Survey
Automated tools that extract competitor website data is called:
Survey
Observational study
Web scraping
Data validation
Which method is best for qualitative insights?
Surveys
Observations
Web scraping
Interviews and Focus Groups
Duplicate entries must be removed because they:
Can skew analysis results
Increase accuracy
Reduce sample size
Enhance performance
Standardizing all dates into MM/DD/YYYY is:
Feature engineering
Data validation
Standardizing formats
Imputation
Which is NOT a data cleaning technique?
Removing duplicates
Correcting errors
Outlier detection
Regression modeling
A dataset shows “Manila” and “MNL.” This requires:
Outlier detection
Correcting errors
Deletion
Encoding
Which technique is best for predicting a continuous outcome?
Decision Tree
Regression
Clustering
Time-Series
Logistic regression is commonly used for:
Predicting sales
Forecasting electricity demand
Binary outcomes (Yes/No)
Customer segmentation
Which regression method assumes a straight-line relationship?
Linear Regression
Logistic Regression
Hierarchical Clustering
Decision Tree
Predicting house prices using size, location, and age is an example of:
Logistic Regression
Multiple Regression
Clustering
ARIMA
Which model visually represents decision-making with branches?
Regression
Decision Tree
Clustering
Exponential Smoothing
What is the biggest disadvantage of decision trees?
High interpretability
Handles both numerical and categorical data
Overfitting
Uses probabilities
A bank wants to classify loan applicants into approve/deny categories. The best model is:
Clustering
Regression
Decision Tree
ARIMA
Which model groups customers with similar behaviors?
Regression
Decision Trees
Clustering
Time-Series
K-Means clustering works by:
Assigning points to the nearest cluster mean
Building a tree of nodes
Drawing a regression line
Using exponential smoothing
Hierarchical clustering is best for:
Forecasting demand
Predicting binary outcomes
Understanding relationships between clusters
Pruning decision trees
Which metric is most misleading in imbalanced datasets?
Precision
Recall
Accuracy
F1 Score
A hospital needs to ensure that patients with a disease are detected. Which metric is most important?
Precision
Recall
Accuracy
Silhouette Score
The F1 Score is the:
Arithmetic mean of precision and recall
Weighted sum of precision and recall
Harmonic mean of precision and recall
Ratio of true positives to total samples
ROC-AUC measures:
Regression variance explained
Trade-off between true positive rate and false positive rate
Model interpretability
Cluster separation
K-fold cross-validation divides the dataset into:
k subsets used for training and testing in rotation
Equal halves
Training only
Testing only
Which is most interpretable?
Linear regression
Random forests
Neural networks
Gradient boosting
Fairness in models means:
Highest accuracy possible
Equal treatment of different groups
Minimum computation cost
No retraining needed
A/B testing helps in:
Balancing datasets
Feature selection
Comparing old and new models
Removing outliers
Model performance drift means:
Accuracy suddenly increases
Performance degrades over time due to changing data
Dataset duplication
Model stops training
Retraining a model is done to:
Save storage space
Adapt to new data patterns
Increase interpretability
Avoid testing
