WorksheetsDATA SCIENCE 3 QUIZZ
Total questions: 15
Worksheet time: 11mins
Q1. Which of the following best defines the role of a Data Scientist?
A. Writing SQL scripts only
B. Collecting, cleaning, analyzing, and interpreting large datasets to drive decision-making
C. Creating user interfaces
D. Developing operating systems
Data Science can be seen as the intersection of:
A. Programming, marketing, and ethics
B. Statistics, domain knowledge, and mathematics
C. Mathematics, statistics, and computer science
D. Engineering, biology, and art
The term Big Data is often associated with:
A. Small-scale databases
B. Only structured datasets
C. Real-time streaming and high-dimensional data
D. Low-volume transactional logs
Which scenario best represents the concept of datafication?
A. A paper logbook being scanned into a PDF
B. Recording biometric signals via wearable devices and analyzing behavior patterns
C. Posting on social media
D. Watching a movie
Which of the following is a major ethical concern with datafication?
A. Server storage costs
B. UI/UX design flaws
C. Personal data privacy
D. File format conversions
Which of the following is a discrete probability distribution?
A. Normal Distribution
B. Poisson Distribution
C. Exponential Distribution
D. Log-Normal Distribution
What does it mean to “fit” a model in machine learning?
A. Choose the model with the most features
B. Apply statistical formulas randomly
C. Optimize model parameters to minimize prediction error
D. Remove all outliers from the data
Which of the following can reduce overfitting in machine learning models?
A. Using more features
B. Increasing the model depth
C. Applying regularization techniques like L1/L2
D. Decreasing the dataset size
Overfitting typically occurs when:
A. The model is too simple
B. The data is too clean
C. The model captures noise instead of signal
D. The model uses cross-validation
Data Science is an interdisciplinary field that uses scientific methods, algorithms, and systems to extract (a) from structured and unstructured data.
Statistical inference involves using a sample to draw conclusions about a larger (a) .
A (a) is a subset of a population used to estimate characteristics of the whole population.
The normal distribution is a (a) -shaped curve that is symmetric about the mean
Model fitting involves adjusting model parameters to minimize (a) between predicted and actual values.
Techniques like cross-validation and (a) can help prevent overfitting in machine learning.
