Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

DATASET'24 Quizz (Round 3)

Total questions: 20

Worksheet time: 30mins

Name
Class
Date
1.

Which of the following is the primary goal of "Feature Scaling" in machine learning?

a)

To normalize data

b)

To improve model accuracy

c)

To reduce overfitting

d)

To adjust features to a common scale so that the model doesn't give undue importance to one feature

2.

Which technique is used for dimensionality reduction in Data Science?

a)

K-Nearest Neighbors (KNN)

b)

Principal Component Analysis (PCA)

c)

Naive Bayes

d)

Decision Trees

3.

Which of these is a type of unsupervised learning?

a)

Linear Regression

b)

K-Means Clustering

c)

Logistic Regression

d)

Decision Trees

4.

What does the head() function in Pandas do?

a)

Deletes the first row of the DataFrame

b)

Displays the first few rows of the DataFrame

c)

Returns the last few rows of the DataFrame

d)

Sorts the DataFrame

5.

Which of the following is a Python library for data visualization?

a)

Seaborn

b)

Pandas

c)

Scikit-learn

d)

NumPy

6.

In a Random Forest model, what does "out-of-bag error" refer to?

a)

The error generated by the validation set

b)

The error computed from the data that wasn't used in each individual tree's training

c)

The error from test data

d)

The error generated due to overfitting

7.

Which of the following is a method for feature selection that penalizes large coefficients?

a)

Ridge Regression

b)

Lasso Regression

c)

Random Forests

d)

Decision Trees

8.

In deep learning, which activation function is known to alleviate the vanishing gradient problem?

a)

Sigmoid

b)

Tanh

c)

ReLU

d)

Softmax

9.

What is the key difference between Lasso and Ridge regression?

a)

Lasso uses L2 regularization, while Ridge uses L1 regularization

b)

Lasso can shrink coefficients to zero, while Ridge cannot

c)

Lasso performs better on highly correlated data

d)

Ridge regression penalizes the intercept term, while Lasso does not

10.

What is the role of "Bias" in a machine learning model?

a)

To make predictions more accurate

b)

To improve the model’s ability to fit the data

c)

To add a constant value to the prediction

d)

To eliminate noise from the data

11.

What is "Dropout" in a neural network?

a)

A regularization technique where random units are ignored during training to prevent overfitting

b)

A method of reducing the number of features

c)

A method to speed up model training by skipping certain layers

d)

A technique for early stopping of model training

12.

In the context of decision trees, what is the “Gini Impurity”?

a)

A measure of the variability in the data

b)

A measure of how often a randomly chosen element would be incorrectly classified

c)

A measure of the accuracy of the model

d)

A method of scaling features

13.

Which of the following algorithms is not commonly used for clustering?

a)

K-Means

b)

DBSCAN

c)

Random Forest

d)

Hierarchical Clustering

14.

What is the purpose of a confusion matrix in machine learning?

a)

To identify data cleaning errors

b)

To evaluate the performance of a classification model

c)

To optimize the model’s parameters

d)

To visualize data correlations

15.

Which of the following is true for principal component analysis (PCA)?

a)

PCA increases the number of dimensions in a dataset

b)

PCA minimizes variance in the dataset

c)

PCA creates new features that are linear combinations of the original features

d)

PCA uses a decision tree to identify principal components

16.

What is a key advantage of using decision trees in data science?

a)

High accuracy for all types of data

b)

Easy interpretation and visualization

c)

Robustness to overfitting

d)

Support for continuous updating of the model

17.

Which evaluation metric is most appropriate for imbalanced datasets?

a)

Accuracy

b)

Precision-Recall AUC

c)

Mean Squared Error

d)

Adjusted R-Squared

18.

In natural language processing, what does TF-IDF stand for?

a)

Term Frequency - Inverse Data Frequency

b)

Term Frequency - Inverse Document Frequency

c)

Token Frequency - Inverse Density Frequency

d)

Text Frequency - Inverse Data Frequency

19.

What is the curse of dimensionality?

a)

The model’s inability to process large datasets

b)

The tendency of algorithms to perform poorly as the number of features increases

c)

The difficulty in scaling algorithms for distributed systems

d)

The challenge of cleaning datasets with missing values

20.

Which of these techniques can be used for feature selection?

a)

Forward Selection

b)

Principal Component Analysis

c)

Lasso Regression

d)

All of the above