WorksheetsDATASET'24 Quizz (Round 3)
Total questions: 20
Worksheet time: 30mins
Which of the following is the primary goal of "Feature Scaling" in machine learning?
To normalize data
To improve model accuracy
To reduce overfitting
To adjust features to a common scale so that the model doesn't give undue importance to one feature
Which technique is used for dimensionality reduction in Data Science?
K-Nearest Neighbors (KNN)
Principal Component Analysis (PCA)
Naive Bayes
Decision Trees
Which of these is a type of unsupervised learning?
Linear Regression
K-Means Clustering
Logistic Regression
Decision Trees
What does the head() function in Pandas do?
Deletes the first row of the DataFrame
Displays the first few rows of the DataFrame
Returns the last few rows of the DataFrame
Sorts the DataFrame
Which of the following is a Python library for data visualization?
Seaborn
Pandas
Scikit-learn
NumPy
In a Random Forest model, what does "out-of-bag error" refer to?
The error generated by the validation set
The error computed from the data that wasn't used in each individual tree's training
The error from test data
The error generated due to overfitting
Which of the following is a method for feature selection that penalizes large coefficients?
Ridge Regression
Lasso Regression
Random Forests
Decision Trees
In deep learning, which activation function is known to alleviate the vanishing gradient problem?
Sigmoid
Tanh
ReLU
Softmax
What is the key difference between Lasso and Ridge regression?
Lasso uses L2 regularization, while Ridge uses L1 regularization
Lasso can shrink coefficients to zero, while Ridge cannot
Lasso performs better on highly correlated data
Ridge regression penalizes the intercept term, while Lasso does not
What is the role of "Bias" in a machine learning model?
To make predictions more accurate
To improve the model’s ability to fit the data
To add a constant value to the prediction
To eliminate noise from the data
What is "Dropout" in a neural network?
A regularization technique where random units are ignored during training to prevent overfitting
A method of reducing the number of features
A method to speed up model training by skipping certain layers
A technique for early stopping of model training
In the context of decision trees, what is the “Gini Impurity”?
A measure of the variability in the data
A measure of how often a randomly chosen element would be incorrectly classified
A measure of the accuracy of the model
A method of scaling features
Which of the following algorithms is not commonly used for clustering?
K-Means
DBSCAN
Random Forest
Hierarchical Clustering
What is the purpose of a confusion matrix in machine learning?
To identify data cleaning errors
To evaluate the performance of a classification model
To optimize the model’s parameters
To visualize data correlations
Which of the following is true for principal component analysis (PCA)?
PCA increases the number of dimensions in a dataset
PCA minimizes variance in the dataset
PCA creates new features that are linear combinations of the original features
PCA uses a decision tree to identify principal components
What is a key advantage of using decision trees in data science?
High accuracy for all types of data
Easy interpretation and visualization
Robustness to overfitting
Support for continuous updating of the model
Which evaluation metric is most appropriate for imbalanced datasets?
Accuracy
Precision-Recall AUC
Mean Squared Error
Adjusted R-Squared
In natural language processing, what does TF-IDF stand for?
Term Frequency - Inverse Data Frequency
Term Frequency - Inverse Document Frequency
Token Frequency - Inverse Density Frequency
Text Frequency - Inverse Data Frequency
What is the curse of dimensionality?
The model’s inability to process large datasets
The tendency of algorithms to perform poorly as the number of features increases
The difficulty in scaling algorithms for distributed systems
The challenge of cleaning datasets with missing values
Which of these techniques can be used for feature selection?
Forward Selection
Principal Component Analysis
Lasso Regression
All of the above
