WorksheetsModule-III
Total questions: 15
Worksheet time: 33mins
Suppose, you are working on a binary classification problem. And there are 3 models each with 70% accuracy. You ensembled these models using majority voting method. What will be the maximum accuracy you can get?
70%
40%
100%
78.38%
How can we assign the weights to output of different models in an ensemble?
- Use an algorithm to return the optimal weights
- Choose the weights using cross validation
- Give high weights to more accurate models
1
1 and 2
2 and 3
1, 2, and 3
What is the assumption on accuracy of base model for an emsemble?
<50%
>50%
<75%
>75%
Generally, an ensemble method works better, if the individual base models have ____________?
Independent
Dependent
dependency does not play any role
None of above
If you use an ensemble of different base models, is it necessary to tune the hyper parameters of all base models to improve the ensemble performance?
No
Yes
Cannot say
Which of the following is / are true about weak learners used in ensemble model?
- They have low variance and they don’t usually overfit
- They have high bias, so they can not solve hard learning problems
- They have high variance and they don’t usually overfit
1 and 2
1 and 3
2 and 3
None of these
Ensembles will yield bad results when there is significant diversity among the models.
True
False
Which of the following can be true for selecting base learners for an ensemble?
- Different learners can come from same algorithm with different hyper parameters
- Different learners can come from different algorithms
- Different learners can come from different training spaces
1
2
1 and 3
1, 2, and 3
Which of the following option is / are correct regarding benefits of ensemble model?
- Better performance
- Generalized models
- Better interpretability
1 and 3
2 and 3
1 and 2
1, 2, and 3
The less correlation among base models of an ensemble is necessary.
Yes
No
Some Times
How can Clustering (Unsupervised Learning) be used to improve the accuracy of Linear Regression model (Supervised Learning)?
- Creating different models for different cluster groups.
- Creating an input feature for cluster ids as an ordinal variable.
- Creating an input feature for cluster centroids as a continuous variable.
- Creating an input feature for cluster size as a continuous variable.
1 only
2 and 4
1 and 4
all of the above
Assume, you want to cluster 7 observations into 3 clusters using K-Means clustering algorithm. After first iteration clusters, C1, C2, C3 has following observations:
C1: {(2,2), (4,4), (6,6)}
C2: {(0,4), (4,0)}
C3: {(5,5), (9,9)}
What will be the cluster centroids if you want to proceed for second iteration?
C1: (6,6), C2: (4,4), C3: (9,9)
C1: (2,2), C2: (0,0), C3: (5,5)
C1: (4,4), C2: (2,2), C3: (7,7)
None of these
If two variables V1 and V2, are used for clustering. Which of the following are true for K means clustering with k =3?
- If V1 and V2 has a correlation of 1, the cluster centroids will be in a straight line
- If V1 and V2 has a correlation of 0, the cluster centroids will be in straight line
1 only
2 only
1 and 2
None of the above
Feature scaling is an important step before applying K-Mean algorithm. What is reason behind this?
In distance calculation it will give the same weights for all features
You always get the same clusters. If you use or don’t use feature scaling
In Manhattan distance it is an important step but in Euclidian it is not
None of these
Which of the following can be applied to get good results for K-means algorithm corresponding to global minima?
- Try to run algorithm for different centroid initialization
- Adjust number of iterations
- Find out the optimal number of clusters
1 and 3
2 and 3
1 and 2
All of the above
