Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Module-III

Total questions: 15

Worksheet time: 33mins

Name
Class
Date
1.

Suppose, you are working on a binary classification problem. And there are 3 models each with 70% accuracy. You ensembled these models using majority voting method. What will be the maximum accuracy you can get?

a)

70%

b)

40%

c)

100%

d)

78.38%

2.

How can we assign the weights to output of different models in an ensemble?


  1. Use an algorithm to return the optimal weights
  2. Choose the weights using cross validation
  3. Give high weights to more accurate models
a)

1

b)

1 and 2

c)

2 and 3

d)

1, 2, and 3

3.

What is the assumption on accuracy of base model for an emsemble?

a)

<50%

b)

>50%

c)

<75%

d)

>75%

4.

Generally, an ensemble method works better, if the individual base models have ____________?

a)

Independent

b)

Dependent

c)

dependency does not play any role

d)

None of above

5.

If you use an ensemble of different base models, is it necessary to tune the hyper parameters of all base models to improve the ensemble performance?

a)

No

b)

Yes

c)

Cannot say

6.

Which of the following is / are true about weak learners used in ensemble model?

  1. They have low variance and they don’t usually overfit
  2. They have high bias, so they can not solve hard learning problems
  3. They have high variance and they don’t usually overfit
a)

1 and 2

b)

1 and 3

c)

2 and 3

d)

None of these

7.

Ensembles will yield bad results when there is significant diversity among the models.

a)

True

b)

False

8.

Which of the following can be true for selecting base learners for an ensemble?

  1. Different learners can come from same algorithm with different hyper parameters
  2. Different learners can come from different algorithms
  3. Different learners can come from different training spaces
a)

1

b)

2

c)

1 and 3

d)

1, 2, and 3

9.

Which of the following option is / are correct regarding benefits of ensemble model?

  1. Better performance
  2. Generalized models
  3. Better interpretability
a)

1 and 3

b)

2 and 3

c)

1 and 2

d)

1, 2, and 3

10.

The less correlation among base models of an ensemble is necessary.

a)

Yes

b)

No

c)

Some Times

11.

How can Clustering (Unsupervised Learning) be used to improve the accuracy of Linear Regression model (Supervised Learning)?

  1. Creating different models for different cluster groups.
  2. Creating an input feature for cluster ids as an ordinal variable.
  3. Creating an input feature for cluster centroids as a continuous variable.
  4. Creating an input feature for cluster size as a continuous variable.
a)

1 only

b)

2 and 4

c)

1 and 4

d)

all of the above

12.

Assume, you want to cluster 7 observations into 3 clusters using K-Means clustering algorithm. After first iteration clusters, C1, C2, C3 has following observations:


C1: {(2,2), (4,4), (6,6)}

C2: {(0,4), (4,0)}

C3: {(5,5), (9,9)}


What will be the cluster centroids if you want to proceed for second iteration?

a)

C1: (6,6), C2: (4,4), C3: (9,9)

b)

C1: (2,2), C2: (0,0), C3: (5,5)

c)

C1: (4,4), C2: (2,2), C3: (7,7)

d)

None of these

13.

If two variables V1 and V2, are used for clustering. Which of the following are true for K means clustering with k =3?


  1. If V1 and V2 has a correlation of 1, the cluster centroids will be in a straight line
  2. If V1 and V2 has a correlation of 0, the cluster centroids will be in straight line
a)

1 only

b)

2 only

c)

1 and 2

d)

None of the above

14.

Feature scaling is an important step before applying K-Mean algorithm. What is reason behind this?

a)

In distance calculation it will give the same weights for all features

b)

You always get the same clusters. If you use or don’t use feature scaling

c)

In Manhattan distance it is an important step but in Euclidian it is not

d)

None of these

15.

Which of the following can be applied to get good results for K-means algorithm corresponding to global minima?


  1. Try to run algorithm for different centroid initialization
  2. Adjust number of iterations
  3. Find out the optimal number of clusters
a)

1 and 3

b)

2 and 3

c)

1 and 2

d)

All of the above