wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Machine Learning Quiz

Total questions: 100

Worksheet time: 50mins

Name
Class
Date
1.

is a field of computer science that deals with system programming to learn and improve with experience.

a)

Artifical Intelligence

b)

Machine learning

c)

Data Mining

d)

Model Selection

2.

The process of choosing models among diverse mathematical models, which are used to define the same data set is known as

a)

Model Selection

b)

Data Mining

c)

Artifical Intelligence

d)

Machine learning

3.

Machine learning is The autonomous acquisition of knowledge through the use of computer programs

a)

The autonomous acquisition of knowledge through the use of manual

b)

The selective acquisition of knowledge through the use of computer programs

c)

The selective acquisition of knowledge through the use of manual programs

d)

The autonomous acquisition of knowledge through the use of computer programs

4.

Who provided a formal definition of machine learning?

a)

Tom A. Mitchell

b)

Tom M. Mitchell

c)

Tom K. Mitchell

d)

Tom J. Mitchell

5.

Machine learning is divided into how many categories?

a)

2

b)

3

c)

4

d)

5

6.

Inputs are divided into how many in Classifications?

a)

2

b)

3

c)

4

d)

5

7.

What happens in clustering?

a)

An exponential of the inputs is found

b)

Inputs are multiplied

c)

Inputs are divided into groups

d)

Sets are multiplied

8.

Computers are best at learning ______

a)

facts

b)

concepts

c)

procedures

d)

principles

9.

Data used to build a data mining model.

a)

validation data

b)

training data

c)

test data

d)

hidden data

10.

Data used to optimize the parameter settings of a supervised learner model.

a)

training

b)

test

c)

verification

d)

validation

11.

The average squared difference between classifier predicted output and actual output.

a)

mean squared error

b)

root mean squared error

c)

mean absolute error

d)

mean relative error

12.

The process of forming general concept definitions from examples of concepts to be learned.

a)

deduction

b)

abduction

c)

induction

d)

conjunction

13.

Data mining is best described as the process of __________ identifying patterns in data.

a)

deducing relationships in data.

b)

representing data.

c)

simulating trends in data.

d)

identifying patterns in data.

14.

Computers are best at learning ______

a)

facts

b)

concepts

c)

procedures

d)

principles

15.

Like the probabilistic view, the ______ view allows us to associate a probability of membership with each classification.

a)

exemplar

b)

deductive

c)

classical

d)

inductive

16.

______ used to build a data mining model.

a)

validation data

b)

training data

c)

test data

d)

hidden data

17.

Supervised learning and unsupervised clustering both require at least one hidden attribute.

a)

output attribute

b)

input attribute

c)

categorical attribute

d)

input attribute

18.

Supervised learning differs from unsupervised clustering in that supervised learning requires ______ at least one input attribute.

a)

input attributes to be categorical.

b)

at least one output attribute.

c)

output attributes to be categorical.

d)

at least one output attribute.

19.

Database query is used to uncover this type of knowledge________.

a)

deep

b)

hidden

c)

shallow

d)

multidimensional

20.

______ is a statement to be tested.

a)

theory

b)

procedure

c)

principle

d)

hypothesis

21.

______ is a person trained to interact with a human expert in order to capture their knowledge.

a)

knowledge programmer

b)

knowledge developer

c)

knowledge engineer

d)

knowledge extractor

22.

Which of the following is not a characteristic of a data warehouse?

a)

contains historical data

b)

designed for decision support

c)

stores data in normalized tables

d)

promotes data redundancy

23.

______ is a structure designed to store data for decision support.

a)

operational database

b)

flat file

c)

decision tree

d)

data warehouse

24.

A nearest neighbor approach is best used _______ with large-sized datasets.

a)

when irrelevant attributes have been removed from the data.

b)

when a generalized model of the data is desirable.

c)

when an explanation of what has been found is of primary importance.

d)

when irrelevant attributes have been removed from the data.

25.

If a customer is spending more than expected, the customer’s intrinsic value is

a)

their actual value.

b)

greater than

c)

less than

d)

less than or equal to

26.

can be any unprocessed fact, value, text, sound or picture that is not being interpreted and analyze

a)

data

b)

knowledge

c)

information

d)

machine

27.

has been interpreted and manipulated

a)

data

b)

knowledge

c)

information

d)

machine

28.

________is the combination of inferred information and learning.

a)

data

b)

knowledge

c)

information

d)

machine

29.

model

a)

training data

b)

testing data

c)

validation data

d)

knowledge

30.

How we split data in Machine Learning?

a)

3

b)

5

c)

6

d)

8

31.

means scale of data.

a)

volume

b)

value

c)

velocity

d)

veracity

32.

Which defined correctness in data?

a)

volume

b)

value

c)

velocity

d)

veracity

33.

Supervised Learning is also called as .

a)

Inductive Learning

b)

semi-supervised

c)

regression

d)

labeled

34.

Which dataset is one which has both input and output parameters?

a)

labeled

b)

unlabelled

c)

function model

d)

labeled

35.

What is another name for meaningless data?

a)

unstructured data

b)

structured data

c)

labeled data

d)

value

36.

The data which contains only an input parameters.

a)

unstructured data

b)

structured data

c)

unlabeled data

d)

value

37.

is a classification algorithm for binary and multi class classification problems.

a)

naïve bayes

b)

bayes

c)

stephen

d)

bias

38.

is a sub-field of mathematics concerned with vectors, matrices, and linear

a)

linear algebra

b)

linear graphs

c)

linear arrays

d)

linear matrix

39.

_____is a method of teaching and learning in a logical manner.

a)

Machine learning

b)

PAC Learning

c)

Artifical Intelligence

d)

Sequence learning

40.

is about identifying group membership while regression technique involves predicting a response

a)

Classification

b)

Association

c)

Regression models

d)

Clustering

41.

Which training data includes a few desired outputs?

a)

Inductive Learning

b)

semi-supervised

c)

regression

d)

labeled

42.

The way candidate programs are generated known as the process.

a)

search

b)

hypothesis

c)

evaluation

d)

knowledge

43.

Which is called as idiot?

a)

naive

b)

naive bayes

c)

bias

d)

flemming bayes

44.

The field of study that gives computers the capability to learn without being explicitly programmed is known as .

a)

machine learning

b)

data learning

c)

testing learning

d)

type learning

45.

defined labels.

a)

classification

b)

regression

c)

supervised learning

d)

data warehouse

46.

Predictive models having target attribute having discrete values can be termed as

a)

Regression models

b)

Classification models

c)

supervised learning

d)

data warehouse

47.

When was the name coined?

a)

1987

b)

1959

c)

1978

d)

1990

48.

is based on an assumption that all of the features in the data set are important, equal and independent.

a)

naive

b)

naive bayes

c)

bias

d)

flemming bayes

49.

is a process or a study whether it closely relates to design, development of the algorithms that provide an ability to the machines to capacity to learn.

a)

Model Selection

b)

Data Mining

c)

Artifical Intelligence

d)

Machine learning

50.

technique is a rule based ML technique which finds out some very useful relations between parameters of a large data set.

a)

Classification models

b)

Association

c)

Regression models

d)

data warehouse

51.

technique is mostly applicable in case of image data-sets where usually all images are not labeled.

a)

Inductive Learning

b)

semi-supervised

c)

regression

d)

labeled

52.

model keeps on increasing its performance using a Reward Feedback to learn the behavior or pattern.

a)

supervised

b)

semi-supervised

c)

reinforcement

d)

unsupervised

53.

Which is example of supervised learning algorithm?

a)

K-Means Clustering

b)

Decision Trees

c)

Temporal Difference (TD)

d)

Q-Learning

54.

With Bayes classifier, missing data items are treated as equal

a)

compares.

b)

treated as unequal

c)

replaced with a default value.

d)

ignored.

55.

This unsupervised clustering algorithm terminates when mean values computed for the current iteration of the algorithm are identical to the computed mean values for the previous iteration.

a)

agglomerative clustering

b)

conceptual clustering

c)

K-Means clustering

d)

expectation maximization

56.

MATLAB stands for

a)

Maths Laboratory

b)

Matrix Laboratory

c)

Mathematical Lab

d)

Maths Lab

57.

MATLAB was developed by

a)

MathsWorks

b)

Intel

c)

Microsoft

d)

IBM

58.

In MATLAB the matrix is defined as an

a)

vector

b)

scalar

c)

array

d)

integer

59.

________acts as an outstanding tool for visulaizing technical data

a)

C

b)

C++

c)

Java

d)

MATLAB

60.

In command window the are entered

a)

data

b)

values

c)

commands

d)

files

61.

window displays plots and graphs

a)

command

b)

Edit

c)

Figure

d)

Command history

62.

The term is used to describe an array with two or more dimensions

a)

array

b)

vector

c)

matrix

d)

scale

63.

In MATLAB, the process of replacing loops by vectorized statements is known as

a)

scalarization

b)

vectorization

c)

looping

d)

branching

64.

What do you mean by a hard margin?

a)

The SVM allows very low error in classification

b)

The SVM allows high amount of error in classification

c)

The SVM is not allow very low error in classification

d)

The SVM allows medium amount of error in classification

65.

The minimum time complexity for training an SVM is O(n2). According to this fact, what sizes of datasets are not best suited for SVM’s?

a)

Large datasets

b)

Small datasets

c)

Medium sized datasets

d)

Size does not matter

66.

The effectiveness of an SVM depends upon:

a)

Selection of Kernel

b)

Kernel Parameters

c)

Soft Margin Parameter C

d)

All of the above

67.

Suppose you are using RBF kernel in SVM with high Gamma value. What does this signify?

a)

The model would consider even far away points from hyperplane for modeling

b)

The model would consider only the points close to the hyperplane for modeling

c)

The model would not be affected by distance of points from hyperplane for modeling

d)

None of the above

68.

What would happen when you use very small C (C~0)?

a)

Misclassification would happen

b)

Data will be correctly classified

c)

Can’t say

d)

None of these

69.

Suppose you gave the correct answer in previous question. What do you think that is actually happening?

a)

1. We are lowering the bias

b)

2. We are lowering the variance

c)

3. We are increasing the bias

d)

4. We are increasing the variance

70.

How many times we need to train our SVM model in such case?

a)

1

b)

2

c)

3

d)

4

71.

Linear Regression is a machine learning algorithm based on _____.

a)

unsupervised learning

b)

supervised learning.

c)

reinforcement learning

d)

none of these

72.

Regression models a target prediction value based on _____.

a)

dependent variable

b)

independent variables

c)

independent value

d)

dependent value

73.

regression technique finds out a linear relationship between x (input) and y(output) hence it is called as _________.

a)

Hypothesis function

b)

Related regression

c)

Linear Regression

d)

none of these

74.

In Linear Regression RMSE stands for_________.

a)

Root Mean Squared Error

b)

Read Mean Squared Error

c)

Root Mode Squared Error

d)

none of these

75.

Root Mean Squared error give difference between_________.

a)

original value and wrong value

b)

predict value and true value

c)

True value and false value

d)

none of these

76.

A decision tree has low training error and a large test error. What is the possible problem?

a)

Decision tree is too shallow

b)

Learning rate too high

c)

There is too much training data

d)

Decision tree is overfitting

77.

Which of the following is a disadvantage of non-parametric machine learning algorithms?

a)

Capable of fitting a large number of functional forms (Flexibility)

b)

Very fast to learn (Speed)

c)

More of a risk to overfit the training data (Overfitting)

d)

They do not require much training data

78.

For Ridge Regression, if the regularization parameter = 0, what does it mean?

a)

Large coefficients are not penalized

b)

Overfitting problems are not accounted for

c)

The loss function is as same as the ordinary least square loss function

d)

All of the above

79.

Which of the following is a disadvantage of non-parametric machine learning algorithms?

a)

Capable of fitting a large number of functional forms (flexibility)

b)

Very fast to learn (speed)

c)

More of a risk to overfit the training data (overfitting)

d)

They do not require much training data

80.

Which of the following is a true statement for regression methods in the case of feature selection?

a)

Ridge regression uses subset selection of features

b)

Lasso regression uses subset selection of features

c)

Both use subset selection of features

d)

None of above

81.

To check the linear relationship of dependent and independent continuous variables, which of the following plots are best suited?

a)

Scatter plot

b)

Bar chart

c)

Histograms

d)

All of the above

82.

Which of the following of the coefficients is added as the penalty term to the loss function in Lasso regression?

a)

Squared magnitude

b)

Absolute value of magnitude

c)

Number of non-zero entries

d)

None of the above

83.

What type of penalty is used on regression weights in Ridge regression?

a)

L0

b)

L1

c)

L2

d)

None of the above

84.

If two variables, x and y, have a very strong linear relationship, then_____

a)

There is evidence that x causes a change in y

b)

There is evidence that y causes a change in x

c)

There might not be any causal relationship between x and y

d)

None of these alternatives is correct

85.

Lasso Regression uses which norm?

a)

L1

b)

L2

c)

L1 & L2 both

d)

None of the above

86.

In Ridge regression, A hyper parameter is used called “_____________” that controls the weighting of the penalty to the loss function.

a)

Alpha

b)

Gamma

c)

Lambda

d)

None of above

87.

Is the logistic regression model a Generalized Linear Model (GLM)? What is a possible reason?

a)

No, it possible the shape function highly logistic

b)

Yes, it's a GLM but only by convention. It's properties don't match with the current definition

c)

Yes, it's a GLM since it's parameterized by a linear combination of its features, and

d)

No, it's not a GLM since it belongs to the 'exponential' family of distributions

e)

Yes, it's a GLM since it's parameterized by a linear combination of its features,

88.

Which of the following is false regarding logistic regression?

a)

We can solve the maximum log-

b)

We can solve the maximum log- likelihood

c)

We can solve the maximum log-

d)

The conditional

89.

A regression model in which more than one independent variable is used to predict the

a)

a simple linear regression model

b)

a multiple regression model

c)

none of the above

90.

A term used to describe the case when the independent variables in a multiple regression

a)

regression

b)

correlation

c)

multicollinearity

d)

none of the above

91.

Which of the following methods do we use to best fit the data in Logistic Regression?

a)

Least Square Error

b)

Maximum Likelihood

c)

Jaccard distance

d)

Both A and B

92.

Which of the following evaluation metrics can not be applied in case of logistic

a)

AUC-ROC

b)

Accuracy

c)

Logloss

d)

Mean-Squared-Error

93.

Which of the following algorithms do we use for Variable Selection?

a)

LASSO

b)

Ridge

c)

Both

d)

None of these

94.

Logistic regression is used to predict valued output?

a)

Continuous

b)

Categorical

c)

Discrete

d)

Ordinal

95.

How many types in logistic regression

a)

4

b)

3

c)

2

d)

43

96.

Logistic regression is when the observed outcome of dependent variable can

a)

Binomial

b)

Multinomial

c)

Ordinal

d)

Discrete

97.

Which of the following is link function in logistic regression __________

a)

Identity

b)

Logit

c)

Error

d)

Logit

98.

The odds of the dependent variable equaling a case (given some linear combination x of the

a)

the linear regression

b)

function of the linear regression

c)

calculations

d)

estimation method

99.

Regression coefficients in logistic regression are estimated using.

a)

Ordinary least squares method

b)

Maximum likelihood estimation method

c)

Dependent variable equalling a case

d)

calculations

100.

Which of the following is analogous to R-Squared for logistic regression.

a)

Likelihood ratio R-squared

b)

McFadden R-squared

c)

Cox and Snell R-Squared

d)

All of the above