wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Data Robot

Total questions: 126

Worksheet time: 1hrs 6mins

Name
Class
Date
1.

Machines that can preform tasks that are characteristic of human intelligence.

a)

Artificial Intelligence

b)

Machine Learning

c)

Target

d)

Features

2.

The practice of using algorithms to parse data, learn from it, and then make a determination or prediction about something in the word

a)

Artificial Intelligence

b)

Machine Learning

c)

Target

d)

Features

3.

The variable we are trying to predict and gain insights about

a)

Artificial Intelligence

b)

Machine Learning

c)

Target

d)

Features

4.

Can be thought of as the independent variables we will use to predict the target

a)

Unsupervised ML

b)

Machine Learning

c)

Target

d)

Features

e)

Supervised ML

5.

Data scientist tells the machine what it wants it learn (identifies target)

a)

Unsupervised ML

b)

Machine Learning

c)

Target

d)

Features

e)

Supervised ML

6.

Up to the machine to decide what it wants to learn

a)

Unsupervised ML

b)

Machine Learning

c)

Target

d)

Features

e)

Supervised ML

7.

“A field of study that gives computers the ability to learn without being explicitly programmed”-McClendon & Meghanathan

a)

Unsupervised ML

b)

Machine Learning

c)

Target

d)

Features

e)

Supervised ML

8.

The practice of using algorithms to parse data, learn from it, and then make a determination or prediction about something in the word

a)

Unsupervised ML

b)

Machine Learning

c)

Target

d)

Features

e)

Supervised ML

9.

about predicting the future based on the past.

a)

Unsupervised ML

b)

Machine Learning

c)

Target

d)

Features

e)

Supervised ML

10.

Makes ML possible with out extensive math/stat/programming

a)

Unsupervised ML

b)

automl

c)

Target

d)

Features

e)

Supervised ML

11.

Model diagnosis- Evaluation/ranking of the models• Models can be run simultaneously• Combinations of models can be run (Blenders)

a)

Unsupervised ML

b)

automl

c)

Target

d)

Features

e)

Supervised ML

12.

what are the parts of automated machine learning?

a)

Exploratory data analysis

b)

Feature engineering

c)

Algorithm selection and hyper-parameter tuning

d)

Model diagnostics

13.

The process of examining the descriptive statistics for all features as well as their relationship with the target variable

a)

Exploratory data analysis

b)

Feature engineering

c)

Algorithm selection and hyper-parameter tuning

d)

Model diagnostics

14.

Cleaning data, combining features, splitting features into multiple features, handling missing values, and dealing with text, etc.

a)

Exploratory data analysis

b)

Feature engineering

c)

Algorithm selection and hyper-parameter tuning

d)

Model diagnostics

15.

Keeping up with the “dizzying number” of available algorithms and their quadrillions of parameter combinations

a)

Exploratory data analysis

b)

Feature engineering

c)

Algorithm selection and hyper-parameter tuning

d)

Model diagnostics

16.

Keeping up with the “dizzying number” of available algorithms and their quadrillions of parameter combinations

a)

Exploratory data analysis

b)

Feature engineering

c)

Algorithm selection and hyper-parameter tuning

d)

Model diagnostics

17.

Evaluation of top models

a)

Exploratory data analysis

b)

Feature engineering

c)

Algorithm selection and hyper-parameter tuning

d)

Model diagnostics

18.

what is step one of the machine learning life cycle?

a)

define project objectives

b)

acquire and explore data

c)

model data

d)

interpret and communicate

e)

implement, document, and maintain

19.

what is step two of the machine learning life cycle?

a)

define project objectives

b)

acquire and explore data

c)

model data

d)

interpret and communicate

e)

implement, document, and maintain

20.

what is step three of the machine learning life cycle?

a)

define project objectives

b)

acquire and explore data

c)

model data

d)

interpret and communicate

e)

implement, document, and maintain

21.

what is step four of the machine learning life cycle?

a)

define project objectives

b)

acquire and explore data

c)

model data

d)

interpret and communicate

e)

implement, document, and maintain

22.

what is step five of the machine learning life cycle?

a)

define project objectives

b)

acquire and explore data

c)

model data

d)

interpret and communicate

e)

implement, document, and maintain

23.

Set up batch or API prediction system ￿ Document modeling process for reproducibility ￿ Create model monitoring and maintenance plan

a)

define project objectives

b)

acquire and explore data

c)

model data

d)

interpret and communicate

e)

implement, document, and maintain

24.

Interpret model ￿ Communicate model insights

a)

define project objectives

b)

acquire and explore data

c)

model data

d)

interpret and communicate

e)

implement, document, and maintain

25.

￿ Variable selection ￿ Build candidate models ￿ Model validation and selection

a)

define project objectives

b)

acquire and explore data

c)

model data

d)

interpret and communicate

e)

implement, document, and maintain

26.

￿ Variable selection ￿ Build candidate models ￿ Model validation and selection

a)

define project objectives

b)

acquire and explore data

c)

model data

d)

interpret and communicate

e)

implement, document, and maintain

27.

￿ Find appropriate data ￿ Merge data into single table ￿ Conduct exploratory data analysis ￿ Find and remove any target leakage ￿ Feature engineering

a)

define project objectives

b)

acquire and explore data

c)

model data

d)

interpret and communicate

e)

implement, document, and maintain

28.

￿ Specify business problem ￿ Acquire subject matter expertise ￿ Define unit of analysis and prediction target ￿ Prioritize modeling criteria ￿ Consider risks and success criteria ￿ Decide whether to continue

a)

define project objectives

b)

acquire and explore data

c)

model data

d)

interpret and communicate

e)

implement, document, and maintain

29.

what are considered part of the 8 criteria of AutoML excellence

a)

accuracy

b)

productivity

c)

ease of use

d)

understanding and learning

30.

what are considered part of the 8 criteria of AutoML excellence

a)

resource availability

b)

process transparency- effects understanding and learning

c)

generalizability across contexts

d)

recommended actions

31.

• Anything a company would want to know in order to increase sales or reduce costs

a)

business problem

b)

subject matter expertise

c)

unit of analysis

d)

modeling criteria

32.

State problem in language of business (not language of modeling) • What actions might result from this modeling project• Specify actions that might result • Include specifics (number of customers affected, costs etc.)• Explain impact to the bottom line

a)

business problem

b)

subject matter expertise

c)

unit of analysis

d)

modeling criteria

33.

qualitative

a)

categorical data

b)

numerical

34.

quantitative

a)

categorical data

b)

numerical

35.

Described by words rather than numbers

a)

categorical data

b)

numerical

36.

i.e. Freshmen, Sophomore, Junior, Seniori.e. Train, Plane, Bus, etc.

a)

categorical data

b)

numerical

37.

Arise from counting, measuring, or some kind of mathematical operation

a)

categorical data

b)

numerical

38.

i.e. 20 of you visited the course website since last class

a)

categorical data

b)

numerical

39.

discrete and continuous

a)

categorical data

b)

numerical

40.

nominal and binary

a)

categorical data

b)

numerical

41.

• Infinite number of possible responses• Like any point on a number line

a)

continuous

b)

discrete

42.

• Finite number of options• Examples: course letter grade, country of origin, or Likert scale

a)

continuous

b)

discrete

43.

You can identify groups are different, but no meaningful rankingExamples–Occupation {teacher, dentist, data scientist, student...}–Marital status {single, married, divorced, widowed}–CustomerId {1, 2, 3,..., n}

a)

categorical

b)

binary

44.

You can identify groups are different, but no meaningful rankingExamples–Occupation {teacher, dentist, data scientist, student...}–Marital status {single, married, divorced, widowed}–CustomerId {1, 2, 3,..., n}

a)

categorical

b)

binary

45.

Nominal attribute with only two categories/states0 -> no 1 -> yesBoolean (when states are represented as true or false]

a)

categorical

b)

binary

46.

Usually stored internally as an unsigned integer (number of seconds since 1970) ______ lead to conversion nightmares because of how many formats there are

a)

date and time

b)

text strings

47.

You specify a number of characters• This is either exact or a maximum depending on the data type• Varchar, variable length string, text varying, VSTR, and other names for variable lengthYou will almost always have to convert text or extract text to make something useful out of it

a)

date and time

b)

text strings

48.

subsection of a dataset from which the machine learning algorithm uncovers or “learns” relationships between the features and the target variable.

a)

training set

b)

validation (test) set

c)

holdout set

49.

subsection of a dataset to which we apply the machine learning algorithm to see how accurately it identifies relationships between the known outcomes for the target variable and the dataset’s other features

a)

training set

b)

validation (test) set

c)

holdout set

50.

a subsection of a dataset to provide a final estimate of the machine learning model’s performance after if has been trained and validated. Holdout sets should never be used to make decisions about which algorithms to use for improving tuning algorithms.

a)

training set

b)

validation (test) set

c)

holdout set

51.

The goal of machine learning is to build computational models with high prediction and generalization capabilities

a)

split data

b)

over training (over fitting)

c)

cross validation

52.

At the end of the training process the final model should predict correct outputs for the input samples, but it should also be able to generalize to unforeseen data

a)

split data

b)

over training (over fitting)

c)

cross validation

53.

High prediction and generalization capabilities are conflicting• Improper splitting can lead to excessively high variance in model performance

a)

split data

b)

over training (over fitting)

c)

cross validation

54.

Poor generalization can be classified a

a)

split data

b)

over training (over fitting)

c)

cross validation

55.

Poor generalization can be classified a

a)

split data

b)

over training (over fitting)

c)

cross validation

56.

If the original validation partition is not representative of the overall population, then the resulting model may appear to have a high accuracy when in reality it just happens to fit the unusual validation set well

a)

split data

b)

over training (over fitting)

c)

cross validation

57.

Randomize the order of the data and select a holdout sample. • For us this will be 20% of the data (5-fold cross validation)• Set aside, keep locked

a)

step 1

b)

step 2

c)

step 3

d)

step 4

e)

step 5

58.

Randomize the order of the data and select a holdout sample. • For us this will be 20% of the data (5-fold cross validation)• Set aside, keep locked

a)

step 1

b)

step 2

c)

step 3

d)

step 4

e)

step 5

59.

Randomize the order of the data and select a holdout sample. • For us this will be 20% of the data (5-fold cross validation)• Set aside, keep locked

a)

step 1

b)

step 2

c)

step 3

d)

step 4

e)

step 5

60.

Split the remaining data into a set of folds• For us this will be 5-folds (16% each)

a)

step 1

b)

step 2

c)

step 3

d)

step 4

e)

step 5

61.

Set aside Fold 1 as the validation set, combine the rows in the remaining four folds (2-5) and use these rows to create a model of which the features explain the target

a)

step 1

b)

step 2

c)

step 3

d)

step 4

e)

step 5

62.

Use the model to try to predict the target from Fold 1

a)

step 1

b)

step 2

c)

step 3

d)

step 4

e)

step 5

63.

Reveal true target and compare to predicted target• Evaluated on a variety of metrics- Validation Score

a)

step 1

b)

step 2

c)

step 3

d)

step 4

e)

step 5

64.

Set aside Fold 2 as the new validation set, the remaining folds (1,3,4,5) are now used for training.

a)

step 6

b)

step 7

c)

step 8

d)

step 9

e)

step 10

65.

Set aside Fold 2 as the new validation set, the remaining folds (1,3,4,5) are now used for training.

a)

step 6

b)

step 7

c)

step 8

d)

step 9

e)

step 10

66.

Use the model to try to predict the target from Fold 2

a)

step 6

b)

step 7

c)

step 8

d)

step 9

e)

step 10

67.

Reveal the true target from Fold 2 and compare to predicted target. Success metrics are again calculated.

a)

step 6

b)

step 7

c)

step 8

d)

step 9

e)

step 10

68.

Reveal the true target from Fold 2 and compare to predicted target. Success metrics are again calculated.

a)

step 6

b)

step 7

c)

step 8

d)

step 9

e)

step 10

69.

The process is repeated with Folds 3,4, and 5 in turn being set aside as the validation set.

a)

step 6

b)

step 7

c)

step 8

d)

step 9

e)

step 10

70.

Overall accuracy for cross validation is calculated

a)

step 6

b)

step 7

c)

step 8

d)

step 9

e)

step 10

71.

Estimate the accuracy of your machine learning model by averaging the accuracies derived in all k cases across cross validation

a)

business problem

b)

subject matter expertise

c)

unit of analysis

d)

k-fold cross validation

72.

Train your machine learning model using the training set and calculate the accuracy of your model by validating the predicted results against the validation set

a)

business problem

b)

subject matter expertise

c)

unit of analysis

d)

k-fold cross validation

73.

Train your machine learning model using the training set and calculate the accuracy of your model by validating the predicted results against the validation set

a)

business problem

b)

subject matter expertise

c)

unit of analysis

d)

k-fold cross validation

74.

Keep the fold fi as the validation set and keep all the remaining k-1 folds in the training set.

a)

business problem

b)

subject matter expertise

c)

unit of analysis

d)

k-fold cross validation

75.
a)

business problem

b)

subject matter expertise

c)

unit of analysis

d)

k-fold cross validation

76.

Answers the question “How rapidly will the model evaluate new cases after being put into production?”

a)

speed

b)

accuracy

77.

Use cases/sec

a)

speed

b)

accuracy

78.

additional features (more data)

a)

learning curves will not tell you this

b)

learning curves will tell you this

79.

additional cases (more data)

a)

learning curves will not tell you this

b)

learning curves will tell you this

80.

• Shows how the models predicative ability changes with ‘sample size’• Answers the question “Will more data help our model?” or “Would more data improve the model’s predicative ability?”

a)

learning curves will not tell you this

b)

learning curves will tell you this

81.

• Shows how the models predicative ability changes with ‘sample size’• Answers the question “Will more data help our model?” or “Would more data improve the model’s predicative ability?”

a)

learning curves will not tell you this

b)

learning curves will tell you this

82.

how do we determine if we should add more data

a)

cost

b)

time

83.

The overall impact of a feature without consideration of the impact of other features• The overall impact of a feature adjusted for the impact of the other features• Feature impact for specific feature values

a)

important features

b)

learning curves

84.

Moving from creating and selecting a model to understanding what features drive the target •What (statistical) relationships exist between the target and the other features•Clarifies why it works in some cases in fails in others

a)

important features

b)

learning curves

85.

The overall impact of a feature without consideration of the impact of other features

a)

importance

b)

feature impact

c)

feature effects

86.

The overall impact of a feature adjusted for the impact of the other features

a)

importance

b)

feature impact

c)

feature effects

87.

Feature Impact for specific feature values

a)

importance

b)

feature impact

c)

feature effects

88.

green bar

a)

importance

b)

feature impact

c)

feature effects

89.
a)

importance

b)

feature impact

c)

feature effects

90.
a)

importance

b)

feature impact

c)

feature effects

91.
a)

importance

b)

feature impact

c)

feature effects

92.

may not be the best model but many top models will come from tree based algorithms• Random Forest• ExtremeGradient Boosted Trees

a)

learning curves

b)

importance levels

c)

decisions tree

93.

what algorithms come from tree based algorithms

a)

random forest

b)

extremegradient boosted trees

94.

biggest problems with decision trees

a)

overfitting

b)

target

95.

how to determine if model was good

a)

Best case scenario order of models did not change

b)

Failing that, still a good scenario if the top two models stay at the top of the list

96.

Half of the data from the remaining four folds is randomly selected – 3,200 rows, 32%• 14 algorithms in this case• Validated against the 16% validation set, 1,600 rows• The 7 “best”/top performing algorithms from the 32% make it to round 2

a)

round 1

b)

round 2

c)

round 3

d)

round 4

97.

7 Algorithms from the previous stage• Algorithms are run on 64% of the data, 6,400 rows• Validated against the 16% validation set, 1,600 rows• The results appear on your “Leader board” as validation scores

a)

round 1

b)

round 2

c)

round 3

d)

round 4

98.

64% sample

a)

round 1

b)

round 2

c)

round 3

d)

round 4

99.

Run cross-validation on the top four models• Only if validation set is <=10,000 rows• Be sure to sort leaderboard by cross validation column

a)

round 1

b)

round 2

c)

round 3

d)

round 4

100.

cross validation

a)

round 1

b)

round 2

c)

round 3

d)

round 4

101.

After cross validation has run models are internally sorted by cross validation score and then the best models are blended

a)

round 1

b)

round 2

c)

round 3

d)

round 4

102.

blending

a)

round 1

b)

round 2

c)

round 3

d)

round 4

103.

what are the modeling modes

a)

autopilot

b)

quick autopilot (fewer blueprints)

c)

Manual mode (user chooses blueprints)

104.

how should be run a model

a)

informative features

b)

all features

105.

A measure of accuracy

a)

logloss

b)

holdout

106.

Lower scores are ‘better’

a)

logloss

b)

holdout

107.

Rather than evaluating the model directly on whether it assigns cases (rows) to the correct “label”, the model is evaluated based on probabilities generated by the model and their distance from the correct answer

a)

logloss

b)

holdout

108.

Rather than evaluating the model directly on whether it assigns cases (rows) to the correct “label”, the model is evaluated based on probabilities generated by the model and their distance from the correct answer

a)

logloss

b)

holdout

109.

Open data and examine• Also open the Data Dictionary to be sure you have a good grasp on the terms

a)

logloss

b)

holdout

c)

before we load the data

110.

subsection of a dataset from which the machine learning algorithm uncovers or “learns” relationships between the features and the target variable.

a)

training set

b)

validation test set

c)

holdout set

d)

over training (over fitting)

e)

boolean variables

111.

subsection of a dataset to which we apply the machine learning algorithm to see how accurately it identifies relationships between the known outcomes for the target variable and the dataset’s other features

a)

training set

b)

validation test set

c)

holdout set

d)

over training (over fitting)

e)

boolean variables

112.

subsection of a dataset to which we apply the machine learning algorithm to see how accurately it identifies relationships between the known outcomes for the target variable and the dataset’s other features

a)

training set

b)

validation test set

c)

holdout set

d)

over training (over fitting)

e)

boolean variables

113.

a subsection of a dataset to provide a final estimate of the machine learning model’s performance after if has been trained and validated. Holdout sets should never be used to make decisions about which algorithms to use for improving tuning algorithms.

a)

training set

b)

validation test set

c)

holdout set

d)

over training (over fitting)

e)

boolean variables

114.

Poor generalization can be classified as over training. The model simply memorizes the training examples and is not able to give correct outputs also for patterns that were not in the training dataset

a)

training set

b)

validation test set

c)

holdout set

d)

over training (over fitting)

e)

boolean variables

115.

Poor generalization can be classified as over training. The model simply memorizes the training examples and is not able to give correct outputs also for patterns that were not in the training dataset

a)

training set

b)

validation test set

c)

holdout set

d)

over training (over fitting)

e)

boolean variables

116.

Yes/No, True/False, 1/0 (binary)

a)

training set

b)

validation test set

c)

holdout set

d)

over training (over fitting)

e)

boolean variables

117.

Average of these 5 validation scores is the

a)

importance

b)

feature impact

c)

feature effects

d)

cross validation score

118.

why bother with missing data

a)

Not as important for Decision Trees, they do fine with missing values.

b)

However, any row with even one missing value will get that whole row kicked out from the analysis when......using Regression

c)

However, any row with even one missing value will get that whole row kicked out from the analysis when......using neural networks

119.

when choosing a module what should you consider

a)

adding more data

b)

predictive accuracy and prediction speed

c)

speed and to build a model

d)

familiarity with model

e)

insights

120.

what should the prediction distribution be under

a)

density

b)

frequency

121.

what are strengths of the confusion matrix

a)

straight forwards

b)

interpretation

c)

imprecise

d)

quick view of model and frequently maps well to business use case (once threshold is adjusted)

122.

what are limitations of the confusion matrix

a)

straight forwards

b)

interpretations

c)

imprecise

d)

quick view of model and frequently maps well to business use case (once threshold is adjusted)

123.

what are use cases of the confusion matrix

a)

straight forwards

b)

interpretations

c)

imprecise

d)

quick view of model and frequently maps well to business use case (once threshold is adjusted)

124.

TP+TN/ All cases

a)

accuracy

b)

true positive rate

c)

true negative rate

125.

TP/(TP+FN)

a)

accuracy

b)

true positive rate

c)

true negative rate

126.

TN/(TN+FP)

a)

accuracy

b)

true positive rate

c)

true negative rate