Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

FINAL EXAM

Total questions: 25

Worksheet time: 4hrs 10mins

Name
Class
Date
1.

It builds classification models in the form of a tree structure that represents rules that can be easily understood.

(a)  

2.

This predicts the class of previously unseen records by aggregating predictions made by multiple classifiers.

 

(a)  

3.

A relatively modern algorithm that is essentially an ensemble of decision trees.

(a)  

4.

Assumes that there is a linear relationship present between dependent and independent variables. In simple words, it finds the best fitting line/plane that describes two or more variables.

(a)  

5.

As a statistical approach, it is used to predict the relationship between two variables the dependent variable, Y, and the independent variable, X.

(a)  

6.

A numerical variable used in regression analysis to represent subgroups of the sample in your study.

(a)  

7.

A technique used to measure the degree to which the various independent variable and various dependent variables are linearly related to each other.

(a)  

8.

The size of the weight indicates the precision of the information contained in the associated observation.

(a)  

9.

The linear model assumes that the conditional expectation of the dependent variable Y is equal to a linear combination of the explanatory variables X.

(a)  

10.

Can refer to manipulation or dropping of data before it is used in order to ensure or enhance performance.

(a)  

11.

Is the process of detecting and correcting (or removing) corrupt or inaccurate records from a record set, table, or database and refers to identifying incomplete, incorrect, inaccurate or irrelevant parts of the data and then replacing, modifying, or deleting the dirty or coarse data.

(a)  

12.

To obtain a reduced representation of the data set that is much smaller in volume, yet closely maintains the integrity of the original data.

(a)  

13.

Traditionally, data transformation has been a (a)   or batch process, whereby developers write code or implement transformation rules in a data integration tool, and then execute that code or those rules on large volumes of data. This process can follow the linear set of steps as described in the data transformation process above.

14.

There are companies that provide self-service (a)   . They are aiming to efficiently analyze, map and transform large volumes of data without the technical and process complexity that currently exists. While these companies use traditional batch transformation, their tools enable more interactivity for users through visual platforms and easily repeated scripts.

15.

Provide an integrated visual interface that combines the previously disparate steps of data analysis, data mapping and code generation/execution and data inspection. IDT interfaces incorporate visualization to show the user patterns and anomalies in the data so they can identify erroneous or outlying values.

(a)  

16.

(a)   Following the data's transformation into a machine-readable format, it can also be utilized to train machine learning models. Computers can learn without being explicitly programmed thanks to the discipline of machine learning.

17.

A type of artificial intelligence (AI) that allows software applications to become more accurate at predicting outcomes without being explicitly programmed to do so. Machine learning algorithms use historical data as input to predict new output values. Machine learning is frequently used in recommendation engines.

(a)  

18.

The model adjusts its weights as input data is fed into it until it is properly fitted.

(a)  

19.

The ability of this method to identify similarities and differences in data makes it ideal for exploratory data analysis, cross-selling strategies, consumer segmentation, and picture and pattern recognition.

(a)  

20.

The purpose of the statistic is to select the best model using a subset of variables from all available variables.

(a)  

21.

A form of regression that allows multiple linear models to be fitted

to the data for different ranges of X. The regression function at the breakpoint may be discontinuous, but it is possible to specify the model such that the model is continuous at all

points.

(a)  

22.

A model approach to regression

modeling effectively uncovers important data patterns and relationships that are

difficult, if not possible, for other regression methods to reveal.

(a)  

23.

A simple extension of binary logistic regression that allows

for more than two categories of the dependent or outcome variable. Like binary logistic

regression, multinomial logistic regression uses maximum likelihood estimation to evaluate

the probability of categorical membership, you have multiple possible outcomes instead of

just one.

(a)  

24.

A classification tree is an algorithm where the target variable is fixed or categorical. The algorithm is then used to identify the “class” within which a target variable would most likely

fall.

(a)  

25.

One of the simplest forms of pruning is reduced error pruning. Starting at the leaves, each node is replaced with its most popular class. If the prediction accuracy is not affected then the change is kept. While somewhat naive, reduced error pruning has the advantage of

simplicity and speed.

(a)