wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Data Principles week 2 part 1

Total questions: 5

Worksheet time: 54mins

Name
Class
Date
1-100.
1.

What are the distinctions between the model planning and model building phase?

a)

Model planning involves data collection, while model building involves data analysis.

b)

Model planning is about choosing the right tools, while model building is about implementing the model.

c)

Model planning focuses on the theoretical framework, while model building is the practical application.

d)

There are no distinctions; both phases are the same.

2.

What are some key considerations in model building?

a)

Selecting the right dataset and tools for implementation.

b)

Ensuring the model is overfit to the training data.

c)

Ignoring the problem statement and focusing on the algorithm.

d)

Choosing a complex model for simple tasks.

3.

What software tools (commercial, open source) are typically used at this phase?

a)

Word processors and spreadsheet software.

b)

Web browsers and email clients.

c)

Data analysis and machine learning libraries.

d)

Graphic design and video editing software.

4.

What is Linear Regression model and in what situation is it appropriate?

a)

It is a classification model used for image recognition tasks.

b)

It is a regression model used for predicting numerical values based on independent variables.

c)

It is a clustering model used for grouping similar data points.

d)

It is a reinforcement learning model used for real-time decision making.

5.

How does the Linear Regression model work for predictive modelling tasks?

a)

By finding the median value of the dependent variable.

b)

By clustering data points into different categories.

c)

By establishing a relationship between dependent and independent variables using a straight line.

d)

By using decision trees to make predictions.

6.

How do we prepare our data prior to applying the Linear Regression model?

a)

By converting all data into text format.

b)

By ensuring data is clean, relevant, and properly formatted.

c)

By randomly shuffling the data to create diversity.

d)

By deleting all outliers without analysis.

7.

What is one of the learning outcomes related to the processes within the Data Analytics Lifecycle?

a)

Evaluate suitable techniques and tools for specific data science tasks.

b)

Develop analytics plan for a given business case study.

c)

Describe the processes within the Data Analytics Lifecycle.

d)

Create a marketing strategy for data analytics.

8.

Which learning outcome focuses on the analysis of business and organisational problems?

a)

Evaluate suitable techniques and tools for specific data science tasks.

b)

Analyse business and organisational problems and formulate them into data science tasks.

c)

Develop analytics plan for a given business case study.

d)

Summarize the history of data analytics.

9.

What does the learning outcome related to evaluating techniques and tools pertain to?

a)

Developing a comprehensive understanding of data analytics history.

b)

Evaluating suitable techniques and tools for specific data science tasks.

c)

Describing the processes within the Data Analytics Lifecycle.

d)

Analyzing the effectiveness of data analytics in marketing.

10.

What is the goal of the learning outcome that involves developing an analytics plan?

a)

To create a new data analytics software.

b)

To evaluate the impact of data analytics on global markets.

c)

To develop analytics plan for a given business case study.

d)

To describe the ethical considerations in data analytics.

11.

What is the first phase of the Data Analytics Lifecycle according to the diagram?

a)

Data Prep

b)

Model Planning

c)

Discovery

d)

Operationalize

12.

Which phase in the Data Analytics Lifecycle is directly after Model Planning?

a)

Discovery

b)

Data Prep

c)

Model Building

d)

Communicate Results

13.

What question should be asked during the Model Building phase of the Data Analytics Lifecycle?

a)

Do I have enough information to draft an analytic plan?

b)

Is the model robust enough? Have we failed enough?

c)

Do I have a good idea about the type of model to try? Can I refine the analytic plan?

d)

Do I have enough “good” data to start building the model?

14.

Before moving on from the Data Prep phase, what question must be addressed?

a)

Do I have enough “good” data to start building the model?

b)

Do I have a good idea about the type of model to try? Can I refine the analytic plan?

c)

Is the model robust enough? Have we failed enough?

d)

Do I have enough information to draft an analytic plan?

15.

What is the final phase of the Data Analytics Lifecycle as shown in the diagram?

a)

Model Planning

b)

Data Prep

c)

Communicate Results

d)

Operationalize

16.

What is the main activity of the data science team during the 4th phase of the Data Analytics Lifecycle (DAL)?

a)

Evaluating the results of the models

b)

Developing datasets for testing, training, and production purposes

c)

Publishing the final report

d)

Collecting new data

17.

In the 4th phase of the Data Analytics Lifecycle, what does the team build and execute?

a)

Data collection strategies

b)

Data visualization tools

c)

Models based on the work done in the model planning phase

d)

Marketing campaigns

18.

What does the team consider regarding the tools used to run the models in the 4th phase of the Data Analytics Lifecycle?

a)

The cost of the tools

b)

The sufficiency of the existing tools

c)

The color scheme of the tools' user interface

d)

The brand of the tools

19.

What might the team need to ensure a more robust environment for executing the models?

a)

Slower processing speeds

b)

Basic spreadsheet software

c)

Fast hardware, parallel processing, etc.

d)

A smaller dataset

20.

What is the key activity in Phase 4 – Model Building?

a)

Collecting more data

b)

Developing an analytical model, fitting it on the training data, and evaluating its performance on the test data

c)

Presenting the findings to stakeholders

d)

Cleaning the data

21.

When can the data science team move to the next phase after Phase 4 – Model Building?

a)

When the model is sufficiently robust to solve the problem

b)

After a fixed period of time regardless of model performance

c)

When the model's accuracy is below a certain threshold

d)

If the team has failed

22.

What does the development of an analytical model in Phase 4 typically involve?

a)

Coding an entirely new analytics model from scratch

b)

Selecting and experimenting with various models and fine-tuning their parameters

c)

Solely evaluating the model against the test dataset

d)

Iterating back to data preparation without considering model parameters

23.

Can the model building and model planning phases overlap in the data science process?

a)

Yes, they can overlap and involve iteration between the two phases

b)

No, they are strictly separate phases with no overlap

c)

Yes, but only after the final model has been settled on

d)

No, because model building is always shorter than model planning

24.

Compared to the time spent preparing data and planning the model, how is the actual duration of model building generally described?

a)

Longer than the time spent for preparing data and planning the model

b)

The same as the time spent for preparing data and planning the model

c)

Shorter than the time spent for preparing data and planning the model

d)

Not comparable as they are unrelated tasks

25.

Why is documentation considered important during the model building phase?

a)

It helps in marketing the final product.

b)

It is required for legal compliance.

c)

It prevents the details and decisions made during the modeling from being forgotten.

d)

It is only necessary for training new employees.

26.

What should be recorded during the model building phase to ensure the process is well-documented?

a)

The financial costs of the project.

b)

The results and logic of the model, as well as any operating assumptions made.

c)

The names of the team members involved.

d)

The software tools used for data analysis.

27.

What is the purpose of SAS Enterprise Miner according to the slide?

a)

It provides a high-level language for data analytics.

b)

It offers methods to explore and analyze data through GUI.

c)

It allows users to run predictive and descriptive models based on large volumes of data from across the enterprise.

d)

It provides a GUI front end for users to develop analytic workflows.

28.

Which tool is described as offering methods to explore and analyze data through GUI?

a)

SAS Enterprise Miner

b)

IBM SPSS Modeler

c)

Matlab

d)

Chorus 6

29.

What does Matlab provide according to the information on the slide?

a)

Predictive modeling capabilities

b)

A GUI front end for developing analytic workflows

c)

A high-level language for performing a variety of data analytics, algorithms, and data exploration

d)

Methods to analyze data through GUI

30.

What functionality does Chorus 6 offer?

a)

Predictive and descriptive model running

b)

High-level language for data analytics

c)

GUI front end for developing analytic workflows and interaction with Big Data tools

d)

Exploration and analysis of data through GUI

31.

What is PL/R?

a)

A data mining package

b)

A GUI software for data processing

c)

A procedural language for PostgreSQL with R

d)

A programming language for computational modeling

32.

Which open-source tool is described as a GUI ready software for easier data processing?

a)

Python

b)

KNIME

c)

WEKA

d)

MADlib

33.

What does WEKA offer?

a)

Machine learning library for PostgreSQL

b)

Computational modeling functionalities

c)

An analytic workbench and rich Java API

d)

A procedural language for database commands

34.

Which of the following is not a feature of Python as mentioned in the image?

a)

Data mining package

b)

Machine learning and data visualization packages

c)

NumPy and SciPy

d)

pandas and matplotlib

35.

What are predictive models used for?

a)

Finding specific patterns or structures within the data

b)

Predicting certain attributes of a given object

c)

Guessing the weather

d)

Calculating the speed of an object

36.

How does the goal of a predictive model differ from that of unsupervised models like k-Means Clustering?

a)

Predictive models are used for calculating probabilities, while unsupervised models are not

b)

Predictive models are used for guessing whether a customer will subscribe to a service, while unsupervised models are used for customer segmentation

c)

Predictive models are used for predicting outcomes, while unsupervised models are limited to finding specific patterns or structures within the data

d)

There is no difference between the goals of predictive models and unsupervised models

37.

Which of the following is an example of a predictive model application?

a)

Predicting whether a customer will subscribe to a product or service

b)

Finding the average height of a population

c)

Organizing books in a library

d)

Calculating the area of a farm

38.

What type of attribute are predictive models typically used to predict in classification problems?

a)

Numerical

b)

Categorical

c)

Continuous

d)

Binary

39.

What is the term used to describe the dataset that a model learns from in classification problems?

a)

Evaluation dataset

b)

Test dataset

c)

Training dataset

d)

Validation dataset

40.

What kind of models are most classification models categorized as?

a)

Unsupervised models

b)

Supervised models

c)

Reinforcement models

d)

Semi-supervised models

41.

What is the purpose of the test dataset in classification problems?

a)

To train the model

b)

To validate the model during training

c)

To test the model's predictions

d)

To provide additional examples for the model to learn from

42.

What are the class labels in the given example?

a)

'yes', 'maybe'

b)

'true', 'false'

c)

'yes', 'no'

d)

'1', '0'

43.

What is the purpose of a training dataset?

a)

To assess the strength and utility of a predictive relationship

b)

To minimize the possible overfitting of a model

c)

To discover a predictive relationship

d)

To partition the data randomly

44.

What is the test dataset used for?

a)

To discover a predictive relationship

b)

To assess the strength and utility of a predictive relationship

c)

To minimize the possible overfitting of a model

d)

To partition the data randomly

45.

How are training and test datasets usually related to each other?

a)

They are the same dataset

b)

They are overlapping datasets

c)

They are independent from each other

d)

They are sequentially ordered datasets

46.

What is a validation dataset?

a)

A dataset used to discover a predictive relationship

b)

A dataset used to assess the strength and utility of a predictive relationship

c)

A dataset used to minimize the possible overfitting of a model

d)

A dataset used to partition the data randomly

47.

What is the goal of Linear Regression?

a)

To classify data into different categories

b)

To understand the relationship between input and output variables

c)

To estimate the probability of an event

d)

To cluster similar data points together

48.

How old is the Linear Regression model?

a)

More than 100 years old

b)

More than 200 years old

c)

More than 300 years old

d)

More than 400 years old

49.

What does the Linear Regression model assume about the relationship between input variables and the output variable?

a)

A non-linear relationship exists

b)

No relationship exists

c)

A linear relationship exists

d)

A quadratic relationship exists

50.

What type of values is Linear Regression limited to predicting?

a)

Categorical values

b)

Numerical values

c)

Textual values

d)

Boolean values

51.

Which of the following is an advantage of Linear Regression?

a)

Works well for modeling non-linear relationships

b)

Does not require numerical input variables

c)

Simplicity and gives optimal results when relationships are linear

d)

Can predict any type of value

52.

Which of the following is a disadvantage of Linear Regression?

a)

Too complex to understand

b)

Cannot handle large datasets

c)

Will not work for modeling non-linear relationships

d)

Requires a large number of input variables

53.

What does the equation y = f(x) = a · x + b represent in the context of the image?

a)

A) A quadratic equation

b)

B) A polynomial equation

c)

C) A linear regression model

d)

D) A logarithmic function

54.

In the context of the image, what does the variable 'y' represent in the linear regression model?

a)

A) The input variable (height)

b)

B) The coefficient of determination

c)

C) The output variable (weight)

d)

D) The slope of the line

55.

What does the variable 'x' represent in the linear regression model as shown in the image?

a)

A) The output variable (weight)

b)

B) The input variable (height)

c)

C) The slope of the line

d)

D) The y-intercept of the line

56.

Based on the graph shown in the image, what happens to the weight as the height increases?

a)

A) The weight decreases.

b)

B) The weight stays the same.

c)

C) The weight increases.

d)

D) The relationship between weight and height is not clear.

57.

Which of the following best describes the purpose of the different lines in the right graph of the image?

a)

A) They represent different quadratic equations.

b)

B) They show the effect of changing the 'a' value in the equation y = f(x) = a · x + b.

c)

C) They indicate the variability of the data points around the regression line.

d)

D) They represent different categories of data.

58.

Linear Regression belongs to which type of learning approach?

a)

Non-parametric learning

b)

Parametric learning

c)

Supervised learning

d)

Unsupervised learning

59.

What is the objective of building a model in the context of Linear Regression?

a)

To classify data into different categories

b)

To estimate the best values for unspecified numeric parameters from the training data

c)

To cluster data into different groups

d)

To reduce the dimensionality of the data

60.

How can attributes for a Linear Regression model be chosen?

a)

Based on the color of the data points

b)

Randomly without any specific criteria

c)

Based on domain knowledge or attribute selection techniques

d)

Based on the size of the data set

61.

What does y represent in the Linear Regression model equation?

a)

The value of the input variable when all other variables are zero

b)

The predicted output variable

c)

The bias coefficient / intercept

d)

One of the parameters / weights / coefficients of the input values

62.

What is w_0 in the Linear Regression model?

a)

The predicted output variable

b)

One of the input variables

c)

The bias coefficient / intercept

d)

A parameter that needs to be estimated from the training data

63.

What do w_1, w_2, ... represent in the Linear Regression model?

a)

The predicted output variable

b)

The bias coefficient / intercept

c)

The input variables

d)

The parameters / weights / coefficients of the input values

64.

What are x_1, x_2, ... in the context of the Linear Regression model?

a)

The predicted output variables

b)

The bias coefficients / intercepts

c)

The values of the input variables

d)

The parameters that need to be estimated from the training data

65.

What is the purpose of using different values for w0 and w1 in a linear regression model?

a)

To adjust the slope and intercept of the regression line

b)

To calculate the area under the curve

c)

To determine the correlation coefficient

d)

To measure the variance of the dependent variable

66.

Based on the left graph, which equation represents the steepest line?

a)

y = x

b)

y = 2x

c)

y = 0.5x

d)

y = x + 1

67.

Based on the right graph, which equation represents a line with a positive y-intercept?

a)

y = x

b)

y = x + 1

c)

y = x - 2

d)

y = 0.5x

68.

What does the parameter w1 represent in the context of the graphs provided?

a)

The y-intercept of the line

b)

The slope of the line

c)

The correlation between x and y

d)

The maximum value of y

69.

What does the parameter w0 represent in the context of the graphs provided?

a)

The slope of the line

b)

The maximum value of y

c)

The y-intercept of the line

d)

The correlation between x and y

70.

What is the task described in the example for Linear Regression Model?

a)

Calculate the mean of the Y values.

b)

Build a simple Linear Regression model that predicts the value of Y when the value of X is known.

c)

Determine the correlation coefficient between X and Y.

d)

Plot the data points on a graph.

71.

What type of plot is shown in Figure 1?

a)

Line plot

b)

Bar chart

c)

Pie chart

d)

Scatter plot

72.

Based on Table 1, what is the Y value when X is 2.00?

a)

1.00

b)

2.00

c)

1.30

d)

3.75

73.

According to the example data provided, which of the following statements is true?

a)

The value of Y increases as X increases.

b)

The value of Y is always equal to the value of X.

c)

The value of Y decreases as X increases.

d)

The value of Y is independent of the value of X.

74.

What is the purpose of the regression line in a linear regression model?

a)

To represent the maximum values of Y for each X

b)

To predict the value of Y for any given value of X

c)

To connect all the data points in the scatter plot

d)

To minimize the errors of prediction for each point

75.

Why does the regression line not need to pass exactly over all the actual points on the scatterplot?

a)

Because it is only a rough estimation

b)

Because it represents the minimum values of Y

c)

To avoid an overfitting problem

d)

To ensure that the line is perfectly straight

76.

What do the vertical lines between the data points and the regression line represent?

a)

The predicted values of Y

b)

The actual values of X

c)

The errors of prediction

d)

The maximum errors allowed

77.

What is the purpose of finding the best-fitting regression line in Linear Regression?

a)

To maximize the prediction error

b)

To categorize the data points into different groups

c)

To minimize the prediction error

d)

To calculate the mean value of the data points

78.

What does SSE stand for in the context of Linear Regression?

a)

Sum of Squared Estimates

b)

Sum of Squared Errors

c)

Sum of Standard Errors

d)

Sum of Squared Exponents

79.

What does the 'r' value indicate in a Linear Regression model?

a)

The range of the data points

b)

The ratio of the dependent variable to the independent variable

c)

How well the model fits the data

d)

The residual value of the regression

80.

What is the absolute error value represented by in the table?

a)

Y

b)

Y'

c)

Y-Y'

d)

(Y-Y')^2

81.

How is the squared error value calculated for each data point in the table?

a)

By squaring the Y value

b)

By squaring the Y' value

c)

By squaring the absolute error value (Y-Y')

d)

By adding the Y and Y' values

82.

What is the value of y when x = 1 in the given linear regression model equation?

a)

1.21

b)

1.64

c)

2.10

d)

0.785

83.

What is the value of y when x = 3 using the regression model equation y = 0.785 + 0.425x?

a)

2.07

b)

2.36

c)

1.64

d)

1.21

84.

What are the five statistics required to calculate the Linear Regression equation?

a)

mean of X, mean of Y, standard deviation of X, standard deviation of Y, Pearson's correlation coefficient

b)

median of X, median of Y, variance of X, variance of Y, covariance of X and Y

c)

mode of X, mode of Y, range of X, range of Y, Pearson's correlation coefficient

d)

mean of X, mean of Y, variance of X, variance of Y, Spearman's correlation coefficient

85.

How is the mean of X (μx) calculated in the context of Linear Regression?

a)

μx = ∑i=1N xi / N

b)

μx = ∑i=1N xi^2 / N

c)

μx = ∑i=1N xi / N^2

d)

μx = ∑i=1N (xi - x̄) / N

86.

What does the standard deviation measure in a set of random numbers?

a)

The average value of the numbers

b)

How far a set of random numbers are spread out from their average value (mean)

c)

The sum of all numbers in the set

d)

The difference between the highest and lowest values in the set

87.

What is the term used to describe the part of the equation \(\sum_{i=1}^{N} (y_i - \mu_y)^2\) in the context of the linear regression model example?

a)

Sample mean

b)

Sample median

c)

Sample variance

d)

Sample range

88.

What percentage of data falls within one standard deviation from the mean in a normal distribution, according to the diagram?

a)

68%

b)

95%

c)

99.7%

d)

34%

89.

What does Pearson's correlation coefficient measure?

a)

The mean of the data points

b)

The strength of association between two variables

c)

The total number of data points

d)

The slope of the linear regression line

90.

Which of the following represents the formula for Pearson's correlation coefficient?

a)

r_xy = Σ(x_i - μ_x)(y_i - μ_y) / sqrt(Σ(x_i - μ_x)^2 * Σ(y_i - μ_y)^2)

b)

r_xy = Σ(x_i - μ_x)^2 / N

c)

r_xy = Σ(x_i * y_i) / N

d)

r_xy = Σ(x_i - μ_x)(y_i - μ_y) / N

91.

In the context of Pearson's correlation coefficient, what does 'N' stand for?

a)

The slope of the regression line

b)

The mean of the x variables

c)

The total number of data points

d)

The standard deviation of the y variables

92.

Based on the scatter plots provided, which correlation coefficient value indicates the strongest association between two variables?

a)

r = 0.1

b)

r = 0.5

c)

r = 0.8

d)

r = -0.2

93.

What is the value of the slope (w_x) in the linear regression formula y = 0.785 + 0.425x?

a)

0.785

b)

0.425

c)

2.06

d)

3.00

94.

How is the y-intercept (w_0) calculated in the linear regression example?

a)

w_0 = μ_y + w_xμ_x

b)

w_0 = μ_y - w_xμ_x

c)

w_0 = μ_y / w_xμ_x

d)

w_0 = μ_y * w_xμ_x

95.

What is the value of the y-intercept (w_0) in the linear regression formula?

a)

0.785

b)

0.425

c)

2.06

d)

1.581

96.

What is the correlation coefficient (r_xy) given in the statistics for the linear regression model?

a)

1.581

b)

1.072

c)

0.627

d)

3.00

97.

What does the ordinary least squares regression aim to minimize?

a)

The sum of the absolute errors of each data point

b)

The sum of the squared error of each data point

c)

The product of the squared errors of each data point

d)

The maximum error of any single data point

98.

In the context of ordinary least squares regression, what is the goal when there are multiple input variables?

a)

To maximize the coefficients of each input variable

b)

To minimize the coefficients of each input variable

c)

To estimate the parameter value of each input variable

d)

To eliminate the need for input variables

99.

According to the text, how is the optimization problem of ordinary least squares regression typically solved in practice?

a)

By using a calculator

b)

By doing it manually

c)

By using data science software packages

d)

By ignoring the optimization problem

100.

What assumption does Linear Regression make about the relationships between the input and output variables?

a)

The relationships are non-linear.

b)

The relationships are linear.

c)

The relationships are exponential.

d)

The relationships are logarithmic.

101.

What should be done to prepare data for Linear Regression to ensure it is clean?

a)

Add more noise and outliers

b)

Apply data cleaning techniques to remove noice and outliers

c)

Ignore the noice and outliers

d)

Increase the number of variables

102.

What principle is applied to remove collinearity in the context of preparing data for Linear Regression?

a)

Newton's Third Law

b)

Occam's Razor

c)

Murphy's Law

d)

Pareto Principle

103.

According to Occam's Razor, how should a model be designed for an event?

a)

With the most complex explanation.

b)

With the least correlated variables.

c)

With the simplest explanation.

d)

With the maximum number of variables.

104.

What can be used to solve a non-linear problem as a linear one?

a)

Linear regression

b)

Non-linear transformation

c)

Exponential smoothing

d)

Logarithmic reduction