wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

D@+a

Total questions: 75

Worksheet time: 1hrs 29mins

Name
Class
Date
1.

What is data?

a)

Unprocessed infomation

b)

Set of values

c)

Ideas or objects

d)

All of the choices

2.

What is Big Data?

a)

Data with a large size

b)

Data with the word 'big' in it

c)

Data about people who are big

d)

Data made with a big purpose

3.

What are the 3 main concepts of Data Science?

a)

Data, Science, and Knowledge

b)

Mathematics, Computer Science, and Domain Expertise

c)

Machine Learning, Data Processing, and Statistical Research

4.

What is Data Science comprised of?

a)

Predictive

b)

Machine Learning

c)

Business Intelligence

d)

Personal Opinion

5.

Can be described as huge amount of data

a)

Volume

b)

Variety

c)

Value

d)

Velocity

e)

Veracity

6.

Can be described as different formats of data from various sources

a)

Volume

b)

Value

c)

Veracity

d)

Variety

e)

Velocity

7.

Can be described with extracting useful data

a)

Volume

b)

Value

c)

Veracity

d)

Variety

e)

Velocity

8.

Can be described as the high speed of accumulation of data

a)

Volume

b)

Variety

c)

Value

d)

Velocity

e)

Veracity

9.

Can be described as inconsistencies and uncertainty in data

a)

Volume

b)

Value

c)

Veracity

d)

Variety

e)

Velocity

10.

The process of evaluating data through analytical and statistical tools

a)

Data Mining

b)

Data Analysis

c)

Data Exploration

d)

Data Visualization

11.

Which was not mentioned as a latest trend tool

a)

SPSS

b)

Excel

c)

Pentaho

d)

Notepad

12.

Which of the following is performed by Data Scientist ?

a)

Define the question

b)

Create reproducible code

c)

Challenge results

d)

All of the Mentioned

13.

Which of the following is most important language for Data Science ?

a)

Java

b)

Ruby

c)

R

d)

None of the mentioned

14.

Which of the following approach should be used to ask Data Analysis question ?

a)

Find only one solution for particular problem

b)

Find out the question which is to be answered

c)

Find out answer from dataset without asking question

d)

None of the mentioned

15.

Which of the following is one of the key data science skill ?

a)

Statistics

b)

Machine Learning

c)

Data Visualization

d)

All of the Mentioned

16.

Which of the following is key characteristic of hacker ?

a)

Afraid to say they don’t know the answer

b)

Willing to find answers on their own

c)

Not Willing to find answers on their own

d)

All of the mentioned

17.

Point out the correct statement:

a)

Raw data is original source of data

b)

Preprocessed data is original source of data

c)

Raw data is the data obtained after processing steps

d)

None of the Mentioned

18.

Which of the following is the top most important thing in data science ?

a)

answer

b)

question

c)

data

d)

none of the Mentioned

19.

Which of the following term is appropriate to the above figure ?

a)

Large Data

b)

Big Data

c)

Dark Data

d)

None of the mentioned

20.

Which of the following characteristic of big data is relatively more concerned to data science ?

a)

Velocity

b)

Variety

c)

Volume

d)

None of the Mentioned

21.

Which of the following step is performed by data scientist after acquiring the data ?

a)

Data Integration

b)

Data Replication

c)

Data Cleansing

d)

All of the Mentioned

22.

Which of the following is the most accurate definition of data science?

a)

a) Data science is extracting meaning from large data sets in order to provide insights to support decision-making

b)

b) Data science is using computers to analyse data and to perform calculations on the data to produce information

c)

c) Data science is performing experiments and recording the data produced by those experiments

d)

d) Data science is writing code to make sure that any inaccuracies in data sets are spotted and removed (cleaned)

23.

Which of the following best describes a data visualisation?

a)

A. Presenting related data so that a user can see individual items of data

b)

B. Making sure that data is accessible and that no data is hidden

c)

C. A visual representation that communicates relationships among the data

d)

D. A collection of graphs that tell a story when put together

24.

Is the following a visualisation or an infographic?

a)

A. Visualisation

b)

B. Infographic

25.

What is meant by a correlation?

a)

A. The relationship between two or more variables

b)

B. When there is an upward trend in a graph

c)

C. When there is a set of data that doesn’t lie in the normal or expected range

d)

D. When data is placed in a graph

26.

The visualisation below plots life expectancy (y-axis) against time (x-axis). What type of correlation does this visualisation show from 1949 onwards?

a)

A. Positive

b)

B. Negative

c)

C. Neutral

d)

D. No correlation is visible

27.

The following graph shows the annual average temperatures recorded by a weather station over a period of 20 years. Identify the outlier in the data.

a)

A. Point A

b)

B. Point B

c)

C. Point C

d)

D. Point D

28.

Which of the following is the correct order of the investigative cycle?

a)

A. Problem, data, plan, analysis, conclusion

b)

B. Conclusion, plan, problem, data, Analysis

c)

C. Plan, problem, analysis, data, conclusion

d)

D. Problem, plan, data, analysis, conclusion

29.

In which step of the cycle would you pose the question(s) that you will use data to help you answer?

a)

A. Data

b)

B. Analysis

c)

C. Plan

d)

D. Problem

30.

In which step of the cycle would you cleanse the data?

a)

A. Data

b)

B. Analysis

c)

C. Plan

d)

D. Problem

31.

In which step of the cycle would you work out where the data will come from or how you will collect it?

a)

A. Data

b)

B. Analysis

c)

C. Plan

d)

D. Problem

32.

How many students spent 7 hours doing homework that week?(1 mark)

a)

6

b)

5

c)

7

d)

4

33.

How many total students are represented by the histogram?(1 mark)

a)

32

b)

31

c)

35

d)

8

34.

What is the MEDIAN of this data?(1 mark)

a)

95

b)

90

c)

100

d)

40

35.

What was the difference between the cars sold on Monday and Tuesday than the cars sold on Friday and Saturday? (1 mark)

a)

4

b)

3

c)

2

d)

1

36.

What is name of the chart pictured below? (1 mark)

a)

Graph chart

b)

Pie chart

c)

Line chart

d)

Bar chart

37.

What is name of the chart pictured below? (1 mark)

a)

Graph chart

b)

Pie chart

c)

Line chart

d)

Bar chart

38.

What is name of the chart pictured below? (1 mark)

a)

Graph chart

b)

Pie chart

c)

Line chart

d)

Bar chart

39.

What is Machine Learning? (Choose 3 Answers)

a)

Artificial Intelligence

b)

Machine Learning

c)

Data Statistics

d)

Deep Learning

40.

Which one in the following is not Machine Learning disciplines?

a)

Information Theory

b)

Neurostatistics

c)

Optimization + Control

d)

Physics

41.

from the picture, what kind of programming is it?

a)

Traditional Programming

b)

Modern Programming

c)

Machine Learning

d)

Traditional Learning

42.

from the picture, what kind of programming is it?

a)

Traditional Programming

b)

Machine Learning

c)

Modern Programming

d)

Traditional Learning

43.

What kind of learning algorithm for "Future stock prices or currency exchange rates"?

a)

Recognizing Anomalies

b)

Prediction

c)

Generating Patterns

d)

Recognition Patterns

44.

What kind of learning algorithm for "Facial identities or facial expressions"?

a)

Recognizing Anomalies

b)

Prediction

c)

Generating Patterns

d)

Recognition Patterns

45.

Which of the following is not type of learning?

a)

Semi-unsupervised Learning

b)

Unsupervised Learning

c)

Supervised Learning

d)

Reinforcement Learning

46.

Real-Time decisions, Game AI, Learning Tasks, Skill Aquisition, and Robot Navigation are applications in ...

a)

Unsupervised Learning: Clustering

b)

Supervised Learning: Classification

c)

Reinforcement Learning

d)

Unsupervised Learning: Regression

47.

Targetted marketing, Recommended Systems, and Customer Segmentation are applications in ...

a)

Unsupervised Learning: Clustering

b)

Supervised Learning: Classification

c)

Reinforcement Learning

d)

Unsupervised Learning: Regression

48.

Fraud Detection, Image Classification, Diagnostic, and Customer Retention are applications in ...

a)

Unsupervised Learning: Clustering

b)

Supervised Learning: Classification

c)

Reinforcement Learning

d)

Unsupervised Learning: Regression

49.

This picture shows a result of ...

a)

Supervised Learning: Classification

b)

Unsupervised Learning: Regression

c)

Unsupervised Learning: Prediction

d)

Supervised Learning: Regression

50.

This picture shows an application of ...

a)

Supervised Learning: Classification

b)

Unsupervised Learning: Clustering

c)

Unsupervised Learning: Prediction

d)

Supervised Learning: Regression

51.

Machine Learning has various function representation, which of the following is not function of symbolic?

a)

Decision Trees

b)

Rules in propotional Logic

c)

Hidden-Markov Models (HMM)

d)

Rules in first-order predicate logic

52.

Machine Learning has various function representation, which of the following is not numerical functions?

a)

Linear Regression

b)

Support Vector Machines

c)

Neural Network

d)

Case-based

53.

Machine Learning has various search/ optimization algorithms, which of the following is not evolutionary computation?

a)

Perceptron

b)

Genetic Algorithm (GA)

c)

Neuro Evolution

d)

Genetic Programming (GP)

54.

What type of Machine Learning Algorithm is suitable for predicting the continuous dependent variable?

a)

Logistic Regression

b)

Linear Regression

c)

Decision Tree Classifier

d)

KNN Classifier

55.

What type of Machine Learning Algorithm is suitable for predicting the dependent variable with two different values?

a)

Logistic Regression

b)

Linear Regression

c)

Multiple Linear Regression

d)

Polynomial Regression

56.

The correlation in between mobile usage and exam score of a person found to be -2.2. What is your inference from the above statement.

a)

Mobile usage is positively correlated with exam score

b)

Mobile usage is negatively correlated with exam score

c)

None of the mentioned

d)

Need some other information

57.

The residual is the difference in between ________________

a)

actual value of y and the estimated value of y

b)

actual value of x and the estimated value of x

c)

actual value of y and the estimated value of x

d)

actual value of x and the estimated value of y

58.

Suitable evaluation metric for measuring the performance of a given regression model is

a)

Mean Absolute Error

b)

Root Mean Square Error

c)

Precision

d)

Recall

59.

If we decrease the input variable by one unit in a simple linear regression model. How many units of the output variable will change?

a)

reduced by Intercept

b)

increased by Intercept

c)

increased by Slope

d)

reduced by Slope

60.

Appropriate chart for visualizing the linear relationship between two variables is _________________

a)

Scatter plot

b)

Barchart

c)

Histograms

d)

None of Mentioned

61.

The Number of coefficients required to estimate a simple linear regression?

a)

1

b)

2

c)

0

d)

3

62.

KNN Algorithm can be used for

a)

Only for Classification

b)

Only for Regression

c)

Both Classification and Regression

d)

None of the Mentioned

63.

KNN is ___________ algorithm

a)

Non-parametric and Lazy Learning

b)

Parametric and Lazy Learning

c)

Parametric and Eager Learning

d)

Non-parametric and Eager Learning

64.

What kind of distance metric(s) are suitable for categorical variables to finding the closest neighbors

a)

Euclidean Distance

b)

Manhattan distance

c)

Minkowski distance

d)

Hamming distance

65.

What kind of distance metric(s) are suitable for continuous variables to find the closest neighbors

a)

Euclidean Distance

b)

Manhattan distance

c)

Minkowski distance

d)

Hamming distance

66.

KNN algorithm appropriate for

a)

Lower number of features

b)

Large number of features

c)

No such restriction on number of features

d)

None of the Mentioned

67.

KNN algorithm requires

a)

More time for training

b)

More time for testing

c)

Equal time for training and testing

d)

None of the Mentioned

68.

The entropy of a given dataset is zero. This statement implies what?

a)

further splitting is required

b)

no further splitting is required

c)

Need some other information to decide splitting

d)

None of the Mentioned

69.

If the given dataset contains 100 observations out of 50 belongs to class1 and other 50 belongs to class2. What will be the entropy of the given dataset?

a)

0

b)

1

c)

-1

d)

0.5

70.

How do you choose the root node while constructing a Decision Tree?

a)

An attribute having high entropy

b)

An attribute having largest information gain

c)

An attribute having high entropy and Information gain

d)

None of the Mentioned

71.

Chose the correct criterion for Decision Tree Classifier in sklearn package

a)

Gini

b)

Entropy

c)

Information Gain

d)

Random

72.

In a Decision Tree Leaf Node represents_____________

a)

One of the Class Label

b)

One of the complete observation

c)

One of the attribute

d)

None of the Mentioned

73.

Consider the above Confusion Matrix of a classifier and choose the correct statements

a)

Accuracy is 84%

b)

Misclassification Rate is 16%

c)

Type-I Error is 6

d)

Type-II Error is 10

74.

Artificial Intelligence is superset of ________________________ & ________________________ ,

a)

Machine Learning & Neural Networks

b)

Machine Learning & Deep Learning

c)

Deep Learning & Neural Networks

75.

Machine Learning is a subset of AI. ML deals with developing systems which can improve their performance with _________________.

a)

Experience

b)

Deep Learning

c)

Neural Networks