wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

The MATRIX by DARVIX 2024 - ROUND 1

Total questions: 21

Worksheet time: 12mins

Name
Class
Date
1.

Which of the following is an example of a data visualization software?

a)

SQL

b)

Hadoop

c)

Power BI

d)

Apache Spark

2.

In the context of data analytics, what does 'ETL' stand for?

a)

Extract, Test, Load

b)

Extract, Transform, Load

c)

Execute, Transfer, Load

d)

Extract, Transfer, Load

3.

What is 'overfitting' in the context of machine learning?

a)

When the model performs well on training data but poorly on new data

b)

When the model generalizes well

c)

When the model is too complex and captures noise along with the underlying pattern

d)

When the model uses too little data

4.

What does 'data wrangling' involve?

a)

Data visualization

b)

Cleaning and organizing data for analysis

c)

Storing data in databases

d)

Collecting data from sensors

5.

Identify the person in the picture.

a)

Geoffrey Hinton

b)

Yann LeCun

c)

Sebastian Thrun

d)

Andrew Ng

6.

Which library in Python is commonly used for data manipulation and analysis?

a)

NumPy

b)

SciPy

c)

Pandas

d)

Matplotlib

7.

Which of the following is a type of chart used to show the relationship between two variables?

a)

Bar chart

b)

Line Chart

c)

Pie Chart

d)

Scatter Plot

8.

In data analytics, what does the acronym "SQL" stand for?

a)

Structured Query Language

b)

Statistical Query Language

c)

Statistical Query Logic

d)

System Query Locator

9.

What does V stand for in DARVIX?

(a)  

10.

Which of the following is NOT a primary task in the data analytics process?

a)

Data collection

b)

Data encryption

c)

Data cleaning

d)

Data visualization

11.

Identify the software in the picture.

(a)  

12.

In data visualization, what is a 'choropleth map' used for?

a)

Showing relationships between variables

b)

Displaying time series data

c)

Representing data values using colors on a geographical map

d)

Comparing different categories

13.

Which of the following techniques is used to handle missing data in datasets?

a)

Imputation

b)

Increasing sample size

c)

Data augmentation

d)

Data scaling

14.

Choose the most appropriate option based on the following:


Statement I - Deep learning is a subset of Machine learning.

Statement II - Machine learning is not a subset of Artificial Intelligence.

a)

Only Statement I is true

b)

Only Statement II is true

c)

Both Statements are true

d)

Both Statements are false

15.

Identify the company in the picture.

(a)  

16.

In a survey, 60% of respondents preferred only product A, and 40% preferred only product B. If 20 more respondents had preferred product A, the percentage would have increased to 65%. How many respondents were there in the survey?

(a)  

17.

You have a dataset with a mean of 70 and a standard deviation of 10. If you multiply every data point by 2, what will be the new mean and standard deviation?

a)

Mean = 140, Standard deviation = 20

b)

Mean = 70, Standard deviation = 10

c)

Mean = 140, Standard deviation = 10

d)

Mean = 140, Standard deviation = 40

18.

If the probability density function of a continuous random variable X is given by 𝑓(𝑥)=3𝑥^2

For 0 ≤ 𝑥 ≤ 1, what is the expected value of X?

(a)  

19.

Which statistical test would you use to compare the means of three or more groups?

a)

t-test

b)

Chi-square test

c)

ANOVA

d)

Pearson correlation

20.

Your marketing team has launched a new campaign, and you need to measure its effectiveness. What metrics would you consider, and how would you visualize the results to present to the stakeholders?

a)

Website traffic, conversion rates, and sales data visualized using pie charts

b)

Click-through rates, engagement rates, conversion rates, and sales data visualized using line graphs and bar charts

c)

Social media likes and shares visualized using scatter plots

d)

Customer satisfaction scores visualized using heat maps

21.

You are tasked with reducing customer churn for a subscription-based service. What data would you analyze, and what predictive model would you develop to identify at-risk customers?

a)

Analyze usage patterns, customer feedback, and support tickets, and develop a decision tree model to predict churn

b)

Analyze demographic data and build a linear regression model

c)

Analyze competitors' offerings and use k-means clustering to segment customers

d)

Analyze marketing data and use a time series model to forecast churn