Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Mastering Data Analytics Concepts

Total questions: 30

Worksheet time: 23mins

Name
Class
Date
1.

What is the primary purpose of data cleaning?

a)

The primary purpose of data cleaning is to improve data quality.

b)

To simplify data analysis

c)

To increase data storage capacity

d)

To enhance data visualization

2.

Which technique is commonly used to handle missing values in a dataset?

a)

Substitution

b)

Deletion

c)

Imputation

d)

Normalization

3.

What does EDA stand for in data analysis?

a)

Effective Data Analysis

b)

Exploratory Data Analysis

c)

Explanatory Data Application

d)

Enhanced Data Assessment

4.

Which library in Python is primarily used for data manipulation?

a)

NumPy

b)

Matplotlib

c)

SciPy

d)

Pandas

5.

What is the function of the 'dropna()' method in Pandas?

a)

The 'dropna()' method filters out duplicate values from a DataFrame or Series.

b)

The 'dropna()' method adds missing values to a DataFrame or Series.

c)

The 'dropna()' method removes missing values from a DataFrame or Series.

d)

The 'dropna()' method sorts the values in a DataFrame or Series.

6.

What is the significance of outlier detection in data cleaning?

a)

Outlier detection is only relevant for large datasets.

b)

Outlier detection reduces the need for data analysis.

c)

Outlier detection improves data quality and accuracy in analysis.

d)

Outlier detection is primarily used for data visualization.

7.

Which visualization tool is known for its interactive capabilities?

a)

Tableau

b)

Power BI

c)

Google Data Studio

d)

Excel Charts

8.

What is the purpose of a scatter plot in exploratory data analysis?

a)

To visualize relationships between two variables and identify patterns or correlations.

b)

To summarize data using averages and totals.

c)

To create a timeline of events in the dataset.

d)

To display the distribution of a single variable.

9.

How can you visualize the distribution of a dataset using Matplotlib?

a)

Use histograms, box plots, or density plots with Matplotlib to visualize the distribution of a dataset.

b)

Use line graphs to show trends over time.

c)

Create pie charts to represent categorical data.

d)

Display scatter plots for correlation analysis.

10.

What is the role of NumPy in data analytics?

a)

NumPy is a programming language for data visualization.

b)

NumPy is primarily used for web development.

c)

NumPy is a database management system.

d)

NumPy provides efficient array operations and mathematical functions, crucial for data analytics.

11.

Which function in Pandas is used to read a CSV file?

a)

load_csv

b)

import_csv

c)

read_csv

d)

fetch_csv

12.

What is the difference between a bar chart and a histogram?

a)

A bar chart is used for displaying percentages, while a histogram is for averages.

b)

A bar chart is for categorical data, and a histogram is for numerical data distribution.

c)

A bar chart uses continuous data, and a histogram uses discrete data.

d)

A bar chart displays trends over time, while a histogram shows categories.

13.

How do you handle categorical data in data visualization?

a)

Use line graphs to represent categorical data.

b)

Apply scatter plots for categorical data visualization.

c)

Utilize histograms for displaying categorical information.

d)

Use bar charts, pie charts, or box plots to visualize categorical data.

14.

What is the purpose of using a box plot?

a)

The purpose of using a box plot is to summarize and compare the distribution of data.

b)

To illustrate the correlation between two variables

c)

To show the exact values of each data point

d)

To display only the mean of the data

15.

Which method in Pandas can be used to group data?

a)

filter

b)

groupby

c)

aggregate

d)

sort

16.

What is the advantage of using Seaborn over Matplotlib?

a)

Seaborn simplifies the creation of complex visualizations and improves aesthetics compared to Matplotlib.

b)

Seaborn is slower than Matplotlib for rendering plots.

c)

Seaborn does not support statistical visualizations.

d)

Matplotlib has better default color palettes than Seaborn.

17.

How can you create a line plot using Matplotlib?

a)

Call plt.display() to show the plot

b)

Only import NumPy for data preparation

c)

Import Matplotlib, prepare data, use plt.plot(), and plt.show()

d)

Use plt.scatter() instead of plt.plot()

18.

What is the purpose of data normalization?

a)

To increase data redundancy for faster access.

b)

To create backups of the data regularly.

c)

The purpose of data normalization is to minimize data redundancy and ensure data integrity.

d)

To encrypt sensitive data for security purposes.

19.

Which function in NumPy is used to calculate the mean of an array?

a)

numpy.median()

b)

numpy.mean()

c)

numpy.sum()

d)

numpy.average()

20.

What is the significance of correlation in exploratory data analysis?

a)

Correlation only applies to linear relationships.

b)

Correlation is significant in exploratory data analysis as it reveals relationships between variables.

c)

Correlation measures the strength of causation between variables.

d)

Correlation is irrelevant in data analysis and should be ignored.

21.

What is the primary function of the 'pivot_table()' method in Pandas?

a)

To visualize data in a graphical format.

b)

To merge two DataFrames into one.

c)

To filter rows based on specific conditions.

d)

To create a new DataFrame by reshaping the existing data.

22.

Which type of plot is best suited for visualizing the relationship between three variables?

a)

Line plot

b)

Bar chart

c)

Box plot

d)

3D scatter plot

23.

What is the purpose of using the 'fillna()' method in Pandas?

a)

To group data based on certain criteria.

b)

To replace missing values with a specified value or method.

c)

To remove rows with missing values from a DataFrame.

d)

To sort the DataFrame by a specific column.

24.

What is the primary benefit of using Pandas for data manipulation?

a)

Pandas is primarily used for creating visualizations.

b)

Pandas is slower than using raw Python lists for data manipulation.

c)

Pandas allows for easy handling of large datasets with its DataFrame structure.

d)

Pandas does not support time series data.

25.

Which type of plot is most effective for comparing multiple categories?

a)

Box plot

b)

Line plot

c)

Bar chart

d)

Histogram

26.

What is the purpose of the 'describe()' method in Pandas?

a)

To provide a summary of statistics for numerical columns in a DataFrame.

b)

To visualize data in a graphical format.

c)

To filter rows based on specific conditions.

d)

To merge two DataFrames into one.

27.

What is the main advantage of using a heatmap in data visualization?

a)

Heatmaps are only useful for categorical data.

b)

Heatmaps do not convey any information about data distribution.

c)

Heatmaps provide a visual representation of data density and relationships.

d)

Heatmaps are primarily used for time series analysis.

28.

What is the purpose of the 'groupby()' function in Pandas?

a)

To create a new DataFrame by merging multiple DataFrames.

b)

To aggregate data based on one or more keys.

c)

To visualize data in a graphical format.

d)

To sort the DataFrame by a specific column.

29.

Which method is used to concatenate two DataFrames in Pandas?

a)

merge()

b)

join()

c)

concat()

d)

append()

30.

Which type of plot is most suitable for visualizing time series data?

a)

Scatter plot

b)

Box plot

c)

Line plot

d)

Bar chart