WorksheetsDVB8
Total questions: 10
Worksheet time: 12mins
You have a dataset with five numeric variables and want to see how all variables relate to each other at once. Which visualization is most appropriate to start exploring pairwise relationships?
Box plot of the main outcome variable by one grouping variable
Heatmap of all numeric variables
Single scatter plot between the two most important variables
Histogram for each variable separately
Which method reduces dimensionality by creating orthogonal linear combinations of original features ordered by variance explained?
Decision Tree
Principal Component Analysis
Feature selection using chi-square test
Point-biserial Correlation
You're analyzing a left-skewed distribution of household incomes. Which measure of central tendency will be larger than the median?
Mean
Median will be largest
Mode
Mean and median will be at same point
For two reverse-ordered permutations, Kendall's τ equals:
0
1
-1
Undefined
In Python's matplotlib, which function renders the figure to the screen?
For visualizing hourly stock price movements across five different companies over a single trading day, which visualization approach is most effective?
Stacked area chart with companies as layers
Multiple line plots (one per company) on shared x/y axes
Heatmap with hours as rows and companies as columns
Grouped bar chart with hourly bars for each company
When clustering data with a mix of features measured in dollars (0-1,000,000) and percentages (0-100), which preprocessing step is most critical for k-means using Euclidean distance?
Centering the data by subtracting the mean from each feature
Feature scaling to comparable ranges
Removing outliers beyond 3 standard deviations
Converting percentages to decimal format
A variable with original mean=50 and variance=100 undergoes min-max scaling to the range [0, 1]. What is the effect on its variance?
Divided by 100
Multiplied by 10⁻⁴
Becomes exactly 1
Remains unchanged (scale-invariant)
For displaying the 95th percentile response time of a web server every hour over a 7-day period, the most effective visualization would be:
Box plot for each day
Line chart with hourly data points
Heatmap of hour vs. day
Histogram of all response times
You're analyzing Netflix subscriber churn using a dataset with 4 age groups (18-25, 26-35, 36-50, 50+), 3 usage levels (Low, Medium, High), and churn status. Your goal is to compare churn rates across both age and usage dimensions to identify high-risk segments for targeted retention campaigns. Which visualization approach best supports this analysis?
Create a grouped bar chart: Age groups on the x-axis, churn rate percentage on the y-axis, with three adjacent bars per age group (one for each usage level colored distinctly: Low=blue, Medium=orange, High=red). Include error bars for confidence intervals and a horizontal reference line at the overall churn rate.
Design a 3D pie chart with 12 slices (one per age-usage combination), each labeled with churn percentage, exploded slightly for emphasis, and a color legend mapping age groups to hues (with usage depth controlled by saturation). Add a comprehensive data table beneath the chart.
Build a stacked bar chart: Age groups on the x-axis, total bar height represents total subscribers, with each bar divided into three colored segments showing the proportion of Low, Medium, and High usage levels within churned customers. Overlay churn rate percentages as text labels on each segment.
Generate a scatter plot with subscriber age on the x-axis (treated as continuous), support calls on the y-axis, points colored by churn status (red/blue), and point size scaled to usage frequency. Add a trend line and confidence bands to show age-churn correlation.
