NEW
Font size
WorksheetsIntroduction to Data Analysis
Total questions: 54
Worksheet time: 27mins
Which statement best defines data analysis in engineering contexts?
Visualizing results without transforming data
Designing sensors for measurement systems
Collecting, storing, backing up raw datasets
Inspecting, cleaning, transforming, modeling data
A study summarizes the distribution of material thickness from one sensor. Which type of analysis is being performed?
Univariate analysis of one variable
Time-series forecasting across seasons
Bivariate analysis of two variables
Multivariate analysis of many variables
An engineer examines how temperature relates to pressure in a reactor. Which classification fits this task?
Bivariate analysis between two variables
Univariate analysis of a single metric
Multivariate analysis with many metrics
Descriptive reporting without relationships
A quality engineer models defect rate using temperature, humidity, and line speed. What type of analysis is this?
Exploratory visualization without modeling
Bivariate analysis of paired measurements
Univariate analysis of one measured variable
Multivariate analysis with more than two variables
Which statement best defines univariate analysis as shown in the diagram?
Analysis of a single variable to summarize features
Analysis of two variables to compare relationships
Modeling multiple variables to predict outcomes
Optimization of parameters across several variables
Which goal is part of univariate analysis when examining one variable?
Design experimental factorial treatments
Identify central tendency measures
Compute cross-variable correlation matrices
Estimate multivariate regression coefficients
A variable recorded on a continuous scale should be analyzed using which univariate type?
Multivariate scale analysis
Binary interaction analysis
Categorical univariate analysis
Continuous univariate analysis
Which pair correctly matches variable type and univariate analysis type?
Continuous variable with categorical univariate analysis
Ordinal variable with multivariate univariate analysis
Nominal variable with interaction univariate analysis
Categorical variable with categorical univariate analysis
When exploring outliers in a single variable, which concept is also typically assessed?
Causal inference among variables
Time series autoregressive order
Parameter estimation across models
Distribution shape and spread
Which category of descriptive statistics is most concerned with the typical value of a variable?
Measures of dispersion
Measures of position
Measures of central tendency
Measures of shape
Which category primarily evaluates how spread out the data is?
Measures of dispersion
Measures of position
Measures of central tendency
Measures of shape
In univariate analysis, which category addresses symmetry and peakedness of a distribution?
Measures of shape
Measures of position
Measures of central tendency
Measures of dispersion
When detecting outliers, which descriptive statistics aspect is directly involved?
Typical value estimation
Outlier identification
Skewness assessment
Variability measurement
A quality engineer wants the relative standing of a measurement within a dataset. Which category should be used?
Measures of dispersion
Measures of central tendency
Measures of shape
Measures of position
If a dataset shows a long right tail, which descriptive statistics focus would help characterize this feature?
Measures of dispersion
Measures of central tendency
Measures of position
Measures of shape
For the dataset of marks [45, 52, 60, 68, 70, 75, 78, 82, 88, 95], what is the arithmetic mean?
72.3 marks
71.3 marks
72.5 marks
71.5 marks
Which definition correctly describes the mean in univariate analysis?
Sum of all observations divided by count
Difference between max and min values
Middle value after sorting observations
Most frequently occurring observation
In PSPP, which menu path computes descriptive statistics like the mean?
Analyze → Regression → Linear
Analyze → Descriptive Statistics → Descriptives
Data → Transform → Compute Variable
Graphs → Chart Builder → Summary
If the mark 95 is removed from the dataset, how does the mean change?
It decreases noticeably
It increases slightly
It remains exactly same
It becomes undefined
When moving the variable ‘marks’ into the PSPP Descriptives dialog, what is the next action to obtain the mean?
Select Frequencies instead
Click OK to run
Click Cancel to exit
Open Variable View first
Which statement best defines the median for a dataset?
Central value computed by squaring values
Average of all values in the dataset
Middle value after ascending order sorting
Most frequent value among all observations
Why is the median preferred when data contain extreme outliers?
Unaffected by outliers in position
Converts data into categorical classes
Maximally influenced by extremes
Requires normal distribution assumption
In PSPP, which sequence computes the median using Frequencies?
Analyze → Descriptive Statistics → Frequencies
File → Open → Data → Output
Transform → Compute → Descriptives
Graphs → Chart Builder → Summary
After opening Frequencies in PSPP, which action ensures the median is reported?
Apply Z-scores transformation
Click Statistics and select Median
Choose Mode in variable view
Enable histogram with normal curve
Which statement correctly defines the mode?
Middle value after sorting
Arithmetic mean of observations
Value least influenced by outliers
Value occurring most frequently
A dataset has maximum value 42 and minimum value 17. What is the range of the dataset?
23
26
25
24
Which statement best defines population variance σ²?
Average of absolute deviations from μ
Sum of deviations divided by N
Square root of mean squared deviations
Mean of squared deviations from μ
For population variance σ² = Σ(xi − μ)² / N, what does μ represent?
Population mean
Sample mean of subgroups
Sample median
Population standard score
Given observations {10, 14, 19, 25} with population mean μ = 17, compute σ² using σ² = Σ(xi − μ)² / N.
24.5
23.5
22.5
25.5
Two datasets share the same mean μ. Dataset A has values tightly clustered around μ, while Dataset B has values spread far from μ. Which statement is true?
Dataset B has larger variance
Dataset A has larger variance
Variance cannot be compared
Both have equal variance
Which statement best defines standard deviation in relation to variance?
It is the variance divided by sample size
It is the square root of the variance
It is the square of the variance value
It is variance minus the population mean
Given population data x₁…xN with mean μ, which expression represents population variance used to compute standard deviation?
Σ(μ−xi) / (N−1)
Σ(xi)²−μ / N
Σ(xi−μ) / N
Σ(xi−μ)² / N
You need PSPP to output variance and standard deviation for a variable. Which sequence of menu actions accomplishes this?
Transform → Compute → Statistics → Options
Analyze → Regression → Linear → OK
File → Open → Descriptive Statistics → OK
Analyze → Descriptive Statistics → Descriptives → Options
A dataset has population variance of 25. What is the population standard deviation?
12.5 units
25 units
4 units
5 units
Which statement best describes quartiles in a dataset?
They divide data into four equal parts
They cluster data into three main groups
They sort data into ten equal bins
They average values across all observations
What proportion of observations is at or below Q1 in a sorted dataset?
About twenty-five percent
About ten percent
About fifty percent
About seventy-five percent
In measures of position, which quartile is the median?
Q3, the upper quartile
Q2, the middle quartile
Q4, the highest quartile
Q1, the lower quartile
Which definition correctly describes a percentile?
Proportion of data above the median value
Difference between Q1 and Q3 values
Average of the lowest half of observations
Value below which a given percentage falls
Which step sequence produces quartiles in PSPP Frequencies?
Analyze → Descriptive Statistics → Frequencies
Graphs → Descriptive Statistics → Quartiles
Analyze → Regression → Frequencies
Transform → Compute → Quartiles
After opening Frequencies in PSPP, which action enables quartile output?
Enable Filters, select By Group
Click Charts, choose Boxplot
Click Statistics, select Quartiles
Open Options, check Means
In a sorted dataset, which statement about Q3 is accurate?
Approximately seventy-five percent lie at or below it
Exactly half of observations equal its value
It is computed as the arithmetic mean
It marks the lowest twenty-five percent boundary
Which statement best describes a distribution with skewness equal to zero?
It is perfectly symmetrical around its center
It has a long right tail and short left tail
It has a long left tail and short right tail
It has multiple modes and heavy tails
A positive skewness value most commonly indicates which distribution feature?
Tail extends toward lower values on the left
Center is shifted left without tail change
Peak is flat with light tails
Tail extends toward higher values on the right
In PSPP, which sequence accesses skewness and kurtosis for a dataset?
File → Settings → Statistics → Moments
Transform → Compute → Moments → Shape
Graph → Explore → Shape → Moments
Analyze → Descriptive Statistics → Descriptives
After opening Descriptives in PSPP, what must you do to include shape measures?
Run Normality test under Nonparametrics
Select Means and Standard Deviations only
Choose Graphs and enable Histograms and QQ plots
Click Options and select Skewness and Kurtosis
An engineer observes a dataset with negative skewness. Which interpretation aligns with this value?
Distribution is bimodal with balanced tails
Distribution is symmetrical with equal tails
Distribution is right skewed with a longer right tail
Distribution is left skewed with a longer left tail
Which statistic best describes the count of observations in each category for a categorical variable?
Frequency for each category
Percentage of the sample
Mode of the distribution
Proportion across all categories
A survey records product color choices: Red 40, Blue 30, Green 30 out of 100. What is the percentage for Blue?
Most common category value
Count divided by unique categories
30 percent of the sample
0.30 of the sample
In categorical univariate analysis, which measure identifies the most frequently occurring category?
Frequency per category
Percentage over total
Mode of categories
Proportion by class
Which statement best defines bivariate analysis in engineering data contexts?
Study of three variables and interactions
Study of relationships between two variables
Study of one variable without comparisons
Study of experimental design with many factors
When conducting bivariate analysis, what aspects of the relationship are commonly evaluated?
Cost, schedule, and risk
Mean, median, and mode
Bias, variance, and error
Strength, direction, and nature
Which pairing of variable types typically uses cross-tabulation and chi-square?
Continuous with binary
Categorical with continuous
Categorical with categorical
Continuous with continuous
You have a categorical factor (material grade) and a continuous response (tensile strength). Which technique is most appropriate?
Cross-tabulation
Correlation or regression
Principal component analysis
t-test or ANOVA
For two continuous variables, such as temperature and viscosity, which technique is commonly selected to model their relationship?
Chi-square test
Regression or correlation
Median test
Cross-tabulation
