WorksheetsPredictive Analytics Exam 1
Total questions: 105
Worksheet time: 53mins
Which of the following broad categories is not a type of analytic technique?
descriptive analytics
predictive analytics
prescriptive analytics
manipulative analytics
Which of the following is not related to data privacy?
data transmission
data ethics
data collection
data usage
A massive volume of both structured and unstructured data that is extremely difficult to manage, process, and analyze is known by which catch phrase?
general data
data mining
big data
data wrangling
According to a report in US Today, 38% of young people between the ages of 18-29 have at least one tattoo. What does the 38% represent?
population
random data
categorical data
sample set
______ is a set of data that are organized and processed in a meaningful and purposeful way.
Statistics
Information
Knowledge
Data
The time in hours spent sleeping per day is what kind of variable?
categorical
continuous numerical
discrete numerical
distraction
An instructor hands out course evaluations where students have a rank of 0 to 5. What is the best way for the data to be measured?
numerical
nominal
ordinal
filtered
Mary asks her friends in Facebook for recommendations for the best restaurants in Chicago. The results are then placed in a table for review. What does the data represent?
quantitative data
time-series data
cross-sectional data
numerical data
The ability to use qualitative reasoning with quantitative tools allows management to make decisions to improve business performance.
True
False
The primary purpose of a(n) _____ is to support decision-making and provide a composite view of the organization
data warehouse
attribute
entity
data mart
Which term represents data items, events, or things stored in a database file?
setting
quantitative
entity
instance
Mary has been tasked with reviewing a large data file. She wants to begin by first inspecting the number of values in each cell, both numeric and non-numeric, for any blank entries. The plan is to first find the blank or missing values for first review. Using Excel, what function(s) should she use to complete this task?
COUNT
COUNTIF
COUNTA
Both COUNT and COUNTA
Which of the following is NOT a process of the data management system?
acquire
distribute
store
summarize
In a data set with 20 variables, if 8% of the values, randomly spread across observations, are missing (blank), what is the probable percent of complete and usable observations?
15.29%
8%
92%
18.87%
In the presence of outliers in a data set, extremely small or large values, it is preferred to use the _____ instead of the _______ to impute missing variables.
mean; median
median; mean
subset; total
average; range
Ann is analyzing a data set that contains two variables, Job Title and 401K. 401K contains the name of the three companies that carry the retirement accounts. It is mandatory to have an account, thus no observation is blank. If 401K was transformed to dummy variables, how many should be created?
1
4
2
3
When too many variables are categorized in an analysis, several potential issues may occur. Which of the following is not one of the issues that may occur?
rarely occurring categories may not be captured accurately
an increase in the number of categories as the data set becomes larger
difficulty in differentiating among observations
model performance suffers
The strategy of removing observations with missing data is called omission.
True
False
If the coefficient correlation is computed to be -0.85, this means the relationship between the two variables are ____.
strong, negative
weak, negative
strong, positive
weak, positive
Survey results provided the skewness coefficient is 0.21672 and the (excess) kurtosis coefficient is -1.15926. These values imply that the return value for the survey is _____ skewed, and the distribution has a _____ tail than a normal distribution.
positively; longer
positive; shorter
negatively; shorter
negative; longer
Survey results provided the skewness coefficient is -0.141974 and the (excess) kurtosis coefficient is 1.15926. These values imply that the return value for the survey is _____ skewed, and the distribution has a _____ tail than a normal distribution.
positively; shorter
positively; longer
negatively; longer
negatively; shorter
In analyzing the S&P 500 and the XYZ Incorporated in a five-year study, the covariance (S&P 500, XYZ Incorporated) is 9,107.30. What kind of linear relationship does the S&P 500 and the XYZ Incorporated have?
neutral linear relationship
positive linear relationship
no linear relationship
In a boxplot, the dashed vertical line in the middle of the box represents which of the following measures of location?
median
percentile
mean
mode
Using R, Bart wants to create a bar chart showing the frequency of the color of cars that pass over the I-270 overpass at the Main Street exit. What function should he use?
abline
table
barplot
view
Which of the following would use a contingency table to visualize the results?
daily stock prices
location and house prices
income and purchase price
gender and restaurant ratings
Which of the following would likely be used to report data related to gender and phone model purchased?
scatterplot
contingency table
line chart
frequency distribution
Which visualization method will Jorge want to use if he wants to understand the relationships between study time, screen time, and academic performance (e.g., GPA) data?
scatterplot with a categorical data
bubble plot
line chart
heat map
Simone is a marketing consultant hired to review the product sales for a new high-end barista machine line. The product line has four variations, selling in four specialty store regions. To clearly show where each variation is selling best and in which regions, she plans to provide a color-scaled chart using percentage by type and location. What is the name of the chart she will be using?
heat map
bubble plot
color map
pivot table
Which visualization method will Jorge want to use if he wants to understand daily sales over the course of the year?
bubble plot
scatterplot with a categorical variable
heat map
line chart
Cross-Industry Standard Process for Data Mining (CRISP-DM) consists of six phases. Of the six, which one represents the phase where data wrangling occurs?
data understanding
deployment
modeling
data preparation
Cross-Industry Standard Process for Data Mining (CRISP-DM) consists of six phases. Which of the following is the first phase?
business understanding
data understanding
data preparation
modeling
The calculated average error measuring the average magnitude of errors in predictive performance measures is called _______.
mean absolute deviation
root mean square error
mean percentage error
mean error
Which chart allows for a visual representation to determine a point where a model's predictions become less useful?
sensitivity measure
cumulative lift chart
decile-wise lift chart
ROC curve
Of the following selections, which is not a descriptor of principal component analysis?
The first principal accounts for most of the variability
Principal components are uncorrelated variables
The first principal account is not suitable for analysis
Principal component variables are weighted linear combinations of the original variables
When using the PCA, all the following are disadvantages except
PCA results are difficult to interpret clearly
PCA significantly increases the dimension of the data
PCA only works with numerical data
components are weighted linear combinations and abstract
develop applications for end users
Data Science
Business Analytics
data analyses for business applications
Data Science
Business Analytics
What has happened?
Descriptive Analytics
Predictive Analytics
Prescriptive Analytics
What could happen in the future?
Descriptive Analytics
Predictive Analytics
Prescriptive Analytics
What should we do?
Descriptive Analytics
Predictive Analytics
Prescriptive Analytics
Data that has been organized, analyzed, and processed in a meaningful and purposeful way become _____.
Information
Knowledge
Figures
Formatted Data
Use a blend of data, contextual information, experience, and intuition to derive _____.
Knowledge
Figures
Charts
Compact Data
collected by recording a characteristic of many subjects at the same point in time; Recording a characteristic of many subjects at the same point in time
Cross-Sectional Data
Time Series Data
Structured Data
Unstructured Data
collected over several time periods focusing on certain groups of people, specific events, or objects; Hourly, daily, weekly, monthly, quarterly, or annual observations
Cross-Sectional Data
Time Series Data
Structured Data
Unstructured Data
reside in a pre-defined, row-column format; spreadsheet or database applications; enter, store, query, and analyze; numerical information that is objective and not open to interpretation
Cross-Sectional Data
Time Series Data
Structured Data
Unstructured Data
do not conform to a pre-defined, row-column format; textual; multimedia content; do not conform to database structures
Cross-Sectional Data
Time Series Data
Structured Data
Unstructured Data
price, income, retail sales
Structured Human
Structured Machine
Unstructured Human
Unstructured Machine
sensors, speed cameras, web server logs
Structured Human
Structured Machine
Unstructured Human
Unstructured Machine
email, text, social media, presentations
Structured Human
Structured Machine
Unstructured Human
Unstructured Machine
satellite images, video data, camera images
Structured Human
Structured Machine
Unstructured Human
Unstructured Machine
Three characteristics of big data
Volume
Velocity
Variety
Value
also called qualitative; labels or names to identify distinguishing characteristics; arithmetic operations on the labels/values are not meaningful; coded into numbers for data processing
Categorical Variables
Numerical Variables
also called quantitative; arithmetic operations are meaningful; represent meaningful numbers
Categorical Variables
Numerical Variables
assumes a countable number of values
discrete
continuous
assumes an uncountable number of values within an interval
discrete
continuous
categorical; least sophisticated; values differ by label or name; ex. marital status
nominal
ordinal
interval
ratio
categorical; reflect labels or name, but can be ranked; cannot interpret the difference between the ranked values; ex. reviews from 1 star to 5 stars
nominal
ordinal
interval
ratio
numerical; categorize and rank, differences are meaningful; zero value is arbitrary and does not reflect absence of characteristic; ratios are not meaningful; ex. temperature
nominal
ordinal
interval
ratio
numerical; most sophisticated; a true zero point, reflects absence of characteristics; ratios are meaningful; ex. profits
nominal
ordinal
interval
ratio
each column starts and ends in the same place in every row
fixed-width format
delimited format
something separates fields, typically a comma
fixed-width format
delimited format
structured data, each piece enclosed in a pair of tags, give information in what the data are
Extensible Markup Language (XML)
HyperText Markup Language (HTML)
JavaScript Object Notation (JSON)
structured data with tags, gives information on how to display the data
Extensible Markup Language (XML)
HyperText Markup Language (HTML)
JavaScript Object Notation (JSON)
transmit human-readable data in compact files, supports wide range of data types, parsing is faster
Extensible Markup Language (XML)
HyperText Markup Language (HTML)
JavaScript Object Notation (JSON)
Best for distribution of a single continuous variable
Histogram
Scatterplot
Box Plot
Best for displaying relationship between two continuous variables
Histogram
Scatterplot
Box Plot
Best for side by side comparisons of subgroups on a single continuous variable
Histogram
Scatterplot
Box Plot
the process of retrieving, cleansing, integrating, transforming, and enriching data to support subsequent analysis
Data Wrangling
Data Management
Data Modeling
a process that an organization uses to acquire, organize, store, manipulate, and distribute data
Data Wrangling
Data Management
Data Modeling
the process of defining the structure of a database
Data Wrangling
Data Management
Data Modeling
person, places, things, events
Entity
Instance
Primary Key
Composite Key
Foreign Key
a single occurrence of an entity; represented as a record in a database
Entity
Instance
Primary Key
Composite Key
Foreign Key
attribute that uniquely identifies each instance of the entity; used to create a data structure called an index for fast data retrieval and searches
Entity
Instance
Primary Key
Composite Key
Foreign Key
key that consists of more than one attribute; used when none of the individual attributes alone can uniquely identify each instance of the entity
Entity
Instance
Primary Key
Composite Primary Key
Foreign Key
a primary key of a related entity
Entity
Instance
Primary Key
Composite Primary Key
Foreign Key
specifies the attributes
SELECT
FROM
WHERE
specifies the tables (more than one)
SELECT
FROM
WHERE
specifies selection criteria and/or conditions
SELECT
FROM
WHERE
a small-scale data warehouse
data mart
data store
data storage
data safe
describes business things such as customer, product, location, and time
Dimension table
Fact table
facts about the business operation, often quantitative format
Dimension table
Fact table
complete-case analysis; exclude observations with missing values; appropriate when the amount of missing value is small or concentrated in a small number of observations
Omission
Imputation
replace missing values with some reasonable values
Omission
Imputation
For numerical values, replace missing values with the ______ value across relevant observations.
mean
median
standard deviation
mode
It is noteworthy that in the presence of outliers it is preferred to use the ____ instead of the ____ to impute missing values.
median; mean
mean; median
mode; standard deviation
standard deviation; mode
The process of extracting portions of a data set that are relevant to the analysis is called _____.
subsetting
organizing
subgrouping
ranking
the data conversion process from one format or structure to another
data transformation
data modeling
data shifting
data changing
the process of transforming numerical variables into the categorical variables by grouping the numerical values into a small number of groups or bins
binning
grouping
transitioning
ranking
refers to how numerical data tend to cluster around some middle or central value
central location
clustering
grouping
percentiles
reflect the typical or central value but they fail to describe other characteristics
measures of central location
measures of dispersion
measures of shape
measures of association
gauge the underlying variability of the variable
measures of central location
measures of dispersion
measures of shape
measures of association
reveal whether the distribution of the variable is symmetric or if the tails are more or less extreme than the normal distribution
measures of central location
measures of dispersion
measures of shape
measures of association
show whether two numeric variables have a linear relationship
measures of central location
measures of dispersion
measures of shape
measures of association
a summary measure that tells us whether the tails of the distribution are more or less extreme than the normal distribution
kurtosis coefficient
leptokurtic
platykurtic
skewness
a distribution that has tails that are more extreme than the normal distribution
kurtosis coefficient
leptokurtic
platykurtic
skewness
a distribution that has shorter tails, or tails that are less extreme, than the normal distribution
kurtosis coefficient
leptokurtic
platykurtic
skewness
situational context, specific objectives, project schedule, deliverables; first phase of CRISP-DM
Business Understanding
Data Understanding
Data Preparation
Modeling
Deployment
collecting raw data, preliminary results, potential hypotheses
Business Understanding
Data Understanding
Data Preparation
Modeling
Deployment
record and variable selection, wrangling, cleaning
Business Understanding
Data Understanding
Data Preparation
Modeling
Deployment
selection and execution of data mining techniques, convert or transform data to formats/types needed for certain analyses, document assumptions, cross-validation
Business Understanding
Data Understanding
Data Preparation
Modeling
Deployment
develop a set of actionable insights and a strategy for deployment/monitoring/feedback
Business Understanding
Data Understanding
Data Preparation
Modeling
Deployment
use for developing predictive models
Supervised Data Mining
Unsupervised Data Mining
effective for data exploration, dimension reduction, and pattern recognition
Supervised Data Mining
Unsupervised Data Mining
predict the class memberships of new cases
Classification Model
Prediction Model
predict the target for a new case
Classification Model
Prediction Model
