wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

arcc

Total questions: 142

Worksheet time: 1hrs 11mins

Name
Class
Date
1.
What statement below best describes why we do data analytics in business?
a)
Analytics improve our understanding of how the business works
b)
We must show a return on the investment we make in data & analytical resources
c)
We need specific insights to make business decisions
d)
We have to calculate & report financial results to owners/shareholders
2.

What should you consider as you approach an analytical problem and in which order? Identify the correct order for the following ideas/steps:
A) Sourcing Data
B) Analysis Outputs
C) Execute Analysis
D) Analysis Methods
E) Define Decision
F) Data Needs

a)
ABCDEF
b)
EBDAFC
c)
EBDFAC
d)
BDFACE
3.
Select a source best describes where the following data might come from: "The average temperature of a turbine bearing over the last 8 hours"
a)
Billing System
b)
Usage Tracking System
c)
Customer Relationship Management System
d)
Machine Data System
e)
Enterprise Resource Planning System
4.
Select a source that best describes where the following data might come from: "The number of developers allocated to a company software project"
a)
Billing System
b)
Usage Tracking System
c)
Customer Relationship Management System
d)
Machine Data System
e)
Enterprise Resource Planning System
5.
Select a source best describes where the following data might come from: "Household water consumption by month"
a)
Billing System
b)
Usage Tracking System
c)
Customer Relationship Management System
d)
Machine Data System
e)
Enterprise Resource Planning System
6.
Select a source best describes where the following data might come from: "The dollar amount of unpaid invoices at the end of a month"
a)
Billing System
b)
Usage Tracking System
c)
Customer Relationship Management System
d)
Machine Data System
e)
Enterprise Resource Planning System
7.
Select a source best describes where the following data might come from: "The average age of customers in Madison, Wisconsin"
a)
Billing System
b)
Usage Tracking System
c)
Customer Relationship Management System
d)
Machine Data System
e)
Enterprise Resource Planning System
8.
Identify the correct order of steps in the Information-Action Value Chain:A)Develop Strategy & Plan B)Deliver the Pitch C)Events & Characteristics in the Real World D)Take Action E)Data Capture by Source Systems F)Data Extraction G)Data Storage H)Analytical Methods I) Summarize & Interpret Results
a)
CEFGHIABD
b)
CEGFHIABD
c)
CEGFHIBAD
9.
Why do we bring data together into a common location? (Select all that apply.)
a)
We can establish relationships among data sources
b)
It's more convenient for extraction to have data in one place
c)
Sometimes we can't access source systems directly
d)
Source data may be unstructured or not formatted for analysis
10.
What type of analytics would you use to determine the best way to route delivery trucks to minimize miles driven or gasoline consumed?
a)
Descriptive
b)
Predictive
c)
Transitive
d)
Cognitive
e)
Prescriptive
11.
What type of file normally stores two-dimensional data with column and row breaks, identified using special characters?
a)
XML File
b)
Log File
c)
Delimited Text File
d)
Excel File
12.
What term best describes data storage that is optimized for handling front-end business operations?
a)
Document Store
b)
Hadoop Distributed File System (HDFS)
c)
Online Transactional Processing (OLTP)
d)
Online Analytical Processing (OLAP)
13.
Suppose you are a software developer looking for an online environment to help you rapidly build and scale applications. Which of the following services would best accommodate your needs?
a)
Platform as a Service (PaaS)
b)
Software as a Service (SaaS)
c)
Development as a Service (DaaS)
d)
Infrastructure as a Service (IaaS)
14.
Which of the following statements about Cloud computing are true? (Select all that apply)
a)
Cloud computing is needed for handling Big Data
b)
Cloud computing speaks to where data is stored or manipulated
c)
Cloud computing is more secure than a company's data center
d)
Cloud computing outsources all of a company's data operations
e)
Cloud computing can allow cheaper and more scalable operations
15.
Suppose your objective is to build a predictive model that can be used to recommend products to customers in real-time based on their navigation on your website. Which of these technologies would be most critical in helping you achieve this objective?
a)
In-Database Analytics
b)
In-Memory Computing
c)
Data Federation
d)
Hadoop Distributed File System (HDFS)
e)
Data Virtualization
16.
Suppose you are a data analyst working on a project to show why sales in a particular region are down relative to other regions. Your job is to figure out what's going on, find a good way to show the data, and produce a report that can be automated to go out weekly to track progress on any actions that are taken. You anticipate that only descriptive analytics will be needed for this project, and you're working from a dataset that has been prepared by your partners in IT. Which of the following classes of tools are you most likely to use directly in this project? (Select all that apply)
a)
Dashboarding
b)
Statistical modeling
c)
Data visualization & exploration
d)
Database systems
e)
Standard reporting
17.
Suppose you're a data analyst and you're traveling to a conference. There's a straightforward but critical ad-hoc analysis you need to accomplish, but you're not certain how much internet connectivity you'll have during your trip. You also haven't decided which of your desktop tools you'll use in the analysis. Which of the following process methodologies would work best for your situation?
a)
Intermediate File Approach
b)
Direct Connection Approach
c)
Downstream Integration Approach
18.
You've just completed an analysis that reveals the importance of a few metrics that business leaders would like to see monthly. You need someone to help you productionalize and automate a monthly report containing those metrics. Who should you talk to?
a)
IT Infrastructure Resource
b)
Application Developer
c)
Data Architect
d)
Database Administrator
e)
BI Developer
19.
You've arranged for an external partner to send you data each day. You need someone to help set up a file transfer process that will allow that partner to securely connect to your company through a firewall. Who should you talk to?
a)
IT Infrastructure Resource
b)
Application Developer
c)
Data Architect
d)
Database Administrator
e)
ETL Developer
20.
To ensure that the results of a data analysis can be placed into context, you need someone who can examine how certain business processes work and help you map them out. Who should you talk to?
a)
IT Infrastructure Resource
b)
Application Developer
c)
Data Architect
d)
Database Administrator
e)
Business Analyst
21.
You've done a descriptive analysis that seems to show a correlation between customer defection and several customer characteristics, but you think that a formal statistical procedure would yield more powerful results that can predict churn. You need someone who knows how to do this. Who should you talk to?
a)
IT Infrastructure Resource
b)
Application Developer
c)
Data Architect
d)
Database Administrator
e)
Modeler
22.
You know that a new product is coming online, and you'd like to understand how measurements around that product will be represented in the database model. Who should you talk to?
a)
IT Infrastructure Resource
b)
Application Developer
c)
Data Architect
d)
Database Administrator
e)
ETL Developer
23.
You're finding that the SQL queries you are writing against your data warehouse are taking a long time to run. You need someone who can help you determine if your queries are written in the best way. Who should you talk to?
a)
IT Infrastructure Resource
b)
Application Developer
c)
Data Architect
d)
Database Administrator
e)
ETL Developer
24.
Your company is functionally organized. The data sources and analytical techniques tend to be pretty similar across functions, and most resources are located in a headquarters building in downtown Chicago. The executive team gets along, but they are very protective of their teams and work product. Which structure would fit best in this scenario?
a)
Allocated Model
b)
Centralized Model
c)
Distributed Model
d)
Coordinated Model
25.
Your company is a multinational organization that operates in a number of distinct industries. Each industry uses its own methods and tends to hire somewhat different types of people into analytical organizations. Which structure would fit best in this scenario?
a)
Allocated Model
b)
Centralized Model
c)
Distributed Model
d)
Coordinated Model
26.
Your company is organized by customer groups, which are mostly distinct but have some limited overlaps. Analyses vary in similarity - some are very similar, but others are quite different. They do, however, use most of the same data sources. Currently, resources are located within each customer group organization, but it's pretty typical for there to be only one or two analysts in each area. Which structure would fit best in this scenario (select all that apply)?
a)
Allocated Model
b)
Centralized Model
c)
Distributed Model
d)
Coordinated Model
27.
What term best describes the process of identifying and standardizing an organization's most critical data?
a)
SOX Compliance
b)
Metadata Management
c)
Master Data Management
d)
Data Governance
e)
Data Stewardship
28.
Who is responsible for making sure that a data domain is correctly represented and used within an organization?
a)
Data Steward
b)
Data Architect
c)
Data Governance Council
d)
SOX Compliance Auditor
e)
ETL Developer
29.
A large drugstore chain wants to use prescription data from its pharmacy to make complementary relevant offers to specific customers via custom coupon books, delivered via direct mail.What is the most limiting standard that might be relevant in this case?
a)
Policy Standards
b)
Legal Standards
c)
Good Judgement
d)
Ethical Standards
30.
A Mobile Phone Company wants to construct and sell 'profiles' of customers based on a combination of internet sites visited and location data. The profiles would provide aggregate information that is not considered CPNI.What are the most limiting standards that might be relevant in this case? (select all that apply)
a)
Policy Standards
b)
Legal Standards
c)
Good Judgement
d)
Ethical Standards
31.
A few years ago your company acquired another company and merged summary financial data into a key database. The data looks complete, but there are some peculiarities we can't explain.What is the dominant issue in this case?
a)
Completeness / Uniqueness
b)
Accuracy / Consistency
c)
Conformance / Validity
d)
Timeliness
e)
Provenance
32.
Your company wants a mobile application that allows certain purchases to be made via the application. However, those transactions use the date/time of the user's device as the timestamp of the transaction that is stored in the purchase database.What is the dominant issue in this case?
a)
Accuracy / Consistency
b)
Completeness / Uniqueness
c)
Conformance / Validity
d)
Timeliness
e)
Provenance
33.
We notice that when we join data from two different tables, we need to be careful to convert the time zone in one table from Central Standard Time (CST) to Universal Coordinated Time (UTC) to match the second table, even though the company standard is UTC.What is the dominant issue in this case?
a)
Completeness / Uniqueness
b)
Accuracy / Consistency
c)
Conformance / Validity
d)
Timeliness
e)
Provenance
34.
On your company's website, a customer can accidentally click a purchase button twice, which results in two purchase records being generated. Luckily, these purchases are filtered by the credit card payment processing system and removed from the company's general ledger. However, those records are not removed from the analytical data warehouse.What is the dominant issue in this case?
a)
Completeness / Uniqueness
b)
Accuracy / Consistency
c)
Conformance / Validity
d)
Timeliness
e)
Provenance
35.
At what stage(s) of Data Exploration would you address missing values in a dataset?
a)
Data transformation
b)
Data clean-up
c)
Data reduction
36.
Which of the following statements regarding data transformation and data reduction is correct?
a)
Data transformations work on individual variables, while data reduction works on a set of variables
b)
Only data transformation would create dummy variables
c)
The goal of data transformations is to create larger datasets while the goal of data reduction is to create smaller datasets
d)
Data transformations are out of style; data reduction is the modern man's tool
37.
What does a data value measure after centering and scaling has been applied?
a)
Accuracy
b)
The number of standard deviations between each data point and the median
c)
The number of standard deviations between each data point and the mean
d)
Slope
38.
Why would one want to center and scale a set of data?
a)
So multiple variables in the dataset are on a common scale
b)
To make all data values positive
c)
To remove duplicates
d)
To make data easier to interpret
39.
Match the Box-Cox Transformation associated with the given value of lambda: When Lambda = 0, transformation is
a)
Logarithmic
b)
Cubed polynomial
c)
Inverse
d)
Square root
40.
Match the Box-Cox Transformation associated with the given value of lambda: When Lambda = 0.5, transformation is
a)
Logarithmic
b)
Cubed polynomial
c)
Inverse
d)
Square root
41.
Match the Box-Cox Transformation associated with the given value of lambda: When Lambda = -1, transformation is
a)
Logarithmic
b)
Cubed polynomial
c)
Inverse
d)
Square root
42.
What is the purpose of applying a Data Reduction?
a)
To generate a larger set of variables
b)
To make all variables positively valued
c)
To use a smaller set of variables to capture most of the information in the original variables
43.
What must be done to variables of a dataset before applying principal component analysis and why?
a)
You must scale the variables so that only outliers are considered as principal components
b)
You must scale the variables so that principal components are not dominated by variables of much larger scale
c)
You must make all variables negative to work with values of the same sign
d)
You must take the square root of all data values to reduce the overall magnitudes of the dataset
44.
Which of the following can be an appropriate way to deal with missing values? (Select all that apply.)
a)
Removing the columns or rows with missing values
b)
Imputing a value with averages of all other records
c)
Imputing a value from "similar" data points
d)
Making "missing" its own category
45.
Your organization asks you to analyze a dataset that shows the number of FreeFly ALTA drones sold in 2016. You noticed that only 2 drones were sold the day after Black Friday, while the average number of drones sold in 2016 is around 100 a day. What is the most probable explanation for this small data value?
a)
It's a missing value that someone filled in with a guess
b)
There was a glitch in the system, and the data value was corrupted
c)
It's a censored value that was inputted incorrectly
d)
It's a censored value; drone inventory probably ran out
46.
What are the risks of replacing a missing value with a guess? (Select all that apply.)
a)
None, the database is capable of correcting input mistakes
b)
Introducing biases
c)
Distorting the dataset
d)
Falsifying results
47.
Why is removing all data records with missing values often not a good way to deal with missing values? (Select all that apply.)
a)
Some modeling tools require a data value for each row/column
b)
A dataset is incomplete if there are missing values
c)
We may end up with too little data to conduct meaningful analysis
d)
The pattern of missing values can have high predictive power
48.
What are the characteristics of an outlier? (Select all that apply.)
a)
It is the data point most proximal to the mean
b)
It is the pivot point for the overall pattern that the data follows
c)
It falls far outside the overall data pattern
d)
It is above or below 3 standard deviations of the mean
49.
A data point is not considered an outlier unless it deviates dramatically on either the x-axis or the y-axis.
a)
True
b)
False
50.
Why do outliers exist? (Select all that apply.)
a)
Data recording errors
b)
Legitimate but odd observations
c)
Entropy of a system
d)
Distortion of time
51.
Which statistical measure is more resistant to outliers?
a)
Mean
b)
Median
c)
Standard deviation
d)
Range
52.
To say a variable is degenerate means which of the following? (Select all that apply.)
a)
The variable is immoral and corrupt
b)
The variable can only take on a single value
c)
When plotted, the variable is modeled with an exponential decay
d)
The variable is a zero variance variable
53.
Which of the following is a remedy to collinearity issues in regression analysis?
a)
Adding more dummy variables
b)
Cutting the dataset in half
c)
Removing zero variance and near zero variance variables
d)
Duplicating the dataset
54.
Which type of target variable are we dealing with in linear regression?
a)
Binary
b)
Categorical
c)
Continuous
d)
Imaginary
55.
We cannot perform linear regression unless both the target variable and predictor variables are continuous.
a)
True
b)
False
56.
What is the validation set used for in predictive modeling?
a)
To fit the models
b)
To evaluate the various models
c)
To increase the size of our training set
d)
To average the training set data
57.
Why can multicollinearity cause problems in multiple regression? (Select all that apply.)
a)
It creates unstable estimates
b)
It creates problems in model interpretation
c)
You cannot make predictions based on regression models with multicollinearity issues
d)
It makes estimating the model impossible
58.
How many transformed variables can we create based on one predictor variable?
a)
An unlimited number
b)
2
c)
None
d)
1
59.
What do you achieve when you apply a log transformation to a variable in your data set?
a)
It makes highly skewed distributions less skewed
b)
It compresses data to make big sets more manageable
c)
It removes negative data values
d)
It removes duplicate values
60.
A soccer team is believed to have an 8 to 2 odds of winning. What is the probability of winning for the team?
a)
0.2
b)
0.25
c)
0.8
d)
2
61.
It is estimated that an appointment with a 10-day lag for a male patient has a predicted probability of 0.1372 of canceling. Compare this with the predicted cancellation probability for a female patient who also has an appointment with a 10-day lag. Assume that the value of the gender variable is 1 for male patients and 0 for females. Also, assume that the estimated coefficient for gender is -0.3572, beta-0 is -1.6515, beta-1 is 0.01699.
a)
A female is less likely to cancel by 2.4%
b)
A female is more likely to cancel by 4.8%
c)
A female is equally likely to cancel as a male
d)
A female is more likely to cancel by 6.9%
62.
The bagging procedure can reduce the variance of a predictive model. Check all true statements about the bagging method: (Check all that apply.)
a)
Helps avoid overfitting of the dataset
b)
Helps group similar data outliers
c)
Has access to multiple training sets
d)
Can be applied to tree models
63.
What do the bagging and random forest methods have in common?
a)
Both methods grow multiple numbers of trees
b)
Both methods operate on only 2 trees
c)
Both methods sample the validation set
d)
Both methods increase the variance of a dataset
64.
What sets the random forest algorithm apart from bagging and boosting algorithms?
a)
It operates on bootstrap sets
b)
It focuses on reducing correlation among models
c)
It involves multiple tree models
d)
It predicts the average variance of a set
65.
True or False: Both linear regression and logistic regression can be viewed as a neural network with no hidden layers.
a)
True
b)
False
66.
Which of the following is true of cluster analysis?
a)
It is a data analysis technique to discover trends in time-series data
b)
It is a data mining tool that is used to create homogeneous groups
c)
It is a data visualization tool in market research
d)
It is a model for customer behavior in the organic and natural products industry
67.
Which of the following settings are appropriate applications of cluster analysis? (Select all that apply.)
a)
A recommender system that seeks to predict the rating or preference that a user would give to an item (e.g., a movie, a book, or a restaurant)
b)
A delivery scheduling system that assigns delivery trucks to customers in the same general geographical area
c)
A cable company seeking to identify the number and type of TV packages to offer (e.g., Basic, Sports, Entertainment, or Premium)
d)
An inventory management system for retail pharmacies that attempts to minimize both the probability of running out of stock and the inventory carrying cost
68.
Which of the following statements is true of principal component analysis (PCA) and cluster analysis?
a)
PCA and cluster analysis are incompatible techniques; only one of them can be applied to the same data
b)
PCA is a data reduction technique and cluster analysis is a dimensionality reduction technique
c)
Cluster analysis is a data reduction technique and PCA is a dimensionality reduction technique
d)
The main goal of cluster analysis is to identify redundant variables, and the main goal of PCA is to create homogeneous groups of observations
69.
Cluster analysis is considered an unsupervised learning technique because it operates on historical observations that are not labeled. That is, it is not known to which group historical observations belong, and therefore it is not known how many groups there are.
a)
True
b)
False
70.
Which of the following is the definition of distance between two clusters in a complete linkage clustering?
a)
The average of distances between all pairs of objects, where each pair is made up of one object of each group
b)
The distance between the most distant pair of objects, one from each group
c)
The sum of squares of the distance between clusters
d)
The distance between the value of the shortest link between the clusters
71.
Which of the following is true of hierarchical clustering?
a)
All clusters must have the same number of objects
b)
No single cluster can have all objects
c)
Each step of the procedure consists of merging the two closest clusters
d)
All clusters must have more than one object in them
72.
Which of the following is true of clustering methods?
a)
The k-means method is an exact procedure that finds the optimal (i.e., the best) solution
b)
The best clustering approach when dealing with very large data sets is to solve the optimization problem using Excel's Solver
c)
The k-means method and hierarchical clustering always arrive at the same solution; that is, they always produce the same set of clusters
d)
Finding the best set of clusters is complicated because the number of ways of partitioning the observations into k groups is very large, and this is why approximation methods such as k-means and hierarchical clustering are used
73.
Which of the following best defines Monte Carlo simulation?
a)
It's a tool for building statistical models that characterize relationships among a dependent variable and one or more independent variables.
b)
It's a collection of techniques that seeks to group or segment a collection of objects into subsets.
c)
It's the process of selecting values of decision variables that minimizes or maximizes some quantity of interest.
d)
It's the process of generating random values for uncertain inputs in a model and computing the output variables of interest.
74.
If chance or uncertainty is present in a system, then there is an element of ______ in the decision-making problem.
a)
danger
b)
security
c)
risk
d)
difficulty
75.
Which of the following are weaknesses of manual what-if analysis? (Select all that apply.)
a)
Biased sample values of performance measures
b)
Hard to do many what-if scenarios
c)
Does not provide distribution information
76.
Which of the following is a parameter of the Poisson distribution?
a)
Maximum value
b)
Mean
c)
Minimum value
d)
Most likely value
77.
In the Analytic Solver Platform, "Psi" functions are used to add uncertainty to a spreadsheet model.
a)
True
b)
False
78.
Why would a manager be interested in analyzing risk?
a)
To determine a most likely outcome
b)
To determine a range of outcomes
c)
To determine a distribution of outcomes
d)
To determine a confidence interval on most likely outcomes
79.
The PsiOutput function of the Analytic Solver Platform is used to collect simulation data to create an empirical distribution of an output variable.
a)
True
b)
False
80.
Historical data is used in simulation to:
a)
Perform a worst-case analysis
b)
Optimize the outcomes
c)
Estimate a probability distribution function for critical inputs to the model
d)
Simplify the model
81.
Distribution fitting is the process of gathering historical data.
a)
True
b)
False
82.
Adding a correlation matrix to a simulation model is necessary when:
a)
The uncertain input variables in the model are independent
b)
The model is deterministic (i.e., it does not have any uncertain inputs)
c)
Two or more of the uncertain input variables in the model are not independent
d)
An output variable is related to an uncertain input variable
83.
Which of the following statements is false?
a)
Correlation is a measure of the strength of the relationship between two variables
b)
Correlation values are always positive
c)
The correlation between two variables can be positive or negative
d)
The correlation between two independent variables is zero
84.
The Analytic Solver Platform ________ allows you to determine the influence that each uncertain input variable has on an output variable based on the correlation between the input and the output variable.
a)
Trend chart
b)
Overlay chart
c)
Box-whisker chart
d)
Sensitivity chart
85.
The Analytic Solver Platform ________ allows you to superimpose the frequency distributions of selected output variables in order to compare them.
a)
Trend chart
b)
Overlay chart
c)
Box-whisker chart
d)
Sensitivity chart
86.
The Flaw of Averages typically results when a single number, the average value, is used in a spreadsheet model to represent an uncertain future quantity.
a)
True
b)
False
87.
The average value for an output cell in a deterministic spreadsheet model that uses average values for uncertain input cells is always the same as the average value for the same output cell obtained with a Monte Carlo simulation.
a)
True
b)
False
88.
Which of the following statements are true? (Select all that apply.)
a)
Optimization has been defined as the process of selecting the values of decision variables that minimize or maximize some quantity of interest.
b)
Optimization started in the area of operations management but it is now used in all areas of business.
c)
Optimization models are prescriptive because their outcome is a recommendation of what to do.
89.
In an optimization model, decision variables are:
a)
The unknowns for which the optimization process will find the best values.
b)
The functions to be maximized or minimized.
c)
The restrictions or limitations that are either related to technical and practical considerations or they are imposed by managerial policies.
d)
The parameter values provided by the analyst.
90.
In a linear programming model, both the objective function and the constraints are formulated as linear functions of the decision variables.
a)
True
b)
False
91.
What is the goal in the optimization of the transportation problem?
a)
Find the values of the decision variables that use all supplier capacities.
b)
Find the decision variable values (i.e., the shipment quantities) that result in the best objective function (i.e., lowest total cost) and satisfy all constraints.
c)
Find the values of the decision variables that satisfy all the demand constraints.
d)
None of these.
92.
What does the Excel =SUMPRODUCT(A1:A3, B1:B3) function do?
a)
Sums each range and multiplies the sums. That is, (A1+A2+A3) * (B1+B2+B3).
b)
Sums each pair of cells and multiplies each sum. That is, (A1+B1)(A2+B2)(A3+B3).
c)
Multiplies each range and sums the products. That is, (A1A2A3)+(B1B2B3).
d)
Multiplies each pair of cells and sums the products. That is, (A1B1)+(A2B2)+(A3B3).
93.
Which of the following statements are true about a Sensitivity Report?
a)
It provides very useful information for pricing decisions, the value of resources, and the robustness of the optimal solution.
b)
It's not able to provide answers to what-if questions that involve multiple changes in the model, such as simultaneously changing the coefficient of a decision variable and the right-hand side of a constraint.
c)
It provides information about decision variables (reduced costs) and constraints (shadow prices).
94.
If the shadow price for a resource constraint is 0, the allowable increase is 200 units, and 150 units of the resource are added, what happens to the objective function value?
a)
It increases by 150
b)
It increases by more than 0 but less than 150
c)
No change d. It increases but by an unknown amount
95.
Which of the following approaches provided by the Analytic Solver Platform can automatically run multiple optimizations while varying model parameters (e.g., the right-hand side of a constraint) within a pre-specified range?
a)
Breakdown analysis
b)
Parameter analysis
c)
Uncertainty analysis
d)
Sensitivity analysis
96.
A bar chart is an effective way of visualizing the use of a resource in an optimal solution, where colors represent how the resource is used and the height represents how much of the resource is used.
a)
True
b)
False
97.
Which of the following is not a benefit of using binary variables?
a)
Models are easy to solve (i.e., the solvers can find optimal solutions faster) because the variables can only be zero or one.
b)
Binary variables are useful in selection problems.
c)
Binary variables can be used to model yes/no decisions.
d)
Binary variables can enforce logical conditions.
98.
If optimization model has 5 binary decision variables. How many possible integer solutions are there to this problem?
a)
5
b)
10
c)
25
d)
32
99.
A company wants to select no more than 2 projects from a set of 4 possible projects. Which of the following constraints ensures that no more than 2 will be selected, assuming that the P variables are binary and represent whether a project is selected (value of 1) or not (value of 0)?
a)
P1+P2+P3+P4=2
b)
P1+P2+P3+P4≤2
c)
P1+P2+P3+P4≥2
d)
P1+P2+P3+P4≥0
100.
A company must invest in project 1 in order to invest in project 2. P1P1 is a binary variable representing whether project 1 is chosen (value of 1) or not (value of 0). P2P2 has the same interpretation for project 2. Which of the following constraints ensures that if project 2 is chosen, then project 1 must also be chosen?
a)
P1+P2=0
b)
P1+P2=1
c)
P1−P2≥0
d)
P1−P2≤0
101.
Which of the following statements is not true about metaheuristic optimization?
a)
Metaheuristics provide great modeling flexibility.
b)
Metaheuristics can solve optimization models with nonlinear and/or non-smooth functions.
c)
The metaheuristic solver in the Analytic Solver Platform is called the Evolutionary Engine.
d)
Metaheuristics are exact procedures that guarantee finding an optimal solution.
102.
In market basket analysis, the Lift Ratio tells us how much more likely it is for item Y to be purchased given that item X has been purchased.
a)
True
b)
False
103.
A chance constraint is a special type of constraint that it is satisfied only in a fraction of the trials in a simulation.
a)
True
b)
False
104.
An optimization model includes a chance constraint to satisfy demand for a particular product. The demand is uncertain and is modeled with an integer uniform distribution with parameter values of 0 and 4. That is, the probability that the demand is 0, 1, 2, 3, or 4 is exactly the same. A decision is made to order 2 units of the product from a supplier to satisfy the uncertain demand. What is the value at risk (VaR) for the demand constraint?
a)
30%
b)
40%
c)
50%
d)
60%
105.
Which of the following statements correctly describes what happens in the last stage of an Information-Action Value Chain? Recall: there are 3 stages.
a)
Analyze data set with statistical methods, clustering algorithms, or advanced association to help make better decisions going forward.
b)
Summarize + interpret analytic results ⇒ model results with visual representations ⇒ develop action plan and alternatives ⇒ deliver the pitch.
c)
Apply prescriptive analytics to optimize business rules or financial models to figure out what choices should be made to achieve a specific outcome.
d)
Identify an object or phenomena ⇒ Access data on the object or phenomena ⇒ Logically organize the data.
106.
Recall the concept of a data warehouse as an alternative to directly accessing data from a source system. Select all characteristics that pertain to the concept of a data warehouse:
a)
Data is pulled from multiple data warehouses to create a source system.
b)
Data is organized in the data warehouse before it is used in an analysis.
c)
Data is organized in the source system BEFORE being pulled into a data warehouse.
d)
A data warehouse serves as a common location for data pulled from one or more source systems.
e)
You can selectively choose what type of data to extract from a data warehouse using a query language such as SQL.
107.
What properties are a desirable outcome of segmentation?
a)
"Homogeneity between" groups and "heterogeneity between" groups.
b)
"Homogeneity within" groups and "heterogeneity between" groups.
c)
"Homogeneity between" groups and "heterogeneity within" groups.
d)
"Homogeneity within" groups and "heterogeneity within" groups.
108.
What does a customer lifetime value represent?
a)
The sum of all revenues and costs of a customer over time.
b)
The difference of all revenues and costs of a customer over time.
c)
The frequency of contact a business has with a customer over their lifetime.
d)
The number of years an individual is considered a customer to a particular business.
109.
Your company just broke a sales record. You would like to show the new sales record on a big screen at the company entrance.
a)
Text
b)
Table
c)
Graph
110.
In an internal meeting, you would like to show the monthly sales amount of a few key products in the last ten years.
a)
Text
b)
Table
c)
Graph
111.
In a meeting with a sales team, you would like to show the quarterly units sold, sales amount, and market share for several products in the last year.
a)
Text
b)
Table
c)
Graph
112.
Which of the following are benefits of a table? (Select all that apply.)
a)
Can display very complex relationships
b)
Can display precise values
c)
Speeds up lookup of individual values
d)
Relates different units of measurement
e)
Can present patterns of the data
113.
Select all methods that can be used to represent single numerical values:
a)
Points
b)
Lines
c)
Bars
d)
Shapes
e)
Color Intensity
114.
Which of the following graphs CANNOT be used to show the distribution of data? (Select all that apply.)
a)
Strip plot
b)
Line plot
c)
Histogram
d)
Density curve
e)
Scatter plot
115.
Which of the following decreases the data-ink ratio? (Select all that apply.)
a)
A large amount of points in a scatter plot
b)
Large range intervals in histograms
c)
3D effects
d)
Background images
116.
Visualization experts recommend against using points without trend lines to show time series data. Which of the following statements is the best explanation for this recommendation?
a)
Points are generally not a good way to encode numerical values.
b)
Points cannot show precise values of numerical data.
c)
Points have small visual weight and do not help show the sequential nature of the data.
d)
None of the above.
117.
Which of the following statements regarding small multiple design are correct? (Select all that apply.)
a)
In general, the same chart type should be used in a small multiple design.
b)
In general, small multiple design is preferred whenever it is feasible.
c)
In general, the same axis scale and color scheme should be used in a small multiple design.
d)
In a small multiple design, the graphs should be arranged following some natural order whenever possible.
118.
Which of the following statements best explains why stacked area charts should be avoided?
a)
Stacked area charts are not aesthetically pleasing.
b)
Stacked area charts are difficult to construct with commonly available software tools.
c)
Stacked area charts may mask the trend of data series except for the one in the bottom.
d)
Stacked area charts cannot represent numerical values accurately.
119.
In addition to analyzing data well, a data analyst may also need to:
a)
Ensure isolation between the data and the customer
b)
Take account of bad or questionable data
c)
Organize and process data as efficiently as possible
d)
Pitch their ideas to one or more decision makers
120.
Who or what is the best option to appeal your ideas to?
a)
The process
b)
The people
c)
The place
d)
The period
121.
Telling compelling stories about the data analytic results requires: (Select all that apply.)
a)
The story to be simple
b)
The story to be short
c)
The story to be true
d)
The story to be informative
122.
What is the relevance of knowing your audience for selling your story?
a)
To give input to audience opinion
b)
To present results in a way that resonates with the audience
c)
To present data that is familiar to the audience
d)
To help present abstract concepts in an easy manner
123.
What is an important principle to remember when preparing or giving a presentation?
a)
Slides are the only effective way of presenting ideas
b)
Occasionally use specific materials to keep the presentation concise
c)
The presentation is about the problem you're trying to solve - not about you
d)
Slides should contain as much information as possible
124.
What is the pyramid principle?
a)
Give ideas a sufficient proof-of-concept using deductive analysis
b)
Writing and thinking can be structured to nest supporting ideas under one common point
c)
Always start from underdeveloped ideas and work towards a marketable one
d)
Ideas can be broken into steps and then recombined together
125.
In preparing presentation materials, the materials should be:
a)
Well-rounded slides and graphs
b)
At least 3-5 items
c)
Colorful for visual aid
d)
Correct and good quality
126.
Although delivering the pitch requires conciseness and clarity, it must:
a)
Accommodate all possible audiences
b)
Be delivered to the audience as quickly as possible
c)
Fit into the time allotted
d)
Be delivered via PowerPoint slides
127.
Numerical patterns and trends are best understood when:
a)
Presented in some context
b)
Plotted as concisely as possible
c)
Compared and contrasted against current ideas
d)
Organized and analyzed very thoroughly
128.
A good rule of thumb when presenting with slides is:
a)
To use note cards sparingly
b)
To allot about 5 minutes per slide
c)
Construct a familiar scenario to the audience
d)
To present complex ideas first
129.
True or False: We identify correlation by specifically looking for a linear relationship when two data sets are plotted against each other.
a)
True
b)
False
130.
Consider the following two sets of measures: Temperature in degrees Celsius for each day in November, 2011 in Boulder, Colorado Hot chocolate sales in dollars for each day in November, 2011 in Boulder, Colorado. Say you calculate the Pearson correlation coefficient for the two sets of measures. If your r-value is equal to 0, what can you say about the relationship between sales of hot chocolate versus temperature based off of these two data sets? (Select all that apply.)
a)
The data sets are not correlated at all
b)
The data sets are perfectly negatively correlated
c)
The data sets are perfectly positively correlated
d)
There is no linear relationship between temperature and sales of hot chocolate
131.
Which of the following statements about causality and correlation are true? (Select all that apply.)
a)
Causality is always implied if two data sets are found to be correlated
b)
If correlation is present between two data sets, causality is just one possible explanation for the identified relationship
c)
There can never be both correlation AND causality present between two data sets, they are mutually exclusive concepts
d)
Two sets of measurements can have a causal relationship AND also have no correlation present
132.
Pretend a study is performed on the following two sets of data: Number of engineering degrees awarded in 2012 Number of kittens adopted in 2012. A very high degree of correlation is found between the two almost unrelated data sets. What is the most likely explanation for this high degree of correlation?
a)
A third factor is likely causing the trends in both data sets
b)
No real relationship exists between the two data sets, it is coincidence
c)
There is no arguing with the data, the two types of events are related in the real world
d)
The event of receiving an engineering degree is the cause of the second event, the purchasing of a kitten
133.
Define cognitive biases:
a)
The mental action of acquiring knowledge and understanding through thought, experience, and the senses
b)
A mode of altering information to deviate from reality
c)
A concentration on or interest in one particular area or subject
d)
Using thought or rational judgment
134.
It is a data analyst's responsibility to:
a)
Draw different conclusions from information based on how it's presented
b)
Support the agenda of the sponsor of a study
c)
Take into consideration external pressure to show favorable results
d)
Remain objective
135.
The lie of average distorts data by:
a)
Using summary statistics to conceal data distribution
b)
Using the average of a data set to hide variance
c)
Using correlation to prove causation
d)
Using summary statistics to find correlated data sets
136.
What is the basic idea behind chart myopia?
a)
Zooming in or out on data visualization results to make insignificant things look significant or vice versa
b)
Using only summary statistics can misrepresent the underlying nuances of our data
c)
Leaving out the 'out of how many' value when counting the frequency of a specific event
d)
Showing data in a way that puts it into perspective and shows it in the right context
137.
Select all characteristics of a properly conducted controlled market experiment:
a)
Apply a predetermined treatment to the control group
b)
Only control group members are selected randomly
c)
Both control group and treatment group participants are selected at random
d)
Differences in behavior between groups should be attributable to the treatment
e)
Differences between groups are measurable and quantifiable
138.
What are the advantages of the "Baseball Card" approach to showing options in slides? (Select all that apply.)
a)
Allows different types of information about an option to be displayed in one place
b)
Focuses attention on the speaker by only showing the most relevant bullet points
c)
Facilitates comparisons of options to each other
d)
Is optimal when the analysis involves mostly math and figures
139.
Why might we consider using a segmentation schema provided by a third party instead of building our own? (Select the best answer.)
a)
Third party schemas are usually better because they incorporate information we don't have
b)
They can help us to target customers in the market for which we may not have information
c)
They are the experts and are better at segmentation analytics
d)
It's a better option when we are planning efforts to retain our current customers
140.
Which of the following describes a fully factorialized experiment?
a)
An experiment designed to provide maximum insight with the fewest number of test groups
b)
An experiment that has a separate control group for each test group
c)
An experiment that uses a before and after comparison for each test group
d)
An experiment where all combinations of all factors are tested
141.
In the Customer Acquisition Strategy case study, why did we evaluate so many options? (Select all that apply.)
a)
Different constituencies had different opinions on what the strategy should be
b)
There were significant tradeoffs among different approaches
c)
Because it's always better to have as many options as possible
d)
The preferred number of strategic options is the magic number seven plus or minus two
142.
To make use of a Customer Lifetime Value calculation, it's necessary to have all revenues and costs for each customer readily available.
a)
True
b)
False