Font size
Worksheetsархитектура 1
Total questions: 69
Worksheet time: 35mins
What statement below best describes why we do data analytics in business?
Analytics improve our understanding of how the business works
We must show a return on the investment we make in data & analytical resources
We need specific insights to make business decisions
We have to calculate & report financial results to owners/shareholders
What should you consider as you approach an analytical problem and in which order? Identify the correct order for the following ideas/steps:
А Sourcing Data,
B Analysis Outputs,
C Execute Analysis,
D Analysis Methods,
E Define Decision,
F Data Needs
ABCDEF
EBDAFC
EBDFAC
BDFACE
Select a source best describes where the following data might come from: "The average temperature of a turbine bearing over the last 8 hours"
Machine Data System
Billing System
Enterprise Resource Planning System
Usage Tracking System
Customer Relationship Management System
Select a source that best describes where the following data might come from: "The number of developers allocated to a company software project"
Customer Relationship Management System
Usage Tracking System
Billing System
Machine Data System
Enterprise Resource Planning System
Select a source best describes where the following data might come from: "Household water consumption by month"
Billing System
Usage Tracking System
Enterprise Resource Planning System
Machine Data System
Customer Relationship Management System
Select a source best describes where the following data might come from: "The dollar amount of unpaid invoices at the end of a month"
Billing System
Usage Tracking System
Machine Data System
Enterprise Resource Planning System
Customer Relationship Management System
Select a source best describes where the following data might come from: "The average age of customers in Madison, Wisconsin"
Machine Data System
Usage Tracking System
Billing System
Enterprise Resource Planning System
Customer Relationship Management System
Identify the correct order of steps in the Information-Action Value Chain:
A Develop Strategy & Plan
B Deliver the Pitch
C Events & Characteristics in the Real World
D Take Action
E Data Capture by Source Systems
F Data Extraction
G Data Storage
H Analytical Methods
I Summarize & Interpret Results
CEFGHIABD
CEGFHIABDㅤ
CEGFHIBAD
Why do we bring data together into a common location? (Select all that apply.)
We can establish relationships among data sources
It's more convenient for extraction to have data in one place
Sometimes we can't access source systems directly
Source data may be unstructured or not formatted for analysis
What type of analytics would you use to determine the best way to route delivery trucks to minimize miles driven or gasoline consumed?
Descriptive
Predictive
Transitive
Cognitive
Prescriptive
What type of file normally stores two-dimensional data with column and row breaks, identified using special characters?
XML File
Log File
Delimited Text File
Excel File
What term best describes data storage that is optimized for handling front-end business operations?
Document Store
Hadoop Distributed File System (HDFS)
Online Transactional Processing (OLTP)
Online Analytical Processing (OLAP)
Suppose you are a software developer looking for an online environment to help you rapidly build and scale applications. Which of the following services would best accommodate your needs?
Platform as a Service (PaaS)
Software as a Service (SaaS)
Development as a Service (DaaS)
Infrastructure as a Service (IaaS)
Which of the following statements about Cloud computing are true? (Select all that apply)
Cloud computing is needed for handling Big Data
Cloud computing speaks to where data is stored or manipulated
Cloud computing is more secure than a company's data center
Cloud computing outsources all of a company's data operations
Cloud computing can allow cheaper and more scalable operations
Suppose your objective is to build a predictive model that can be used to recommend products to customers in real-time based on their navigation on your website. Which of these technologies would be most critical in helping you achieve this objective?
In-Database Analytics
In-Memory Computing
Data Federation
Hadoop Distributed File System (HDFS)
Data Virtualization
Suppose you are a data analyst working on a project to show why sales in a particular region are down relative to other regions. Your job is to figure out what's going on, find a good way to show the data, and produce a report that can be automated to go out weekly to track progress on any actions that are taken. You anticipate that only descriptive analytics will be needed for this project, and you're working from a dataset that has been prepared by your partners in IT. Which of the following classes of tools are you most likely to use directly in this project? (Select all that apply)
Dashboarding
Statistical modeling
Data visualization & exploration
Database systems
Standard reporting
Suppose you're a data analyst and you're traveling to a conference. There's a straightforward but critical ad-hoc analysis you need to accomplish, but you're not certain how much internet connectivity you'll have during your trip. You also haven't decided which of your desktop tools you'll use in the analysis. Which of the following process methodologies would work best for your situation?
Intermediate File Approach
Direct Connection Approach
Downstream Integration Approach
You've just completed an analysis that reveals the importance of a few metrics that business leaders would like to see monthly. You need someone to help you productionalize and automate a monthly report containing those metrics. Who should you talk to?
Data Architect
IT Infrastructure Resource
BI Developer
Database Administrator
Application Developer
You've arranged for an external partner to send you data each day. You need someone to help set up a file transfer process that will allow that partner to securely connect to your company through a firewall. Who should you talk to?
IT Infrastructure Resource
Application Developer
Database Administrator
ETL Developer
Data Architect
To ensure that the results of a data analysis can be placed into context, you need someone who can examine how certain business processes work and help you map them out. Who should you talk to?
Business Analyst
IT Infrastructure Resource
Database Administrator
Application Developer
Data Architect
You've done a descriptive analysis that seems to show a correlation between customer defection and several customer characteristics, but you think that a formal statistical procedure would yield more powerful results that can predict churn. You need someone who knows how to do this. Who should you talk to?
Modeler
Database Administrator
IT Infrastructure Resource
Data Architect
Application Developer
You know that a new product is coming online, and you'd like to understand how measurements around that product will be represented in the database model. Who should you talk to?
ETL Developer
Database Administrator
Data Architect
Application Developer
IT Infrastructure Resource
You're finding that the SQL queries you are writing against your data warehouse are taking a long time to run. You need someone who can help you determine if your queries are written in the best way. Who should you talk to?
IT Infrastructure Resource
Database Administrator
Data Architect
Application Developer
ETL Developer
Your company is functionally organized. The data sources and analytical techniques tend to be pretty similar across functions, and most resources are located in a headquarters building in downtown Chicago. The executive team gets along, but they are very protective of their teams and work product. Which structure would fit best in this scenario?
Allocated Model
Centralized Model
Distributed Model
Coordinated Model
Your company is a multinational organization that operates in a number of distinct industries. Each industry uses its own methods and tends to hire somewhat different types of people into analytical organizations. Which structure would fit best in this scenario?
Allocated Model
Centralized Model
Distributed Model
Coordinated Model
Your company is organized by customer groups, which are mostly distinct but have some limited overlaps. Analyses vary in similarity - some are very similar, but others are quite different. They do, however, use most of the same data sources. Currently, resources are located within each customer group organization, but it's pretty typical for there to be only one or two analysts in each area. Which structure would fit best in this scenario (select all that apply)?
Allocated Model
Centralized Model
Distributed Model
Coordinated Model
What term best describes the process of identifying and standardizing an organization's most critical data?
SOX Compliance
Metadata Management
Master Data Management
Data Governance
Data Stewardship
Who is responsible for making sure that a data domain is correctly represented and used within an organization?
Data Steward
Data Architect
Data Governance Council
SOX Compliance Auditor
ETL Developer
A large drugstore chain wants to use prescription data from its pharmacy to make complementary relevant offers to specific customers via custom coupon books, delivered via direct mail. What is the most limiting standard that might be relevant in this case?
Policy Standards
Legal Standards
Good Judgement
Ethical Standards
A Mobile Phone Company wants to construct and sell 'profiles' of customers based on a combination of internet sites visited and location data. The profiles would provide aggregate information that is not considered CPNI. What are the most limiting standards that might be relevant in this case? (select all that apply)
Policy Standards
Legal Standards
Good Judgement
Ethical Standards
A few years ago your company acquired another company and merged summary financial data into a key database. The data looks complete, but there are some peculiarities we can't explain. What is the dominant issue in this case?
Completeness / Uniqueness
Accuracy / Consistency
Conformance / Validity
Timeliness
Provenance
Your company wants a mobile application that allows certain purchases to be made via the application. However, those transactions use the date/time of the user's device as the timestamp of the transaction that is stored in the purchase database. What is the dominant issue in this case?
Completeness / Uniqueness
Accuracy / Consistency
Conformance / Validity
Timeliness
Provenance
We notice that when we join data from two different tables, we need to be careful to convert the time zone in one table from Central Standard Time (CST) to Universal Coordinated Time (UTC) to match the second table, even though the company standard is UTC. What is the dominant issue in this case?
Completeness / Uniqueness
Accuracy / Consistency
Conformance / Validity
Timeliness
Provenance
On your company's website, a customer can accidentally click a purchase button twice, which results in two purchase records being generated. Luckily, these purchases are filtered by the credit card payment processing system and removed from the company's general ledger. However, those records are not removed from the analytical data warehouse. What is the dominant issue in this case?
Completeness / Uniqueness
Accuracy / Consistency
Conformance / Validity
Timeliness
Provenance
At what stage(s) of Data Exploration would you address missing values in a dataset?
Data transformation
Data clean-up
Data reduction
Which of the following statements regarding data transformation and data reduction is correct?
Data transformations work on individual variables, while data reduction works on a set of variables
Only data transformation would create dummy variables
The goal of data transformations is to create larger datasets while the goal of data reduction is to create smaller datasets
Data transformations are out of style; data reduction is the modern man's tool
What does a data value measure after centering and scaling has been applied?
Accuracy
The number of standard deviations between each data point and the median
The number of standard deviations between each data point and the mean
Slope
Why would one want to center and scale a set of data?
So multiple variables in the dataset are on a common scale
To make all data values positive
To remove duplicates
To make data easier to interpret
Match the Box-Cox Transformation associated with the given value of lambda: When Lambda = 0, transformation is
Logarithmic
Cubed polynomial
Inverse
Square root
Match the Box-Cox Transformation associated with the given value of lambda: When Lambda = 0.5, transformation is
Logarithmic
Cubed polynomial
Inverse
Square root
Match the Box-Cox Transformation associated with the given value of lambda: When Lambda = -1, transformation is
Logarithmic
Cubed polynomial
Inverse
Square root
What is the purpose of applying a Data Reduction?
To generate a larger set of variables
To make all variables positively valued
To use a smaller set of variables to capture most of the information in the original variables
What must be done to variables of a dataset before applying principal component analysis and why?
You must scale the variables so that only outliers are considered as principal components
You must scale the variables so that principal components are not dominated by variables of much larger scale
You must make all variables negative to work with values of the same sign
You must take the square root of all data values to reduce the overall magnitudes of the dataset
Which of the following can be an appropriate way to deal with missing values? (Select all that apply.)
Removing the columns or rows with missing values
Imputing a value with averages of all other records
Imputing a value from 'similar' data points
Making 'missing' its own category
Your organization asks you to analyze a dataset that shows the number of FreeFly ALTA drones sold in 2016. You noticed that only 2 drones were sold the day after Black Friday, while the average number of drones sold in 2016 is around 100 a day. What is the most probable explanation for this small data value?
It's a missing value that someone filled in with a guess
There was a glitch in the system, and the data value was corrupted
It's a censored value that was inputted incorrectly
It's a censored value; drone inventory probably ran out
What are the risks of replacing a missing value with a guess? (Select all that apply.)
None, the database is capable of correcting input mistakes
Introducing biases
Distorting the dataset
Falsifying results
Why is removing all data records with missing values often not a good way to deal with missing values? (Select all that apply.)
Some modeling tools require a data value for each row/column
A dataset is incomplete if there are missing values
We may end up with too little data to conduct meaningful analysis
The pattern of missing values can have high predictive power
What are the characteristics of an outlier? (Select all that apply.)
It is the data point most proximal to the mean
It is the pivot point for the overall pattern that the data follows
It falls far outside the overall data pattern
It is above or below 3 standard deviations of the mean
A data point is not considered an outlier unless it deviates dramatically on either the x-axis or the y-axis.
True
False
Why do outliers exist? (Select all that apply.)
Data recording errors
Legitimate but odd observations
Entropy of a system
Distortion of time
Which statistical measure is more resistant to outliers?
Mean
Median
Standard deviation
Range
To say a variable is degenerate means which of the following? (Select all that apply.)
The variable is immoral and corrupt
The variable can only take on a single value
When plotted, the variable is modeled with an exponential decay
The variable is a zero variance variable
Which of the following is a remedy to collinearity issues in regression analysis?
Adding more dummy variables
Cutting the dataset in half
Removing zero variance and near zero variance variables
Duplicating the dataset
Which type of target variable are we dealing with in linear regression?
Binary
Categorical
Continuous
Imaginary
We cannot perform linear regression unless both the target variable and predictor variables are continuous.
True
False
What is the validation set used for in predictive modeling?
To fit the models
To evaluate the various models
To increase the size of our training set
To average the training set data
Why can multicollinearity cause problems in multiple regression? (Select all that apply.)
It creates unstable estimates
It creates problems in model interpretation
You cannot make predictions based on regression models with multicollinearity issues
It makes estimating the model impossible
How many transformed variables can we create based on one predictor variable?
An unlimited number
2
None
1
What do you achieve when you apply a log transformation to a variable in your data set?
It makes highly skewed distributions less skewed
It compresses data to make big sets more manageable
It removes negative data values
It removes duplicate values
A soccer team is believed to have an 8 to 2 odds of winning. What is the probability of winning for the team?
0.2
0.25
0.8
2
It is estimated that an appointment with a 10-day lag for a male patient has a predicted probability of 0.1372 of canceling. Compare this with the predicted cancellation probability for a female patient who also has an appointment with a 10-day lag. Assume that the value of the gender variable is 1 for male patients and 0 for females. Also, assume that the estimated coefficient for gender is -0.3572, beta-0 is -1.6515, beta-1 is 0.01699.
A female is less likely to cancel by 2.4%
A female is more likely to cancel by 4.8%
A female is equally likely to cancel as a male
A female is more likely to cancel by 6.9%
The bagging procedure can reduce the variance of a predictive model. Check all true statements about the bagging method: (Check all that apply.)
Helps avoid overfitting of the dataset
Helps group similar data outliers
Has access to multiple training sets
Can be applied to tree models
What do the bagging and random forest methods have in common?
Both methods grow multiple numbers of trees
Both methods operate on only 2 trees
Both methods sample the validation set
Both methods increase the variance of a dataset
What sets the random forest algorithm apart from bagging and boosting algorithms?
It operates on bootstrap sets
It focuses on reducing correlation among models
It involves multiple tree models
It predicts the average variance of a set
True or False: Both linear regression and logistic regression can be viewed as a neural network with no hidden layers.
True
False
Which of the following is true of cluster analysis?
It is a data analysis technique to discover trends in time-series data
It is a data mining tool that is used to create homogeneous groups
It is a data visualization tool in market research
It is a model for customer behavior in the organic and natural products industry
Which of the following settings are appropriate applications of cluster analysis? (Select all that apply.)
A recommender system that seeks to predict the rating or preference that a user would give to an item (e.g., a movie, a book, or a restaurant)
A delivery scheduling system that assigns delivery trucks to customers in the same general geographical area
A cable company seeking to identify the number and type of TV packages to offer (e.g., Basic, Sports, Entertainment, or Premium)
An inventory management system for retail pharmacies that attempts to minimize both the probability of running out of stock and the inventory carrying cost
Which of the following statements is true of principal component analysis (PCA) and cluster analysis?
PCA and cluster analysis are incompatible techniques; only one of them can be applied to the same data
PCA is a data reduction technique and cluster analysis is a dimensionality reduction technique
Cluster analysis is a data reduction technique and PCA is a dimensionality reduction technique
The main goal of cluster analysis is to identify redundant variables, and the main goal of PCA is to create homogeneous groups of observations
Cluster analysis is considered an unsupervised learning technique because it operates on historical observations that are not labeled. That is, it is not known to which group historical observations belong, and therefore it is not known how many groups there are.
True
False
