wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

архитектура 1

Total questions: 69

Worksheet time: 35mins

Name
Class
Date
1.

What statement below best describes why we do data analytics in business?

a)

Analytics improve our understanding of how the business works

b)

We must show a return on the investment we make in data & analytical resources

c)

We need specific insights to make business decisions

d)

We have to calculate & report financial results to owners/shareholders

2.

What should you consider as you approach an analytical problem and in which order? Identify the correct order for the following ideas/steps:
А Sourcing Data,
B Analysis Outputs,
C Execute Analysis,
D Analysis Methods,
E Define Decision,
F Data Needs

a)

ABCDEF

b)

EBDAFC

c)

EBDFAC

d)

BDFACE

3.

Select a source best describes where the following data might come from: "The average temperature of a turbine bearing over the last 8 hours"

a)

Machine Data System

b)

Billing System

c)

Enterprise Resource Planning System

d)

Usage Tracking System

e)

Customer Relationship Management System

4.

Select a source that best describes where the following data might come from: "The number of developers allocated to a company software project"

a)

Customer Relationship Management System

b)

Usage Tracking System

c)

Billing System

d)

Machine Data System

e)

Enterprise Resource Planning System

5.

Select a source best describes where the following data might come from: "Household water consumption by month"

a)

Billing System

b)

Usage Tracking System

c)

Enterprise Resource Planning System

d)

Machine Data System

e)

Customer Relationship Management System

6.

Select a source best describes where the following data might come from: "The dollar amount of unpaid invoices at the end of a month"

a)

Billing System

b)

Usage Tracking System

c)

Machine Data System

d)

Enterprise Resource Planning System

e)

Customer Relationship Management System

7.

Select a source best describes where the following data might come from: "The average age of customers in Madison, Wisconsin"

a)

Machine Data System

b)

Usage Tracking System

c)

Billing System

d)

Enterprise Resource Planning System

e)

Customer Relationship Management System

8.

Identify the correct order of steps in the Information-Action Value Chain:
A Develop Strategy & Plan

  1. B Deliver the Pitch

  2. C Events & Characteristics in the Real World

  3. D Take Action

    E Data Capture by Source Systems

  4. F Data Extraction

  5. G Data Storage

  6. H Analytical Methods

  7. I Summarize & Interpret Results

a)

CEFGHIABD

b)

CEGFHIABDㅤ

c)

CEGFHIBAD

9.

Why do we bring data together into a common location? (Select all that apply.)

a)

We can establish relationships among data sources

b)

It's more convenient for extraction to have data in one place

c)

Sometimes we can't access source systems directly

d)

Source data may be unstructured or not formatted for analysis

10.

What type of analytics would you use to determine the best way to route delivery trucks to minimize miles driven or gasoline consumed?

a)

Descriptive

b)

Predictive

c)

Transitive

d)

Cognitive

e)

Prescriptive

11.

What type of file normally stores two-dimensional data with column and row breaks, identified using special characters?

a)

XML File

b)

Log File

c)

Delimited Text File

d)

Excel File

12.

What term best describes data storage that is optimized for handling front-end business operations?

a)

Document Store

b)

Hadoop Distributed File System (HDFS)

c)

Online Transactional Processing (OLTP)

d)

Online Analytical Processing (OLAP)

13.

Suppose you are a software developer looking for an online environment to help you rapidly build and scale applications. Which of the following services would best accommodate your needs?

a)

Platform as a Service (PaaS)

b)

Software as a Service (SaaS)

c)

Development as a Service (DaaS)

d)

Infrastructure as a Service (IaaS)

14.

Which of the following statements about Cloud computing are true? (Select all that apply)

a)

Cloud computing is needed for handling Big Data

b)

Cloud computing speaks to where data is stored or manipulated

c)

Cloud computing is more secure than a company's data center

d)

Cloud computing outsources all of a company's data operations

e)

Cloud computing can allow cheaper and more scalable operations

15.

Suppose your objective is to build a predictive model that can be used to recommend products to customers in real-time based on their navigation on your website. Which of these technologies would be most critical in helping you achieve this objective?

a)

In-Database Analytics

b)

In-Memory Computing

c)

Data Federation

d)

Hadoop Distributed File System (HDFS)

e)

Data Virtualization

16.

Suppose you are a data analyst working on a project to show why sales in a particular region are down relative to other regions. Your job is to figure out what's going on, find a good way to show the data, and produce a report that can be automated to go out weekly to track progress on any actions that are taken. You anticipate that only descriptive analytics will be needed for this project, and you're working from a dataset that has been prepared by your partners in IT. Which of the following classes of tools are you most likely to use directly in this project? (Select all that apply)

a)

Dashboarding

b)

Statistical modeling

c)

Data visualization & exploration

d)

Database systems

e)

Standard reporting

17.

Suppose you're a data analyst and you're traveling to a conference. There's a straightforward but critical ad-hoc analysis you need to accomplish, but you're not certain how much internet connectivity you'll have during your trip. You also haven't decided which of your desktop tools you'll use in the analysis. Which of the following process methodologies would work best for your situation?

a)

Intermediate File Approach

b)

Direct Connection Approach

c)

Downstream Integration Approach

18.

You've just completed an analysis that reveals the importance of a few metrics that business leaders would like to see monthly. You need someone to help you productionalize and automate a monthly report containing those metrics. Who should you talk to?

a)

Data Architect

b)

IT Infrastructure Resource

c)

BI Developer

d)

Database Administrator

e)

Application Developer

19.

You've arranged for an external partner to send you data each day. You need someone to help set up a file transfer process that will allow that partner to securely connect to your company through a firewall. Who should you talk to?

a)

IT Infrastructure Resource

b)

Application Developer

c)

Database Administrator

d)

ETL Developer

e)

Data Architect

20.

To ensure that the results of a data analysis can be placed into context, you need someone who can examine how certain business processes work and help you map them out. Who should you talk to?

a)

Business Analyst

b)

IT Infrastructure Resource

c)

Database Administrator

d)

Application Developer

e)

Data Architect

21.

You've done a descriptive analysis that seems to show a correlation between customer defection and several customer characteristics, but you think that a formal statistical procedure would yield more powerful results that can predict churn. You need someone who knows how to do this. Who should you talk to?

a)

Modeler

b)

Database Administrator

c)

IT Infrastructure Resource

d)

Data Architect

e)

Application Developer

22.

You know that a new product is coming online, and you'd like to understand how measurements around that product will be represented in the database model. Who should you talk to?

a)

ETL Developer

b)

Database Administrator

c)

Data Architect

d)

Application Developer

e)

IT Infrastructure Resource

23.

You're finding that the SQL queries you are writing against your data warehouse are taking a long time to run. You need someone who can help you determine if your queries are written in the best way. Who should you talk to?

a)

IT Infrastructure Resource

b)

Database Administrator

c)

Data Architect

d)

Application Developer

e)

ETL Developer

24.

Your company is functionally organized. The data sources and analytical techniques tend to be pretty similar across functions, and most resources are located in a headquarters building in downtown Chicago. The executive team gets along, but they are very protective of their teams and work product. Which structure would fit best in this scenario?

a)

Allocated Model

b)

Centralized Model

c)

Distributed Model

d)

Coordinated Model

25.

Your company is a multinational organization that operates in a number of distinct industries. Each industry uses its own methods and tends to hire somewhat different types of people into analytical organizations. Which structure would fit best in this scenario?

a)

Allocated Model

b)

Centralized Model

c)

Distributed Model

d)

Coordinated Model

26.

Your company is organized by customer groups, which are mostly distinct but have some limited overlaps. Analyses vary in similarity - some are very similar, but others are quite different. They do, however, use most of the same data sources. Currently, resources are located within each customer group organization, but it's pretty typical for there to be only one or two analysts in each area. Which structure would fit best in this scenario (select all that apply)?

a)

Allocated Model

b)

Centralized Model

c)

Distributed Model

d)

Coordinated Model

27.

What term best describes the process of identifying and standardizing an organization's most critical data?

a)

SOX Compliance

b)

Metadata Management

c)

Master Data Management

d)

Data Governance

e)

Data Stewardship

28.

Who is responsible for making sure that a data domain is correctly represented and used within an organization?

a)

Data Steward

b)

Data Architect

c)

Data Governance Council

d)

SOX Compliance Auditor

e)

ETL Developer

29.

A large drugstore chain wants to use prescription data from its pharmacy to make complementary relevant offers to specific customers via custom coupon books, delivered via direct mail. What is the most limiting standard that might be relevant in this case?

a)

Policy Standards

b)

Legal Standards

c)

Good Judgement

d)

Ethical Standards

30.

A Mobile Phone Company wants to construct and sell 'profiles' of customers based on a combination of internet sites visited and location data. The profiles would provide aggregate information that is not considered CPNI. What are the most limiting standards that might be relevant in this case? (select all that apply)

a)

Policy Standards

b)

Legal Standards

c)

Good Judgement

d)

Ethical Standards

31.

A few years ago your company acquired another company and merged summary financial data into a key database. The data looks complete, but there are some peculiarities we can't explain. What is the dominant issue in this case?

a)

Completeness / Uniqueness

b)

Accuracy / Consistency

c)

Conformance / Validity

d)

Timeliness

e)

Provenance

32.

Your company wants a mobile application that allows certain purchases to be made via the application. However, those transactions use the date/time of the user's device as the timestamp of the transaction that is stored in the purchase database. What is the dominant issue in this case?

a)

Completeness / Uniqueness

b)

Accuracy / Consistency

c)

Conformance / Validity

d)

Timeliness

e)

Provenance

33.

We notice that when we join data from two different tables, we need to be careful to convert the time zone in one table from Central Standard Time (CST) to Universal Coordinated Time (UTC) to match the second table, even though the company standard is UTC. What is the dominant issue in this case?

a)

Completeness / Uniqueness

b)

Accuracy / Consistency

c)

Conformance / Validity

d)

Timeliness

e)

Provenance

34.

On your company's website, a customer can accidentally click a purchase button twice, which results in two purchase records being generated. Luckily, these purchases are filtered by the credit card payment processing system and removed from the company's general ledger. However, those records are not removed from the analytical data warehouse. What is the dominant issue in this case?

a)

Completeness / Uniqueness

b)

Accuracy / Consistency

c)

Conformance / Validity

d)

Timeliness

e)

Provenance

35.

At what stage(s) of Data Exploration would you address missing values in a dataset?

a)

Data transformation

b)

Data clean-up

c)

Data reduction

36.

Which of the following statements regarding data transformation and data reduction is correct?

a)

Data transformations work on individual variables, while data reduction works on a set of variables

b)

Only data transformation would create dummy variables

c)

The goal of data transformations is to create larger datasets while the goal of data reduction is to create smaller datasets

d)

Data transformations are out of style; data reduction is the modern man's tool

37.

What does a data value measure after centering and scaling has been applied?

a)

Accuracy

b)

The number of standard deviations between each data point and the median

c)

The number of standard deviations between each data point and the mean

d)

Slope

38.

Why would one want to center and scale a set of data?

a)

So multiple variables in the dataset are on a common scale

b)

To make all data values positive

c)

To remove duplicates

d)

To make data easier to interpret

39.

Match the Box-Cox Transformation associated with the given value of lambda: When Lambda = 0, transformation is

a)

Logarithmic

b)

Cubed polynomial

c)

Inverse

d)

Square root

40.

Match the Box-Cox Transformation associated with the given value of lambda: When Lambda = 0.5, transformation is

a)

Logarithmic

b)

Cubed polynomial

c)

Inverse

d)

Square root

41.

Match the Box-Cox Transformation associated with the given value of lambda: When Lambda = -1, transformation is

a)

Logarithmic

b)

Cubed polynomial

c)

Inverse

d)

Square root

42.

What is the purpose of applying a Data Reduction?

a)

To generate a larger set of variables

b)

To make all variables positively valued

c)

To use a smaller set of variables to capture most of the information in the original variables

43.

What must be done to variables of a dataset before applying principal component analysis and why?

a)

You must scale the variables so that only outliers are considered as principal components

b)

You must scale the variables so that principal components are not dominated by variables of much larger scale

c)

You must make all variables negative to work with values of the same sign

d)

You must take the square root of all data values to reduce the overall magnitudes of the dataset

44.

Which of the following can be an appropriate way to deal with missing values? (Select all that apply.)

a)

Removing the columns or rows with missing values

b)

Imputing a value with averages of all other records

c)

Imputing a value from 'similar' data points

d)

Making 'missing' its own category

45.

Your organization asks you to analyze a dataset that shows the number of FreeFly ALTA drones sold in 2016. You noticed that only 2 drones were sold the day after Black Friday, while the average number of drones sold in 2016 is around 100 a day. What is the most probable explanation for this small data value?

a)

It's a missing value that someone filled in with a guess

b)

There was a glitch in the system, and the data value was corrupted

c)

It's a censored value that was inputted incorrectly

d)

It's a censored value; drone inventory probably ran out

46.

What are the risks of replacing a missing value with a guess? (Select all that apply.)

a)

None, the database is capable of correcting input mistakes

b)

Introducing biases

c)

Distorting the dataset

d)

Falsifying results

47.

Why is removing all data records with missing values often not a good way to deal with missing values? (Select all that apply.)

a)

Some modeling tools require a data value for each row/column

b)

A dataset is incomplete if there are missing values

c)

We may end up with too little data to conduct meaningful analysis

d)

The pattern of missing values can have high predictive power

48.

What are the characteristics of an outlier? (Select all that apply.)

a)

It is the data point most proximal to the mean

b)

It is the pivot point for the overall pattern that the data follows

c)

It falls far outside the overall data pattern

d)

It is above or below 3 standard deviations of the mean

49.

A data point is not considered an outlier unless it deviates dramatically on either the x-axis or the y-axis.

a)

True

b)

False

50.

Why do outliers exist? (Select all that apply.)

a)

Data recording errors

b)

Legitimate but odd observations

c)

Entropy of a system

d)

Distortion of time

51.

Which statistical measure is more resistant to outliers?

a)

Mean

b)

Median

c)

Standard deviation

d)

Range

52.

To say a variable is degenerate means which of the following? (Select all that apply.)

a)

The variable is immoral and corrupt

b)

The variable can only take on a single value

c)

When plotted, the variable is modeled with an exponential decay

d)

The variable is a zero variance variable

53.

Which of the following is a remedy to collinearity issues in regression analysis?

a)

Adding more dummy variables

b)

Cutting the dataset in half

c)

Removing zero variance and near zero variance variables

d)

Duplicating the dataset

54.

Which type of target variable are we dealing with in linear regression?

a)

Binary

b)

Categorical

c)

Continuous

d)

Imaginary

55.

We cannot perform linear regression unless both the target variable and predictor variables are continuous.

a)

True

b)

False

56.

What is the validation set used for in predictive modeling?

a)

To fit the models

b)

To evaluate the various models

c)

To increase the size of our training set

d)

To average the training set data

57.

Why can multicollinearity cause problems in multiple regression? (Select all that apply.)

a)

It creates unstable estimates

b)

It creates problems in model interpretation

c)

You cannot make predictions based on regression models with multicollinearity issues

d)

It makes estimating the model impossible

58.

How many transformed variables can we create based on one predictor variable?

a)

An unlimited number

b)

2

c)

None

d)

1

59.

What do you achieve when you apply a log transformation to a variable in your data set?

a)

It makes highly skewed distributions less skewed

b)

It compresses data to make big sets more manageable

c)

It removes negative data values

d)

It removes duplicate values

60.

A soccer team is believed to have an 8 to 2 odds of winning. What is the probability of winning for the team?

a)

0.2

b)

0.25

c)

0.8

d)

2

61.

It is estimated that an appointment with a 10-day lag for a male patient has a predicted probability of 0.1372 of canceling. Compare this with the predicted cancellation probability for a female patient who also has an appointment with a 10-day lag. Assume that the value of the gender variable is 1 for male patients and 0 for females. Also, assume that the estimated coefficient for gender is -0.3572, beta-0 is -1.6515, beta-1 is 0.01699.

a)

A female is less likely to cancel by 2.4%

b)

A female is more likely to cancel by 4.8%

c)

A female is equally likely to cancel as a male

d)

A female is more likely to cancel by 6.9%

62.

The bagging procedure can reduce the variance of a predictive model. Check all true statements about the bagging method: (Check all that apply.)

a)

Helps avoid overfitting of the dataset

b)

Helps group similar data outliers

c)

Has access to multiple training sets

d)

Can be applied to tree models

63.

What do the bagging and random forest methods have in common?

a)

Both methods grow multiple numbers of trees

b)

Both methods operate on only 2 trees

c)

Both methods sample the validation set

d)

Both methods increase the variance of a dataset

64.

What sets the random forest algorithm apart from bagging and boosting algorithms?

a)

It operates on bootstrap sets

b)

It focuses on reducing correlation among models

c)

It involves multiple tree models

d)

It predicts the average variance of a set

65.

True or False: Both linear regression and logistic regression can be viewed as a neural network with no hidden layers.

a)

True

b)

False

66.

Which of the following is true of cluster analysis?

a)

It is a data analysis technique to discover trends in time-series data

b)

It is a data mining tool that is used to create homogeneous groups

c)

It is a data visualization tool in market research

d)

It is a model for customer behavior in the organic and natural products industry

67.

Which of the following settings are appropriate applications of cluster analysis? (Select all that apply.)

a)

A recommender system that seeks to predict the rating or preference that a user would give to an item (e.g., a movie, a book, or a restaurant)

b)

A delivery scheduling system that assigns delivery trucks to customers in the same general geographical area

c)

A cable company seeking to identify the number and type of TV packages to offer (e.g., Basic, Sports, Entertainment, or Premium)

d)

An inventory management system for retail pharmacies that attempts to minimize both the probability of running out of stock and the inventory carrying cost

68.

Which of the following statements is true of principal component analysis (PCA) and cluster analysis?

a)

PCA and cluster analysis are incompatible techniques; only one of them can be applied to the same data

b)

PCA is a data reduction technique and cluster analysis is a dimensionality reduction technique

c)

Cluster analysis is a data reduction technique and PCA is a dimensionality reduction technique

d)

The main goal of cluster analysis is to identify redundant variables, and the main goal of PCA is to create homogeneous groups of observations

69.

Cluster analysis is considered an unsupervised learning technique because it operates on historical observations that are not labeled. That is, it is not known to which group historical observations belong, and therefore it is not known how many groups there are.

a)

True

b)

False