Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Microsoft 70-773: Analyzing Big Data with Microsoft R Exam

Total questions: 38

Worksheet time: 1hrs 15mins

Name
Class
Date
1.

You have a Microsoft SQL Server instance that has R Services (In-Database) installed.

You need to monitor the R jobs that are sent to SQL Server.

Solution: You create an events trace configuration file and place the file in the same directory as the BXLServer process.

Does this meet the goal?

a)

Yes

b)

No

2.

You have a dataset.

You need to repeatedly split randomly the dataset so that 80 percent of the data is used as a training set and the remaining 20 percent is used as a test set.

Which method should you use?

a)

threshold

b)

binary classification

c)

imputation

d)

cross validation

e)

pruning

3.

You have a Microsoft SQL Server instance that has R Services (In-Database) installed.

You need to monitor the R jobs that are sent to SQL Server.

Solution: You register an Extended Events package.

Does this meet the goal?

a)

Yes

b)

No

4.

You need to calculate a measure of central tendency and variability for the variables in a dataset that is grouped by using another categorical variable.

What should you use?

a)

the rxCrossTabs function

b)

the rxHistogram function

c)

the rxSummary function

d)

the rxQuantile function

e)

the rxCube function

5.

You are developing a Microsoft R Open solution that will leverage the computing power of the database server for some of your datasets.

You are performing feature engineering and data preparation for the datasets.

The following is a sample of the dataset.

You need to analyze the dataset without the missing values. The solution must not remove the missing values from the dataset.

Which R code segment should you use?

a)

rxDataStep(varsToDrop = NULL)

b)

rxDataStep(transforms = `removeMissing')

c)

rxDataStep(transformFunc = `removeMissing')

d)

rxDataStep(removeMissingsOnRead = FALSE, removeMissing = TRUE)

6.

You are developing a Microsoft R Open solution that will leverage the computing power of the database server for some of your datasets.

You are performing feature engineering and data preparation for the datasets.

The following is a sample of the dataset.

You have the following R code.

Which function determines the variable?

a)

transformVars

b)

rxXdfDataFrame

c)

createRandomSample

d)

transformFunc

7.

You need to evaluate the significance of coefficients that are

produced by using a model that was estimated already.

Which function should you use?

a)

rxPredict

b)

rxLogit

c)

summary

d)

rxLinMod

e)

rxTweedie

8.

You use dplyrXdf, and you discover that after you exit the session, the output files that were created were deleted.

You need to prevent the files from being deleted.

Solution: You use rxSetComputeContext with the local parameter before performing operations that save results.

Does this meet the goal?

a)

Yes

b)

No

9.

You need to use the ScaleR distributed processing in an Apache Hadoop environment.

Which data source should you use?

a)

Microsoft SQL Server database

b)

XDF data files

c)

ODBC data

d)

Teradata database

10.

You perform an analysis that produces the decision tree shown in the exhibit.

How many leaf nodes are there on the tree?

a)

2

b)

3

c)

5

d)

7

11.

You plan to read data from an Oracle database table and to store the data in the file system for later processing by dplyrXdf. The size of the data is larger than the memory on the server to be used for modelling.

You need to ensure that the data can be processed by dplyrXdf in the least amount of time possible.

How should you transfer the data from the Oracle database?

a)

Define a data source to the Oracle database server by using RxOdbcData. Use rxImport to save the data to a comma-separated values (CSV) file.

b)

Use the RODBC library, connect to the Oracle database server by using odbcConnect, and then use rxDataStep to export the data to a comma-separated values (CSV) file.

c)

Define a data source to the Oracle database server by using RxOdbcData, and then use rxImport to save the data to an XDF file.

d)

Use the RODBC library, connect to the Oracle database server by using odbcConnect, and then use rxSplit to save the data to multiple comma-separated values (CSV) file.

12.

You need to estimate a model where the outcome variable is continuous, is in the range of [0, inf], and has a substantial mass at an exact value of 0.

Which function should you use?

a)

rxLogit

b)

rxLinMod

c)

rxTweedie

d)

stepAic

13.

You plan to analyze data on a local computer. To improve performance, you plan to alternate the operation between a Microsoft SQL Server and the local computer.

You need to run complex code on the SQL Server, and then revert to the local compute context.

Which R code segment should you use?

a)

sqlCompute <-RxInSqlServer(connectionString = "Driver=SQL Server;Server = myServer;Database = TestDB; Uid = myID; Pwd = myPwd;")

sqlPackagePaths <-RxFindPackage(package =

"RevoScaleR", computeContext = sqlServerCompute)

b)

sqlCompute <-RxInSqlServer(connectionstring = sqlConnString, shareDir = sqlShareDir,wait = sqlWait, consoleOutput =

sqlConsoleOutput)

rxSetComputeContext("local")

x <-1:10

rxExec(print, x, elemType = "cores", timesToRun = 10)

rxSetComputeContext("RxLocalParallel")

c)

sqlCompute <-RxInSqlServer(connectionstring = sqlConnString, shareDir = sqlShareDir,wait = sqlWait, consoleOutput =

sqlConsoleOutput)

rxSetComputeContext("sqlCompute")

x <-1:10

rxExec(print, x, elemType = "cores", timesToRun = 10)

rxSetComputeContext("local")

d)

sqlCompute <-RxInSqlServer(connectionstring = sqlConnString, shareDir = sqlShareDir,wait = sqlWait, consoleOutput =

sqlConsoleOutput)

rxSetComputeContext("local")

x <-1:10

rxExec(print, x, elemType = "cores", timesToRun = 10)

rxSetComputeContext("sqlCompute")

14.

You have a slow Map Reduce job.

You need to optimize the job to control the number of mapper and runner tasks.

Which function should you use?

a)

RxComputeContext

b)

RxHadoopMR

c)

rxExec

d)

RxLocalParallel

15.

You need to build a model that looks at the probability of an outcome. You must regulate between L1 and L2.

Which classification method should you use?

a)

Two-Class Neutral Network

b)

Two-Class Support Vector Machine

c)

Two-Class Decision Forest

d)

Two-Class Logistic Regression

16.

You are planning the compute contexts for your environment. You need to execute rx-function calls in parallel.

What are three possible compute contexts that you can use to achieve this goal? Each correct answer presents a complete solution.

NOTE: Each correct selection is worth one point.

a)

local parallel

b)

Spark

c)

local sequential

d)

Map Reduce

e)

SQL

17.

You use dplyrXdf, and you discover that after you exit the session, the output files that were created were deleted.

You need to prevent the files from being deleted.

Solution: You use dplyrXdf with the persist verb.

Does this meet the goal?

a)

Yes

b)

No

18.

You have a data source that is larger than memory.

You need to visualize the distribution of the values for a variable in the data source.

What should you use?

a)

the rxHistogram function

b)

the rxSummary function

c)

the rxQuantile function

d)

the summary function

e)

the rxCube function

19.

You have a dataset that has multiple blocks and only numeric variables. You are computing in a local compute context.

You plan to lag a variable named x to create a new variable named x_lagged by using a transform function. You will create a new element in the output of the function.

You need to minimize the number of missing values.

Which three actions should you perform?

Each correct answer presents part of the solution.

NOTE: Each correct selection is worth one point.

a)

Assign a value to the first value of x_lagged in the current block.

b)

Use rxSet to store the last value of x_lagged in the current block.

c)

Use rxSet to store the last value of x in the current block.

d)

Use rxGet to retrieve the first value of x in the next block to be processed.

e)

Use rxGet to retrieve a value stored in processing of the prior block.

20.

You have an Apache Hadoop Hive data warehouse. RevoScalerR is not installed.

You need to sort the data according to the variables in the dataset.

What should you do?

a)

Connect to the database by using an ODBC connection, and then use the rxSort function.

b)

Create a table in the ORC file format.

c)

Connect to the database by using an ODBC connection, and then use the rxDataStep function.

d)

Execute a Hive query that sorts the data, and then reads the results.

21.

You need to get all of the deciles for a variable in a data frame.

What should you use?

a)

the Describe package

b)

the rxHistogram function

c)

the rxSummary function

d)

the rxQuantile function

e)

the rxCube function

22.

You need to run a large data tree model by using rxDForest. The model must use cross validation.

Which rxDForest option should you use?

a)

maxSurrogate

b)

maxNumBins

c)

maxDepth

d)

maxCompete

e)

xVal

23.

You build a model that uses xyz regression.

You need to estimate a model that predicts a binary variable.

Which function should you use?

a)

rxLogit

b)

rxLinMod

c)

rxTweedie

d)

stepAic

24.

You have one-class support vector machines (SVMs).

You have a large dataset, but you do not have enough training time to fully test the model.

What is an alternative method to validate the model?

a)

Use Principal Components Analysis (PCA)-Based Anomaly Detection.

b)

Replace the SVMs with two-class SVMs.

c)

Perform feature selection.

d)

Use outlier detection.

25.

You are running a parallel function that uses the following R code segment. (Line numbers are included for reference only.)


01 cp <-0.01 xval <-0 maxdepth <-5

02 [...](form, data = "segmentationDataBig", maxDepth = maxdepth, cp = cp, xval = xval, blocksPerRead = 250)


You need to complete the R code. The solution must support chunking.

Which function should insert at line 02?

a)

rxBTrees

b)

rxExec

c)

rxDForest

d)

rxDTree

26.

You are running a large logistic regression for 1,000 feature variables by using the LogisticRegression() function in the MicrosoftML package. All of the predictor variables are

numeric.

Currently, you specify the input variables separately by using the following formula:

Outcome ~ Feature000 + Feature001 + Feature002 + ... + Feature999

You discover that it takes 20 minutes to estimate each model. You need to reduce the amount of time required to estimate each model without losing any information in the predictors.

What should you do?

a)

Use stepControl() to perform stepwise regression to limit the number of variables that contribute to the model.

b)

Use selectFeatures() to select the features that provide the most information about the outcome variable.

c)

Use princomp() on the correlation matrix of Features, and then use only the first 100 principle components to reduce the number of input variables.

d)

Use concat() to create a single array variable named Features, and then specify a new formula named Outcome ~ Features.

27.

You have a Microsoft SQL Server instance that has R Services (In-Database) installed. The server has a comma-separated values (CSV) file stored in the local file system.

For analytic purposes, you need to read the CSV file into a database table in the SQL Server instance.

You connect to the SQL Server instance by using SQL Server Management Studio.

What should you use from sp_execute_external_script?

a)

RxSqlServerData and specify the CSV file path in the connection string

b)

rxDataStep and specify the CSV file path as the inFile argument

c)

rxImportToXdf and specify specify the CSV file as the input

d)

read.csv and specify the CSV file path as the parameter

28.

You need to prevent the files from being deleted.

Solution: You use dplyrXdf with the outFile parameter and specify a path other than the working directory for dplyrXdf.

Does this meet the goal?

a)

Yes

b)

No

29.

You have a Microsoft SQL Server instance that has R Services (In-Database) installed.

You need to monitor the R jobs that are sent to SQL Server.

Solution: You call a function from the RevoPemaR package.

Does this meet the goal?

a)

Yes

b)

No

30.

You need to generate a residual based on two columns. The solution must build a trend indicator.

Which function should you use?

a)

rxPredict

b)

rxLogit

c)

rxLinMod

d)

rxTweedie

e)

stepAic

31.

You have cloud and on-premises resources that include Microsoft SQL Server and a big data environment in Apache Hadoop.

You have 50 billion fact records.

You need to build time series models to execute forecasting reports on the fact records.

What should you use?

a)

RxSpark on the Hadoop cluster

b)

RxHadoopMR on the Hadoop cluster

c)

RxLocalseq on the SQL Server database

d)

RxLocalParallel on the SQL Server database

32.

You have a dataset that has a character variable.

You need to create a bag of counts of n-grams.

Which function should you use?

a)

featurizeText()

b)

categoricalHash()

c)

concat()

d)

selectFeatures()

e)

categorical()

33.

You have a dataset that contains the physical characteristics of people.

You need to visualize a relationship between height and weight for a subset of observations in the dataset.

What should you use?

a)

the Describe package

b)

the rxQuantile function

c)

the rxCube function

d)

the ggplot2 package

e)

the rxCrossTabs function

34.

You have the following regression forest.

Which variable contributes the most to the dependent variable?

a)

stack.loss

b)

Water.Temp

c)

Air.Flow

d)

Acid.Conc.

35.

You are developing a Microsoft R Open solution that will leverage the computing power of the database server for some of your datasets.

You are performing feature engineering and data preparation for the datasets.

You need to sort the data from the dataset sample and to remove duplicates by using wkswork1.

Which R code segment should you use? to answer, select the appropriate options in the answer area.

a)

removeDupKeys=TRUE; varsToDrop='wkswork1'

b)

removeDupKeys=FALSE; varsToKeep='wkswork1'

c)

removeDupKeys=TRUE; dupFreqVar='wkswork1'

d)

removeDupKeys=FALSE; dupFreqVar='wkswork1'

36.

You need to set the compute context for three different target environments.

Which Statement should you use for each environment? To answer, drag the appropriate statements to the correct execution contexts.

a)

1. RxSpark(); 2. RxHadoopMR(); 3. rxSetComputeContext('localpar')

b)

1. rxSetComputeContext('localpar'); 2. RxHadoopMR(); 3. rxSetComputeContext('local')

c)

1. RxHadoopMR(); 2. RxSpark(); 3. rxSetComputeContext('localpar')

d)

1. RxHadoopMR(); 2. RxSpark(); 3. rxSetComputeContext('local')

e)

1. rxSetComputeContext('local'); 2. rxSetComputeContext('localpar'); 3. RxHadoopMR();

37.

You are developing a Microsoft R Open solution that will leverage the computing power of the database server for some of your datasets.

You are performing feature engineering and data preparation for the datasets. The following is a sample of the dataset.

You plan to score some data to create data features to address empty rows. You have the following R code.

You need to transform the data and overwrite the current dataset. Which R code segment should you use?

a)

rxExec(inData=[sampleInData], outFile=[sampleOutDataIncludingFeatures], transformFunc=computeNonLagFeatures, overwrite=FALSE)

b)

rxDataStep(inData=[sampleInData], outFile=[sampleOutDataIncludingFeatures], transformFunc=computeNonLagFeatures, overwrite=FALSE)

c)

rxDataStep(inData=[sampleInData], outFile=[sampleOutDataIncludingFeatures], transformFunc=computeNonLagFeatures, overwrite=TRUE)

d)

transform(inData=[sampleInData], outFile=[sampleOutDataIncludingFeatures], transformFunc=computeNonLagFeatures, overwrite=FALSE)

38.

You are using rxPredict for a logistic regression model.

You need to obtain prediction standard errors and confidence intervals. Which R code segment should you

use?

a)

glm; covCoef=FALSE; computeStdErr=TRUE; interval='confidence'

b)

rxLogit; covCoef=none; computeStdErr=TRUE; interval='confidence'

c)

glm; covCoef=TRUE; computeStdErr=FALSE; interval='confidence'

d)

rxLogit; covCoef=FALSE; computeStdErr=TRUE; interval='confidence'

e)

rxLogit; covCoef=TRUE; computeStdErr=TRUE; interval='confidence'