WorksheetsMicrosoft 70-773: Analyzing Big Data with Microsoft R Exam
Total questions: 38
Worksheet time: 1hrs 15mins
You have a Microsoft SQL Server instance that has R Services (In-Database) installed.
You need to monitor the R jobs that are sent to SQL Server.
Solution: You create an events trace configuration file and place the file in the same directory as the BXLServer process.
Does this meet the goal?
Yes
No
You have a dataset.
You need to repeatedly split randomly the dataset so that 80 percent of the data is used as a training set and the remaining 20 percent is used as a test set.
Which method should you use?
threshold
binary classification
imputation
cross validation
pruning
You have a Microsoft SQL Server instance that has R Services (In-Database) installed.
You need to monitor the R jobs that are sent to SQL Server.
Solution: You register an Extended Events package.
Does this meet the goal?
Yes
No
You need to calculate a measure of central tendency and variability for the variables in a dataset that is grouped by using another categorical variable.
What should you use?
the rxCrossTabs function
the rxHistogram function
the rxSummary function
the rxQuantile function
the rxCube function
You are developing a Microsoft R Open solution that will leverage the computing power of the database server for some of your datasets.
You are performing feature engineering and data preparation for the datasets.
The following is a sample of the dataset.
You need to analyze the dataset without the missing values. The solution must not remove the missing values from the dataset.
Which R code segment should you use?
rxDataStep(varsToDrop = NULL)
rxDataStep(transforms = `removeMissing')
rxDataStep(transformFunc = `removeMissing')
rxDataStep(removeMissingsOnRead = FALSE, removeMissing = TRUE)
You are developing a Microsoft R Open solution that will leverage the computing power of the database server for some of your datasets.
You are performing feature engineering and data preparation for the datasets.
The following is a sample of the dataset.
You have the following R code.
Which function determines the variable?
transformVars
rxXdfDataFrame
createRandomSample
transformFunc
You need to evaluate the significance of coefficients that are
produced by using a model that was estimated already.
Which function should you use?
rxPredict
rxLogit
summary
rxLinMod
rxTweedie
You use dplyrXdf, and you discover that after you exit the session, the output files that were created were deleted.
You need to prevent the files from being deleted.
Solution: You use rxSetComputeContext with the local parameter before performing operations that save results.
Does this meet the goal?
Yes
No
You need to use the ScaleR distributed processing in an Apache Hadoop environment.
Which data source should you use?
Microsoft SQL Server database
XDF data files
ODBC data
Teradata database
You perform an analysis that produces the decision tree shown in the exhibit.
How many leaf nodes are there on the tree?
2
3
5
7
You plan to read data from an Oracle database table and to store the data in the file system for later processing by dplyrXdf. The size of the data is larger than the memory on the server to be used for modelling.
You need to ensure that the data can be processed by dplyrXdf in the least amount of time possible.
How should you transfer the data from the Oracle database?
Define a data source to the Oracle database server by using RxOdbcData. Use rxImport to save the data to a comma-separated values (CSV) file.
Use the RODBC library, connect to the Oracle database server by using odbcConnect, and then use rxDataStep to export the data to a comma-separated values (CSV) file.
Define a data source to the Oracle database server by using RxOdbcData, and then use rxImport to save the data to an XDF file.
Use the RODBC library, connect to the Oracle database server by using odbcConnect, and then use rxSplit to save the data to multiple comma-separated values (CSV) file.
You need to estimate a model where the outcome variable is continuous, is in the range of [0, inf], and has a substantial mass at an exact value of 0.
Which function should you use?
rxLogit
rxLinMod
rxTweedie
stepAic
You plan to analyze data on a local computer. To improve performance, you plan to alternate the operation between a Microsoft SQL Server and the local computer.
You need to run complex code on the SQL Server, and then revert to the local compute context.
Which R code segment should you use?
sqlCompute <-RxInSqlServer(connectionString = "Driver=SQL Server;Server = myServer;Database = TestDB; Uid = myID; Pwd = myPwd;")
sqlPackagePaths <-RxFindPackage(package =
"RevoScaleR", computeContext = sqlServerCompute)
sqlCompute <-RxInSqlServer(connectionstring = sqlConnString, shareDir = sqlShareDir,wait = sqlWait, consoleOutput =
sqlConsoleOutput)
rxSetComputeContext("local")
x <-1:10
rxExec(print, x, elemType = "cores", timesToRun = 10)
rxSetComputeContext("RxLocalParallel")
sqlCompute <-RxInSqlServer(connectionstring = sqlConnString, shareDir = sqlShareDir,wait = sqlWait, consoleOutput =
sqlConsoleOutput)
rxSetComputeContext("sqlCompute")
x <-1:10
rxExec(print, x, elemType = "cores", timesToRun = 10)
rxSetComputeContext("local")
sqlCompute <-RxInSqlServer(connectionstring = sqlConnString, shareDir = sqlShareDir,wait = sqlWait, consoleOutput =
sqlConsoleOutput)
rxSetComputeContext("local")
x <-1:10
rxExec(print, x, elemType = "cores", timesToRun = 10)
rxSetComputeContext("sqlCompute")
You have a slow Map Reduce job.
You need to optimize the job to control the number of mapper and runner tasks.
Which function should you use?
RxComputeContext
RxHadoopMR
rxExec
RxLocalParallel
You need to build a model that looks at the probability of an outcome. You must regulate between L1 and L2.
Which classification method should you use?
Two-Class Neutral Network
Two-Class Support Vector Machine
Two-Class Decision Forest
Two-Class Logistic Regression
You are planning the compute contexts for your environment. You need to execute rx-function calls in parallel.
What are three possible compute contexts that you can use to achieve this goal? Each correct answer presents a complete solution.
NOTE: Each correct selection is worth one point.
local parallel
Spark
local sequential
Map Reduce
SQL
You use dplyrXdf, and you discover that after you exit the session, the output files that were created were deleted.
You need to prevent the files from being deleted.
Solution: You use dplyrXdf with the persist verb.
Does this meet the goal?
Yes
No
You have a data source that is larger than memory.
You need to visualize the distribution of the values for a variable in the data source.
What should you use?
the rxHistogram function
the rxSummary function
the rxQuantile function
the summary function
the rxCube function
You have a dataset that has multiple blocks and only numeric variables. You are computing in a local compute context.
You plan to lag a variable named x to create a new variable named x_lagged by using a transform function. You will create a new element in the output of the function.
You need to minimize the number of missing values.
Which three actions should you perform?
Each correct answer presents part of the solution.
NOTE: Each correct selection is worth one point.
Assign a value to the first value of x_lagged in the current block.
Use rxSet to store the last value of x_lagged in the current block.
Use rxSet to store the last value of x in the current block.
Use rxGet to retrieve the first value of x in the next block to be processed.
Use rxGet to retrieve a value stored in processing of the prior block.
You have an Apache Hadoop Hive data warehouse. RevoScalerR is not installed.
You need to sort the data according to the variables in the dataset.
What should you do?
Connect to the database by using an ODBC connection, and then use the rxSort function.
Create a table in the ORC file format.
Connect to the database by using an ODBC connection, and then use the rxDataStep function.
Execute a Hive query that sorts the data, and then reads the results.
You need to get all of the deciles for a variable in a data frame.
What should you use?
the Describe package
the rxHistogram function
the rxSummary function
the rxQuantile function
the rxCube function
You need to run a large data tree model by using rxDForest. The model must use cross validation.
Which rxDForest option should you use?
maxSurrogate
maxNumBins
maxDepth
maxCompete
xVal
You build a model that uses xyz regression.
You need to estimate a model that predicts a binary variable.
Which function should you use?
rxLogit
rxLinMod
rxTweedie
stepAic
You have one-class support vector machines (SVMs).
You have a large dataset, but you do not have enough training time to fully test the model.
What is an alternative method to validate the model?
Use Principal Components Analysis (PCA)-Based Anomaly Detection.
Replace the SVMs with two-class SVMs.
Perform feature selection.
Use outlier detection.
You are running a parallel function that uses the following R code segment. (Line numbers are included for reference only.)
01 cp <-0.01 xval <-0 maxdepth <-5
02 [...](form, data = "segmentationDataBig", maxDepth = maxdepth, cp = cp, xval = xval, blocksPerRead = 250)
You need to complete the R code. The solution must support chunking.
Which function should insert at line 02?
rxBTrees
rxExec
rxDForest
rxDTree
You are running a large logistic regression for 1,000 feature variables by using the LogisticRegression() function in the MicrosoftML package. All of the predictor variables are
numeric.
Currently, you specify the input variables separately by using the following formula:
Outcome ~ Feature000 + Feature001 + Feature002 + ... + Feature999
You discover that it takes 20 minutes to estimate each model. You need to reduce the amount of time required to estimate each model without losing any information in the predictors.
What should you do?
Use stepControl() to perform stepwise regression to limit the number of variables that contribute to the model.
Use selectFeatures() to select the features that provide the most information about the outcome variable.
Use princomp() on the correlation matrix of Features, and then use only the first 100 principle components to reduce the number of input variables.
Use concat() to create a single array variable named Features, and then specify a new formula named Outcome ~ Features.
You have a Microsoft SQL Server instance that has R Services (In-Database) installed. The server has a comma-separated values (CSV) file stored in the local file system.
For analytic purposes, you need to read the CSV file into a database table in the SQL Server instance.
You connect to the SQL Server instance by using SQL Server Management Studio.
What should you use from sp_execute_external_script?
RxSqlServerData and specify the CSV file path in the connection string
rxDataStep and specify the CSV file path as the inFile argument
rxImportToXdf and specify specify the CSV file as the input
read.csv and specify the CSV file path as the parameter
You need to prevent the files from being deleted.
Solution: You use dplyrXdf with the outFile parameter and specify a path other than the working directory for dplyrXdf.
Does this meet the goal?
Yes
No
You have a Microsoft SQL Server instance that has R Services (In-Database) installed.
You need to monitor the R jobs that are sent to SQL Server.
Solution: You call a function from the RevoPemaR package.
Does this meet the goal?
Yes
No
You need to generate a residual based on two columns. The solution must build a trend indicator.
Which function should you use?
rxPredict
rxLogit
rxLinMod
rxTweedie
stepAic
You have cloud and on-premises resources that include Microsoft SQL Server and a big data environment in Apache Hadoop.
You have 50 billion fact records.
You need to build time series models to execute forecasting reports on the fact records.
What should you use?
RxSpark on the Hadoop cluster
RxHadoopMR on the Hadoop cluster
RxLocalseq on the SQL Server database
RxLocalParallel on the SQL Server database
You have a dataset that has a character variable.
You need to create a bag of counts of n-grams.
Which function should you use?
featurizeText()
categoricalHash()
concat()
selectFeatures()
categorical()
You have a dataset that contains the physical characteristics of people.
You need to visualize a relationship between height and weight for a subset of observations in the dataset.
What should you use?
the Describe package
the rxQuantile function
the rxCube function
the ggplot2 package
the rxCrossTabs function
You have the following regression forest.
Which variable contributes the most to the dependent variable?
stack.loss
Water.Temp
Air.Flow
Acid.Conc.
You are developing a Microsoft R Open solution that will leverage the computing power of the database server for some of your datasets.
You are performing feature engineering and data preparation for the datasets.
You need to sort the data from the dataset sample and to remove duplicates by using wkswork1.
Which R code segment should you use? to answer, select the appropriate options in the answer area.
removeDupKeys=TRUE; varsToDrop='wkswork1'
removeDupKeys=FALSE; varsToKeep='wkswork1'
removeDupKeys=TRUE; dupFreqVar='wkswork1'
removeDupKeys=FALSE; dupFreqVar='wkswork1'
You need to set the compute context for three different target environments.
Which Statement should you use for each environment? To answer, drag the appropriate statements to the correct execution contexts.
1. RxSpark(); 2. RxHadoopMR(); 3. rxSetComputeContext('localpar')
1. rxSetComputeContext('localpar'); 2. RxHadoopMR(); 3. rxSetComputeContext('local')
1. RxHadoopMR(); 2. RxSpark(); 3. rxSetComputeContext('localpar')
1. RxHadoopMR(); 2. RxSpark(); 3. rxSetComputeContext('local')
1. rxSetComputeContext('local'); 2. rxSetComputeContext('localpar'); 3. RxHadoopMR();
You are developing a Microsoft R Open solution that will leverage the computing power of the database server for some of your datasets.
You are performing feature engineering and data preparation for the datasets. The following is a sample of the dataset.
You plan to score some data to create data features to address empty rows. You have the following R code.
You need to transform the data and overwrite the current dataset. Which R code segment should you use?
rxExec(inData=[sampleInData], outFile=[sampleOutDataIncludingFeatures], transformFunc=computeNonLagFeatures, overwrite=FALSE)
rxDataStep(inData=[sampleInData], outFile=[sampleOutDataIncludingFeatures], transformFunc=computeNonLagFeatures, overwrite=FALSE)
rxDataStep(inData=[sampleInData], outFile=[sampleOutDataIncludingFeatures], transformFunc=computeNonLagFeatures, overwrite=TRUE)
transform(inData=[sampleInData], outFile=[sampleOutDataIncludingFeatures], transformFunc=computeNonLagFeatures, overwrite=FALSE)
You are using rxPredict for a logistic regression model.
You need to obtain prediction standard errors and confidence intervals. Which R code segment should you
use?
glm; covCoef=FALSE; computeStdErr=TRUE; interval='confidence'
rxLogit; covCoef=none; computeStdErr=TRUE; interval='confidence'
glm; covCoef=TRUE; computeStdErr=FALSE; interval='confidence'
rxLogit; covCoef=FALSE; computeStdErr=TRUE; interval='confidence'
rxLogit; covCoef=TRUE; computeStdErr=TRUE; interval='confidence'
