Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Big Data 7: Spark

Total questions: 91

Worksheet time: 46mins

Name
Class
Date
1.
According to the title slide, Apache Spark is described as which of the following?
a)
An unified analytics engine for large-scale data processing
b)
A key-value store for large-scale data processing
c)
A batch-only scheduler for large-scale jobs
d)
A message queue for large-scale data processing
e)
A distributed file system for large-scale data processing
2.
The slide on iterative MapReduce jobs argues the main performance bottleneck comes from what repeated behavior?
a)
Repeated disk I/O in every iteration
b)
Repeated CPU vectorization in every iteration
c)
Repeated metadata lookups in every iteration
d)
Repeated client-side parsing in every iteration
e)
Repeated in-memory caching in every iteration
3.
In the iterative MapReduce diagram, what happens between consecutive iterations that makes the loop expensive?
a)
Each iteration performs an HDFS write followed by an HDFS read for the next iteration
b)
Each iteration keeps intermediate state only in CPU registers
c)
Each iteration broadcasts the entire dataset to every mapper
d)
Each iteration avoids materialization and only updates a local accumulator
e)
Each iteration reuses the same in-memory RDD without reading or writing
4.
Based on the hardware comparison slide, which resource is presented as having much higher throughput but also higher cost per GB?
a)
RAM
b)
Magnetic disk
c)
Network links between racks
d)
A single CPU core
e)
The file system namespace
5.
The slide contrasts network bandwidth by topology. Which statement matches the figure?
a)
Nodes in another rack are shown with lower bandwidth (0.1 Gb/s) than nodes in the same rack (1 Gb/s)
b)
Nodes in another rack are shown with higher bandwidth (1 Gb/s) than nodes in the same rack (0.1 Gb/s)
c)
Same-rack traffic is shown to be limited to 0.1 Gb/s while cross-rack is 10 Gb/s
d)
Both same-rack and cross-rack links are shown as 10 Gb/s
e)
Network bandwidth is not shown; only disk latency is compared
6.
The "RAM is the new disk" slide emphasizes what principle as key to interactive response times?
a)
Memory-locality
b)
Strict serial execution
c)
Mandatory disk persistence between stages
d)
Single-threaded drivers
e)
Eliminating partitions
7.
In the trend plots on the "RAM is the new disk" slide, which trend is contrasted against the other?
a)
RAM throughput increasing exponentially while disk throughput increases slowly
b)
Disk throughput increasing exponentially while RAM throughput increases slowly
c)
Both RAM and disk throughput increasing at the same exponential rate
d)
Both RAM and disk throughput decreasing over time
e)
Network throughput increasing exponentially while CPU throughput is constant
8.
The unified analytics engine slide lists two workload patterns it supports better. Which pair is shown?
a)
Iterative algorithms and interactive data mining
b)
Single-pass ETL and offline backups
c)
Web serving and DNS resolution
d)
Packet routing and log rotation
e)
Key-value caching and file compaction
9.
In the stack diagram, which set of components appears as libraries above Spark Core?
a)
MLlib, Streaming, SQL, GraphX
b)
HDFS, YARN, HBase, Hive
c)
Kafka, Zookeeper, Flume, Sqoop
d)
Docker, Kubernetes, Mesos, Nomad
e)
Map, Reduce, Combine, Shuffle
10.
In the "Memory instead of disk" comparison, what is the key difference in how intermediate results are handled?
a)
Spark can keep intermediate RDDs in memory, while MapReduce materializes tuples on disk between steps
b)
MapReduce keeps intermediate tuples in memory, while Spark forces every step to write to HDFS
c)
Both Spark and MapReduce require every intermediate step to be a reduce stage
d)
Both systems avoid intermediate materialization entirely
e)
Spark forces all intermediates to disk, but MapReduce can use RAM-only pipelines
11.
In the diagram, where is persistent storage shown as the boundary of the workflow for Spark-style processing?
a)
Input data starts on disk and output data ends on disk, with intermediate RDDs potentially in memory
b)
All RDDs are shown as in disk and only the driver is in memory
c)
The workflow starts in memory and ends in memory with no disk involvement
d)
Only the reduce output is on disk; input is always in memory
e)
Disk is used only for shuffle; input and output are not shown
12.
From the Spark vs Hadoop MapReduce table, which execution model combination is listed?
a)
Hadoop MR: batch; Spark: batch, iterative, streaming
b)
Hadoop MR: streaming only; Spark: batch only
c)
Hadoop MR: batch, iterative; Spark: batch only
d)
Hadoop MR: interactive; Spark: batch only
e)
Hadoop MR: batch, streaming; Spark: map-only
13.
According to the differences table, which statement about supported operations is consistent with the slide?
a)
Spark supports many transformations and actions, including Map and Reduce
b)
Spark supports only Map and Reduce, while Hadoop MR supports arbitrary transformations
c)
Both support only Map and Reduce but Spark uses different names
d)
Hadoop MR supports many transformations and actions, while Spark is restricted to Reduce
e)
Spark is limited to SQL operations and cannot express Map or Reduce
14.
In the benchmark table, which statement correctly matches the data size and elapsed time for Spark 100 TB?
a)
100 TB and 23 mins
b)
100 TB and 72 mins
c)
102.5 TB and 23 mins
d)
1000 TB and 23 mins
e)
102.5 TB and 234 mins
15.
Based on the table, which comparison is accurate about the rate per node?
a)
Spark 1 PB shows 22.5 GB/min per node, much higher than Hadoop World Record at 0.67 GB/min
b)
Hadoop World Record shows 22.5 GB/min per node, much higher than Spark 1 PB at 0.67 GB/min
c)
Spark 100 TB and Hadoop World Record have the same rate per node (0.67 GB/min)
d)
Spark 1 PB shows a lower rate per node than Spark 100 TB because it uses more nodes
e)
Rate per node is not reported in the table
16.
Which description best matches how the slides define an RDD?
a)
A fault-tolerant, parallel data structure with a rich set of operators
b)
A single-machine table that is updated in place
c)
A fine-grained shared-memory object for per-record updates
d)
A disk-only storage format used by Hadoop MR reducers
e)
A network protocol for streaming analytics
17.
The RDD slide contrasts "coarse-grained transformations" with "fine-grained updates". What example is given for coarse-grained transformations?
a)
map, filter, and join applied across many items at once
b)
updating a single record by primary key
c)
changing one array element in place
d)
incrementing a shared counter for one event
e)
patching a file block on disk
18.
According to the RDD slide, which capability is explicitly called out as something users can control to optimize data placement?
a)
Partitioning of intermediate results
b)
HDFS block size for input files
c)
TCP congestion control on each worker
d)
Kernel page cache eviction policy
e)
Database indexing strategy
19.
The partitioning diagram implies what relationship between partitions and parallelism?
a)
More partitions can increase parallelism
b)
More partitions always reduce parallelism
c)
Parallelism is independent of partitions
d)
Partitions matter only for actions, not transformations
e)
Partitions apply only to DataFrames, not RDDs
20.
In the diagram showing RDD partitions across multiple nodes, what does each partition conceptually map to during execution?
a)
A unit that can be processed in parallel (e.g., as separate work across workers)
b)
A single reducer process that must run after all maps finish
c)
A driver-only object that cannot be distributed
d)
A fixed HDFS directory containing reducer outputs
e)
A schema definition for DataFrame columns
21.
A base RDD can be created in two ways in the slides. Which pair is listed?
a)
Parallelize a collection, or read data from an external source (S3, Cassandra, HDFS, etc.)
b)
Run a reduce, or run a join
c)
Create a DataFrame, or create a SQL view
d)
Cache a DataFrame, or unpersist it
e)
Start a cluster, or stop a cluster
22.
The "RDD with 4 partitions" slide most directly illustrates which property of distributed datasets?
a)
A single dataset is split into multiple partitions that can hold different subsets of records
b)
All partitions contain identical copies of every record by default
c)
Each partition must contain exactly the same number of records
d)
Partitions exist only after calling collect()
e)
Partitions are created only when using DataFrames, not RDDs
23.
Why does the slide say parallelize() is not generally used outside prototyping and testing?
a)
It requires the entire dataset to be in memory on one machine before distributing it
b)
It cannot create more than one partition
c)
It always writes the input to HDFS first
d)
It forces eager execution of all transformations
e)
It can only be used from Java and not from Python or Scala
24.
Which API call on the slide demonstrates parallelizing an in-memory collection into an RDD in Python?
a)
sc.parallelize(["fish", "cats", "dogs"])
b)
sc.textFile("/path/to/README.md")
c)
sqlContext.createDataFrame(["fish", "cats", "dogs"])
d)
df.write.partitionBy("fish").saveAsTable("cats")
e)
rdd.reduceByKey(lambda x, y: x + y)
25.
The "Read from Text File" slide shows which SparkContext call to create an RDD from a local text file path?
a)
sc.textFile("/path/to/README.md")
b)
sc.parallelize("/path/to/README.md")
c)
sqlContext.read.text("/path/to/README.md").collect()
d)
df.load("/path/to/README.md").format("text")
e)
sc.hadoopFile("/path/to/README.md").saveAsTextFile()
26.
The same slide notes other sources besides local files. Which set is explicitly mentioned?
a)
HDFS, Cassandra, S3, HBase
b)
PostgreSQL, Redis, Memcached, DNS
c)
Kafka, Zookeeper, Flume, Oozie
d)
Git, SVN, Mercurial, CVS
e)
TensorFlow, PyTorch, ONNX, CUDA
27.
According to the "Operations on Distributed Data" slide, what two categories of operations exist?
a)
Transformations and actions
b)
Maps and reducers
c)
Stages and jobs
d)
Blocks and files
e)
Drivers and workers
28.
The slide states transformations are lazy. What does that imply in this deck?
a)
Transformations are executed only when an action is run
b)
Transformations execute immediately and return concrete results
c)
Actions are lazy and run only when cached
d)
Transformations require disk materialization after each step
e)
Transformations can only be applied to DataFrames, not RDDs
29.
The same slide mentions persisting distributed data. Which storage locations are explicitly listed?
a)
In memory or on disk
b)
In CPU registers only
c)
In the cluster manager metadata store only
d)
In the driver heap only
e)
In the reducer output directory only
30.
In the filter transformation example, what is the relationship between logLinesRDD and errorsRDD?
a)
errorsRDD is created by filtering logLinesRDD with a predicate
b)
logLinesRDD is created by filtering errorsRDD with a predicate
c)
errorsRDD is created by collecting logLinesRDD to the driver
d)
errorsRDD is created by coalescing logLinesRDD into one partition
e)
errorsRDD is created by saving logLinesRDD as a text file
31.
Based on the slide, which statement about the filter transformation is consistent with the illustrated behavior?
a)
It keeps only the records that satisfy the predicate, producing a new RDD
b)
It modifies the input RDD in place by deleting records
c)
It executes only on the driver because it requires collect()
d)
It always changes the number of partitions to one
e)
It is an action because it triggers the DAG immediately
32.
In the collect action diagram, what is the destination of the collected data?
a)
The driver
b)
HDFS output directories
c)
The cluster manager
d)
The block manager only
e)
The SparkContext configuration store
33.
The same slide shows cleanedRDD being produced using coalesce(2). What does that operation suggest in the figure?
a)
Reducing the number of partitions to 2 before collecting
b)
Increasing the number of partitions to 2 for more parallelism
c)
Sorting the records into 2 groups by key
d)
Persisting the RDD in memory twice for safety
e)
Writing two output files regardless of partitions
34.
The "DAG execution" slide visually connects a specific action to triggering execution. Which action is shown as the trigger?
a)
collect()
b)
filter()
c)
coalesce(2)
d)
map()
e)
distinct()
35.
In the "Logical" pipeline diagram, which sequence of operations is shown from input RDD to the final action?
a)
logLinesRDD -> filter -> errorsRDD -> coalesce(2) -> cleanedRDD -> collect()
b)
logLinesRDD -> collect() -> filter -> errorsRDD -> coalesce(2) -> cleanedRDD
c)
errorsRDD -> filter -> logLinesRDD -> coalesce(2) -> cleanedRDD -> collect()
d)
logLinesRDD -> coalesce(2) -> filter -> cleanedRDD -> errorsRDD -> collect()
e)
logLinesRDD -> join -> groupBy -> reduceByKey -> collect()
36.
Which element in the logical diagram indicates that collect() produces a driver-side result rather than another distributed dataset?
a)
The arrow from collect() goes to a single output displayed near the Driver label
b)
The diagram shows a new partitioned RDD labeled collectRDD
c)
collect() is drawn as a transformation between two RDDs
d)
collect() is placed before filter() in the flow
e)
collect() is shown as a disk write labeled HDFS
37.
In the "Physical" diagram, which compute step number is shown closest to the driver output?
a)
1. compute
b)
2. compute
c)
3. compute
d)
4. compute
e)
0. compute
38.
The "Physical" diagram suggests what execution direction when building results from cleanedRDD back to logLinesRDD?
a)
Compute proceeds upward through dependencies, culminating in computing logLinesRDD partitions
b)
Compute proceeds top-down starting from logLinesRDD and never revisits dependencies
c)
Compute skips errorsRDD because it is a logical-only object
d)
Compute runs only on the driver without worker participation
e)
Compute is performed only after saveAsTextFile(), not after collect()
39.
The slide labeled "DAG" shows multiple actions that can be invoked on derived RDDs. Which set of actions is explicitly shown?
a)
collect(), saveAsTextFile(), count()
b)
filter(), map(), coalesce()
c)
distinct(), sort(), describe()
d)
groupBy(), agg(), toDF()
e)
format(), option(), load()
40.
In the DAG illustration, what is the relationship between errorsRDD, cleanedRDD, and errorMsg1RDD?
a)
errorsRDD is derived from logLinesRDD; cleanedRDD and errorMsg1RDD are further derived from errorsRDD
b)
errorMsg1RDD is the base input; errorsRDD and cleanedRDD are independent copies
c)
cleanedRDD is the base input; errorsRDD is created by collecting cleanedRDD
d)
All three are base RDDs created independently from external sources
e)
errorsRDD is created from cleanedRDD using saveAsTextFile()
41.
The "Cache" slide adds cache() into the pipeline. Where is cache() shown being applied?
a)
On logLinesRDD before downstream filters/actions
b)
On the driver after collect()
c)
On HDFS outputs after saveAsTextFile()
d)
On the cluster manager before launching tasks
e)
On the output of describe() only
42.
What change in behavior is the cache slide intended to enable for pipelines with repeated actions?
a)
Reuse previously computed distributed data instead of recomputing lineage each time
b)
Force all transformations to run eagerly and sequentially
c)
Disable fault recovery by removing lineage
d)
Convert every transformation into a shuffle stage
e)
Guarantee that collect() returns results in sorted order
43.
The "Partition >>> Task >>> Partition" slide indicates what mapping in Spark execution?
a)
Each partition is processed by a corresponding task
b)
Each task always processes all partitions
c)
Partitions are created only after tasks complete
d)
Tasks exist only in the driver and do not run on workers
e)
Partitions and tasks are unrelated concepts in Spark
44.
In the same slide, logLinesRDD is labeled as HadoopRDD and errorsRDD as filteredRDD. What does that labeling imply?
a)
errorsRDD is a filtered derivative of an input HadoopRDD
b)
logLinesRDD is derived by filtering errorsRDD
c)
Both RDDs are outputs of a SQL query
d)
errorsRDD is created by writing logLinesRDD to disk
e)
HadoopRDD is a DataFrame type and filteredRDD is a table type
45.
In the RDD lineage diagram, which operation is shown producing ShuffledRDD from two branches?
a)
join()
b)
filter()
c)
map()
d)
distinct()
e)
collect()
46.
The lineage figure separates "Transformations (Lazy)" from "Action (Execute Transformations)". Which call is shown as the action that produces HDFS output?
a)
SparkContext.saveAsHadoopFile()
b)
SparkContext.hadoopFile()
c)
SparkContext.textFile()
d)
SparkContext.parallelize()
e)
SparkContext.cache()
47.
The RDD recap slide indicates where the initial RDD is typically stored. What location is listed?
a)
On disks (e.g., HDFS)
b)
Only in CPU caches
c)
Only in the driver heap
d)
Only inside the cluster manager
e)
Only in the task scheduler queue
48.
According to the same recap slide, fault recovery for RDDs is based on what?
a)
Lineage
b)
A global write-ahead log
c)
Synchronous replication of every partition
d)
Checkpointing every record on each transformation
e)
Manual re-running of failed tasks by the user
49.
Which statement matches how the slides describe a DataFrame in Spark 2.0?
a)
A primary abstraction that is immutable once constructed
b)
A mutable container that supports in-place updates by row index
c)
A disk-only format for reducer outputs
d)
A network stream of events processed by receivers
e)
A cluster manager component for allocating executors
50.
The DataFrame slide lists three construction routes. Which option matches one of them exactly?
a)
Transforming an existing Spark or pandas DataFrame
b)
Calling collect() to create a DataFrame on the driver
c)
Calling saveAsTextFile() to convert text output into a DataFrame
d)
Running reduceByKey() to cast an RDD into a DataFrame
e)
Setting the master parameter to local[K] to enable DataFrames
51.
In the createDataFrame example, what are the two inputs provided to sqlContext.createDataFrame(data, ...)?
a)
A list of tuples and a list of column names
b)
A file path and a schema DDL string
c)
A DataFrame and a partitioning key
d)
A reducer function and a combiner function
e)
A cluster URL and an executor count
52.
The displayed result of createDataFrame is shown as a list of what objects?
a)
Row(...) objects
b)
HDFS blocks
c)
RDD partitions
d)
Stages and tasks
e)
Kafka records
53.
Which statement about DataFrame transformations is consistent with the "Transformations" slide?
a)
They create a new DataFrame from an existing one and are evaluated lazily
b)
They execute immediately and return results to the driver
c)
They require calling collect() after each step
d)
They mutate the source DataFrame in place
e)
They can only be applied to cached DataFrames
54.
On the transformation list slide, which pair is explicitly described as aliases?
a)
where(func) is an alias for filter
b)
select(*cols) is an alias for sort
c)
drop(col) is an alias for distinct
d)
sort(*cols) is an alias for describe
e)
distinct() is an alias for collect()
55.
According to the transformation table, what does distinct() return?
a)
A new DataFrame that contains the distinct rows of the source DataFrame
b)
A new DataFrame that sorts rows by a specified column
c)
A new DataFrame that drops the specified column
d)
A list of unique Row objects on the driver
e)
A count of distinct values for each numeric column
56.
In the example pipeline df1 -> df2 -> df3, what is df2 produced by applying?
a)
distinct()
b)
collect()
c)
count()
d)
describe()
e)
saveAsTable()
57.
The example uses sort("age", ascending=False). What effect is being demonstrated?
a)
Sorting rows by age in descending order
b)
Filtering out rows where age is false
c)
Dropping the age column from the DataFrame
d)
Computing the average age per name
e)
Coalescing partitions down to two
58.
The "Actions" slide describes actions as doing what in Spark?
a)
Causing Spark to execute the recipe to transform the source
b)
Saving a transformation recipe without executing it
c)
Only changing metadata about partitions
d)
Defining new columns without touching data
e)
Replacing the SparkContext master parameter
59.
Which action is explicitly marked with an asterisk on the actions list?
a)
collect()
b)
take(n)
c)
count()
d)
show(n, truncate)
e)
describe(*cols)
60.
According to the actions table, which action is described as an exploratory data analysis function returning summary statistics for numeric columns?
a)
describe(*cols)
b)
distinct()
c)
drop(col)
d)
sort(*cols, **kw)
e)
where(func)
61.
In the usage example, what value does df.count() return for a DataFrame created from two rows?
a)
2
b)
1
c)
0
d)
It prints the rows but returns nothing
e)
It returns a list of Row objects
62.
In the shown output, df.show() is used for what kind of result?
a)
Printing a tabular view of the DataFrame rows
b)
Returning all rows as a Python list
c)
Saving the DataFrame to HDFS
d)
Changing the partition count to match the number of rows
e)
Computing summary statistics for numeric columns
63.
In the caching example, what object is cached first?
a)
linesDF
b)
commentsDF
c)
The result of linesDF.count()
d)
The SparkContext
e)
The cluster manager
64.
The example then computes counts and caches commentsDF. Why is this ordering meaningful given Spark execution semantics described earlier?
a)
Counts are actions that trigger computation; caching allows reuse across subsequent actions
b)
Caching automatically triggers computation even without actions
c)
Caching turns the DataFrame into an RDD and disables actions
d)
Caching makes transformations eager so actions are unnecessary
e)
Caching forces results to be written to HDFS before counting
65.
Which step order best matches the "Spark Programming Routine" slide?
a)
Create DataFrames, lazily transform them, cache some for reuse, then run actions
b)
Run actions first, then cache outputs, then create DataFrames
c)
Only create RDDs; DataFrames are not used
d)
Always cache every DataFrame before any transformation
e)
Skip actions and rely on transformations to produce results automatically
66.
In the described routine, what is the stated purpose of cache()?
a)
Reuse some DataFrames in later computations
b)
Convert DataFrames into SQL tables permanently
c)
Force immediate execution of transformations
d)
Guarantee a single partition for all outputs
e)
Disable lineage tracking for faster failure recovery
67.
The "DataFrames versus RDDs" slide claims DataFrames improve performance through what mechanisms?
a)
Intelligent optimizations and code-generation
b)
Mandatory disk writes after every transformation
c)
Eliminating the need for partitions
d)
Replacing the cluster manager with a single driver thread
e)
Disabling fault tolerance to reduce overhead
68.
According to the same slide, DataFrames are positioned as easier to program primarily for which audience(s)?
a)
Both new users familiar with data frames and existing Spark users
b)
Only Java developers migrating from Hadoop MR
c)
Only users of GraphX
d)
Only Python users running locally
e)
Only cluster administrators configuring masters
69.
In the unified I/O example, which method call sequence belongs to reading data rather than writing it?
a)
read.format("json").option("samplingRatio", "0.1").load(...)
b)
write.format("parquet").mode("append").partitionBy("year").saveAsTable(...)
c)
write.partitionBy("year").option("samplingRatio", "0.1").load(...)
d)
read.mode("append").saveAsTable(...).format("json")
e)
read.saveAsTable(...).partitionBy("year").format("parquet")
70.
In the same example, which builder method is explicitly shown as controlling how existing data is handled during write?
a)
mode("append")
b)
option("samplingRatio", "0.1")
c)
format("json")
d)
load("/Users/spark/data/stuff.json")
e)
toDF("year")
71.
The second I/O slide states that read and write functions create what for doing I/O?
a)
New builders
b)
New partitions
c)
New reducers
d)
New DAG schedulers
e)
New HadoopRDDs
72.
In the shown code pattern, which calls are described as finishing the I/O specification?
a)
load(...), save(...), or saveAsTable(...)
b)
format(...), option(...), or mode(...)
c)
filter(...), where(...), or distinct()
d)
collect(), count(), or show()
e)
join(...), groupBy(...), or agg(...)
73.
On the builder-methods slide, which group is explicitly listed as being specified by builder methods?
a)
format, partitioning, handling of existing data
b)
keys, values, reducers
c)
threads, block manager, cluster manager
d)
schemas, indexes, transactions
e)
mappers, combiners, reducers
74.
Based on the example, which method best fits the "partitioning" category called out in the builder slide?
a)
partitionBy("year")
b)
option("samplingRatio", "0.1")
c)
format("json")
d)
load("/Users/spark/data/stuff.json")
e)
collect()
75.
The I/O slides frame builder usage as a two-step process. Which statement matches that framing?
a)
Builder methods specify options, then a terminal call like load(...) or saveAsTable(...) completes the I/O
b)
A terminal call like collect() completes the I/O, then builder methods specify options
c)
I/O is performed only through actions like count() and describe()
d)
I/O requires explicit reducers and cannot be expressed with builders
e)
I/O is executed automatically when format(...) is called
76.
Which data source types are explicitly listed as supported by DataFrames on the "Data Sources" slide?
a)
JSON, built-in external sources, JDBC
b)
Only HDFS and local text files
c)
Only Cassandra and HBase
d)
Only Kafka topics and sockets
e)
Only Parquet tables
77.
The "Data Sources supported" slide ends with an open-ended phrase. What is it conveying?
a)
Support extends beyond the listed examples ("and more")
b)
Support is limited strictly to the listed examples
c)
Only built-in sources are supported; external connectors are excluded
d)
Only JDBC is supported; file formats are not supported
e)
Support depends only on the master parameter value
78.
The "High-Level Operations" slide lists several common problems solvable with DataFrame functions. Which one is explicitly included?
a)
Joining different data sources
b)
Configuring the cluster manager
c)
Replacing SparkContext with Hadoop FileSystem
d)
Managing HDFS replication factors
e)
Implementing fine-grained record updates
79.
Which operation is given as an example of integrating with another tool for visualization?
a)
Plotting results (e.g., with Pandas)
b)
Plotting results (e.g., with HDFS)
c)
Plotting results (e.g., with YARN)
d)
Plotting results (e.g., with SparkContext)
e)
Plotting results (e.g., with GraphX only)
80.
In the RDD-based average computation shown, what is the purpose of mapping each record to (key, (value, 1)) before reduceByKey?
a)
To carry both a running sum component and a count component per key
b)
To force a global sort of the dataset by key
c)
To eliminate keys so only values are averaged
d)
To make the computation run only on the driver
e)
To convert the RDD into a DataFrame without schema
81.
In the reduceByKey step shown, what is combined across records that share the same key?
a)
Both the numeric value sum and the count are added together
b)
Only the count is added; the numeric value is ignored
c)
Only the numeric value is added; the count is ignored
d)
Keys are concatenated into a longer string
e)
Records are filtered to keep only one value per key
82.
After reduceByKey, the example applies a final map. What does that final map compute?
a)
The average as sum divided by count for each key
b)
The maximum value observed for each key
c)
A distinct list of keys only
d)
A count of partitions per key
e)
A sorted list of all values per key
83.
In the DataFrame version of the average computation, which sequence matches the slide?
a)
groupBy("key").agg(avg("value")).collect()
b)
groupBy("value").agg(avg("key")).collect()
c)
filter("key").distinct().collect()
d)
sort("value").drop("key").collect()
e)
describe("key").count().collect()
84.
The slide contrasts using RDDs vs DataFrames for the same task. What intermediate step bridges from RDDs to DataFrames in the shown snippet?
a)
Converting an RDD of tuples to a DataFrame with toDF("key", "value")
b)
Saving the RDD as text and re-reading it as a DataFrame
c)
Calling collect() on the RDD and reconstructing a DataFrame on the driver
d)
Using SparkContext.master to infer the schema
e)
Calling cache() to automatically change an RDD into a DataFrame
85.
The architecture slide characterizes Spark as what kind of architecture?
a)
A master-worker type architecture
b)
A peer-to-peer architecture with no coordinator
c)
A ring topology where each node forwards tasks
d)
A single-node architecture that scales vertically only
e)
A serverless architecture with no long-lived processes
86.
According to the same slide, what may the master instruct workers to do regarding data access?
a)
Pull data from memory or from hard disk (or another source like S3 or HDFS)
b)
Always pull data only from the driver heap
c)
Always pull data only from reducers
d)
Always pull data only from the cluster manager
e)
Never pull data; workers only receive fully materialized datasets
87.
What is the first object a Spark program creates, according to the "Architecture(2)" slide?
a)
A SparkContext object
b)
A HadoopRDD object
c)
A DAG Scheduler object
d)
A BlockManager object
e)
A JDBC connection pool
88.
The slide states the master parameter for SparkContext determines what?
a)
Which type and size of cluster to use
b)
Which DataFrame columns are numeric
c)
Which transformations are eager
d)
How many reducers exist in Hadoop
e)
Whether RDDs track lineage
89.
In the master parameter table, which value is described as ideal to set K to the number of cores?
a)
local[K]
b)
local
c)
yarn
d)
spark://HOST:PORT
e)
mesos://HOST:PORT
90.
In the "Lifetime of a Job in Spark" diagram, what is the DAG Scheduler shown to do before tasks are launched?
a)
Split the DAG into stages of tasks and submit each stage and its tasks as ready
b)
Store and serve blocks directly to clients
c)
Change the master parameter from local to yarn
d)
Convert DataFrames into RDDs for execution
e)
Write intermediate results to HDFS after each operator
91.
According to the same diagram, which responsibilities are shown under the Worker component?
a)
Execute tasks and store and serve blocks
b)
Split DAG into stages and retry straggler tasks
c)
Allocate executors and schedule stages
d)
Build the operator DAG from rdd.join/groupBy/filter calls
e)
Run only the cluster manager communication loop