WorksheetsMapReduce Programming Model: Introduction and Core Mechanics
Total questions: 10
Worksheet time: 5mins
Which statement best defines MapReduce in the context of big data processing?
A programming model for processing large datasets across multiple nodes using distributed and parallel computation
A single-node database engine optimized for transactional queries
A visualization framework for plotting distributed systems
A network protocol for fault-tolerant data transfer
In the MapReduce workflow, what is the primary role of the Map phase?
Merge values with the same key to produce the final output
Convert input key/value pairs into intermediate key/value pairs
Schedule tasks across the cluster
Provide automatic replication of input data
Which combination correctly pairs the function with its executor?
Mapper executes Reduce; Reducer executes Map
Mapper executes Map; Reducer executes Reduce
Cluster executes Map; Mapper executes Reduce
Reducer executes Map; Cluster executes Reduce
Given the pseudocode: def map(key, value): for word in value.split(): emit(word, 1). What intermediate key/value pairs are emitted for the input value "data data scale"?
(data, 2), (scale, 1)
(data, 1), (data, 1), (scale, 1)
(data, 3), (scale, 0)
(data scale, 1)
Which feature contributes to MapReduce fault tolerance as described in the material?
Manual retry by developers for failed tasks
Replicating the entire cluster state after each job
Automatically reassigning failed tasks to other machines
Pausing the job until the failed node is repaired
Which advantage is explicitly stated for MapReduce?
It guarantees real-time analytics for all workloads
It handles massive data and provides parallelism automatically
It eliminates the need for key/value pairs
It replaces the need for clusters
In the Reduce phase pseudocode def reduce(key, values): emit(key, sum(values)), what output would be produced for key="word" and values=[1,1,3]?
emit(word, 1)
emit(word, 3)
emit(word, 5)
emit(word, [1,1,3])
Which statement best describes a limitation of MapReduce for small jobs?
It processes small jobs with very low latency
It is not suitable for real-time processing
It has high latency for small jobs
It cannot shuffle and sort intermediate results
According to the quick recap process flow, what immediately follows the Map phase?
Final Output
Shuffle & Sort
Reduce
Intermediate Results
What is the primary purpose of the Map function in MapReduce?
Aggregate values across keys
Filter out all duplicate records
Transform input data into key–value pairs producing intermediate results
Sort keys globally before reducing
