Wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

MapReduce Programming Model: Introduction and Core Mechanics

Total questions: 10

Worksheet time: 5mins

Name
Class
Date
1.

Which statement best defines MapReduce in the context of big data processing?

a)

A programming model for processing large datasets across multiple nodes using distributed and parallel computation

b)

A single-node database engine optimized for transactional queries

c)

A visualization framework for plotting distributed systems

d)

A network protocol for fault-tolerant data transfer

2.

In the MapReduce workflow, what is the primary role of the Map phase?

a)

Merge values with the same key to produce the final output

b)

Convert input key/value pairs into intermediate key/value pairs

c)

Schedule tasks across the cluster

d)

Provide automatic replication of input data

3.

Which combination correctly pairs the function with its executor?

a)

Mapper executes Reduce; Reducer executes Map

b)

Mapper executes Map; Reducer executes Reduce

c)

Cluster executes Map; Mapper executes Reduce

d)

Reducer executes Map; Cluster executes Reduce

4.

Given the pseudocode: def map(key, value): for word in value.split(): emit(word, 1). What intermediate key/value pairs are emitted for the input value "data data scale"?

a)

(data, 2), (scale, 1)

b)

(data, 1), (data, 1), (scale, 1)

c)

(data, 3), (scale, 0)

d)

(data scale, 1)

5.

Which feature contributes to MapReduce fault tolerance as described in the material?

a)

Manual retry by developers for failed tasks

b)

Replicating the entire cluster state after each job

c)

Automatically reassigning failed tasks to other machines

d)

Pausing the job until the failed node is repaired

6.

Which advantage is explicitly stated for MapReduce?

a)

It guarantees real-time analytics for all workloads

b)

It handles massive data and provides parallelism automatically

c)

It eliminates the need for key/value pairs

d)

It replaces the need for clusters

7.

In the Reduce phase pseudocode def reduce(key, values): emit(key, sum(values)), what output would be produced for key="word" and values=[1,1,3]?

a)

emit(word, 1)

b)

emit(word, 3)

c)

emit(word, 5)

d)

emit(word, [1,1,3])

8.

Which statement best describes a limitation of MapReduce for small jobs?

a)

It processes small jobs with very low latency

b)

It is not suitable for real-time processing

c)

It has high latency for small jobs

d)

It cannot shuffle and sort intermediate results

9.

According to the quick recap process flow, what immediately follows the Map phase?

a)

Final Output

b)

Shuffle & Sort

c)

Reduce

d)

Intermediate Results

10.

What is the primary purpose of the Map function in MapReduce?

a)

Aggregate values across keys

b)

Filter out all duplicate records

c)

Transform input data into key–value pairs producing intermediate results

d)

Sort keys globally before reducing