wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

Introduction to Big Data

Total questions: 50

Worksheet time: 38mins

Name
Class
Date
1.

Big Data refers to data that is:

a)

Small and structured

b)

Difficult to store using traditional databases

c)

Only numerical

d)

Stored only in spreadsheets

2.

Which of the following best defines Big Data?

a)

Data stored in files

b)

Large volume of data generated at high speed

c)

Data in databases

d)

Backup data

3.

Big Data is mainly generated from:

a)

Typewriters

b)

Social media, sensors, and transactions

c)

Paper records

d)

Calculators

4.

Which era led to the rapid growth of Big Data?

a)

Mechanical era

b)

Digital and internet era

c)

Industrial revolution

d)

Pre-computer era

5.

Big Data processing generally requires:

a)

Single computer systems

b)

Distributed computing systems

c)

Manual processing

d)

Local storage only

6.

Which of the following is NOT a characteristic of Big Data?

a)

Volume

b)

Velocity

c)

Variety

d)

Validity

7.

Volume in Big Data refers to:

a)

Speed of data generation

b)

Size of data

c)

Accuracy of data

d)

Data format

8.

Velocity refers to:

a)

Data quality

b)

Data storage method

c)

Speed of data generation and processing

9.

Variety in Big Data means:

a)

Only text data

b)

Structured data only

c)

Different data formats like text, images, videos

d)

Duplicate data

10.

Veracity refers to:

a)

Data size

b)

Data trustworthiness

c)

Data speed

d)

Data type

11.

Structured data is best stored in:

a)

Text files

b)

Relational databases

c)

Images

d)

Videos

12.

Which is an example of unstructured data?

a)

Tables

b)

CSV files

c)

Images

d)

Excel sheets

13.

Semi-structured data example is:

a)

Relational table

b)

XML or JSON

c)

Image file

d)

Video file

14.

Data stored in rows and columns is called:

a)

Unstructured

b)

Semi-structured

c)

Structured

d)

Raw data

15.

Social media posts are an example of:

a)

Structured data

b)

Semi-structured data

c)

Unstructured data

d)

Metadata

16.

Big Data is widely used in:

a)

Healthcare

b)

Finance

c)

Education

17.

Which application uses Big Data for recommendation systems?

a)

Banking

b)

E-commerce

c)

Agriculture

d)

Manufacturing

18.

Big Data helps in healthcare mainly for:

a)

Entertainment

b)

Disease prediction and diagnosis

c)

Gaming

d)

Networking

19.

Which sector uses Big Data for fraud detection?

a)

Banking

b)

Education

c)

Sports

d)

Tourism

20.

Traffic management systems use Big Data for:

a)

File storage

b)

Route optimization

c)

Database backup

d)

Image editing

21.

Big Data helps organizations to:

a)

Increase paperwork

b)

Make data-driven decisions

c)

Reduce data

d)

Avoid automation

22.

One major benefit of Big Data analytics is:

a)

Increased cost

b)

Faster decision making

c)

Data loss

d)

Manual reporting

23.

Big Data improves business performance by:

a)

Ignoring customer data

b)

Predicting trends and behavior

c)

Reducing storage

d)

Limiting data access

24.

Which of the following is an advantage of Big Data?

a)

Better customer insights

b)

High hardware cost only

c)

Complex manual processing

25.

Big Data is important because it:

a)

Replaces computers

b)

Handles large and complex datasets

c)

Reduces internet usage

d)

Eliminates databases

26.

Serialization is the process of:

a)

Encrypting data

b)

Converting objects into a byte stream

c)

Deleting objects

d)

Compressing files

27.

Serialization is mainly used for:

a)

Object storage and transmission

b)

Data deletion

c)

Data visualization

d)

Image processing

28.

In Java, which interface is used for serialization?

a)

Cloneable

b)

Serializable

c)

Runnable

d)

Comparable

29.

Deserialization means:

a)

Converting object to byte stream

b)

Converting byte stream back to object

c)

Encrypting data

d)

Deleting data

30.

Serialization is useful in:

a)

Distributed systems

b)

Single-user systems only

c)

Offline applications

d)

Manual processing

31.

Wrapper classes are used to:

a)

Convert objects to files

b)

Convert primitive data types into objects

c)

Delete data

d)

Compress data

32.

Which is a wrapper class for int?

a)

Integer

b)

Int

c)

Number

33.

Wrapper classes belong to which package in Java?

a)

java.io

b)

java.lang

c)

java.util

d)

java.sql

34.

Which of the following is NOT a wrapper class?

a)

Double

b)

Character

c)

String

d)

Boolean

35.

Wrapper classes support:

a)

Object-oriented features

b)

Only primitive operations

c)

Hardware execution

d)

Compilation only

36.

Scaling out means:

a)

Increasing CPU speed

b)

Adding more machines

c)

Increasing memory of one system

d)

Reducing nodes

37.

Distributed File System stores data:

a)

On a single machine

b)

Across multiple machines

c)

Only in memory

d)

On external drives

38.

Scaling out is preferred over scaling up because it:

a)

Is cheaper and flexible

b)

Uses one system

c)

Reduces availability

d)

Limits growth

39.

A Distributed File System improves:

a)

Data availability

b)

Single-user access

c)

Manual processing

d)

Local storage

40.

Which is a feature of Distributed File Systems?

a)

Centralized storage only

b)

Fault tolerance

c)

Low scalability

41.

GFS was developed by:

a)

Microsoft

b)

Google

c)

Amazon

d)

IBM

42.

GFS is designed mainly for:

a)

Small files

b)

Large data-intensive applications

c)

Desktop applications

d)

Mobile apps

43.

Hadoop Distributed File System (HDFS) is inspired by:

a)

NTFS

b)

FAT

c)

GFS

d)

EXT4

44.

Hadoop is mainly used for:

a)

Image editing

b)

Big Data storage and processing

c)

Gaming

d)

Web browsing

45.

Hadoop ecosystem includes:

a)

HDFS, MapReduce, YARN

b)

Word, Excel

c)

HTML, CSS

d)

C, C++

46.

HDFS stores data in the form of:

a)

Tables

b)

Blocks

c)

Rows

d)

Objects

47.

Hadoop is best suited for:

a)

Small datasets

b)

Large-scale data processing

c)

Real-time gaming

d)

Image design

48.

Which Hadoop component manages resources?

a)

HDFS

b)

MapReduce

c)

YARN

49.

Hadoop works on the principle of:

a)

Centralized computing

b)

Distributed computing

c)

Sequential processing

d)

Manual processing

50.

The main advantage of Hadoop ecosystem is:

a)

Low fault tolerance

b)

High cost

c)

Scalability and fault tolerance

d)

Limited storage