wayground logo

Free Printable Worksheets

Font size

S
M
L
XL
Worksheets

BDA-Quiz-Odd2025-NITA

Total questions: 40

Worksheet time: 21mins

Name
Class
Date
1.

In _________________ computing, the computing task is divided across several computers.   

a)

Parallel

b)

IMC

c)

Distributed

d)

Grid

2.

_____________________ database is provided for distributed data stores where there is  a need for large scale of data storing.

a)

HBase

b)

HDFS

c)

NoSQL

d)

SQL

3.

------ is a tool used for data transfer between Hadoop and relational databases.

a)

Zookeeper

b)

Oozie

c)

Sqoop

d)

Pig

4.

----- element of big data describes the rate at which data is generated, captured and shared.

a)

Volume

b)

Variety

c)

Velocity

d)

Veracity

5.

.......... storage is spread across a cluster of nodes; a single large file could be stored across multiple nodes in the cluster.

a)

HDFS

b)

Mapreduce

c)

Network FileSystem

d)

None of the above

6.

Unstructured data has a faster growth rate than ........... data.

a)

Semi structured

b)

Unstructured

c)

Structured

7.

XML representation is an example of..............

a)

Unstructured data

b)

Structured data

c)

Semi-Structured data

d)

None

8.

Abbreviation of YARN.........

a)

Yet another reason notation

b)

Yet another resource navigator

c)

Yet another resource notes

d)

None

9.

..... is not a tool of Hadoop Eco system.

a)

Mahout

b)

Lucene

c)

HBase

d)

Apache Ant

10.

What are the two core components of HDFS?

a)

Data Node

b)

Resource manager

c)

Application Manager

d)

Name node

11.

Which tool of Hadoop is responsible for performing synchronisation, inter- component based communication, grouping and maintenance.

a)

Oozie

b)

Zookeeper

c)

Apache HBase

d)

Mahout

12.

Which tool of Hadoop Eco system provides machine learning libraries of functionalities such as collaborative filtering, clustering, classification etc.

a)

Apache Spark

b)

Solr, Lucene

c)

Mahout

d)

HIVE

13.

.............. is designed to process web-scale data in the order of hundreds of GB to 100s of TB, even to several PB.

a)

Hadoop

b)

MapReduce

c)

NFS

d)

Multi-processor CPU cores

14.

Name three key technologies used in Big Data.

a)

Excel, PowerPoint, Word

b)

Hadoop, Spark, NoSQL databases

c)

MySQL, PostgreSQL, Oracle

d)

Java, Python, C++

15.

What is the role of YARN in the Hadoop ecosystem?

a)

YARN manages resources and scheduling in the Hadoop ecosystem.

b)

YARN provides a user interface for Hadoop applications.

c)

YARN is responsible for data processing in Hadoop.

d)

YARN is a data storage system in Hadoop.

16.

How does HDFS ensure data reliability and fault tolerance?

a)

HDFS ensures data reliability and fault tolerance by replicating data blocks across multiple nodes.

b)

HDFS uses a single node to store all data blocks.

c)

HDFS automatically deletes corrupted data blocks without replication.

d)

Data blocks are encrypted to ensure reliability.

17.

If you have data on two or more tables in a RDBMS you would need to tell SQL to...?

a)

JOIN them.

b)

MERGE them.

c)

CONNECT them.

d)

LINK them.

18.

Output of the mapper is first written on the local disk for sorting and _________ process.

a)

shuffling

b)

secondary sorting

c)

forking

d)

reducing

19.

A method of storing data within a system that facilitates the collocation of data in various schemata and structural forms.

a)

Data Visualization

b)

Data Lake

c)

Big Data Management

d)

Deep Analytics

20.

The Four Vs of big data are:

a)

volume

variety

visual

velocity

b)

Volume

velocity

variety

veracity

c)

There are actually 5 Vs. Volume, variety, veracity and validity.

21.

The main goal of a Recommendation System is to:

a)

Classify data into categories

b)

Predict user preferences and suggest items

c)

Reduce dataset size

d)

Increase training time

22.

It enables the parallel processing required to perform Big Data tasks

a)

MapReduce

b)

Kafka

c)

Storm

d)

Sqoop

23.

A computing infrastructure similar to the Data Warehouse that provides query and aggregation services for very large volumes of data stored

a)

Hive

b)

Storm

c)

kafka

d)

Pig

24.

It is a computational engine that performs distributed processing in memory on a cluster

a)

Spark

b)

Oozie

c)

Kafka

d)

Storm

25.

________ NameNode is used when the Primary NameNode goes down.

a)

Rack

b)

Data

c)

Secondary

d)

Name

26.

HDFS is implemented in _____________ programming language.

a)

Scala

b)

C++

c)

Java

d)

C

27.

The daemons associated with the MapReduce phase are ________ and task-trackers.

a)

job-tracker

b)

map-tracker

c)

reduce-tracker

d)

all of the mentioned

28.

HDFS works in a __________ fashion

a)

master-worker

b)

master-slave

c)

worker/slave

d)

worker/master

29.

20. There is a single _______ per slave node.

a)

JobTracker

b)

Sqoop

c)

TaskTracker

d)

All the above

30.

7.Which ecosystem project is ideal for use when we have multiple MapReduce and Pig programs to run in a sequence?

a)

Pig

b)

Oozie

c)

Sqoop

d)

Hive

31.

5.Who created the hadoop?

a)

dennis ritchie

b)

james goasling

c)

Doug Cutting

d)

Carlo strozzi

32.

_____________is the slave/worker node and holds the user data in the form of Data Blocks.

a)

NameNode

b)

DataNode

c)

Data block

d)

Replication

33.

17. DataNode is responsible for ________ file operation

a)

Read only

b)

Write only

c)

Read / Write

d)

View only

34.

HDFS is based on ____ file system

a)

IBM

b)

Android

c)

Google

d)

iOS

35.

Columnar storage of Hadoop is

a)

Hive

b)

Pig

c)

Hbase

d)

HDFS

36.

The following are advantages of Big data processing

a)

Cost efficency

b)

Speed

c)

Fault tolerance

d)

All of the above

37.

Which of the following is NOT a common use case for Big Data analytics?

a)

Fraud detection

b)

Recommendation systems

c)

Sentiment analysis

d)

Online transaction processing (OLTP)

38.

Which of the following is NOT a common technique used for data cleansing in Big Data?

a)

Data deduplication

b)

Handling missing values

c)

Normalization

d)

Database indexing

39.

Cassandara is made for

a)

Parallel Processing

b)

SQL Data Base

c)

NoSQL Data Base

d)

RDD

40.

By the time going we mostly save the Data in order of time in this way ..........

a)

Data Ware House- Hadoop - Cloud

b)

Data WareHouse - Cloud- Hadoop

c)

Hadoop- Data Warehouse - Cloud

d)

Hadoop- Cloud-Data Warehouse