Font size
WorksheetsBDA-Quiz-Odd2025-NITA
Total questions: 40
Worksheet time: 21mins
In _________________ computing, the computing task is divided across several computers.
Parallel
IMC
Distributed
Grid
_____________________ database is provided for distributed data stores where there is a need for large scale of data storing.
HBase
HDFS
NoSQL
SQL
------ is a tool used for data transfer between Hadoop and relational databases.
Zookeeper
Oozie
Sqoop
Pig
----- element of big data describes the rate at which data is generated, captured and shared.
Volume
Variety
Velocity
Veracity
.......... storage is spread across a cluster of nodes; a single large file could be stored across multiple nodes in the cluster.
HDFS
Mapreduce
Network FileSystem
None of the above
Unstructured data has a faster growth rate than ........... data.
Semi structured
Unstructured
Structured
XML representation is an example of..............
Unstructured data
Structured data
Semi-Structured data
None
Abbreviation of YARN.........
Yet another reason notation
Yet another resource navigator
Yet another resource notes
None
..... is not a tool of Hadoop Eco system.
Mahout
Lucene
HBase
Apache Ant
What are the two core components of HDFS?
Data Node
Resource manager
Application Manager
Name node
Which tool of Hadoop is responsible for performing synchronisation, inter- component based communication, grouping and maintenance.
Oozie
Zookeeper
Apache HBase
Mahout
Which tool of Hadoop Eco system provides machine learning libraries of functionalities such as collaborative filtering, clustering, classification etc.
Apache Spark
Solr, Lucene
Mahout
HIVE
.............. is designed to process web-scale data in the order of hundreds of GB to 100s of TB, even to several PB.
Hadoop
MapReduce
NFS
Multi-processor CPU cores
Name three key technologies used in Big Data.
Excel, PowerPoint, Word
Hadoop, Spark, NoSQL databases
MySQL, PostgreSQL, Oracle
Java, Python, C++
What is the role of YARN in the Hadoop ecosystem?
YARN manages resources and scheduling in the Hadoop ecosystem.
YARN provides a user interface for Hadoop applications.
YARN is responsible for data processing in Hadoop.
YARN is a data storage system in Hadoop.
How does HDFS ensure data reliability and fault tolerance?
HDFS ensures data reliability and fault tolerance by replicating data blocks across multiple nodes.
HDFS uses a single node to store all data blocks.
HDFS automatically deletes corrupted data blocks without replication.
Data blocks are encrypted to ensure reliability.
If you have data on two or more tables in a RDBMS you would need to tell SQL to...?
JOIN them.
MERGE them.
CONNECT them.
LINK them.
Output of the mapper is first written on the local disk for sorting and _________ process.
shuffling
secondary sorting
forking
reducing
A method of storing data within a system that facilitates the collocation of data in various schemata and structural forms.
Data Visualization
Data Lake
Big Data Management
Deep Analytics
The Four Vs of big data are:
volume
variety
visual
velocity
Volume
velocity
variety
veracity
There are actually 5 Vs. Volume, variety, veracity and validity.
The main goal of a Recommendation System is to:
Classify data into categories
Predict user preferences and suggest items
Reduce dataset size
Increase training time
It enables the parallel processing required to perform Big Data tasks
MapReduce
Kafka
Storm
Sqoop
A computing infrastructure similar to the Data Warehouse that provides query and aggregation services for very large volumes of data stored
Hive
Storm
kafka
Pig
It is a computational engine that performs distributed processing in memory on a cluster
Spark
Oozie
Kafka
Storm
________ NameNode is used when the Primary NameNode goes down.
Rack
Data
Secondary
Name
HDFS is implemented in _____________ programming language.
Scala
C++
Java
C
The daemons associated with the MapReduce phase are ________ and task-trackers.
job-tracker
map-tracker
reduce-tracker
all of the mentioned
HDFS works in a __________ fashion
master-worker
master-slave
worker/slave
worker/master
20. There is a single _______ per slave node.
JobTracker
Sqoop
TaskTracker
All the above
7.Which ecosystem project is ideal for use when we have multiple MapReduce and Pig programs to run in a sequence?
Pig
Oozie
Sqoop
Hive
5.Who created the hadoop?
dennis ritchie
james goasling
Doug Cutting
Carlo strozzi
_____________is the slave/worker node and holds the user data in the form of Data Blocks.
NameNode
DataNode
Data block
Replication
17. DataNode is responsible for ________ file operation
Read only
Write only
Read / Write
View only
HDFS is based on ____ file system
IBM
Android
iOS
Columnar storage of Hadoop is
Hive
Pig
Hbase
HDFS
The following are advantages of Big data processing
Cost efficency
Speed
Fault tolerance
All of the above
Which of the following is NOT a common use case for Big Data analytics?
Fraud detection
Recommendation systems
Sentiment analysis
Online transaction processing (OLTP)
Which of the following is NOT a common technique used for data cleansing in Big Data?
Data deduplication
Handling missing values
Normalization
Database indexing
Cassandara is made for
Parallel Processing
SQL Data Base
NoSQL Data Base
RDD
By the time going we mostly save the Data in order of time in this way ..........
Data Ware House- Hadoop - Cloud
Data WareHouse - Cloud- Hadoop
Hadoop- Data Warehouse - Cloud
Hadoop- Cloud-Data Warehouse
