Skip to content
Research Article Open access CC BY 4.0

A Comprehensive Survey for Hadoop Distributed File System

Karwan Jameel Merceedi, Nareen Abdulla Sabry

Asian Journal of Research in Computer Science · pp. 46–57 · Published 23 Aug 2021

10.9734/ajrcos/2021/v11i230260

Abstract

In the last few days, data and the internet have become increasingly growing, occurring in big data. For these problems, there are many software frameworks used to increase the performance of the distributed system. This software is used for available ample data storage. One of the most beneficial software frameworks used to utilize data in distributed systems is Hadoop. This software creates machine clustering and formatting the work between them. Hadoop consists of two major components: Hadoop Distributed File System (HDFS) and Map Reduce (MR). By Hadoop, we can process, count, and distribute each word in a large file and know the number of affecting for each of them. The HDFS is designed to effectively store and transmit colossal data sets to high-bandwidth user applications. The differences between this and other file systems provided are relevant. HDFS is intended for low-cost hardware and is exceptionally tolerant to defects. Thousands of computers in a vast cluster both have directly associated storage functions and user programmers. The resource scales with demand while being cost-effective in all sizes by distributing storage and calculation through numerous servers. Depending on the above characteristics of the HDFS, many researchers worked in this field trying to enhance the performance and efficiency of the addressed file system to be one of the most active cloud systems. This paper offers an adequate study to review the essential investigations as a trend beneficial for researchers wishing to operate in such a system. The basic ideas and features of the investigated experiments were taken into account to have a robust comparison, which simplifies the selection for future researchers in this subject. According to many authors, this paper will explain what Hadoop is and its architectures, how it works, and its performance analysis in a distributed systems. In addition, assessing each Writing and compare with each other.

Hadoop HDFS distributed file system

Cited by 32

Design of Multi-priority Message Sending Method Based on Bandwidth State Control

Zhong-Jun-Dong-Zhong-Jun Dong · Journal of Science & Technology · 2023

Quantification and analysis of performance fluctuation in distributed file system

Yuchen Zhang, Xiao Zhang, Jin-Yang Yu · Cluster Computing · 2023

Survey of Distributed Computing Frameworks for Supporting Big Data Analysis

Xudong Sun, Yulin He, Dingming Wu · Big Data Mining and Analytics · 2023

Research of the methods of creating content aggregation systems

Denis Aleksandrovich Kiryanov · Программные системы и вычислительные методы · 2022

Research on Mass Image Data Storage Method for Data Center

Sen Pan, Jing Jiang, Hongbin Qiu · Smart Innovation, Systems and Technologies · 2023

Article metrics

Real usage data collected on this platform.

0

Page views

0

PDF downloads

0

Outbound clicks

32

Citations

Views by country

Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".

No views recorded yet.

Traffic sources

Referring site, by host.

No traffic recorded yet.

Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.