A Comprehensive Survey for Hadoop Distributed File System
Karwan Jameel Merceedi, Nareen Abdulla Sabry
Asian Journal of Research in Computer Science · pp. 46–57 · Published 23 Aug 2021
10.9734/ajrcos/2021/v11i230260Abstract
In the last few days, data and the internet have become increasingly growing, occurring in big data. For these problems, there are many software frameworks used to increase the performance of the distributed system. This software is used for available ample data storage. One of the most beneficial software frameworks used to utilize data in distributed systems is Hadoop. This software creates machine clustering and formatting the work between them. Hadoop consists of two major components: Hadoop Distributed File System (HDFS) and Map Reduce (MR). By Hadoop, we can process, count, and distribute each word in a large file and know the number of affecting for each of them. The HDFS is designed to effectively store and transmit colossal data sets to high-bandwidth user applications. The differences between this and other file systems provided are relevant. HDFS is intended for low-cost hardware and is exceptionally tolerant to defects. Thousands of computers in a vast cluster both have directly associated storage functions and user programmers. The resource scales with demand while being cost-effective in all sizes by distributing storage and calculation through numerous servers. Depending on the above characteristics of the HDFS, many researchers worked in this field trying to enhance the performance and efficiency of the addressed file system to be one of the most active cloud systems. This paper offers an adequate study to review the essential investigations as a trend beneficial for researchers wishing to operate in such a system. The basic ideas and features of the investigated experiments were taken into account to have a robust comparison, which simplifies the selection for future researchers in this subject. According to many authors, this paper will explain what Hadoop is and its architectures, how it works, and its performance analysis in a distributed systems. In addition, assessing each Writing and compare with each other.
Cited by 32
Helmi Tlich, Hanene Chettaoui, T. Hamrouni · Cluster Computing · 2026
R. Gandhi, Jimmy Singla · 2026 5th International Conference on Sentiment Analysis and Deep Learning (ICSADL) · 2026
Serik Aliaskarov, Orazmukhamed Bekmurat, Vassily Serbin · HighTech and Innovation Journal · 2025
Yuan Yao, Yang Zhang · Future Internet · 2025
A. Khaleel, Ahmed Adnan Mohammed Al-Azzawi, Abas Wisam Mahdi Abas · 2025 International Conference on Big Data, Knowledge and Control Systems Engineering (BdKCSE) · 2025
Muhammad Babar, Sarah Kaleem, M. El-Affendi · Engineering, Technology & Applied Science Research · 2025
Youness Filaly, Nisrine Berros, Fatna El Mendili · Results in Engineering · 2025
Rana Ghazali, Douglas G. Down · EAI Endorsed Transactions on Scalable Information Systems · 2025
F. Ishengoma · Technological Sustainability · 2025
Yang Han · Scientific Reports · 2025
Related research
- Service Recommendation Based on Ranking Using Keywords in Hadoop — shares topic coverage
- Improved FTWeightedHashT Apriori Algorithm for Big Data using Hadoop-MapReduce Model — shares topic coverage
- Using Hadoop Technology to Overcome Big Data Problems by Choosing Proposed Cost-efficient Scheduler Algorithm for Heterogeneous Hadoop System (BD3) — shares topic coverage
Article metrics
Real usage data collected on this platform.
0
Page views
0
PDF downloads
0
Outbound clicks
32
Citations
Views by country
Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".
No views recorded yet.
Traffic sources
Referring site, by host.
No traffic recorded yet.
Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.