Skip to content
Research Article Open access CC BY 4.0

Using Hadoop Technology to Overcome Big Data Problems by Choosing Proposed Cost-efficient Scheduler Algorithm for Heterogeneous Hadoop System (BD3)

Abou_el_ela Abdou Hussein

Journal of Scientific Research and Reports · pp. 58–84 · Published 28 Nov 2020

10.9734/jsrr/2020/v26i930310

Abstract

Day by day advanced web technologies have led to tremendous growth amount of daily data generated volumes. This mountain of huge and spread data sets leads to phenomenon that called big data which is a collection of massive, heterogeneous, unstructured, enormous and complex data sets. Big Data life cycle could be represented as, Collecting (capture), storing, distribute, manipulating, interpreting, analyzing, investigate and visualizing big data. Traditional techniques as Relational Database Management System (RDBMS) couldn’t handle big data because it has its own limitations, so Advancement in computing architecture is required to handle both the data storage requisites and the weighty processing needed to analyze huge volumes and variety of data economically. There are many technologies manipulating a big data, one of them is hadoop. Hadoop could be understand as an open source spread data processing that is one of the prominent and well known solutions to overcome handling big data problem. Apache Hadoop was based on Google File System and Map Reduce programming paradigm. Through this paper we dived to search for all big data characteristics starting from first three V's that have been extended during time through researches to be more than fifty six V's and making comparisons between researchers to reach to best representation and the precise clarification of all big data V’s characteristics. We highlight the challenges that face big data processing and how to overcome these challenges using Hadoop and its use in processing big data sets as a solution for resolving various problems in a distributed cloud based environment. This paper mainly focuses on different components of hadoop like Hive, Pig, and Hbase, etc. Also we institutes absolute description of Hadoop Pros and cons and improvements to face hadoop problems by choosing proposed Cost-efficient Scheduler Algorithm for heterogeneous Hadoop system.

Big data traditional techniques hadoop hadoop distributed file system MapReduce improvements Scheduler.

Cited by 15

Advanced NLP Based Entity Key Phrase Extraction and Text-Based Similarity Measures in Hadoop Environment

Adyasha Dash, Aryan Mohanty, Sohini Ghosh · 2023 6th International Conference on Information Systems and Computer Networks (ISCON) · 2023

Load Balancing Algorithms for Hadoop Cluster in Unbalanced Environment

Weiyu Fu, Lixia Wang · Computational Intelligence and Neuroscience · 2022

Comparative Study of Leveraging Big Data Processing Techniques for Sentiment Analysis

Chris-Ern-Zer Wong, Lee-Yeng Ong, Meng-Chew Leow · 2023 10th International Conference on Electrical Engineering, Computer Science and Informatics (EECSI) · 2023

Domain and Challenges of Big Data and Archaeological Photogrammetry With Blockchain

Omer Aziz, Muhammad Shoaib Farooq, Adel Khelifi · IEEE Access · 2022

Raif-Semantics: A Robust Automated Interlinking Framework for Semantic Web Using MapReduce and Multi-node Data Processing

Shweta S. Aladakatti, S. Senthil Kumar · Journal of Interconnection Networks · 2022

Application Analysis of Hadoop in Big Data Processing

Zhiming Zheng · 2021 International Conference on Aviation Safety and Information Technology · 2021

Improving Big Data Visualization and Association Rule Mining through Splunk Database Gap Analysis

S.M Moshiuzzaman Shatil, Ahsan Habib, Tahmid Enam Shrestha · Proceedings of the 3rd International Conference on Computing Advancements · 2024

Showing 14 of 15 known citations — external sources report more than can currently be individually listed.

Article metrics

Real usage data collected on this platform.

0

Page views

0

PDF downloads

0

Outbound clicks

15

Citations

Views by country

Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".

No views recorded yet.

Traffic sources

Referring site, by host.

No traffic recorded yet.

Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.