Skip to content
Research Article Open access CC BY 4.0

Using Hadoop Technology to Overcome Big Data Problems by Choosing Proposed Cost-efficient Scheduler Algorithm for Heterogeneous Hadoop System (BD3)

Abou_el_ela Abdou Hussein

Journal of Scientific Research and Reports · pp. 58–84 · Published 28 Nov 2020

10.9734/jsrr/2020/v26i930310

Abstract

Day by day advanced web technologies have led to tremendous growth amount of daily data generated volumes. This mountain of huge and spread data sets leads to phenomenon that called big data which is a collection of massive, heterogeneous, unstructured, enormous and complex data sets. Big Data life cycle could be represented as, Collecting (capture), storing, distribute, manipulating, interpreting, analyzing, investigate and visualizing big data. Traditional techniques as Relational Database Management System (RDBMS) couldn’t handle big data because it has its own limitations, so Advancement in computing architecture is required to handle both the data storage requisites and the weighty processing needed to analyze huge volumes and variety of data economically. There are many technologies manipulating a big data, one of them is hadoop. Hadoop could be understand as an open source spread data processing that is one of the prominent and well known solutions to overcome handling big data problem. Apache Hadoop was based on Google File System and Map Reduce programming paradigm. Through this paper we dived to search for all big data characteristics starting from first three V's that have been extended during time through researches to be more than fifty six V's and making comparisons between researchers to reach to best representation and the precise clarification of all big data V’s characteristics. We highlight the challenges that face big data processing and how to overcome these challenges using Hadoop and its use in processing big data sets as a solution for resolving various problems in a distributed cloud based environment. This paper mainly focuses on different components of hadoop like Hive, Pig, and Hbase, etc. Also we institutes absolute description of Hadoop Pros and cons and improvements to face hadoop problems by choosing proposed Cost-efficient Scheduler Algorithm for heterogeneous Hadoop system.

Big data traditional techniques hadoop hadoop distributed file system MapReduce improvements Scheduler.

Cited by 15

An Information Fusion Technology for Statistical Analysis of Computer Professional Certificates

Jianlai Liao, Bo Li, Zhiyue Ren · 2022 International Conference on 3D Immersion, Interaction and Multi-sensory Experiences (ICDIIME) · 2022

Large-scale Information Management using Hadoop Platform

Vinay Kumar Mishra, Avdesh Singh Pundir · 2022 Second International Conference on Advanced Technologies in Intelligent Control, Environment, Computing & Communication Engineering (ICATIECE) · 2022

Japanese Translation Teaching Reform System in Internet Era on Account of Big Data Algorithm

Fengxian Yin · 2022 International Conference on Education, Network and Information Technology (ICENIT) · 2022

Fast-ICA Algorithm in Industrial Control Network Anomaly Detection System

Yuanyuan Ma · Lecture Notes on Data Engineering and Communications Technologies · 2023

Showing 14 of 15 known citations — external sources report more than can currently be individually listed.

Article metrics

Real usage data collected on this platform.

0

Page views

0

PDF downloads

0

Outbound clicks

15

Citations

Views by country

Approximate, from request IP at view time — not citizenship or institution. Countries with fewer than 5 views are grouped as "Other".

No views recorded yet.

Traffic sources

Referring site, by host.

No traffic recorded yet.

Views and downloads exclude known bots/crawlers. Citations combines this platform's own DOI-resolved index with each external source's own reported total — see Cited by above for individually listed citing works. Last refreshed 0 seconds ago.