2017•2017 International conference of Electronics, Communication and Aerospace Technology (ICECA)Requires access

Mapreduce scheduler: A bird eye view

Ashwani Sheoran, D. Malathi, K. Senthil Kumar

Open publisher page 1 citations

Abstract

MapReduce Scheduling problem has been an active area of research in Computer Science field. MapReduce is a programming model used by Google to process large amount of data in a distributed computing environment. The Apache Hadoop software library is a framework that allows the distributed processing of large data set across clusters of computers using programming model. The programming model automatic handle of node failures hiding the complexity of fault tolerance. It is most widely adopted framework for distributed data processing because of open source and allowing commodity hardware. MapReduce scheduling has become an important factor to achieve high performance in Hadoop cluster. There are many MapReduce scheduling algorithms have been developed for Hadoop. This paper provides an overview of six different scheduling algorithms for MapReduce namely; Scheduling algorithm in Hadoop, First In First Out(FIFO) MapReduce Scheduling algorithm, Fair MapReduce scheduling algorithm, Capacity MapReduce scheduling algorithm, Delay MapReduce scheduling algorithm, MatchMaking MapReduce scheduling algorithm, longest Approximate Time To End(LATE) MapReduce scheduling algorithm. An overview of these techniques is provided through this paper. Advantages and disadvantages of these algorithms are identified. This paper is helpful for the beginners and researchers for understanding the scheduling in big data processing.

About this research paper

What this paper is about

MapReduce Scheduling problem has been an active area of research in Computer Science field. MapReduce is a programming model used by Google to process large amount of data in a distributed computing environment. The Apache Hadoop software library is a framework that allows the distributed processing of large data set across clusters of computers using programming model. The programming model automatic handle of node failures hiding the complexity of fault tolerance. It is most widely adopted framework for distributed data processing because of open source and allowing commodity hardware. MapReduce scheduling has become an important factor to achieve high performance in Hadoop cluster. There are many MapReduce scheduling algorithms have been developed for Hadoop. This paper provides an overview of six different scheduling algorithms for MapReduce namely; Scheduling algorithm in Hadoop, First In First Out(FIFO) MapReduce Scheduling algorithm, Fair MapReduce scheduling algorithm, Capacity MapReduce scheduling algorithm, Delay MapReduce scheduling algorithm, MatchMaking MapReduce scheduling algorithm, longest Approximate Time To End(LATE) MapReduce scheduling algorithm. An overview of these techniques is provided through this paper. Advantages and disadvantages of these algorithms are identified. This paper is helpful for the beginners and researchers for understanding the scheduling in big data processing.

Why it matters

OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

MapReduce Scheduling problem has been an active area of research in Computer Science field. MapReduce is a programming model used by Google to process large amount of data in a distributed computing environment. The Apache Hadoop software library is a framework that allows the distributed processing of large data set across clusters of computers using programming model. The programming model automatic handle of node failures hiding the complexity of fault tolerance. It is most widely adopted framework for distributed data processing because of open source and allowing commodity hardware. MapReduce scheduling has become an important factor to achieve high performance in Hadoop cluster. There are many MapReduce scheduling algorithms have been developed for Hadoop. This paper provides an overview of six different scheduling algorithms for MapReduce namely; Scheduling algorithm in Hadoop, First In First Out(FIFO) MapReduce Scheduling algorithm, Fair MapReduce scheduling algorithm, Capacity MapReduce scheduling algorithm, Delay MapReduce scheduling algorithm, MatchMaking MapReduce scheduling algorithm, longest Approximate Time To End(LATE) MapReduce scheduling algorithm. An overview of these techniques is provided through this paper. Advantages and disadvantages of these algorithms are identified. This paper is helpful for the beginners and researchers for understanding the scheduling in big data processing.

Key concepts: Computer science, Scheduling (production processes), Distributed computing, Fair-share scheduling, Big data, Parallel computing, Programming paradigm, Two-level scheduling

Related papers

Back to paper searchBrowse research topicsOriginal source
Mapreduce scheduler: A bird eye view — Research Paper | ScholarLens