2013International Journal of Computer ApplicationsOpen access

Verification and Validation of MapReduce Program model for Parallel K-Means algorithm on Hadoop Cluster

Amresh Kumar, M. Kiran, Saikat Mukherjee, Ravi Prakash G

Open full text 39 citations

Abstract

With the development of information technology, a large volume of data is growing and getting stored electronically.Thus, the data volumes processing by many applications will routinely cross the petabyte threshold range, in that case it would increase the computational requirements.Efficient processing algorithms and implementation techniques are the key in meeting the scalability and performance requirements in such scientific data analyses.So for the same here, it has been analyzed with the various MapReduce Programs and a parallel clustering algorithm (PKMeans) on Hadoop cluster, using the Concept of MapReduce.Here, in this experiment we have verified and validated various MapReduce applications like wordcount, grep, terasort and parallel K-Means Clustering Algorithm.It has been found that as the number of nodes increases the execution time decreases, but also some of the interesting cases has been found during the experiment and recorded the various performance change and drawn different performance graphs.This experiment is basically a research study of above MapReduce applications and also to verify and validate the MapReduce Program model for Parallel K-Means algorithm on Hadoop Cluster having four nodes.

Open-access reader

About this research paper

What this paper is about

With the development of information technology, a large volume of data is growing and getting stored electronically.Thus, the data volumes processing by many applications will routinely cross the petabyte threshold range, in that case it would increase the computational requirements.Efficient processing algorithms and implementation techniques are the key in meeting the scalability and performance requirements in such scientific data analyses.So for the same here, it has been analyzed with the various MapReduce Programs and a parallel clustering algorithm (PKMeans) on Hadoop cluster, using the Concept of MapReduce.Here, in this experiment we have verified and validated various MapReduce applications like wordcount, grep, terasort and parallel K-Means Clustering Algorithm.It has been found that as the number of nodes increases the execution time decreases, but also some of the interesting cases has been found during the experiment and recorded the various performance change and drawn different performance graphs.This experiment is basically a research study of above MapReduce applications and also to verify and validate the MapReduce Program model for Parallel K-Means algorithm on Hadoop Cluster having four nodes.

Why it matters

OpenAlex reports 39 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

With the development of information technology, a large volume of data is growing and getting stored electronically.Thus, the data volumes processing by many applications will routinely cross the petabyte threshold range, in that case it would increase the computational requirements.Efficient processing algorithms and implementation techniques are the key in meeting the scalability and performance requirements in such scientific data analyses.So for the same here, it has been analyzed with the various MapReduce Programs and a parallel clustering algorithm (PKMeans) on Hadoop cluster, using the Concept of MapReduce.Here, in this experiment we have verified and validated various MapReduce applications like wordcount, grep, terasort and parallel K-Means Clustering Algorithm.It has been found that as the number of nodes increases the execution time decreases, but also some of the interesting cases has been found during the experiment and recorded the various performance change and drawn different performance graphs.This experiment is basically a research study of above MapReduce applications and also to verify and validate the MapReduce Program model for Parallel K-Means algorithm on Hadoop Cluster having four nodes.

Key concepts: Computer science, Cluster (spacecraft), Algorithm, Parallel computing, Data mining, Operating system

Related papers

Back to paper searchBrowse research topicsOriginal source
Verification and Validation of MapReduce Program model for Parallel K-Means algorithm on Hadoop Cluster — Research Paper | ScholarLens