2010Unpublished venueRequires access

A novel approach to multiple sequence alignment using hadoop data grids

G. Sudha Sadasivam, G. Baktavatchalam

Open publisher page 42 citations

Abstract

Multiple alignment of protein sequences is an essential tool in molecular biology. It aids to determine evolutionary linkage and to predict molecular structures. The factors to be considered while aligning multiple sequences are speed and accuracy of alignment. Dynamic programming algorithms like Needleman-Wunsch and Smith-Waterman produce accurate alignments. But these algorithms are computation intensive and are limited to a small number of short sequences. In this paper we propose a time efficient approach to sequence alignment that produces quality alignment. The dynamic nature of the algorithm coupled with data and computational parallelism of hadoop data grids improves the accuracy and speed of sequence alignment. Further due to the scalability of hadoop framework, the proposed multiple sequence alignment is also highly suited for large scale alignment problems.

About this research paper

What this paper is about

Multiple alignment of protein sequences is an essential tool in molecular biology. It aids to determine evolutionary linkage and to predict molecular structures. The factors to be considered while aligning multiple sequences are speed and accuracy of alignment. Dynamic programming algorithms like Needleman-Wunsch and Smith-Waterman produce accurate alignments. But these algorithms are computation intensive and are limited to a small number of short sequences. In this paper we propose a time efficient approach to sequence alignment that produces quality alignment. The dynamic nature of the algorithm coupled with data and computational parallelism of hadoop data grids improves the accuracy and speed of sequence alignment. Further due to the scalability of hadoop framework, the proposed multiple sequence alignment is also highly suited for large scale alignment problems.

Why it matters

OpenAlex reports 42 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Multiple alignment of protein sequences is an essential tool in molecular biology. It aids to determine evolutionary linkage and to predict molecular structures. The factors to be considered while aligning multiple sequences are speed and accuracy of alignment. Dynamic programming algorithms like Needleman-Wunsch and Smith-Waterman produce accurate alignments. But these algorithms are computation intensive and are limited to a small number of short sequences. In this paper we propose a time efficient approach to sequence alignment that produces quality alignment. The dynamic nature of the algorithm coupled with data and computational parallelism of hadoop data grids improves the accuracy and speed of sequence alignment. Further due to the scalability of hadoop framework, the proposed multiple sequence alignment is also highly suited for large scale alignment problems.

Key concepts: Multiple sequence alignment, Scalability, Computer science, Sequence alignment, Alignment-free sequence analysis, Smith–Waterman algorithm, Sequence (biology), Computation

Related papers

Back to paper searchBrowse research topicsOriginal source
A novel approach to multiple sequence alignment using hadoop data grids — Research Paper | ScholarLens