INFORMATION ABOUT SECONDARY STRUCTURE IMPROVES QUALITY OF PROTEIN ALIGNMENT
I. I. Litvinov, A A Mironov, Alexei V. Finkelstein, Mikhail Roytberg
Abstract
I. I. Litvinov, A A Mironov, Alexei V. Finkelstein, Mikhail Roytberg
Abstract
Summary Motivation: The Smith-Waterman (SW) alignment algorithm is known as the most accurate algorithm for pair wise alignment of amino acid sequences.. It means, that SW alignments are more similar to alignments of corresponding 3D-structures, than FASTA, BLAST, etc. alignments. But even SW algorithm is unable to restore alignment of proteins’ 3D-structures if the sequence identity is less than 30 % (“twilight zone”). Our goal is to design a new alignment method, which is significantly more accurate, than SW algorithm. Results: We propose to modify SW alignment score to take into account protein secondary structure. We give bonus for alignment of residues belonging to the regions of same secondary structure type. We have shown that alignments maximizing the improved score are much more accurate, than SW alignments (57 % accuracy vs. 31 % for the twilight zone sequence identity; both experimentally determined and theoretically predicted secondary structure can be used). The dynamic programming algorithm to find the optimal secondary structure alignment was designed and implemented as C++ program STRUSWER.
OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Summary Motivation: The Smith-Waterman (SW) alignment algorithm is known as the most accurate algorithm for pair wise alignment of amino acid sequences.. It means, that SW alignments are more similar to alignments of corresponding 3D-structures, than FASTA, BLAST, etc. alignments. But even SW algorithm is unable to restore alignment of proteins’ 3D-structures if the sequence identity is less than 30 % (“twilight zone”). Our goal is to design a new alignment method, which is significantly more accurate, than SW algorithm. Results: We propose to modify SW alignment score to take into account protein secondary structure. We give bonus for alignment of residues belonging to the regions of same secondary structure type. We have shown that alignments maximizing the improved score are much more accurate, than SW alignments (57 % accuracy vs. 31 % for the twilight zone sequence identity; both experimentally determined and theoretically predicted secondary structure can be used). The dynamic programming algorithm to find the optimal secondary structure alignment was designed and implemented as C++ program STRUSWER.
Key concepts: Multiple sequence alignment, Structural alignment, Protein secondary structure, Sequence alignment, Alignment-free sequence analysis, Smith–Waterman algorithm, Dynamic programming, Computer science