Protein structure prediction: sequence to structure alignment generation and atomic potential development
Jian Qiu, Ron Elber
Abstract
Jian Qiu, Ron Elber
Abstract
In template-based modeling of protein structures, the generation of the alignment between the target and the template is a critical step that significantly affects the accuracy of the final model. In this study, I propose an alignment algorithm SSALN that learns substitution matrices and position-specific gap penalties from a database of structurally aligned protein pairs. In addition to the amino acid sequence information, secondary structure and solvent accessibility information of a position are used to derive substitution scores and position-specific gap penalties. In the CASP5 and CASP6 test sets, SSALN generates more accurate alignments than sequence alignment methods. LOOPP server prediction based on an SSALN alignment is ranked the best for target T0280_1 in CASP6. SSALN is also compared with several threading methods and sequence alignment methods on the ProSup benchmark. SSALN has the highest alignment accuracy among the methods compared. On the Fischer's benchmark, SSALN performs better than CLUSTALW and GenTHREADER, and generates more alignments with >50%, >60% or >70% accuracy than FUGUE. In protein structure prediction, multiple models are often generated and a potential is used to distinguish native-like (homologous) models from misfolded (decoy) models. Here I present a procedure to develop atomic potentials for recognition of protein folds. The potentials are based on pairwise interactions (close in distance in the structure) between heavy (non-hydrogen) atoms. One or three distance steps are used to describe the range of interactions between a pair. Training is carried out with the mathematical programming approach on the decoy sets of Baker, of Levitt, and of our own design. Recognition is required not only for decoy-native structural pairs but also for pairs of decoy and homologous structures. Performance is tested on the template-based modeling test sets of CASP5 targets, and on two ab initio decoy test sets from the Skolnick's laboratory and from the Moult's laboratory. The newly derived potentials have significant recognition capacity that is comparable to the best atomic potential compared using a significantly smaller number of parameters.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
In template-based modeling of protein structures, the generation of the alignment between the target and the template is a critical step that significantly affects the accuracy of the final model. In this study, I propose an alignment algorithm SSALN that learns substitution matrices and position-specific gap penalties from a database of structurally aligned protein pairs. In addition to the amino acid sequence information, secondary structure and solvent accessibility information of a position are used to derive substitution scores and position-specific gap penalties. In the CASP5 and CASP6 test sets, SSALN generates more accurate alignments than sequence alignment methods. LOOPP server prediction based on an SSALN alignment is ranked the best for target T0280_1 in CASP6. SSALN is also compared with several threading methods and sequence alignment methods on the ProSup benchmark. SSALN has the highest alignment accuracy among the methods compared. On the Fischer's benchmark, SSALN performs better than CLUSTALW and GenTHREADER, and generates more alignments with >50%, >60% or >70% accuracy than FUGUE. In protein structure prediction, multiple models are often generated and a potential is used to distinguish native-like (homologous) models from misfolded (decoy) models. Here I present a procedure to develop atomic potentials for recognition of protein folds. The potentials are based on pairwise interactions (close in distance in the structure) between heavy (non-hydrogen) atoms. One or three distance steps are used to describe the range of interactions between a pair. Training is carried out with the mathematical programming approach on the decoy sets of Baker, of Levitt, and of our own design. Recognition is required not only for decoy-native structural pairs but also for pairs of decoy and homologous structures. Performance is tested on the template-based modeling test sets of CASP5 targets, and on two ab initio decoy test sets from the Skolnick's laboratory and from the Moult's laboratory. The newly derived potentials have significant recognition capacity that is comparable to the best atomic potential compared using a significantly smaller number of parameters.
Key concepts: Structural alignment, Protein structure prediction, Alignment-free sequence analysis, Multiple sequence alignment, Loop modeling, Pairwise comparison, Threading (protein sequence), Sequence alignment