Analysis of the Distributions of 5' Splice Site-Like Sequences in mRNA Precursors by a Position-Tree Method
Sumie Kitamura, Nobuyuki Takahashi, Mineichi Kudo, Masaru Shimbo, Akihiro Tsutsumi
Abstract
Sumie Kitamura, Nobuyuki Takahashi, Mineichi Kudo, Masaru Shimbo, Akihiro Tsutsumi
Abstract
Several methods to predict 5’ splice sites of mRNA precursors (pre-mRNAs) have been proposed [1, 2]. In our previous paper, we proposed a subclass method, and could predict correctly 5’ splice sites from unknown sequences with a discrimination rate of over ninety percent [3]. However, there were some sequences that were not correct 5’ splice site sequences but were similar to 5’ splice site sequences. On the other hand, positional characterization of false positives from computational prediction have been presented [5]. The study indicated that one in every three false positive splice sites, as predicted by programs that use coding information as well as splice signals, was located in the vicinity of real splice sites. In the present study, using a position-tree method, to analyze the distributions of the 5’ splice site-like sequences in pre-mRNAs, we obtained splice site sequences whose lengths were minimal but were sufficient for specifying 5’ splice sites. Then we investigated the distributions of the 5’ splice site-like sequences that were one nucleotide shorter than the obtained splice site sequences.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Several methods to predict 5’ splice sites of mRNA precursors (pre-mRNAs) have been proposed [1, 2]. In our previous paper, we proposed a subclass method, and could predict correctly 5’ splice sites from unknown sequences with a discrimination rate of over ninety percent [3]. However, there were some sequences that were not correct 5’ splice site sequences but were similar to 5’ splice site sequences. On the other hand, positional characterization of false positives from computational prediction have been presented [5]. The study indicated that one in every three false positive splice sites, as predicted by programs that use coding information as well as splice signals, was located in the vicinity of real splice sites. In the present study, using a position-tree method, to analyze the distributions of the 5’ splice site-like sequences in pre-mRNAs, we obtained splice site sequences whose lengths were minimal but were sufficient for specifying 5’ splice sites. Then we investigated the distributions of the 5’ splice site-like sequences that were one nucleotide shorter than the obtained splice site sequences.
Key concepts: splice, Splice site mutation, RNA splicing, False positive paradox, Tree (set theory), Sequence (biology), Coding region, Computational biology