Interactions Between Introns and Corresponding Protein Coding Sequences of Ribosomal Protein Genes in C. elegans
Zhao Xiao
Abstract
Zhao Xiao
Abstract
Intron as a kind of non-coding DNA is rich in eukaryote genomes. The functions and involution mechanisms are not very clear besides the splicing. It was thought that introns play a very important role in maintaining and regulating the functional mRNA structure after splicing in the process of mRNA export and translation elongation, etc. Moreover, intron sequence and its corresponding coding sequence are existed interaction or co-evolution relations. The relations between intron sequences and its corresponding coding sequences were studied. For the C. elegans ribosomal protein genes, 85 genes were selected from RPG (http: //www.cbi.pku.edu.cn/chinese/mirrors.html). The intron sequences were divided into first introns, second introns, other introns, short introns, and long introns and the corresponding coding sequences were divided into exons and all protein coding sequences (CDS), then the matching local alignment between introns and the corresponding coding sequences were done with Smith-Waterman local alignment software. The results show that there are really the interaction regions in introns when it is aligned with coding sequences. When intron sequences are aligned with CDSs, the significant interaction regions for the first intron and the other intron are located in about 15%~55% of intron length and it is located in about 30% ~80% of intron length for the second intron. The distribution of interaction regions for short introns is similar to the distribution of the first introns. For long introns, there are two significant interaction regions. The first peak region is located about 15%~30% of intron sequence and the second peak region is located about 54%~78% of intron sequence. When long introns are aligned with exons, there is only one peak region. It is located in about 5%~20% of intron upstream region. When CDS are aligned with every kind of introns, it was found that there are many interaction regions and forbidden regions in CDSs. It was also found that there are two common forbidden regions in the CDSs, they are located at the 10% and 80% of coding sequence. The distribution of interaction regions for the first introns is different from the second introns. When compared the distributions of long introns aligned with CDS and aligned with exons, it can be concluded that the segment of the first peak region are acted on the inner exon segment, the segment of the second peak region are acted mainly on the exon-exon junction regions. Furthermore, there are many peak regions and forbidden regions which are distributed in protein coding sequences. It is speculated that the forbidden regions may be the combined regions of protein complex. In a word, all of the intron sequences besides the 5' end and 3' end correlate closely with their corresponding coding sequences or the two kinds of sequence segments are existed co-evolution relation.
OpenAlex reports 3 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Intron as a kind of non-coding DNA is rich in eukaryote genomes. The functions and involution mechanisms are not very clear besides the splicing. It was thought that introns play a very important role in maintaining and regulating the functional mRNA structure after splicing in the process of mRNA export and translation elongation, etc. Moreover, intron sequence and its corresponding coding sequence are existed interaction or co-evolution relations. The relations between intron sequences and its corresponding coding sequences were studied. For the C. elegans ribosomal protein genes, 85 genes were selected from RPG (http: //www.cbi.pku.edu.cn/chinese/mirrors.html). The intron sequences were divided into first introns, second introns, other introns, short introns, and long introns and the corresponding coding sequences were divided into exons and all protein coding sequences (CDS), then the matching local alignment between introns and the corresponding coding sequences were done with Smith-Waterman local alignment software. The results show that there are really the interaction regions in introns when it is aligned with coding sequences. When intron sequences are aligned with CDSs, the significant interaction regions for the first intron and the other intron are located in about 15%~55% of intron length and it is located in about 30% ~80% of intron length for the second intron. The distribution of interaction regions for short introns is similar to the distribution of the first introns. For long introns, there are two significant interaction regions. The first peak region is located about 15%~30% of intron sequence and the second peak region is located about 54%~78% of intron sequence. When long introns are aligned with exons, there is only one peak region. It is located in about 5%~20% of intron upstream region. When CDS are aligned with every kind of introns, it was found that there are many interaction regions and forbidden regions in CDSs. It was also found that there are two common forbidden regions in the CDSs, they are located at the 10% and 80% of coding sequence. The distribution of interaction regions for the first introns is different from the second introns. When compared the distributions of long introns aligned with CDS and aligned with exons, it can be concluded that the segment of the first peak region are acted on the inner exon segment, the segment of the second peak region are acted mainly on the exon-exon junction regions. Furthermore, there are many peak regions and forbidden regions which are distributed in protein coding sequences. It is speculated that the forbidden regions may be the combined regions of protein complex. In a word, all of the intron sequences besides the 5' end and 3' end correlate closely with their corresponding coding sequences or the two kinds of sequence segments are existed co-evolution relation.
Key concepts: Intron, RNA splicing, Exon, Biology, Group I catalytic intron, Genetics, Coding region, Group II intron