2002Proceedings Genome Informatics Workshop/Genome informaticsRequires access

Improvement in Performance of Homology-Based Prediction of Human Gene Organizations

Osamu Gotoh, Kaoru Sugimori

Open publisher page 0 citations

Abstract

The rst and one of the most important processes in the eld of genome annotation is to nd allgenes encoded in a genome and to identify all variation of transcripts. Although great progress incollection of a large amount of cDNA and EST sequences has been achieved, the goal is not yet close.A promising approach toward solution of this problem is to use comparative analyses of genomes andcDNA or protein sequences [1, 4]. These \homology-based gene-prediction programs perform wellwhen one or more closely related homologous sequence is available. However, accuracy in predictingexons drops sharply with a decrease in similaritybetween the reference and target sequences [5, 7]. Forexample, the performance of GeneWise [1], which is currently most popular and used for constructionof Ensembl annotation, falls behind that of ab initio methods, such as Genescan [2], when amino-acididentities between the reference sequence and translated target are less than ˘60%.It is natural to expect that combination of both approaches of homology-based and ab initiomethods may lead to better performance than that achieved by individual approaches. With thisexpectation, we developed the program aln [3] which adopts a dynamic programming algorithm tooptimizea totalscore derived from severallines of informationon similarityto known cDNA or proteinsequences, intrinsic statistical properties of coding and non-coding parts of genomic sequences, andsignal strengths around translational start sites and intron-exon boundaries. Although aln proved tosigni cantly outperform GeneWise for prediction of nematode genes [3], it had several shortcomingswhen applied to the human genome. We report here our attempt to adapt aln to human genome.

About this research paper

What this paper is about

The rst and one of the most important processes in the eld of genome annotation is to nd allgenes encoded in a genome and to identify all variation of transcripts. Although great progress incollection of a large amount of cDNA and EST sequences has been achieved, the goal is not yet close.A promising approach toward solution of this problem is to use comparative analyses of genomes andcDNA or protein sequences [1, 4]. These \homology-based gene-prediction programs perform wellwhen one or more closely related homologous sequence is available. However, accuracy in predictingexons drops sharply with a decrease in similaritybetween the reference and target sequences [5, 7]. Forexample, the performance of GeneWise [1], which is currently most popular and used for constructionof Ensembl annotation, falls behind that of ab initio methods, such as Genescan [2], when amino-acididentities between the reference sequence and translated target are less than ˘60%.It is natural to expect that combination of both approaches of homology-based and ab initiomethods may lead to better performance than that achieved by individual approaches. With thisexpectation, we developed the program aln [3] which adopts a dynamic programming algorithm tooptimizea totalscore derived from severallines of informationon similarityto known cDNA or proteinsequences, intrinsic statistical properties of coding and non-coding parts of genomic sequences, andsignal strengths around translational start sites and intron-exon boundaries. Although aln proved tosigni cantly outperform GeneWise for prediction of nematode genes [3], it had several shortcomingswhen applied to the human genome. We report here our attempt to adapt aln to human genome.

Why it matters

A significance statement is not available in the OpenAlex record.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

The rst and one of the most important processes in the eld of genome annotation is to nd allgenes encoded in a genome and to identify all variation of transcripts. Although great progress incollection of a large amount of cDNA and EST sequences has been achieved, the goal is not yet close.A promising approach toward solution of this problem is to use comparative analyses of genomes andcDNA or protein sequences [1, 4]. These \homology-based gene-prediction programs perform wellwhen one or more closely related homologous sequence is available. However, accuracy in predictingexons drops sharply with a decrease in similaritybetween the reference and target sequences [5, 7]. Forexample, the performance of GeneWise [1], which is currently most popular and used for constructionof Ensembl annotation, falls behind that of ab initio methods, such as Genescan [2], when amino-acididentities between the reference sequence and translated target are less than ˘60%.It is natural to expect that combination of both approaches of homology-based and ab initiomethods may lead to better performance than that achieved by individual approaches. With thisexpectation, we developed the program aln [3] which adopts a dynamic programming algorithm tooptimizea totalscore derived from severallines of informationon similarityto known cDNA or proteinsequences, intrinsic statistical properties of coding and non-coding parts of genomic sequences, andsignal strengths around translational start sites and intron-exon boundaries. Although aln proved tosigni cantly outperform GeneWise for prediction of nematode genes [3], it had several shortcomingswhen applied to the human genome. We report here our attempt to adapt aln to human genome.

Key concepts: Ensembl, Genome, Gene prediction, Genome project, Computational biology, Human genome, Annotation, Gene

Related papers

Back to paper searchBrowse research topicsOriginal source
Improvement in Performance of Homology-Based Prediction of Human Gene Organizations — Research Paper | ScholarLens