2020•Nucleic Acids ResearchOpen access

GENCODE 2021

Adam Frankish, Mark E. Diekhans, Irwin Jungreis, Julien Lagarde, Jane E. Loveland, Jonathan M. Mudge, Cristina Sisu, James C. Wright, Joel C. Armstrong, If Habib Ahmed Barnes, Andrew E. Berry, Alexandra Bignell, Carles A. Boix, Sílvia Carbonell Sala, Fiona M Cunningham, Tomás Di Domenico, Sarah M. Donaldson, Ian T. Fiddes, Carlos García Girón, José M. González, Tiago Grego, Matthew Philip Hardy, Thibaut Hourlier, Kevin Howe, Toby Hunt, Osagie G. Izuogu, Rory Johnson, Fergal J. Martin, Laura Martínez Gómez, Shamika Mohanan, Paul Muir, Fábio C. P. Navarro, Anne Parker, Baikang Pei, Fernando Campo del Pozo, Ferriol Calvet, Magali Ruffier, Bianca M. Schmitt, Eloise Stapleton, Marie‐Marthe Suner, Irina Nikolaevna Sycheva, Barbara Uszczyńska-Ratajczak, Maxim Y. Wolf, Jinrui Xu, Yucheng Yang, Andrew D. Yates, Daniel R. Zerbino, Yan Zhang, Jyoti Sharma Choudhary, Mark Gerstein, Roderic Guigó, Tim Hubbard, Manolis Kellis, Benedict Paten, Michael L. Tress, Paul Flicek

Open full text 1,516 citations

Abstract

The GENCODE project annotates human and mouse genes and transcripts supported by experimental data with high accuracy, providing a foundational resource that supports genome biology and clinical genomics. GENCODE annotation processes make use of primary data and bioinformatic tools and analysis generated both within the consortium and externally to support the creation of transcript structures and the determination of their function. Here, we present improvements to our annotation infrastructure, bioinformatics tools, and analysis, and the advances they support in the annotation of the human and mouse genomes including: the completion of first pass manual annotation for the mouse reference genome; targeted improvements to the annotation of genes associated with SARS-CoV-2 infection; collaborative projects to achieve convergence across reference annotation databases for the annotation of human and mouse protein-coding genes; and the first GENCODE manually supervised automated annotation of lncRNAs. Our annotation is accessible via Ensembl, the UCSC Genome Browser and https://www.gencodegenes.org.

Open-access reader

About this research paper

What this paper is about

The GENCODE project annotates human and mouse genes and transcripts supported by experimental data with high accuracy, providing a foundational resource that supports genome biology and clinical genomics. GENCODE annotation processes make use of primary data and bioinformatic tools and analysis generated both within the consortium and externally to support the creation of transcript structures and the determination of their function. Here, we present improvements to our annotation infrastructure, bioinformatics tools, and analysis, and the advances they support in the annotation of the human and mouse genomes including: the completion of first pass manual annotation for the mouse reference genome; targeted improvements to the annotation of genes associated with SARS-CoV-2 infection; collaborative projects to achieve convergence across reference annotation databases for the annotation of human and mouse protein-coding genes; and the first GENCODE manually supervised automated annotation of lncRNAs. Our annotation is accessible via Ensembl, the UCSC Genome Browser and https://www.gencodegenes.org.

Why it matters

OpenAlex reports 1516 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

The GENCODE project annotates human and mouse genes and transcripts supported by experimental data with high accuracy, providing a foundational resource that supports genome biology and clinical genomics. GENCODE annotation processes make use of primary data and bioinformatic tools and analysis generated both within the consortium and externally to support the creation of transcript structures and the determination of their function. Here, we present improvements to our annotation infrastructure, bioinformatics tools, and analysis, and the advances they support in the annotation of the human and mouse genomes including: the completion of first pass manual annotation for the mouse reference genome; targeted improvements to the annotation of genes associated with SARS-CoV-2 infection; collaborative projects to achieve convergence across reference annotation databases for the annotation of human and mouse protein-coding genes; and the first GENCODE manually supervised automated annotation of lncRNAs. Our annotation is accessible via Ensembl, the UCSC Genome Browser and https://www.gencodegenes.org.

Key concepts: Ensembl, Annotation, Biology, Genome project, Genome, Genome browser, Gene Annotation, Computational biology

Related papers

Back to paper searchBrowse research topicsOriginal source
GENCODE 2021 — Research Paper | ScholarLens