GENCODE

GENCODE: An Overview

GENCODE is a pivotal scientific project in the realm of genome research, integral to the larger ENCODE (ENCyclopedia Of DNA Elements) initiative. Initially established during the pilot phase of the ENCODE project, GENCODE aims to meticulously identify and map all protein-coding genes within designated ENCODE regions, which encompass approximately 1% of the human genome. The success of the initial phases has propelled GENCODE into a broader mission: to create an “Encyclopedia of genes and gene variants.” This ambitious goal encompasses the comprehensive annotation of protein-coding loci, non-coding loci with transcript evidence, and pseudogenes.

Current Progress in GENCODE

As GENCODE advances towards its objectives, it is currently operating in Phase 2 of the project. The most recent release of the Human gene set annotations is Gencode 36, which was frozen in December 2020. This release utilizes the latest GRCh38 human reference genome assembly, ensuring that researchers have access to the most up-to-date genomic data available. Furthermore, for mouse gene set annotations, Gencode M25 was also released with a freeze date in December 2020. Since September 2009, GENCODE has served as the human gene set employed by the Ensembl project, with each new GENCODE release corresponding directly to an Ensembl release.

A Brief History of GENCODE

The inception of GENCODE dates back to September 2003 when the project was structured into three distinct phases: Pilot, Technology Development, and Production. The primary aim during the pilot stage was to conduct an in-depth investigation—both computationally and experimentally—of 44 regions totaling 30 megabases of sequence that represent around 1% of the human genome. It was during this stage that the GENCODE consortium was formed to systematically identify and map all protein-coding genes within these ENCODE regions.

The first significant milestone occurred in April 2005 when Release 1 of the annotation for these regions was completed. This initial dataset included 416 known loci, along with several novel coding DNA sequence (CDS) loci and putative loci. A subsequent update (Release 02) followed in October 2005, refining these initial findings based on experimental validation results.

In June 2007, key findings from the pilot project were published, confirming the project’s capability to develop a viable platform for characterizing functional elements within the human genome. This laid essential groundwork for future genome-wide studies. By October 2007, new funding facilitated a scaling up of the ENCODE Project to encompass a comprehensive analysis of the entire human genome.

Significant progress continued through subsequent years, with major releases such as GENCODE Release 7 published in September 2012. By 2018, innovations like the CRISPR/Cas9 track were added to assist researchers in identifying potential binding sites for genome editing applications.

Key Participants and Contributions

The success of GENCODE can be attributed to a diverse consortium of participants and institutions committed to advancing genomic science. The Wellcome Trust Sanger Institute leads these efforts, supported by notable contributors including EMBL European Bioinformatics Institute, Massachusetts Institute of Technology (MIT), Yale University, and others from across Europe and North America.

The collaborative framework ensures that insights from various fields converge to enhance gene annotation accuracy and depth. Key participants include:

  • Paul Flicek (Lead PI), EMBL European Bioinformatics Institute
  • Roderic Guigo (PI), Centre de Regulació Genòmica (CRG)
  • Manolis Kellis (PI), MIT
  • Mark B. Gerstein (PI), Yale University
  • Benedict Paten (PI), University of California, Santa Cruz
  • Michael Tress, Spanish National Cancer Research Centre (CNIO)
  • Jyoti Choudhary, Institute of Cancer Research (ICR)

Methodology Behind GENCODE Annotations

The methodology employed by GENCODE integrates both manual and automatic annotation processes to ensure comprehensive genomic coverage. Putative loci are verified through wet-lab experiments while computational predictions undergo meticulous manual review.

The automatic annotation is conducted via the Ensembl gene build system, which relies on experimental evidence derived from mRNAs and protein sequences available in public databases. Conversely, manual annotations are performed by dedicated teams who utilize advanced pipelines to identify unannotated regions and rectify any inaccuracies in existing annotations.

Additionally, quality assessment mechanisms are integral to ensuring high standards for transcript models. For instance, transcript models are assigned support levels based on rigorous scoring methods developed specifically for this purpose.

Challenges Faced by GENCODE

The journey towards creating an exhaustive encyclopedia of genes is fraught with challenges—most notably regarding the definition of what constitutes a “gene.” Historically evolving definitions have led to complexities surrounding alternative splicing, non-genic conservation, and the roles played by noncoding RNA genes. These elements add layers of difficulty as researchers strive for clarity amid ongoing discoveries that complicate long-standing notions about genetic coding.

The Future Direction of GENCODE

As GENCODE continues its work into future phases, it remains committed to enhancing genomic data accessibility and utility for researchers worldwide. With ongoing advancements in sequencing technologies and collaborations with other databases like RefSeq and Uniprot for annotation convergence, there is optimism surrounding its potential contributions toward understanding human genetics profoundly.

Conclusion

The GENCODE project stands as a testament to the collaborative spirit inherent in modern science. By focusing on comprehensive gene mapping and annotation through meticulous methodology combined with innovative technology applications such as CRISPR/Cas9 tracking systems, it not only advances our understanding of genetics but also supports broader initiatives aimed at addressing health challenges globally. As research progresses within this dynamic field, projects like GENCODE will undoubtedly play an essential role in shaping our comprehension of genomics today and into the future.


Artykuł sporządzony na podstawie: Wikipedia (EN).