Phylogenetic-based propagation of functional annotations within the Gene Ontology consortium
Top Cited Papers
Open Access
- 27 August 2011
- journal article
- research article
- Published by Oxford University Press (OUP) in Briefings in Bioinformatics
- Vol. 12 (5), 449-462
- https://doi.org/10.1093/bib/bbr042
Abstract
The goal of the Gene Ontology (GO) project is to provide a uniform way to describe the functions of gene products from organisms across all kingdoms of life and thereby enable analysis of genomic data. Protein annotations are either based on experiments or predicted from protein sequences. Since most sequences have not been experimentally characterized, most available annotations need to be based on predictions. To make as accurate inferences as possible, the GO Consortium's Reference Genome Project is using an explicit evolutionary framework to infer annotations of proteins from a broad set of genomes from experimental annotations in a semi-automated manner. Most components in the pipeline, such as selection of sequences, building multiple sequence alignments and phylogenetic trees, retrieving experimental annotations and depositing inferred annotations, are fully automated. However, the most crucial step in our pipeline relies on software-assisted curation by an expert biologist. This curation tool, Phylogenetic Annotation and INference Tool (PAINT) helps curators to infer annotations among members of a protein family. PAINT allows curators to make precise assertions as to when functions were gained and lost during evolution and record the evidence (e.g. experimentally supported GO annotations and phylogenetic information including orthology) for those assertions. In this article, we describe how we use PAINT to infer protein function in a phylogenetic context with emphasis on its strengths, limitations and guidelines. We also discuss specific examples showing how PAINT annotations compare with those generated by other highly used homology-based methods.Keywords
This publication has 16 references indexed in Scilit:
- The what, where, how and why of gene ontology--a primer for bioinformaticiansBriefings in Bioinformatics, 2011
- Formalization of taxon-based constraints to detect inconsistencies in annotation and ontology developmentBMC Bioinformatics, 2010
- GIGA: a simple, efficient algorithm for gene tree inference in the genomic ageBMC Bioinformatics, 2010
- PANTHER version 7: improved phylogenetic trees, orthologs and collaboration with the Gene Ontology ConsortiumNucleic Acids Research, 2009
- The Gene Ontology in 2010: extensions and refinementsNucleic Acids Research, 2009
- The Gene Ontology's Reference Genome Project: A Unified Framework for Functional Annotation across SpeciesPLoS Computational Biology, 2009
- How confident can we be that orthologs are similar, but paralogs differ?Trends in Genetics, 2009
- EnsemblCompara GeneTrees: Complete, duplication-aware phylogenetic trees in vertebratesGenome Research, 2008
- The Princeton Protein Orthology Database (P-POD): A Comparative Genomics Analysis Tool for BiologistsPLOS ONE, 2007
- Protein Molecular Function Prediction by Bayesian PhylogenomicsPLoS Computational Biology, 2005