NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins

Top Cited Papers

Open Access

3 January 2007

journal article
Published by Oxford University Press (OUP) in Nucleic Acids Research

Vol. 35 (Database), D61-D65
https://doi.org/10.1093/nar/gkl842

Abstract

NCBI's reference sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) is a curated non-redundant collection of sequences representing genomes, transcripts and proteins. The database includes 3774 organisms spanning prokaryotes, eukaryotes and viruses, and has records for 2,879,860 proteins (RefSeq release 19). RefSeq records integrate information from multiple sources, when additional data are available from those sources and therefore represent a current description of the sequence and its features. Annotations include coding regions, conserved domains, tRNAs, sequence tagged sites (STS), variation, references, gene and protein product names, and database cross-references. Sequence is reviewed and features are added using a combined approach of collaboration and other input from the scientific community, prediction, propagation from GenBank and curation by NCBI staff. The format of all RefSeq records is validated, and an increasing number of tests are being applied to evaluate the quality of sequence and annotation, especially in the context of complete genomic sequence.

Keywords

This publication has 11 references indexed in Scilit:

Entrez Gene: gene-centered information at NCBI
Nucleic Acids Research, 2006
The Mouse Genome Database (MGD): updates and enhancements
Nucleic Acids Research, 2006
WormBase: better software, richer content
Nucleic Acids Research, 2005
FlyBase: genes and gene models
Nucleic Acids Research, 2004
Regulation of gene expression by stop codon recoding: selenocysteine
Gene, 2003
Generation of protein isoform diversity by alternative initiation of translation at non‐AUG codons
Biology of the Cell, 2003
The Arabidopsis Information Resource (TAIR): a model organism database providing a centralized, curated gateway to Arabidopsis biology, research materials and community
Nucleic Acids Research, 2003
Complete genomes in WWW Entrez: data representation and analysis.
Bioinformatics, 1999
Gapped BLAST and PSI-BLAST: a new generation of protein database search programs
Nucleic Acids Research, 1997
Basic local alignment search tool
Journal of Molecular Biology, 1990

Cited by 2949 articles