Learning to Count: Robust Estimates for Labeled Distances between Molecular Sequences

6 January 2009

journal article
Published by Oxford University Press (OUP) in Molecular Biology and Evolution

Vol. 26 (4), 801-814
https://doi.org/10.1093/molbev/msp003

Abstract

Researchers routinely estimate distances between molecular sequences using continuous-time Markov chain models. We present a new method, robust counting, that protects against the possibly severe bias arising from model misspecification. We achieve this robustness by generalizing the conventional distance estimation to incorporate the empirical distribution of site patterns found in the observed pairwise sequence alignment. Our flexible framework allows for computing distances based only on a subset of possible substitutions. From this, we show how to estimate labeled codon distances, such as expected numbers of synonymous or nonsynonymous substitutions. We present two simulation studies. The first compares the relative bias and variance of conventional and robust labeled nucleotide estimators. In the second simulation, we demonstrate that robust counting furnishes accurate synonymous and nonsynonymous distance estimates based only on easy-to-fit models of nucleotide substitution, bypassing the need for computationally expensive codon models. We conclude with three empirical examples. In the first two examples, we investigate the evolutionary dynamics of the influenza A hemagglutinin gene using labeled codon distances. In the final example, we demonstrate the advantages of using robust synonymous distances to alleviate the effect of convergent evolution on phylogenetic analysis of an HIV transmission network.

Keywords

This publication has 48 references indexed in Scilit:

Historical contingency and the evolution of a key innovation in an experimental population ofEscherichia coli
Proceedings of the National Academy of Sciences, 2008
Counting labeled transitions in continuous-time Markov models of evolution
Journal of Mathematical Biology, 2007
FLAN: a web server for influenza virus genome annotation
Nucleic Acids Research, 2007
PAML 4: Phylogenetic Analysis by Maximum Likelihood
Molecular Biology and Evolution, 2007
Phylogenetic Reconstruction of Orthology, Paralogy, and Conserved Synteny for Dog and Human
PLoS Computational Biology, 2006
Statistical Inference in Evolutionary Models of DNA Sequences via the EM Algorithm
Statistical Applications in Genetics and Molecular Biology, 2005
An expectation maximization algorithm for training hidden substitution models 1 1Edited by F. Cohen
Journal of Molecular Biology, 2002
Convergent evolution: the need to be explicit
Trends in Biochemical Sciences, 1994
Evaluation of the maximum likelihood estimate of the evolutionary tree topologies from DNA sequence data, and the branching order in hominoidea
Journal of Molecular Evolution, 1989
A simple method for estimating evolutionary rates of base substitutions through comparative studies of nucleotide sequences
Journal of Molecular Evolution, 1980

Cited by 137 articles