Modelling and simulating generic RNA-Seq experiments with the flux simulator

Open Access

7 September 2012

journal article
research article
Published by Oxford University Press (OUP) in Nucleic Acids Research

Vol. 40 (20), 10073-10083
https://doi.org/10.1093/nar/gks666

Abstract

High-throughput sequencing of cDNA libraries constructed from cellular RNA complements (RNA-Seq) naturally provides a digital quantitative measurement for every expressed RNA molecule. Nature, impact and mutual interference of biases in different experimental setups are, however, still poorly understood—mostly due to the lack of data from intermediate protocol steps. We analysed multiple RNA-Seq experiments, involving different sample preparation protocols and sequencing platforms: we broke them down into their common—and currently indispensable—technical components (reverse transcription, fragmentation, adapter ligation, PCR amplification, gel segregation and sequencing), investigating how such different steps influence abundance and distribution of the sequenced reads. For each of those steps, we developed universally applicable models, which can be parameterised by empirical attributes of any experimental protocol. Our models are implemented in a computer simulation pipeline called the Flux Simulator, and we show that read distributions generated by different combinations of these models reproduce well corresponding evidence obtained from the corresponding experimental setups. We further demonstrate that our in silico RNA-Seq provides insights about hidden precursors that determine the final configuration of reads along gene bodies; enhancing or compensatory effects that explain apparently controversial observations can be observed. Moreover, our simulations identify hitherto unreported sources of systematic bias from RNA hydrolysis, a fragmentation technique currently employed by most RNA-Seq protocols.

Keywords

This publication has 37 references indexed in Scilit:

Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation
Nature Biotechnology, 2010
Biases in Illumina transcriptome sequencing caused by random hexamer priming
Nucleic Acids Research, 2010
FRT-seq: amplification-free, strand-specific transcriptome sequencing
Nature Methods, 2010
RNA-Seq: a revolutionary tool for transcriptomics
Nature Reviews Genetics, 2009
A large genome center's improvements to the Illumina sequencing system
Nature Methods, 2008
Substantial biases in ultra-short read data sets from high-throughput DNA sequencing
Nucleic Acids Research, 2008
Mapping and quantifying mammalian transcriptomes by RNA-Seq
Nature Methods, 2008
Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nanostructures
Proceedings of the National Academy of Sciences, 2008
The Arabidopsis Information Resource (TAIR): gene structure and function annotation
Nucleic Acids Research, 2007
NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins
Nucleic Acids Research, 2007

Cited by 267 articles