Probabilistic retrieval and visualization of biologically relevant microarray experiments
Open Access
- 27 May 2009
- journal article
- research article
- Published by Oxford University Press (OUP) in Bioinformatics
- Vol. 25 (12), i145-i153
- https://doi.org/10.1093/bioinformatics/btp215
Abstract
Motivation: As ArrayExpress and other repositories of genome-wide experiments are reaching a mature size, it is becoming more meaningful to search for related experiments, given a particular study. We introduce methods that allow for the search to be based upon measurement data, instead of the more customary annotation data. The goal is to retrieve experiments in which the same biological processes are activated. This can be due either to experiments targeting the same biological question, or to as yet unknown relationships. Results: We use a combination of existing and new probabilistic machine learning techniques to extract information about the biological processes differentially activated in each experiment, to retrieve earlier experiments where the same processes are activated and to visualize and interpret the retrieval results. Case studies on a subset of ArrayExpress show that, with a sufficient amount of data, our method indeed finds experiments relevant to particular biological questions. Results can be interpreted in terms of biological processes using the visualization techniques. Availability: The code is available from http://www.cis.hut.fi/projects/mi/software/ismb09. Contact:jose.caldas@tkk.fiKeywords
This publication has 30 references indexed in Scilit:
- ArrayExpress update--from an archive of functional genomics experiments to the atlas of gene expressionNucleic Acids Research, 2009
- GEOmetadb: powerful alternative search engine for the Gene Expression OmnibusBioinformatics, 2008
- Gene set enrichment analysis using linear models and diagnosticsBioinformatics, 2008
- GenMAPP 2: new features and resources for pathway analysisBMC Bioinformatics, 2007
- Reactome: a knowledge base of biologic pathways and processesGenome Biology, 2007
- Model based analysis of real-time PCR data from DNA binding dye protocolsBMC Bioinformatics, 2007
- Gene set enrichment analysis: A knowledge-based approach for interpreting genome-wide expression profilesProceedings of the National Academy of Sciences, 2005
- Integrative analysis of genome‐wide experiments in the context of a large high‐throughput data compendiumMolecular Systems Biology, 2005
- PGC-1α-responsive genes involved in oxidative phosphorylation are coordinately downregulated in human diabetesNature Genetics, 2003
- KEGG: Kyoto Encyclopedia of Genes and GenomesNucleic Acids Research, 2000