Strong Feature Sets from Small Samples

1 January 2002

journal article
research article
Published by Mary Ann Liebert Inc in Journal of Computational Biology

Vol. 9 (1), 127-146
https://doi.org/10.1089/10665270252833226

Abstract

For small samples, classifier design algorithms typically suffer from overfitting. Given a set of features, a classifier must be designed and its error estimated. For small samples, an error estimator may be unbiased but, owing to a large variance, often give very optimistic estimates. This paper proposes mitigating the small-sample problem by designing classifiers from a probability distribution resulting from spreading the mass of the sample points to make classification more difficult, while maintaining sample geometry. The algorithm is parameterized by the variance of the spreading distribution. By increasing the spread, the algorithm finds gene sets whose classification accuracy remains strong relative to greater spreading of the sample. The error gives a measure of the strength of the feature set as a function of the spread. The algorithm yields feature sets that can distinguish the two classes, not only for the sample data, but for distributions spread beyond the sample data. For linear classifiers, the topic of the present paper, the classifiers are derived analytically from the model, thereby providing an enormous savings in computation time. The algorithm is applied to cancer classification via cDNA microarrays. In particular, the genes BRCA1 and BRCA2 are associated with a hereditary disposition to breast cancer, and the algorithm is used to find gene sets whose expressions can be used to classify BRCA1 and BRCA2 tumors.

Keywords

This publication has 32 references indexed in Scilit:

Molecular classification of cutaneous malignant melanoma by gene expression profiling
Nature, 2000
Tissue Classification with Gene Expression Profiles
Journal of Computational Biology, 2000
Broad patterns of gene expression revealed by clustering analysis of tumor and normal colon tissues probed by oligonucleotide arrays
Proceedings of the National Academy of Sciences, 1999
The Lymphochip: A Specialized cDNA Microarray for the Genomic-scale Analysis of Gene Expression in Normal and Malignant Lymphocytes
Cold Spring Harbor Symposia on Quantitative Biology, 1999
The Molecular Chaperone αA-Crystallin Enhances Lens Epithelial Cell Growth and Resistance to UVA Stress
Published by Elsevier ,1998
Cytokeratin expression in breast cancer: Phenotypic changes associated with disease progression
Cytometry, 1998
Exploring the Metabolic and Genetic Control of Gene Expression on a Genomic Scale
Science, 1997
Cyclin D1 protein expression and function in human breast cancer
International Journal of Cancer, 1994
Modified controlled random search algorithms
International Journal of Computer Mathematics, 1994
On the capabilities of multilayer perceptrons
Journal of Complexity, 1988

Cited by 68 articles