<?xml version="1.0" encoding="UTF-8"?><xml><records><record><source-app name="Biblio" version="7.x">Drupal-Biblio</source-app><ref-type>17</ref-type><contributors><authors><author><style face="normal" font="default" size="100%">Sakellariou, A.</style></author><author><style face="normal" font="default" size="100%">Sanoudou, D.</style></author><author><style face="normal" font="default" size="100%">Spyrou, G</style></author></authors></contributors><titles><title><style face="normal" font="default" size="100%">Combining multiple hypothesis testing and affinity propagation clustering leads to accurate, robust and sample size independent classification on gene expression data</style></title><secondary-title><style face="normal" font="default" size="100%">BMC BioinformaticsBMC BioinformaticsBMC Bioinformatics</style></secondary-title><alt-title><style face="normal" font="default" size="100%">BMC bioinformatics</style></alt-title><short-title><style face="normal" font="default" size="100%">BMC bioinformaticsBMC bioinformatics</style></short-title></titles><keywords><keyword><style  face="normal" font="default" size="100%">*Cluster Analysis</style></keyword><keyword><style  face="normal" font="default" size="100%">Algorithms</style></keyword><keyword><style  face="normal" font="default" size="100%">computer simulation</style></keyword><keyword><style  face="normal" font="default" size="100%">Gene Expression Profiling/*methods</style></keyword><keyword><style  face="normal" font="default" size="100%">Humans</style></keyword><keyword><style  face="normal" font="default" size="100%">Multigene Family</style></keyword><keyword><style  face="normal" font="default" size="100%">Neoplasms/genetics</style></keyword><keyword><style  face="normal" font="default" size="100%">Neuromuscular Diseases/genetics</style></keyword><keyword><style  face="normal" font="default" size="100%">Oligonucleotide Array Sequence Analysis/methods</style></keyword><keyword><style  face="normal" font="default" size="100%">Sample Size</style></keyword><keyword><style  face="normal" font="default" size="100%">Support Vector Machine</style></keyword></keywords><dates><year><style  face="normal" font="default" size="100%">2012</style></year><pub-dates><date><style  face="normal" font="default" size="100%">Oct 17</style></date></pub-dates></dates><volume><style face="normal" font="default" size="100%">13</style></volume><pages><style face="normal" font="default" size="100%">270</style></pages><isbn><style face="normal" font="default" size="100%">1471-2105 (Electronic)1471-2105 (Linking)</style></isbn><language><style face="normal" font="default" size="100%">eng</style></language><abstract><style face="normal" font="default" size="100%">BACKGROUND: A feature selection method in microarray gene expression data should be independent of platform, disease and dataset size. Our hypothesis is that among the statistically significant ranked genes in a gene list, there should be clusters of genes that share similar biological functions related to the investigated disease. Thus, instead of keeping N top ranked genes, it would be more appropriate to define and keep a number of gene cluster exemplars. RESULTS: We propose a hybrid FS method (mAP-KL), which combines multiple hypothesis testing and affinity propagation (AP)-clustering algorithm along with the Krzanowski &amp; Lai cluster quality index, to select a small yet informative subset of genes. We applied mAP-KL on real microarray data, as well as on simulated data, and compared its performance against 13 other feature selection approaches. Across a variety of diseases and number of samples, mAP-KL presents competitive classification results, particularly in neuromuscular diseases, where its overall AUC score was 0.91. Furthermore, mAP-KL generates concise yet biologically relevant and informative N-gene expression signatures, which can serve as a valuable tool for diagnostic and prognostic purposes, as well as a source of potential disease biomarkers in a broad range of diseases. CONCLUSIONS: mAP-KL is a data-driven and classifier-independent hybrid feature selection method, which applies to any disease classification problem based on microarray data, regardless of the available samples. Combining multiple hypothesis testing and AP leads to subsets of genes, which classify unknown samples from both, small and large patient cohorts with high accuracy.</style></abstract><accession-num><style face="normal" font="default" size="100%">23075381</style></accession-num><notes><style face="normal" font="default" size="100%">Sakellariou, ArgirisSanoudou, DespinaSpyrou, GeorgeengResearch Support, Non-U.S. Gov'tEngland2012/10/19 06:00BMC Bioinformatics. 2012 Oct 17;13:270. doi: 10.1186/1471-2105-13-270.</style></notes><custom2><style face="normal" font="default" size="100%">3542193</style></custom2><auth-address><style face="normal" font="default" size="100%">Biomedical Informatics Unit, Biomedical Research Foundation of the Academy of Athens, Athens, Greece.</style></auth-address></record></records></xml>