Joint Data Harmonization and Group Cardinality Constrained Classification
To boost the power of classifiers, studies often increase the size of existing samples through the addition of independently collected data sets. Doing so requires harmonizing the data for demographic and acquisition differences based on a control cohort before performing disease specific classification. The initial harmonization often mitigates group differences negatively impacting classification accuracy. To preserve cohort separation, we propose the first model unifying linear regression for data harmonization with a logistic regression for disease classification. Learning to harmonize data is now an adaptive process taking both disease and control data into account. Solutions within that model are confined by group cardinality to reduce the risk of overfitting (via sparsity), to explicitly account for the impact of disease on the inter-dependency of regions (by grouping them), and to identify disease specific patterns (by enforcing sparsity via the \(l_0\)-‘norm’). We test those solutions in distinguishing HIV-Associated Neurocognitive Disorder from Mild Cognitive Impairment of two independently collected, neuroimage data sets; each contains controls and samples from one disease. Our classifier is impartial to acquisition difference between the data sets while being more accurate in diseases seperation than sequential learning of harmonization and classification parameters, and non-sparsity based logistic regressors.
KeywordsMild Cognitive Impairment Disease Accuracy Disease Cohort Group Sparsity Block Coordinate Descent
This research was supported in part by the NIH grants U01 AA017347, AA010723, K05-AA017168, K23-AG032872, and P30 AI027767. We thank Dr. Valcour for giving us access to the UHES data set. With respect to the ADNI data, collection and sharing for this project was funded by the NIH Grant U01 AG024904 and DOD Grant W81XWH-12-2-0012. Please see https://adni.loni.usc.edu/wp-content/uploads/how_to_apply/ADNI_DSP_Policy.pdf for further details.
- 2.Moradi, E., et al.: Predicting symptom severity in autism spectrum disorder based on cortical thickness measures in agglomerative data. bioRxiv (2016)Google Scholar
- 4.Zhang, Y., et al.: Computing group cardinality constraint solutions for logistic regression problems. Medical Image Analysis (2016, in press)Google Scholar
- 5.Sanmarti, M., et al.: HIV-associated neurocognitive disorders. J.M.P. 2(2) (2014)Google Scholar