Evaluating Misclassifications in Imbalanced Data
Evaluating classifier performance with ROC curves is popular in the machine learning community. To date, the only method to assess confidence of ROC curves is to construct ROC bands. In the case of severe class imbalance with few instances of the minority class, ROC bands become unreliable. We propose a generic framework for classifier evaluation to identify a segment of an ROC curve in which misclassifications are balanced. Confidence is measured by Tango’s 95%-confidence interval for the difference in misclassification in both classes. We test our method with severe class imbalance in a two-class problem. Our evaluation favors classifiers with low numbers of misclassifications in both classes. Our results show that the proposed evaluation method is more confident than ROC bands.
Unable to display preview. Download preview PDF.
- 1.Ling, C.X., Huang, J., Zang, H.: Auc: a better measure than accuracy in comparing learning algorithms. In: Canadian Conference on AI, pp. 329–341 (2003)Google Scholar
- 2.Provost, F., Fawcett, T.: Analysis and visualization f classifier performance: Comparison under imprecise class and cost distributions. In: The Third International Conference on Knowledge Discovery and Data Mining, pp. 34–48 (1997)Google Scholar
- 3.Cohen, W.W., Schapire, R.E., Singer, Y.: Learning to order things. Journal of Artificial Intelligence Research (10), 243–270 (1999)Google Scholar
- 4.Swets, J.: Measuring the accuracy of diagnostic systems. Science (240), 1285–1293 (1988)Google Scholar
- 5.Drummond, C., Holte, R.C.: Explicitly representing expected cost: An alternative to roc representation. In: The Sixth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 198–207 (2000)Google Scholar
- 6.Drummond, C., Holte, R.C.: What roc curves can’t do (and cost curves can). In: ECAI 2004 Workshop on ROC Analysis in AI (2004)Google Scholar
- 7.Macskassy, S.A., Provost, F., Rosset, S.: Roc confidence bands: An empirical evaluation. In: Proceedings of the 22nd International Conference on Machine Learning (ICML 2005), pp. 537–544 (2005)Google Scholar
- 8.Macskassy, S.A., Provost, F.: Confidence bands for roc curves: Methods and empirical study. In: Proceedings of the 1st Workshop on ROC Analasis in AI (ROCAI-2004) at ECAI-2004 (2004)Google Scholar
- 9.Drummond, C., Holte, R.C.: Severe class imbalance: Why better algorithms aren’t the answer. In: Proceedings of the 16th European Conference of Machine Learning, pp. 539–546 (2005)Google Scholar
- 10.Motulsky, H.: Intuitive Biostatistics. Oxford University Press, Oxford (1995)Google Scholar
- 13.Newman, D.J., Hettich, S., Blake, C.L., Merz, C.J.: UCI repository of machine learning databases, University of California, Irvine, Dept. of Information and Computer Sciences (1998), http://www.ics.uci.edu/~mlearn/MLRepository.html
- 17.Everitt, B.S.: The analysis of contingency tables. Chapman-Hall, Boca Raton (1992)Google Scholar