Chi-Square Classifier for Document Categorization
The problem of document categorization is considered. The set of domains and the keywords specific for these domains is supposed to be selected beforehand as initial data. We apply the well-known statistical hypothesis test that considers images of documents and domains as normalized vectors. In comparison with existing methods, such approach allows to take into account a random character of initial data. The classifier is developed in the framework of Document Investigator software package.
Unable to display preview. Download preview PDF.
- 1.Alexandrov, M., Gelbukh, A., and Makagonov, P. Some keyword-based characteristics for evaluation of thematic structure of multidisciplinary documents. Proc. of 1st Int. Conf. on Intelligent Text Processing and Computational Linguistics, Mexico City, 2000, pp. 390–401.Google Scholar
- 2.Cramer, H. Mathematical methods of statistics. Cambridge, 1946.Google Scholar
- 3.Guzman-Arenas, A. Finding the main themes in a Spanish documents. Intern. J. of Expert Systems with Applications, 1998, v. 14, N 1/2, pp. 139–148.Google Scholar
- 4.Mitchel, T. Machine learning. New-York, McGraw Hill, 1997.Google Scholar