Chi-Square Classifier for Document Categorization

  • Mikhail Alexandrov
  • Alexander Gelbukh
  • George Lozovoi
Conference paper

DOI: 10.1007/3-540-44686-9_45

Part of the Lecture Notes in Computer Science book series (LNCS, volume 2004)
Cite this paper as:
Alexandrov M., Gelbukh A., Lozovoi G. (2001) Chi-Square Classifier for Document Categorization. In: Gelbukh A. (eds) Computational Linguistics and Intelligent Text Processing. CICLing 2001. Lecture Notes in Computer Science, vol 2004. Springer, Berlin, Heidelberg

Abstract

The problem of document categorization is considered. The set of domains and the keywords specific for these domains is supposed to be selected beforehand as initial data. We apply the well-known statistical hypothesis test that considers images of documents and domains as normalized vectors. In comparison with existing methods, such approach allows to take into account a random character of initial data. The classifier is developed in the framework of Document Investigator software package.

Preview

Unable to display preview. Download preview PDF.

Unable to display preview. Download preview PDF.

Copyright information

© Springer-Verlag Berlin Heidelberg 2001

Authors and Affiliations

  • Mikhail Alexandrov
    • 1
  • Alexander Gelbukh
    • 1
  • George Lozovoi
    • 2
  1. 1.Center for Computing Research, IPNMexico
  2. 2.DatagisticsCanada

Personalised recommendations