A Study on Optimal Parameter Tuning for Rocchio Text Classifier

  • Alessandro Moschitti
Conference paper
Part of the Lecture Notes in Computer Science book series (LNCS, volume 2633)

Abstract

Current trend in operational text categorization is the designing of fast classification tools. Several studies on improving accuracy of fast but less accurate classifiers have been recently carried out. In particular, enhanced versions of the Rocchio text classifier, characterized by high performance, have been proposed. However, even in these extended formulations the problem of tuning its parameters is still neglected. In this paper, a study on parameters of the Rocchio text classifier has been carried out to achieve its maximal accuracy. The result is a model for the automatic selection of parameters. Its main feature is to bind the searching space so that optimal parameters can be selected quickly. The space has been bound by giving a feature selection interpretation of the Rocchio parameters. The benefit of the approach has been assessed via extensive cross evaluation over three corpora in two languages. Comparative analysis shows that the performances achieved are relatively close to the best TC models (e.g. Support Vector Machines).

Preview

Unable to display preview. Download preview PDF.

Unable to display preview. Download preview PDF.

References

  1. [1]
    Pivoted document length normalization. Technical Report TR95-1560, Cornell University, Computer Science, 1995.Google Scholar
  2. [2]
    Avi Arampatzis, Jean Beney, C. H. A. Koster, and T. P. van der Weide. Incrementality, half-life, and threshold optimization for adaptive document filtering. In the Nineth Text REtrieval Conference (TREC-9), Gaithersburg, Maryland, 2000.Google Scholar
  3. [3]
    Christopher Buckley and Gerald Salton. Optimization of relevance feedback weights. In Proceedings of SIGIR-95, pages 351–357, Seattle, US, 1995.Google Scholar
  4. [4]
    Wesley T. Chuang, Asok Tiyyagura, Jihoon Yang, and Giovanni Giuffrida. A fast algorithm for hierarchical text classification. In Proceedings of DaWaK-00, 2000.Google Scholar
  5. [5]
    William W. Cohen and Yoram Singer. Context-sensitive learning methods for text categorization. ACM Transactions on Information Systems, 17(2):141–173, 1999.CrossRefGoogle Scholar
  6. [6]
    Harris Drucker, Vladimir Vapnik, and Dongui Wu. Automatic text categorization and its applications to text retrieval. IEEE Transactions on Neural Networks, 10(5), 1999.Google Scholar
  7. [7]
    Norbert Gövert, Mounia Lalmas, and Norbert Fuhr. A probabilistic description-oriented approach for categorising Web documents. In Proceedings of CIKM-99.Google Scholar
  8. [8]
    David J. Ittner, David D. Lewis, and David D. Ahn. Text categorization of low quality images. In Proceedings of SDAIR-95, pages 301–315, Las Vegas, US, 1995.Google Scholar
  9. [9]
    T. Joachims. Text categorization with support vector machines: Learning with many relevant features. In In Proceedings of ECML-98, pages 137–142, 1998.Google Scholar
  10. [10]
    Thorsten Joachims. A probabilistic analysis of the rocchio algorithm with tfidf for text categorization. In Proceedings of ICML97 Conference. Morgan Kaufmann, 1997.Google Scholar
  11. [11]
    Ron Kohavi and George H. John. Wrappers for feature subset selection. Artificial Intelligence, 97(1–2):273–324, 1997.MATHCrossRefGoogle Scholar
  12. [12]
    Wai Lam and Chao Y. Ho. Using a generalized instance set for automatic text categorization. In Proceedings of SIGIR-98, 1998.Google Scholar
  13. [13]
    G: Salton and C. Buckley. Term-weighting approaches in automatic text retrieval. Information Processing and Management, 24(5):513–523, 1988.CrossRefGoogle Scholar
  14. [14]
    Robert E. Schapire, Yoram Singer, and Amit Singhal. Boosting and Rocchio applied to text filtering. In W. Bruce Croft, A. Moffat, C. J. van Rijsbergen, R. Wilkinson, and J. Zobel, editors, Proceedings of SIGIR-98, pages 215–223, Melbourne, AU, 1998. ACM Press, New York, US.CrossRefGoogle Scholar
  15. [15]
    Fabrizio Sebastiani. Machine learning in automated text categorization. ACM Computing Surveys, 34(1):1–47, 2002.CrossRefGoogle Scholar
  16. [16]
    Amit Singhal, John Choi, Donald Hindle, and Fernando C. N. Pereira. ATT at TREC-6: SDR track. In Text REtrieval Conference, pages 227–232, 1997.Google Scholar
  17. [17]
    Amit Singhal, Mandar Mitra, and Christopher Buckley. Learning routing queries in a query zone. In Proceedings of SIGIR-97, pages 25–32, Philadelphia, US, 1997.Google Scholar
  18. [18]
    K. Tzeras and S. Artman. Automatic indexing based on bayesian inference networks. In SIGIR 93, pages 22–34, 1993.Google Scholar
  19. [19]
    Y. Yang. An evaluation of statistical approaches to text categorization. Information Retrieval Journal, 1999.Google Scholar
  20. [20]
    Yiming Yang and Jan O. Pedersen. A comparative study on feature selection in text categorization. In Proceedings of ICML-97, pages 412–420, Nashville, US, 1997.Google Scholar

Copyright information

© Springer-Verlag Berlin Heidelberg 2003

Authors and Affiliations

  • Alessandro Moschitti
    • 1
  1. 1.Department of Computer Science Systems and ProductionUniversity of Rome Tor VergataRome(Italy)

Personalised recommendations