International Journal on Digital Libraries

, Volume 3, Issue 2, pp 115–130

Automatic recognition of multi-word terms:. the C-value/NC-value method

Authors

  • Katerina Frantzi
    • Centre for Computational Linguistics, UMIST, P.O. Box 88, Manchester, M60 1QD, UK; E-mail: katerina,effie@ccl.umist.ac.uk
  • Sophia Ananiadou
    • Centre for Computational Linguistics, UMIST, P.O. Box 88, Manchester, M60 1QD, UK; E-mail: katerina,effie@ccl.umist.ac.uk
  • Hideki Mima
    • Dept. of Information Science, University of Tokyo, Hongo 7-3-1, Bunkyo-ku, Tokyo 113, Japan; E-mail: mima@is.s.u-tokyo.ac.jp
Natural language processing for digital libraries

DOI: 10.1007/s007999900023

Cite this article as:
Frantzi, K., Ananiadou, S. & Mima, H. Int J Digit Libr (2000) 3: 115. doi:10.1007/s007999900023

Abstract.

Technical terms (henceforth called terms ), are important elements for digital libraries. In this paper we present a domain-independent method for the automatic extraction of multi-word terms, from machine-readable special language corpora. The method, (C-value/NC-value ), combines linguistic and statistical information. The first part, C-value, enhances the common statistical measure of frequency of occurrence for term extraction, making it sensitive to a particular type of multi-word terms, the nested terms. The second part, NC-value, gives: 1) a method for the extraction of term context words (words that tend to appear with terms); 2) the incorporation of information from term context words to the extraction of terms.

Key words: Terms – Automatic extraction – Domain independence – Automatic Term Recognition (ATR) – Linguistic and statistical information

Copyright information

© Springer-Verlag 2000