Compilation of a Spanish Representative Corpus
- Cite this paper as:
- Gelbukh A., Sidorov G., Chanona-Hernández L. (2002) Compilation of a Spanish Representative Corpus. In: Gelbukh A. (eds) Computational Linguistics and Intelligent Text Processing. CICLing 2002. Lecture Notes in Computer Science, vol 2276. Springer, Berlin, Heidelberg
Due to the Zipf law, even a very large corpus contains very few occurrences (tokens) for the majority of its different words (types). Only a corpus containing enough occurrences of even rare words can provide necessary statistical information for the study of contextual usage of words. We call such corpus representative and suggest to use Internet for its compilation. The corresponding algorithm and its application to Spanish are described. Different concepts of a representative corpus are discussed.
Unable to display preview. Download preview PDF.