Using Literal and Grammatical Statistics for Authorship Attribution
- Cite this article as:
- Kukushkina, O.V., Polikarpov, A.A. & Khmelev, D.V. Problems of Information Transmission (2001) 37: 172. doi:10.1023/A:1010478226705
Markov chains are used as a formal mathematical model for sequences of elements of a text. This model is applied for authorship attribution of texts. As elements of a text, we consider sequences of letters or sequences of grammatical classes of words. It turns out that the frequencies of occurrences of letter pairs and pairs of grammatical classes in a Russian text are rather stable characteristics of an author and, apparently, they could be used in disputed authorship attribution. A comparison of results for various modifications of the method using both letters and grammatical classes is given. Experimental research involves 385 texts of 82 writers. In the Appendix, the research of D.V. Khmelev is described, where data compression algorithms are applied to authorship attribution.