Plagiarism Detection of Paraphrases in Text Documents with Document Retrieval
Retrieval of documents is used for finding relevant documents to user queries and plagiarism is the act of copying the contents of one’s work without any acknowledgement. Paraphrasing is a type of plagiarism where the contents from source may be changed. This paper proposes a new document retrieval system and paraphrase plagiarism detection of text documents using multi-layered self organizing map (MLSOM). In the proposed system tree structure is extracted for the document that hierarchically represents the document features as document, pages and paragraphs. To handle the tree-structured documents in an efficient way, MLSOM is used as a clustering algorithm. Using MLSOM the documents can be compared for detecting plagiarism and it finds out the local similarity. Paraphrased plagiarism can be detected by finding the similarity between sentences of two documents which is a kind of local similarity detection.
KeywordsPlagiarism detection paraphrase multi-layer self organizing map sentence similarity
Unable to display preview. Download preview PDF.
- 1.Yates, Neto: Modern Information Retrieval, vol. 15, pp. 750–780. Addison- Wesley/Longman, Reading, MA (1999)Google Scholar
- 3.Liu, Croft: Statistical Language Modeling for Information Retrieval. In: Cronin, B. (ed.) Annual Review of Information Science & Technology, vol. 38, pp. 556–567 (2004)Google Scholar
- 4.Lin, Y., Ye, H.: Input Data Representation for Self-Organizing Map in Software Classification. In: Second International Conference on Knowledge Acquisition and Modeling, Callaghan, Australia, pp. 163–195 (2009)Google Scholar
- 5.Kappe, Zaka: Plagiarism—A survey. Journal of Universal Computing 12(8), 1050–1084 (2006)Google Scholar
- 7.Heintze: Scalable document fingerprinting. In: Proc. 2nd USENIX Workshop Electron. Commerce, Oakland, CA, pp. 18–21 (November 2007) Google Scholar
- 8.Monostori, Zaslavsky, Schmidt: MatchDetectReveal: Finding overlapping and similar digital documents. In: Proc. of 21st Century Inf. Resources Manage. Assoc. Int. Conf. Challenges Inf. Technol. Manage., Anchorage, AK, pp. 955–957 (2000) Google Scholar
- 9.Weir, G.R.S., Gordon, M.A., Macgregor, G.: Technology in plagiarism detection and management. In: 34th ASEE/IEEE Frontiers in Education Conference, Savannah, GA, vol. 13, pp. 351–370 (2004)Google Scholar