Advertisement

Creation of Textual Versions of Historical Documents from Polish Digital Libraries

  • Adam Dudczak
  • Miłosz Kmieciak
  • Marcin Werla
Part of the Lecture Notes in Computer Science book series (LNCS, volume 7489)

Abstract

This paper describes the results of initial work aimed at increasing the number and improving the quality of textual versions of the historical documents available in Polish digital libraries. Digital libraries community is missing tools that integrate existing digitisation workflow with customizable OCR engine and crowd–based text correction, this paper describes work on providing such a solution. Apart from today’s state of the art in this field, this paper includes a description of the Virtual Transcription Laboratory (VTL) prototype, a crowdsourcing platform that utilize the Tesseract OCR engine. The last chapter outlines results of the prototype’s evaluation on real life dataset of historical documents from the IMPACT project. Results prove the applicability of the proposed solution as an enhancement of the digitisation workflow.

Keywords

Ground Truth Digital Library Historical Document Textual Version Real Life Dataset 
These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.

Preview

Unable to display preview. Download preview PDF.

Unable to display preview. Download preview PDF.

References

  1. 1.
    Lewandowska, A., Werla, M.: Pionier network digital libraries federation interoperability of advanced network services implemented on a country scale. Computational Methods in Science and Technology, 119–124 (2010)Google Scholar
  2. 2.
    Mazurek, C., Sielski, K., Stroiński, M., Walkowska, J., Werla, M., Węglarz, J.: Transforming a Flat Metadata Schema to a Semantic Web Ontology: The Polish Digital Libraries Federation and CIDOC CRM Case Study. In: Bembenik, R., Skonieczny, L., Rybiński, H., Niezgodka, M. (eds.) Intelligent Tools for Building a Scient. Info. Plat. SCI, vol. 390, pp. 153–177. Springer, Heidelberg (2012)CrossRefGoogle Scholar
  3. 3.
    Dudczak, A., Kmieciak, M., Werla, M.: Country scale infrastructure for creation of full text versions of historical documents from Polish Digital Libraries. Presented at Interedition Symposium: Scholarly Digital Editions, Tools and Infrastructure, The Hague, Netherlands (2012)Google Scholar
  4. 4.
    Neudecker, C., Tzadok, A.: User Collaboration for Improving Access to Historical Texts. Liber Quarterly 20(1), 119–128 (2010)Google Scholar
  5. 5.
    Alexandrov, V.: Error evaluation and applicability of ocr systems. In: Proceedings of the 4th International Conference Conference on Computer Systems and Technologies: e-Learning. CompSysTech 2003, pp. 308–313. ACM, New York (2003)CrossRefGoogle Scholar

Copyright information

© Springer-Verlag Berlin Heidelberg 2012

Authors and Affiliations

  • Adam Dudczak
    • 1
  • Miłosz Kmieciak
    • 1
  • Marcin Werla
    • 1
  1. 1.Poznań Supercomputing and Networking CenterPoznańPoland

Personalised recommendations