A Complete OCR System for Gurmukhi Script
Recognition of Indian language scripts is a challenging problem. Work for the development of complete OCR systems for Indian language scripts is still in infancy. Complete OCR systems have recently been developed for Devanagri and Bangla scripts. Research in the field of recognition of Gurmukhi script faces major problems mainly related to the unique characteristics of the script like connectivity of characters on the headline, characters in a word present in both horizontal and vertical directions, two or more characters in a word having intersecting minimum bounding rectangles along horizontal direction, existence of a large set of visually similar character pairs, multi-component characters, touching characters which are present even in clean documents and horizontally overlapping text segments. This paper addresses the problems in the various stages of the development of a complete OCR for Gurmukhi script and discusses potential solutions.
- 6.Bansal, V.: Integrating knowledge sources in Devanagri text recognition. Ph.D. thesis. IIT Kanpur (1999).Google Scholar
- 7.Lehal, G. S., Singh, C.: Text segmentation of machine printed Gurmukhi script. Document Recognition and Retrieval VIII. Paul B. Kantor, Daniel P. Lopresti, Jiangying Zhou (eds.), Proceedings SPIE, USA. Vol. 4307. (2001) 223–231.Google Scholar
- 8.Lehal, G. S., Singh, C.: A shape based post processor for Gurmukhi OCR. Proceedings 6th International Conference on Document Analysis and Recognition, Seattle, USA. (2001) 1105–1109.Google Scholar