๐ค AI Summary
This study addresses the challenge of automated interpretation of Ashokan Brahmi inscription images, which is hindered by severe image degradation and scarce digital resources. We propose an end-to-end artificial intelligence framework that integrates image enhancement, optical character recognition (OCR), transliteration, and neural machine translation to enable automatic conversion from low-quality inscription images into English translations. Furthermore, this work constructs the largest Brahmi OCR dataset and PrakritโEnglish parallel corpus to date, bridging a critical gap in the digitization of ancient scripts. By establishing a reusable benchmark for low-resource ancient script analysis, this project significantly advances interdisciplinary research at the intersection of computer vision and digital epigraphy.
๐ Abstract
Ancient script image restoration is a fundamental problem in computer vision, as it directly affects the reliable analysis and interpretation of historical documents and inscriptions. Ashokan Brahmi is an ancient script extensively used during the reign of Emperor Ashoka in the 3rd century BC, primarily for inscriptions in Prakrit. These inscriptions, including major and minor rock and pillar edicts, constitute a valuable yet largely unexplored source of data for computational analysis. The degraded nature of inscription imagery and the lack of standardized digital resources pose significant challenges for automated processing.
We present an end-to-end AI-based framework for understanding ancient inscriptions that encompasses image enhancement, optical character recognition (OCR), transliteration, and neural machine translation (NMT). The proposed pipeline processes low-quality images captured directly from stone inscriptions, performs image restoration and Brahmi script character recognition, maps the recognized characters to the Roman script, and finally translates the resulting Prakrit text into English. We also introduce two new datasets: (i) InscriptionOCR Dataset: the largest publicly usable digital OCR dataset for Brahmi script to date, consisting of over 200,000 character images across about 600 classes, and (ii) a bilingual Prakrit-English parallel corpus comprising over 2,000 sentence pairs for NMT. We believe that the proposed framework and datasets will facilitate future research in ancient script analysis, low-resource OCR, and digital epigraphy.