Score
Designs and implements optical character recognition systems specialized for a particular writing script, producing models and pipelines that segment and classify individual script characters from images (including cropped character images) and map them to canonical character labels. Includes developing script-specific preprocessing and augmentation strategies, label sets, and evaluation methods to improve recognition accuracy and generalization.
Addressing the challenges of cross-script generalization (e.g., Latin, Chinese, cipher scripts) and high annotation cost for character-level supervision in text-line recognition (OCR/HTR), this paper proposes DTLR—an end-to-end detection-based text-line recognition framework. DTLR reformulates text-line recognition as a parallel character detection task, abandoning autoregressive decoding. It is trained solely with line-level supervision, eliminating the need for costly character-level annotations. Key technical contributions include: (1) a Transformer-based multi-instance detector for simultaneous character localization and classification; (2) synthetic data pretraining; (3) dynamic masked consistency learning to enhance robustness; and (4) cross-script transfer strategies. DTLR achieves new state-of-the-art results on CASIA v2 (Chinese), Borg, and Copiale (cipher scripts), demonstrating significantly improved generalization across multilingual and low-resource scripts while drastically reducing annotation dependency.
This study addresses the challenge that existing handwritten text recognition methods lack interpretable visual metrics suitable for paleographic analysis. The authors propose a novel architecture requiring only line-level transcription supervision, integrating a Transformer-based detection model, prototype-driven character representation learning, and a line-level reconstruction module to enable weakly supervised character localization and deformation modeling. This approach is the first to support automated paleographic measurements of individual characters, bigrams, and inter-character spacing. Evaluated on 160 pages from the 14th-century manuscript BnF fr. 2813, the method effectively distinguishes scribal styles and reveals subtle writing variations using only single-column text, significantly outperforming baselines such as Learnable Typewriter. Code and data are publicly released.
To address the poor robustness of text-image classification in uncontrolled environments, this paper proposes an edge-deployable document type classification system capable of distinguishing invoices, tables, letters, and reports. Methodologically, it introduces a novel collaborative architecture integrating DBNet++ and BART: DBNet++ enhances text detection robustness under challenging conditions—including curved text, low resolution, illumination variations, and partial occlusion—while BART performs semantic-driven, fine-grained classification based on detected text. A multimodal preprocessing pipeline and a cross-platform PyQt5-based UI further enable flexible input modalities, including USB/hard-disk file loading and real-time camera streaming. Evaluated on the Total-Text dataset, the system achieves a text recognition accuracy of 94.62%. It demonstrates stable, uninterrupted operation for over 10 hours, significantly improving both classification accuracy and practical usability for document-type identification in real-world scenarios.
To address low OCR accuracy caused by degradation in historical document images, this paper proposes a two-stage end-to-end optimization framework. In the first stage, a U-Net–based image restoration model is trained on a synthetically generated multi-degradation dataset to jointly optimize visual clarity and linguistic consistency. In the second stage, a semantic-aware ByT5 model performs post-OCR error correction, enhanced by a multi-directional block extraction and fusion mechanism tailored for large-format documents. The key innovations include the first joint optimization of image restoration quality and text semantic consistency, and the construction of the first cross-lingual (English/French/Spanish) synthetic dataset for historical text. Evaluated on 13,831 pages of real historical documents, the framework reduces character error rate by 63.9–70.3% over baseline OCR systems, demonstrating substantial improvement.
This work addresses the limited generalization of current OCR models beyond mainstream scripts, despite their strong performance on widely used writing systems. To this end, we introduce GlotOCR Bench, a comprehensive evaluation benchmark covering over 100 Unicode scripts, constructed from authentic multilingual text. The dataset includes both clean and degraded images rendered using Google Fonts, HarfBuzz, and FreeType, and supports bidirectional writing directions. We present the first systematic assessment of leading open- and closed-source OCR models on this benchmark, revealing that most models perform effectively on fewer than 10 scripts, with even the strongest failing to generalize reliably beyond 30. Models frequently produce noisy or character-confused outputs on unseen scripts. All code, data, and rendering pipelines are publicly released.
This work addresses the challenges of diverse historical Manchu handwriting styles—such as regular script, running script, and semi-cursive memorial script—and the scarcity of annotated data in optical character recognition (OCR). The authors propose a mixture-of-experts routing system that innovatively repurposes model checkpoints from iterative fine-tuning as domain-specific experts. A lightweight page-level visual style classifier enables highly accurate expert selection, achieving 99.3% routing accuracy, and dynamically instantiates new experts when no suitable one exists. Remarkably, without access to ground-truth style labels, the system attains character error rates of 0.30%, 1.57%, and 4.83% on three test sets, matching the performance upper bound achievable with oracle style labels and substantially improving cross-style Manchu text recognition under low-resource conditions.
This study addresses the challenging task of character detection and recognition in handwritten forms, which remains difficult due to structural complexity and heavy reliance on extensive manual annotations. The authors propose an end-to-end deep neural network approach that unifies character detection and classification into a single task, thereby eliminating the dependency on manually annotated data inherent in conventional two-stage pipelines. By synthesizing training data using the EMNIST dataset combined with realistic form layouts, the model achieves strong generalization without requiring real-world labeled examples. Evaluated on actual handwritten examination forms, the method attains an overall character recognition accuracy of 88.28%, substantially outperforming existing two-stage approaches.
This work addresses the limitation of existing Korean pre-trained language models, which predominantly rely on subword tokenization and thus fail to explicitly capture the compositional structure of Hangul characters—formed from Jamo sub-characters—and their morphophonological regularities. To overcome this, the authors propose SCRIPT, a model-agnostic module that injects Jamo-level structural knowledge into existing Korean pre-trained models without requiring architectural modifications or re-pretraining. SCRIPT enhances subword embeddings through compositional Jamo representations and an embedding augmentation mechanism, thereby refining the granularity of linguistic representation. Experimental results demonstrate that SCRIPT consistently outperforms baseline models across a range of Korean natural language understanding and generation tasks, while also yielding embedding spaces that more accurately reflect underlying grammatical and semantic patterns.
This work addresses the challenge of accurately and efficiently digitizing complex documents containing handwritten content, irregular tables, and heterogeneous layouts—tasks that remain difficult for conventional OCR systems and current large language models. The authors propose an interactive document digitization system that integrates layout-aware parsing, OCR, and a large language model, enhanced by a user-in-the-loop correction propagation mechanism. Leveraging layout-aware inference, the system automatically generalizes user edits or natural language instructions applied to a local region to structurally similar regions across the document. In a user study (n=12), this approach significantly improved correction efficiency, reduced repetitive manual operations, and enabled more controllable and effective reconstruction of document structure and content.