Score
Designs and builds classifiers, feature representations, and evaluation protocols that detect and distinguish very small stylistic differences in visual script or handwriting (e.g., different calligraphers or dynasty scripts). Also analyzes methods to isolate the style signal from textual content and to evaluate model performance on fine-grained style labels.
This study addresses the critical challenge of generating visual error feedback in handwritten documents to enhance handwriting skill training, particularly in educational settings. We present the first systematic analysis of the core difficulties inherent in this task, establish a unified evaluation framework, and conduct a comparative assessment of modular versus end-to-end deep learning approaches across handwriting recognition, error detection, and feedback generation. Experimental results reveal significant bottlenecks in current systems, particularly in precise character localization, semantic understanding, and alignment between detected errors and generated feedback, indicating that overall performance remains below practical usability. Beyond exposing these limitations, our work provides a clear foundation and direction for future research in this emerging domain.
Existing Chinese calligraphy datasets are scarce and predominantly annotated only at the character level, lacking cultural metadata—such as stylistic category, dynasty, and calligrapher—which hinders both accurate recognition and historical analysis. Method: We introduce MCCD, the first fine-grained, triple-attribute (style/dynasty/calligrapher) annotated dataset of individual Chinese calligraphic characters, comprising 329,715 images across 7,765 classes. We propose a novel annotation paradigm integrating expert verification with historical textual scholarship, and establish comprehensive single-task and multi-task deep learning benchmarks. Contribution/Results: Experiments reveal that stroke complexity and attribute coupling significantly degrade recognition performance; we report baseline results across multiple challenging subsets. MCCD bridges the gap from character-level recognition to cultural-attribute understanding, providing a foundational resource for calligraphic AI and diachronic studies of Chinese character evolution.
This work addresses the challenges of diverse historical Manchu handwriting styles—such as regular script, running script, and semi-cursive memorial script—and the scarcity of annotated data in optical character recognition (OCR). The authors propose a mixture-of-experts routing system that innovatively repurposes model checkpoints from iterative fine-tuning as domain-specific experts. A lightweight page-level visual style classifier enables highly accurate expert selection, achieving 99.3% routing accuracy, and dynamically instantiates new experts when no suitable one exists. Remarkably, without access to ground-truth style labels, the system attains character error rates of 0.30%, 1.57%, and 4.83% on three test sets, matching the performance upper bound achievable with oracle style labels and substantially improving cross-style Manchu text recognition under low-resource conditions.
Handwritten text recognition suffers from ambiguity arising from diverse writing styles, yet existing methods fail to explicitly model writer-specific stylistic variations, limiting recognition accuracy. To address this, we propose the Writer Style Block (WSB), a writer-conditioned instance normalization layer built upon text-block embeddings, which incorporates writer identity as an explicit input to enable style-adaptive modeling and zero-shot estimation of unseen writer embeddings. Our approach integrates contrastive pretraining, writer-conditioned representation learning, and domain-adaptive fine-tuning. Experiments demonstrate that WSB significantly outperforms style-agnostic baselines in writer-dependent settings and exhibits strong cross-writer generalization; however, it slightly underperforms in writer-independent configurations, suggesting a need for improved embedding regularization and training stability. The core contribution lies in the first integration of learnable writer embeddings with conditional normalization—yielding a lightweight, scalable architecture for style-adaptive recognition.
Historical handwritten Arabic manuscripts—particularly cursive scripts—are notoriously difficult to recognize, and high-quality annotated datasets are severely lacking. Method: This work introduces Muharaf, the first large-scale, open-source dataset of historical Arabic manuscript images, comprising over 1,600 pages from diverse document types (e.g., letters, poetry, legal texts), each annotated with expert-verified text-line polygon coordinates and page-structure labels. A novel data construction pipeline is proposed, integrating transcription alignment with spatial annotation. A CNN-based baseline model is trained and evaluated to validate the dataset’s utility. Contribution/Results: Muharaf is the first systematically curated Arabic manuscript dataset featuring fine-grained spatial annotations, thereby filling a critical gap in the Handwritten Text Recognition (HTR) community’s benchmark resources. It substantially strengthens data support for Arabic and other connected-script recognition tasks, and establishes a reproducible, community-accessible benchmark for historical document digitization and general cursive script analysis.
This study addresses the challenge that existing handwritten text recognition methods lack interpretable visual metrics suitable for paleographic analysis. The authors propose a novel architecture requiring only line-level transcription supervision, integrating a Transformer-based detection model, prototype-driven character representation learning, and a line-level reconstruction module to enable weakly supervised character localization and deformation modeling. This approach is the first to support automated paleographic measurements of individual characters, bigrams, and inter-character spacing. Evaluated on 160 pages from the 14th-century manuscript BnF fr. 2813, the method effectively distinguishes scribal styles and reveals subtle writing variations using only single-column text, significantly outperforming baselines such as Learnable Typewriter. Code and data are publicly released.
This study addresses the inefficiency of manual methods in large-scale typographic comparison of 17th-century Spanish playbills by proposing a character prototype–based statistical framework. The approach automatically extracts, clusters, and aligns character images to compute inter-book font distances and introduces, for the first time, an a contrario significance test to assess the reliability of observed typographic differences, enabling robust automated comparison of both roman and italic typefaces. By overcoming the scalability limitations of traditional analyses, the method—validated by domain experts—successfully uncovers new printer attributions and revises existing conclusions, thereby advancing digital bibliographical scholarship toward large-scale, automated investigation.
Existing vision-language models struggle to achieve fine-grained perception of calligraphic styles due to modality entanglement and flattened labeling in current datasets. To address this, this work introduces HCSU—the first dataset dedicated to historical Chinese calligraphy style understanding—comprising 39,307 character images from 49 calligraphers across 10 dynasties. HCSU systematically disentangles the two primary modalities of ink-on-paper (tie) and stone inscriptions (bei) and incorporates expert-authored hierarchical aesthetic descriptions. This dataset enables fine-grained style discrimination and interpretable aesthetic reasoning, establishing a new benchmark for the field. Evaluations reveal that while contemporary large models exhibit preliminary style awareness, they remain susceptible to interference from font type, textual content, and source provenance, hindering accurate aesthetic judgments grounded in nuanced brushstroke details.
This study investigates whether visual text styling—such as font, color, and size—can still interfere with the conceptual attribute descriptions generated by large vision-language models (LVLMs), even when the underlying concepts are correctly recognized. Through controlled experiments, the work systematically compares the effects of functional versus decorative text styles on LVLM outputs and reveals, for the first time, a phenomenon termed “style leakage”: despite accurate semantic recognition, decorative styling significantly distorts the generated attribute descriptions. This finding underscores the subtle yet substantial influence of visual appearance on semantic reasoning in multimodal systems and calls for the integration of style-aware evaluation and mitigation mechanisms in future model development and deployment.