OmniHandwritingOCR: A Diagnostic Benchmark for Evaluating Multimodal LLMs in Handwritten OCR Scenarios

📅 2026-08-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决现有OCR基准对真实手写场景覆盖不足的问题,本文通过OmniHandwritingOCR基准测试多模态大语言模型在手写识别上的表现,涵盖六项子任务和十二个数据集。
📝 Abstract
Multimodal large language models (MLLMs) are increasingly used as OCR systems in document and knowledge-processing pipelines, but their ability to faithfully read real handwriting remains underexplored. Existing OCR benchmarks focus largely on printed text or clean single-line inputs, leaving limited coverage of realistic handwritten OCR scenarios such as multilingual handwriting, writer errors, and structurally complex mathematical expressions. We introduce OmniHandwritingOCR, a diagnostic benchmark for evaluating MLLMs and OCR systems on handwritten OCR. It covers handwritten text recognition and handwritten mathematical expression recognition across six subtasks and twelve subsets, totaling 77.57K labeled images from public datasets and newly collected student writings. A key component is a difficulty-stratified multi-line formula corpus designed to test robustness under increasing structural complexity. We evaluate thirteen open- and closed-source systems with five complementary metrics under a unified protocol. Results show that current systems remain far from faithful transcription: performance drops sharply on complex multi-line formulas, model rankings vary across language and formula settings, and several generative models hallucinate plausible but visually unsupported corrections. OmniHandwritingOCR provides a challenging testbed for diagnosing language, content, structural, and visual-grounding failure modes of multimodal models in handwritten OCR scenarios.
Problem

Research questions and friction points this paper is trying to address.

Handwritten OCR
Multimodal Large Language Models
Realistic Handwriting Recognition
Structural Complexity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Large Language Models
Handwritten OCR
Difficulty-Stratified Corpus
Unified Evaluation Protocol
Structural Complexity
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Z
Zinuo Guo
East China Normal University
Min Zhang
Min Zhang
East China Normal University
B
Bo Jiang
East China Normal University