dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model

📅 2025-12-02
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing document layout parsing methods rely on multi-stage pipelines, which suffer from error propagation and lack task-coordinated optimization. To address these limitations, we propose UniDoc—the first end-to-end multilingual vision-language model—unifying layout detection, text recognition, and relational understanding within a single architecture. We further design a scalable synthetic data engine to generate large-scale, high-quality training data spanning 126 languages. Through joint multi-task training and cross-modal alignment, UniDoc significantly enhances robustness and generalization on complex documents. It achieves state-of-the-art performance on OmniDocBench and outperforms the strongest baseline by 7.4 percentage points on our newly constructed multilingual benchmark, XDocParse, demonstrating superior multilingual comprehension and parsing capability.

Technology Category

Natural Language Processing: Language Grounding & Multi-modal NLPMachine Learning: Large Multimodal Models (LMMs)Computer Vision: Multi-modal Vision

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Document Layout Parsing serves as a critical gateway for Artificial Intelligence (AI) to access and interpret the world's vast stores of structured knowledge. This process,which encompasses layout detection, text recognition, and relational understanding, is particularly crucial for empowering next-generation Vision-Language Models. Current methods, however, rely on fragmented, multi-stage pipelines that suffer from error propagation and fail to leverage the synergies of joint training. In this paper, we introduce dots.ocr, a single Vision-Language Model that, for the first time, demonstrates the advantages of jointly learning three core tasks within a unified, end-to-end framework. This is made possible by a highly scalable data engine that synthesizes a vast multilingual corpus, empowering the model to deliver robust performance across a wide array of tasks, encompassing diverse languages, layouts, and domains. The efficacy of our unified paradigm is validated by state-of-the-art performance on the comprehensive OmniDocBench. Furthermore, to catalyze research in global document intelligence, we introduce XDocParse, a challenging new benchmark spanning 126 languages. On this testbed, dots.ocr establishes a powerful new baseline, outperforming the next-best competitor by a remarkable +7.4 point margin and proving its unparalleled multilingual capabilities.
Problem

Research questions and friction points this paper is trying to address.

Unifies layout detection, text recognition, and relational understanding in a single model
Overcomes fragmented multi-stage pipelines that cause error propagation
Enables robust multilingual document parsing across diverse layouts and domains
Innovation

Methods, ideas, or system contributions that make the work stand out.

Single Vision-Language Model for unified end-to-end layout parsing
Scalable data engine synthesizes vast multilingual training corpus
Unified framework jointly learns layout detection, text recognition, and relational understanding
🔎 Similar Papers
2024-07-17European Conference on Computer VisionCitations: 2
💼 Related Jobs
No related jobs found.
Xiaohongshu Inc
Y
Yumeng Li
hi lab, Xiaohongshu Inc
G
Guang Yang
hi lab, Xiaohongshu Inc
H
Hao Liu
hi lab, Xiaohongshu Inc
B
Bowen Wang
hi lab, Xiaohongshu Inc
C
Colin Zhang
hi lab, Xiaohongshu Inc