End-to-End Text Line Detection and Ordering

📅 2026-06-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of traditional historical document recognition approaches, which treat text line detection and reading order prediction as separate tasks relying on handcrafted rules and struggle with complex layouts such as marginalia, multi-column texts, and tables. The authors propose Orli, a novel model that unifies these tasks into an end-to-end image-to-sequence framework, autoregressively generating text line baselines directly in reading order. Orli represents baselines using chord-based parameterization and incorporates an iterative refinement head with a local visual refinement module. Trained on a diverse corpus of 196,691 pages spanning ten writing systems, Orli slightly surpasses state-of-the-art performance on cBAD text line detection without fine-tuning and achieves near-perfect zero-shot generalization on reading order prediction, while requiring only minimal fine-tuning to adapt to specialized out-of-domain layouts.
📝 Abstract
Practical text-recognition pipelines for historical documents typically decompose layout analysis into line detection followed by a separate reading-order step, with the latter most often handled by a hand-coded geometric heuristic that struggles with marginalia, multiple columns, tables, and source-specific editorial conventions. This article introduces Orli (Ordered Regression of Lines), an end-to-end model that casts both sub-tasks as a single image-to-sequence problem: from a page image, Orli autoregressively generates text-line baselines directly in reading order. Baselines are represented in a chord-frame parameterization that anchors a line's position, orientation, and extent while encoding local geometry through perpendicular offsets; an iterative refinement head and a local visual refiner produce the final curve. Trained on a heterogeneous corpus of 196,691 pages spanning ten writing systems, Orli marginally exceeds the previously reported state of the art for cBAD line detection without dataset-specific training, reaches near perfect coverage and ordering on multiple reading-order benchmarks zero-shot, and adapts to more specialized out-of-domain layouts with limited fine-tuning. The method's source code and model weights are available under an open license at https://github.com/mittagessen/orli.
Problem

Research questions and friction points this paper is trying to address.

text line detection
reading order
historical documents
layout analysis
geometric heuristics
Innovation

Methods, ideas, or system contributions that make the work stand out.

end-to-end
text line detection
reading order
chord-frame parameterization
autoregressive generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Benjamin Kiessling
ALMAnaCH, Inria, France