Score
Design and evaluate algorithms that infer the sequential reading order of textual regions in a page or layout, producing an ordered sequence or directed transition graph and supporting tasks labeled as prediction, detection, recovery, or sequence modeling. This includes building methods that construct candidate-transition graphs, recover a global reading order from region-level inputs, can operate without training data, and handle non-rectangular or wrap-around layouts.
This study addresses the challenge of accurately inferring reading order in complex historical manuscripts—such as the Glossa Ordinaria—where annotations irregularly surround the main text, confounding existing layout analysis methods. The authors propose a training-free graph-based reasoning approach that constructs a directed candidate transition graph from OCR text lines, weights edges using signals from a causal language model and BERT’s next-sentence prediction, and recovers the global reading sequence via degree-constrained path cover combined with a max-regret inference rule to avoid greedy errors. This method uniquely integrates lightweight language model signals with max-regret reasoning, achieving 95% average recovery of true successor edges on Glossa layouts (versus 50% for XY-cut) and 88% macro-edge accuracy on the OmniDocBench multi-column subset—substantially outperforming LayoutReader (25%) and XY-cut (75%)—while exhibiting mirror invariance.
This work addresses the limitations of traditional historical document recognition approaches, which treat text line detection and reading order prediction as separate tasks relying on handcrafted rules and struggle with complex layouts such as marginalia, multi-column texts, and tables. The authors propose Orli, a novel model that unifies these tasks into an end-to-end image-to-sequence framework, autoregressively generating text line baselines directly in reading order. Orli represents baselines using chord-based parameterization and incorporates an iterative refinement head with a local visual refinement module. Trained on a diverse corpus of 196,691 pages spanning ten writing systems, Orli slightly surpasses state-of-the-art performance on cBAD text line detection without fine-tuning and achieves near-perfect zero-shot generalization on reading order prediction, while requiring only minimal fine-tuning to adapt to specialized out-of-domain layouts.
To address the scarcity of real-world data and the difficulty small language models (SLMs) face in effectively modeling spatial structures for semi-structured document layout understanding, this paper proposes a spatial information integration framework tailored for SLMs. Methodologically: (1) it introduces a coordinate-based synthetic layout generation mechanism to alleviate annotation scarcity; (2) it designs a bounding-box-aware text encoder to enable lightweight joint modeling of spatial and semantic information. Contributions include: (i) the first spatial information fusion paradigm specifically customized for SLMs; and (ii) empirical validation that synthetic layout data significantly improves downstream performance—achieving superior layout generation metrics compared to LayoutTransformer and substantially boosting multi-class document classification accuracy through bounding-box integration.
This study addresses the challenge of automatically reordering pages in Dutch Freedom of Information (WOO) PDF documents, which suffer from shuffled page sequences and heterogeneous content—including emails, legal texts, and tables—rendering semantic cues unreliable. The authors propose a page-embedding-based approach and systematically compare Pointer Networks, seq2seq Transformers, and a specialized pairwise ranking model. Their findings reveal that short and long documents require fundamentally different reordering strategies. By tailoring models to document length, they achieve a substantial performance gain for long documents (Kendall’s τ improves by 0.21), whereas generic seq2seq models severely degrade on longer inputs (τ drops to 0.014). The best-performing method attains τ scores ranging from 0.95 for 2–5-page documents to 0.72 for 15-page documents, while also uncovering why curriculum learning fails in this context.
This study investigates how graph description ordering affects large language models’ (LLMs) performance on graph reasoning tasks. We systematically evaluate four structured graph representations—adjacency list, edge list, adjacency matrix, and node-neighbor list—across six canonical graph tasks (e.g., shortest path, connectivity) and six state-of-the-art LLMs, using standardized prompting templates and controlled-variable experiments. Our results reveal, for the first time, that description ordering significantly impacts LLMs’ structural understanding of graphs, with pronounced task-specific sensitivity (e.g., high for shortest path, low for connectivity) and strong cross-model consistency. Based on these findings, we propose a *task-driven description ordering optimization paradigm*—a lightweight, zero-parameter, and transferable technique. Empirical evaluation shows that optimized ordering yields an average accuracy improvement of 12.7% across tasks and models, establishing it as an effective, implementation-efficient enhancement for graph reasoning with LLMs.
This study addresses the challenge of reconstructing reading order in historical Armenian newspapers, which is hindered by complex layouts and scarce linguistic resources. The authors propose a hybrid approach that integrates semantic region detection with a generative large language model (LLM). Their method combines geometric heuristics, YOLO-based layout parsing, the ECLAIR end-to-end model, and a semantic-LLM architecture, alongside a newly released Tesseract OCR model tailored for this domain. This framework effectively handles noisy OCR outputs and multi-page scenarios. As the first work to incorporate semantic guidance with LLMs for reading order reconstruction in low-resource settings, it significantly improves robustness, reducing ordering errors by up to 76% compared to the strongest geometric baseline, and offers a data-bootstrapping strategy to accelerate annotation workflows.
This work addresses the challenges of low accuracy and inefficiency in layout analysis caused by the heterogeneity of document elements, geometric deformations, and the entanglement of reading order in complex documents. To this end, we propose an end-to-end, single-stage multi-task framework based on RT-DETR that jointly models geometric and structural information within a unified query decoder. The model simultaneously performs element classification, bounding box detection, pixel-level segmentation, and reading order prediction. Evaluated on public benchmarks, our approach achieves state-of-the-art performance with an inference speed of 132.1 FPS and significantly enhances the quality of full-text reconstruction in downstream OCR tasks.
研究使用冻结的知识图谱辅助9B本地语言模型解决长篇小说推理问题,通过图引导证据导航提高多选题回答准确率。
本文通过双向图文转换方法,解决了图标题生成中结构信息保留与简洁性之间的矛盾,提出了一种轻量级结构化提示协议以生成更紧凑且一致的图标题。
This work addresses a critical limitation in existing reading order detection methods, which apply uniform supervision and overlook the inherent “positional disparity” in documents—where middle regions are significantly harder to model than beginning or ending segments—leading to degraded performance on complex layouts. To tackle this, we propose FocalOrder, the first framework to explicitly reveal and model this phenomenon. Our approach introduces Focal Preference Optimization (FPO) to dynamically identify challenging sequence transitions, combined with an exponential moving average mechanism and a difficulty-calibrated pairwise ranking loss to enhance global logical consistency. Evaluated on OmniDocBench v1.0 and Comp-HRDoc, FocalOrder achieves new state-of-the-art results; notably, its lightweight variant not only outperforms specialized baselines but also substantially surpasses large-scale general-purpose vision-language models.