Score
Designs and builds algorithms and systems that detect, localize, segment, and classify structural elements and regions on document pages—such as headings, paragraphs, tables, figures, and form fields—across single- and multi-page documents. Analyzes spatial relationships and grouping between detected regions to produce segmented document representations used for extraction, rendering, or interactive form handling.
Existing PDF parsing tools are primarily designed for academic papers and struggle to accurately process pedagogical documents—such as legal textbooks—that exhibit complex, implicitly structured hierarchies. To address this, we propose a hierarchical text segmentation framework integrating structure-aware preprocessing with large language models (LLMs). Our method jointly leverages OCR-based heading detection, XML structural feature extraction, and contextual semantic modeling to infer implicit heading hierarchies without requiring explicit table-of-contents input. Compared to pure LLM–based or traditional rule-based approaches, our framework significantly reduces false positives and improves segmentation accuracy. When high-quality metadata is available, a supplementary table-of-contents–driven strategy further enhances performance. The source code and benchmark dataset are publicly released to support reproducible research.
Existing document chunking methods rely solely on semantic similarity while ignoring spatial layout, leading to suboptimal segmentation in complex documents (e.g., multi-column or image-text interleaved layouts) and poor controllability of chunk length for LLM input constraints. To address this, we propose a structure-aware adaptive chunking method: it jointly models textual bounding boxes, semantic embeddings (BERT/MPNet), and spatial relationships to construct a weighted heterogeneous graph, then applies spectral clustering for semantic-structural co-optimization; additionally, a dynamic length truncation strategy enforces strict token limits. This is the first document chunking framework that explicitly integrates spatial structure modeling with semantic coherence. Experiments demonstrate a 12.7% F1-score improvement on multi-layout benchmarks, 98.3% intra-chunk semantic consistency, and 100% compliance with prescribed token constraints.
To address the scarcity of real-world data and the difficulty small language models (SLMs) face in effectively modeling spatial structures for semi-structured document layout understanding, this paper proposes a spatial information integration framework tailored for SLMs. Methodologically: (1) it introduces a coordinate-based synthetic layout generation mechanism to alleviate annotation scarcity; (2) it designs a bounding-box-aware text encoder to enable lightweight joint modeling of spatial and semantic information. Contributions include: (i) the first spatial information fusion paradigm specifically customized for SLMs; and (ii) empirical validation that synthetic layout data significantly improves downstream performance—achieving superior layout generation metrics compared to LayoutTransformer and substantially boosting multi-class document classification accuracy through bounding-box integration.
This study addresses the common practice in web page segmentation of naively fusing visual and structural cues without a systematic analysis of DOM coordinate representations and their efficacy. Through a comprehensive evaluation of various DOM coordinates—such as tree position and visual layout—the work compares the performance of single versus composite coordinate vectors combined with multiple clustering algorithms. The findings reveal that single-coordinate representations consistently outperform complex fused vectors, and that visual coordinates underperform DOM-based coordinates by 20–30%, challenging the prevailing reliance on visual information. Experimental results demonstrate that, when coordinate representations, clustering algorithms, and page types are appropriately aligned, segmentation accuracy reaches 74%, a 20% improvement over baseline methods, with 68.2% of optimal configurations employing a single coordinate vector.
This work addresses high-accuracy conversion of unstructured/semi-structured documents (e.g., contracts, academic papers, invoices) into structured, machine-readable data. Method: We systematically survey and empirically compare modular pipeline approaches against end-to-end multimodal large models, proposing a unified framework integrating OCR, layout analysis (LayoutParser), graph neural networks, vision-language models (VLMs), and specialized formula/table recognition. We identify and characterize core bottlenecks—layout understanding, dense text recognition, and cross-modal alignment—for the first time. Contribution/Results: We establish a comprehensive analytical framework covering methodology, challenges, and benchmarks, revealing >32% performance gaps of current SOTA on complex layouts (e.g., multi-column, nested tables). We propose a “dual-driven” evolution path emphasizing both data diversity and scale, and open-source a larger annotated dataset to significantly advance knowledge base construction and training-data generation for large models.
This work addresses the challenge of accurately segmenting handwritten and printed text in document digitization under the computational constraints of edge devices, where existing deep learning approaches incur high computational costs and are difficult to deploy in lightweight settings. To overcome this, the authors propose a novel lightweight segmentation framework that operates without deep neural networks. The method first extracts semantically coherent text regions via sentence-level connected component segmentation, then introduces a region-aware handwriting descriptor (RHD) to effectively capture the variability inherent in handwriting. Finally, a conventional classifier is employed for efficient discrimination between handwritten and printed text. Evaluated on both the newly curated MAD-HPTS dataset and the public PHD-AS benchmark, the approach outperforms state-of-the-art methods, achieving over 8× inference speedup with only a 1.4% drop in accuracy, thereby substantially reducing computational overhead and enabling practical edge deployment.
This work addresses the limitations of existing methods in question answering over semi-structured documents containing tables, figures, and hierarchical paragraphs, which often suffer from fragmented OCR outputs, inadequate hierarchical modeling, and insufficient cross-region alignment. To overcome these challenges, we propose MoDora, a novel system that reconstructs OCR results through local alignment aggregation, explicitly models inter-component hierarchical and spatial relationships via a Component Correlation Tree (CCTree), and introduces a question-type-aware hybrid retrieval mechanism that fuses semantic and layout information. Evaluated on multiple benchmarks, MoDora significantly outperforms current state-of-the-art approaches, achieving absolute accuracy gains of 5.97% to 61.07% and enabling precise, layout- and semantics-driven document understanding.
This work addresses the challenge of accurately and efficiently digitizing complex documents containing handwritten content, irregular tables, and heterogeneous layouts—tasks that remain difficult for conventional OCR systems and current large language models. The authors propose an interactive document digitization system that integrates layout-aware parsing, OCR, and a large language model, enhanced by a user-in-the-loop correction propagation mechanism. Leveraging layout-aware inference, the system automatically generalizes user edits or natural language instructions applied to a local region to structurally similar regions across the document. In a user study (n=12), this approach significantly improved correction efficiency, reduced repetitive manual operations, and enabled more controllable and effective reconstruction of document structure and content.
This study addresses the challenge of high-precision, non-invasive reconstruction of fragile paper fragments in cultural heritage by proposing a human–robot collaborative real-time reconstruction system. The system integrates a vacuum-based collaborative robot with the detector-free feature matching algorithm SE2-LoFTR, enabling vision-guided fragment alignment and assembly in either manual or fully automatic modes. It innovatively combines AI-driven analysis—leveraging image segmentation and local feature matching—with a safe vacuum gripper mechanism and high-accuracy robotic control. Experimental results demonstrate a repeatability positioning accuracy of 0.57 mm on fragments as small as 8 cm² and confirm the superior robustness of SE2-LoFTR under conditions involving rotation, scaling, and partial damage.
This study addresses the challenges of multimodal fusion in visually rich documents and the lack of standardized evaluation in existing approaches by conducting controlled comparative and ablation experiments on representative models—including LayoutLMv3, Donut, Qwen3-VL-32B-Instruct, and Qwen3-32B—within a unified experimental framework. For the first time, it systematically compares OCR-dependent and OCR-free multimodal methods, revealing that specialized multimodal Transformers significantly outperform large language models. The findings further demonstrate that visual information plays a dominant role in layout-intensive document classification, whereas OCR-derived text provides only auxiliary value. These results offer empirical guidance for model selection and feature composition in multimodal document understanding.