Source-preserving alignment for robust evidence localization in scientific PDFS

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of precisely localizing evidence text within scientific PDFs, where line breaks and formatting discrepancies disrupt alignment with the original layout. To overcome this, we propose a source-preserving alignment framework that achieves standardized matching through text normalization and introduces a line-break-aware token alignment algorithm to suppress noise. Furthermore, character-level span mapping is employed to project matched results back onto the geometric coordinates of the source characters, enabling precise highlighting. Experiments conducted on one thousand chemistry papers demonstrate that the proposed method attains an automatic localization rate of 92.6%, significantly outperforming existing baselines. These findings establish the framework as an efficient and reliable solution for fine-grained information extraction from scientific literature.
📝 Abstract
Scientific information-extraction systems often return a claim with an evidence string, which users must locate in the original PDF. This is challenging because the extracted evidence and PDF text layer are different representations: line wrapping, Unicode variants, superscripts, citation markers, and fragmented items alter text sequences and geometry. We present a source-preserving alignment framework: normalize text for robust matching while preserving provenance for accurate localization. It aligns evidence with normalized page text, maps matches back to source-character spans, and renders only their geometry. When exact alignment fails, line-break-aware token alignment recovers supported spans while excluding unmatched noise. Experiments on 1,020 chemistry papers show that the framework achieves a 92.6\% quote-level automatic localization rate, compared with 43.6\% for text search and 19.1\% for a precomputed bounding-box baseline. Component ablation confirms distinct contributions from normalization and approximate token alignment, while human verification assesses the visual correctness of returned highlights. Overall, these results demonstrate that reliable evidence verification requires robust matching and precise localization within a shared source-preserving alignment representation.
Problem

Research questions and friction points this paper is trying to address.

evidence localization
scientific PDFs
information extraction
text alignment
source-preserving
Innovation

Methods, ideas, or system contributions that make the work stand out.

Source-preserving alignment
Evidence localization
Text normalization
Line-break-aware token alignment
Scientific PDF
🔎 Similar Papers
Z
Zihao Liu
School of Computer Science and Technology, University of Science and Technology of China; State Key Laboratory of Cognitive Intelligence
W
Wei Yang
University of Science and Technology of China; State Key Laboratory of Cognitive Intelligence
Z
Zixiao Dong
School of Computer Science and Technology, University of Science and Technology of China; State Key Laboratory of Cognitive Intelligence
C
Chenshu Li
School of Computer Science and Technology, University of Science and Technology of China; State Key Laboratory of Cognitive Intelligence
L
Longzhang Liu
School of Computer Science and Technology, University of Science and Technology of China; State Key Laboratory of Cognitive Intelligence
Tao Tan
Tao Tan
FCA MPU
Medical Imaging AI
Hong Xie
Hong Xie
University of Science and Technology of China (USTC)
Data Science/MiningOnline Learning