🤖 AI Summary
This study addresses the challenge of precisely localizing evidence text within scientific PDFs, where line breaks and formatting discrepancies disrupt alignment with the original layout. To overcome this, we propose a source-preserving alignment framework that achieves standardized matching through text normalization and introduces a line-break-aware token alignment algorithm to suppress noise. Furthermore, character-level span mapping is employed to project matched results back onto the geometric coordinates of the source characters, enabling precise highlighting. Experiments conducted on one thousand chemistry papers demonstrate that the proposed method attains an automatic localization rate of 92.6%, significantly outperforming existing baselines. These findings establish the framework as an efficient and reliable solution for fine-grained information extraction from scientific literature.
📝 Abstract
Scientific information-extraction systems often return a claim with an evidence string, which users must locate in the original PDF. This is challenging because the extracted evidence and PDF text layer are different representations: line wrapping, Unicode variants, superscripts, citation markers, and fragmented items alter text sequences and geometry. We present a source-preserving alignment framework: normalize text for robust matching while preserving provenance for accurate localization. It aligns evidence with normalized page text, maps matches back to source-character spans, and renders only their geometry. When exact alignment fails, line-break-aware token alignment recovers supported spans while excluding unmatched noise. Experiments on 1,020 chemistry papers show that the framework achieves a 92.6\% quote-level automatic localization rate, compared with 43.6\% for text search and 19.1\% for a precomputed bounding-box baseline. Component ablation confirms distinct contributions from normalization and approximate token alignment, while human verification assesses the visual correctness of returned highlights. Overall, these results demonstrate that reliable evidence verification requires robust matching and precise localization within a shared source-preserving alignment representation.