Adapt Semantics, Not Structure: Few-Instance Schema Calibration for Scientific PDF Extraction

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the dual challenges of costly manual schema tuning and unreliable large language model execution in few-shot information extraction from scientific PDFs. To this end, it proposes CPSE, a novel framework that introduces a structure-preserving semantic calibration paradigm. By decoupling structural contracts from field semantics, CPSE jointly optimizes prompts and descriptions while decomposing the task into identity discovery and record completion stages to effectively circumvent blind trial-and-error. This work further integrates prompt engineering with checklist-conditioned parsing to achieve efficient low-resource extraction. Experimental results on polymer documents demonstrate a 9.93-point improvement over baselines, with effectiveness rigorously validated through independent review and blind testing.
📝 Abstract
A well-designed extraction schema is not necessarily ready for reliable LLM execution. When only limited verified extractions are available, manually tuning hundreds of field definitions through trial and error is costly. We frame this problem as few-instance schema calibration: adapting the operational semantics of an existing schema from a few annotated documents while preserving its structural contract. We introduce CPSE, a contract-preserving semantic extraction framework that jointly calibrates extraction prompts and field-level semantic descriptions from a few gold annotations. CPSE decomposes the schema into an invariant structural contract and mutable field semantics, and further separates identity discovery from record completion using manifest-conditioned resolution. On expert-annotated polymer-science documents, CPSE improves extraction by 9.93 points over an execution-matched baseline, with consistent gains under an independent judge and in a blinded expert audit. These results show that CPSE enables low-resource schema execution while preserving the output structure required downstream.
Problem

Research questions and friction points this paper is trying to address.

Schema Calibration
Scientific PDF Extraction
Few-Instance Learning
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Schema Calibration
Contract-Preserving Extraction
Few-Instance Learning
Scientific PDF Extraction
Manifest-Conditioned Resolution
🔎 Similar Papers
2024-06-08Annual Meeting of the Association for Computational LinguisticsCitations: 2
Z
Zixiao Dong
School of Computer Science and Technology, University of Science and Technology of China; State Key Laboratory of Cognitive Intelligence
W
Wei Yang
University of Science and Technology of China; State Key Laboratory of Cognitive Intelligence
Z
Zihao Liu
School of Computer Science and Technology, University of Science and Technology of China; State Key Laboratory of Cognitive Intelligence
C
Chenshu Li
School of Computer Science and Technology, University of Science and Technology of China; State Key Laboratory of Cognitive Intelligence
L
Longzhang Liu
School of Computer Science and Technology, University of Science and Technology of China; State Key Laboratory of Cognitive Intelligence
Tao Tan
Tao Tan
FCA MPU
Medical Imaging AI
Hong Xie
Hong Xie
University of Science and Technology of China (USTC)
Data Science/MiningOnline Learning