🤖 AI Summary
This study addresses the dual challenges of costly manual schema tuning and unreliable large language model execution in few-shot information extraction from scientific PDFs. To this end, it proposes CPSE, a novel framework that introduces a structure-preserving semantic calibration paradigm. By decoupling structural contracts from field semantics, CPSE jointly optimizes prompts and descriptions while decomposing the task into identity discovery and record completion stages to effectively circumvent blind trial-and-error. This work further integrates prompt engineering with checklist-conditioned parsing to achieve efficient low-resource extraction. Experimental results on polymer documents demonstrate a 9.93-point improvement over baselines, with effectiveness rigorously validated through independent review and blind testing.
📝 Abstract
A well-designed extraction schema is not necessarily ready for reliable LLM execution. When only limited verified extractions are available, manually tuning hundreds of field definitions through trial and error is costly. We frame this problem as few-instance schema calibration: adapting the operational semantics of an existing schema from a few annotated documents while preserving its structural contract. We introduce CPSE, a contract-preserving semantic extraction framework that jointly calibrates extraction prompts and field-level semantic descriptions from a few gold annotations. CPSE decomposes the schema into an invariant structural contract and mutable field semantics, and further separates identity discovery from record completion using manifest-conditioned resolution. On expert-annotated polymer-science documents, CPSE improves extraction by 9.93 points over an execution-matched baseline, with consistent gains under an independent judge and in a blinded expert audit. These results show that CPSE enables low-resource schema execution while preserving the output structure required downstream.