Evidence-Ledger Adjudication for Claim-Evidence Traceability

📅 2026-07-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that AI-generated claims often lack reliable evidential support, while manual verification remains inefficient and poorly traceable. To tackle this, the authors propose an Evidence Ledger Adjudication pipeline—the first to integrate an evidence ledger mechanism into AI-assisted writing—automatically pairing each claim with a heterogeneous evidence bundle and classifying their relationship as supporting, contradicting, or mixed. Problematic claims are then routed back for revision. By constructing a blind-test benchmark based on external annotations and leveraging relation classification alongside evidence provenance techniques, the system enables automated consistency assessment and routing between claims and evidence. Evaluated on 2,335 samples, the approach achieves a relation accuracy of 0.676 and a macro F1 of 0.601, significantly outperforming baselines; it correctly identifies 89.2% of problematic claims while misrouting only 32.8% of valid ones, substantially enhancing the management of claim credibility.
📝 Abstract
AI agents can draft claims faster than authors can check whether the cited or retrieved evidence supports them. We study evidence-ledger adjudication: a claim-evidence traceability workflow that pairs each claim with an evidence packet, assigns a support relation, and routes unsupported, contradicted, or mixed-evidence claims back to the author. The empirical core is a 2,335-row blind benchmark built from independent external labels in AVeriTeC, CLIMATE-FEVER, and SciFact. Gold relations and source evidence labels are hidden during prediction and joined only for scoring. On this benchmark, the agent evidence-ledger condition achieves 0.676 relation accuracy and 0.601 macro-F1, compared with 0.383 accuracy and 0.303 macro-F1 for the best non-agent baseline. It also routes 1270/1435 claims whose gold labels indicate contradiction, missing evidence, or mixed evidence, while routing 295/900 supported claims. These results show that evidence-ledger adjudication can turn heterogeneous evidence packets into an auditable traceability layer for AI-assisted writing.
Problem

Research questions and friction points this paper is trying to address.

claim-evidence traceability
evidence adjudication
AI-assisted writing
support relation
evidence verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

evidence-ledger adjudication
claim-evidence traceability
AI-assisted writing
support relation classification
auditable evidence layer
🔎 Similar Papers
No similar papers found.
G
Gengyu Chen
Carnegie Mellon University
Y
Yongjie Yu
Carnegie Mellon University
W
Weiling Wang
Syracuse University