Evidence Absence Is Not Evidence Insufficiency: Diagnosing NEI Construction Artifacts in Fact Verification

📅 2026-05-26
📈 Citations: 0
Influential: 0
📄 PDF

career value

151K/year
🤖 AI Summary
This work addresses a critical flaw in fact verification models, which often conflate “missing evidence” with “insufficient evidence,” thereby learning constructional biases rather than genuine reasoning capabilities. To systematically disentangle these two types of Not Enough Information (NEI) instances, the authors propose NEI-CAP, a diagnostic protocol that categorizes NEI samples by their construction family, incorporates human arbitration, evaluates cross-construction generalization, and analyzes claim-confidence alignment under fixed statements. Experiments reveal that models trained solely on shortcut-based constructions fail to recognize semantically grounded cases of insufficient evidence, while mixed-construction training only partially mitigates this issue. Moreover, aggregate NEI scores obscure model performance on specific subtasks. This study establishes a transferable, construction-aware diagnostic framework, offering a new paradigm for robust evaluation of fact verification systems.
📝 Abstract
Evidence absence is not evidence insufficiency, but fact verification benchmarks can make them observationally similar. The Not Enough Information (NEI) label is often operationalized through different evidence conditions, and that choice silently determines what a verifier learns and what its score can hide. We introduce NEI-CAP, a construction-aware diagnostic protocol for insufficient-evidence evaluation. Each NEI example carries the construction family that produced it; NEI-CAP audits shortcut cues, validates hard cases through human adjudication, and tests whether competence transfers across constructions. We instantiate the protocol in SciFact-style scientific verification, with FEVER and HoVer as bounded external controls. Across these settings, NEI competence does not transfer reliably: models trained on shortcut-prone constructions fail to recognize semantically related insufficient evidence, and mixed-construction training narrows but does not close the gap. Fixed-claim diagnostics further show that the evidence condition shifts confidence in the reference Support/Refute label, not only NEI recall, so an aggregate NEI score can hide which problem a model has actually solved.
Problem

Research questions and friction points this paper is trying to address.

fact verification
Not Enough Information
evidence insufficiency
construction artifacts
evaluation bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

NEI-CAP
construction artifacts
fact verification
evidence insufficiency
diagnostic protocol
🔎 Similar Papers