Locally Sound, Globally Insufficient: The Local-Global Gap in Multi-Hop Reasoning

๐Ÿ“… 2026-09-26
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Multi-hop reasoning frequently suffers from locally plausible yet globally invalid inferences, a discrepancy that traditional verifiers struggle to detect. To address this challenge, this work proposes E-Closure, a method that bridges the local-global gap by formalizing the dependencies among evidence, questions, and answers. During training, E-Closure supervises reasoning dependencies and integrates counterfactual generation with bidirectional switching constraints to achieve end-to-end optimization, thereby reinforcing the model's understanding of reasoning integrity. Experimental evaluations across three benchmarks and three backbone architectures demonstrate that the proposed approach achieves an average accuracy of 92.8% and a trajectory reliability of 89.0%, while reducing the local-global discrepancy rate to 6.2%.
๐Ÿ“ Abstract
Reliable multi-hop reasoning requires more than locally supported steps: a trace can be sound at every reasoning step yet still fail to answer the question as a whole. We call this failure regime the local-global gap (LGG), in which the trace is locally sound yet globally insufficient. Local soundness requires each step to be supported by the available evidence and preceding steps, whereas global sufficiency requires the reasoning trace to align with the question and establish the submitted answer. In a human-adjudicated diagnostic of 2,598 responses across three multi-hop QA benchmarks and three models, we find that the LGG occurs in every benchmark-model combination and accounts for nearly half of globally insufficient responses overall. However, conventional faithfulness verifiers that check claims against the evidence largely miss these failures: at thresholds retaining at least 95% of reliable traces, recall for LGG cases is substantially lower than that for locally unsound traces. To address these failures, we formalize three dependencies for reliable reasoning: evidence-to-step support, question-to-trace alignment, and trace-to-answer closure. Instead of post-hoc diagnosis, we introduce E-Closure to supervise these dependencies during training, combining generation supervision on supported original and counterfactual responses with bidirectional switching constraints. Averaged over three benchmarks and three backbones, existing fine-tuning baselines improve accuracy and local soundness over the base models, but at the cost of global sufficiency. E-Closure improves both: among all fine-tuned methods, it achieves the highest average accuracy (92.8%) and trace reliability (89.0%) while yielding the lowest LGG rate (6.2%).
Problem

Research questions and friction points this paper is trying to address.

multi-hop reasoning
local-global gap
global sufficiency
faithfulness verification
trace reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Hop Reasoning
Local-Global Gap
E-Closure
Faithfulness Verification
Counterfactual Supervision
๐Ÿ”Ž Similar Papers
2024-02-26Annual Meeting of the Association for Computational LinguisticsCitations: 97
๐Ÿ’ผ Related Jobs
No related jobs found.