π€ AI Summary
This study addresses the limitation of existing benchmarks that evaluate multi-hop reasoning solely by answer correctness, which fails to verify whether models genuinely rely on prescribed evidence chains and remains susceptible to shortcut learning and memorization biases. We propose a "diagnose-construct-verify" framework that compels models to reason along complete evidence chains through entity rebinding, relation factorization, and competing path injection. Furthermore, we introduce the Behavioral Necessity Rate (BNR) to quantify evidence dependency, integrating evidence removal interventions with structural-semantic dual verification to shift the evaluation paradigm from outcome-oriented to process-verifiable. Experiments demonstrate that BNR increases dramatically from 27.4% to 94.4%, significantly widening performance gaps among models and establishing evidence dependency as a core metric for evaluating multi-hop reasoning.
π Abstract
Complex questions often require multi-hop reasoning that connects facts distributed across sources or distant regions of a long context through intermediate steps. Benchmarks commonly evaluate this ability with questions built around predefined reasoning chains, treating a correct answer as evidence that the intended composition was used. Yet answer correctness alone leaves open whether success depends on the evidence associated with each intended step: models may instead rely on memorized associations, shorter paths, or partial evidence. We examine this dependence using the Behavioral Necessity Rate (BNR), which measures how often targeted evidence removal prevents answer recovery on initially correct instances. Across five existing benchmarks, panel-mean BNR ranges from 16.6% to 48.9%, exposing a substantial gap between annotated structure and observed dependence. Guided by this diagnosis, we introduce REALHOP, a diagnose-construct-verify framework that rebinds entities, factorizes selected relations, adds complete competing paths, and places evidence at traceable locations. Structural and semantic checks precede freezing; behavioral interventions follow. On 790 paired MuSiQue questions, REALHOP raises panel-mean BNR from 27.4% to 94.4% while retaining high Full accuracy. It also yields high BNR on REALHOP-FRAMES and REALHOP-LONGBENCH. On 216 long-context questions, the matched multiple-choice spread across 16 models grows from 13.9 to 59.2 points and persists under repeated open-ended evaluation. Together, these results show that a conceptually coherent chain and a correct final answer do not by themselves establish multi-hop reasoning. Verifying that success depends on every intended hop is therefore as fundamental to multi-hop evaluation as measuring answer accuracy itself.