REALHOP: Rethinking Multi-Hop Reasoning Evaluation via Behavioral Auditing

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of existing benchmarks that evaluate multi-hop reasoning solely by answer correctness, which fails to verify whether models genuinely rely on prescribed evidence chains and remains susceptible to shortcut learning and memorization biases. We propose a "diagnose-construct-verify" framework that compels models to reason along complete evidence chains through entity rebinding, relation factorization, and competing path injection. Furthermore, we introduce the Behavioral Necessity Rate (BNR) to quantify evidence dependency, integrating evidence removal interventions with structural-semantic dual verification to shift the evaluation paradigm from outcome-oriented to process-verifiable. Experiments demonstrate that BNR increases dramatically from 27.4% to 94.4%, significantly widening performance gaps among models and establishing evidence dependency as a core metric for evaluating multi-hop reasoning.
πŸ“ Abstract
Complex questions often require multi-hop reasoning that connects facts distributed across sources or distant regions of a long context through intermediate steps. Benchmarks commonly evaluate this ability with questions built around predefined reasoning chains, treating a correct answer as evidence that the intended composition was used. Yet answer correctness alone leaves open whether success depends on the evidence associated with each intended step: models may instead rely on memorized associations, shorter paths, or partial evidence. We examine this dependence using the Behavioral Necessity Rate (BNR), which measures how often targeted evidence removal prevents answer recovery on initially correct instances. Across five existing benchmarks, panel-mean BNR ranges from 16.6% to 48.9%, exposing a substantial gap between annotated structure and observed dependence. Guided by this diagnosis, we introduce REALHOP, a diagnose-construct-verify framework that rebinds entities, factorizes selected relations, adds complete competing paths, and places evidence at traceable locations. Structural and semantic checks precede freezing; behavioral interventions follow. On 790 paired MuSiQue questions, REALHOP raises panel-mean BNR from 27.4% to 94.4% while retaining high Full accuracy. It also yields high BNR on REALHOP-FRAMES and REALHOP-LONGBENCH. On 216 long-context questions, the matched multiple-choice spread across 16 models grows from 13.9 to 59.2 points and persists under repeated open-ended evaluation. Together, these results show that a conceptually coherent chain and a correct final answer do not by themselves establish multi-hop reasoning. Verifying that success depends on every intended hop is therefore as fundamental to multi-hop evaluation as measuring answer accuracy itself.
Problem

Research questions and friction points this paper is trying to address.

multi-hop reasoning
evaluation benchmark
behavioral auditing
evidence dependence
reasoning chain
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-hop Reasoning
Behavioral Necessity Rate
Behavioral Auditing
Evaluation Benchmark
Diagnose-Construct-Verify Framework
πŸ”Ž Similar Papers
2024-02-26Annual Meeting of the Association for Computational LinguisticsCitations: 97
J
Jiawen Tao
Hunyuan Team, Tencent
X
Xiaokun Yuan
Hunyuan Team, Tencent
Y
Yaoming Li
Peking University
C
Chenxu Liu
Hunyuan Team, Tencent
Mengzhou Wu
Mengzhou Wu
Peking University
Software EngineeringLarge Language Model
Tong Yang
Tong Yang
Peking University, Beijing, China. PKU. εŒ—δΊ¬ε€§ε­¦
SketchNetwork measurementBloom filterIP lookupHash Table
M
Maxm Pan
Hunyuan Team, Tencent