Score
Designs methods that take a multi-hop or complex question and produce a hierarchy of atomic subquestions with explicit dependency relations and ordering. Builds the assignment and scheduling of those subquestions to retrieval or processing stages to minimize retrieval scope and initial ambiguity while preserving dependencies across subquestions.
This work addresses the lack of a unified analytical framework for retrieval and reasoning pipelines in multi-hop question answering, which has hindered systematic comparison across methods. We propose, for the first time, a four-axis design framework that treats the execution process as the fundamental unit of analysis, encompassing execution plans, index structures, control strategies, and termination criteria. Through a comprehensive literature review and ablation studies, we structurally map prominent approaches—including RAG and agent-based systems—onto this framework using benchmarks such as HotpotQA, revealing consistent trade-offs among effectiveness, efficiency, and evidence faithfulness. Our framework systematically organizes existing design choices, identifies reproducible empirical trends, and highlights key challenges, including structure-aware planning and transferable control strategies.
This work addresses the limitations of existing multi-hop question answering methods, which are prone to lexical ambiguity and often neglect the logical dependencies among reasoning steps, leading to incoherent inference. To overcome these issues, the authors propose the STRIDE framework, which decouples reasoning structure planning from entity grounding: a meta-planner first generates an abstract, entity-agnostic reasoning skeleton, and a dependency-aware supervision module dynamically schedules subtasks, adaptively selecting between retrieval and reasoning while integrating cross-branch information to enhance robustness. Furthermore, they introduce STRIDE-FT, a modular fine-tuning approach that requires no human annotations and leverages self-generated reasoning trajectories to optimize individual components. Experimental results demonstrate that STRIDE substantially improves multi-hop QA accuracy, and STRIDE-FT effectively enhances the reasoning capabilities of open-source large language models.
Existing RAG benchmarks largely overlook query difficulty, leading to inflated and unreliable evaluations. Robust assessment necessitates jointly considering answer quality, response diversity, and query difficulty. Method: We propose a fine-grained difficulty modeling framework based on multi-hop tree structures. Specifically: (1) we design a logically coherent multi-step query synthesis mechanism; (2) we formulate a difficulty metric integrating evidence distribution and reasoning depth; and (3) we establish a controllable data synthesis pipeline enabling difficulty-stratified dataset generation. Contribution/Results: To our knowledge, this is the first work to introduce a difficulty estimation algorithm that jointly evaluates retrieval and generation capabilities within RAG. Experiments show strong correlation (r > 0.85) between our estimated query difficulty and end-to-end RAG performance, significantly enhancing evaluation robustness and interpretability.
This work addresses the limitation of existing approaches in multi-hop question generation, which overlook the intrinsic duality between question generation and question answering, thereby constraining generation quality. To overcome this, the paper proposes the QQ framework, which explicitly models this dual relationship for the first time by jointly training multi-hop question generation and question answering within a unified architecture. The framework incorporates bidirectional alignment constraints and a contrastive learning mechanism to strengthen semantic correspondence between generated questions and their answers. Experimental results demonstrate that the proposed method significantly improves question quality on the HotpotQA and MuSiQue datasets, with both automatic metrics and human evaluations confirming its superiority over baseline approaches.
This work addresses the lack of rigorous mathematical formalization in decomposition-based reasoning for multi-hop question answering. It introduces operad theory into the analysis of large language model reasoning by constructing a question operad $Q$, where question templates are treated as operations and sub-answers as inputs to be composed. The question-answering model is thereby interpreted as an algebra over this operad, providing a formal framework for problem decomposition and composition. Building on this foundation, the paper proposes operadic consistency—a novel metric for evaluating the coherence of multi-step reasoning. Experiments across 12 large language models and 4 multi-hop question answering benchmarks demonstrate that this metric exhibits strong correlation with answer accuracy and significantly outperforms temperature-based self-consistency baselines.
This work addresses the limitations of existing approaches in multi-hop question answering, which often produce inaccurate answers due to deviated reasoning paths or retrieved information lacking practical utility assessment. To overcome these issues, the authors propose an end-to-end reasoning path navigation framework that leverages a fine-tuned Llama3.1-8B model for controllable sub-question decomposition and introduces a dependency-tree-driven, structured entity-aware retrieval mechanism. This mechanism explicitly quantifies the informational contribution of retrieved documents to the reasoning process, thereby transcending the constraints of conventional similarity-based retrieval. Evaluated on three standard multi-hop QA benchmarks, the proposed method significantly outperforms current state-of-the-art models, demonstrating superior accuracy, robustness, and generalization capability.
This work addresses a critical limitation in existing retrieval-augmented reasoning methods, which often neglect dependencies among sub-skills and rely excessively on strong-model distillation, leading to early derailment in multi-hop retrieval due to initial noise. To mitigate this, the authors propose a structured planning (Plan) mechanism that decomposes the original question into an ordered sequence of sub-questions prior to retrieval, ensuring each retrieval step targets a well-defined objective. They further introduce a distillation-free bootstrapping framework, wherein a small-scale seed model generates high-quality reasoning trajectories to activate the planning capability of larger models. Notably, the study reveals for the first time that identical reward signals induce heterogeneous reinforcement learning failure modes across models of different scales. The approach consistently activates the Plan mechanism across models ranging from 3B to 14B parameters and achieves sustained improvements over current baselines on multi-hop question answering benchmarks.
Current retrieval-augmented generation (RAG) systems exhibit fragility in multi-hop, knowledge-intensive question answering due to implicit reasoning that often leads to retrieval drift and unreliable self-reflection. This work proposes reframing multi-hop RAG as a program synthesis and execution task, wherein reasoning is explicitly modeled through the generation of executable Python programs. Intermediate reasoning states are exposed as program variables, and deterministic feedback from program execution drives self-correction and adaptive retrieval. This approach achieves, for the first time, explicit and verifiable reasoning traces without requiring additional training. It significantly outperforms strong baselines across five benchmarks—including PopQA and HotpotQA—with particularly pronounced gains on compositional multi-hop tasks, and supports both zero-shot and reinforcement learning settings.