Score
Design and build evaluation protocols, question templates, and metrics for multi-hop question answering that require chaining information across multiple documents or time steps. Implement and analyze end-to-end systems by combining complementary retrievers with language models, measure performance on cross-source and temporal reasoning, and diagnose failure modes in multi-step retrieval and reasoning.
This work addresses the lack of a unified analytical framework for retrieval and reasoning pipelines in multi-hop question answering, which has hindered systematic comparison across methods. We propose, for the first time, a four-axis design framework that treats the execution process as the fundamental unit of analysis, encompassing execution plans, index structures, control strategies, and termination criteria. Through a comprehensive literature review and ablation studies, we structurally map prominent approaches—including RAG and agent-based systems—onto this framework using benchmarks such as HotpotQA, revealing consistent trade-offs among effectiveness, efficiency, and evidence faithfulness. Our framework systematically organizes existing design choices, identifies reproducible empirical trends, and highlights key challenges, including structure-aware planning and transferable control strategies.
To address the challenges of poor coordination between reasoning and external knowledge retrieval, as well as limited interpretability in multi-hop question answering, this paper proposes a dynamic knowledge graph construction framework that integrates question decomposition with breadth-first search (BFS)-guided reasoning. The method retrieves and structurally organizes external knowledge on-demand during inference, explicitly generating multi-hop evidence chains to enable synchronous evolution of reasoning paths and their supporting knowledge. Core techniques include retrieval-augmented generation, dynamic knowledge subgraph construction, stepwise question decomposition, and BFS-driven iterative reasoning. Evaluated on MuSiQue, 2WikiMultiHopQA, and HotpotQA, the approach achieves state-of-the-art performance: average exact match (EM) improves by 2.57% and F1 by 2.13%; on HotpotQA specifically, EM and F1 increase by 4.70% and 3.44%, respectively—demonstrating substantial gains in both accuracy and interpretability.
Multi-hop question answering (QA) demands identifying multi-hop reasoning requirements, performing reading comprehension and logical inference, and integrating knowledge across documents—posing significant challenges for both humans and large language models. This study conducts a crowdsourced experiment to quantitatively assess human performance across constituent subtasks: while knowledge integration achieves high accuracy (97%), multi-hop requirement identification drops to 67%, and semantic mismatches (e.g., answering “where” with “when”) persist in both single- and multi-hop QA. Results reveal that humans excel at cross-document knowledge fusion but struggle with reasoning path planning and semantic alignment. Based on these findings, we propose a “complementary human-AI collaboration” design paradigm: AI handles reasoning trigger identification and semantic constraint enforcement, while humans lead knowledge integration. This work provides empirical grounding and architectural guidance for building interpretable, robust hybrid intelligence systems.
To address the “lost-in-retrieval” problem in retrieval-augmented multi-hop question answering—where incomplete entity coverage during sub-question decomposition leads to retrieval failure and broken reasoning chains—this paper proposes ChainRAG. The framework introduces a sentence-level graph structure for hop-wise entity completion and establishes a closed-loop, chain-like mechanism integrating retrieval, query rewriting, and feedback to ensure cross-hop information propagation and retrieval completeness. It further designs entity-aware sub-question rewriting and multi-hop answer aggregation strategies. ChainRAG is compatible with mainstream large language models, including GPT-4o-mini, Qwen2.5-72B, and GLM-4-Plus. Extensive experiments on MuSiQue, 2Wiki, and HotpotQA demonstrate consistent superiority over state-of-the-art baselines, achieving significant improvements in both answer accuracy and retrieval efficiency.
Multi-hop question answering faces challenges including difficulty in acquiring all necessary evidence in a single retrieval step, information overload from iterative retrieval, and lack of explicit process tracking. This paper proposes ReSP, an iterative retrieval-augmented generation framework. ReSP introduces a novel dual-objective summarizer that concurrently compresses information relevant to both the global question and dynamically generated sub-questions. It incorporates a retrieval trajectory memory mechanism to explicitly model and record retrieval paths, thereby suppressing redundant planning. Additionally, it leverages LLM-driven sub-question decomposition and verification for lightweight, efficient iterative information integration. Evaluated on HotpotQA and 2WikiMultihopQA, ReSP achieves significant improvements over state-of-the-art methods, demonstrates strong robustness to long contexts, improves inference efficiency by 23%, and attains a summary compression rate of 68%.
Current retrieval-augmented generation (RAG) systems exhibit fragility in multi-hop, knowledge-intensive question answering due to implicit reasoning that often leads to retrieval drift and unreliable self-reflection. This work proposes reframing multi-hop RAG as a program synthesis and execution task, wherein reasoning is explicitly modeled through the generation of executable Python programs. Intermediate reasoning states are exposed as program variables, and deterministic feedback from program execution drives self-correction and adaptive retrieval. This approach achieves, for the first time, explicit and verifiable reasoning traces without requiring additional training. It significantly outperforms strong baselines across five benchmarks—including PopQA and HotpotQA—with particularly pronounced gains on compositional multi-hop tasks, and supports both zero-shot and reinforcement learning settings.
This work addresses the limitations of existing approaches in multi-hop question answering, which often produce inaccurate answers due to deviated reasoning paths or retrieved information lacking practical utility assessment. To overcome these issues, the authors propose an end-to-end reasoning path navigation framework that leverages a fine-tuned Llama3.1-8B model for controllable sub-question decomposition and introduces a dependency-tree-driven, structured entity-aware retrieval mechanism. This mechanism explicitly quantifies the informational contribution of retrieved documents to the reasoning process, thereby transcending the constraints of conventional similarity-based retrieval. Evaluated on three standard multi-hop QA benchmarks, the proposed method significantly outperforms current state-of-the-art models, demonstrating superior accuracy, robustness, and generalization capability.