Score
Design, build, or analyze retrieval systems that generate and score candidate evidence paths by expanding and matching from both the topic/query side and the answer/candidate side, using simulation of retrieval steps and prefix expansions to retrieve compact candidate paths while suppressing noisy mixed-type expansions. This competence includes implementing retrieval-forward simulation, answer-side final-hop relation matching, topic-side prefix expansion, and unified bidirectional modeling to combine forward and backward signals into a single retrieval pipeline.
This work addresses the current lack of a systematic integrative framework in reasoning-intensive retrieval (RIR) research. It proposes the first structured taxonomy that organizes existing RIR benchmarks according to knowledge domains and modalities, and establishes a methodological framework elucidating how reasoning capabilities are integrated into the retrieval pipeline and the associated trade-offs. By synergistically combining the reasoning power of large language models with conventional retrievers and re-rankers, the study systematically evaluates performance across multimodal and multidomain scenarios. This effort not only unifies the fragmented landscape of RIR research but also constructs a comprehensive research map, identifies critical challenges, and provides a clear roadmap for future exploration.
To address error propagation and answer bias arising from unidirectional retrieval-then-reasoning in multi-hop question answering, this paper proposes RetroRAG, the first framework introducing backtracking-style reasoning. Its core is an evidence backtracking mechanism: inferring entity-centric queries to dynamically revise retrieved evidence and reconstruct reasoning paths, enabling iterative refinement and dynamic reorganization of trustworthy evidence through coordinated multi-round retrieval-generation-evaluation cycles. This establishes a closed-loop “evidence curation–discovery–verification” process, substantially enhancing robustness and interpretability for complex reasoning. On mainstream multi-hop QA benchmarks, RetroRAG consistently outperforms existing RAG methods, achieving significant gains in answer accuracy—particularly under challenging conditions involving long reasoning chains and noisy evidence.
In multi-hop graph retrieval, the original query often fails to fully capture the complete information need distributed across multiple reasoning steps, resulting in insufficient retrieval signals. This work proposes an evidence-guided query reformulation mechanism that decouples query refinement from evidence aggregation on the graph: a residual query is generated from already retrieved passages to characterize unmet information needs, and the retrieval signals from the original and residual queries are separately normalized and then fused, propagating through shared entities across propositions. By moving beyond the conventional reliance solely on the initial query, the method achieves substantial gains, improving Recall@5 by up to 5.59 points and F1 by up to 4.50 points on 2WikiMultiHopQA, HotpotQA, and MuSiQue.
This work addresses the challenge in LongEval-RAG tasks where responses must be strictly grounded in a given set of candidate documents. To this end, the authors propose a candidate-constrained retrieval-augmented generation (RAG) system that integrates rule-based chunking, query expansion, pseudo-relevance feedback, reciprocal rank fusion, MiniLM sentence-level reranking, and citation-aware evidence aggregation, complemented by deterministic provenance tracing and a neural sentence selection mechanism. Experimental results demonstrate that the proposed rule-MiniLM variant significantly outperforms baselines across multiple metrics—including BERTScore, retrieval precision, information point coverage, and human evaluation—thereby validating the effectiveness of combining rule-based chunking with neural sentence selection. The study further underscores the critical role of multi-metric evaluation in diagnosing and advancing RAG system performance.