QReason: Query-Focused Decoupled Chain-of-Thought for Efficient Passage Reranking

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high inference redundancy and latency in existing listwise LLM rerankers caused by sliding window strategies. To mitigate this, we propose a query-focused decoupled Chain-of-Thought (CoT) framework that extracts complex reasoning from individual windows for global reuse. Specifically, a dedicated rewriter generates a single reusable reasoning chain for non-reasoning rerankers. We employ a two-stage training paradigm: supervised fine-tuning on semantic evidence to produce deeply grounded CoTs, followed by reinforcement learning to align reasoning with reranking objectives. On the BRIGHT benchmark, our approach significantly reduces computational redundancy while achieving ranking performance comparable to or surpassing strong reasoning-based rerankers, outperforming existing query rewriting models.
📝 Abstract
Passage reranking plays a crucial role in information retrieval by refining the ordering of candidate passages to better reflect relevance. Existing listwise LLM rerankers with Chain-of-Thought (CoT) reasoning can handle complex queries effectively, but they suffer from substantial redundancy and high latency due to sliding-window strategies, which repeatedly generate highly similar CoTs. To address this, we propose QReason, a decoupled framework that separates query-focused reasoning from window-specific passage relevance assessment. Specifically, QReason introduces a dedicated rewriter that generates a ranking-oriented reasoning query once, capturing the query's core intent while avoiding redundant reasoning, and then reuses it across all windows with a non-reasoning reranker. The rewriter is trained via a two-stage process that first uses supervised fine-tuning with relevant-passage guidance through semantic evidence to produce deeply grounded, query-focused CoTs. It then applies reinforcement learning to align CoT generation with both the inference-time setting and the reranking objective, optimizing listwise metrics and passage-level discrimination to produce reusable reasoning chains for reranking. Experiments on the BRIGHT benchmark demonstrate that QReason significantly reduces redundant reasoning, achieves ranking performance comparable to or better than strong reasoning-based rerankers, and outperforms existing query rewriting models.
Problem

Research questions and friction points this paper is trying to address.

Passage Reranking
Chain-of-Thought
Redundancy
Latency
Listwise LLM Rerankers
Innovation

Methods, ideas, or system contributions that make the work stand out.

Query-Focused Reasoning
Decoupled Chain-of-Thought
Passage Reranking
Two-Stage Training
Reinforcement Learning