🤖 AI Summary
This study addresses the limitation of traditional conversational query rewriting (CQR), which lacks retrieval feedback and struggles to correct coreference and lexical errors. We reformulate CQR as a sequential retrieval problem and propose an agent-based iterative optimization framework that dynamically refines queries through a multi-step rewrite-retrieve-feedback interaction mechanism. The method introduces a typed rewriting space and a pseudo-document synthesis strategy, enabling retrieval-reward-driven learning without manual annotation. Training is conducted via supervised fine-tuning combined with reinforcement learning. Experiments demonstrate that our approach significantly outperforms existing baselines on benchmarks such as TopiOCQA, while exhibiting strong generalization across different retrieval backends and the CAsT task.
📝 Abstract
Conversational Query Rewriting (CQR) turns a context dependent user turn into a standalone query for a retriever, and most methods do this in a single step from the dialogue history before retrieving once. The rewrite is therefore fixed before any corpus evidence is available to correct its reference resolution or its vocabulary. We recast CQR as a sequential retrieval problem: an agent rewrites the current turn, retrieves, and conditions its next rewrite on the returned passages. The agent acts in a typed space of three rewriting operations, resolving conversational intent into a standalone query, generating lexical reformulations, or synthesizing pseudo-documents for document-to-document matching, together with a stop action that ends the episode. We train the policy with supervised fine-tuning followed by reinforcement learning against a single retrieval-quality reward, using no human rewrite annotations. Across TopiOCQA and QReCC, the agent outperforms several retrieval-aligned baselines, while remaining effective across retrieval backends and generalizing to the CAsT benchmarks without additional training. Further analysis shows that, through retrieval-reward optimization alone, the learned policy develops a behavior of grounding pseudo-documents in passages retrieved by earlier steps, substantially improving retrieval.