Training Documents Reranker with Search Rubrics for Deep Research Agent

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current retrieval systems struggle to simultaneously satisfy the demands of deep research agents for document collections that are diverse, concise, and authoritative. To address this challenge, this work proposes Search Rubrics—a structured scoring framework that explicitly defines multidimensional criteria for high-quality document sets—and introduces RubricRanker, a two-stage trained document reranker. RubricRanker first aligns with the rubrics through supervised fine-tuning and then optimizes the overall quality of document subsets via rubric-guided reinforcement learning. This approach is the first to incorporate hierarchical scoring rubrics into document reranking, achieving a 2.6-point improvement over the strongest baseline across four deep research benchmarks and demonstrating strong generalization performance on five RAG benchmarks.
📝 Abstract
Retrieval systems help deep research agents generate high-quality answers by providing relevant documents. However, existing retrievers typically select documents through relevance matching, while individually well-matched top-$k$ documents may not form a \textit{set} that satisfies the complex information needs of an agent query (\eg, diverse, concise and authoritative documents). In this paper, we propose search-oriented rubrics that \textit{explicitly} define the requirements that high-quality document sets should satisfy for each agent query. Our search rubrics are organized into a hierarchical structure and synthesized using a powerful LLM. Based on these search rubrics, we further train a document reranker \textbf{RubricRanker} to select a high-quality subset from retrieved documents. We design a two-stage training framework that consists of rubrics-guided supervised fine-tuning and rubric-based reinforcement learning. Extensive experiments demonstrate that RubricRanker outperforms the strongest baseline by 2.6 points on four deep research benchmarks and generalizes well to five RAG benchmarks.
Problem

Research questions and friction points this paper is trying to address.

document reranking
deep research agent
retrieval systems
information needs
search rubrics
Innovation

Methods, ideas, or system contributions that make the work stand out.

search rubrics
document reranking
deep research agent
reinforcement learning
retrieval-augmented generation
Wenhan Liu
Wenhan Liu
Gaoling School of Artificial Intelligence, Renmin University of China
Information RetrievalLarge Language Models
Y
Yu Lu
Tencent, Beijing, China
Q
Qiaolin Xia
Tencent, Beijing, China
H
Hui Xu
Tencent, Beijing, China
T
Tong Zhao
Gaoling School of Artificial Intelligence, Renmin University of China
J
Jian Xi
Tencent, Beijing, China
Y
Yutao Zhu
Gaoling School of Artificial Intelligence, Renmin University of China
H
Haijin Liang
Tencent, Beijing, China
H
Haibo Shi
Tencent, Beijing, China
H
Hao Wang
Tencent, Beijing, China
Zhicheng Dou
Zhicheng Dou
Renmin University of China
Information RetrievalRetrieval Augmented GenerationLarge Language ModelsGenerative IR