AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of ambiguous credit assignment and evidence redundancy caused by sparse supervision in RAG reranking. We propose a set-level reranker based on adaptive tutoring optimization. Methodologically, we construct a three-tier scoring system to provide dense labels and rewards, and introduce a novel adaptive tutoring mechanism that dynamically matches multi-granularity prompts according to policy quality to distinguish contributors from free-riders. For training, we integrate reinforcement learning with online policy distillation, leveraging a frozen teacher model to compute token-level advantages for efficient learning. The proposed approach achieves state-of-the-art performance across ten benchmarks while significantly reducing retrieval invocations, thereby effectively enhancing the efficiency of deep research pipelines.
📝 Abstract
Document rerankers determine what evidence reaches the downstream model in RAG and deep research, yet mainstream rerankers select by relevance matching, and individually relevant documents rarely constitute the complete, complementary, non-redundant set a complex information need demands. Prior work rewards a set by its aggregate rubric score, shifting the objective from ranking documents to composing sets. Yet that score is one scalar shared by every document in the set, so the supervision is sparse: a redundant document is rewarded with the rest whenever the set scores well, and a decisive one penalized with the rest whenever it does not; credit assignment leaves contributors indistinguishable from free riders. On-policy distillation could densify this supervision, but existing methods give every rollout the same fixed guidance, too prescriptive for strong rollouts and too abstract for weak ones. We therefore propose AdaTutoRank, a setwise reranker trained with Adaptive Tutoring Optimization (ATO) under a three-level hierarchy of nine rubric dimensions, which supplies silver labels for the cold start, rewards for reinforcement learning, and hints for distillation. ATO draws three hint forms of increasing specificity from the policy's own frozen snapshot: the rubrics alone, a self-selector's sibling-set chosen under rubrics, and a self-reflector's reflection contrasting the rollout with that sibling-set; each rollout receives the form matched to its quality. Re-scoring that rollout under the hint-conditioned frozen teacher and the hint-free snapshot distills the hint's effect into a token-level advantage that complements the group-relative outcome advantage. Across ten benchmarks spanning RAG, deep research, and setwise evaluation, AdaTutoRank attains the best overall performance while issuing fewer retrieval calls.
Problem

Research questions and friction points this paper is trying to address.

Document Reranking
Retrieval-Augmented Generation
Setwise Evaluation
Credit Assignment
Sparse Supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Setwise Reranking
Adaptive Tutoring Optimization
Retrieval-Augmented Generation
Credit Assignment
On-policy Distillation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kailin Jiang
University of Science and Technology of China, Yuanbao Team, Tencent
Lei Liu
Lei Liu
Anhui University of Science & Technology
CV
J
Jian Xi
Yuanbao Team, Tencent
Y
Yangqi Chen
Yuanbao Team, Tencent
H
Hui Xu
Yuanbao Team, Tencent
H
Hongwei Zhao
University of Science and Technology of China
B
Bin Li
University of Science and Technology of China
Y
Yu Lu
Yuanbao Team, Tencent
H
Haibo Shi
Yuanbao Team, Tencent