Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of sparse rewards and the absence of intermediate process credit assignment in reinforcement learning for deep research agents. We propose a rubric-based fine-grained credit assignment method that leverages task requirements as a shared reference, evaluating the incremental contribution of tool calls by contrasting historical evidence. This approach eliminates reliance on ground-truth answers, enabling process supervision for open-ended tasks. Furthermore, we construct a hybrid reinforcement learning framework that integrates the outcome advantages of GRPO with rubric-based process advantages to guide research decision-making. Experimental results demonstrate that our method comprehensively outperforms open-source baselines across multiple benchmarks. Notably, an 8B-parameter model achieves performance comparable to frontier proprietary models while significantly improving evidence acquisition efficiency.
📝 Abstract
Rubric-based tasks are increasingly addressed through reinforcement learning (RL), with rubric scores used as training rewards. However, these rewards typically supervise final answers without distinguishing the contributions of intermediate decisions. Many existing credit assignment methods rely on ground-truth answers to define process rewards, limiting their applicability to open-ended tasks without canonical solutions. To address this limitation, the proposed rubric-grounded credit uses task requirements as a shared reference for final answer evaluation and process supervision. The information returned by tools is assessed for the additional support it provides toward satisfying each rubric relative to that rubric's history of accepted support. By referencing these histories, credit distinguishes new support from evidence already present in the trajectory while recognizing partial support for each rubric. Dr.Credit uses rubric-grounded credit to supervise intermediate tool turns in an RL framework for deep research agents. The resulting process advantages are combined with GRPO outcome advantages to guide research decisions while retaining supervision of final-report quality. Evaluations on four in-domain and out-of-domain benchmarks show that Dr.Credit outperforms the evaluated open deep research baselines on every primary metric and submetric. Meanwhile, with an 8B-parameter backbone, the trained agent achieves average performance competitive with the evaluated frontier proprietary models. Further analyses suggest more efficient evidence acquisition and higher-quality reports under limited research-turn budgets, motivating the extension of rubric-grounded process supervision to a broader range of rubric-based tasks.
Problem

Research questions and friction points this paper is trying to address.

Credit Assignment
Process Supervision
Deep Research Agents
Reinforcement Learning
Rubric-based Tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Rubric-Grounded Credit Assignment
Process Supervision
Deep Research Agents
Reinforcement Learning
GRPO