UniOPSD: Unifying Outcome and Hindsight Feedback for Agentic Reinforcement Learning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the credit assignment challenge in agent reinforcement learning arising from sparse rewards and locally inconsistent feedback. To this end, we propose UniOPSD, a framework that adaptively arbitrates between environmental returns and hindsight supervision based on historical consistency and real-time signal precision. The method constructs comparable credit estimates through online self-distillation and shared interaction anchors, combined with bounded token modulation to enable fine-grained policy optimization. Experimental results demonstrate that UniOPSD significantly improves success rates on benchmarks such as ALFWorld and WebShop. Notably, a 3B-parameter model achieves a 7-percentage-point improvement over SDAR on WebShop, highlighting the effectiveness of the proposed approach in complex interactive environments.
📝 Abstract
Reinforcement learning has become an effective approach to training language model agents, but sparse and delayed outcome rewards provide limited guidance for credit assignment across long interaction sequences. Recent work on on-policy self-distillation (OPSD) offers complementary supervision by evaluating a policy's sampled responses under privileged training-time context. However, our diagnostics show that positive average agreement between outcome and hindsight feedback coexists with substantial local disagreement, raising the question of how to allocate influence between them at each decision. We introduce UniOPSD (Unified On-Policy Self-Distillation), which unifies these feedback sources through adaptive local credit arbitration. UniOPSD constructs comparable credit estimates from environmental returns and successful-peer hindsight at shared interaction anchors. Historical agreement determines the global mixing level, while current signal availability and relative precision adjust each source's influence at individual decisions. The episode-level outcome contribution is retained, and bounded token modulation refines the fused step credit for policy optimization. With Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct, UniOPSD achieves ALFWorld success rates of $82.8\%$ and $83.6\%$, WebShop success rates of $75.0\%$ and $82.0\%$, and Search-QA aggregate accuracies of $45.3\%$ and $49.8\%$, respectively. On 3B WebShop, UniOPSD improves over SDAR by $7.0$ percentage points. Our code is available at https://github.com/Zenghuang-Fu/Uniopsd
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Credit Assignment
Language Model Agents
Outcome Feedback
Hindsight Feedback
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Reinforcement Learning
On-Policy Self-Distillation
Adaptive Credit Arbitration
Hindsight Feedback
Credit Assignment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.