🤖 AI Summary
This study addresses the challenge of precisely localizing errors within trajectory-level rewards in reinforcement learning for agents. To this end, we propose a segment-level post-hoc advantage reweighting method that introduces a novel segment-level credit assignment mechanism. Specifically, it computes bounded multipliers based on the probability divergence between teacher and student policies to fine-tune GRPO advantages, while integrating online self-distillation with large language models to optimize policy gradients. Experimental results demonstrate that our approach significantly outperforms existing baselines on the ALFWorld and WebShop benchmarks, effectively enhancing both credit assignment precision and policy optimization efficiency in complex interactive tasks.
📝 Abstract
Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitation, we introduce Segment-level Hindsight Advantage Reweighting for Policy Optimization (SHARPO), a credit-assignment mechanism that refines Group Relative Policy Optimization (GRPO) at the level of environment-facing segments. Inspired by the existing on-policy self-distillation (OPSD) method, SHARPO computes teacher-student log-probability gaps within each segment and uses the resulting signal to compute a bounded multiplier on the GRPO advantage. This multiplier is shared by all tokens within the segment, allowing credit to vary across different segments. With Qwen2.5-7B-Instruct, SHARPO outperforms existing baselines on the ALFWorld and WebShop benchmarks, including GRPO, SDAR, RLSD, and StepOPSD.