🤖 AI Summary
This study addresses the challenges of applying reinforcement learning in scenarios lacking explicit rewards and the underutilization of environmental feedback. To this end, we propose SELF, a framework built upon Qwen3-8B that jointly optimizes environment feedback modeling and hindsight self-distillation. By revealing and leveraging the mutually reinforcing mechanism between these two components, SELF reconstructs the agent's feedback processing paradigm, enabling policy enhancement through the prediction of environmental responses. Experimental results demonstrate that by integrating supervised fine-tuning with reinforcement learning techniques, our method outperforms baselines by 4.1 to 10.7 percentage points on the τ-bench and AppWorld benchmarks, respectively. These findings indicate that SELF significantly improves agent decision-making performance in reward-free settings.
📝 Abstract
Reinforcement learning is commonly used to train language agents in interactive environments, but cannot be directly applied when rewards are unavailable. Recent methods use environmental feedback as privileged context for hindsight self-distillation, but our analysis suggests that simply conditioning the teacher on feedback is insufficient, motivating us to rethink how environmental feedback is used in agentic self-distillation. Given that environmental feedback contains rich supervision for modeling how the environment responds to agent actions, we introduce \textit{agentic SElf-distilLation with environmental Feedback modeling} (SELF), a framework that jointly optimizes environmental feedback modeling and hindsight self-distillation. SELF learns to predict environmental responses while distilling guidance from a feedback-conditioned self-teacher into the policy. Our analysis reveals a mutually reinforcing mechanism: environmental feedback modeling strengthens hindsight supervision and policy learning, while self-distillation enhances the model's ability to model environmental feedback. With Qwen3-8B, SELF outperforms SDPO and GRPO by 6.4 and 4.1 percentage points in $τ$-bench success rate, and by 10.71 and 3.57 percentage points in AppWorld task goal completion, respectively. These results show that SELF uses environmental feedback more effectively within agentic self-distillation, improving agent capabilities.