Environmental Feedback Modeling Matters: Rethinking Feedback Treatment in Agentic Hindsight Self-Distillation

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of applying reinforcement learning in scenarios lacking explicit rewards and the underutilization of environmental feedback. To this end, we propose SELF, a framework built upon Qwen3-8B that jointly optimizes environment feedback modeling and hindsight self-distillation. By revealing and leveraging the mutually reinforcing mechanism between these two components, SELF reconstructs the agent's feedback processing paradigm, enabling policy enhancement through the prediction of environmental responses. Experimental results demonstrate that by integrating supervised fine-tuning with reinforcement learning techniques, our method outperforms baselines by 4.1 to 10.7 percentage points on the τ-bench and AppWorld benchmarks, respectively. These findings indicate that SELF significantly improves agent decision-making performance in reward-free settings.
📝 Abstract
Reinforcement learning is commonly used to train language agents in interactive environments, but cannot be directly applied when rewards are unavailable. Recent methods use environmental feedback as privileged context for hindsight self-distillation, but our analysis suggests that simply conditioning the teacher on feedback is insufficient, motivating us to rethink how environmental feedback is used in agentic self-distillation. Given that environmental feedback contains rich supervision for modeling how the environment responds to agent actions, we introduce \textit{agentic SElf-distilLation with environmental Feedback modeling} (SELF), a framework that jointly optimizes environmental feedback modeling and hindsight self-distillation. SELF learns to predict environmental responses while distilling guidance from a feedback-conditioned self-teacher into the policy. Our analysis reveals a mutually reinforcing mechanism: environmental feedback modeling strengthens hindsight supervision and policy learning, while self-distillation enhances the model's ability to model environmental feedback. With Qwen3-8B, SELF outperforms SDPO and GRPO by 6.4 and 4.1 percentage points in $τ$-bench success rate, and by 10.71 and 3.57 percentage points in AppWorld task goal completion, respectively. These results show that SELF uses environmental feedback more effectively within agentic self-distillation, improving agent capabilities.
Problem

Research questions and friction points this paper is trying to address.

environmental feedback modeling
agentic self-distillation
language agents
reinforcement learning
hindsight supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Environmental Feedback Modeling
Hindsight Self-Distillation
Language Agents
Joint Optimization
Reinforcement Learning
🔎 Similar Papers
No similar papers found.