Shockingly Simple Self-retrospection Improves Agentic Models Without RL

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of efficient mechanisms for agents to self-improve from their own experiences, as traditional reinforcement learning relies on costly reward signals. We propose Retrospective Online Fine-Tuning (ROFT), which fine-tunes a model’s retrospective interpretations of its task trajectories using only next-token prediction loss, without external teachers or rewards. Based on Qwen3.5-4B, this work provides the first demonstration that “learning to explain” directly facilitates “learning to act,” establishing self-generated retrospection as an effective training objective for reinforcement-learning-free self-evolution. On SWE-bench, ROFT surpasses GRPO baselines with fewer update steps, significantly accelerates early convergence, and successfully resolves tasks that initially failed entirely.
📝 Abstract
People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO's 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer.
Problem

Research questions and friction points this paper is trying to address.

language-model agent
self-retrospection
explanation-to-action transfer
fine-tuning
software engineering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-retrospection
Retrospection-Only Fine-Tuning
Explanation-to-action transfer
Agentic models
Credit assignment
🔎 Similar Papers
No similar papers found.