RLTL;DR: Self-improvement by Internalizing Self-generated Feedback

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the self-improvement bottleneck in reinforcement learning for extremely difficult tasks, where vanishingly low success rates hinder agent progress. To overcome this limitation, we propose a novel paradigm that compresses failure experiences into single-sentence insights for internalization. Building upon the GRPO framework, our approach integrates in-context learning, backpropagation, and supervised fine-tuning, enabling agents to extract critical insights from failures to guide subsequent attempts without relying on complete trajectories. Empirically, on extremely challenging problems where the Pass@128 metric is zero, our method increases the pass rate from 1% to 31%. Notably, even under zero-prompt conditions, it achieves 12–13%, significantly outperforming existing baselines and successfully breaking through the zero-shot learning barrier.
📝 Abstract
The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own feedback, in the form of a single TL;DR insight. The next rollout is conditioned on all previous insights, and we sequentially sample rollouts until a solution is found. Moreover, we enable backpropagation on the in-context insights to internalize a direct task to insight mapping. On challenging tool-calling and coding datasets (filtered to Pass@128=0), standard GRPO training of a Qwen 3.5 9B Thinking policy stays flat at a Pass@1 of 0% to 1%. RLTL;DR breaks through this learning barrier, achieving a Pass@1 of 14-31% with insights in context during training and, crucially, 12-13% when no insight is in context at eval time. We identify that the key is the task to insight internalization. To study this further, we reduce our approach to SFTL;DR, training only on (task, insight) tuples, without showing or backpropagating on any rollouts. Training on only 4k of these tuples recovers almost the full performance of RLTL;DR and classical SFT on full rollouts. This demonstrates a promising compacted training paradigm of the form "on this sort of task, keep this sort of thing in mind", which we hope to inspire future research on.
Problem

Research questions and friction points this paper is trying to address.

self-improvement
reinforcement learning
verifiable rewards
hard tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-improvement
Reinforcement Learning
Feedback Internalization
Insight Generation
Compacted Training
🔎 Similar Papers
No similar papers found.