R$^2$ Flow: Recursive Self-Improvement via Recursive Skill Evolution

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses policy collapse, misleading non-negative flow rewards, and the reliance on single-reward signals for skill editing in LLM agent self-improvement. We propose a recursive self-improvement framework that orchestrates shared-state graphs by merging equivalent histories to alternate between policy learning and independent verification. The method introduces invariant flow-share readouts and signed utility ranking to optimize training, while leveraging verification evidence and residual variance plateaus to precisely govern versioned skill updates. Experiments demonstrate that this framework significantly improves accuracy and editing precision across question answering, mathematical reasoning, interactive decision-making, and code generation tasks. It outperforms multiple baselines and exhibits strong transferability across different executors.
📝 Abstract
LLM-based agents can improve themselves across tasks by reusing and revising the skills they orchestrate into executable procedures. Flow-based training fits this loop: it samples procedures in proportion to reward, and the flow through each skill credits it for the next library revision. Three obstacles stand in the way of making this self-improvement reliable: flow training suffers strategy collapse over tree-structured histories; nonnegative flow-based credit rewards frequent use as if it were benefit; and library edits rest on the task reward the policy optimizes. We introduce R$^2$ Flow, a recursive self-improvement framework that alternates policy learning, independent verification, and versioned skill-library updates on a shared-state orchestration graph. The graph merges histories that differ only in the order of independent steps, allowing flow training to pool evidence across equivalent executions. A flow-share readout of the trained flow, invariant to the backward policy, and a separate signed utility rank which skills to change, verifier evidence decides whether an edit is warranted, and a residual-variance plateau sets when to update. Committed edits reshape the graph the next policy learns on, realizing recursive skill evolution. Across question answering, mathematical reasoning, interactive decision making, and code generation, R$^2$ Flow improves task accuracy and library-edit precision over heuristic orchestration, reinforcement learning, and skill-evolution baselines, and transfers across executors. Code is available at https://github.com/beita6969/r2flow.
Problem

Research questions and friction points this paper is trying to address.

self-improvement
skill evolution
strategy collapse
credit assignment
LLM-based agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recursive Self-Improvement
Shared-State Orchestration Graph
Flow Training
Signed Utility
Skill Library Evolution