🤖 AI Summary
This study addresses the misalignment between static training data and evolving capabilities in self-evolving large language models, where tasks become either too easy or too difficult, thereby hindering continuous improvement. To overcome this, we propose a multi-agent collaborative reinforcement learning framework that jointly optimizes a synthesizer and a reasoner. Through a complementary reward mechanism, the framework guides both agents to co-evolve synchronously, enabling dynamic adaptation between task generation and solving strategies so that the task distribution autonomously evolves alongside model capabilities. Experiments across eight mathematical benchmarks demonstrate that our approach significantly outperforms existing baselines. These performance gains primarily stem from effectively solving previously unresolved problems, validating the superiority of synergistically training online feedback with data synthesis.
📝 Abstract
Self-evolving LLM agents promise to improve autonomously through continual interaction and learning, reducing their dependence on manually curated supervision. Realizing this promise requires not only updating the agent, but also evolving its training experience as its capabilities change. However, most existing pipelines rely on static datasets or separately updated synthesis models, causing previously useful tasks to become trivial while overly difficult tasks remain uninformative. This growing mismatch between agent capability and training experience limits sustained self-improvement. To address this problem, we propose SynCo, an agentic data synthesis co-training framework for self-evolving LLMs based on multi-agent reinforcement learning. SynCo jointly optimizes two independently parameterized agents: a Synthesizer that constructs training tasks from the Reasoner's evolving capability state, and a Reasoner that learns from the resulting experience. Each synthesized task induces multiple Reasoner rollouts whose outcomes provide complementary rewards to both agents. Correctness feedback improves the Reasoner, while task quality, answer reliability, and outcome-grounded teachability guide the Synthesizer. Their updates are fed back into subsequent synthesis rounds, allowing the task-solving policy and its training distribution to evolve together. Extensive experiments across eight mathematical reasoning benchmarks demonstrate that SynCo substantially outperforms a broad range of existing synthetic-data methods and controlled baselines, achieving the strongest overall performance while deriving most of its gains from previously unsolved problems.