🤖 AI Summary
This study addresses the bottleneck of recursive self-improvement in non-verifiable tasks, where models lack external supervision. To overcome this limitation, we propose a multi-agent self-supervised framework that pioneers the use of structural constraints to guide evolutionary search. By alternating between workflow evolution and trajectory-supervised fine-tuning, our approach achieves joint optimization of orchestration logic and sub-agent execution, thereby transcending the homogeneity constraints inherent in single-model paradigms. The primary contribution lies in enabling agents to autonomously discover efficient collaboration topologies that enhance complex reasoning capabilities. Extensive evaluations demonstrate that our method yields 1.2× to 1.6× performance improvements across four open benchmarks, while exhibiting significantly superior training efficiency compared to conventional single-agent paradigms.
📝 Abstract
Recursive self-improvement (RSI) of a model on non-verifiable tasks, such as open-ended research, faces a supervision bottleneck when its outputs exceed what even human experts can reliably assess, leaving the model itself (optimizee) as the best available optimizer and evaluator. However, a single model instance struggles to critique and improve its own complex reasoning under this homogeneous loop. To address this, we propose Multi-Agent Self-Supervision (MASS), an RSI method that alternates between evolutionary workflow optimization and supervised fine-tuning on self-generated trajectories. Guided by early findings that multi-agent topologies excel at complex reasoning, MASS prompts a single base model to iteratively propose, execute, and self-evaluate multi-agent workflows. Through an evolutionary search constrained by structural guardrails, the model optimizes these computational-graph-like orchestrations, discovering the most effective distinct roles and information routing for a given task. Over two MASS cycles with Qwen3.6-27B, the model achieves 1.2-1.6x higher performance per output tokens on four open-ended public benchmarks. Because the improved model subsequently acts as a better optimizer and evaluator, this alternating framework enables a continuous, recursive bootstrapping of the model's capabilities. Moreover, multi-agent traces are also more training-efficient: a student trained on them outperforms a single-agent student trained on 1.4x more training tokens. These findings suggest that jointly learning orchestration and bounded subagent execution from multi-agent trajectories can provide an effective signal for RSI.