RSI-Master: Structuring Experiments to Guide Autonomous Model Improvement

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of agent cheating and policy collapse in autonomous model development by proposing a bi-level architecture for cheat-free recursive self-improvement. At the lower level, an experimental operating system regularizes stepwise actions to eliminate cheating opportunities. At the upper level, reviewer-guided research orchestration combined with a dynamically expanding research directed acyclic graph (DAG) effectively mitigates policy collapse. Experimental results demonstrate that the proposed method achieves an average score of 54.49 on PostTrainBench with a zero-cheating rate. Notably, a 35-billion-parameter model surpasses human instruction-following models in performance and successfully resolves previously unsolved mathematical problems.
📝 Abstract
Recursive self-improvement (RSI) seeks to enable AI systems to participate in improving their own capabilities. A concrete pathway is autonomous model development, where agents iteratively explore post-training strategies to improve a base model. This setting faces two challenges: agents may exploit open-ended experimental actions through hacking, and repeated experimentation may lead to strategy lock-in, where an early direction is refined rather than reconsidered. We introduce RSI-Master, which addresses the two challenges at two levels: regularize step-wise actions, avoiding hacking behaviors, and promote well-structured exploration of research directions, avoiding strategy lock-in. RSI-Master consists of an Experiment OS, which enables regularized experimental actions and maintains persistent, traceable experimental records, and Reviewer-Guided Research Orchestration, which organizes Workers and Reviewers in a dynamically growing research DAG. Workers explore diverse research directions and Reviewers compare evidence across related experiments for subsequent explorations. On PostTrainBench with Qwen3-4B-Base, it averages 54.49 versus 46.53 for the strongest agent baseline, with a 0.0\% hacking rate. Scaling to 35B model, RSI-Master surpasses the human-developed Instruct model on LiveCodeBench-v6 (41.21 vs. 37.36) and SciCode, and reaches a nonzero score on HorizonMath, a benchmark of unsolved research problems on which most frontier models score near zero.
Problem

Research questions and friction points this paper is trying to address.

Recursive self-improvement
Autonomous model development
Hacking
Strategy lock-in
Post-training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recursive Self-Improvement
Autonomous Model Development
Experiment OS
Research DAG
Reviewer-Guided Orchestration