🤖 AI Summary
This study addresses the challenges of redundant infrastructure development and constrained exploration in recursive self-improvement (RSI) research by proposing an agent-native research environment grounded in an "Everything-as-a-Service" paradigm. Methodologically, it achieves joint optimization of data, training, and execution frameworks through reusable services, incorporating shared budget control, sandboxed execution, and automated evaluation pipelines. Furthermore, the RSI-Index metric is introduced to facilitate both individual interventions and multi-component iterative evaluations within a unified environment. Experimental results demonstrate that the Opus 5 model attains the highest RSI index, yielding significant performance improvements on the SWE-bench and AIME benchmarks. All code and experimental results have been released as open source.
📝 Abstract
Recursive self-improvement requires carrying accepted changes into later improvement cycles, while studying agent-proposed changes also requires substantial research infrastructure. Existing settings often leave agents to rebuild routine infrastructure or restrict exploration to individual components. We introduce RSIGym, an agent-native research environment based on Everything as a Service (EaaS). RSIGym exposes training, inference, rollout, evaluation, and sandbox execution through reusable services, with shared budget and permission controls supporting Data, Harness, and Joint improvement tracks. This design enables agents to investigate individual interventions and jointly optimize data, training settings, and execution harnesses within the same environment. We define RSI-Index as the mean fraction of the remaining performance gap closed across five benchmarks covering software engineering, terminal interaction, mathematics, scientific reasoning, and skill-based tasks. Comparing six frontier research models in independent Joint runs, Opus 5 achieves the highest RSI-Index of 0.4809 under a $500 platform-service budget per benchmark run. Its selected systems improve all five benchmarks, raising SWE-bench Verified from 17.67% to 50.33% and AIME from 31.67% to 97.78%. Additional experiments examine DSH-harness refinement, budget variation, and restricted network access, while recorded trajectories reveal how agents diagnose failures and select candidates. We open-source the full RSIGym codebase and results to support reproducibility and further research.