ReSI: Recursive Safety Improvement toward Resistant and Resilient AI

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that safety alignment struggles to adapt to frequent iterations and unforeseen risks during AI self-improvement. To this end, we propose a Recursive Safety Improvement framework that introduces the recursive self-improvement paradigm into the safety domain for the first time. By integrating automated red teaming, training recipe generation, and a Mixture-of-Experts-based Pareto gating mechanism, the framework dynamically balances defenses against known threats with resilience to unknown risks, thereby iteratively enhancing model safety. Experimental results demonstrate that our approach significantly reduces the attack success rate to 31.45%, surpassing existing frontier models in safety while effectively preserving general-purpose capabilities.
📝 Abstract
Recursive self-improvement, the participation of AI systems in improving their own capabilities, is beginning to move from theoretical prospect to practice, posing both challenges and opportunities for safety alignment. Models evolve through frequent updates, and their safety alignment requires continual adaptation to each new checkpoint. Meanwhile, with evolving red-teaming methods exposing new vulnerabilities, safety improvement for each checkpoint needs to mitigate exposed vulnerabilities and generalize to risks not yet revealed. Following R$^2$AI, we term these goals resistance to known threats and resilience to unforeseen risks. Recursive self-improvement, in turn, inspires an approach to both goals: safety alignment could likewise advance through successive rounds of evaluation and update. We therefore introduce ReSI, a recursive safety improvement framework that implements this approach through automated research. In each round, ReSI applies diverse red-teaming methods to identify vulnerabilities in the current target model, develops training recipes, and promotes the update with the largest safety gain among those passing a Pareto gate on capability retention as the next target model. Across four dense and mixture-of-experts models, ReSI matches or exceeds evaluated frontier models on in-distribution and out-of-distribution safety benchmarks, and outperforms alignment baselines on nearly all safety evaluations while largely preserving general capabilities. In particular, ReSI reduces the mean X-Teaming attack success rate across the four models from 86.01% to 31.45%, well below GPT-5.6-Luna's leading frontier result of 56.69%, indicating stronger resilience to attacks unseen during training. These findings support recursive safety improvement as a practical path toward resistant and resilient AI.
Problem

Research questions and friction points this paper is trying to address.

recursive self-improvement
safety alignment
red-teaming
resistance
resilience
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recursive Safety Improvement
Automated Red-teaming
Pareto Gate
Safety Alignment
Resilience