Evolutionary Safety of Recursive Self-Improving AI: Taxonomy, Risk Discovery, and Evaluation

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the safety challenges of recursively self-improving (RSI) AI during continuous evolution, including intent drift, error accumulation, and risk propagation. It pioneers an "evolutionary safety" theoretical framework that shifts the paradigm from static protection to dynamic governance. By constructing a multidimensional risk taxonomy encompassing agent states, model updates, and feedback mechanisms, combined with cross-generational trajectory analysis and automated governance derivation techniques, this work systematically reveals novel risks and their propagation pathways within RSI processes. Furthermore, it establishes a comprehensive risk assessment framework and proposes key governance principles—modification, selection, and provenance tracing—while open-sourcing relevant resources. Ultimately, this research provides both a theoretical foundation and practical toolset for ensuring the safe evolution of RSI systems.
📝 Abstract
Artificial intelligence is advancing rapidly, with increasingly capable systems taking larger roles in reasoning, decision-making, scientific discovery, and autonomous development. As AI begins to participate in its own improvement, from model training and experience accumulation to agent evolution and automated AI development, the prospect of recursive self-improvement (RSI) is becoming increasingly relevant. This transition raises a fundamental safety question: how can safety be maintained when the system, its accumulated experience, and even the process producing its successors continue to change? We introduce Evolutionary Safety as a perspective for studying safety under persistent and recursive self-improvement. It concerns not only whether an AI system is safe at a particular moment, but how safety properties change, persist, accumulate, and propagate throughout evolution. We characterize recurring manifestations, including intent drift, error accumulation, experience contamination, safety-property erosion, evaluator drift, and risk propagation. We then develop a taxonomy spanning persistent agent state, model state, evaluation and environmental feedback, computational substrate, and meta-level update mechanisms. Building on this taxonomy, we examine how evolutionary risks can be discovered and evaluated across states, updates, trajectories, and lineages, and derive governance principles for modification, selection, authorization, provenance, and recovery. Finally, we outline open problems toward maintaining safety guarantees as AI systems become increasingly persistent, adaptive, and recursively self-improving. Project resources and proposed evaluation systems are available at https://chaunceykung.github.io/evolutionary-safety-rsi.
Problem

Research questions and friction points this paper is trying to address.

Recursive Self-Improvement
Evolutionary Safety
AI Safety
Intent Drift
Risk Propagation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evolutionary Safety
Recursive Self-Improvement
Risk Taxonomy
Intent Drift
Safety Evaluation
Chang Gong
Chang Gong
AstraZeneca
Computational BiologyImmuno-oncology
J
Jingping Bi
Institute of Computing Technology, Chinese Academy of Sciences, China
Di Yao
Di Yao
Institute of Computing Technology, Chinese Academy of Sciences
Spatial-Temporal Data MiningTrajectory Data MiningGraph Neural NetworkTime-series Analysis
X
Xinjian Liang
Institute of Computing Technology, Chinese Academy of Sciences, China
Chao Xiang
Chao Xiang
University of Hong Kong
silicon photonicssemiconductor lasersphotonic integrated circuits
R
Ruijie Guo
Institute of Computing Technology, Chinese Academy of Sciences, China