🤖 AI Summary
This study addresses the challenges of unsafe behavior persistence and failure recovery in recursive self-improvement by proposing a stateful authorized task testbed that decouples revision verification, program execution, and edit selection. Methodologically, we design a paired intervention framework to reveal how historical scoring leads to the retention of unsafe programs. We employ a fixed LLM editor to optimize executable components while independently tracking its effects, implementing comprehensive verification and rollback strategies. Results demonstrate that core trajectories successfully generate fully correct programs while preserving over 43% in deployment overhead savings. Furthermore, this work establishes principled guidelines for checking and edit selection under current conditions, offering a novel paradigm for the safe self-improvement of AI systems.
📝 Abstract
Recursive self-improvement (RSI) allows agents to carry useful changes across generations. Maintaining safety across these generations involves both preventing unsafe behavior from persisting and enabling recovery when failures occur. We study these challenges through a controlled testbed of stateful authorization tasks, where fixed LLM editors optimize executable agent components and independent traces record their effects. Paired interventions separate which revisions pass validation, which program continues running, and which program the editor revises next. After a new authorization dependency invalidates previously tested optimizations, historical scores preserve the same unsafe programs in 22 of 48 framework histories despite a correct alternative in every affected archive. Refreshing scores restores correctness on the original suite, with residual failures on independently composed tests. Failures also persist under an unchanged contract when all proposals are rejected and the failed incumbent remains active. Starting from shared failures, editing the initial implementation instead of the failed one improves recovery, although the advantage varies across editors. Full validation and validated rollback end with fully correct programs in the core trajectory study while retaining over 43% deployment savings. Preserving agent safety requires checking what will run under current conditions and choosing which implementation to edit next.