🤖 AI Summary
This study addresses the challenge that LLM agents struggle to continuously adapt to emerging safety risks from experience in real-world environments due to the absence of fixed task distributions. To overcome this limitation, we propose SafeCoEvo, a framework introducing a novel test-time Harness-Guard dual-timescale co-evolution mechanism. Specifically, S-Harness rapidly externalizes recent experiences to facilitate short-term adaptation, while GuardVPO internalizes risk judgments over the long term to consolidate safety capabilities, thereby transcending the constraints of traditional optimization methods reliant on fixed distributions. Experimental results demonstrate that, compared to the strongest baseline, our approach reduces the unsafe outcome rate by 10.05% and improves the task success rate by 12.15%, achieving simultaneous gains in both safety and utility.
📝 Abstract
LLM agents deployed in real-world environments continually encounter new tasks and safety risks, while execution feedback typically becomes available only after each task is completed. However, existing self-evolving approaches commonly rely on multiple rounds of optimization over fixed and repeatedly accessible task distributions, fundamentally differing from test-time adaptation in real-world deployment, where only experience accumulated from past tasks can be used to improve safety decisions on future unseen tasks. To address this limitation, we propose SafeCoEvo, a test-time Harness-Guard co-evolution framework for LLM agent safety that enables the external safety system to continually adapt from accumulated runtime experience. SafeCoEvo jointly improves two complementary safety capabilities at different timescales: S-Harness rapidly externalizes recent runtime experience into updatable explicit safety knowledge that can promptly influence subsequent tasks, while GuardVPO internalizes accumulated runtime safety experience over a longer timescale into parametric risk-judgment capabilities. By combining short-term rapid adaptation with long-term capability consolidation, SafeCoEvo continually improves the agent's safety capabilities, reducing the unsafe outcome rate by 10.05% while improving the task success rate by 12.15% over the strongest baseline, thereby achieving simultaneous gains in safety and task utility.