Compositional Safety Failures in Harness Evolution: Identification and Runtime Monitoring

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of efficiently verifying compositional security failures arising from component interactions during the evolution of agent frameworks. To this end, it proposes a runtime monitoring mechanism based on typed hypergraphs. By constructing a local-update hypergraph model and an incremental neighborhood update algorithm, this work overcomes the combinatorial complexity bottleneck inherent in traditional global verification, enabling incremental interaction checking across component updates. Experimental results demonstrate that the proposed approach effectively mitigates security risks while preserving task utility and significantly reducing interaction-checking overhead. Furthermore, this research elucidates the intrinsic trade-off among security, utility, and cost, offering both theoretical insights and practical guidance for securing evolving multi-component agent systems.
📝 Abstract
Self-evolving agent harnesses continually update persistent components such as memory, prompts, skills, and tools. We call this process harness evolution. However, such evolution could introduce unexpected safety risks. Existing work studies harness misevolution and validates candidate harnesses or attributed individual component updates, leaving safety analysis of cross-component update interactions largely unexamined. To address this gap, we study compositional safety failures in harness evolution, where interactions among individually safe and utility-preserving component updates can produce undesirable or unsafe agent behavior, revealing a safety risk intrinsic to harness evolution. Across three safety-related benchmarks, we identify 43 pairwise and 18 irreducible 3-way compositional safety failures. Conventional solution incurs combinatorial complexity in validating cross-component interactions, leaving the safety checking impractical as the harness evolves. To solve this, we introduced a typed hypergraph that represents component states as nodes and safety-relevant higher-order interactions as hyperedges. When the harness changes, the hypergraph updates only the interaction neighborhood of the changed states rather than reconstructing the global composition space. Building on that, we develop a hypergraph-guided runtime monitoring mechanism. Experiments show that our method effectively mitigates compositional safety risks while preserving task utility and reducing interaction-checking costs, and further reveal an empirical safety-utility-cost trade-off across different safety mechanisms.
Problem

Research questions and friction points this paper is trying to address.

compositional safety failures
harness evolution
self-evolving agents
cross-component interactions
runtime monitoring
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compositional Safety Failures
Harness Evolution
Typed Hypergraph
Runtime Monitoring
Self-evolving Agents
🔎 Similar Papers
No similar papers found.
Z
Zhixiang Zhang
The Hong Kong University of Science and Technology
Zesen Liu
Zesen Liu
Ph.D. Student, HKUST
Security
W
Wai Ip Lai
The Hong Kong University of Science and Technology
H
Hongxu Chen
The Hong Kong University of Science and Technology
Dongdong She
Dongdong She
Hong Kong University of Science and Technology
SecurityMachine LearningProgram AnalysisFuzzing