HASTE: Evolving Agent Harnesses Against Emerging Attacks Using Sparse Evidence

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that existing agent defense mechanisms struggle to counter rapidly emerging attacks under conditions of sparse threat evidence. To this end, we propose an adversarial evolution paradigm grounded in sparse evidence. The method constructs a multi-agent adversarial framework that integrates adversarial generation techniques with a closed-loop feedback evaluation mechanism, enabling defense strategies to dynamically balance safety compliance and attack exploration for continuous evolution beyond initial observational data. Experimental results demonstrate that the proposed framework significantly reduces attack success rates across diverse models and attack scenarios while effectively preserving utility on benign tasks.
📝 Abstract
Agent harnesses play a critical role in defenses by enforcing safety constraints to prevent unsafe actions. However, rapidly emerging attacks outpace manual harness adaptation, motivating automated harness evolution. Yet the signals available for harness evolution are often sparse, such as brief descriptions or a few attack examples in threat reports and preprints. To address this limitation, we introduce HASTE, a multi-agent framework that evolves agent harnesses from sparse threat evidence through an adversarial interplay between safety-specification generation and attack-case generation. Safety specifications guide harness updates toward addressing identified safety vulnerabilities, while attack cases probe for remaining safety vulnerabilities after each update. By feeding evaluation outcomes back into both processes, HASTE enables harness evolution against emerging attacks beyond the initially observed evidence. Experimental results across multiple backbone models, attack types, and evidence forms show that HASTE consistently reduces attack success rates while preserving benign-task utility. The code is available at https://github.com/xxiqiao/HASTE.
Problem

Research questions and friction points this paper is trying to address.

Agent Harnesses
Emerging Attacks
Sparse Evidence
Automated Evolution
Safety Constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-agent framework
adversarial interplay
sparse evidence
agent harness evolution
safety specification
🔎 Similar Papers
No similar papers found.