🤖 AI Summary
This study addresses the vulnerability of large language model (LLM) agents to injection attacks and the lack of real-time mitigation capabilities in existing defenses. To this end, we propose a risk-aware framework coupled with a two-stage optimization method. The framework integrates triggering, monitoring, and feedback modules, leveraging LLM-based dynamic monitoring and modular execution control to enable precise security interventions. Furthermore, by generating guidelines through isolated probing and refining them iteratively, our approach effectively balances safety and task utility. Experimental results demonstrate that the proposed method significantly enhances runtime safety across multiple benchmarks while preserving high task utility, consistently outperforming manually designed baselines.
📝 Abstract
Large language model (LLM) agents are vulnerable to safety risks such as injected malicious instructions or misleading information, motivating runtime defenses that prevent unsafe action in execution across diverse risks while preserving benign-task utility. Existing system-level defenses either focus on risk detection rather than timely prevention or rely on predefined rules with limited flexibility across diverse risks. We propose a risk-aware harness that integrates LLM-based monitoring for flexible risk detection and structures monitor-guided execution around three core modules: trigger, monitor, and feedback, enabling targeted safety interventions while limiting disruption to benign task execution. To adapt the harness to different risks and deployment settings, we introduce HARDE, a two-stage harness optimization framework that first performs isolated probing of each module to derive an optimization guide, then uses this guide to iteratively optimize the harness based on safety and utility feedback. Experiments across three attack benchmarks show that HARDE improves runtime safety while preserving utility, outperforming manually designed harnesses and naive optimization baselines. Our analysis shows that effective runtime defense benefits from complementary safety mechanisms, attack-aware harness optimization, and harness designs matched to monitor capabilities. Our code is available at https://github.com/Liuz233/HARDE.