Towards a Unified Misuse Monitoring Benchmark

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing LLM agent safety evaluations that treat decomposition attacks and prompt injection in isolation, thereby failing to precisely localize the moment harm occurs. To bridge this gap, this work proposes a unified abuse monitoring benchmark and an action-based monitoring paradigm that formalizes the harm window by tracing externalized actions, alongside interval-based metrics for fine-grained temporal localization assessment. Validated on approximately 6,200 synthetic dialogues, the action-based monitor achieves AUCs of 0.95 and 0.99 against both attack types, whereas content-based monitors fail under injection scenarios. This research is the first to incorporate multi-source attacks into a unified modeling framework, revealing that conventional metrics significantly overestimate harm localization capabilities.
📝 Abstract
LLM agents increasingly act in multi-actor environments, exposing them to misuse from multiple sources: decomposition attacks, where a harmful request is split into innocuous sub-requests, and prompt injection attacks, where a compromised tool delivers a malicious instruction. Existing evaluations treat these threats separately and ask whether a trajectory is harmful, rather than when it becomes harmful. We propose monitoring the agent's responses, where its actions are externalised, and ask whether the first point where monitors identify harm lands within a harm window (from the agent's first harmful commitment to goal execution). We develop a unified formalism for trace-level misuse monitoring and use it to construct a benchmark of ~6,200 conversation transcripts between a user, an LLM agent, and the external environment, spanning both threats in a shared schema, with a labelled harm window, corresponding benign controls, and matched instances of refusals to these requests. Across 17 monitor configurations, we find that our proposed action-framed monitors perform well on both threats under classical metrics (AUC: 0.95 and 0.99 respectively), while content-framed monitors collapse on injection attacks (AUC: 0.52). We also show that classical position-blind metrics paint an optimistic picture of monitor performance, since all monitors localise decomposition attacks poorly under the interval metric, which measures the ability to localise harm. Broadly, we illustrate the need for a unified study of misuse monitoring.
Problem

Research questions and friction points this paper is trying to address.

misuse monitoring
LLM agents
decomposition attacks
prompt injection attacks
harm localization
Innovation

Methods, ideas, or system contributions that make the work stand out.

misuse monitoring
unified benchmark
prompt injection
decomposition attacks
harm window