🤖 AI Summary
This study addresses the challenge of formally specifying, auditing, and reusing safety policies for LLM agents, particularly in multi-step data-dependent scenarios. We propose an offline runtime verification framework based on Metric First-Order Temporal Logic (MFOTL). Leveraging the MonPoly monitor to replay and analyze AgentDojo and STAC traces, we evaluate its capability to detect generic obligation-violating attack chains. Our investigation reveals that existing benchmarks lack sufficient precision due to missing critical context, prompting us to introduce a binding provenance mechanism to defend against evasion attacks. Experiments demonstrate that the framework detects 71.8% of STAC and 70.1% of AgentDojo attack chains, confirming that exposing trusted execution histories is a prerequisite for formal monitoring to be effective.
📝 Abstract
Guardrails for tool-using LLM agents are usually application-specific rules, which makes multi-step, data-dependent safety policies hard to specify, audit and reuse. As a declarative alternative, we evaluate metric first-order temporal logic (MFOTL), replaying the recorded trajectories that AgentDojo, STAC and R-Judge already ship through the unmodified MonPoly monitor, offline and without running an agent. On these corpora, five generic obligations flag 71.8% of STAC attack chains and 70.1% of successful AgentDojo attacks, but also fire on 29.3% of benign runs. This imprecision stems from the corpora rather than the logic: they rarely record approvals and never record timestamps, so history-dependent obligations reduce to detecting risky action types. Where the trace does carry relational context, provenance-aware policies discriminate better; that context, however, is itself attackable, and one planted line defeats a naive provenance check on 94-99% of the runs it would otherwise flag. Binding provenance to the lookup that produced it closes this evasion at no cost in detection or benign firing. Taken together, these results show that formal temporal monitoring adds value exactly when the trace exposes trustworthy history. We therefore quantify how far current benchmarks are from that point and propose a twelve-field enforcement-ready trace schema.