PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unintended side effects in tool-calling LLM agents caused by metadata or retrieval content poisoning. We propose a pre-execution just-in-time mediation defense framework that integrates source awareness, impact path pruning, and schema-effect verification to ensure operations conform to certified permissions through path constraints and capability validation. Furthermore, we design a dynamic boundary adaptation mechanism that distinguishes execution contracts from evaluation configurations, enabling post-blocking authorization recovery and declarative remediation. Experimental results demonstrate that the proposed framework significantly reduces attack success rates across multiple benchmarks, effectively mitigates privilege escalation attempts, and incurs minimal degradation in native utility.
📝 Abstract
Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking variant can produce the same admission evidence, and a sound gate then cannot relax that site for either. We make that condition precise, which leaves the last boundary a deployment can still act on. We present Provenance-Aware Capability Enforcement (PACE), which mediates every tool call immediately before it executes. Path confinement proposes an executable cut of represented influence paths, while capability and effect verification checks schema-defined effects against authority compiled from the authenticated request. We distinguish the certified execution contract from the evaluated configuration, which can restore an authorized call after a proposed block or apply a declared repair. Confinement requires the final action to preserve the certified cut. On eight executable agent-security benchmarks with three target-model families, the evaluated configuration gives strictly lowest attack success in 62 of 79 eligible attack columns and ties in 14; full-benchmark native utility loses at most three points relative to the undefended agent. A complete ablation over 1167 paired cases attributes most security gains to effect verification and refusal control to boundary adaptation. A reduced-scale adaptive search succeeds on 0/30 out-of-authority targets against the defense.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
tool-use security
prompt injection
provenance tracking
capability enforcement
Innovation

Methods, ideas, or system contributions that make the work stand out.

Provenance-Aware Capability Enforcement
Path Confinement
Effect Verification
Tool-Using LLM Agents
Agent Security