ToolFence: Fine-Grained Authorization for Secure Tool-Using LLM Agents

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the risks of indirect prompt injection and tool parameter tampering in LLM agents arising from context confusion. To mitigate these threats, we propose a fine-grained, provenance-aware defense framework grounded in typed authorization blueprints. The approach prevents intra-tool attacks by strictly isolating user-authorized values from untrusted observational data. It further enables fast-path execution through statically compiled blueprints paired with deterministic monitors, while dynamically provisioning runtime capabilities when blueprints are unavailable. Evaluated on the AgentDojo benchmark, the proposed method reduces the attack success rate to near zero with only a 3.8% utility loss. By maintaining low latency alongside robust security guarantees, this framework demonstrates substantial practical value for real-world deployment.
📝 Abstract
Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allowing malicious content to steer consequential input-filtering defenses. Multi-path consensus defenses still leave a high attack success rate because they examine content or aggregated outputs rather than authorizing effects, especially for the within-tool attack, which preserves the intended tool but manipulates its arguments. Data-Flow Control such as CaMeL provides stronger guarantees, but incurs substantial time latency that limits practical deployment. We introduce ToolFence, which compiles a typed authorization blueprint before execution, enforces it through a deterministic monitor, and when the blueprint is incomplete asks a judge to grant new capabilities rather than adjudicate each concrete call. ToolFence provides two key advantages. First, its fine-grained provenance-aware authorization enables the system to distinguish user-authorized values from untrusted observations, effectively addressing the within-tool attack. Second, its deterministic fast path and capability-level runtime grants substantially reduce the frequency of expensive judge calls, improving runtime efficiency. On AgentDojo with Qwen3-max, ToolFence reduces overall ASR to near zero with only a 3.80 percentage-point clean-utility drop and practical runtime overhead.
Problem

Research questions and friction points this paper is trying to address.

indirect prompt injection
tool-using LLM agents
within-tool attack
fine-grained authorization
runtime efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Indirect Prompt Injection
Fine-Grained Authorization
Provenance-Aware
Deterministic Monitor
Tool-Using LLM Agents