Pincer: Resource Authorization for Agents using a Digital Twin

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the risks of external attacks during the long-term operation of autonomous coding agents and the imbalance between security and functionality in existing defense mechanisms. To this end, we propose Pincer, a defense framework that introduces a novel digital twin technique based on an isolated context model. Pincer automatically learns and dynamically enforces user-level least-privilege policies at the resource layer, orchestrating large language models, sandboxing mechanisms, and user behavior modeling to achieve multi-layered defense. Furthermore, we construct a user-centric dataset comprising multi-day interactions. Experimental results demonstrate that Pincer significantly outperforms baseline methods, including LLM-as-a-Judge and Conseca, in both security and utility, exhibiting exceptional protective capabilities against specific attack vectors.
📝 Abstract
Coding agents have become increasingly long-horizon, autonomous, reliant on general-purpose shell and maintain their own persistent memory for self-improvement. While these capabilities have made the agents powerful, they have also made them harder to defend against external adversaries. Defenses that restrict this architecture --- typed tools, information-flow control, or policy prediction engines --- give up too much functionality to be adopted. Agents deployed today (e.g. Claude, Codex) rely on a combination of user-mediated and automode sandboxing as their primary defense. In user-mediated sandboxing, user-maintained policies decay over time and repeated permission requests cause user fatigue, while auto mode's tool-call classifiers learn no user-specific policy and are not meant to defend against adversarial setups. Pincer is a new defense that operates at the resource layer and works alongside existing defenses at the tool-call layer like the auto mode. At the core of Pincer lies a digital twin, an isolated-context model that automatically learns and enforces dynamic user-specific least-privilege policies. The digital twin keeps continually learning the user's preferences allowing it to act as the user's proxy for the agent's permission requests. To emulate the learning phase, we propose a new usercentric dataset with examples following a multi-day transcript of user-agent interaction. Our evaluation shows that Pincer performs strongly on both security and utility in comparison to several baselines which includes variants of LLM judges and adaptations of Conseca (HotOS '25). We highlight attack types where Pincer's design leads to a significant security improvement compared to all other baselines, while outperforming the baselines even for other types of attacks.
Problem

Research questions and friction points this paper is trying to address.

coding agents
resource authorization
sandboxing
adversarial defense
least-privilege policies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Digital Twin
Resource Authorization
Least-Privilege Policy
Coding Agents
User-Centric Dataset
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.