🤖 AI Summary
This study addresses the security risks of operating system-level side effects in AI agents caused by prompt injection or planning errors. To mitigate these threats, we propose a Linux reference monitor built upon the eBPF Linux Security Module (LSM) framework. This approach pioneers binding human final decision-making authority to a kernel-level contract mechanism enforced prior to task execution. By introducing a three-state asset contract model, it deterministically enforces damage boundaries for untrusted code, thereby achieving robust asset isolation. Experimental evaluations demonstrate that the proposed system passes all 19 security tests while maintaining file I/O overhead lower than the ActPlane baseline. These results validate both the feasibility and efficiency of kernel-level enforcement for securing autonomous AI agents against unintended system interactions.
📝 Abstract
Large language model agents can execute commands, create subprocesses, and directly access files and networks, allowing prompt injection or planning errors to become operating-system side effects. We present ContractWarden, a Linux reference monitor that enforces a human-authorized damage boundary without trusting the agent or its policy suggestions. A model may propose a tri-state asset contract - allow, deny, or no_egress - but a human makes the final choice. An execution gate binds the contract to a concrete task before untrusted code runs. An extended Berkeley Packet Filter (eBPF) Linux Security Modules (LSM) data plane then enforces file and network decisions and monotonically propagates no_egress through processes, regular files, pipes, FIFOs, and supported Unix-domain sockets. All 570 runs across 19 security tests satisfy predefined return-value and side-effect criteria. On three co-located file-I/O workloads, median overhead is 11.96-12.89% in a Linux 6.15 virtual machine and 35.79-61.54% on a Linux 6.15 physical platform, lower than the evaluated frozen ActPlane baseline. The results demonstrate deterministic kernel enforcement for declared assets, supported paths, and controlled object lifecycles.