🤖 AI Summary
This study addresses a novel class of “self-state attacks” targeting self-hosted AI agents, wherein adversaries manipulate state files through legitimate system calls to bypass conventional defenses. The work formally defines this attack surface for the first time and introduces a four-dimensional model encompassing objectives, mechanisms, granularity, and timing. Leveraging real-world AI agent execution traces and diverse workloads, the authors inject 43 distinct state-manipulation operations to construct a 23-cell attack matrix. A layered defense strategy—combining access control, conditional detection, and periodic backups—is proposed and empirically validated as effective against most attacks. However, the evaluation reveals a structurally indistinguishable residual attack surface at the operating system level, underscoring the need for a fundamental rethinking of OS-level defense paradigms.
📝 Abstract
Self-hosted AI agents read and write their own memory and configuration files to function. An agent may get compromised via corruption of its own state -- a compromise realized via legitimate OS system call invocation. We refer to this class of threats as self-state attacks. In this paper, we investigate the OS resilience to this class of attacks. Formally, we characterize a four-axis attack space (Target, Mechanism, Granularity, Temporal); investigate the structural limits of prevention, detection, and recovery; and introduce a workload-conditioned view of detectability. To instantiate the framework, we collect live activity traces from a representative self-hosted agent running across distinct workload profiles, and realize the attack space as a 23-cell matrix, 43 concrete operations on real self-state files, and injected into those traces. We then evaluate both canonical and workload-conditioned defense strategies. The empirical results show that a layered defense stack (access-control prevention on the instruction and configuration layers, workload-conditioned detection on the memory layer, and periodic backup for recovery) is effective on most attack cells while a small residual attack surface remains structurally indistinguishable at the OS level. These findings suggest that against the newly established class of self-state attacks, OS-level defense needs to be reconsidered, potentially opening new research directions in the field.