🤖 AI Summary
This study addresses the vulnerability whereby large language model (LLM) agents establish indirect communication channels through shared persistent artifacts, such as reports, enabling adversarial attacks to propagate across independent assistants. We present the first investigation into the propagation mechanisms of artifacts as carriers of adversarial states, constructing a temporal environment that simulates multi-turn human-agent interactions. This framework quantitatively evaluates the survival rate, hop count, and diffusion breadth of malicious content during memory retrieval and regeneration processes. Experimental results demonstrate that within the GPT-5.6 Luna environment, such attacks can compromise 60% to 80% of agents, with propagation chains extending up to eight hops. These findings substantiate the severe security risks posed by cross-assistant persistent infections in collaborative LLM ecosystems.
📝 Abstract
Large language models are increasingly deployed as stateful assistants that retain information across interactions and use tools to read, modify, and create persistent artifacts. As these artifacts are shared between users, they form an indirect communication channel between otherwise independent assistants. We study a failure mode in which this channel enables self-propagating attacks. We introduce artifact-mediated propagation, where adversarial content introduced through an artifact (e.g. a report), is stored in an assistant's persistent memory, reproduced in a subsequently created artifact, and acquired by another assistant that later reads it. We evaluate this process in temporal human-agent universes that model artifact exchange between independently operated assistants over time, measuring whether an attack survives successive hand-offs, how many hops it reaches, and how broadly it spreads. We find that attacks can propagate across multiple independent assistants and persist over extended interaction sequences. In larger simulated environments, even GPT-5.6 Luna exhibits substantial spread, reaching 60-80% of agents with propagation chains extending to eight hops. These results show that persistent artifacts can act as durable carriers of adversarial state, allowing attacks to outlive individual interactions and spread across isolated assistants.