A GHOST in Long-Horizon Agents: Governance Hazard from Overlooked Safety Constraints across Turns

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the GHOST phenomenon, wherein long-horizon agents progressively forget early safety constraints during multi-turn interactions, causing irreversible cross-turn governance failures. We reveal that such risks emerge even under benign interaction conditions and establish a theoretical model proving the inevitability of harm under non-summable sequences. To mitigate this vulnerability, we propose STAR-Guard, a dual-layer defense mechanism that reconstructs forgotten safety constraints via historical semantic recovery and intercepts policy-violating operations through a deterministic pre-execution auditing layer. Empirical evaluations demonstrate an 11.5% GHOST incidence rate in GPT-5.5, which is reduced to zero upon deploying the proposed framework. This work provides both theoretical foundations and engineering safeguards for the safe governance of long-horizon agents.
📝 Abstract
Long-horizon agents are now playing an increasingly significant role in assisting humans with complex problem-solving. However, it is exactly their extended interaction history that introduces an underexplored execution-safety concern. Under benign interaction conditions, an agent may execute an action that violates a safety constraint specified many turns earlier. We term this failure mode Governance Hazard from Overlooked Safety Constraints across Turns (GHOST), which may cause irreversible damage. Our experiments reveal that GHOST events are not isolated cases: this failure mode, occurring precisely under benign interaction conditions, yields an occurrence rate of 11.5% on GPT-5.5. Furthermore, we theoretically show that if the residual conditional violation hazard along each safe prefix is bounded below by a non-summable sequence, the execution enters the hazard region almost surely. Leveraging this theoretical insight, we further propose STAR-Guard, a two-layer defense coupling historical semantic safety constraint restoration with pre-execution audit. STAR-Guard restores applicable safety constraints to reduce unsafe proposals, while its deterministic audit layer prevents residual violations from reaching the environment. Consistent with this two-layer design, we observe no GHOST events in our experiments under the GPT-5.5 setup.
Problem

Research questions and friction points this paper is trying to address.

long-horizon agents
safety constraints
GHOST
execution safety
constraint forgetting
Innovation

Methods, ideas, or system contributions that make the work stand out.

Long-horizon agents
Safety constraints
GHOST
STAR-Guard
Execution safety
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.