Agent Safety From Within: Detecting Harmful Trajectories from LLM Internal States

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of traditional content moderation in detecting harmful content and unsafe tool calls within LLM agent trajectories. We propose TACIT, a method that analyzes internal model representations to decode trajectory safety directly from a frozen backbone using linear probes. This work is the first to demonstrate that these two categories of harm are orthogonal and linearly separable in representation space, enabling safety detection without token generation. Experimental results show that TACIT improves the macro F1 score from 62.3% to 86.2%, reduces trainable parameters by six orders of magnitude, and decreases inference latency to one-sixth of the baseline. These findings establish TACIT as an efficient, low-latency approach for trajectory-level safety monitoring in autonomous agents.
📝 Abstract
Language model agents can now perform sophisticated sequences of actions via tools and harnesses, which has increased the scope of the damage they can cause. Guard models, however, are mainly built for content moderation and thus are not well-suited to detecting this agentic risk. To address this, we proceed by first conducting a representational analysis, then use the resulting insights to build a solution. In our analysis, we focus on two types of trajectory-level agentic harms: harmful content, which is expressed directly, and unsafe tool use, which depends on whether an action is consistent with the interaction that produced it. We investigate how open-source guard models represent these two types of harm and find that they are linearly readable inside the model, even though guard models predict no better than chance on pairs that differ only in the called tool's schema. The two harm types also follow nearly orthogonal internal directions, and neither reliably serves as a proxy for the other. These results motivate reading trajectory safety directly from internal states. We introduce TACIT, a readout of a frozen backbone's internal states that decodes no tokens. Trained on six trajectory-safety benchmarks, a linear probe raises mean macro-F1 from 62.3 for the strongest open guard to 80.7, and refined readouts reach 86.2. With each benchmark held out of training entirely, the refined readouts still lead the strongest guard (65.7 vs. 61.1). With the same backbone, training data and test split, the frozen readout is on par with full safety fine-tuning, and it improves the fine-tuned model further when applied on top. The probe trains about one millionth as many parameters as full fine-tuning in about a sixth of the time, and TACIT has the lowest latency of the guards we evaluate.
Problem

Research questions and friction points this paper is trying to address.

agent safety
trajectory-level harm
unsafe tool use
guard models
internal states
Innovation

Methods, ideas, or system contributions that make the work stand out.

Internal States
Agent Safety
Linear Probing
Trajectory-level Harm
TACIT
🔎 Similar Papers
No similar papers found.