🤖 AI Summary
This study addresses the security vulnerability arising from erroneous trust in large language model (LLM) assistants when they fail to verify user intent. To mitigate this, the authors construct contrastive dialogue datasets and identify linear directions within the activation space, enabling causal intervention on the model’s trust decisions through activation steering matrices while keeping parameters frozen. This work provides the first demonstration that LLM trust behavior can be controlled both monotonically and causally via such a mechanism. Furthermore, it achieves bidirectional trust regulation across multiple model families, effectively mitigating critical security threats, including harmful requests and prompt injection attacks.
📝 Abstract
Large Language Model (LLM) assistants routinely decide whether they can trust users and third parties whose competence, intentions, and integrity they cannot verify. This uncertainty matters for safety, as trusting the wrong party can lead an agent to comply with harmful requests or act on malicious instructions encountered during tool use. To study this problem, we define trust as an assistant's willingness to accept vulnerability to the actions of another party and ask whether such behavior can be causally controlled through model activations. We build 2,000 contrastive conversations spanning ability, benevolence, and integrity, where paired responses complete the same request but differ in whether the assistant trusts the user. From these pairs, we learn steering matrices while keeping the model parameters frozen and test them across six instruction-tuned models from three families, finding that steering changes trust decisions monotonically in both directions. We then ask whether this effect extends to several safety-related agent settings involving harmful requests, prompt injections, and insider threats, while using benign-task and reasoning as controls. Our findings provide evidence that trust in the user can be causally controlled along linear directions in model activations and provide a way to study how trust shapes safety-relevant behavior in language models.