TrustMI: Causally controlling how assistants trust their users

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the security vulnerability arising from erroneous trust in large language model (LLM) assistants when they fail to verify user intent. To mitigate this, the authors construct contrastive dialogue datasets and identify linear directions within the activation space, enabling causal intervention on the model’s trust decisions through activation steering matrices while keeping parameters frozen. This work provides the first demonstration that LLM trust behavior can be controlled both monotonically and causally via such a mechanism. Furthermore, it achieves bidirectional trust regulation across multiple model families, effectively mitigating critical security threats, including harmful requests and prompt injection attacks.
📝 Abstract
Large Language Model (LLM) assistants routinely decide whether they can trust users and third parties whose competence, intentions, and integrity they cannot verify. This uncertainty matters for safety, as trusting the wrong party can lead an agent to comply with harmful requests or act on malicious instructions encountered during tool use. To study this problem, we define trust as an assistant's willingness to accept vulnerability to the actions of another party and ask whether such behavior can be causally controlled through model activations. We build 2,000 contrastive conversations spanning ability, benevolence, and integrity, where paired responses complete the same request but differ in whether the assistant trusts the user. From these pairs, we learn steering matrices while keeping the model parameters frozen and test them across six instruction-tuned models from three families, finding that steering changes trust decisions monotonically in both directions. We then ask whether this effect extends to several safety-related agent settings involving harmful requests, prompt injections, and insider threats, while using benign-task and reasoning as controls. Our findings provide evidence that trust in the user can be causally controlled along linear directions in model activations and provide a way to study how trust shapes safety-relevant behavior in language models.
Problem

Research questions and friction points this paper is trying to address.

Large Language Model trust
AI safety
causal control
prompt injection
model steering
Innovation

Methods, ideas, or system contributions that make the work stand out.

Causal Control
Trust Steering
Activation Engineering
LLM Safety
Contrastive Pairs
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Théo Lasnier
Inria Paris, Sorbonne Université
R
Romain Froger
Inria Paris, Sorbonne Université, Meta SuperIntelligence Labs
Maxence Lasbordes
Maxence Lasbordes
ENS-Ulm / Dauphine / Télécom SudParis
IANLPLLMMachine Learning
Djamé Seddah
Djamé Seddah
Inria (Almanach)
LLMsdata set developmentlow-resource languagesArabic dialectsUGC