Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind

📅 2026-04-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of how large language models can leverage partial prior knowledge to protect privacy and mislead adversaries in adversarial dialogues. We propose a Theory of Mind (ToM)-based double-agent defense framework that introduces the novel ToM-SB task, revealing a bidirectional emergent relationship between belief modeling and deceptive capability. By jointly optimizing these two components, our approach enhances defensive performance. We employ reinforcement learning with a reward function that integrates ToM consistency and deception success rate to train AI agents capable of belief-guided reasoning. Experimental results demonstrate that our method significantly outperforms state-of-the-art models such as Gemini3-Pro and GPT-5.4 across diverse in-distribution and out-of-distribution attack scenarios, exhibiting superior generalization and proactive deception capabilities.

Technology Category

Multiagent Systems: Adversarial AgentsMachine Learning: Adversarial Learning & RobustnessGame Theory and Economic Paradigms: Adversarial Learning

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Agentic searchUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systems
📝 Abstract
As large language models (LLMs) become the engine behind conversational systems, their ability to reason about the intentions and states of their dialogue partners (i.e., form and use a theory-of-mind, or ToM) becomes increasingly critical for safe interaction with potentially adversarial partners. We propose a novel privacy-themed ToM challenge, ToM for Steering Beliefs (ToM-SB), in which a defender must act as a Double Agent to steer the beliefs of an attacker with partial prior knowledge within a shared universe. To succeed on ToM-SB, the defender must engage with and form a ToM of the attacker, with a goal of fooling the attacker into believing they have succeeded in extracting sensitive information. We find that strong frontier models like Gemini3-Pro and GPT-5.4 struggle on ToM-SB, often failing to fool attackers in hard scenarios with partial attacker prior knowledge, even when prompted to reason about the attacker's beliefs (ToM prompting). To close this gap, we train models on ToM-SB to act as AI Double Agents using reinforcement learning, testing both fooling and ToM rewards. Notably, we find a bidirectionally emergent relationship between ToM and attacker-fooling: rewarding fooling success alone improves ToM, and rewarding ToM alone improves fooling. Across four attackers with different strengths, six defender methods, and both in-distribution and out-of-distribution (OOD) evaluation, we find that gains in ToM and attacker-fooling are well-correlated, highlighting belief modeling as a key driver of success on ToM-SB. AI Double Agents that combine both ToM and fooling rewards yield the strongest fooling and ToM performance, outperforming Gemini3-Pro and GPT-5.4 with ToM prompting on hard scenarios. We also show that ToM-SB and AI Double Agents can be extended to stronger attackers, demonstrating generalization to OOD settings and the upgradability of our task.
Problem

Research questions and friction points this paper is trying to address.

Theory of Mind
Belief Steering
Privacy
Adversarial Dialogue
Double Agent
Innovation

Methods, ideas, or system contributions that make the work stand out.

Theory of Mind
belief steering
double-agent defense
reinforcement learning
adversarial dialogue