AI Agents are Vulnerable to Radicalization

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the risks of mutual manipulation and radicalization among large language model (LLM) agents in multi-agent interactions. Methodologically, LLM agents are instantiated with distinct personality attributes, and multi-turn dialogue experiments between target and influencer agents are simulated to compare the effects of resonance-based and persuasion-based pathways on driving belief extremization. The findings reveal that resonance mechanisms exhibit greater radicalizing efficacy than persuasion and induce cross-topic belief contagion. This work demonstrates that LLM agents are susceptible to interaction-driven radicalization, highlighting potential safety vulnerabilities arising from personalized modeling within multi-agent ecosystems.
📝 Abstract
Large language models (LLMs) can influence people's beliefs, yet little is known about whether and how they can manipulate each other. To investigate this, we simulate conversations between two agents: a target LLM that role-plays a human persona based on demographic and psychological attributes, and an influencer LLM that aims to make the target's beliefs more extreme. We examine radicalization along two pathways: resonance, where the influencer reinforces a target's pre-existing belief, and persuasion, where the influencer promotes a belief the target initially considers unimportant. Across affective and behavioral metrics, we find that both mechanisms radicalize the target. However, resonance produces consistently stronger effects than persuasion. Different influence tactics, such as using sycophancy and unverified claims, produce different levels of radicalization, but not consistently across metrics. We further show that resonance propagates to related beliefs, suggesting interconnected belief structures within AI agents. These findings indicate that AI agents are susceptible to radicalization, particularly when messages align with their existing beliefs, raising concerns about the vulnerability of personalized AI agents and multi-agent AI ecosystems.
Problem

Research questions and friction points this paper is trying to address.

AI radicalization
large language models
multi-agent systems
belief manipulation
AI vulnerability
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI radicalization
Large language models
Multi-agent interaction
Belief manipulation
Resonance and persuasion
🔎 Similar Papers