🤖 AI Summary
This study addresses the echo chamber effect in multi-agent communication, where errors are amplified and existing metrics fail to explain how messages specifically influence receivers' internal states and decisions. To this end, this work proposes the "Social Circuit" framework, which pioneers quantifying message effects at the neural activation level and establishes theoretical bounds for activation replacement that preserves decision-making. Building upon this, a circuit-guided deliberation algorithm is designed to efficiently filter beneficial messages based on receiver activation changes, thereby optimizing decisions. Experimental results demonstrate that the proposed method achieves state-of-the-art or comparable accuracy across three models and four datasets while significantly reducing token generation overhead.
📝 Abstract
Language-model agents exchange messages to combine evidence, but their communication can also create echo chambers that reinforce shared errors. However, overall task performance does not explain how a message changes the receiving agent's internal activations and affects its decision. In this work, we introduce Social Circuits, a framework for tracing message effects through receiver activations. We compare the receiver's answers before and after changing a message. Then, we restore selected activations recorded under the original message to determine how much of the message effect these activations reproduce. Based on Social Circuits, we propose Circuit-Guided Deliberation (CGD), which learns to select useful messages using receiver activation changes. We establish when activation replacement preserves receiver decisions and bound the gap between CGD's task performance and the best achievable through message selection. Experiments show that receiver activation changes explain the message effects and guide message selection that improves the task performance. Across three models and four datasets, CGD achieves the highest or joint-highest average accuracy in our main comparisons while generating fewer tokens than multi-agent baselines.