🤖 AI Summary
This study addresses the limitation that aggregate consensus alone fails to reveal internal coordination mechanisms within multi-agent large language model (LLM) systems. To investigate this, we construct a system of fifty stateless LLM agents, integrating reasoning effort with communication topology. Through graph-theoretic analysis and cross-model comparative experiments, we identify collective states including synchronization, distortion, and chimera. Notably, this work introduces the concept of "distortion"—a state characterized by local order yet global incoherence—demonstrating for the first time that low spatial heterogeneity does not necessarily imply global consensus. Furthermore, our findings reveal that increased reasoning effort induces distorted states, whereas high network connectivity drives global synchronization. Collectively, these results provide profound insights into the inherent complexity of coordination mechanisms in multi-agent LLM systems.
📝 Abstract
Multi-agent LLM systems are increasingly used for deliberation and evaluation, often under the assumption that greater peer interaction leads to more reliable consensus. Existing work largely evaluates these systems through final accuracy or aggregate agreement. However, such measures do not reveal how agreement is organized in the panel. In this paper, we study \(N=50\) stateless LLM agents that update their predictions from locally visible peers, and characterize their behavior using both global and local measurements of agreement. We identify three collective regimes: synchronised, twisted (locally ordered but globally incoherent) and chimera-like, where coherent and incoherent subpopulations coexist. Increasing reasoning effort in gpt-5-mini shifts panels from variable, often fragmented outcomes toward locally ordered twisted states, and a small follow-up shows such states can also form from permuted initial conditions, whereas increasing communication connectivity drives them toward global synchronisation. Fragmentation collapses faster as algebraic connectivity increases across rewired graphs. The topology effect also appears on a non-circular judging task and across models from three providers. Finally, low spatial heterogeneity does not guarantee global consensus: 40\% of trials with $ΔZ$ below 0.03 retain a twisted configuration through the final 20 turns. These results show that reasoning effort and communication topology control different aspects of multi-agent coordination, and that aggregate agreement alone is insufficient to characterize collective LLM behavior.