Consensus and Factual Dynamics in Large Populations of Interacting Language Models

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study investigates how group consensus emerges among large language models (LLMs) in multi-agent interactions and its relationship with factual correctness. To this end, it proposes the RHEON framework, which models LLM populations as a physics-inspired O(n) spin system. Leveraging Glauber asynchronous dynamics, multi-geometric topology simulations, and a large-scale evolutionary corpus, the framework systematically analyzes consensus dynamics across diverse network topologies. The results reveal nonlinear regulatory mechanisms through which temperature and coupling topology govern hallucination-driven consensus. Furthermore, the study demonstrates that an increased number of neighboring nodes significantly accelerates convergence. While semantic consistency is positively correlated with factual correctness, the former does not constitute a sufficient guarantee for the latter.
πŸ“ Abstract
Large Language Model (LLM) agents are increasingly deployed as populations of interacting entities, in which consensus --agreement on a shared answer-- emerges as a collective, unengineered behaviour. Prior work on LLM consensus shows that agents can cross-verify their answers and converge towards more factual responses, treating agreement as a proxy for correctness. However, these studies usually fix a single interaction structure, leaving open how consensus depends on how agents interact. We address this gap by introducing RHEON, a physics-inspired framework that recasts a population drawn from a single frozen model as an evolving $O(n)$ spin system on a ladder of interaction geometries of increasing effective dimension --from a 1D ring to a full-coupling mean-field graph-- with the sampling temperature $T$ as the tunable source of thermal disorder, evolved through a Glauber-like asynchronous dynamics. Sweeping RHEON across $432$ configurations of prompt, population size, communication topology, and sampling temperature yields Eraclitus-4.7M, a tagged evolutionary corpus of $4.7$ million responses. We find that agents reach their strongest consensus gain within the first few update sweeps and that increasing the number of neighbours per agent accelerates convergence on average. We further show that whether a configuration settles on factually correct or hallucinated consensus is not predictable from its initial state alone, and that the hallucination-minimising temperature depends on how the agents are coupled, so the common near-greedy default is not automatically the safest. Finally, semantic agreement correlates positively with factual convergence, and interaction strengthens the association, yet never enough for unanimity to certify correctness.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Consensus dynamics
Interaction topology
Hallucination
Factual correctness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interacting LLM agents
Spin system dynamics
Communication topology
Consensus emergence
Hallucination mitigation
E
Emanuele Ricco
Computer, Electrical and Mathematical Sciences and Engineering (CEMSE) Division, King Abdullah University of Science and Technology (KAUST), Thuwal 23955, Saudi Arabia
Elia Onofri
Elia Onofri
CNR - Institute for Applied Mathematics "Mauro Picone" (IAC), Italy
CryptographyAlgorithmsNumerical SimulationsGraph TheoryComputational Biology
V
Vincenzo Sammartino
Computer, Electrical and Mathematical Sciences and Engineering (CEMSE) Division, King Abdullah University of Science and Technology (KAUST), Thuwal 23955, Saudi Arabia; Dipartimento di Informatica, UniversitΓ  di Pisa, Pisa, Italy
Roberto Di Pietro
Roberto Di Pietro
IEEE Fellow; ACM Distinguished Scientist; Full Professor of Cybersecurity, KAUST
AI driven CybersecurityDistributed Systems SecurityWireless SecurityOSN Security and Privacy