π€ AI Summary
This study investigates how different message types influence cooperative equilibria among large language model (LLM) agents in multi-agent games. By constructing a simulation environment in which four agents engage in four distinct game classes, the authors compare the effects of natural language, numerical signals, and random sequences on strategic behavior. The findings reveal that although structured messages alter payoffs, they lack stable convergence patterns. Furthermore, numerical signals deviate from randomness and correlate with payoff structures, exhibiting monitorable, message-level fingerprints. Based on these insights, this work proposes a novel paradigm for AI agent coordination: prioritizing the monitoring of message fingerprints over behavioral decisions.
π Abstract
Large language model (LLM)-based agents increasingly operate in multi-agent systems (MAS) characterised by strategic interaction. However, little is known about whether, and to what extent, different types of messages affect the outcomes of strategic games. By investigating AI agents based on four popular LLMs, playing four games with different cooperation equilibria, we study whether messages of different kinds (natural language, numerical signals, or random sequences) significantly modify the levels of cooperation in each game, also depending on the agents' assigned personalities. We observe that structured messages alter the final payoffs for most games and LLMs, but without a predictable pattern; this challenges the assumption that AI agents can converge to stable equilibria regardless of additional capabilities. Moreover, we observe that agent-generated numerical messages depart from randomness, most strongly and consistently when agents are explicitly instructed to communicate; however, they introduce an additional interpretability challenge, as their symbol distributions are mostly associated with the payoff structure and typically become more concentrated with repetition, but are overall difficult for humans to interpret. Monitoring for coordination of AI agents through restricted channels should thus prioritise message-level fingerprints, which generalise across models, over behavioural decisions, which do not.