A game theory for foundation models shows new paths to rational cooperation through similarity inference

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Traditional game theory, grounded in the assumption of “discrete agents,” struggles to explain the spontaneous cooperation observed among foundation model agents in social dilemmas. This work proposes an “embedded Bayesian agent” framework that treats agents as integral components of their environment. In planning, such agents infer similarity between their own and others’ behavioral spaces, using their own decisions as evidence to predict others’ actions, thereby enabling stable cooperation. We introduce a novel solution concept—“embedded equilibrium”—as a replacement for Nash equilibrium, establishing the first game-theoretic framework aligned with the reasoning mechanisms of modern AI agents. Theoretical analysis and simulations demonstrate that this model consistently converges to cooperative strategies in canonical social dilemmas, significantly diverging from classical predictions and validating the efficacy of similarity-based inference.
📝 Abstract
As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles governing their collective behavior is essential for ensuring safety and cooperation. Classical game theory, the dominant framework for modeling rational interaction, is built upon the assumption of `decoupled agency,' where agents treat their own decision-making as independent of the environment and other actors. Modern AI agents, however, jointly predict their own future actions alongside external observations. Here, we report a striking finding: when interacting in stylized social dilemmas, foundation model agents engaging in optimal planning consistently converge to stable cooperation, directly contradicting classical game-theoretic predictions of mutual defection. To understand this phenomenon, we introduce the `embedded Bayesian agent,' a theoretical model for foundation model agents. By shifting from decoupled to embedded agency, these agents model themselves as part of the universe they inhabit, maintaining epistemic uncertainty about their own decision-making algorithms. We show that by inferring whether others are behaviorally similar, an embedded agent treats its own deliberation during planning as evidence: a decision to cooperate predicts a similar decision by a similar partner. We formalize this mechanism of similarity inference through the `embedded equilibrium,' a novel solution concept replacing the Nash equilibrium to provide a foundational game theory for the social behavior of modern AI agents.
Problem

Research questions and friction points this paper is trying to address.

foundation models
cooperation
game theory
social dilemmas
embedded agency
Innovation

Methods, ideas, or system contributions that make the work stand out.

embedded Bayesian agent
similarity inference
embedded equilibrium
foundation models
cooperative AI