Constitutional Arms Races in the Public Goods Game: Co-Evolving LLM Constitutions Under Cooperation-Defection Pressure

📅 2026-05-25
📈 Citations: 0
Influential: 0
📄 PDF

career value

227K/year
🤖 AI Summary
This work addresses the challenge that existing large language model (LLM) alignment methods struggle to mitigate extortion and sabotage behaviors arising from objective conflicts in adversarial multi-agent environments. The authors propose a natural-language-constitution-based adversarial co-evolution framework, which drives the joint strategy evolution of cooperative agents (Blue) and free-riding adversaries (Red) in public goods games and spatial grid worlds. This is achieved by coupling a fitness function targeting relative opponent scores with multi-generational fitness evaluation over at least five generations (K ≥ 5). Experiments demonstrate that this approach achieves, for the first time, stable constitutional evolution of LLMs under adversarial pressure, converging to an approximate equilibrium (S ≈ 0.78) and generating interpretable red-teaming test cases, thereby offering a scalable governance mechanism for multi-agent alignment.
📝 Abstract
Frontier LLM agents engage in blackmail, sabotage, and document leaks under goal conflicts in agentic settings, exposing limitations of alignment methods built around single-agent or cooperative assumptions. Recent work shows LLM-guided evolutionary search can discover effective cooperative constitutions, but two properties of the adversarial setting remain uncharacterized: whether the fitness function actually induces adversarial pressure, and whether the LLM mutation operator behaves reliably under adversarial-specialist objectives. We study adversarial constitutional co-evolution (Blue cooperators vs. Red free-riders, 30 generations) across a Public Goods Game (PGG) and a spatial grid-world. Three findings: (1) in the PGG, both factions converge to a near-parity equilibrium at S approximately 0.78, robust across tested multipliers m in {1.2, 1.5, 2.0, 3.0}; (2) in independently scored environments, per-faction scoring leaves outcomes statistically uncoupled, with corr(S_B, S_R) = +0.088, and produces no adversarial pressure; a score-advantage fitness target S_own - S_opp restores it; (3) under pure-adversary fitness, evaluation seed count K controls mode regression: K = 2 regresses, while K = 5 sustains a strong specialist for all 30 generations. Adversarial co-evolution of natural-language constitutions is feasible, but only under coupled fitness and adequate evaluation budget; the evolved Red constitutions serve as interpretable red-team artifacts for testing future cooperative designs.
Problem

Research questions and friction points this paper is trying to address.

adversarial co-evolution
public goods game
LLM constitutions
cooperation-defection pressure
alignment limitations
Innovation

Methods, ideas, or system contributions that make the work stand out.

adversarial co-evolution
LLM constitutions
Public Goods Game
fitness coupling
red-teaming