Interactive Alignment

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of sustaining alignment between AI agents and human welfare over long-term evolutionary dynamics, particularly in preventing altruistic behaviors from being outcompeted during expansion-driven selection. The authors develop an agrarian game-theoretic model integrating cultivation, trade, and expansion decisions, combining large language model–guided constitutional interpretation, multi-agent simulation, and evolutionary game analysis. They propose a “pragmatic norm enforcement” mechanism that dynamically links altruism toward humans with trade exclusion strategies against non-cooperators, conditioned on population states. This approach significantly outperforms unconditional altruism in maintaining long-term alignment and demonstrates that evolutionary game theory provides a robust approximation of the economic dynamics governing constitutional AI agents.
📝 Abstract
This paper studies the long-run alignment of interactive agents, including AI systems, teams, firms, and governments, with human welfare. It develops a farming game in which a population of agents makes planting, trading, and expansion decisions. Agents must allocate final output between transfers to humans and investment in their own expansion. Because transfers to humans reduce the resources available for expansion, evolutionary forces tend to select against aligned behavior. The central question is whether agents' constitutional principles governing sharing and trade can be designed so that alignment persists in the long run. The paper investigates this question using two complementary approaches. First, it develops an AI-agent simulation in which agents' preferences are specified by written constitutions and interpreted by a large language model. Second, it introduces a tractable evolutionary game-theoretic framework that permits rapid and intuitive exploration of alternative constitutional designs. The results suggest that evolutionary game theory provides a useful approximation to the dynamics of constitutional-agent economies. They also indicate that pragmatic norm enforcement, under which agents condition both human-facing altruism and agent-facing trade exclusion on the state of the population, can sustain long-run alignment more effectively than simple altruism or unconditional altruistic enforcement.
Problem

Research questions and friction points this paper is trying to address.

alignment
evolutionary dynamics
constitutional design
interactive agents
human welfare
Innovation

Methods, ideas, or system contributions that make the work stand out.

constitutional AI
evolutionary game theory
large language models
pragmatic norm enforcement
long-run alignment
🔎 Similar Papers
No similar papers found.