CEO Arena: Evaluating Long-Horizon Multi-Agent Decision-Making in Competitive Markets

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of evaluating agents’ long-term strategic decision-making and competitive adaptability in uncertain markets. To this end, it constructs an eight-firm simulated market benchmark and introduces a “match-replacement evaluation” mechanism, wherein large language model-driven CEO agents engage in long-horizon multi-agent games involving pricing and R&D decisions, effectively disentangling individual returns from market externalities. The findings reveal that most agents yield negative average returns, indicating that private gains frequently entail aggregate market losses. Furthermore, the analysis identifies critical interaction patterns, such as demand capture and competitor pricing responses. Collectively, this work establishes a rigorous, controlled evaluation framework for investigating the economic behavior of large language models.
📝 Abstract
Long-horizon competition tests agents'ability to coordinate business decisions under uncertainty and adapt to changing rival strategies. We introduce CEO Arena, a benchmark that uses matched replacement evaluation to assess operating returns alongside an agent's effects on rivals and the market. Each CEO agent is compared with a reference policy in the same company under the same economic seed, holding other agents'identities and assignments fixed while all agents adapt. In a shared eight-company market spanning 500 simulated days, CEOs make sequential decisions on pricing, procurement, marketing, research and development, and service using private company information and noisy market signals, under resource constraints and delayed feedback. We evaluate eight LLM-based CEO agents in 27 main runs and 26 robustness runs. In the main evaluation, most agents have negative mean returns, and private gains can accompany market losses. Robustness analyses suggest that aggregate patterns extend beyond the original rule-based baseline; four of the 56 directed pairs show relatively stable effects. Memory, action, and accounting traces suggest demand capture and rivals'pricing and spending responses as possible explanations. CEO Arena provides a controlled testbed for studying long-horizon agent competition, strategic interaction, and market externalities.
Problem

Research questions and friction points this paper is trying to address.

long-horizon decision-making
multi-agent competition
benchmark evaluation
strategic interaction
market externalities
Innovation

Methods, ideas, or system contributions that make the work stand out.

Long-Horizon Multi-Agent Decision-Making
Matched Replacement Evaluation
Competitive Market Simulation
LLM-based Agents
Market Externalities
🔎 Similar Papers
No similar papers found.