Game Arena: Strategic LLM Evaluation in Competitive Environments

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of static benchmarks, which often lead to performance saturation in large language models (LLMs) and fail to adequately evaluate dynamic strategic capabilities. To overcome this, we construct an open competitive platform based on chess, poker, and Werewolf. Methodologically, we replace static test sets with dynamically evolving game environments spanning perfect-information, imperfect-information, and multi-player scenarios to surpass existing performance ceilings. Supported by robust infrastructure and large-scale ground-truth evaluation techniques, our platform enables reproducible, transparent, and scalable automated head-to-head adversarial benchmarking. This work effectively validates LLMs' proficiency in strategic planning, environmental adaptation, and robustness under uncertainty, establishing a novel paradigm for evaluating the dynamic capabilities of foundation models.
📝 Abstract
We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where the gameplay strength naturally increases as models evolve, preventing performance saturation. This technical report details the infrastructure behind Game Arena and describes the three pilot game environments: Chess, Poker, and Werewolf. These environments span perfect information, imperfect information, and multiplayer game settings, enabling a systematic study of models' strategic planning, adaptation, and robustness under uncertainty. For each game, we provide a detailed description of the environment, evaluation metrics, and results from running full competitions across models. Through robust infrastructure and large-scale ground-truth based evaluation, Game Arena ensures reproducibility, transparency and generalizability to new games and variants over time.
Problem

Research questions and friction points this paper is trying to address.

LLM evaluation
static benchmarks
performance saturation
strategic planning
competitive environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

Game Arena
LLM Evaluation
Competitive Games
Dynamic Benchmark
Strategic Planning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Bovard Doerschuk-Tiberi
Yao Yan
Yao Yan
University of Electronic Science and Technology of China
cutting chattercapsule robotexoskeleton robot
J
Justin Chiu
H
Hann Wang
T
Timothy Chung
M
Martyna Plomecka
J
John Schultz
Jon Lipovetz
Jon Lipovetz
Software Engineer, Google
C
Clayton Drazner
Yuchen Zhuang
Yuchen Zhuang
Google DeepMind
Reinforcement LearningLarge Language ModelsAgentic Coding
J
Jaimie Hwang
N
Nate Keating
R
Riley Jones
A
Andrew Lee
O
Oran Kelly
Ian Gemp
Ian Gemp
Google DeepMind
machine learningartificial intelligenceoptimizationgame theorydynamical systems
M
Michael Aaron
L
Laurel Prince
Kate Larson
Kate Larson
University of Waterloo
Artificial IntelligenceMultiagent Systems
J
Jeff Moser
H
Harrison Jobe
C
Chad Woodford
S
Siqi Liu
Andrew Wang
Andrew Wang
University of Toronto, Vector Institute
AI Safety
B
Bo Chang