🤖 AI Summary
This study addresses the limitations of static benchmarks, which often lead to performance saturation in large language models (LLMs) and fail to adequately evaluate dynamic strategic capabilities. To overcome this, we construct an open competitive platform based on chess, poker, and Werewolf. Methodologically, we replace static test sets with dynamically evolving game environments spanning perfect-information, imperfect-information, and multi-player scenarios to surpass existing performance ceilings. Supported by robust infrastructure and large-scale ground-truth evaluation techniques, our platform enables reproducible, transparent, and scalable automated head-to-head adversarial benchmarking. This work effectively validates LLMs' proficiency in strategic planning, environmental adaptation, and robustness under uncertainty, establishing a novel paradigm for evaluating the dynamic capabilities of foundation models.
📝 Abstract
We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where the gameplay strength naturally increases as models evolve, preventing performance saturation. This technical report details the infrastructure behind Game Arena and describes the three pilot game environments: Chess, Poker, and Werewolf. These environments span perfect information, imperfect information, and multiplayer game settings, enabling a systematic study of models' strategic planning, adaptation, and robustness under uncertainty. For each game, we provide a detailed description of the environment, evaluation metrics, and results from running full competitions across models. Through robust infrastructure and large-scale ground-truth based evaluation, Game Arena ensures reproducibility, transparency and generalizability to new games and variants over time.