Can AI Agents Learn Their Way to the Top? Evaluating Heuristic Learning in a Long-Running Game Agent Competition

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of policy learning under limited samples in adversarial games by proposing an Adversarial Heuristic Learning paradigm. Rather than fine-tuning large language model weights, this approach iteratively optimizes external executable policy code through rule parsing, replay analysis, and opponent selection. To support this investigation, we construct AAArena, a real-world game benchmark comprising 1,920 programs, and design an evaluation protocol that simulates authentic competitive settings. Experimental results demonstrate that the Opus5.5 configuration secures six gold medals, validating the significant effectiveness of learning from both self-play and opponent replays. Furthermore, our findings reveal that comprehending complex rules and developing long-horizon strategies remain substantial challenges. This work highlights the potential of heuristic-driven, weight-frozen LLM frameworks for sample-efficient strategic reasoning in competitive environments.
📝 Abstract
Adversarial games have driven advances from heuristic search to reinforcement learning, yet learning and adapting strategies from limited samples remain challenging. AI agents offer an alternative by turning game experience into revisions of executable policies. Building on heuristic learning (HL), we formalize Adversarial Heuristic Learning (AHL), a paradigm that uses AI agents as learning engines to refine game policies and supporting software while keeping model weights fixed. We introduce AAArena, a benchmark comprising 12 authentic adversarial games and 1,920 archived human programs, with an evaluation protocol modeled on real-world game competitions. Agents interpret rules, choose opponents, analyze replays, and revise game agents to achieve their highest ranking within fixed match and evaluation budgets. We evaluate \val{completedmodels} model and harness configurations: Opus5.5 with Claude Code earns 6 gold medals, while no evaluated configuration tops the remaining 6 human ladders. Performance is generally weaker in games with more complex rule specifications. Further experiments show that opponent selection and dense feedback support policy improvement, and that agents learn from both on-policy replays of their own matches and off-policy replays of other players' matches. These results highlight HL's potential in adversarial games and identify persistent challenges in game understanding, strategy implementation, and long-horizon policy development.
Problem

Research questions and friction points this paper is trying to address.

Adversarial games
Heuristic learning
AI agents
Policy refinement
Long-horizon planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adversarial Heuristic Learning
AI Agents
Benchmark
Heuristic Learning
Policy Revision
🔎 Similar Papers
No similar papers found.
K
Kaisen Yang
Department of Computer Science and Technology, Tsinghua University
Q
Qingle Liu
Department of Computer Science and Technology, Tsinghua University
Kejin Wang
Kejin Wang
Department of Computer Science and Technology, Tsinghua University
Y
Yicheng Zhao
Department of Computer Science and Technology, Tsinghua University
J
Jieming Li
Department of Computer Science and Technology, Tsinghua University
S
Shenghan Zheng
Department of Computer Science and Technology, Tsinghua University
R
Ruize Yang
Department of Computer Science and Technology, Tsinghua University
B
Bojun Yang
Department of Computer Science and Technology, Tsinghua University
Heng Gong
Heng Gong
Department of Computer Science and Technology, Tsinghua University
X
Xiang Gao
Department of Computer Science and Technology, Tsinghua University
L
Lanyue Zhang
Department of Computer Science and Technology, Tsinghua University
K
Kaiyu Zhong
Department of Computer Science and Technology, Tsinghua University
Z
Zhuo Liu
Department of Computer Science and Technology, Tsinghua University
S
Shaoxuan Li
Department of Computer Science and Technology, Tsinghua University
Chengxi Li
Chengxi Li
Department of Computer Science and Technology, Tsinghua University
Yong Yan
Yong Yan
Department of Computer Science and Technology, Tsinghua University
W
Weixuan Zhang
Department of Computer Science and Technology, Tsinghua University
T
Tianwei Luo
Department of Computer Science and Technology, Tsinghua University
S
Situ Wang
Department of Computer Science and Technology, Tsinghua University
Y
Youjie Zheng
Department of Computer Science and Technology, Tsinghua University
Sihan Zhao
Sihan Zhao
Tsinghua University
Shengyuan Wang
Shengyuan Wang
Tsinghua University
Huan-ang Gao
Huan-ang Gao
Ph.D. student, Tsinghua University
AgentVision & Robotics
Jiazheng Xu
Jiazheng Xu
Tsinghua University
Vision Language ModelsAlignmentGenerative ModelsMachine Learning
Xiaohui Xie
Xiaohui Xie
Professor of Computer Science, University of California, Irvine
AIMachine LearningGenomicsNeural Computation