Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the overestimation of large language model (LLM) chess capabilities by existing static evaluations, which fail to verify complete plan execution under adversarial conditions. We construct an executable benchmark for Xiangqi endgames utilizing an interactive REPL environment, wherein agents must achieve checkmate through multi-turn gameplay against engine-level defense. Furthermore, we propose three gap metrics—transformation, consistency, and simulation—to elucidate the mechanism underlying the observation that identifying the correct first move does not guarantee victory. Experiments quantify the low win rates and unstable performance of frontier LLMs in genuine adversarial settings, underscoring the necessity of closed-loop outcome verification and reliability assessment.
📝 Abstract
Static evaluations credit a language model for naming the right move, but an agent must carry a plan through to a verified outcome while an opponent responds. We introduce XiangqiBench, an executable benchmark that measures this difference in Chinese chess: starting from 119 tactical endgames with forced mates supported by engine or checks-only search, an LLM agent must deliver checkmate against an engine defender. An interactive REPL interface separates real moves, state queries, and forward simulation, and we record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols. Three signals that look like competence each overstate closed-loop success. (i) The Conversion Gap: models play the stored reference first move in 26.1\% of Sighted trials, yet only 13.9\% of these trials end in a win. (ii) The Consistency Gap: the leading model reaches 38.7\% pass@3 but only 5.9\% pass^3, winning all three trials on 7 of the 46 positions it ever wins. (iii) The Simulation Gap: 32.3\% of accepted simulation calls stop on an illegal move, and in 49.3\% of comparable cases the real defender replies differently from the line the agent simulated; self-authored rollouts check legality but cannot anticipate the opponent. Finding the move is not winning the game: agent evaluations should score closed-loop outcomes and report reliability alongside coverage.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
closed-loop evaluation
Chinese chess (Xiangqi)
benchmark
multi-turn interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Closed-Loop Evaluation
Executable Benchmark
LLM Agents
XiangqiBench
Interactive REPL
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.