π€ AI Summary
This study investigates whether large language model (LLM) agents can autonomously develop complex strategies that transcend established human solutions. Using game speedrunning as a testbed, we explore agentsβ capacity for reflection and strategy optimization in long-horizon tasks. To this end, we propose an anti-saturation evaluation benchmark that continuously challenges the reasoning limits of models through dynamically refreshed optimal records, combined with a multi-round self-iterative optimization algorithm driven by game mechanics analysis to guide frontier LLM agents. Experimental results demonstrate that our approach enables agents to approximate human records in simple games, validating the feasibility of autonomous strategy generation. However, these agents still significantly underperform top-tier human players in complex, long-horizon tasks, thereby identifying critical bottlenecks for future research in this domain.
π Abstract
Frontier LLM agents have been shown to be capable of solving increasingly complex tasks for which humans have measurable solutions. This begs the pertinent question of whether LLM agents can go beyond what humans have already solved. The ability to develop sophisticated strategies to tackle consequential problems becomes paramount as well-trodden, human-developed solutions become insufficient for problems for which we lack context or enough training data. We study agents' capability of such strategy formation through the communal practice of video game speedrunning. In speedrunning, practitioners compete to find the fastest way to complete a video game under certain conditions, and in so doing uncovering interesting unorthodox play styles that require a thorough understanding and mastery of the underlying game mechanics. We introduce SPEEDRUNBENCH, a benchmark that evaluates frontier LLM agents across 9 different games. To perform well in this benchmark, agents must repeatedly improve their strategy, reflect on their performance, exploit their gained knowledge, and reason across a long-horizon of actions to improve on an increasingly difficult problem: being faster than themselves and everyone else. Our experiments show that while frontier agents approach human world records in simple platformer games, they remain behind human performance on longer, more complex games under practical budgets. These results suggest that SPEEDRUNBENCH is a useful testbed for studying agents' strategy formation capabilities as well as being a saturation-resistant evaluation measure, as there is almost always a faster completion time waiting to be discovered.