OptiArena: Can LLMs Improve Executable Algorithms under Fixed Resource Budgets?

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of evaluating large language models (LLMs) in optimizing executable game algorithms under fixed resource budgets. To this end, it constructs a controlled testbed incorporating computational budget constraints, surface confounding elimination, and calibrated reference mechanisms. Specifically, the experimental protocol employs five rounds of code-editing iterations, a fixed minimal scaffolding setup, and local evaluator feedback to systematically examine LLMs’ algorithmic refinement capabilities and robustness. Experimental results demonstrate that LLMs exhibit significantly greater consistency when improving weak baselines compared to optimizing strong ones. These findings validate the effectiveness and practical utility of the proposed evaluation framework for measuring the algorithmic optimization capacity of LLMs under bounded resource conditions.
📝 Abstract
Static QA and code-generation benchmarks only partially capture the role that large language models (LLMs) now play as coding agents and research tools. We introduce OptiArena, a budget-controlled testbed for studying whether LLMs can improve executable game-playing algorithms through five rounds of code edits within a fixed minimal scaffold and under bounded evaluator feedback and fixed resource budgets. The testbed uses two optimization regimes, surface obfuscation controls, calibrated references, held-out/stress splits, and diagnostics for degradation and exceptional failures, with LLM API cost reported separately from local evaluator wall-clock. The empirical study asks three questions: whether models can close the calibrated gap between a designated weak starter and an editable competent baseline, whether they can refine editable competent baselines without damaging them, and whether gains survive surface obfuscation controls. Across twelve frontier LLMs and five games, models improve designated weak starters more consistently than they refine editable competent baselines, with substantial variation across games and models. OptiArena provides a practical testbed for measuring bounded-resource algorithm optimization within the five-edit, fixed-scaffold setting studied here. Code is available at https://github.com/WJ-Peng/OptiArena.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Algorithm Optimization
Resource Budgets
Code Editing
Benchmark Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Budget-controlled testbed
Algorithm optimization
Code editing
Surface obfuscation controls
Large language models
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Wenjun Peng
Wenjun Peng
University of Science and Technology of China
NLP
X
Xinyu Wang
Adelaide University, Australia