🤖 AI Summary
This study addresses the limited capacity of large language model agents for iterative self-improvement during test-time inference. To overcome this, the authors fine-tune a Qwen3.8-27B model on synthetically generated long-horizon reflection trajectories in machine learning and algorithmic programming, thereby endowing the agent with deep reflective reasoning and persistent execution capabilities. Furthermore, they propose a domain-agnostic hypothesis that leverages verifiable feedback for data synthesis to facilitate cross-domain transferability. The resulting approach achieves state-of-the-art performance on benchmarks including MLE-bench and Frontier-CS. Notably, performance scales consistently as the iterative computation budget increases, empirically validating the effectiveness of long-horizon reflection for test-time scaling in autonomous agents.
📝 Abstract
We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test time. This ability rests on two complementary capabilities: reflection, which produces a solution better than the current one, and long-horizon execution, which keeps the iteration effective over many rounds. We hypothesize that both capabilities are domain-agnostic, and can therefore be learned in scenarios that are well suited for supervision. Accordingly, we synthesize long-horizon improvement trajectories from machine learning and algorithmic programming tasks, two domains that offer verifiable feedback and reward sustained iteration. Trained on this data, our agent, built on Qwen3.8-27B, achieves strong results on MLE-bench Lite (81.8) and Frontier-CS (70.7), transfers to deep research with 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA, and keeps improving as its budget of rounds grows. These results show that long-horizon reflective data is an effective route toward self-improving agents.