๐ค AI Summary
This work addresses the challenge of training large language model (LLM) agents in long-horizon games, where conventional reinforcement learning struggles due to reliance on sparse terminal rewards. The authors propose CAST, a novel method that leverages state-value trajectories from game solvers to generate episode-level credit assignment signals, which serve as dense supervision for reinforcement learning. CAST is theoretically equivalent to policy distillation under a soft-optimal solver and requires only scalar state valuesโwithout needing access to a full teacher policy. Empirical results demonstrate that CAST significantly outperforms existing baselines on Sokoban, Minesweeper, and Rush Hour, while achieving state-of-the-art zero-shot average performance on ALFWorld and WebShop.
๐ Abstract
Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR) relies on sparse final rewards that reveal little about which decisions determine success. Denser process signals could supply this missing turn-level credit, but existing sources are hard to keep both cheap and accurate. We observe that changes in a game solver's state value reveal whether an action advances the state toward success. Building on this insight, we propose CAST (Credit Assignment from Solver Teachers), which converts these value changes into solver advantages and injects them into RLVR as turn-level signals. We further show that, under a soft-optimal solver assumption, maximizing the solver advantage is equivalent to on-policy distillation from the solver, requiring only scalar values rather than teacher logits. Across Sokoban, Minesweeper, and Rush Hour, CAST outperforms all trained baselines on every game under both in-domain and unseen-difficulty evaluation and achieves the highest average zero-shot performance on ALFWorld and WebShop. Our code is available at https://github.com/Wloner0809/CAST.