🤖 AI Summary
This study addresses the challenge in automatic curriculum learning for reinforcement learning, where insufficient local signal information renders the consequences of decisions difficult to predict. To overcome this limitation, the proposed method introduces large language models (LLMs) for the first time to evaluate both the downstream utility and training feasibility of tasks, integrating these assessments with online progress estimation to optimize curriculum selection strategies. Experimental evaluations on the Craftax benchmark demonstrate that this approach significantly accelerates learning efficiency for specific text-conditioned tasks while achieving sustained performance gains across diverse task scenarios. By effectively leveraging LLM-driven evaluations to guide curriculum design, this work establishes a robust paradigm for automatic curriculum learning, mitigating the shortcomings of conventional methods and highlighting the potential of foundation models in sequential decision-making contexts.
📝 Abstract
Automatic curriculum learning can improve the effectiveness of reinforcement learning by selecting the training experiences presented to the agent over time. Predicting the consequences of such decisions can, however, be difficult. We analyze automatic curriculum learning as a sequential decision-making problem, highlighting a gap between the quantities that determine the value of curriculum decisions and the information captured by local learning signals commonly used to guide them. We then investigate whether Large Language Models (LLMs) can exploit richer information about the learning problem to better anticipate the consequences of curriculum decisions. We introduce a method that combines online learning-progress estimates with LLM-informed estimates of (i) the potential downstream benefits of learning on each task and (ii) whether direct training on a task is currently likely to produce progress. We evaluate the approach on a custom benchmark of 256 textual goals in Craftax under different curriculum objectives. We observe the strongest gains when optimizing for individual target tasks. When optimizing across the full task set, the benefits vary across learners with different mechanisms for cross-task transfer, ranging from modest improvements in learning speed to larger gains that persist through the end of training.