🤖 AI Summary
This study investigates whether embodied agents can achieve self-improvement through autonomous exploration without demonstrations or external rewards. To this end, we construct a video game testbed integrating multimodal large language models, curiosity-driven exploration, hierarchical subgoal planning, and multimodal memory mechanisms, requiring agents to infer knowledge solely from their own experience to accomplish navigation, interaction, and level-completion tasks. This work proposes the first unsupervised benchmark for embodied self-improvement, revealing critical limitations in existing world models and skill discovery methods. Experimental results demonstrate that state-of-the-art frontier models achieve task completion rates below 50%, while complex pipelines fail to reach even the initial milestone, thereby establishing a rigorous evaluation standard for this challenging research direction.
📝 Abstract
Powered by expert guidance, agents can operate in interactive environments; however, it is unclear whether they can learn autonomously from their own experience. To evaluate such self-improvement methods, we introduce GameBoyWorlds, a testbed for agentic self-improvement in video games. GameBoyWorlds-Execution evaluates task execution on a collection of 5 distinct game series. Agents are allowed access to dedicated training games but are provided no demonstrations, documentation, or rewards. Agents must ground themselves in the environment through self-directed exploration and by inferring actionable knowledge from their own experience. At test time, agents must complete short-horizon tasks that evaluate their ability to navigate, interact, and engage with game-specific mechanics in unseen games. Out-of-the-box frontier models complete fewer than 50% of the 500 tasks due to failures in multimodal grounding, establishing that self-improvement methods have room to push performance. We demonstrate that contemporary approaches to self-improvement are lacking, with world modelling and autonomous skill discovery failing, and a novel strategy that uses curiosity-based exploration to write guides achieving only partial success. GameBoyWorlds-Playthrough tests end-to-end game completion in two fan-made Pok\'emon games. We show that while frontier models have been pre-exposed to official releases such as Pok\'emon Red, they lack essential information on the games in our testbed. Instead of relying on their parametric knowledge to succeed, agents must learn from their own experience and autonomously improve over the course of the playthrough. We show that a sophisticated agentic pipeline with multimodal memory and hierarchical subgoals fails to reach even the first major milestone in both games, establishing GameBoyWorlds as an ambitious target for self-improving agents.