🤖 AI Summary
This study addresses the challenges of convergence difficulty and state representation rank collapse in deep reinforcement learning under sparse reward environments. We propose STRAT, an auxiliary task inspired by human navigation that leverages environment rules to automatically generate textual state trajectories containing landmark and path information. This approach enables interpretable agent belief tracking without manual annotation. Through a multi-task learning architecture, a single auxiliary head is introduced into the policy network for text prediction, thereby enhancing spatial cognition. Experimental results demonstrate that STRAT significantly outperforms existing baselines on XLand-MiniGrid tasks, effectively overcoming learning difficulties in complex sparse-reward scenarios while achieving state representation compression.
📝 Abstract
We propose STRAT, an auxiliary task that trains deep reinforcement learning (RL) agents to predict a short textual trace of their own state. Inspired by human spatial navigation, the description combines landmark, route, and survey knowledge, tracking the agent's position, inventory, goals, and immediate progress. Environment rules generate this text online without human labelling. Our method adds a single auxiliary head to a standard policy. Across 60 sparse-reward XLand-MiniGrid tasks, STRAT solves complex environments where standard RL fails outright, while compacting state representations and preventing rank collapse. Beyond performance gains, the predicted trace provides a readable account of agent beliefs at every step for no extra cost.