State Trace Rationale As Auxiliary Task in Reinforcement Learning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of convergence difficulty and state representation rank collapse in deep reinforcement learning under sparse reward environments. We propose STRAT, an auxiliary task inspired by human navigation that leverages environment rules to automatically generate textual state trajectories containing landmark and path information. This approach enables interpretable agent belief tracking without manual annotation. Through a multi-task learning architecture, a single auxiliary head is introduced into the policy network for text prediction, thereby enhancing spatial cognition. Experimental results demonstrate that STRAT significantly outperforms existing baselines on XLand-MiniGrid tasks, effectively overcoming learning difficulties in complex sparse-reward scenarios while achieving state representation compression.
📝 Abstract
We propose STRAT, an auxiliary task that trains deep reinforcement learning (RL) agents to predict a short textual trace of their own state. Inspired by human spatial navigation, the description combines landmark, route, and survey knowledge, tracking the agent's position, inventory, goals, and immediate progress. Environment rules generate this text online without human labelling. Our method adds a single auxiliary head to a standard policy. Across 60 sparse-reward XLand-MiniGrid tasks, STRAT solves complex environments where standard RL fails outright, while compacting state representations and preventing rank collapse. Beyond performance gains, the predicted trace provides a readable account of agent beliefs at every step for no extra cost.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Sparse Reward
Rank Collapse
State Representation
XLand-MiniGrid
Innovation

Methods, ideas, or system contributions that make the work stand out.

Auxiliary Task
Reinforcement Learning
State Trace Rationale
Sparse Reward
Rank Collapse
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Muhammad U. Nasir
University of the Witwatersrand, Johannesburg, South Africa
A
Alex Vogt
University of the Witwatersrand, Johannesburg, South Africa
S
Steven D. James
University of the Witwatersrand, Johannesburg, South Africa
Julian Togelius
Julian Togelius
Associate Professor of Computer Science and Engineering, New York University; co-founder, modl.ai
Artificial IntelligenceGamesEvolutionary ComputationGame AIProcedural Content Generation