Lucid Dreaming for World Models: Learning to Doubt Imagination and Decide by Trust
This study addresses the unreliability of imagined planning in world models, where overconfidence in unseen state-action pairs misleads decision-making. To mitigate this, we propose LucidWM, a framework that introduces subjective logic to assign degrees of doubt to categorical latent transitions, enabling uncertainty estimation without additional parameters. By accumulating trust over multi-step trajectories to reweight returns, the method integrates doubt into the imagination process, thereby guiding reinforcement learning policy optimization. The effectiveness of this approach is validated against four baseline models and seventeen uncertainty readouts. In navigation tasks, LucidWM reduces the number of steps required for goal attainment from 362 to 190, significantly enhancing the robustness of trust-based decision-making.