🤖 AI Summary
This study addresses the prediction-planning mismatch in latent world models, where low prediction errors nonetheless yield unreliable planning due to misalignment between actions and their consequences. Inspired by the neuroscience principle of self-tickling, this work proposes an action-consequence alignment training objective that corrects model biases by penalizing the advantages of local alternatives. Furthermore, a principled guided data collection strategy is designed to enable self-improvement, bridging predictive learning and reliable planning without requiring additional components or interactions. The proposed approach significantly reduces true goal errors and enhances planning performance across diverse environments, outperforming random sampling baselines. These findings establish a new paradigm for constructing highly reliable world models.
📝 Abstract
Latent world models learn to predict observed transitions, yet low prediction error alone does not guarantee reliable planning. Inspired by self tickling experiments in neuroscience showing that disrupting motor sensory correspondence increases prediction mismatch, we examine whether learned world models preserve an analogous action consequence correspondence.The results show nearby alternatives can receive lower prediction errors despite producing physical outcomes farther from the recorded target. With that future treated as a goal, this reveals a concrete prediction planning mismatch: the model assigns a lower cost to an action that achieves the target less accurately. To mitigate this gap, we introduce Action Consequence Alignment (ACA), a training objective that complements forward prediction by penalizing the prediction error advantage of locally searched alternatives over factual actions without additional model components or environment interactions during training. The same principle can also guide additional data collection for self improvement. We demonstrate that across diverse environments and evaluation settings, ACA improves planning performance and reduces real goal error, while ACA guided data collection outperforms random local sampling. These results support action consequence alignment as a practical principle for bridging predictive learning and reliable planning.