🤖 AI Summary
This study addresses the limitation of latent-space world models in effectively distinguishing successful from failed action sequences due to encoding distortions of critical physical quantities. To overcome this, we propose an auxiliary loss function that leverages success criteria as a training objective. Unlike prior approaches that merely incorporate such criteria as model inputs, this work pioneers their use as direct supervisory signals. Specifically, a linear regression head enforces the encoder and predictor to retain essential information during training, while being removed at test time to ensure zero inference overhead. This approach precisely constrains the representational capacity of latent states. Experimental results demonstrate significant improvements, increasing success rates by 3.5% and 3.4% on the PushT and Cube manipulation tasks, respectively, thereby validating the effectiveness of the proposed method.
📝 Abstract
Latent world models plan by scoring candidate action sequences with distances in latent space. However, task success is judged by physical quantities, which we call the success-criterion quantities. In all four latent world models we examine, the end-effector position is encoded in the latent state with an error larger than the success criterion allows. Such a latent state cannot separate successful candidates from failing ones. We propose an auxiliary loss that uses success-criterion quantities as training targets, whereas existing latent world models take them only as inputs. During training, a linear head on the encoder and predictor outputs regresses the success-criterion quantities, and the regression error is added to the training loss. The head is discarded after training, so the model, its cost, and its inputs at test time are unchanged. This loss alone improves the success rate on PushT and cube by 3.5% and 3.4% (absolute), respectively, and both improvements are statistically significant. A success criterion thus specifies what a world model must retain in its latent state, and we show that it can serve directly as a training target.