WorldGuide: Learning Success-Failure Boundaries in Latent World Models for Vision-Language-Action Policies

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottleneck of Vision-Language-Action (VLA) policies in distinguishing between successful and failed interactions by proposing an optimization framework based on success-failure boundary learning in latent space. The method leverages contrastive learning over success-failure sample pairs to construct differentiable reward signals within a latent world model, guiding the joint optimization of the policy without introducing additional inference overhead during deployment. Experimental results demonstrate that this framework substantially enhances the reliability and generalization capability of VLA policies, achieving success rates of 96.8% and 72.0% on the LIBERO-100 and SimplerEnv benchmarks, respectively.
📝 Abstract
Latent world models offer a promising way to improve Vision-Language-Action policies by capturing the consequences of actions. However, models trained primarily on expert demonstrations have limited exposure to failure outcomes and may struggle to distinguish visually similar successful and failed interactions. We propose \textbf{WorldGuide}, a framework that learns these distinctions in latent space and uses them to guide policy training. WorldGuide combines predictive pretraining on successful and failed trajectories with contrastive learning on matched success--failure pairs. The learned predictor then provides a differentiable reward to guide joint optimization of the policy and visual encoder. The predictor is discarded after training, so deployment requires no additional world-model inference. Extensive experiments show that WorldGuide substantially improves VLA reliability and achieves state of the art performance on LIBERO 100 and SimplerEnv, reaching \textbf{96.8\%} and \textbf{72.0\%}, respectively. Code will be publicly available.
Problem

Research questions and friction points this paper is trying to address.

Latent World Models
Vision-Language-Action Policies
Success-Failure Boundaries
Expert Demonstrations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent World Models
Vision-Language-Action Policies
Contrastive Learning
Differentiable Reward
Success-Failure Boundaries
🔎 Similar Papers
No similar papers found.