🤖 AI Summary
This work systematically identifies three distinct modes of representation collapse in JEPA-based world models—physical invariance, identifiability, and counterfactual dynamics—even when global latent collapse is avoided. To address these failure modes, the paper introduces PhyLatent, a novel training objective that jointly optimizes dynamics-relevant representations through physical state anchoring, future representation alignment, static visual invariance constraints, counterfactual branch disentanglement, and latent denoising. Moving beyond reliance on global non-collapse assumptions alone, PhyLatent significantly reduces the three collapse rates to 7.53%, 0.95%, and 4.62% on OGBench-Cube, yielding a model-predictive control (MPC) success rate of 78.1%. It further achieves a 98.0% success rate on the TwoRooms task and maintains state-of-the-art performance on Reacher and PushT benchmarks.
📝 Abstract
We propose PhyLatent, a dynamics-relevant training objective for JointEmbedding Predictive Architecture (JEPA) world models. Our key observation is that preventing global latent collapse does not ensure that a representation preserves physical states and action consequences. We identify three failure modes in JEPA world models: physical invariance collapse, physical identifiability collapse, and counterfactual dynamics collapse. PhyLatent addresses them through three training pathways: physical invariance, physical identifiability, and counterfactual dynamics, implemented with physical state grounding, future representation alignment, static visual invariance, counterfactual branch separation, and latent denoising. On OGBench-Cube, PhyLatent reduces the three failure rates from 15.60%, 6.71%, and 8.41% to 7.53%, 0.95%, and 4.62%, respectively, and improves model predictive control (MPC) success from 70.0% to 78.1%. With the same architecture and planner, it further improves success from 81.0% to 98.0% on TwoRooms and remains competitive on Reacher and PushT. These results show that global non-collapse alone is insufficient for learning a reliable JEPA worldmodel state space.