CF-JEPA: Improving Robustness of JEPA World Models via Controllability Factorization
This study addresses the vulnerability of JEPA-based world models to distractor signals in visual control, which often leads to latent space collapse. To this end, we propose CF-JEPA, a framework built upon a JEPA-style latent world model with a pixel-reconstruction-free self-supervised architecture. By introducing a controllability factorization mechanism, CF-JEPA explicitly decomposes the latent space into controllable and uncontrollable subspaces, effectively isolating distracting information while extracting task-relevant features for control. Experimental results demonstrate that CF-JEPA achieves performance comparable to conventional methods under nominal conditions while significantly outperforming them in the presence of distractors. Notably, it is the only approach capable of preventing latent space collapse. Furthermore, its practical utility is validated through simulated robotic tasks.