Preserving Unstable Modes Through Inverse Dynamics in JEPA World Models

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the tendency of Joint-Embedding Predictive Architecture (JEPA) world models to compress unstable modes, which compromises system controllability. To mitigate this, the authors introduce an inverse dynamics loss that preserves control-relevant features and ensures encoder injectivity over the reachable subspace. By integrating JEPA with linear systems theory, they theoretically demonstrate that action reconstruction prevents the loss of unstable modes and reveal a convergence relationship between the controllability Gramian and the unstable subspace. The proposed approach is validated on nonlinear visual control simulation tasks such as CartPole, where it effectively retains unstable modes and significantly improves downstream control performance. Overall, this work provides a rigorous theoretical foundation for the synergy between representation learning and control.
📝 Abstract
Robotic systems often exhibit unstable modes, along which small perturbations and disturbances can cause unbounded growth unless corrected through feedback. Controlling such systems from high-dimensional visual observations requires representations that preserve these modes. Joint-embedding predictive architectures (JEPAs) provide a natural framework for learning such representations and their dynamics from visual data. However, we demonstrate that next step prediction combined with anti-collapse regularization does not guarantee that controllable unstable modes are preserved: the training loss can be minimized while these modes are collapsed, making stabilization from the learned representation impossible. To address this, we augment world-model training with an action reconstruction objective (i.e., an inverse dynamics loss) that encourages control-aware representations, namely, visual representations that preserve crucial features for control. We prove that exact action reconstruction makes the encoder injective on the finite-horizon reachable subspace. Thus, the encoder cannot discard any state direction reachable by an action sequence within $H$ steps. Moreover, we show that, as $H$ grows, the dominant eigenspace of the finite-horizon controllability Gramian converges to the controllable unstable subspace. We establish our theoretical results for linear systems and demonstrate empirically that our findings extend to nonlinear visual control tasks (CartPole, Walker2D, and PointMaze), highlighting the benefits of control-aware representation learning.
Problem

Research questions and friction points this paper is trying to address.

JEPA world models
unstable modes
visual representation learning
robotic control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Joint-Embedding Predictive Architectures
Inverse Dynamics
World Models
Control-Aware Representations
Unstable Modes
🔎 Similar Papers
No similar papers found.