🤖 AI Summary
This study addresses the challenge of learning invariant and variant factor structures within latent states without relying on observation reconstruction. To this end, it proposes SplitJEPA, a framework built upon joint-embedding predictive architectures and Gaussian predictive dynamics. This approach jointly recovers latent states along with their invariant and variant organization directly in the representation space, entirely eliminating the need for an observation decoder. The primary contribution lies in achieving, for the first time, the identification of invariant and variant subspaces under reconstruction-free conditions, supported by theoretical guarantees of block-isomorphic structure recovery. Furthermore, experiments conducted on both synthetic systems and robotic manipulation tasks validate the robustness and efficiency advantages of the proposed method.
📝 Abstract
Understanding a dynamical world calls for more than a latent state that summarizes its observations: the state should also be organized into the factors that stay shared across related observations and the factors that vary between them. For example, a robot pushing a cube to a goal should take the same action when the camera shifts or the lights dim, since nothing in the scene has moved. Existing approaches to this decomposition commonly obtain it through reconstruction, so the latent variables must first explain the entire observational world before their organization can be trusted. Joint embedding predictive architectures (JEPAs) model the latent state directly and never reconstruct, yet no existing result recovers the invariant and variant parts of the state they learn. How to learn the invariant-variant structure of the latent world without paying for its reconstruction therefore remains open. To close this gap, we introduce SplitJEPA, a JEPA that jointly recovers the latent state and its invariant and variant organization directly in representation space, without any reconstruction. We prove that, under stationary Gaussian predictive dynamics and a full-rank variation condition, SplitJEPA identifies the invariant and variant subspaces up to independent block-wise isometries, without introducing an observation decoder. Since the guarantee needs no decoder, the result extends reconstruction-free latent recovery to invariant-variant block identification. Experiments on synthetic nonlinear systems and robotic manipulation tasks support the theoretical results and show their practical value for both robustness and efficiency.