🤖 AI Summary
This study addresses the limitation of existing methods such as JEPA, which can only identify linearly mixed latent variables and fail to recover individual world latents. To overcome this, we propose Dependency-Sparse Regularization (DSReg), which leverages a structural diversity condition strictly weaker than those in prior work to achieve latent disentanglement without requiring decoders or labels. Our primary contribution is establishing LeJEPA, the first fully identifiable theoretical framework for JEPA, proving that individual latent variables can be recovered up to permutation without reconstruction. Experiments on synthetic data and visual encoders demonstrate that our approach preserves dense prediction performance while significantly enhancing both latent recovery accuracy and downstream task effectiveness.
📝 Abstract
Methods that recover individual latent variables of the world, from nonlinear ICA to dictionary learning and causal representation learning, anchor the latents to observations through reconstruction, auxiliary supervision, or distributional asymmetries such as non-Gaussianity. Methods without these anchors, including joint-embedding predictive architectures (JEPAs), identify the latent state only up to a linear transformation, so individual latents remain mixed. We close this gap: individual world latents can be provably recovered with no reconstruction, no decoder, and no labels. The key condition is Structural Diversity: different latents leave distinct dependency footprints on observations, just as no two snowflakes are alike. Building on the linear identifiability that LeJEPA provides, we prove that under Structural Diversity, DSReg (Dependency-Sparsity Regularization) recovers individual world latents up to signed permutation, without reconstruction or a decoder. It applies post hoc to any linearly identified representation, reusing trained checkpoints at no loss over joint training, and establishes the first fully identifiable JEPA that recovers every world latent. Moreover, as a condition on dependency footprints, Structural Diversity is strictly weaker than all structural conditions of prior identifiable latent variable models. Across synthetic regimes, world model probes, learned visual encoders, and external renderers, DSReg preserves dense prediction while improving individual-latent recovery and downstream use with scales.