๐ค AI Summary
This work addresses the tendency of deterministic models in self-supervised representation learning to collapse to conditional means and exhibit asymmetric dependency structures under multimodal inverse problems. To overcome these limitations, the authors propose a probabilistic approach based on generative joint modeling. By explicitly modeling the joint density of context and target variables through Gaussian Joint Embedding (GJE) and its multimodal extension GMJE, the method replaces black-box prediction with closed-form conditional inference. A covariance-aware objective is introduced to shape the geometry of the latent space, revealing and avoiding the โMahalanobis trace trapโ in batch optimization, while unifying contrastive learning as a degenerate nonparametric limit of GMJE. Integrating Gaussian mixture embeddings, prototypical mechanisms, conditional mixture density networks, and topology-adaptive memory, the approach successfully recovers complex multimodal structures on synthetic and visual benchmarks, yielding representations that are highly discriminative and exhibit superior unconditional generative performance compared to existing baselines.
๐ Abstract
Self-supervised representation learning often relies on deterministic predictive architectures to align context and target views in latent space. While effective in many settings, such methods are limited in genuinely multi-modal inverse problems, where squared-loss prediction collapses towards conditional averages, and they frequently depend on architectural asymmetries to prevent representation collapse. In this work, we propose a probabilistic alternative based on generative joint modeling. We introduce Gaussian Joint Embeddings (GJE) and its multi-modal extension, Gaussian Mixture Joint Embeddings (GMJE), which model the joint density of context and target representations and replace black-box prediction with closed-form conditional inference under an explicit probabilistic model. This yields principled uncertainty estimates and a covariance-aware objective for controlling latent geometry. We further identify a failure mode of naive empirical batch optimization, which we term the Mahalanobis Trace Trap, and develop several remedies spanning parametric, adaptive, and non-parametric settings, including prototype-based GMJE, conditional Mixture Density Networks (GMJE-MDN), topology-adaptive Growing Neural Gas (GMJE-GNG), and a Sequential Monte Carlo (SMC) memory bank. In addition, we show that standard contrastive learning can be interpreted as a degenerate non-parametric limiting case of the GMJE framework. Experiments on synthetic multi-modal alignment tasks and vision benchmarks show that GMJE recovers complex conditional structure, learns competitive discriminative representations, and defines latent densities that are better suited to unconditional sampling than deterministic or unimodal baselines.