🤖 AI Summary
This study addresses the theoretical challenge in self-supervised learning (SSL) of distinguishing prediction-relevant stochastic signals from irrelevant nuisance factors. We demonstrate that predictive SSL implicitly constructs latent variable models incorporating stochastic dynamics, achieving signal identification through latent-space prediction, mutual information maximization, and latent distribution matching. This work provides the first rigorous theoretical proof that mainstream SSL methods can effectively disentangle stochastic signals from nuisances, elucidating their underlying implicit modeling mechanisms. Furthermore, experiments utilizing Gaussian predictors empirically validate the proposed framework, confirming its effectiveness in recovering affine-transformed representations of true signals under complex nuisance conditions.
📝 Abstract
Self-supervised learning (SSL) by predicting in latent space, without generating the input data itself, learns highly abstract, useful representations. Intuitively, this success is often attributed to its ability to discard nuisance information that is irrelevant to prediction. However, this poses a conundrum: both stochastic variation in a prediction-relevant latent signal and true nuisance make observations partly unpredictable; how could they be distinguished? Surprisingly, we prove that common SSL methods can achieve exactly this, by implicitly instantiating a latent-variable model with stochastic dynamics and observation-private nuisance. We trace their ability to recover the stochastic signal to two complementary principles: Predictive mutual information maximization ensures that representations retain the information needed for prediction, while latent distribution matching constrains how this information is encoded, thereby making the retained signal identifiable. We confirm this identifiability result in simulations for Gaussian predictors, which recover the true signal up to an affine transformation even in dynamic, nuisance-laden environments.