Score
Design and evaluate representation-learning methods and transformations that split a model's latent space into distinct, minimally entangled components so each latent factor corresponds to an independent generative or control parameter. Build analyses and algorithms to isolate control-parameter-related features, prevent spurious feature entanglement, and enable stable extrapolation and manipulation by enforcing factor independence and interpretability.
This work addresses the identifiability of disentangled representations without relying on strong distributional assumptions—such as statistical independence—thereby establishing necessary and sufficient conditions for unique recovery of latent factors under nonlinear, non-invertible mixing. Method: We introduce “mechanism independence” as a novel paradigm, modeling latent variables via their generative mechanisms rather than latent distributions. We develop a spectrum of identifiability criteria grounded in support-set structure, sparsity, and higher-order conditional independence, and employ graph-theoretic analysis to characterize connected components of latent subspaces. Contribution/Results: We prove that each mechanism independence condition guarantees uniqueness of the latent subspace. Our framework transcends classical assumptions in causal discovery and representation learning—namely linearity, invertibility, or statistical independence—and provides the first rigorous, assumption-free theoretical foundation for unsupervised disentanglement.
This work investigates whether sparse autoencoders (SAEs) and sparse linear probes can reliably disentangle and localize causally relevant semantic concepts—such as sentiment, domain, or tense—when concepts exhibit controlled inter-concept correlations. Method: We introduce the first evaluation framework that explicitly manipulates multi-concept correlations, integrating subspace projection analysis, feature steering interventions, and quantitative disentanglement metrics. Contribution/Results: We find that (1) the mapping from concepts to features is many-to-one, rendering conventional correlation-based disentanglement metrics insufficient for guaranteeing steering independence; (2) while individual features lack concept selectivity, their causal effects are confined to orthogonal subspaces; and (3) reliable interpretability assessment requires combinatorial, intervention-driven evaluation rather than isolated metrics. Our framework establishes a new paradigm and empirical benchmark for rigorously validating the reliability of interpretability methods in language models.
This work identifies a fundamental limitation of KL-divergence-based prior regularization in variational autoencoders (VAEs): it fails to reliably enforce posterior aggregation toward a factorized Gaussian prior, resulting in entangled latent representations. To address this, we propose a programmable prior framework grounded in the Maximum Mean Discrepancy (MMD), enabling flexible modeling of complex, semantically aligned priors within the VAE architecture. Furthermore, we introduce an unsupervised Latent Predictability Score (LPS) to quantitatively assess disentanglement. Experiments on CIFAR-10 and Tiny ImageNet demonstrate that our method achieves state-of-the-art mutual information-based disentanglement performance while preserving high-fidelity reconstructions—thereby resolving the inherent reconstruction–disentanglement trade-off prevalent in conventional disentanglement approaches.
This paper addresses unsupervised representation learning for sequential data. We propose a novel probabilistic flow decomposition framework that disentangles the latent-space dynamics into two orthogonal vector fields: a sparse curl-free field (corresponding to an irrotational potential field) and a divergence-free field (corresponding to a solenoidal rotational field), with sparsity priors newly imposed on both components. Within a variational autoencoder framework, our method jointly optimizes representation encoding, velocity field estimation, and field-structure inference, implicitly learning approximately equivariant representations. Compared to prior approaches, our model simultaneously achieves static representation disentanglement and independence of dynamic transformation primitives, yielding significant improvements in data likelihood and unsupervised equivariance error across multiple sequence transformation benchmarks—achieving state-of-the-art performance. Crucially, the learned vector fields admit clear physical interpretations grounded in classical vector calculus.
Existing disentanglement definitions and metrics assume mutual independence among latent factors, failing to capture inherent statistical dependencies among real-world factors—leading to poor generalization in practical scenarios. Method: We propose the first information-theoretic, generalized disentanglement definition that explicitly accommodates non-independent factors and establish its theoretical connection to the information bottleneck principle. Building upon this, we design the first computable, robust disentanglement metric for non-independent factors—the Generalized Disentanglement Score (G-Disentanglement Score)—integrating mutual information, conditional mutual information, and statistical dependence modeling. Results: Evaluated on controlled synthetic experiments and a unified benchmark, our metric consistently outperforms existing measures across multiple non-independent factor settings, achieving an average improvement of 23.6%. It exhibits strong theoretical grounding and empirical consistency, providing a principled, generalizable evaluation standard for representation learning in realistic settings.
This study addresses the inherent non-identifiability of latent variables in factor models—manifested as non-uniqueness and distributional shifts—by systematically elucidating their nature in linear factor models and their implications for representation learning, drawing on an interdisciplinary perspective spanning psychometrics, statistics, and artificial intelligence. It establishes a theoretical connection between this identifiability issue and posterior collapse in variational autoencoders. By integrating factor analysis, linear autoencoders, and variational inference within a high-dimensional asymptotic framework, the work proves that latent factors become fully identifiable as the observation dimension tends to infinity. Building on this result, the authors propose a nearly distribution-free estimation method for high-dimensional settings, effectively bridging the theoretical gap between classical factor analysis and modern deep generative models, particularly well-suited for representation learning with ultra-high-dimensional data.
This study addresses the ambiguous definition and limited interpretability of discrete latent dimensions in Quantum Variational Autoencoders (QVAEs). To overcome the high-dimensional complexity of Hilbert space, this work proposes a mechanism that designates individual qubits as independent semantic factors. By introducing a quantum regularization strategy and combining theoretical analysis with experiments on synthetic datasets such as MNIST, we systematically investigate the composition of quantum latent dimensions and their factor disentanglement capabilities. Our findings demonstrate that QVAEs can effectively discover interpretable, factorized representations, thereby establishing both theoretical and empirical foundations for structured quantum representation learning.
This work proposes a novel approach to unsupervised disentangled representation learning that circumvents the need for statistical independence or causal assumptions traditionally required for identifiability. By imposing a local orthogonality constraint on the Jacobian of the generative mapping, the authors introduce “functional orthogonality” as a disentanglement principle and theoretically establish the identifiability of nonlinear generative models under this condition. The method is implemented using normalizing flows trained with an orthogonality regularizer. Experiments demonstrate successful recovery of the true underlying latent factor structure, thereby validating the theoretical claims and offering new insights into the mechanisms enabling disentanglement in variational autoencoders. These findings challenge the prevailing view that unsupervised disentanglement is fundamentally unattainable.
This work addresses the challenge of disentangling underlying factors of variation in unsupervised representation learning by introducing Holographic Reduced Representations (HRR) for the first time into this domain. The proposed method models latent variables as vector superpositions of symbol–value pairs and leverages HRR’s unbinding operation as an inductive bias to encourage approximately independent factorized representations. Theoretical analysis derives an upper bound on the information capacity per slot, offering an information-theoretic interpretation of disentanglement. Empirical results demonstrate that the approach outperforms existing baselines in terms of latent traversability and standard disentanglement metrics, while also exhibiting superior robustness to noise and consistently stable reconstruction performance across varying signal-to-noise ratios.
This study addresses the inherent trade-off between covariate dependence and latent structure in disentangled representation learning by proposing a unified supervised framework that elucidates the constraints linking latent independence with covariate alignment. By establishing disentanglement orderliness and deriving a closed-form transformation for realignment, combined with informed Factor Analysis (iFA), this work enables precise regulation of structured representations in pretrained models. Extensive experiments on both simulated and real-world multi-omics datasets validate the method’s effectiveness, demonstrating significant improvements in the controllability and interpretability of learned representations. Ultimately, this research establishes a novel paradigm for disentangling complex data structures, offering a theoretically grounded solution to balance statistical independence with semantic alignment in high-dimensional representation learning.