Score
Designs, builds, and analyzes representation spaces and model components that partition high-dimensional embeddings into distinct factor subspaces and select or allocate those subspaces for different latent factors. This includes constructing disentangled or factorized embeddings, modeling residual stochastic variation within subspaces, and composing or recombining factor embeddings to represent or generate novel combinations.
This work investigates whether semantic factors in neural network representation spaces can be unsupervisedly decomposed into interpretable, orthogonal subspaces. We propose Neighborhood Distance Minimization (NDM), the first fully unsupervised method—without basis alignment assumptions—to learn class-variable-directed, interpretable subspaces, revealing structured, “functional-circuit”-like organization within model internals. Qualitative analysis and quantitative validation against known GPT-2 circuits confirm strong correlations between discovered subspaces and specific semantic variables (e.g., grammatical roles, factual knowledge). Furthermore, we successfully separate contextual representation from knowledge routing in a 2-billion-parameter GPT-2 model, demonstrating the method’s scalability and practical utility. Our approach advances interpretability by enabling decomposition of high-dimensional representations into semantically meaningful, disentangled subspaces without supervision or architectural constraints.
This study investigates how Transformers organize their internal representations during next-token prediction pretraining to reflect the underlying structure of the world. By constructing synthetic sequential data with known latent factors and employing geometric activation analysis, subspace dimensionality estimation, and modeling of contextual embedding distributions, the authors find that Transformers exhibit an inductive bias toward decomposing inputs into orthogonal low-dimensional subspaces. When conditional independence holds among latent factors, the model learns lossless factorized representations. Remarkably, even in the presence of noise or hidden dependencies, such structured representations are prioritized during early training stages. These findings reveal fundamental principles governing representational formation in Transformers and demonstrate their inherent preference for factorized structures that mirror compositional aspects of the environment.
Disentangling reusable semantic factors from complex data and achieving high-quality recomposition under unsupervised conditions remains a key challenge in generative modeling. This work proposes an unsupervised factor decomposition and recombination method based on diffusion models, introducing for the first time a discriminator-guided adversarial mechanism that enhances disentanglement by distinguishing between single-source samples and cross-source recomposed samples. This approach improves both semantic coherence and physical consistency in synthesized outputs. Empirical evaluations demonstrate superior performance across multiple benchmarks: it achieves lower FID scores and higher MIG and MCC metrics on CelebA-HQ, Virtual KITTI, CLEVR, and Falcor3D datasets, and significantly increases state-space exploration coverage in the LIBERO robotic benchmark.
This work addresses the challenge of detecting generative AI–produced content. We propose an unsupervised, interpretable embedding-space analysis method: semantic embeddings of text or images are extracted using pre-trained large language or multimodal models; subsequently, dimensionality reduction (e.g., PCA) uncovers an intrinsic, low-dimensional distributional shift between AI-generated and human-created samples—rendering them highly separable without supervision. This phenomenon is systematically validated for the first time and endowed with human-interpretable semantic meaning (e.g., topic coherence, syntactic redundancy). Experiments across diverse generative models—including ChatGPT, Gemini, and Stable Diffusion—demonstrate that high-accuracy separation is achieved solely from raw embeddings and unsupervised projection, without fine-tuning, labeled data, or model-specific detectors. Our approach thus significantly enhances both generalizability and interpretability of AI-content detection.
This work addresses the joint problem of intrinsic dimension estimation and geometry-invariant embedding learning for nonlinear manifold-structured data. We propose an autoencoder framework incorporating orthogonality constraints on hidden-layer gradients. Methodologically, we establish, for the first time, a theoretical connection between gradient orthogonality in neural network latent spaces and the local tangent space dimension of the underlying manifold; this enables simultaneous intrinsic dimension estimation, learning of invertible embedding mappings, and construction of coordinate-invariant representations under local Lie group actions on low-dimensional submanifolds. Our key contribution lies in unifying gradient orthogonality with differential-geometric structure, thereby extending invariant representation learning to continuous group actions. Experiments on standard benchmarks demonstrate accurate intrinsic dimension estimation, disentangled representations, and robust group-invariant embeddings, validating both theoretical soundness and algorithmic robustness.
Current approaches struggle to extract interpretable low-dimensional representations from sparse or incomplete similarity data, limiting our understanding of representational structures in neural, behavioral, and artificial intelligence systems. This work proposes Similarity Representation Factorization (SRF), a novel method that integrates non-negative matrix factorization with low-dimensional embedding to enable, for the first time, generalizable and interpretable extraction of representational dimensions. SRF effectively recovers task-specific model dimensions, accurately predicts independent behavioral attributes, and substantially enhances both exploratory analysis capabilities and statistical power in hypothesis testing. The method is broadly applicable to heterogeneous, multi-source similarity data, offering a robust framework for uncovering latent structure across diverse domains.
This work addresses the lack of a unified evaluation framework for subspace representations of human-interpretable concepts in neural models. The authors propose a dual-axis framework grounded in “inclusiveness” and “disentanglement” to systematically integrate and compare five subspace estimation methods—including linear subspace modeling, concept probing, and LEACE—through cross-task empirical analyses on both text and speech models such as HuBERT. Their findings reveal the non-uniqueness of concept subspaces and demonstrate that the choice of estimator substantially influences subspace properties. While LEACE excels across both axes, its generalization remains limited. Moreover, phonetic information in HuBERT can be effectively captured and disentangled, whereas speaker-related information resists compact representation.
This work addresses the limited semantic interpretability of embeddings in conventional latent factor recommender models, such as matrix factorization, which undermines system transparency and controllability. The authors propose the first application of Matryoshka Sparse Autoencoders (MSAE) to collaborative filtering for learning hierarchical, sparse, and disentangled latent representations. Evaluated on the Amazon Fashion dataset, the approach demonstrates strong semantic coherence of extracted features through alignment with item metadata and automatic annotation via large language models. Furthermore, neuron-level interventions—such as manipulating gender-related latent variables—validate the model’s interpretability and controllability. The method effectively mitigates feature splitting and composition issues commonly observed in traditional sparse autoencoders when scaled to larger embedding dimensions.