disentangled representation learning

Designs, builds, or analyzes representation-learning methods and model internals to produce factorized or disentangled embeddings and circuitry in which independent latent factors, features, or tasks map to separate latent dimensions or pathways. This includes techniques and evaluations for discovering, enforcing, or measuring such disentanglement (architectural constraints, loss terms, and circuit analysis) to reduce cross-task/pathway interference and enable localized, modular interventions.

disentangledrepresentationlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Defining and Measuring Disentanglement for non-Independent Factors of Variation

Aug 13, 2024
AA
Antonio Almudévar
🏛️ ViV oLab | Aragón Institute for Engineering Research | University of Zaragoza

Existing disentanglement definitions and metrics assume mutual independence among latent factors, failing to capture inherent statistical dependencies among real-world factors—leading to poor generalization in practical scenarios. Method: We propose the first information-theoretic, generalized disentanglement definition that explicitly accommodates non-independent factors and establish its theoretical connection to the information bottleneck principle. Building upon this, we design the first computable, robust disentanglement metric for non-independent factors—the Generalized Disentanglement Score (G-Disentanglement Score)—integrating mutual information, conditional mutual information, and statistical dependence modeling. Results: Evaluated on controlled synthetic experiments and a unified benchmark, our metric consistently outperforms existing measures across multiple non-independent factor settings, achieving an average improvement of 23.6%. It exhibits strong theoretical grounding and empirical consistency, providing a principled, generalizable evaluation standard for representation learning in realistic settings.

Defining disentanglement for dependent factors of variationMeasuring disentanglement under non-independent factor scenariosProposing information theory-based disentanglement metrics

Knowledge distillation’s impact on internal computational mechanisms remains poorly understood, particularly regarding how student models restructure, compress, or discard teacher components. Method: Using GPT2-small and DistilGPT2, we introduce an influence-weighted component alignment metric to quantify functional module alignment post-distillation. We integrate mechanistic interpretability, circuit analysis, activation tracing, and influence functions to assess alignment across multiple tasks. Contribution/Results: We find that distilled students rely on fewer—but more influential—components, challenging the “black-box equivalence” assumption. Although output behavior remains similar, internal computation undergoes significant shifts, degrading robustness and generalization. Our framework provides the first interpretable, quantitative diagnostic tool for assessing functional fidelity in model compression, advancing trustworthy and explainable knowledge distillation.

Comparing teacher and student model circuits and representationsQuantifying functional alignment beyond output similarity in distilled modelsUnderstanding internal computational transformations in knowledge distillation

This study investigates whether internal circuits in language models exhibit task-specificity and consistency, and how such properties inform our understanding of—and ability to intervene on—model behavior. Employing edge attribution patching and component ablation, the authors systematically evaluate causally critical subgraphs within attention heads and MLP layers across six tasks and seven models. Their analysis reveals, for the first time, that circuits within a single task are highly reused and essential for performance, yet circuits across different tasks substantially overlap, with task-exclusive components contributing minimally. This finding challenges the prevailing assumption of task-dedicated circuits and offers a new perspective on model interpretability and targeted intervention.

circuitsconsistencylanguage models

This work addresses the challenge of disentangling underlying factors of variation in unsupervised representation learning by introducing Holographic Reduced Representations (HRR) for the first time into this domain. The proposed method models latent variables as vector superpositions of symbol–value pairs and leverages HRR’s unbinding operation as an inductive bias to encourage approximately independent factorized representations. Theoretical analysis derives an upper bound on the information capacity per slot, offering an information-theoretic interpretation of disentanglement. Empirical results demonstrate that the approach outperforms existing baselines in terms of latent traversability and standard disentanglement metrics, while also exhibiting superior robustness to noise and consistently stable reconstruction performance across varying signal-to-noise ratios.

disentanglementfactors of variationholographic reduced representations

Disentangling Representations through Multi-task Learning

Jul 15, 2024
PV
Pantelis Vafidis
🏛️ Caltech

This study investigates how multi-task learning drives agents to spontaneously develop disentangled representations—orthogonal, generalizable internal coordinate systems that separate latent factors of the world. Methodologically, it integrates recurrent neural networks (RNNs), which implement continuous attractor dynamics for disentanglement, with Transformer architectures, complemented by latent-variable decoding and an out-of-distribution (OOD) zero-shot generalization evaluation framework. Theoretically and empirically, it establishes for the first time that optimal multi-task evidence accumulation implicitly induces disentanglement; it further identifies critical conditions—governing noise level, task cardinality, and decision time—under which disentanglement emerges. Results show that RNNs achieve zero-shot OOD prediction of latent factors, while Transformers exhibit superior disentanglement, deeper world modeling, and enhanced conceptual interpretability, providing a novel mechanistic account and empirical foundation for feature-based generalization.

Demonstrates zero-shot generalization in AI models using disentangled representations.Explores how multi-task learning enables disentangled representations in AI systems.Investigates conditions for disentangled representations in terms of noise and task complexity.

Latest Papers

What's happening recently
View more

This study investigates how network architecture influences the stability-plasticity trade-off and the interplay between interference and transfer in continual learning. By systematically comparing modular and monolithic recurrent networks under controlled task similarity and weight initialization scales—and integrating effective dimensionality analysis—the work identifies representational dimensionality as a critical factor determining the efficacy of architectural separation. The findings reveal that in low-dimensional (representationally rich) regimes, modular networks substantially outperform baselines by adaptively shaping a hierarchical representational geometry aligned with task similarity, thereby achieving alignment and orthogonality of task-specific subspaces. In contrast, architectural differences exhibit negligible effects in high-dimensional regimes.

continual learningdimensionalitymodularity

Existing disentangled representation learning methods rely on factor-specific architectures or objectives, limiting generalizability to novel factor structures—e.g., non-independent or co-occurring factors—and necessitating frequent model redesign. Method: We propose modular compositional priors, enabling unified disentanglement at attribute-level, object-level, and their joint combinations (e.g., global style + objects) without modifying network architecture or loss functions. Our approach leverages factor-specific latent recombination rules and tunable mixing strategies, guided by a prior loss and a composition consistency loss to encourage the encoder to autonomously discover underlying factor structures. Contribution/Results: Our method achieves competitive performance on standard attribute- and object-disentanglement benchmarks and, for the first time, successfully realizes joint disentanglement of global style and objects. This demonstrates both broad applicability across diverse factor configurations and empirical effectiveness.

Achieves joint disentanglement of global style and objects simultaneouslyEnables disentanglement of attributes and objects without architecture changesProposes modular compositional bias for disentangled representation learning

Current approaches struggle to extract interpretable low-dimensional representations from sparse or incomplete similarity data, limiting our understanding of representational structures in neural, behavioral, and artificial intelligence systems. This work proposes Similarity Representation Factorization (SRF), a novel method that integrates non-negative matrix factorization with low-dimensional embedding to enable, for the first time, generalizable and interpretable extraction of representational dimensions. SRF effectively recovers task-specific model dimensions, accurately predicts independent behavioral attributes, and substantially enhances both exploratory analysis capabilities and statistical power in hypothesis testing. The method is broadly applicable to heterogeneous, multi-source similarity data, offering a robust framework for uncovering latent structure across diverse domains.

dimensionsinterpretabilityneural data

This work addresses the challenge that modern neural networks struggle with selective forgetting and long-range extrapolation in tasks exhibiting algebraic structure, such as modular arithmetic, cyclic reasoning, and Lie group dynamics. To overcome this limitation, the authors propose the Bilinear Multilayer Perceptron (Bilinear MLP), which explicitly incorporates multiplicative interactions as an inductive bias to encourage the learning of structurally disentangled internal representations. Theoretical analysis reveals that this architecture possesses a “non-mixing” property under gradient flow, causing functional components to separate into orthogonal subspaces—a characteristic that facilitates precise model editing. Empirical results demonstrate that, compared to conventional pointwise nonlinear networks, the Bilinear MLP recovers operators aligned with the underlying true algebraic structures, significantly improving performance in targeted forgetting and generalization tasks.

algebraic structurelong-horizon extrapolationmodel editability

From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?

Dec 17, 2025
AM
Aaron Mueller
🏛️ Boston University | Harvard University | Mila – Quebec AI Institute | Goodfire | University of Tübingen

This work investigates whether sparse autoencoders (SAEs) and sparse linear probes can reliably disentangle and localize causally relevant semantic concepts—such as sentiment, domain, or tense—when concepts exhibit controlled inter-concept correlations. Method: We introduce the first evaluation framework that explicitly manipulates multi-concept correlations, integrating subspace projection analysis, feature steering interventions, and quantitative disentanglement metrics. Contribution/Results: We find that (1) the mapping from concepts to features is many-to-one, rendering conventional correlation-based disentanglement metrics insufficient for guaranteeing steering independence; (2) while individual features lack concept selectivity, their causal effects are confined to orthogonal subspaces; and (3) reliable interpretability assessment requires combinatorial, intervention-driven evaluation rather than isolated metrics. Our framework establishes a new paradigm and empirical benchmark for rigorously validating the reliability of interpretability methods in language models.

Assesses feature independence under controlled concept correlationsEvaluates concept disentanglement in neural network interpretability methodsExamines concept selectivity and manipulation in steering experiments

Hot Scholars

BG

Ben Glocker

Imperial College London
Medical Image AnalysisComputer VisionMachine Learning
GR

Gaël Richard

Professor, Télécom Paris, Institut polytechnique de Paris
Audio signal processingMachine listeningMusic ProcessingMusic Information Retrieval
LS

Linlin Shen

Shenzhen University
Deep LearningComputer VisionFacial Analysis/RecognitionMedical Image Analysis
IR

Imran Razzak

MBZUAI, Abu Dhabi
Human-Centered AIMedical Image AnalysisMedical Artificial IntelligenceComputational Biology