Score
Designs, builds, or analyzes representation-learning methods and model internals to produce factorized or disentangled embeddings and circuitry in which independent latent factors, features, or tasks map to separate latent dimensions or pathways. This includes techniques and evaluations for discovering, enforcing, or measuring such disentanglement (architectural constraints, loss terms, and circuit analysis) to reduce cross-task/pathway interference and enable localized, modular interventions.
Existing disentanglement definitions and metrics assume mutual independence among latent factors, failing to capture inherent statistical dependencies among real-world factors—leading to poor generalization in practical scenarios. Method: We propose the first information-theoretic, generalized disentanglement definition that explicitly accommodates non-independent factors and establish its theoretical connection to the information bottleneck principle. Building upon this, we design the first computable, robust disentanglement metric for non-independent factors—the Generalized Disentanglement Score (G-Disentanglement Score)—integrating mutual information, conditional mutual information, and statistical dependence modeling. Results: Evaluated on controlled synthetic experiments and a unified benchmark, our metric consistently outperforms existing measures across multiple non-independent factor settings, achieving an average improvement of 23.6%. It exhibits strong theoretical grounding and empirical consistency, providing a principled, generalizable evaluation standard for representation learning in realistic settings.
Knowledge distillation’s impact on internal computational mechanisms remains poorly understood, particularly regarding how student models restructure, compress, or discard teacher components. Method: Using GPT2-small and DistilGPT2, we introduce an influence-weighted component alignment metric to quantify functional module alignment post-distillation. We integrate mechanistic interpretability, circuit analysis, activation tracing, and influence functions to assess alignment across multiple tasks. Contribution/Results: We find that distilled students rely on fewer—but more influential—components, challenging the “black-box equivalence” assumption. Although output behavior remains similar, internal computation undergoes significant shifts, degrading robustness and generalization. Our framework provides the first interpretable, quantitative diagnostic tool for assessing functional fidelity in model compression, advancing trustworthy and explainable knowledge distillation.
This study investigates whether internal circuits in language models exhibit task-specificity and consistency, and how such properties inform our understanding of—and ability to intervene on—model behavior. Employing edge attribution patching and component ablation, the authors systematically evaluate causally critical subgraphs within attention heads and MLP layers across six tasks and seven models. Their analysis reveals, for the first time, that circuits within a single task are highly reused and essential for performance, yet circuits across different tasks substantially overlap, with task-exclusive components contributing minimally. This finding challenges the prevailing assumption of task-dedicated circuits and offers a new perspective on model interpretability and targeted intervention.
This work addresses the challenge of disentangling underlying factors of variation in unsupervised representation learning by introducing Holographic Reduced Representations (HRR) for the first time into this domain. The proposed method models latent variables as vector superpositions of symbol–value pairs and leverages HRR’s unbinding operation as an inductive bias to encourage approximately independent factorized representations. Theoretical analysis derives an upper bound on the information capacity per slot, offering an information-theoretic interpretation of disentanglement. Empirical results demonstrate that the approach outperforms existing baselines in terms of latent traversability and standard disentanglement metrics, while also exhibiting superior robustness to noise and consistently stable reconstruction performance across varying signal-to-noise ratios.
This study investigates how multi-task learning drives agents to spontaneously develop disentangled representations—orthogonal, generalizable internal coordinate systems that separate latent factors of the world. Methodologically, it integrates recurrent neural networks (RNNs), which implement continuous attractor dynamics for disentanglement, with Transformer architectures, complemented by latent-variable decoding and an out-of-distribution (OOD) zero-shot generalization evaluation framework. Theoretically and empirically, it establishes for the first time that optimal multi-task evidence accumulation implicitly induces disentanglement; it further identifies critical conditions—governing noise level, task cardinality, and decision time—under which disentanglement emerges. Results show that RNNs achieve zero-shot OOD prediction of latent factors, while Transformers exhibit superior disentanglement, deeper world modeling, and enhanced conceptual interpretability, providing a novel mechanistic account and empirical foundation for feature-based generalization.
This study investigates how network architecture influences the stability-plasticity trade-off and the interplay between interference and transfer in continual learning. By systematically comparing modular and monolithic recurrent networks under controlled task similarity and weight initialization scales—and integrating effective dimensionality analysis—the work identifies representational dimensionality as a critical factor determining the efficacy of architectural separation. The findings reveal that in low-dimensional (representationally rich) regimes, modular networks substantially outperform baselines by adaptively shaping a hierarchical representational geometry aligned with task similarity, thereby achieving alignment and orthogonality of task-specific subspaces. In contrast, architectural differences exhibit negligible effects in high-dimensional regimes.
Existing disentangled representation learning methods rely on factor-specific architectures or objectives, limiting generalizability to novel factor structures—e.g., non-independent or co-occurring factors—and necessitating frequent model redesign. Method: We propose modular compositional priors, enabling unified disentanglement at attribute-level, object-level, and their joint combinations (e.g., global style + objects) without modifying network architecture or loss functions. Our approach leverages factor-specific latent recombination rules and tunable mixing strategies, guided by a prior loss and a composition consistency loss to encourage the encoder to autonomously discover underlying factor structures. Contribution/Results: Our method achieves competitive performance on standard attribute- and object-disentanglement benchmarks and, for the first time, successfully realizes joint disentanglement of global style and objects. This demonstrates both broad applicability across diverse factor configurations and empirical effectiveness.
Current approaches struggle to extract interpretable low-dimensional representations from sparse or incomplete similarity data, limiting our understanding of representational structures in neural, behavioral, and artificial intelligence systems. This work proposes Similarity Representation Factorization (SRF), a novel method that integrates non-negative matrix factorization with low-dimensional embedding to enable, for the first time, generalizable and interpretable extraction of representational dimensions. SRF effectively recovers task-specific model dimensions, accurately predicts independent behavioral attributes, and substantially enhances both exploratory analysis capabilities and statistical power in hypothesis testing. The method is broadly applicable to heterogeneous, multi-source similarity data, offering a robust framework for uncovering latent structure across diverse domains.
This work addresses the challenge that modern neural networks struggle with selective forgetting and long-range extrapolation in tasks exhibiting algebraic structure, such as modular arithmetic, cyclic reasoning, and Lie group dynamics. To overcome this limitation, the authors propose the Bilinear Multilayer Perceptron (Bilinear MLP), which explicitly incorporates multiplicative interactions as an inductive bias to encourage the learning of structurally disentangled internal representations. Theoretical analysis reveals that this architecture possesses a “non-mixing” property under gradient flow, causing functional components to separate into orthogonal subspaces—a characteristic that facilitates precise model editing. Empirical results demonstrate that, compared to conventional pointwise nonlinear networks, the Bilinear MLP recovers operators aligned with the underlying true algebraic structures, significantly improving performance in targeted forgetting and generalization tasks.
This work investigates whether sparse autoencoders (SAEs) and sparse linear probes can reliably disentangle and localize causally relevant semantic concepts—such as sentiment, domain, or tense—when concepts exhibit controlled inter-concept correlations. Method: We introduce the first evaluation framework that explicitly manipulates multi-concept correlations, integrating subspace projection analysis, feature steering interventions, and quantitative disentanglement metrics. Contribution/Results: We find that (1) the mapping from concepts to features is many-to-one, rendering conventional correlation-based disentanglement metrics insufficient for guaranteeing steering independence; (2) while individual features lack concept selectivity, their causal effects are confined to orthogonal subspaces; and (3) reliable interpretability assessment requires combinatorial, intervention-driven evaluation rather than isolated metrics. Our framework establishes a new paradigm and empirical benchmark for rigorously validating the reliability of interpretability methods in language models.