Score
Designs and trains autoencoder models tailored to each data modality that produce sparse latent codes; implements modality-specific encoders/decoders, sparsity regularizers, and reconstruction objectives to preserve per-modality feature geometry and improve within-modality reconstruction fidelity. Engineers the latent space to decompose embeddings into monosemantic latents and to prevent unwanted mixing between modalities by enforcing per-modality sparsity and separation constraints.
This work identifies and formally names a previously uncharacterized phenomenon in vision-language models termed “cross-modal feature heterogeneity,” wherein semantically equivalent concepts activate sparse features along inconsistent directions across modalities, leading to modality fragmentation that undermines interpretability and controllability. The study demonstrates that mere alignment of activations is insufficient to resolve this feature mismatch. To address this, the authors propose a two-stage strategy: first preserving each modality’s intrinsic feature geometry using modality-specific sparse autoencoders, followed by post-hoc alignment of corresponding cross-modal features. This approach significantly improves reconstruction fidelity and achieves superior performance in cross-modal retrieval and concept-guided generation tasks.
This study investigates how sparse coding organizes the representational structure of language model activation vectors and uncovers its intrinsic links to feature disentanglement and reconstruction fidelity. Method: We propose SAEMA to empirically validate representational hierarchy; formally define local and global representations; establish a causal relationship between their separability and reconstruction quality; and reinterpret sparsity principles from a geometric perspective. Technically, we integrate rank analysis of symmetric positive semi-definite matrices, modal tensor decomposition, noise-robustness evaluation, optimization-driven representation intervention, and joint modeling of sparse coding and feature merging. Contributions/Results: Empirical results demonstrate that sparse coding not only enhances feature discriminability but also introduces orthogonal redundant dimensions; crucially, representation separability—rather than sparsity alone—is the decisive factor governing reconstruction performance. These findings provide both theoretical foundations and empirical evidence for representation disentanglement and tool design in interpretable AI.
Classical sparse autoencoders (SAEs) and variational autoencoders (VAEs) suffer from structural limitations in modeling multi-manifold data—specifically, theoretical constraints under manifold union assumptions and biased intrinsic dimension estimation. Method: We propose a novel hybrid autoencoder integrating deterministic sparse regularization with stochastic encoding, augmented by manifold-aware optimization. Contribution/Results: We formally characterize, for the first time, the theoretical limitations of SAEs/VAEs under multi-manifold union structures. We prove that our model’s global optimum exactly recovers the underlying multi-manifold geometry and yields unbiased intrinsic dimension estimates. Guided by theory, we design an end-to-end loss function. Empirically, our method outperforms same-capacity SAEs, VAEs, and several diffusion models on synthetic benchmarks, real-world images, and large-language-model activation data—achieving lower reconstruction error, sparser latent representations, and more accurate manifold dimension estimation.
This paper identifies “feature absorption” as a critical failure mode in sparse autoencoders (SAEs) when decomposing large language model activations: monosemantic high-level features (e.g., “mathematics”) are competitively suppressed by their fine-grained subfeatures (e.g., “algebra”, “geometry”), leading to interpretability collapse. Method: We design a first-letter identification synthetic task and introduce a ground-truth–driven diagnostic framework with quantitative activation interpretability evaluation. Contribution/Results: We provide the first systematic empirical confirmation that feature absorption is pervasive, highly robust, and non-monotonically alleviated by increasing SAE scale or sparsity. Crucially, hyperparameter tuning alone cannot resolve it, necessitating foundational reformulation of SAE theory. Our work establishes the first benchmark and diagnostic paradigm for feature splitting/absorption failures in interpretable AI—offering both a standardized testbed and methodological framework for diagnosing representational pathologies in dictionary learning-based interpretability methods.
This work addresses the joint problem of intrinsic dimension estimation and geometry-invariant embedding learning for nonlinear manifold-structured data. We propose an autoencoder framework incorporating orthogonality constraints on hidden-layer gradients. Methodologically, we establish, for the first time, a theoretical connection between gradient orthogonality in neural network latent spaces and the local tangent space dimension of the underlying manifold; this enables simultaneous intrinsic dimension estimation, learning of invertible embedding mappings, and construction of coordinate-invariant representations under local Lie group actions on low-dimensional submanifolds. Our key contribution lies in unifying gradient orthogonality with differential-geometric structure, thereby extending invariant representation learning to continuous group actions. Experiments on standard benchmarks demonstrate accurate intrinsic dimension estimation, disentangled representations, and robust group-invariant embeddings, validating both theoretical soundness and algorithmic robustness.
Standard sparse autoencoders tend to learn “split dictionaries” in multimodal embedding spaces, where features activate exclusively for a single modality, thereby disrupting cross-modal semantic alignment. To address this issue, this work proposes the first autoencoder framework that integrates group sparsity regularization with cross-modal random masking, explicitly promoting cross-modal consistency within multimodal embedding spaces such as those of CLIP or CLAP. The proposed approach effectively mitigates modality splitting, substantially reduces the occurrence of dead neurons, and enhances the semantic meaningfulness, cross-modal alignment, interpretability, and controllability of the learned features in multimodal tasks.
This work addresses the challenge that sparse autoencoders in current vision-language models struggle to learn cross-modal consistent concepts, particularly suffering from fragmented visual representations. To overcome this, the authors propose the Structured Sparse Autoencoder (S²AE), which uniquely integrates semantic attention similarity with spatial proximity to group image patches. S²AE further introduces structured sparsity regularization—combining group sparsity and exclusive sparsity—to enforce intra-group conceptual consistency and inter-group disentanglement. Evaluated on Qwen2.5-VL-7B-Instruct, the method achieves a 6.06% improvement in mIoU, reduces the l₀ norm to 60.81, explains over 99% of variance, and enhances cross-modal semantic consistency and neuron monosemanticity by 3.08% and 2.37%, respectively.
This work addresses the issue that latent features in sparse autoencoders (SAEs) often conflate multiple semantic subspaces, compromising atomicity and interpretability. To mitigate this, the authors propose a joint training approach that introduces a small meta-SAE to sparsely reconstruct the columns of the main SAE’s decoder. They further impose a decomposability penalty on directions that are easily reconstructible by the meta dictionary, directly optimizing feature atomicity during training. This method is the first to explicitly treat decomposability as a training objective, effectively reducing cross-semantic aliasing. Experiments show a 7.5% reduction in average |φ| and a 7.6% improvement in automatic interpretability scores on GPT-2 Large, along with an 8.6% gain in Fuzz scores on Gemma 2 9B. Qualitative analysis confirms that polysemantic features are successfully disentangled into semantically coherent atomic subfeatures.
This work addresses the instability in downstream readout performance of sparse autoencoders, which can retain different linearly decodable signals despite identical reconstruction error and sparsity. To resolve this, the authors propose the Decoder-Preserving Sparse Autoencoder (DPSAE), which introduces a matrix-valued distortion metric to explicitly disentangle reconstruction quality from readout capability by embedding the optimal ridge regression predictor directly into the reconstruction loss. Furthermore, DPSAE incorporates task-prior-guided rank-relaxation optimization to modulate feature pattern selection. Evaluated on layer 8 of GPT-2 Small, DPSAE reduces held-out readout distortion by 10.6–11.4% while maintaining constant reconstruction NMSE, and its representational fidelity is confirmed through a KL non-inferiority test on natural text outputs.
Existing representation-based encoder generative paradigms face two key challenges: (1) discriminative feature spaces lack compact regularization, causing diffusion sampling to deviate from the data manifold and yield structural distortions; and (2) encoders exhibit weak pixel-level reconstruction capability, limiting geometric and textural fidelity. To address these, we propose a semantic-pixel joint reconstruction objective, achieving—within a compact 16×16, 96-dimensional latent space—the first unified high-semantic and high-fidelity pixel reconstruction. Our method integrates dual reconstruction losses, a compact latent-space design, a unified representation-encoder-based T2I and editing diffusion architecture, and a VAE feature-space adaptation mechanism. Experiments demonstrate significant improvements in reconstruction quality, text-to-image generation, and image editing—achieving state-of-the-art performance—along with accelerated convergence. This validates the feasibility of efficiently transferring understanding-oriented encoders into robust, generative latent spaces.