Score
Design and implement masked autoencoder architectures that partition input tokens into modality-specific groups and perform group-aware masking and reconstruction. This includes building per-group tokenizers, a shared encoder, masking strategies that account for missing groups (e.g., attention masking and coverage-adaptive masking), and reconstruction/regularization terms that enforce continuity or coherence across modality boundaries.
Standard sparse autoencoders tend to learn “split dictionaries” in multimodal embedding spaces, where features activate exclusively for a single modality, thereby disrupting cross-modal semantic alignment. To address this issue, this work proposes the first autoencoder framework that integrates group sparsity regularization with cross-modal random masking, explicitly promoting cross-modal consistency within multimodal embedding spaces such as those of CLIP or CLAP. The proposed approach effectively mitigates modality splitting, substantially reduces the occurrence of dead neurons, and enhances the semantic meaningfulness, cross-modal alignment, interpretability, and controllability of the learned features in multimodal tasks.
This work investigates the roles of masked-patch self-attention and masked-to-visible cross-attention in the MAE decoder for representation learning, revealing that image reconstruction primarily relies on global semantic representations extracted by the encoder—not on intra-masked-patch interactions within the decoder. Motivated by this finding, we propose CrossMAE: a streamlined framework that retains only the cross-attention mechanism while entirely removing self-attention among masked tokens in the decoder. This design is the first to empirically demonstrate that MAE’s effectiveness stems from the encoder’s strong global modeling capacity, challenging the prevailing assumption that decoder-side modeling of dependencies among masked patches is essential. Evaluated across ViT-S to ViT-H architectures, CrossMAE matches or surpasses standard MAE in performance while reducing GPU memory consumption by 37% and FLOPs by 42%. Code and pretrained models are publicly available.
Existing Masked Autoencoders (MAEs) employ random masking, disregarding inter-patch information content variability and downstream task requirements, thereby limiting representation discriminability and generalization. To address this, we propose an end-to-end differentiable, downstream-aware mask learning framework that, for the first time, backpropagates downstream task gradients into the MAE pretraining masking selection process. Our method jointly optimizes task-oriented dynamic masking policies across multiple levels, enabling gradient-driven mask scheduling without requiring additional annotations. It supports plug-and-play integration of arbitrary downstream task feedback signals. Extensive experiments demonstrate consistent and significant improvements over MAE and other baselines across diverse vision benchmarks—including image classification, object detection, and semantic segmentation—validating both the effectiveness and generality of task-driven masking for self-supervised representation learning.
This work investigates the intrinsic learning mechanisms of Masked Autoencoders (MAEs) and discovers that, early in pretraining, MAEs spontaneously develop pattern-driven clustering capabilities over image patches. Leveraging this insight, we propose Self-Guided Masking—a novel masking strategy that dynamically generates semantic-aware masks from the model’s intermediate features, eliminating reliance on external supervision or handcrafted priors. Unlike conventional random masking, our approach formulates mask selection as an adaptive clustering process in feature space, enabling internally driven optimization of the reconstruction objective. Evaluated on ImageNet-1K and downstream tasks including classification, detection, and segmentation, the method consistently improves transfer performance across diverse vision benchmarks. These results empirically validate both the discovered mechanistic principle and the effectiveness, robustness, and generalizability of the proposed self-guided masking paradigm.
To address input noise and pretraining-finetuning mismatch caused by masked modeling in vision transformer pretraining, this paper proposes MaPeT. Methodologically, MaPeT jointly models structural dependencies among image patches via autoregressive masking and random block permutation—eliminating distributional shift induced by conventional random masking. It further introduces auxiliary positional embeddings to mitigate positional information inconsistency between pretraining and finetuning. Additionally, we design a k-CLIP visual tokenizer that maps image patches to discrete CLIP-aligned semantic tokens. Experiments demonstrate that MaPeT achieves state-of-the-art (SOTA) performance on ImageNet among models of comparable parameter count. The code and pretrained models are publicly released.
Traditional supervised learning for image classification often discards masked pixels outright, leading to contextual information loss and degradation of fine-grained discriminative features. To address this, we propose a novel “mask-as-knowledge” paradigm that explicitly treats masked regions as semantically rich auxiliary supervision signals—rather than mere occlusions. Our method employs a dual-branch architecture: one branch processes visible pixels, while the other reconstructs masked regions; both branches are jointly optimized via classification loss and mask reconstruction loss, thereby enforcing local–global contextual consistency. This relearning mechanism is architecture-agnostic, seamlessly integrating with both CNNs and Transformers. Extensive experiments on multiple fine-grained visual recognition benchmarks demonstrate significant performance gains, validating the approach’s effectiveness in enhancing feature diversity and preserving discriminative details without requiring architectural modifications.
This work addresses the challenge that sparse autoencoders in current vision-language models struggle to learn cross-modal consistent concepts, particularly suffering from fragmented visual representations. To overcome this, the authors propose the Structured Sparse Autoencoder (S²AE), which uniquely integrates semantic attention similarity with spatial proximity to group image patches. S²AE further introduces structured sparsity regularization—combining group sparsity and exclusive sparsity—to enforce intra-group conceptual consistency and inter-group disentanglement. Evaluated on Qwen2.5-VL-7B-Instruct, the method achieves a 6.06% improvement in mIoU, reduces the l₀ norm to 60.81, explains over 99% of variance, and enhances cross-modal semantic consistency and neuron monosemanticity by 3.08% and 2.37%, respectively.
This work addresses the inefficiency of conventional random masking strategies in language model pretraining, which often fail to identify tokens most valuable for learning. The authors propose a dynamic masking approach that prioritizes tokens with high information content and prediction uncertainty, as measured by the model’s own predictive entropy. Notably, this method introduces a self-masking mechanism that operates without requiring an external reference model. By further integrating knowledge distillation into the training process, the approach significantly enhances both training efficiency and downstream performance. Evaluated on the GLUE benchmark, the proposed method achieves an average performance gain of 5% over the baseline, with the combined use of dynamic masking and knowledge distillation yielding state-of-the-art overall results.
This work addresses the limitations of existing speech modeling approaches, which rely heavily on explicit attributes such as pitch, content, and speaker identity and thus struggle to capture implicit factors like timbre, emotion, and background noise. To overcome this, the authors propose RT-MAE, a novel framework that introduces trainable, unsupervised residual tokens into a masked autoencoder architecture. These residual tokens jointly model implicit speech characteristics alongside explicit attributes, enabling the encoding of complex, unannotated information in speech signals. The method significantly improves reconstruction quality and expressive naturalness while preserving content fidelity and speaker similarity. Furthermore, RT-MAE demonstrates strong performance in speech enhancement tasks, effectively balancing noise suppression with the retention of natural speech characteristics.
This study investigates the intrinsic mechanisms underlying the robustness of Masked Autoencoders (MAE) to image degradations such as blur and occlusion in classification tasks. Through layer-wise analysis of token embeddings, combined with subspace separability assessment and global attention visualization, the authors find that MAE establishes persistent global attention early in the encoder and progressively enhances class separability with depth. To quantify this robustness, they introduce two novel metrics: directional alignment between clean and perturbed embeddings, and head-level retention rate of active features under degradation. The results demonstrate that MAE-learned latent representations maintain high classification performance despite image corruption, offering both theoretical insight and quantitative tools for understanding its robustness.