Score
Designs and trains masked autoencoder architectures that ingest two or more data modalities and learn joint representations by masking and reconstructing inputs, performing cross-modal reconstruction, and optionally aligning representations contrastively or fusing modalities late in the network. Implements encoder/decoder components, masking strategies, contrastive and reconstruction losses, and training regimes to enable multimodal imputation, robust joint embeddings, and downstream representation learning.
To address the common issue of missing sequences in multimodal brain MRI data—which degrades model robustness—this paper proposes a pretraining framework based on a Multimodal Masked Autoencoder (MultiMAE). Methodologically, each MRI sequence is treated as an independent modality; a 3D Transformer encoder with late-fusion is employed to capture cross-sequence dependencies, and a multi-decoder reconstruction architecture enables self-supervised learning and cross-modal inference under incomplete inputs. The core contributions are a modality-aware masking mechanism and a decoupled multi-task reconstruction objective, allowing the model to infer missing modalities from available ones. In downstream segmentation and classification tasks, MultiMAE achieves a +10.1 absolute Dice score improvement and a +0.46 Matthews Correlation Coefficient gain over the MAE-ViT baseline under missing-input conditions, demonstrating significantly enhanced generalization and flexibility in downstream adaptation.
This work investigates the roles of masked-patch self-attention and masked-to-visible cross-attention in the MAE decoder for representation learning, revealing that image reconstruction primarily relies on global semantic representations extracted by the encoder—not on intra-masked-patch interactions within the decoder. Motivated by this finding, we propose CrossMAE: a streamlined framework that retains only the cross-attention mechanism while entirely removing self-attention among masked tokens in the decoder. This design is the first to empirically demonstrate that MAE’s effectiveness stems from the encoder’s strong global modeling capacity, challenging the prevailing assumption that decoder-side modeling of dependencies among masked patches is essential. Evaluated across ViT-S to ViT-H architectures, CrossMAE matches or surpasses standard MAE in performance while reducing GPU memory consumption by 37% and FLOPs by 42%. Code and pretrained models are publicly available.
Multimodal variational autoencoders (MVAEs) suffer from overly rigid cross-modal representation coupling, making it difficult to simultaneously ensure high-quality shared representations and modality-specific fidelity. Method: We propose a soft-constrained Mixture-of-Experts (Soft-MoE) prior that replaces hard parameter sharing with learnable gating weights, enabling flexible alignment of modality-specific latent distributions under a unified posterior. This decouples modality-invariant and modality-specific representations while preserving information integrity via variational inference and a soft alignment loss. Contribution/Results: Experiments on multiple benchmarks and real-world multimodal datasets demonstrate that our approach significantly outperforms existing shared-architecture MVAEs. It achieves state-of-the-art performance in both latent representation quality—measured by disentanglement and downstream task accuracy—and missing modality imputation accuracy.
Deep neural networks are prone to overfitting in visual classification tasks, and conventional image augmentation techniques—largely relying on linear geometric or photometric transformations—struggle to generate semantically consistent yet discriminative hard examples. To address this, we propose Mask-Reconstruct Augmentation (MRA), the first method to integrate masked autoencoders (built upon Vision Transformer architectures) into supervised, semi-supervised, and few-shot classification pipelines. MRA employs stochastic block masking coupled with joint pixel-level reconstruction and classification training, yielding nonlinear, semantically coherent distorted views. Crucially, it enables model-driven hard-example generation, overcoming the limitations of hand-crafted augmentation heuristics. Extensive experiments across benchmarks—including ImageNet—demonstrate that MRA consistently improves classification accuracy and generalization performance across supervised, semi-supervised, and 5-shot settings, validating its robustness and broad applicability.
Existing research on masked image modeling (MIM) for self-supervised visual representation learning lacks a unified formalization of pretraining paradigms and standardized, comparable evaluation. Method: We formally categorize MIM into two principal paradigms—reconstruction-based and contrastive-based—and construct an interpretable, hierarchical taxonomy via expert curation and agglomerative clustering. We conduct systematic, unified benchmarking of over 20 state-of-the-art models on ImageNet and other standard datasets. Contribution/Results: Our analysis identifies critical open challenges—including cross-paradigm integration, long-tailed masking strategies, and compute-accuracy trade-offs. To foster reproducibility and standardization, we publicly release a structured literature repository and an extensible evaluation framework on GitHub. This work establishes foundational infrastructure for rigorous, comparable advancement in MIM research.
Standard sparse autoencoders tend to learn “split dictionaries” in multimodal embedding spaces, where features activate exclusively for a single modality, thereby disrupting cross-modal semantic alignment. To address this issue, this work proposes the first autoencoder framework that integrates group sparsity regularization with cross-modal random masking, explicitly promoting cross-modal consistency within multimodal embedding spaces such as those of CLIP or CLAP. The proposed approach effectively mitigates modality splitting, substantially reduces the occurrence of dead neurons, and enhances the semantic meaningfulness, cross-modal alignment, interpretability, and controllability of the learned features in multimodal tasks.
This study investigates the intrinsic mechanisms underlying the robustness of Masked Autoencoders (MAE) to image degradations such as blur and occlusion in classification tasks. Through layer-wise analysis of token embeddings, combined with subspace separability assessment and global attention visualization, the authors find that MAE establishes persistent global attention early in the encoder and progressively enhances class separability with depth. To quantify this robustness, they introduce two novel metrics: directional alignment between clean and perturbed embeddings, and head-level retention rate of active features under degradation. The results demonstrate that MAE-learned latent representations maintain high classification performance despite image corruption, offering both theoretical insight and quantitative tools for understanding its robustness.
This work identifies and formally names a previously uncharacterized phenomenon in vision-language models termed “cross-modal feature heterogeneity,” wherein semantically equivalent concepts activate sparse features along inconsistent directions across modalities, leading to modality fragmentation that undermines interpretability and controllability. The study demonstrates that mere alignment of activations is insufficient to resolve this feature mismatch. To address this, the authors propose a two-stage strategy: first preserving each modality’s intrinsic feature geometry using modality-specific sparse autoencoders, followed by post-hoc alignment of corresponding cross-modal features. This approach significantly improves reconstruction fidelity and achieves superior performance in cross-modal retrieval and concept-guided generation tasks.
This work addresses the challenges of feature entanglement and degraded out-of-distribution (OOD) performance in sparse autoencoders, which arise from under-constrained training objectives and undermine interpretability. To mitigate these issues, the authors propose a mask-based regularization method that randomly replaces input tokens during training to disrupt co-occurring feature patterns. This approach effectively alleviates feature absorption, enhances the stability and robustness of latent representations, and narrows the performance gap between in-distribution and OOD settings. The method is architecture-agnostic and compatible with various sparsity levels, demonstrating consistent improvements across different sparse autoencoder configurations. Furthermore, it leads to better performance on probing tasks, indicating more disentangled and semantically meaningful representations.