multi-modal masked autoencoder

Designs and trains masked autoencoder architectures that ingest two or more data modalities and learn joint representations by masking and reconstructing inputs, performing cross-modal reconstruction, and optionally aligning representations contrastively or fusing modalities late in the network. Implements encoder/decoder components, masking strategies, contrastive and reconstruction losses, and training regimes to enable multimodal imputation, robust joint embeddings, and downstream representation learning.

multi-modalmaskedautoencoder

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.66
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

MultiMAE for Brain MRIs: Robustness to Missing Inputs Using Multi-Modal Masked Autoencoder

Sep 14, 2025
AC
Ayhan Can Erdur
🏛️ Technical University of Munich (TUM) | Technical University of Munich (TUM) and TUM University Hospital | Munich Center for Machine Learning (MCML) | Imperial College London | Deutsches Konsortium für Translationale Krebsforschung (DKTK) | Helmholtz Center Munich

To address the common issue of missing sequences in multimodal brain MRI data—which degrades model robustness—this paper proposes a pretraining framework based on a Multimodal Masked Autoencoder (MultiMAE). Methodologically, each MRI sequence is treated as an independent modality; a 3D Transformer encoder with late-fusion is employed to capture cross-sequence dependencies, and a multi-decoder reconstruction architecture enables self-supervised learning and cross-modal inference under incomplete inputs. The core contributions are a modality-aware masking mechanism and a decoupled multi-task reconstruction objective, allowing the model to infer missing modalities from available ones. In downstream segmentation and classification tasks, MultiMAE achieves a +10.1 absolute Dice score improvement and a +0.46 Matthews Correlation Coefficient gain over the MAE-ViT baseline under missing-input conditions, demonstrating significantly enhanced generalization and flexibility in downstream adaptation.

Enabling cross-sequence reasoning to infer missing inputsHandling missing MRI sequences in medical imagingLearning robust multi-modal representations for brain MRIs

Rethinking Patch Dependence for Masked Autoencoders

Jan 25, 2024
LF
Letian Fu
🏛️ UC Berkeley | UCSF

This work investigates the roles of masked-patch self-attention and masked-to-visible cross-attention in the MAE decoder for representation learning, revealing that image reconstruction primarily relies on global semantic representations extracted by the encoder—not on intra-masked-patch interactions within the decoder. Motivated by this finding, we propose CrossMAE: a streamlined framework that retains only the cross-attention mechanism while entirely removing self-attention among masked tokens in the decoder. This design is the first to empirically demonstrate that MAE’s effectiveness stems from the encoder’s strong global modeling capacity, challenging the prevailing assumption that decoder-side modeling of dependencies among masked patches is essential. Evaluated across ViT-S to ViT-H architectures, CrossMAE matches or surpasses standard MAE in performance while reducing GPU memory consumption by 37% and FLOPs by 42%. Code and pretrained models are publicly available.

Challenges mask token interaction necessity in masked pretrainingExamines inter-patch dependencies in MAE decoders for representation learningProposes CrossMAE using only cross-attention to reduce computation

Unity by Diversity: Improved Representation Learning in Multimodal VAEs

Mar 08, 2024
TM
Thomas M. Sutter
🏛️ ETH Zurich | UC Irvine

Multimodal variational autoencoders (MVAEs) suffer from overly rigid cross-modal representation coupling, making it difficult to simultaneously ensure high-quality shared representations and modality-specific fidelity. Method: We propose a soft-constrained Mixture-of-Experts (Soft-MoE) prior that replaces hard parameter sharing with learnable gating weights, enabling flexible alignment of modality-specific latent distributions under a unified posterior. This decouples modality-invariant and modality-specific representations while preserving information integrity via variational inference and a soft alignment loss. Contribution/Results: Experiments on multiple benchmarks and real-world multimodal datasets demonstrate that our approach significantly outperforms existing shared-architecture MVAEs. It achieves state-of-the-art performance in both latent representation quality—measured by disentanglement and downstream task accuracy—and missing modality imputation accuracy.

Missing Data ImputationMultimodal Data FusionVariational Autoencoder

Masked Autoencoders are Robust Data Augmentors

Jun 10, 2022
HX
Haohang Xu
🏛️ Shanghai Jiao Tong University

Deep neural networks are prone to overfitting in visual classification tasks, and conventional image augmentation techniques—largely relying on linear geometric or photometric transformations—struggle to generate semantically consistent yet discriminative hard examples. To address this, we propose Mask-Reconstruct Augmentation (MRA), the first method to integrate masked autoencoders (built upon Vision Transformer architectures) into supervised, semi-supervised, and few-shot classification pipelines. MRA employs stochastic block masking coupled with joint pixel-level reconstruction and classification training, yielding nonlinear, semantically coherent distorted views. Crucially, it enables model-driven hard-example generation, overcoming the limitations of hand-crafted augmentation heuristics. Extensive experiments across benchmarks—including ImageNet—demonstrate that MRA consistently improves classification accuracy and generalization performance across supervised, semi-supervised, and 5-shot settings, validating its robustness and broad applicability.

Generating hard augmented examples beyond linear transformationsImproving image classification via model-based nonlinear augmentationOvercoming over-fitting in deep neural networks

Masked Image Modeling: A Survey

Aug 13, 2024
VH
Vlad Hondru
🏛️ University of Bucharest | Amazon | University of Trento

Existing research on masked image modeling (MIM) for self-supervised visual representation learning lacks a unified formalization of pretraining paradigms and standardized, comparable evaluation. Method: We formally categorize MIM into two principal paradigms—reconstruction-based and contrastive-based—and construct an interpretable, hierarchical taxonomy via expert curation and agglomerative clustering. We conduct systematic, unified benchmarking of over 20 state-of-the-art models on ImageNet and other standard datasets. Contribution/Results: Our analysis identifies critical open challenges—including cross-paradigm integration, long-tailed masking strategies, and compute-accuracy trade-offs. To foster reproducibility and standardization, we publicly release a structured literature repository and an extensible evaluation framework on GitHub. This work establishes foundational infrastructure for rigorous, comparable advancement in MIM research.

Automatic LearningComputer VisionMasked Image Modeling

Latest Papers

What's happening recently
View more

Standard sparse autoencoders tend to learn “split dictionaries” in multimodal embedding spaces, where features activate exclusively for a single modality, thereby disrupting cross-modal semantic alignment. To address this issue, this work proposes the first autoencoder framework that integrates group sparsity regularization with cross-modal random masking, explicitly promoting cross-modal consistency within multimodal embedding spaces such as those of CLIP or CLAP. The proposed approach effectively mitigates modality splitting, substantially reduces the occurrence of dead neurons, and enhances the semantic meaningfulness, cross-modal alignment, interpretability, and controllability of the learned features in multimodal tasks.

interpretable semanticsmodality alignmentmultimodal embeddings

This study investigates the intrinsic mechanisms underlying the robustness of Masked Autoencoders (MAE) to image degradations such as blur and occlusion in classification tasks. Through layer-wise analysis of token embeddings, combined with subspace separability assessment and global attention visualization, the authors find that MAE establishes persistent global attention early in the encoder and progressively enhances class separability with depth. To quantify this robustness, they introduce two novel metrics: directional alignment between clean and perturbed embeddings, and head-level retention rate of active features under degradation. The results demonstrate that MAE-learned latent representations maintain high classification performance despite image corruption, offering both theoretical insight and quantitative tools for understanding its robustness.

feature robustnessimage classificationMasked Autoencoders

This work identifies and formally names a previously uncharacterized phenomenon in vision-language models termed “cross-modal feature heterogeneity,” wherein semantically equivalent concepts activate sparse features along inconsistent directions across modalities, leading to modality fragmentation that undermines interpretability and controllability. The study demonstrates that mere alignment of activations is insufficient to resolve this feature mismatch. To address this, the authors propose a two-stage strategy: first preserving each modality’s intrinsic feature geometry using modality-specific sparse autoencoders, followed by post-hoc alignment of corresponding cross-modal features. This approach significantly improves reconstruction fidelity and achieves superior performance in cross-modal retrieval and concept-guided generation tasks.

cross-modal feature heterogeneityfeature alignmentmodality split

This work addresses the challenges of feature entanglement and degraded out-of-distribution (OOD) performance in sparse autoencoders, which arise from under-constrained training objectives and undermine interpretability. To mitigate these issues, the authors propose a mask-based regularization method that randomly replaces input tokens during training to disrupt co-occurring feature patterns. This approach effectively alleviates feature absorption, enhances the stability and robustness of latent representations, and narrows the performance gap between in-distribution and OOD settings. The method is architecture-agnostic and compatible with various sparsity levels, demonstrating consistent improvements across different sparse autoencoder configurations. Furthermore, it leads to better performance on probing tasks, indicating more disentangled and semantically meaningful representations.

feature absorptioninterpretabilityout-of-distribution

Hot Scholars

BD

Begüm Demir

Professor, BIFOLD and Faculty of EECS, Technische Universität Berlin
Remote SensingMachine LearningImage AnalysisSignal Processing
EB

Egor Bondarev

Associate Professor, Eindhoven University of Technology
computer visionAI3D reconstructionreal-time architectures
BS

Bernt Schiele

Professor, Max Planck Institute for Informatics, Saarland University, Saarland Informatics Campus
Computer VisionMachine LearningArtificial IntelligenceAutonomous Driving