Score
Design and implement pretraining pipelines and encoder/decoder architectures that learn channel-conditional latent representations by reconstructing intentionally masked channels (masked autoencoders), including mechanisms that explicitly represent channel availability via masks; and analyze reconstruction losses, masking strategies, and model robustness to channel dropouts to enable graceful estimation when channels are missing or unreliable.
Existing research on masked image modeling (MIM) for self-supervised visual representation learning lacks a unified formalization of pretraining paradigms and standardized, comparable evaluation. Method: We formally categorize MIM into two principal paradigms—reconstruction-based and contrastive-based—and construct an interpretable, hierarchical taxonomy via expert curation and agglomerative clustering. We conduct systematic, unified benchmarking of over 20 state-of-the-art models on ImageNet and other standard datasets. Contribution/Results: Our analysis identifies critical open challenges—including cross-paradigm integration, long-tailed masking strategies, and compute-accuracy trade-offs. To foster reproducibility and standardization, we publicly release a structured literature repository and an extensible evaluation framework on GitHub. This work establishes foundational infrastructure for rigorous, comparable advancement in MIM research.
Existing masked autoencoders (MAEs) based on random patch masking struggle to model cross-channel dependencies in multi-channel imaging (MCI), where channels exhibit strong complementarity and low redundancy. To address this, we propose the Channel-Aware Masked Autoencoder (CA-MAE). Its key contributions are: (1) a dynamic channel-block joint masking strategy that explicitly accounts for channel heterogeneity; (2) a memory token mechanism coupled with a hybrid token fusion module, jointly integrating fine-grained local structures and global cross-channel relationships; and (3) a channel-conditioned lightweight decoder enabling structural-aware reconstruction. Evaluated on multimodal remote sensing and microscopy benchmarks—including CHAMMI, JUMP-CP, and So2Sat—CA-MAE outperforms state-of-the-art multi-channel Vision Transformers by 3.0–21.5%, achieving significant gains in cross-channel reconstruction fidelity and downstream task performance.
Existing latent diffusion models (LDMs) suffer from a fundamental trade-off among latent-space smoothness, perceptual compression quality, and reconstruction fidelity—largely attributable to suboptimal autoencoder design. Method: We identify this architectural limitation and propose LDMAEs, a novel framework integrating Variational Masked Autoencoders (VMAEs) into the LDM paradigm. LDMAEs are the first to incorporate hierarchical feature modeling—inspired by Masked Autoencoders (MAEs)—into a variational autoencoding structure, enabling joint optimization of smoothness, perceptual quality, and fidelity directly in the latent space. VMAEs employ layered masking and reconstruction to learn compact, perception-driven representations, thereby improving prior consistency and decoder stability. Results: Extensive experiments demonstrate that LDMAEs significantly outperform state-of-the-art LDM baselines on key metrics—including FID and LPIPS—while reducing sampling computational overhead by over 30%. The framework achieves superior generative quality, inference efficiency, and training stability.
Existing self-supervised learning methods naively adapt image/text paradigms to wireless channel representation, overlooking intrinsic channel properties—including spatiotemporal-frequency correlation, hardware constraints, and noise sensitivity. To address this, we propose WiMAE, the first foundation model for wireless channels: a Transformer-based masked autoencoder. We further extend it into ContraWiMAE, a novel multi-task framework that unifies masked reconstruction with noise-augmented contrastive learning—marking the first such integration. Our method introduces a channel-structure-aware masking strategy and pretraining paradigm tailored to multi-antenna systems, enhancing representation discriminability and cross-scenario generalization. Evaluated on diverse downstream tasks, WiMAE achieves a 12.3% improvement in linear separability and reduces few-shot adaptation error by 19.7%, significantly boosting data efficiency and environmental robustness.
This work investigates the roles of masked-patch self-attention and masked-to-visible cross-attention in the MAE decoder for representation learning, revealing that image reconstruction primarily relies on global semantic representations extracted by the encoder—not on intra-masked-patch interactions within the decoder. Motivated by this finding, we propose CrossMAE: a streamlined framework that retains only the cross-attention mechanism while entirely removing self-attention among masked tokens in the decoder. This design is the first to empirically demonstrate that MAE’s effectiveness stems from the encoder’s strong global modeling capacity, challenging the prevailing assumption that decoder-side modeling of dependencies among masked patches is essential. Evaluated across ViT-S to ViT-H architectures, CrossMAE matches or surpasses standard MAE in performance while reducing GPU memory consumption by 37% and FLOPs by 42%. Code and pretrained models are publicly available.
Existing masked diffusion models suffer from overly complex modeling, suboptimal training objectives, and reliance on heuristic hyperparameter tuning. This paper proposes a simplified, general-purpose masked diffusion framework for discrete data generation. First, it establishes that the continuous-time variational lower bound of such models is mathematically equivalent to a weighted integral of cross-entropy loss—enabling a unified, principled training objective. Second, it introduces a state-dependent masking schedule, eliminating redundant parameterization and ad hoc corrections. The method integrates continuous-time diffusion dynamics, discrete variational inference, and weighted cross-entropy optimization. Experiments demonstrate that our approach outperforms same-scale diffusion language models on OpenWebText; achieves state-of-the-art zero-shot performance on 4 out of 5 language modeling benchmarks; and attains 2.75/3.40 bits per dimension on CIFAR-10 and ImageNet 64×64 image modeling—surpassing same-scale autoregressive baselines.
This work addresses the limitation of conventional variational autoencoders (VAEs), whose encoders—constrained by the reparameterization trick—struggle to model complex posterior distributions. To overcome this, the authors propose a novel encoder that, for the first time, integrates a diffusion model into the VAE encoding process. They further introduce an alternating training strategy inspired by the Expectation-Maximization (EM) algorithm, which effectively aligns the optimization objectives of the encoder and decoder, thereby ensuring reliable synchronization in the latent space. The proposed approach preserves the simplicity and efficiency of standard diffusion model training while substantially enhancing the model’s capacity to capture complex data distributions and improving reconstruction quality.
This work investigates channel redundancy in latent diffusion models and proposes a lightweight approach that employs a learnable channel importance predictor—a two-layer MLP—to score globally pooled features and generate a soft mask that dynamically suppresses less informative channels. The method jointly optimizes diffusion loss, reconstruction loss, and a sparsity regularizer, revealing for the first time a “sparsity collapse” phenomenon under strong sparsity constraints. Notably, the denoising UNet maintains or even improves performance under extreme channel suppression. Experiments on CIFAR-10 demonstrate reduced diffusion loss (0.0236 vs. 0.0240) and lower VAE reconstruction MSE (22.59 vs. 24.67), confirming the model’s robustness to channel sparsification.
This work addresses the trade-off between event understanding and generalization capability in existing mask prediction–based audio self-supervised learning methods, which often incur high computational costs. To reconcile efficiency and effectiveness, the authors propose a lightweight Dispersion-Weighted Masking (DWM) strategy that leverages the spectral sparsity of audio spectrograms to dynamically adjust the masking distribution, thereby enhancing representation quality. By prioritizing informative yet sparse regions in the time–frequency domain, DWM significantly reduces computational complexity while consistently improving performance across multiple audio event understanding benchmarks. The approach effectively mitigates the longstanding tension between model efficiency and representational power in audio self-supervised learning.