masked channel pretraining

Design and implement pretraining pipelines and encoder/decoder architectures that learn channel-conditional latent representations by reconstructing intentionally masked channels (masked autoencoders), including mechanisms that explicitly represent channel availability via masks; and analyze reconstruction losses, masking strategies, and model robustness to channel dropouts to enable graceful estimation when channels are missing or unreliable.

maskedchannelpretraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.62
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing masked autoencoders (MAEs) based on random patch masking struggle to model cross-channel dependencies in multi-channel imaging (MCI), where channels exhibit strong complementarity and low redundancy. To address this, we propose the Channel-Aware Masked Autoencoder (CA-MAE). Its key contributions are: (1) a dynamic channel-block joint masking strategy that explicitly accounts for channel heterogeneity; (2) a memory token mechanism coupled with a hybrid token fusion module, jointly integrating fine-grained local structures and global cross-channel relationships; and (3) a channel-conditioned lightweight decoder enabling structural-aware reconstruction. Evaluated on multimodal remote sensing and microscopy benchmarks—including CHAMMI, JUMP-CP, and So2Sat—CA-MAE outperforms state-of-the-art multi-channel Vision Transformers by 3.0–21.5%, achieving significant gains in cross-channel reconstruction fidelity and downstream task performance.

Addressing limitations of random patch masking in MAEsEnhancing feature reconstruction across diverse MCI channelsImproving cross-channel learning in Multi-Channel Imaging (MCI)

Latent Diffusion Models with Masked AutoEncoders

Jul 14, 2025
JL
Junho Lee
🏛️ Seoul National University

Existing latent diffusion models (LDMs) suffer from a fundamental trade-off among latent-space smoothness, perceptual compression quality, and reconstruction fidelity—largely attributable to suboptimal autoencoder design. Method: We identify this architectural limitation and propose LDMAEs, a novel framework integrating Variational Masked Autoencoders (VMAEs) into the LDM paradigm. LDMAEs are the first to incorporate hierarchical feature modeling—inspired by Masked Autoencoders (MAEs)—into a variational autoencoding structure, enabling joint optimization of smoothness, perceptual quality, and fidelity directly in the latent space. VMAEs employ layered masking and reconstruction to learn compact, perception-driven representations, thereby improving prior consistency and decoder stability. Results: Extensive experiments demonstrate that LDMAEs significantly outperform state-of-the-art LDM baselines on key metrics—including FID and LPIPS—while reducing sampling computational overhead by over 30%. The framework achieves superior generative quality, inference efficiency, and training stability.

Addressing limitations in latent smoothness and reconstruction qualityExploring optimal autoencoder design for Latent Diffusion ModelsImproving image generation efficiency with Variational Masked AutoEncoders

A Multi-Task Foundation Model for Wireless Channel Representation Using Contrastive and Masked Autoencoder Learning

May 14, 2025
BG
Berkay Guler
🏛️ University of California | Nokia | Universitat Pompeu Fabra

Existing self-supervised learning methods naively adapt image/text paradigms to wireless channel representation, overlooking intrinsic channel properties—including spatiotemporal-frequency correlation, hardware constraints, and noise sensitivity. To address this, we propose WiMAE, the first foundation model for wireless channels: a Transformer-based masked autoencoder. We further extend it into ContraWiMAE, a novel multi-task framework that unifies masked reconstruction with noise-augmented contrastive learning—marking the first such integration. Our method introduces a channel-structure-aware masking strategy and pretraining paradigm tailored to multi-antenna systems, enhancing representation discriminability and cross-scenario generalization. Evaluated on diverse downstream tasks, WiMAE achieves a 12.3% improvement in linear separability and reduces few-shot adaptation error by 19.7%, significantly boosting data efficiency and environmental robustness.

Addressing unique wireless communication characteristics in self-supervised learningEnhancing channel representation via multi-task contrastive and reconstruction learningImproving adaptability and separability in diverse wireless environments

Rethinking Patch Dependence for Masked Autoencoders

Jan 25, 2024
LF
Letian Fu
🏛️ UC Berkeley | UCSF

This work investigates the roles of masked-patch self-attention and masked-to-visible cross-attention in the MAE decoder for representation learning, revealing that image reconstruction primarily relies on global semantic representations extracted by the encoder—not on intra-masked-patch interactions within the decoder. Motivated by this finding, we propose CrossMAE: a streamlined framework that retains only the cross-attention mechanism while entirely removing self-attention among masked tokens in the decoder. This design is the first to empirically demonstrate that MAE’s effectiveness stems from the encoder’s strong global modeling capacity, challenging the prevailing assumption that decoder-side modeling of dependencies among masked patches is essential. Evaluated across ViT-S to ViT-H architectures, CrossMAE matches or surpasses standard MAE in performance while reducing GPU memory consumption by 37% and FLOPs by 42%. Code and pretrained models are publicly available.

Challenges mask token interaction necessity in masked pretrainingExamines inter-patch dependencies in MAE decoders for representation learningProposes CrossMAE using only cross-attention to reduce computation

Simplified and Generalized Masked Diffusion for Discrete Data

Jun 06, 2024
JS
Jiaxin Shi
🏛️ Google DeepMind

Existing masked diffusion models suffer from overly complex modeling, suboptimal training objectives, and reliance on heuristic hyperparameter tuning. This paper proposes a simplified, general-purpose masked diffusion framework for discrete data generation. First, it establishes that the continuous-time variational lower bound of such models is mathematically equivalent to a weighted integral of cross-entropy loss—enabling a unified, principled training objective. Second, it introduces a state-dependent masking schedule, eliminating redundant parameterization and ad hoc corrections. The method integrates continuous-time diffusion dynamics, discrete variational inference, and weighted cross-entropy optimization. Experiments demonstrate that our approach outperforms same-scale diffusion language models on OpenWebText; achieves state-of-the-art zero-shot performance on 4 out of 5 language modeling benchmarks; and attains 2.75/3.40 bits per dimension on CIFAR-10 and ImageNet 64×64 image modeling—surpassing same-scale autoregressive baselines.

Complexity OptimizationMasked Diffusion ModelsPerformance Enhancement

Latest Papers

What's happening recently
View more

This work addresses the limitation of conventional variational autoencoders (VAEs), whose encoders—constrained by the reparameterization trick—struggle to model complex posterior distributions. To overcome this, the authors propose a novel encoder that, for the first time, integrates a diffusion model into the VAE encoding process. They further introduce an alternating training strategy inspired by the Expectation-Maximization (EM) algorithm, which effectively aligns the optimization objectives of the encoder and decoder, thereby ensuring reliable synchronization in the latent space. The proposed approach preserves the simplicity and efficiency of standard diffusion model training while substantially enhancing the model’s capacity to capture complex data distributions and improving reconstruction quality.

decoder pressurediffusion encoderencoder-decoder synchronization

This work investigates channel redundancy in latent diffusion models and proposes a lightweight approach that employs a learnable channel importance predictor—a two-layer MLP—to score globally pooled features and generate a soft mask that dynamically suppresses less informative channels. The method jointly optimizes diffusion loss, reconstruction loss, and a sparsity regularizer, revealing for the first time a “sparsity collapse” phenomenon under strong sparsity constraints. Notably, the denoising UNet maintains or even improves performance under extreme channel suppression. Experiments on CIFAR-10 demonstrate reduced diffusion loss (0.0236 vs. 0.0240) and lower VAE reconstruction MSE (22.59 vs. 24.67), confirming the model’s robustness to channel sparsification.

channel redundancygenerative performancelatent diffusion models

This work addresses the trade-off between event understanding and generalization capability in existing mask prediction–based audio self-supervised learning methods, which often incur high computational costs. To reconcile efficiency and effectiveness, the authors propose a lightweight Dispersion-Weighted Masking (DWM) strategy that leverages the spectral sparsity of audio spectrograms to dynamically adjust the masking distribution, thereby enhancing representation quality. By prioritizing informative yet sparse regions in the time–frequency domain, DWM significantly reduces computational complexity while consistently improving performance across multiple audio event understanding benchmarks. The approach effectively mitigates the longstanding tension between model efficiency and representational power in audio self-supervised learning.

audio self-supervised learningcomputational overheadgeneralization trade-off

Hot Scholars

TK

Taejoon Kim

Professor of The School of ECEE, Arizona State University
Wireless CommunicationsSignal ProcessingOptimization
CW

Chunwei Wang

Researcher, Huawei Noah's Ark Lab
Computer VisionAutonomous DrivingMultimodality
AB

Anindya Bijoy Das

Assistant Professor, ECE, The University of Akron
Federated LearningLarge Language ModelsInformation TheorySignal Processing
AR

Arun Ravindran

University of North Carolina at Charlotte
Computing systemsEdge computingAI
CG

Christopher G. Brinton

Elmore Associate Professor of ECE, Purdue University
NetworkingMachine LearningCommunicationsEdge Computing