continual masked image modeling

Designs and implements continual masked image modeling systems that apply masked-image reconstruction as a self-supervised pretext task at each incremental learning step. This work builds the masking and reconstruction targets, selects losses and training schedules, and engineers gradient routing and regularization through the backbone to encourage task-agnostic features while preserving long-term discriminability across tasks.

continualmaskedimagemodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.58
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

This work presents the first systematic survey of continual self-supervised learning (CSSL) in vision, addressing the challenge of enabling models to learn continuously from unlabeled data streams while mitigating catastrophic forgetting. By analyzing existing evaluation protocols, investigating the mechanisms through which self-supervised objectives confer robustness to forgetting, and integrating insights from loss landscape geometry and methodological taxonomies, the study establishes a unified classification framework encompassing six major anti-forgetting strategies—namely distillation, replay, regularization, and others. The survey clarifies the current state of CSSL research, highlights inconsistencies in evaluation practices, reveals a pathway toward large-scale continual pre-training, and identifies key challenges such as scalability and rapid adaptation.

catastrophic forgettingcontinual learninglifelong learning

Must-Read Papers

Most classic and influential ideas
View more

MINR: Implicit Neural Representations with Masked Image Modelling

Jul 30, 2025
SL
Sua Lee
🏛️ Seoul National University

Existing masked autoencoders (e.g., MAE) rely heavily on predefined masking strategies and exhibit limited generalization to out-of-distribution (OOD) data. Method: This paper proposes MINR, the first framework to integrate implicit neural representations (INRs) into masked image modeling. MINR models images as continuous mappings from spatial coordinates to pixel values, enabling geometrically aware structural modeling and enforcing smoothness priors. This continuous functional formulation inherently reduces dependence on specific masking schemes, enhances reconstruction stability, improves OOD robustness, and lowers parameter count. Contribution/Results: Experiments demonstrate that MINR consistently outperforms MAE both in-domain and across diverse OOD scenarios, achieving superior reconstruction fidelity, generalization capability, and transfer performance on downstream tasks. These results validate the effectiveness and universality of continuous function modeling for self-supervised visual representation learning.

Enhances generalization for out-of-distribution data in image reconstructionImproves robustness to varying masking strategies in self-supervised learningReduces model complexity while maintaining or improving performance

Existing Masked Autoencoders (MAEs) employ random masking, disregarding inter-patch information content variability and downstream task requirements, thereby limiting representation discriminability and generalization. To address this, we propose an end-to-end differentiable, downstream-aware mask learning framework that, for the first time, backpropagates downstream task gradients into the MAE pretraining masking selection process. Our method jointly optimizes task-oriented dynamic masking policies across multiple levels, enabling gradient-driven mask scheduling without requiring additional annotations. It supports plug-and-play integration of arbitrary downstream task feedback signals. Extensive experiments demonstrate consistent and significant improvements over MAE and other baselines across diverse vision benchmarks—including image classification, object detection, and semantic segmentation—validating both the effectiveness and generality of task-driven masking for self-supervised representation learning.

Addresses uniform patch masking limitation in self-supervised learningEnhances visual representation learning via task-guided multi-level optimizationOptimizes masking strategy for downstream tasks in MAE

Masked Autoencoders are Robust Data Augmentors

Jun 10, 2022
HX
Haohang Xu
🏛️ Shanghai Jiao Tong University

Deep neural networks are prone to overfitting in visual classification tasks, and conventional image augmentation techniques—largely relying on linear geometric or photometric transformations—struggle to generate semantically consistent yet discriminative hard examples. To address this, we propose Mask-Reconstruct Augmentation (MRA), the first method to integrate masked autoencoders (built upon Vision Transformer architectures) into supervised, semi-supervised, and few-shot classification pipelines. MRA employs stochastic block masking coupled with joint pixel-level reconstruction and classification training, yielding nonlinear, semantically coherent distorted views. Crucially, it enables model-driven hard-example generation, overcoming the limitations of hand-crafted augmentation heuristics. Extensive experiments across benchmarks—including ImageNet—demonstrate that MRA consistently improves classification accuracy and generalization performance across supervised, semi-supervised, and 5-shot settings, validating its robustness and broad applicability.

Generating hard augmented examples beyond linear transformationsImproving image classification via model-based nonlinear augmentationOvercoming over-fitting in deep neural networks

Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training

Jun 12, 2023
LB
L. Baraldi
🏛️ University of Modena and Reggio Emilia | NVIDIA AI Technology Center | IIT-CNR

To address input noise and pretraining-finetuning mismatch caused by masked modeling in vision transformer pretraining, this paper proposes MaPeT. Methodologically, MaPeT jointly models structural dependencies among image patches via autoregressive masking and random block permutation—eliminating distributional shift induced by conventional random masking. It further introduces auxiliary positional embeddings to mitigate positional information inconsistency between pretraining and finetuning. Additionally, we design a k-CLIP visual tokenizer that maps image patches to discrete CLIP-aligned semantic tokens. Experiments demonstrate that MaPeT achieves state-of-the-art (SOTA) performance on ImageNet among models of comparable parameter count. The code and pretrained models are publicly released.

Image UnderstandingPre-trainingVisual Transformers

Latest Papers

What's happening recently
View more

This work addresses the poor generalization of existing AI-generated image detection methods in the face of rapidly evolving generative models. The authors propose the first three-stage continual learning framework tailored for this task: first, a parameter-efficient fine-tuning strategy is employed to build a strong offline detector with enhanced generalization; second, catastrophic forgetting is mitigated through progressive-complexity data augmentation combined with K-FAC–approximated Hessian regularization; third, linear mode connectivity interpolation is leveraged to improve cross-model transferability. Evaluated on a comprehensive benchmark encompassing 27 generative models, the proposed offline detector achieves a 5.51% mAP improvement over baseline methods, and the continual learning phase attains an average accuracy of 92.20%, substantially outperforming current state-of-the-art approaches.

adaptabilityAI-generated image detectioncatastrophic forgetting

MaskAnyNet: Rethinking Masked Image Regions as Valuable Information in Supervised Learning

Nov 16, 2025
JH
Jingshan Hong
🏛️ Zhejiang University of Technology | Zhejiang Normal University

Traditional supervised learning for image classification often discards masked pixels outright, leading to contextual information loss and degradation of fine-grained discriminative features. To address this, we propose a novel “mask-as-knowledge” paradigm that explicitly treats masked regions as semantically rich auxiliary supervision signals—rather than mere occlusions. Our method employs a dual-branch architecture: one branch processes visible pixels, while the other reconstructs masked regions; both branches are jointly optimized via classification loss and mask reconstruction loss, thereby enforcing local–global contextual consistency. This relearning mechanism is architecture-agnostic, seamlessly integrating with both CNNs and Transformers. Extensive experiments on multiple fine-grained visual recognition benchmarks demonstrate significant performance gains, validating the approach’s effectiveness in enhancing feature diversity and preserving discriminative details without requiring architectural modifications.

Addresses underutilization of discarded pixels in supervised image maskingExploits masked regions as semantic diversity sources rather than ignored dataSolves loss of fine-grained features caused by traditional masking methods

Continual Unlearning for Text-to-Image Diffusion Models: A Regularization Perspective

Nov 11, 2025
JL
Justin Lee
🏛️ The Ohio State University | Michigan State University | Texas A&M University

This work addresses the rapid utility collapse of text-to-image diffusion models under continual unlearning—i.e., sequential processing of multiple forgetting requests—caused by parameter drift. We propose a semantic-aware gradient projection regularization method that projects parameter update directions onto the orthogonal complement of the gradient subspace spanned by retained tasks, coupled with a pre-trained weight preservation mechanism to effectively suppress cumulative parameter deviation. As the first systematic study of continual unlearning for diffusion models, our approach is compatible with existing unlearning algorithms and significantly improves post-unlearning image generation quality and semantic fidelity across multiple forgetting rounds. Quantitative evaluation shows consistent superiority over baselines in FID, CLIP Score, and human assessments. Our work establishes a novel paradigm for secure maintenance and auditable updating of generative models.

Addressing sequential unlearning requests in text-to-image diffusion modelsMitigating cumulative parameter drift from pre-training weights via regularizationPreserving retained knowledge while removing designated concepts continually

Sharing the Learned Knowledge-base to Estimate Convolutional Filter Parameters for Continual Image Restoration

Nov 07, 2025
AK
Aupendu Kar
🏛️ Dolby Laboratories, Inc | Indian Institute of Technology, Kharagpur

Continual image restoration suffers from catastrophic forgetting of previous tasks, while existing solutions often require backbone modifications or incur substantial computational overhead. Method: This paper proposes a lightweight convolutional layer enhancement method that operates without altering the backbone network. Its core innovation is a dynamic filter parameter generation mechanism based on a shared knowledge base, which decomposes convolutional weights into task-invariant bases and task-specific increments, updated efficiently via a lightweight adaptation module. Contribution/Results: The method enables dynamic injection of new-task parameters without significantly increasing inference latency, preserving performance on historical tasks while adapting effectively to new ones. Experiments on multiple continual image restoration benchmarks demonstrate substantial mitigation of catastrophic forgetting: average PSNR improvements of 1.2–2.3 dB on new tasks, with inference speed nearly identical to the original model.

Addresses continual learning for image restoration without forgetting previous tasksEliminates architectural modifications by sharing knowledge across restoration tasksMaintains computational efficiency while adapting to new image degradation types

The mechanistic principles and theoretical limits of masked pretraining in multimodal representation learning remain poorly understood. To address this, we propose Randomly Random Mask Autoencoding (R²MAE), which dynamically randomizes the masking ratio during pretraining—departing from conventional fixed-ratio schemes—to compel models to learn multiscale features. Leveraging minimum-norm regression theory in high-dimensional linear models, we systematically characterize the behavior of masked autoencoding across diverse modalities—including language, vision, DNA sequences, and single-cell data—and validate its architectural generality across MLPs, CNNs, and Transformers. Extensive experiments demonstrate that R²MAE consistently outperforms standard and state-of-the-art masking strategies on cross-modal downstream tasks, yielding substantial gains in representation quality and generalization. Our work establishes the first unified theoretical framework for masked pretraining and introduces a scalable, principled paradigm for multimodal self-supervised learning.

Characterizing mask-based pretraining behavior via linear regressionProposing improved masking scheme for universal representation learningUncovering novel aspects of mask pretraining through linear models

Hot Scholars

QZ

Qihang Zhang

The Chinese University of Hong Kong
computer visionrobotics
JG

Jiatao Gu

UPenn CIS / Apple MLR
machine learninggenerative modelsnatural language processingcomputer vision
SJ

SouYoung Jin

Dartmouth College
Computer VisionMachine LearningCognitive Science
YT

Yujin Tang

Dartmouth College
Computer VisionVideo UnderstandingVideo GenerationMLLM
YH

Yifan Hu

Tsinghua University
Spiking Neural NetworksComputational NeuroscienceDeep Learning