Score
Designs and implements continual masked image modeling systems that apply masked-image reconstruction as a self-supervised pretext task at each incremental learning step. This work builds the masking and reconstruction targets, selects losses and training schedules, and engineers gradient routing and regularization through the backbone to encourage task-agnostic features while preserving long-term discriminability across tasks.
Existing research on masked image modeling (MIM) for self-supervised visual representation learning lacks a unified formalization of pretraining paradigms and standardized, comparable evaluation. Method: We formally categorize MIM into two principal paradigms—reconstruction-based and contrastive-based—and construct an interpretable, hierarchical taxonomy via expert curation and agglomerative clustering. We conduct systematic, unified benchmarking of over 20 state-of-the-art models on ImageNet and other standard datasets. Contribution/Results: Our analysis identifies critical open challenges—including cross-paradigm integration, long-tailed masking strategies, and compute-accuracy trade-offs. To foster reproducibility and standardization, we publicly release a structured literature repository and an extensible evaluation framework on GitHub. This work establishes foundational infrastructure for rigorous, comparable advancement in MIM research.
This work presents the first systematic survey of continual self-supervised learning (CSSL) in vision, addressing the challenge of enabling models to learn continuously from unlabeled data streams while mitigating catastrophic forgetting. By analyzing existing evaluation protocols, investigating the mechanisms through which self-supervised objectives confer robustness to forgetting, and integrating insights from loss landscape geometry and methodological taxonomies, the study establishes a unified classification framework encompassing six major anti-forgetting strategies—namely distillation, replay, regularization, and others. The survey clarifies the current state of CSSL research, highlights inconsistencies in evaluation practices, reveals a pathway toward large-scale continual pre-training, and identifies key challenges such as scalability and rapid adaptation.
Existing masked autoencoders (e.g., MAE) rely heavily on predefined masking strategies and exhibit limited generalization to out-of-distribution (OOD) data. Method: This paper proposes MINR, the first framework to integrate implicit neural representations (INRs) into masked image modeling. MINR models images as continuous mappings from spatial coordinates to pixel values, enabling geometrically aware structural modeling and enforcing smoothness priors. This continuous functional formulation inherently reduces dependence on specific masking schemes, enhances reconstruction stability, improves OOD robustness, and lowers parameter count. Contribution/Results: Experiments demonstrate that MINR consistently outperforms MAE both in-domain and across diverse OOD scenarios, achieving superior reconstruction fidelity, generalization capability, and transfer performance on downstream tasks. These results validate the effectiveness and universality of continuous function modeling for self-supervised visual representation learning.
Existing Masked Autoencoders (MAEs) employ random masking, disregarding inter-patch information content variability and downstream task requirements, thereby limiting representation discriminability and generalization. To address this, we propose an end-to-end differentiable, downstream-aware mask learning framework that, for the first time, backpropagates downstream task gradients into the MAE pretraining masking selection process. Our method jointly optimizes task-oriented dynamic masking policies across multiple levels, enabling gradient-driven mask scheduling without requiring additional annotations. It supports plug-and-play integration of arbitrary downstream task feedback signals. Extensive experiments demonstrate consistent and significant improvements over MAE and other baselines across diverse vision benchmarks—including image classification, object detection, and semantic segmentation—validating both the effectiveness and generality of task-driven masking for self-supervised representation learning.
Deep neural networks are prone to overfitting in visual classification tasks, and conventional image augmentation techniques—largely relying on linear geometric or photometric transformations—struggle to generate semantically consistent yet discriminative hard examples. To address this, we propose Mask-Reconstruct Augmentation (MRA), the first method to integrate masked autoencoders (built upon Vision Transformer architectures) into supervised, semi-supervised, and few-shot classification pipelines. MRA employs stochastic block masking coupled with joint pixel-level reconstruction and classification training, yielding nonlinear, semantically coherent distorted views. Crucially, it enables model-driven hard-example generation, overcoming the limitations of hand-crafted augmentation heuristics. Extensive experiments across benchmarks—including ImageNet—demonstrate that MRA consistently improves classification accuracy and generalization performance across supervised, semi-supervised, and 5-shot settings, validating its robustness and broad applicability.
To address input noise and pretraining-finetuning mismatch caused by masked modeling in vision transformer pretraining, this paper proposes MaPeT. Methodologically, MaPeT jointly models structural dependencies among image patches via autoregressive masking and random block permutation—eliminating distributional shift induced by conventional random masking. It further introduces auxiliary positional embeddings to mitigate positional information inconsistency between pretraining and finetuning. Additionally, we design a k-CLIP visual tokenizer that maps image patches to discrete CLIP-aligned semantic tokens. Experiments demonstrate that MaPeT achieves state-of-the-art (SOTA) performance on ImageNet among models of comparable parameter count. The code and pretrained models are publicly released.
This work addresses the poor generalization of existing AI-generated image detection methods in the face of rapidly evolving generative models. The authors propose the first three-stage continual learning framework tailored for this task: first, a parameter-efficient fine-tuning strategy is employed to build a strong offline detector with enhanced generalization; second, catastrophic forgetting is mitigated through progressive-complexity data augmentation combined with K-FAC–approximated Hessian regularization; third, linear mode connectivity interpolation is leveraged to improve cross-model transferability. Evaluated on a comprehensive benchmark encompassing 27 generative models, the proposed offline detector achieves a 5.51% mAP improvement over baseline methods, and the continual learning phase attains an average accuracy of 92.20%, substantially outperforming current state-of-the-art approaches.
Traditional supervised learning for image classification often discards masked pixels outright, leading to contextual information loss and degradation of fine-grained discriminative features. To address this, we propose a novel “mask-as-knowledge” paradigm that explicitly treats masked regions as semantically rich auxiliary supervision signals—rather than mere occlusions. Our method employs a dual-branch architecture: one branch processes visible pixels, while the other reconstructs masked regions; both branches are jointly optimized via classification loss and mask reconstruction loss, thereby enforcing local–global contextual consistency. This relearning mechanism is architecture-agnostic, seamlessly integrating with both CNNs and Transformers. Extensive experiments on multiple fine-grained visual recognition benchmarks demonstrate significant performance gains, validating the approach’s effectiveness in enhancing feature diversity and preserving discriminative details without requiring architectural modifications.
This work addresses the rapid utility collapse of text-to-image diffusion models under continual unlearning—i.e., sequential processing of multiple forgetting requests—caused by parameter drift. We propose a semantic-aware gradient projection regularization method that projects parameter update directions onto the orthogonal complement of the gradient subspace spanned by retained tasks, coupled with a pre-trained weight preservation mechanism to effectively suppress cumulative parameter deviation. As the first systematic study of continual unlearning for diffusion models, our approach is compatible with existing unlearning algorithms and significantly improves post-unlearning image generation quality and semantic fidelity across multiple forgetting rounds. Quantitative evaluation shows consistent superiority over baselines in FID, CLIP Score, and human assessments. Our work establishes a novel paradigm for secure maintenance and auditable updating of generative models.
Continual image restoration suffers from catastrophic forgetting of previous tasks, while existing solutions often require backbone modifications or incur substantial computational overhead. Method: This paper proposes a lightweight convolutional layer enhancement method that operates without altering the backbone network. Its core innovation is a dynamic filter parameter generation mechanism based on a shared knowledge base, which decomposes convolutional weights into task-invariant bases and task-specific increments, updated efficiently via a lightweight adaptation module. Contribution/Results: The method enables dynamic injection of new-task parameters without significantly increasing inference latency, preserving performance on historical tasks while adapting effectively to new ones. Experiments on multiple continual image restoration benchmarks demonstrate substantial mitigation of catastrophic forgetting: average PSNR improvements of 1.2–2.3 dB on new tasks, with inference speed nearly identical to the original model.
The mechanistic principles and theoretical limits of masked pretraining in multimodal representation learning remain poorly understood. To address this, we propose Randomly Random Mask Autoencoding (R²MAE), which dynamically randomizes the masking ratio during pretraining—departing from conventional fixed-ratio schemes—to compel models to learn multiscale features. Leveraging minimum-norm regression theory in high-dimensional linear models, we systematically characterize the behavior of masked autoencoding across diverse modalities—including language, vision, DNA sequences, and single-cell data—and validate its architectural generality across MLPs, CNNs, and Transformers. Extensive experiments demonstrate that R²MAE consistently outperforms standard and state-of-the-art masking strategies on cross-modal downstream tasks, yielding substantial gains in representation quality and generalization. Our work establishes the first unified theoretical framework for masked pretraining and introduces a scalable, principled paradigm for multimodal self-supervised learning.