Score
Designs and implements masking strategies that determine which input elements to hide during self-supervised masked reconstruction tasks, including complementary and self-guided schemes that choose masks adaptively or jointly to cover information efficiently. Builds algorithms to compute and apply these masks (e.g., from gradients or model signals), integrate them into training pipelines, and evaluate their impact on masking ratios, sample efficiency, and training throughput.
Existing research on masked image modeling (MIM) for self-supervised visual representation learning lacks a unified formalization of pretraining paradigms and standardized, comparable evaluation. Method: We formally categorize MIM into two principal paradigms—reconstruction-based and contrastive-based—and construct an interpretable, hierarchical taxonomy via expert curation and agglomerative clustering. We conduct systematic, unified benchmarking of over 20 state-of-the-art models on ImageNet and other standard datasets. Contribution/Results: Our analysis identifies critical open challenges—including cross-paradigm integration, long-tailed masking strategies, and compute-accuracy trade-offs. To foster reproducibility and standardization, we publicly release a structured literature repository and an extensible evaluation framework on GitHub. This work establishes foundational infrastructure for rigorous, comparable advancement in MIM research.
This study addresses why masked prediction captures features overlooked by unmasked reconstruction and how standard data augmentation obscures the advantages of dynamic masking. To investigate this, we construct a high-dimensional theoretical framework that decouples masking objectives from diversity benefits, demonstrating that masked linear reconstruction remains effective even when PCA fails, while quantifying the impact of mask diversity on sample complexity. These theoretical findings are validated through MAE-based modeling and experiments with CNN/ViT and BERT architectures. Our contributions reveal the statistical advantages of mask resampling, confirm that dynamic masking outperforms static alternatives and substantially enhances downstream task performance, and provide both theoretical and empirical guidance for optimizing pretraining pipelines.
This work investigates how masking patterns affect the self-supervised pretraining performance of SparK, revealing that conventional random masking struggles to jointly model local details and global hierarchical structures. To address this, we propose Structured Mesh Masking: an image is partitioned into multi-scale grids, and tokens are masked in a hierarchical, coarse-to-fine manner—enabling joint optimization of sparsity and hierarchy. We integrate Mesh Masking into the SparK framework, jointly optimizing hierarchical feature reconstruction and contrastive representation learning. On ImageNet-1K linear evaluation, our method achieves a +1.8% top-1 accuracy gain over the baseline. To our knowledge, this is the first work to incorporate explicit grid-based geometric structure into masking design, demonstrating that mask geometry serves as a critical inductive bias for visual representation quality. Our approach establishes a new paradigm for sparse, hierarchical self-supervised learning.
Existing unsupervised domain adaptation (UDA) methods treat masked image modeling (MIM) merely as input perturbation, lacking theoretical grounding and thus limiting its potential for feature extraction and representation learning. To address this, we propose MaskTwins—a novel framework that, for the first time, reformulates MIM from the perspective of sparse signal recovery. We introduce complementary mask dualities and theoretically prove that they enhance domain-invariant feature learning and explicitly model cross-domain structural consistency. Our method employs a dual-branch network that jointly optimizes complementary mask reconstruction and feature alignment, enabling end-to-end UDA for semantic segmentation without requiring pretraining. Extensive experiments on both natural and biomedical image segmentation benchmarks demonstrate significant improvements over state-of-the-art UDA baselines, validating the generalizability and effectiveness of MaskTwins.
本文通过引入新的理论框架分析了掩码预训练(MPT)的工作机制,提出了均匀性增强的MPT损失(U-MPT)以解决维度坍缩问题,并提出了一种新的掩码策略来提高下游任务性能。
To address input noise and pretraining-finetuning mismatch caused by masked modeling in vision transformer pretraining, this paper proposes MaPeT. Methodologically, MaPeT jointly models structural dependencies among image patches via autoregressive masking and random block permutation—eliminating distributional shift induced by conventional random masking. It further introduces auxiliary positional embeddings to mitigate positional information inconsistency between pretraining and finetuning. Additionally, we design a k-CLIP visual tokenizer that maps image patches to discrete CLIP-aligned semantic tokens. Experiments demonstrate that MaPeT achieves state-of-the-art (SOTA) performance on ImageNet among models of comparable parameter count. The code and pretrained models are publicly released.
To address the limited performance gains of conventional self-supervised pretraining for semantic segmentation under constrained model capacity and computational resources, this paper proposes Selective Masking Self-supervision (SMS). SMS replaces random masking with a dynamic, error-driven strategy: it identifies hard sample regions during training by leveraging reconstruction errors and iteratively masks and reconstructs high-loss image patches, thereby aligning pretraining more closely with downstream segmentation challenges. A stepwise iterative reconstruction scheme enables improved segmentation accuracy—particularly for low-performing classes—without increasing inference overhead. On general-purpose and weed segmentation benchmarks, SMS achieves absolute mIoU improvements of +2.9% and +2.5%, respectively. These results demonstrate its effectiveness and generalizability for low-budget self-supervised pretraining in resource-constrained scenarios.
Existing text removal methods primarily target simple outdoor scenes and struggle with real-world images containing high-density, complex text layouts; moreover, their performance is highly sensitive to mask shape, necessitating costly manual parameter tuning. To address dense text images, this paper proposes an automated mask shape learning framework integrating deformable mask modeling and Bayesian optimization. First, we construct character-level deformable contour masks and empirically demonstrate that minimal covering masks are suboptimal, highlighting the critical role of fine-grained contour adjustment. Second, we formulate mask shape optimization as a black-box problem, using restoration quality as the feedback objective and leveraging Bayesian optimization to automatically determine optimal shape parameters. Experiments show significant improvements in restoration quality on high-density text images, empirically validating the existence of an optimal mask shape. Our approach delivers an interpretable, reusable, and fully automated solution for industrial-grade text removal.
Medical imaging suffers from scarce annotated data, leading to poor semantic alignment and weak generalization in self-supervised masked image modeling (MIM). To address this, we propose Text-guided Controllable Masking (TCM), a novel framework that leverages vision-language models to parse diagnostic text prompts and dynamically localize anatomically or pathologically salient regions. TCM performs region-aware masking at a low mask ratio (40%) and integrates contrastive learning—eliminating reliance on either supervised signals or reconstruction-based heuristics. By innovatively unifying prompt learning with MIM, TCM significantly enhances representation quality across diverse modalities: on brain MRI, chest CT, and pulmonary X-ray datasets, it improves classification accuracy by up to 3.1%, and boosts object detection performance by +1.3 BoxAP and +1.1 MaskAP. These results demonstrate TCM’s strong cross-modal and cross-task adaptability and generalization capability.
This study addresses the high sensitivity of Joint Embedding Predictive Architectures (JEPA) to masking strategies and the absence of theoretical explanations for the performance disparity between block-wise and scattered masking. By introducing a wavelet-basis linear measurement perspective, this work reveals that mask geometry effectively prevents representational collapse in the target encoder by preserving irrecoverable coarse-scale information. Validated through 151 pretraining runs, the proposed theory quantifies the mechanism of irrecoverable content, confirming that block-wise masking significantly outperforms random and stripe masking. Furthermore, we demonstrate that freezing the target encoder narrows the performance gap across masking strategies, whereas exposing partial targets substantially enhances representation fidelity.
Traditional supervised learning for image classification often discards masked pixels outright, leading to contextual information loss and degradation of fine-grained discriminative features. To address this, we propose a novel “mask-as-knowledge” paradigm that explicitly treats masked regions as semantically rich auxiliary supervision signals—rather than mere occlusions. Our method employs a dual-branch architecture: one branch processes visible pixels, while the other reconstructs masked regions; both branches are jointly optimized via classification loss and mask reconstruction loss, thereby enforcing local–global contextual consistency. This relearning mechanism is architecture-agnostic, seamlessly integrating with both CNNs and Transformers. Extensive experiments on multiple fine-grained visual recognition benchmarks demonstrate significant performance gains, validating the approach’s effectiveness in enhancing feature diversity and preserving discriminative details without requiring architectural modifications.