Score
Design, build, and analyze methods that compute per-instance masks over model tokens to dynamically drop, replace, or sparsify tokens during encoding or intermediate processing. This includes learning mask tokens or embeddings to fill masked positions, producing binary or continuous keep-masks for transmission or later decoding, and balancing bitrate/compute reduction against reconstruction or downstream-task quality.
To address input noise and pretraining-finetuning mismatch caused by masked modeling in vision transformer pretraining, this paper proposes MaPeT. Methodologically, MaPeT jointly models structural dependencies among image patches via autoregressive masking and random block permutation—eliminating distributional shift induced by conventional random masking. It further introduces auxiliary positional embeddings to mitigate positional information inconsistency between pretraining and finetuning. Additionally, we design a k-CLIP visual tokenizer that maps image patches to discrete CLIP-aligned semantic tokens. Experiments demonstrate that MaPeT achieves state-of-the-art (SOTA) performance on ImageNet among models of comparable parameter count. The code and pretrained models are publicly released.
In masked image modeling (MIM) pretraining of Vision Transformers (ViTs), tokenization and local masking induce spatially inconsistent reconstruction supervision, degrading representation discriminability. This work is the first to systematically identify and address this spatial inconsistency issue, proposing Dynamic Token Morphing (DTM): a context-aware, dynamic token aggregation mechanism that generates spatially coherent reconstruction targets. DTM introduces no additional parameters or computational overhead and is plug-and-play across diverse MIM frameworks. On ImageNet-1K and ADE20K, DTM achieves significant gains over state-of-the-art MIM methods—yielding lower training loss and more stable convergence. When transferred to downstream tasks such as iNaturalist, it delivers consistent performance improvements. The core contribution is the first lightweight, parameter-free, and framework-agnostic solution specifically designed to resolve the spatial inconsistency problem in MIM.
This work proposes TokenMask, a novel segmentation framework that departs from conventional query-based Vision Transformer approaches which rely on explicit reconstruction of image-space feature maps—a process that incurs substantial computational redundancy and hinders deployment. Instead, TokenMask operates entirely in the query token space, generating mask logits directly through token affinity and performing interpolation in logit space. By integrating a ViT backbone, a token-space mask head, and TensorRT FP16 inference, the method significantly reduces both computational and memory overhead across multiple datasets and segmentation tasks while preserving accuracy. Notably, it achieves substantial acceleration on the Jetson AGX Orin platform, offering an efficient and streamlined architecture well-suited for embedded vision applications.
Discrete visual generative models have long underperformed continuous counterparts due to limitations in codebook size and insufficient compression. This work proposes BAR (masked Bit AutoRegressive), a scalable autoregressive framework that decomposes discrete tokens into binary bit sequences and employs bit-wise masked modeling to enable efficient training and sampling with arbitrarily large codebooks. BAR overcomes the longstanding scalability and computational bottlenecks of discrete generative models, achieving a state-of-the-art 0.99 gFID on ImageNet-256—surpassing both leading continuous and discrete approaches—while significantly accelerating convergence and reducing sampling cost.
Masked Diffusion Models (MDMs) suffer from redundant computation in discrete sequence generation due to binary masking, which causes tokens to remain unchanged across many sampling steps. To address this, we propose Partial Masking (Prime), the first framework to introduce continuous-interpolated intermediate mask states into discrete diffusion, enabling token-level fine-grained denoising and overcoming the rigid all-or-nothing masking paradigm. Methodologically, we formulate a variational training objective and design a dedicated architecture that eliminates reliance on autoregressive structures. Experiments demonstrate state-of-the-art performance: perplexity of 15.36 on OpenWebText for text generation, and FID scores of 3.26 (CIFAR-10) and 6.98 (ImageNet-32) for image generation—surpassing existing MDMs and hybrid models. Our core contribution is a differentiable intermediate masking mechanism that unifies discrete token representation with continuous denoising dynamics.
Existing image compression methods struggle to simultaneously satisfy the requirements of human vision and diverse machine vision tasks within a single model, and they often lack dynamic adaptation to the semantic importance and complexity of different image regions. To address this, this work proposes MoECodec—a token-aware Mixture-of-Experts image compression framework that replaces conventional feed-forward network (FFN) layers in a Transformer architecture with a dynamic, token-level expert mixture mechanism. The framework employs a content- and task-aware routing strategy, stabilizes expert assignment via spatial total variation regularization, and introduces a lightweight Group Shuffle MLP as the expert structure to enable efficient and coherent computational resource allocation. Experiments demonstrate that MoECodec significantly outperforms existing approaches in both image reconstruction quality and performance across multiple downstream machine vision tasks, confirming the effectiveness and generalization capability of a unified multi-task compression model.
This study addresses the unpredictable robustness of CLIP under mask pruning by identifying "spurious inversion" as the critical factor underlying unstable masking performance. To this end, we introduce the Spurious Inversion Metric (SIM) and propose a label-free, pre-deployment diagnostic framework that integrates semantic masking, text similarity analysis, and asynchronous GPU batch partitioning for efficient evaluation. Experimental results demonstrate that SIM significantly predicts masking efficacy, enabling optimized models to match or surpass baseline performance. This work thereby offers a reliable solution for the robust deployment of compressed CLIP architectures.
本文针对JEPA训练效率低的问题,提出了一种名为M-JEPA的执行架构,通过分离与掩码相关的计算和路由,减少了计算、内存流量和同步开销,提高了训练速度。
Existing token pruning methods for vision Transformers struggle to generalize effectively across diverse tasks such as image classification, semantic segmentation, and object detection. This work addresses this limitation by first revealing through probing experiments the varying sensitivity of different tasks to pruning strategies. Building on this insight, we propose Task-Adaptive Pruning (TAP), which introduces task-specific registers to dynamically guide token pruning and feature restoration at each layer. TAP jointly optimizes the pruning criterion, depth budget allocation, and restoration scale in a task-aware manner. Under a token retention ratio of ρ = 0.5 and without compromising ImageNet-1K classification accuracy, TAP achieves 47.0 mIoU on ADE20K (1.30× encoder throughput gain) and 53.7 box AP on COCO (1.32× throughput gain).
This study addresses the severe storage and computational redundancy caused by independently duplicated expert weights when upgrading dense models to Mixture-of-Experts (MoE) architectures. To this end, it proposes MASKerade, which for the first time defines experts as learned subnetworks over a frozen pretrained FFN. Specifically, experts are instantiated via learned sparse binary masks and dynamically composed through token-level routing. By jointly optimizing the routing and mask parameters, the method supports various sparsity patterns, including 2:4 semi-structured sparsity, without introducing separate expert parameter matrices. Evaluated on Qwen and Gemma backbones, MASKerade achieves MoE-equivalent capacity at the arithmetic cost of a single dense forward pass. Furthermore, it outperforms baselines across five vision-language benchmarks, demonstrating both the efficiency and practicality of this mask-learning paradigm.