Score
Designs and implements token-selection and attention-sparsification methods that choose subsets of tokens or attention connections across scales while preserving spatial or positional correspondence. Builds scale-aligned sampling masks and algorithms (e.g., mask-based alignment, token-subset alignment) and analyzes trade-offs between reduced token count, computational cost, and alignment fidelity for tasks requiring maintained spatial correspondence.
This work addresses the instability in representation alignment during diffusion model training, which arises from the mismatch between noisy inputs and clean image features, leading models to over-rely on complete token sets. To mitigate this alignment discrepancy, the authors propose MaskAlign, the first approach that operates from the perspective of token subsets. MaskAlign dynamically aligns representations using randomly masked token subsets and introduces a lightweight pre-mask token mixing module to encourage cross-token information sharing prior to masking. Integrated with a self-supervised visual encoder and diffusion Transformer training, MaskAlign significantly enhances generation quality and alignment robustness while maintaining computational efficiency, thereby improving the model’s generalization under perturbations of varying token subsets.
This work addresses the limitation of existing linear attention methods, which rely on predefined spatial layouts to compress image tokens, thereby constraining information aggregation to coordinate positions rather than semantic content. To overcome this, the paper introduces Representative Attention (RPAttention), a novel representation-driven token compression mechanism that dynamically generates semantic representative tokens to enable spatially agnostic global interactions. RPAttention adopts a lightweight Gather-Interact-Distribute paradigm, integrating competitive similarity-based routing, interaction among representative tokens in a compact latent space, and query-driven cross-attention. This design maintains linear computational complexity while substantially enhancing semantic alignment. Experimental results demonstrate consistent and significant performance gains across image classification, object detection, and semantic segmentation tasks.
本文提出S^2Prune方法,通过保留图像的空间覆盖并根据局部结构调整视觉令牌密度,解决了多模态大语言模型推理开销大的问题。
Traditional vision Transformers rely on discrete image patching, which constrains the effective deployment of sparsity mechanisms and hinders simultaneous optimization of accuracy and efficiency. To address this, we propose subpixel tokenization—a novel, differentiable tokenization scheme that enables dynamic, subpixel-precise token placement in continuous image space, thereby breaking free from rigid discrete grid constraints. Our approach integrates an oracle-guided search mechanism to adaptively optimize token spatial distribution, enhancing representational capacity while preserving sparsity. Experiments demonstrate that our method significantly reduces the number of tokens required during inference—by 30–50% on average—while improving classification and detection accuracy. It achieves superior accuracy–computation trade-offs on standard benchmarks including ImageNet and COCO. Moreover, the resulting models exhibit enhanced interpretability and architectural flexibility, offering a principled pathway toward efficient, high-fidelity vision representation learning.
This study addresses the prohibitive computational cost of high-resolution tokens in vision-language models by introducing foveal compression, which for the first time incorporates the human foveal mechanism into token compression. The proposed method interleaves native- and compressed-resolution tokens, employing a self-distillation merger for feature alignment and a lightweight selector to dynamically determine high-fidelity preservation strategies for individual spatial units under a fixed budget. Experiments demonstrate that this approach matches baseline performance under extremely tight budgets and outperforms random allocation under moderate budgets. Furthermore, the analysis reveals a complementary bottleneck between regional selection and compression fidelity, indicating that localized high-fidelity preservation is not universally optimal.
To address the prohibitive inference cost of large vision-language models (VLMs) under long-context, multi-image inputs, this paper proposes an adaptive visual token pruning method. Unlike prior approaches, it uniquely decouples redundancy modeling into two orthogonal dimensions: intra-image diversity and inter-image dissimilarity, and introduces a Pareto-optimal selection mechanism to jointly optimize both against text alignment. The proposed two-stage framework—comprising global candidate pool construction, diversity quantification, and greedy subset selection—dynamically allocates token budgets and identifies the most representative visual tokens without fine-tuning or supervision. Evaluated on multi-image long-context benchmarks, our method reduces visual tokens by up to 67% while maintaining or even improving performance on VQA and image captioning tasks. This achieves content-aware, cross-image cooperative inference with significant efficiency gains.
为解决高分辨率图像在视觉-语言模型中产生的高昂预填充成本问题,提出基于PAQ的视觉令牌修剪方法,通过跨模态对齐更精准地保留重要信息。
This work proposes TokenMask, a novel segmentation framework that departs from conventional query-based Vision Transformer approaches which rely on explicit reconstruction of image-space feature maps—a process that incurs substantial computational redundancy and hinders deployment. Instead, TokenMask operates entirely in the query token space, generating mask logits directly through token affinity and performing interpolation in logit space. By integrating a ViT backbone, a token-space mask head, and TensorRT FP16 inference, the method significantly reduces both computational and memory overhead across multiple datasets and segmentation tasks while preserving accuracy. Notably, it achieves substantial acceleration on the Jetson AGX Orin platform, offering an efficient and streamlined architecture well-suited for embedded vision applications.
This work addresses the severe inference latency in high-resolution multimodal large language models caused by the explosion in visual token count. Existing pruning methods rely on iterative optimization, hindering efficient acceleration. To overcome this limitation, the authors propose SFPruner (Single-Forward Pruner), which, for the first time, integrates structured redundancy modeling into visual token scoring. SFPruner employs a semantics-guided ridge leverage mechanism to suppress covariance-dominant directions and incorporates a ranking-based directional mask to enable asymmetric similarity competition. This enables non-iterative pruning in a single forward pass, effectively balancing semantic diversity and instruction relevance. Evaluated on Qwen2.5-VL, the method reduces the selection time for 512 tokens from 112.4 ms to 2.5 ms, achieving substantial inference speedup while maintaining performance comparable to state-of-the-art approaches.
Visual-language models (VLMs) suffer from high inference overhead due to excessive visual tokens, hindering mobile deployment. Existing pruning methods either rely solely on token importance while ignoring redundancy, or neglect spatial structure—resulting in sparse, discontinuous token retention and incomplete target coverage. This paper proposes a training-free, efficient token pruning framework. We introduce an *eccentric pruning paradigm* coupled with a *spatially sparse buffering criterion* to eliminate redundancy while preserving spatial continuity of target regions. Further, we integrate *importance-based parallel greedy selection* with a *salient-information fusion mechanism* for discarded tokens, jointly ensuring fine-grained central semantic fidelity and global contextual integrity. Evaluated on five mainstream VLMs, our method achieves an 88.9% token pruning rate—significantly outperforming state-of-the-art baselines—and delivers end-to-end inference acceleration.