alignment-aware token sampling

Designs and implements token-selection and attention-sparsification methods that choose subsets of tokens or attention connections across scales while preserving spatial or positional correspondence. Builds scale-aligned sampling masks and algorithms (e.g., mask-based alignment, token-subset alignment) and analyzes trade-offs between reduced token count, computational cost, and alignment fidelity for tasks requiring maintained spatial correspondence.

alignment-awaretokensampling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the instability in representation alignment during diffusion model training, which arises from the mismatch between noisy inputs and clean image features, leading models to over-rely on complete token sets. To mitigate this alignment discrepancy, the authors propose MaskAlign, the first approach that operates from the perspective of token subsets. MaskAlign dynamically aligns representations using randomly masked token subsets and introduces a lightweight pre-mask token mixing module to encourage cross-token information sharing prior to masking. Integrated with a self-supervised visual encoder and diffusion Transformer training, MaskAlign significantly enhances generation quality and alignment robustness while maintaining computational efficiency, thereby improving the model’s generalization under perturbations of varying token subsets.

clean-image featuresdiffusion modelsnoisy inputs

This work addresses the limitation of existing linear attention methods, which rely on predefined spatial layouts to compress image tokens, thereby constraining information aggregation to coordinate positions rather than semantic content. To overcome this, the paper introduces Representative Attention (RPAttention), a novel representation-driven token compression mechanism that dynamically generates semantic representative tokens to enable spatially agnostic global interactions. RPAttention adopts a lightweight Gather-Interact-Distribute paradigm, integrating competitive similarity-based routing, interaction among representative tokens in a compact latent space, and query-driven cross-attention. This design maintains linear computational complexity while substantially enhancing semantic alignment. Experimental results demonstrate consistent and significant performance gains across image classification, object detection, and semantic segmentation tasks.

global attentionlinear attentionsemantic organization

SPoT: Subpixel Placement of Tokens in Vision Transformers

Jul 02, 2025
MH
Martine Hjelkrem-Tan
🏛️ University of Oslo

Traditional vision Transformers rely on discrete image patching, which constrains the effective deployment of sparsity mechanisms and hinders simultaneous optimization of accuracy and efficiency. To address this, we propose subpixel tokenization—a novel, differentiable tokenization scheme that enables dynamic, subpixel-precise token placement in continuous image space, thereby breaking free from rigid discrete grid constraints. Our approach integrates an oracle-guided search mechanism to adaptively optimize token spatial distribution, enhancing representational capacity while preserving sparsity. Experiments demonstrate that our method significantly reduces the number of tokens required during inference—by 30–50% on average—while improving classification and detection accuracy. It achieves superior accuracy–computation trade-offs on standard benchmarks including ImageNet and COCO. Moreover, the resulting models exhibit enhanced interpretability and architectural flexibility, offering a principled pathway toward efficient, high-fidelity vision representation learning.

Enhances sparsity as a strategic advantageOvercomes grid constraints in Vision TransformersReduces tokens needed for accurate predictions

This study addresses the prohibitive computational cost of high-resolution tokens in vision-language models by introducing foveal compression, which for the first time incorporates the human foveal mechanism into token compression. The proposed method interleaves native- and compressed-resolution tokens, employing a self-distillation merger for feature alignment and a lightweight selector to dynamically determine high-fidelity preservation strategies for individual spatial units under a fixed budget. Experiments demonstrate that this approach matches baseline performance under extremely tight budgets and outperforms random allocation under moderate budgets. Furthermore, the analysis reveals a complementary bottleneck between regional selection and compression fidelity, indicating that localized high-fidelity preservation is not universally optimal.

Foveated CompressionRegion SelectionToken Budget

Latest Papers

What's happening recently
View more

TrimTokenator-LC: Towards Adaptive Visual Token Pruning for Large Multimodal Models with Long Contexts

Dec 27, 2025
HZ
Hao Zhang
🏛️ Beijing Academy of Artificial Intelligence (BAAI)

To address the prohibitive inference cost of large vision-language models (VLMs) under long-context, multi-image inputs, this paper proposes an adaptive visual token pruning method. Unlike prior approaches, it uniquely decouples redundancy modeling into two orthogonal dimensions: intra-image diversity and inter-image dissimilarity, and introduces a Pareto-optimal selection mechanism to jointly optimize both against text alignment. The proposed two-stage framework—comprising global candidate pool construction, diversity quantification, and greedy subset selection—dynamically allocates token budgets and identifies the most representative visual tokens without fine-tuning or supervision. Evaluated on multi-image long-context benchmarks, our method reduces visual tokens by up to 67% while maintaining or even improving performance on VQA and image captioning tasks. This achieves content-aware, cross-image cooperative inference with significant efficiency gains.

Adaptive visual token pruning for long multimodal contextsBalancing token diversity with text alignment efficientlyReducing inference cost in multi-image LMM scenarios

This work proposes TokenMask, a novel segmentation framework that departs from conventional query-based Vision Transformer approaches which rely on explicit reconstruction of image-space feature maps—a process that incurs substantial computational redundancy and hinders deployment. Instead, TokenMask operates entirely in the query token space, generating mask logits directly through token affinity and performing interpolation in logit space. By integrating a ViT backbone, a token-space mask head, and TensorRT FP16 inference, the method significantly reduces both computational and memory overhead across multiple datasets and segmentation tasks while preserving accuracy. Notably, it achieves substantial acceleration on the Jetson AGX Orin platform, offering an efficient and streamlined architecture well-suited for embedded vision applications.

efficient segmentationembedded visionmask prediction

This work addresses the severe inference latency in high-resolution multimodal large language models caused by the explosion in visual token count. Existing pruning methods rely on iterative optimization, hindering efficient acceleration. To overcome this limitation, the authors propose SFPruner (Single-Forward Pruner), which, for the first time, integrates structured redundancy modeling into visual token scoring. SFPruner employs a semantics-guided ridge leverage mechanism to suppress covariance-dominant directions and incorporates a ranking-based directional mask to enable asymmetric similarity competition. This enables non-iterative pruning in a single forward pass, effectively balancing semantic diversity and instruction relevance. Evaluated on Qwen2.5-VL, the method reduces the selection time for 512 tokens from 112.4 ms to 2.5 ms, achieving substantial inference speedup while maintaining performance comparable to state-of-the-art approaches.

inference latencymultimodal large language modelsstructured redundancy

VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm

Dec 02, 2025
ZW
Zhenkai Wu
🏛️ Zhejiang University | Huawei Noah's Ark Lab

Visual-language models (VLMs) suffer from high inference overhead due to excessive visual tokens, hindering mobile deployment. Existing pruning methods either rely solely on token importance while ignoring redundancy, or neglect spatial structure—resulting in sparse, discontinuous token retention and incomplete target coverage. This paper proposes a training-free, efficient token pruning framework. We introduce an *eccentric pruning paradigm* coupled with a *spatially sparse buffering criterion* to eliminate redundancy while preserving spatial continuity of target regions. Further, we integrate *importance-based parallel greedy selection* with a *salient-information fusion mechanism* for discarded tokens, jointly ensuring fine-grained central semantic fidelity and global contextual integrity. Evaluated on five mainstream VLMs, our method achieves an 88.9% token pruning rate—significantly outperforming state-of-the-art baselines—and delivers end-to-end inference acceleration.

Addresses redundancy and spatial sparsity in token pruning methodsPreserves fine-grained object details while achieving high pruning ratesReduces computational cost of vision-language models for mobile deployment

Hot Scholars

JT

Jacek Tabor

Profesor informatyki, Uniwersytet Jagielloński
mathematicscomputer science
ZD

Zicheng Duan

Ph.D.@ University of Adelaide; Former Leonardo.AI | ANU | CASIA
Computer VisionGenerative ModelsMultiview Detection
YG

Yuxian Gu

Tsinghua University
Natural Language Processing
AM

Anas Mahmoud

Applied Research Scientist @ Mila
Multimodal RepresentationsData & Model DistillationVision-Language Modelling