adaptive token merging

Designs and implements algorithms and model components that dynamically merge, prune, or compress discrete tokens inside token-based neural architectures (e.g., transformer-style models), producing a reduced token sequence in intermediate layers or at inference time. These methods adapt per-input or per-layer token selection and merging (including object- or group-aware schemes) to lower attention complexity and compute cost while preserving semantically important content and maintaining task accuracy.

adaptivetokenmerging

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.5
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Efficient Token Compression for Vision Transformer with Spatial Information Preserved

Mar 30, 2025
JM
Junzhu Mao
🏛️ Nanjing University of Science and Technology | Beihang University | Terminus Group

To address the high computational and memory overhead of Vision Transformers—hindering their deployment in resource-constrained scenarios—this paper proposes a hierarchical “prune-merge” token compression framework. Our method introduces three key innovations: (1) a gradient-weighted attention scoring mechanism that dynamically evaluates token importance during training; (2) learnable merge/reconstruction matrices coupled with residual connections, enabling structured reconstruction of pruned tokens; and (3) end-to-end joint optimization guided by global gradient sensitivity, automatically discovering optimal compression architectures. Evaluated on ImageNet-1K, our approach accelerates DeiT-Small inference by 1.64× with only a 0.2% top-1 accuracy drop. On ADE20K semantic segmentation, it significantly outperforms existing token compression methods. The framework achieves efficient yet accurate vision modeling without architectural modification, offering a principled pathway toward lightweight ViT deployment.

Enables efficient deployment in resource-limited environmentsPreserves spatial information during token compressionReduces computational and memory needs for vision transformers

Token Reduction Should Go Beyond Efficiency in Generative Models -- From Vision, Language to Multimodality

May 23, 2025
ZK
Zhenglun Kong
🏛️ Harvard University | Northeastern University | Chinese Academy of Sciences | Wuhan University | Peking University | Massachusetts Institute of Technology

Large language and multimodal models suffer from pervasive hallucinations, long-range inconsistency, weak cross-modal alignment, and training instability. Method: This paper pioneers token reduction as a foundational principle of generative modeling—beyond mere computational efficiency—and introduces four mechanisms: (1) dynamic token pruning and reweighting, (2) cross-modal alignment constraints, (3) reinforcement learning–guided token optimization, and (4) context-aware scheduling. Contribution/Results: Theoretical analysis and empirical evaluation demonstrate that this paradigm significantly improves multimodal semantic consistency and reasoning robustness: hallucination rates decrease by over 35%, effective context window length increases by 40%, and training stability is markedly enhanced. The framework provides principled foundations and a systematic methodology for building lightweight, trustworthy, and long-context-controllable generative architectures.

Token reduction enhances multimodal integration and alignmentToken reduction mitigates hallucinations and maintains input coherenceToken reduction should transcend efficiency in generative models

Retrofitting (Large) Language Models with Dynamic Tokenization

Nov 27, 2024
DF
Darius Feher
🏛️ University of Cambridge

Existing pre-trained language models rely on static subword tokenizers, leading to suboptimal multilingual efficiency and imbalanced cross-lingual performance. To address this, we propose the first dynamic tokenization framework tailored for pre-trained language models: it dynamically identifies high-frequency subword sequences in each input and merges them on-the-fly; a lightweight hypernetwork then instantaneously generates token embeddings, enabling input-adaptive subword boundary decisions. Our method integrates a BPE-inspired intra-batch merging algorithm and is compatible with both encoder (e.g., XLM-R) and decoder (e.g., Mistral-7B) architectures. Evaluated across 14 languages, it achieves over 20% average sequence length reduction for XLM-R with less than 2% performance degradation; for English decoding, it shortens sequences by 6%, accelerates inference, and significantly improves multilingual fairness.

Dynamic tokenization improves inference speed and fairnessMethod reduces token sequence lengths without major performance lossStatic tokenizers reduce efficiency and language capabilities

Efficient Time Series Processing for Transformers and State-Space Models through Token Merging

May 28, 2024
LG
Leon Götz
🏛️ Volkswagen AG | Technical University of Munich

To address the high computational overhead of Transformers and state-space models (SSMs) when modeling long time series, this paper introduces token merging—previously unexplored in time-series analysis—for the first time. We propose a domain-adapted local merging paradigm that enforces subsequence neighborhood constraints, employs linear weighted aggregation, and integrates lightweight attention/SSM modules to jointly preserve local dependency modeling and computational efficiency. Our method achieves up to 5400% inference speedup on state-of-the-art time-series foundation models (e.g., Chronos), with negligible accuracy degradation. It demonstrates robust performance across diverse architectures and benchmark datasets, significantly improving throughput for long sequences. This work establishes a novel pathway toward efficient deployment of large-scale time-series models.

Developing local merging for scalable and causal token reductionEfficient processing of long token sequences in time series analysisPredicting merging benefits via spectral properties without task evaluation

Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.

Compression-optimized tokenization benefits low-resource languagesImproved performance in multilingual NLP tasksOptimal BPE configuration reduces token count

Latest Papers

What's happening recently
View more

Existing token pruning methods for vision Transformers struggle to generalize effectively across diverse tasks such as image classification, semantic segmentation, and object detection. This work addresses this limitation by first revealing through probing experiments the varying sensitivity of different tasks to pruning strategies. Building on this insight, we propose Task-Adaptive Pruning (TAP), which introduces task-specific registers to dynamically guide token pruning and feature restoration at each layer. TAP jointly optimizes the pruning criterion, depth budget allocation, and restoration scale in a task-aware manner. Under a token retention ratio of ρ = 0.5 and without compromising ImageNet-1K classification accuracy, TAP achieves 47.0 mIoU on ADE20K (1.30× encoder throughput gain) and 53.7 box AP on COCO (1.32× throughput gain).

multi-task transferspatial demandstask adaptation

This work addresses the severe inference latency in high-resolution multimodal large language models caused by the explosion in visual token count. Existing pruning methods rely on iterative optimization, hindering efficient acceleration. To overcome this limitation, the authors propose SFPruner (Single-Forward Pruner), which, for the first time, integrates structured redundancy modeling into visual token scoring. SFPruner employs a semantics-guided ridge leverage mechanism to suppress covariance-dominant directions and incorporates a ranking-based directional mask to enable asymmetric similarity competition. This enables non-iterative pruning in a single forward pass, effectively balancing semantic diversity and instruction relevance. Evaluated on Qwen2.5-VL, the method reduces the selection time for 512 tokens from 112.4 ms to 2.5 ms, achieving substantial inference speedup while maintaining performance comparable to state-of-the-art approaches.

inference latencymultimodal large language modelsstructured redundancy

This work addresses the high computational and memory costs incurred by large language models when processing long prompts, stemming from the quadratic complexity of self-attention. While existing compression methods operate solely in token space and overlook redundancy in the embedding space, this paper introduces K-Token Merging—a novel framework that, for the first time, merges every K consecutive tokens into a single embedding within the latent embedding space via a lightweight encoder. The compressed representation is then processed by a LoRA-finetuned large language model, while generation still employs the original vocabulary. By transcending conventional token-space compression, the method achieves highly efficient input-length reduction with minimal performance degradation. It establishes a Pareto frontier between compression ratio and task performance on Textualized Tree, Amazon Reviews, and CommitPackFT benchmarks, attaining up to 75% compression with negligible loss in accuracy.

computational efficiencyinput length reductionlarge language models

Hot Scholars

WC

Wenhao Chai

Princeton University
Machine LearningComputer Vision
YL

Yixuan Liu

AMD, Tsinghua University
Generative AI
DL

Ding Liang

vast, tsinghua university
3d generation
YP

Yan-Pei Cao

VAST
Computer Graphics3D Computer Vision
OA

Omar Alhussein

Khalifa University
Networking and AINetwork OptimizationEdge IntelligenceQuantum Computing