Score
Designs and implements algorithms and model components that dynamically merge, prune, or compress discrete tokens inside token-based neural architectures (e.g., transformer-style models), producing a reduced token sequence in intermediate layers or at inference time. These methods adapt per-input or per-layer token selection and merging (including object- or group-aware schemes) to lower attention complexity and compute cost while preserving semantically important content and maintaining task accuracy.
Existing token compression methods for Vision Transformers (ViTs) in edge AI lack a systematic taxonomy and cross-architecture evaluation. Method: This work proposes the first unified classification framework, organizing approaches by core strategy (pruning, merging, hybrid) and deployment paradigm (fine-tuning vs. plug-in). We conduct empirical evaluation across standard ViTs (ViT-B/L) and compact variants (e.g., MobileViT, LeViT), enabling the first rigorous assessment of token compression on lightweight architectures. Contribution/Results: Results show that token compression significantly improves inference efficiency on standard ViTs but yields sharply diminished gains on compact models—revealing a strong structural coupling between compression efficacy and underlying architecture. This finding challenges the assumption of method portability across ViT families and provides critical theoretical insights and practical guidelines for designing efficient ViTs tailored to resource-constrained edge devices.
To address the high computational and memory overhead of Vision Transformers—hindering their deployment in resource-constrained scenarios—this paper proposes a hierarchical “prune-merge” token compression framework. Our method introduces three key innovations: (1) a gradient-weighted attention scoring mechanism that dynamically evaluates token importance during training; (2) learnable merge/reconstruction matrices coupled with residual connections, enabling structured reconstruction of pruned tokens; and (3) end-to-end joint optimization guided by global gradient sensitivity, automatically discovering optimal compression architectures. Evaluated on ImageNet-1K, our approach accelerates DeiT-Small inference by 1.64× with only a 0.2% top-1 accuracy drop. On ADE20K semantic segmentation, it significantly outperforms existing token compression methods. The framework achieves efficient yet accurate vision modeling without architectural modification, offering a principled pathway toward lightweight ViT deployment.
Large language and multimodal models suffer from pervasive hallucinations, long-range inconsistency, weak cross-modal alignment, and training instability. Method: This paper pioneers token reduction as a foundational principle of generative modeling—beyond mere computational efficiency—and introduces four mechanisms: (1) dynamic token pruning and reweighting, (2) cross-modal alignment constraints, (3) reinforcement learning–guided token optimization, and (4) context-aware scheduling. Contribution/Results: Theoretical analysis and empirical evaluation demonstrate that this paradigm significantly improves multimodal semantic consistency and reasoning robustness: hallucination rates decrease by over 35%, effective context window length increases by 40%, and training stability is markedly enhanced. The framework provides principled foundations and a systematic methodology for building lightweight, trustworthy, and long-context-controllable generative architectures.
Existing pre-trained language models rely on static subword tokenizers, leading to suboptimal multilingual efficiency and imbalanced cross-lingual performance. To address this, we propose the first dynamic tokenization framework tailored for pre-trained language models: it dynamically identifies high-frequency subword sequences in each input and merges them on-the-fly; a lightweight hypernetwork then instantaneously generates token embeddings, enabling input-adaptive subword boundary decisions. Our method integrates a BPE-inspired intra-batch merging algorithm and is compatible with both encoder (e.g., XLM-R) and decoder (e.g., Mistral-7B) architectures. Evaluated across 14 languages, it achieves over 20% average sequence length reduction for XLM-R with less than 2% performance degradation; for English decoding, it shortens sequences by 6%, accelerates inference, and significantly improves multilingual fairness.
To address the high computational overhead of Transformers and state-space models (SSMs) when modeling long time series, this paper introduces token merging—previously unexplored in time-series analysis—for the first time. We propose a domain-adapted local merging paradigm that enforces subsequence neighborhood constraints, employs linear weighted aggregation, and integrates lightweight attention/SSM modules to jointly preserve local dependency modeling and computational efficiency. Our method achieves up to 5400% inference speedup on state-of-the-art time-series foundation models (e.g., Chronos), with negligible accuracy degradation. It demonstrates robust performance across diverse architectures and benchmark datasets, significantly improving throughput for long sequences. This work establishes a novel pathway toward efficient deployment of large-scale time-series models.
Traditional Byte-Pair Encoding (BPE) tokenization introduces token redundancy in low-resource languages, degrading the performance of small-scale models. Method: This paper proposes a BPE configuration method integrating hyperparameter optimization and compressed sensing. It systematically searches key BPE hyperparameters—including vocabulary size and merge iterations—and jointly evaluates configurations using intrinsic metrics (e.g., token count) and extrinsic task performance (generation and classification). Contribution/Results: The study provides the first empirical evidence that BPE configuration significantly impacts multilingual modeling for low-resource languages. Experiments across diverse languages and model scales show that optimal configurations reduce token counts by 12.7% on average and improve downstream task accuracy by 1.8–3.4 percentage points for small models. These gains substantially enhance modeling efficiency and generalization capability in low-resource settings.
本文提出了一种通过合并模块压缩输入序列的方法,以减少Transformer模型的计算成本,同时保持准确性。
Existing token pruning methods for vision Transformers struggle to generalize effectively across diverse tasks such as image classification, semantic segmentation, and object detection. This work addresses this limitation by first revealing through probing experiments the varying sensitivity of different tasks to pruning strategies. Building on this insight, we propose Task-Adaptive Pruning (TAP), which introduces task-specific registers to dynamically guide token pruning and feature restoration at each layer. TAP jointly optimizes the pruning criterion, depth budget allocation, and restoration scale in a task-aware manner. Under a token retention ratio of ρ = 0.5 and without compromising ImageNet-1K classification accuracy, TAP achieves 47.0 mIoU on ADE20K (1.30× encoder throughput gain) and 53.7 box AP on COCO (1.32× throughput gain).
本文针对视觉令牌修剪问题,提出CoRePrune框架,通过两阶段方法考虑删除上下文和表示深度对令牌可移除性的影响,有效保持模型性能。
This work addresses the severe inference latency in high-resolution multimodal large language models caused by the explosion in visual token count. Existing pruning methods rely on iterative optimization, hindering efficient acceleration. To overcome this limitation, the authors propose SFPruner (Single-Forward Pruner), which, for the first time, integrates structured redundancy modeling into visual token scoring. SFPruner employs a semantics-guided ridge leverage mechanism to suppress covariance-dominant directions and incorporates a ranking-based directional mask to enable asymmetric similarity competition. This enables non-iterative pruning in a single forward pass, effectively balancing semantic diversity and instruction relevance. Evaluated on Qwen2.5-VL, the method reduces the selection time for 512 tokens from 112.4 ms to 2.5 ms, achieving substantial inference speedup while maintaining performance comparable to state-of-the-art approaches.
This work addresses the high computational and memory costs incurred by large language models when processing long prompts, stemming from the quadratic complexity of self-attention. While existing compression methods operate solely in token space and overlook redundancy in the embedding space, this paper introduces K-Token Merging—a novel framework that, for the first time, merges every K consecutive tokens into a single embedding within the latent embedding space via a lightweight encoder. The compressed representation is then processed by a LoRA-finetuned large language model, while generation still employs the original vocabulary. By transcending conventional token-space compression, the method achieves highly efficient input-length reduction with minimal performance degradation. It establishes a Pareto frontier between compression ratio and task performance on Textualized Tree, Amazon Reviews, and CommitPackFT benchmarks, attaining up to 75% compression with negligible loss in accuracy.