OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high memory and computational overhead faced by OmniLLMs when processing long audio-video sequences, a challenge exacerbated by existing compression methods that neglect cross-modal budget allocation. The authors propose OmniDelta, a training-free framework that dynamically optimizes token retention through intent-aware, skill-pool-driven cross-modal budgeting, coupled with intra-modal fine-grained reallocation based on local complexity and temporal redundancy—thereby overcoming fixed-budget constraints. Compatible with existing pruning strategies, OmniDelta achieves substantial efficiency gains across four audio-video benchmarks: for instance, Qwen2.5-Omni-7B retains only 25% of tokens while reducing GPU memory usage by 22.0% and accelerating end-to-end inference by 1.64×.
📝 Abstract
Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs. Existing compression methods mainly focus on selecting important tokens under fixed budgets, leaving the preceding budget-allocation problem underexplored. We show that direct query-to-audio/video similarity is unreliable for inter-modal budget allocation, and that uniform intra-modal budgets can miss key evidence while retaining redundant content. To address these limitations, we propose OmniDelta, a training-free, skill-driven framework that couples intent-aware inter-modal allocation with content-aware intra-modal allocation. OmniDelta first constructs audio and video skill pools to shift the fixed retained-token budget according to query demand, then reallocates modality budgets over audio segments and video frames using local complexity and temporal redundancy. The resulting local budgets can be combined with existing pruning strategies, preserving the total retained-token ratio while changing where the budget is spent. Experiments on four audio-video benchmarks with two Qwen2.5-Omni models show that OmniDelta establishes a new accuracy-efficiency Pareto frontier across pruning ratios. At 25% token retention on Qwen2.5-Omni-7B, OmniDelta reduces GPU memory by 22.0% and achieves a 1.64x end-to-end speedup over full-token inference.
Problem

Research questions and friction points this paper is trying to address.

token compression
budget allocation
Omni-modal LLMs
multi-modal inference
memory efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

token compression
budget allocation
OmniLLMs
skill-driven
multi-modal pruning
Haoyang Huang
Haoyang Huang
JD Explore Academy (present) | StepFun | Microsoft Research
Multimodal & Multilingual Foundation Model
Wenjie Huang
Wenjie Huang
Shanghai Jiao Tong University
点云压缩视频压缩图像压缩
Tianqi Xu
Tianqi Xu
Tokyo Institute of Technology
CloudBurst BufferDistributed File SystemUser-Level File SystemHPC
H
Hongyaoxing Gu
University of Chinese Academy of Sciences, Qwen Application, Alibaba
K
Kang Tan
Zhejiang University
Y
Yikai Fu
Zhejiang University
Y
Yuhao Shen
Zhejiang University, Qwen Application, Alibaba
Tianyu Liu
Tianyu Liu
Hongkong University of Science and Technology
B
Baolin Zhang
Qwen Application, Alibaba
J
Jun Zhang
Qwen Application, Alibaba
X
Xinyi Hu
Qwen Application, Alibaba
J
Jun Dai
Qwen Application, Alibaba
S
Shuang Ge
Qwen Application, Alibaba
L
Lei Chen
Qwen Application, Alibaba
Y
Yue Li
Qwen Application, Alibaba
M
Mingchen Wang
Qwen Application, Alibaba
M
Meng Zhang
Zhejiang University