OmniRoute: Mapping Temporal Semantic Evidence to Audio-Visual Token Budgets for Efficient Omnimodal Large Language Models

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive prefilling costs and the neglect of dynamic temporal semantics when omni-modal large models process long audio-video sequences. To this end, we propose a training-free, two-stage compression framework. Specifically, this work introduces a pioneering temporal evidence-guided budget allocation mechanism that dynamically distributes computational resources through master-slave modality co-calibration. Furthermore, by integrating spatiotemporal grouped selection with query-guided attention merging, the proposed method achieves efficient token compression while preserving local continuity. Extensive evaluations across four benchmarks demonstrate that our approach yields superior trade-offs between inference efficiency and performance compared to existing baselines.
📝 Abstract
Omnimodal large language models (Omni-LLMs) encode audio and visual streams into temporally interleaved token sequences for multimodal reasoning. However, processing long audio-visual token sequences incurs substantial prefill costs. Existing compression methods have made progress, but often overlook temporal changes in audio-visual semantic relevance. Motivated by temporal variation and local continuity, we propose OmniRoute, a training-free, two-stage compression framework. First, Temporal Evidence-Guided Budgeting (TEGB) derives chunk-wise modality preferences and initial leading-modality budgets from semantic relevance and local content variation. Second, Budget-Constrained Semantic Compression (BCSC) compresses the leading modality and then calibrates the follower's retention target using the actual retained fraction. For video, it combines spatiotemporal grouping with query-guided selection; for audio, it selects tokens based on encoder attention and query relevance, then merges residual tokens into context anchors under visual guidance. Experiments on four representative benchmarks demonstrate a better trade-off between inference efficiency and performance than competitive baselines. The code and interface will be released to facilitate further research.
Problem

Research questions and friction points this paper is trying to address.

Omnimodal Large Language Models
Token Compression
Audio-Visual Reasoning
Inference Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Omnimodal Large Language Models
Training-free Compression
Temporal Semantic Evidence
Token Budgeting
Audio-Visual Reasoning
💼 Related Jobs
No related jobs found.
Y
Yuchen Deng
Shenzhen International Graduate School, Tsinghua University, China; Pengcheng Laboratory, China
Z
Zidang Cai
Shenzhen International Graduate School, Tsinghua University, China; Pengcheng Laboratory, China
F
Feidiao Yang
Pengcheng Laboratory, China
Y
Yufei Wang
Shenzhen International Graduate School, Tsinghua University, China; Pengcheng Laboratory, China
Jie Wang
Jie Wang
China Agricultural University
Ecological Remote SensingRemote sensing of agroecosystemsLand Use and Land Cover Change
H
Hai-Tao Zheng
Shenzhen International Graduate School, Tsinghua University, China; Pengcheng Laboratory, China
Yuxing Han
Yuxing Han
Tsinghua University
Smart AgricultureArtificial IntelligenceVideoCommunication