🤖 AI Summary
This study addresses the prohibitive prefilling costs and the neglect of dynamic temporal semantics when omni-modal large models process long audio-video sequences. To this end, we propose a training-free, two-stage compression framework. Specifically, this work introduces a pioneering temporal evidence-guided budget allocation mechanism that dynamically distributes computational resources through master-slave modality co-calibration. Furthermore, by integrating spatiotemporal grouped selection with query-guided attention merging, the proposed method achieves efficient token compression while preserving local continuity. Extensive evaluations across four benchmarks demonstrate that our approach yields superior trade-offs between inference efficiency and performance compared to existing baselines.
📝 Abstract
Omnimodal large language models (Omni-LLMs) encode audio and visual streams into temporally interleaved token sequences for multimodal reasoning. However, processing long audio-visual token sequences incurs substantial prefill costs. Existing compression methods have made progress, but often overlook temporal changes in audio-visual semantic relevance. Motivated by temporal variation and local continuity, we propose OmniRoute, a training-free, two-stage compression framework. First, Temporal Evidence-Guided Budgeting (TEGB) derives chunk-wise modality preferences and initial leading-modality budgets from semantic relevance and local content variation. Second, Budget-Constrained Semantic Compression (BCSC) compresses the leading modality and then calibrates the follower's retention target using the actual retained fraction. For video, it combines spatiotemporal grouping with query-guided selection; for audio, it selects tokens based on encoder attention and query relevance, then merges residual tokens into context anchors under visual guidance. Experiments on four representative benchmarks demonstrate a better trade-off between inference efficiency and performance than competitive baselines. The code and interface will be released to facilitate further research.