Learning to Reason with Compressed Context: Ground-Truth-Free Adaptation of OmniLLMs via Self-Distillation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive computational costs of processing long sequences with omni-modal large models, where aggressive compression often incurs significant accuracy degradation due to the absence of effective mechanisms guiding model adaptation to compressed contexts. To this end, this work proposes the CAFD framework, introducing a novel ground-truth-free self-teacher distillation paradigm. By leveraging a full-context teacher model to supply privileged information, CAFD employs online policy trajectory supervision to guide a fixed-architecture student model in adapting to multimodal token compression pipelines. Extensive validation on Qwen2.5-Omni demonstrates that the proposed framework yields performance improvements in 120 out of 125 experimental conditions, achieving an average accuracy gain of 1.44 points and recovering 26.9% of compression-induced accuracy loss, thereby effectively resolving the trade-off between inference efficiency and predictive precision.
📝 Abstract
Omni-modal large language models (OmniLLMs) enable unified audio-video understanding, but their long multimodal token sequences make deployment computationally expensive. Token compression reduces this cost, yet aggressive compression often lowers accuracy. Existing works predominantly focus on designing better compression mechanisms; however, adapting the underlying language model to reason effectively over the remaining compressed context remains under-explored. To address this, we propose CAFD (Compressed-Context Adaptation via Full-Context Distillation), a ground-truth-free self-distillation framework that adapts OmniLLMs to fixed compression pipelines without requiring reference answers, rationales, or correctness rewards. CAFD leverages the full-token view of the same multimodal sample as a source of privileged information: a full-context self-teacher provides soft target supervision to a compressed-context student along the student's on-policy trajectory. Evaluated on Qwen2.5-Omni-7B across five audio-video benchmarks, five compression pipelines, and five deployment budgets, CAFD demonstrates consistent gains, improving 120 out of 125 conditions with an average accuracy boost of 1.44 points and recovering 26.9% of the accuracy gap on average. These results demonstrate that the proposed ground-truth-free adaptation offers an effective and practical route to improving the accuracy-efficiency trade-off in deployed OmniLLMs.
Problem

Research questions and friction points this paper is trying to address.

Omni-modal large language models
token compression
compressed context reasoning
accuracy-efficiency trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Distillation
Omni-modal LLMs
Token Compression
Ground-Truth-Free Adaptation
On-Policy Trajectory
💼 Related Jobs
No related jobs found.
Jianghao Wang
Jianghao Wang
School of Computer Science, Wuhan University
Ke Meng
Ke Meng
Didi Chuxing
J
Jian Li
Didi Chuxing
C
Chi Cheng
Didi Chuxing
L
Longyu Qi
Didi Chuxing
L
Liyin Liang
Didi Chuxing
Y
Yifeng Qian
Didi Chuxing
C
Chunbo Lai
Didi Chuxing
Y
Yutian Lin
School of Computer Science, Wuhan University
Zeyu Wang
Zeyu Wang
The Hong Kong University of Science and Technology (Guangzhou)
Computer GraphicsHuman-Computer InteractionCreative IntelligenceCultural Heritage