Text-Centric Post-Training for Omni-Modal Reasoning

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive inference costs of omni-modal models in audio-visual reasoning and the challenge of decoupled local optimization between perception and reasoning objectives. We propose a text-centric post-training paradigm that first strengthens reasoning capabilities using purely textual data through supervised fine-tuning, reinforcement learning, and synthetic data, followed by lightweight native audio-visual reinforcement learning to refine perceptual alignment. This multi-stage strategy enables efficient training via text-dominant pretraining coupled with audio-visual refinement. Experimental results demonstrate that our approach improves reasoning scores by 25.83% while reducing GPU computation time by 56.6%. Furthermore, under extremely limited resources, it effectively restores perceptual capabilities while retaining 93.5% of the reasoning gains, offering a promising new pathway for efficient multimodal reasoning.
📝 Abstract
Improving joint audio-visual reasoning in Omni Large Language Models typically incurs substantial data construction and training costs. Our diagnostics reveal multi-hop reasoning difficulties despite correct answers to all corresponding single-hop questions and suggest partial decoupling in the local optimization of perception and reasoning objectives. This motivates post-training with different emphases on these capabilities. Text-only reasoning training yields gains across data sources, model scales, and families. With the best-performing text-only configuration, supervised fine-tuning followed by reinforcement learning (RL) raises Qwen2.5-Omni-7B's geometric mean of nine reasoning scores by 25.83% over the base model, outperforming the complete native audio-visual route with 56.6% fewer GPU-hours. Training on data synthesized entirely by a text-only LLM raises this geometric mean by 21.01% without audio-visual data in construction or training. However, text-only training degrades perception. We therefore propose a text-centric post-training paradigm: text-only training provides the main reasoning optimization, and reduced-data native audio-visual RL then refines perception. Refinement uses about 90% fewer input tokens than full-data audio-visual RL, restores perception above the base level, and retains 93.5% of the best-performing text-only pipeline's reasoning gain.
Problem

Research questions and friction points this paper is trying to address.

Omni-modal reasoning
audio-visual reasoning
multi-hop reasoning
perception-reasoning decoupling
training efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Omni-modal reasoning
Text-centric post-training
Reinforcement learning
Perception-reasoning decoupling
Multi-hop reasoning
🔎 Similar Papers
No similar papers found.