OP-CAD: On-Policy Clean-Audio Distillation for Robust Audio-Visual Reasoning

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of omni-modal models to interference during audio-visual understanding in noisy environments. To this end, we propose a curriculum learning-based privileged self-distillation framework. Specifically, a clean-audio teacher model guides student response generation, while selective token-level supervision is introduced by dynamically weighting critical positions through contrastive multi-context sensitivity. This approach effectively integrates online policy distillation with privileged information exploitation mechanisms. Experimental results demonstrate that the proposed method significantly outperforms baselines, robustly preserving answer accuracy under severe interference without compromising performance on clean data.
📝 Abstract
Omni-modal large language models deployed in real-world environments encounter external noise that can interfere with their perception and understanding of multimodal inputs. We study their robustness in audio-visual understanding, focusing on question answering under environmental noise and competing speech. The challenge is to resist acoustic interference while preserving useful audio evidence. On-policy distillation provides dense teacher feedback on student-generated responses, but uniform token weighting does not explicitly prioritize positions affected by acoustic interference. We introduce OP-CAD (On-Policy Clean-Audio Distillation), a curriculum-based privileged self-distillation framework for robust audio-visual understanding. Training progresses from mild to severe environmental noise and competing speech, with selective token-level supervision at each stage. The student generates responses from corrupted audio-visual input, while a frozen teacher uses clean audio and the verified answer to supervise the same response prefixes. To allocate this supervision, OP-CAD compares teacher predictions under clean, corrupted, and visual-only contexts without revealing the answer. These matched comparisons measure sensitivity to audio removal and corruption; a bounded weighting rule emphasizes positions identified by either signal while retaining supervision throughout the response. OP-CAD outperforms the compared methods across all evaluated noise conditions. Paired analyses further show improved preservation of clean-correct answers under strong interference, with no observed aggregate clean-accuracy penalty. These results demonstrate the value of directing clean-teacher supervision toward acoustically sensitive predictions for robust audio-visual reasoning.
Problem

Research questions and friction points this paper is trying to address.

audio-visual reasoning
robustness
environmental noise
competing speech
knowledge distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Audio-Visual Reasoning
Curriculum Learning
Token-level Supervision
Noise Robustness
💼 Related Jobs
No related jobs found.
X
Xingming Shui
Tsinghua Shenzhen International Graduate School, Tsinghua University
Dapeng Chen
Dapeng Chen
Huawei
Computer VisionMachine Learning
B
Bowei Liu
Tsinghua Shenzhen International Graduate School, Tsinghua University
J
Jingqi Tian
Tsinghua Shenzhen International Graduate School, Tsinghua University
M
Minfu Li
Independent Researcher
Kun Yi
Kun Yi
State Information Center
deep learning in the frequency domaintime series analysis
J
Jiapeng Hong
Independent Researcher
Y
Yansong Tang
Tsinghua Shenzhen International Graduate School, Tsinghua University