EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs
Large audio language models often lose critical acoustic evidence under severe noise, leading to unreliable responses. This work proposes EchoDistill, a novel noisy-to-clean self-distillation framework that leverages clean audio as privileged information to guide a teacher model. The student model is optimized through masked token distillation and task-gated consistency shaping, while the backbone network remains frozen to ensure zero additional inference overhead. Experimental results demonstrate that this approach improves average accuracy by 1.63% under strong noise conditions at −10 dB, with Qwen2.5-Omni achieving 62.94%. These findings confirm that EchoDistill effectively enhances robustness against acoustic degradation without incurring extra computational costs during inference.