Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding

πŸ“… 2026-07-23
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the "perceptual bypass" issue in Omni models, where reliance on textual prompts undermines robust speech understanding in overlapping and noisy environments. To mitigate this, we propose Audio-Grounded Scaffolding Context (AGSC), which constructs non-answering auditory cues from the input audio to guide the model’s attention toward acoustic information and internalizes this capability during training. Our approach enables robust comprehension of complex speech even without external prompts. We are the first to systematically identify and tackle perceptual bypass in multimodal speech models, introducing an internalizable audio-anchoring mechanism combined with scaffolding-based training and a joint GDPO optimization framework that co-learns gating, formatting, and transcription rewards. Experiments show that AGSC reduces unprompted mpWER on overlapping noisy speech from 25%–71% to 9%–15% across three heterogeneous Omni models, with negligible inference overhead.
πŸ“ Abstract
Omni models transcribe clean, single-speaker speech well, but their accuracy drops sharply when speakers overlap and the scene is noisy, exactly where knowing who said what matters most. A natural fix is a short scene description. We show why this is risky: answer-bearing text lets the model copy instead of listen, so the score rises although nothing has been heard; a silent test exposes this shortcut at once. We call this failure mode perception bypass and address it with Audio-Grounded Scaffold Context (AGSC). AGSC links three steps: first, we build clues from audio to guide listening without giving the answer; second, answer-overlap and silence tests probe them for leakage and audio dependence; finally, those clues scaffold training but vanish at test time, yielding no-clue capability. Across three heterogeneous Omni models, training on AGSC lowers no-clue capped mean permutation word error rate (mpWER) on overlapping, noisy speech from 25%-71% to 9%-15%. For streaming control, we formulate a joint GDPO task in which the model learns when to use a clue and how to produce a speaker-attributed transcript from separately normalized format, gate, and transcript rewards. After internalization, AGSC adds almost no inference overhead.
Problem

Research questions and friction points this paper is trying to address.

speech understanding
overlapping speech
perception bypass
noisy environments
speaker attribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audio-Grounded Scaffold Context
perception bypass
omni-model speech understanding
speaker-attributed transcription
GDPO
πŸ”Ž Similar Papers
No similar papers found.