🤖 AI Summary
This study addresses the inability of large audio models to assess their own transcription reliability, which often leads to erroneous responses under degraded inputs. We demonstrate that reliability signals are inherently encoded within frozen audio encoder representations and accordingly propose a plug-and-play detection mechanism that requires no modifications to the underlying model. Specifically, a lightweight predictor is trained to evaluate transcription quality and trigger clarification requests to optimize interaction. Combined with cross-model label transfer, this mechanism achieves macro-F1 scores of 81.10% and 78.09% in in-domain and cross-domain settings, respectively. These results significantly outperform baseline methods while exhibiting strong generalization across diverse model families.
📝 Abstract
Audio large language models allow users to interact with the model through speech. When an input recording is too degraded, the model may misinterpret the user's query and respond based on an incorrect transcription. In this paper, we study model-conditional transcription reliability: whether an Audio LLM can recognize when its own transcription is unreliable. We first prompt the Audio LLM to assess whether its own transcription would be reliable, and find that the model is a poor judge of its own transcription reliability: in most cases, it predicts that its transcription will be reliable. We find that existing approaches, including speech quality predictors, audio LLM generation uncertainty, and transcript-conditioned WER estimation, provide limited signals for detecting transcription failures. In contrast, we discover that transcription reliability is strongly represented in the model's audio-encoder representations. Based on this observation, we devise a lightweight reliability predictor that operates on representations extracted by the frozen audio encoder and predicts the reliability class before generation. The reliability predictor can trigger a clarification request from the user when their voice query is predicted to be unreliable, while allowing reliable queries to proceed without modifying the underlying Audio LLM. Our predictor achieves 81.10% in-domain and 78.09% cross-domain macro-F1 scores, outperforming the strongest baselines by 10.33 and 11.93 points, respectively. Finally, we show that reliability labels can transfer across Audio LLM families, and that transfer performance is closely related to the alignment of their model-specific reliability boundaries.