🤖 AI Summary
This study addresses the inaccuracy and inconsistency of domain-specific term transcription in long-form speech recognition. We propose an LLM agent-based error correction mechanism that integrates global context awareness with selective re-transcription. Specifically, a large language model agent analyzes the global discourse to identify suspicious terms and subsequently directs the automatic speech recognition system to perform targeted re-transcription for error rectification. As the first framework to unify global contextual perception with local selective re-transcription, our approach yields substantial improvements in terminology accuracy for both Chinese and English. Notably, it achieves a 36.8% relative reduction in the Chinese character error rate for terminology (B-CER), demonstrating its effectiveness in enhancing domain-specific transcription fidelity within extended audio contexts.
📝 Abstract
Recent advances in speech language models have improved automatic speech recognition (ASR) for long-form audio. However, accurately and consistently transcribing domain-specific terminology remains challenging. Motivated by the world knowledge and contextual capability of large language models (LLMs), we propose Agentic-GER, an LLM-based agent for terminology correction in long-form speech. The agent uses global context from the full transcript to identify suspicious terms and resolve ambiguous hypotheses. It selectively re-transcribes the source speech to check candidate corrections, and uses accepted edits to guide subsequent decisions. Experiments with four LLMs and two ASR systems on GigaSpeechBench show consistent terminology improvements in both Chinese and English, with and without thinking. On Chinese speech, Agentic-GER achieves up to a 36.8% relative reduction in biased character error rate (B-CER) over the Whisper baseline.