🤖 AI Summary
Idiom translation constitutes a critical bottleneck in speech-to-text translation (SLT): idiomatic meanings are non-compositional and cannot be derived literally, yet prevailing end-to-end SLT models lack sufficient semantic abstraction capability, frequently producing erroneous literal translations. This work is the first to systematically characterize the idiom translation challenge in SLT. We benchmark state-of-the-art SLT systems—including end-to-end models (SeamlessM4T, Whisper v3), cascade SLT, text-based machine translation (MT), and large language models (LLaMA, DeepSeek)—on German–English and Russian–English idiom-rich datasets. Results show that SLT systems suffer an average BLEU drop of 12.7 points on idiomatic content compared to MT and LLMs; even when leveraging high-level acoustic or linguistic representations, they remain prone to literal translation errors. Our analysis reveals that robust idiom handling requires idiom-aware architectural designs and dedicated training strategies. This work establishes foundational insights and paves a new direction toward semantically robust SLT systems.
📝 Abstract
Idioms are defined as a group of words with a figurative meaning not deducible from their individual components. Although modern machine translation systems have made remarkable progress, translating idioms remains a major challenge, especially for speech-to-text systems, where research on this topic is notably sparse. In this paper, we systematically evaluate idiom translation as compared to conventional news translation in both text-to-text machine translation (MT) and speech-to-text translation (SLT) systems across two language pairs (German to English, Russian to English). We compare state-of-the-art end-to-end SLT systems (SeamlessM4T SLT-to-text, Whisper Large v3) with MT systems (SeamlessM4T SLT-to-text, No Language Left Behind), Large Language Models (DeepSeek, LLaMA) and cascaded alternatives. Our results reveal that SLT systems experience a pronounced performance drop on idiomatic data, often reverting to literal translations even in higher layers, whereas MT systems and Large Language Models demonstrate better handling of idioms. These findings underscore the need for idiom-specific strategies and improved internal representations in SLT architectures.