Listen, Attend, Understand: a Regularization Technique for Stable E2E Speech Translation Training on High Variance labels

📅 2026-01-03
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the instability and performance limitations of end-to-end speech translation when target transcriptions exhibit high variance and semantic ambiguity. The authors propose LAU, a method that introduces a directional auxiliary loss by freezing textual embeddings to impose semantic regularization on the acoustic encoder’s latent space, prioritizing semantic fidelity over literal phonetic details. To further enforce this constraint, the encoder’s weight structure is reconfigured. Additionally, the Total Parameter Drift metric is introduced to quantify the impact of regularization. Evaluated on only 30 hours of non-expertly annotated Bambara–French data, the LAU model achieves or surpasses the performance of systems pretrained with twice the amount of data, demonstrating significant improvements in both semantic fidelity and training stability.

Technology Category

Natural Language Processing: Language Grounding & Multi-modal NLPMachine Learning: Large Multimodal Models (LMMs)Computer Vision: Language and Vision

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
End-to-End Speech Translation often shows slower convergence and worse performance when target transcriptions exhibit high variance and semantic ambiguity. We propose Listen, Attend, Understand (LAU), a semantic regularization technique that constrains the acoustic encoder's latent space during training. By leveraging frozen text embeddings to provide a directional auxiliary loss, LAU injects linguistic groundedness into the acoustic representation without increasing inference cost. We evaluate our method on a Bambara-to-French dataset with 30 hours of Bambara speech translated by non-professionals. Experimental results demonstrate that LAU models achieve comparable performance by standard metrics compared to an E2E-ST system pretrained with 100\% more data and while performing better in preserving semantic meaning. Furthermore, we introduce Total Parameter Drift as a metric to quantify the structural impact of regularization to demonstrate that semantic constraints actively reorganize the encoder's weights to prioritize meaning over literal phonetics. Our findings suggest that LAU is a robust alternative to post-hoc rescoring and a valuable addition to E2E-ST training, especially when training data is scarce and/or noisy.
Problem

Research questions and friction points this paper is trying to address.

End-to-End Speech Translation
high variance labels
semantic ambiguity
acoustic encoder
regularization
Innovation

Methods, ideas, or system contributions that make the work stand out.

semantic regularization
end-to-end speech translation
latent space constraint
frozen text embeddings
Total Parameter Drift
🔎 Similar Papers
No similar papers found.