🤖 AI Summary
This study addresses the significant degradation of speech intelligibility in noisy environments. Inspired by the human Lombard effect, it proposes a text-to-speech optimization framework based on dynamic activation guidance. The core innovation lies in the design of a prompt-relative guidance mechanism that enables dynamic adjustment of guidance intensity during generation without requiring model retraining, thereby effectively suppressing cumulative deviation effects. Experimental results demonstrate that the proposed method substantially reduces word error rates under low signal-to-noise ratio conditions while maintaining high speaker similarity. Overall, this work presents an efficient, training-free enhancement paradigm for robust speech synthesis.
📝 Abstract
Speech becomes less intelligible in noisy environments, and humans naturally adapt their voice to compensate. Inspired by this behavior, we investigate whether a text-to-speech (TTS) model can be guided to produce more intelligible speech using activation steering, without retraining. We focus on two characteristics of the Lombard effect: increased vocal effort and hyper-articulation. We introduce a prompt-relative steering mechanism that prevents steering effects from accumulating during generation while allowing their strength to be adjusted dynamically. Across seen and unseen speakers and multiple languages, our method produces systematic changes in Lombard-related acoustic features, preserves speaker similarity (89-95%), and reduces WER under background noise by 7-22% at 1 dB SNR. These results show that pretrained TTS models can be dynamically controlled to generate more intelligible speech without retraining.