🤖 AI Summary
This work addresses the challenge that existing text-to-speech (TTS) systems struggle to generate speech that naturally blends with ambient audio, primarily due to significant discrepancies in acoustic characteristics and temporal dynamics between speech and environmental sounds. To overcome this limitation, the authors propose an environment-aware TTS approach based on a multimodal diffusion Transformer that explicitly models cross-modal interactions between speech and environmental audio. A novel text-conditional joint attention mechanism is introduced to effectively fuse speech latent representations with environmental context. Furthermore, a domain-specific representation alignment objective tailored for environment-aware TTS is innovatively incorporated, leveraging self-supervised speech and general-purpose audio encoders to enhance semantic consistency. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art approaches in both objective metrics and subjective listening tests, markedly improving the naturalness, intelligibility, and audio fidelity of the generated speech.
📝 Abstract
Recent advancements in text-guided audio generation have yielded promising results in diverse domains, including sound effects, speech, and music. However, jointly generating speech with environmental audio remains challenging due to the inherent disparities in their acoustic patterns and temporal dynamics. We propose ImmersiveTTS, an environment-aware text-to-speech (TTS) model that generates natural speech seamlessly integrated within environmental contexts by explicitly modeling cross-modal interactions. Our model builds on a multimodal diffusion transformer and fuses transcript-aligned speech latent with text-conditioned environmental context via joint attention. To enhance semantic consistency, we introduce a domain-specific representation alignment objective tailored to environment-aware TTS, leveraging complementary self-supervised representations from speech and audio encoders. Experimental results show that ImmersiveTTS achieves higher naturalness, intelligibility, and audio fidelity than existing approaches across objective metrics and human listening tests.