ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

📅 2026-05-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that existing text-to-speech (TTS) systems struggle to generate speech that naturally blends with ambient audio, primarily due to significant discrepancies in acoustic characteristics and temporal dynamics between speech and environmental sounds. To overcome this limitation, the authors propose an environment-aware TTS approach based on a multimodal diffusion Transformer that explicitly models cross-modal interactions between speech and environmental audio. A novel text-conditional joint attention mechanism is introduced to effectively fuse speech latent representations with environmental context. Furthermore, a domain-specific representation alignment objective tailored for environment-aware TTS is innovatively incorporated, leveraging self-supervised speech and general-purpose audio encoders to enhance semantic consistency. Experimental results demonstrate that the proposed method significantly outperforms current state-of-the-art approaches in both objective metrics and subjective listening tests, markedly improving the naturalness, intelligibility, and audio fidelity of the generated speech.
📝 Abstract
Recent advancements in text-guided audio generation have yielded promising results in diverse domains, including sound effects, speech, and music. However, jointly generating speech with environmental audio remains challenging due to the inherent disparities in their acoustic patterns and temporal dynamics. We propose ImmersiveTTS, an environment-aware text-to-speech (TTS) model that generates natural speech seamlessly integrated within environmental contexts by explicitly modeling cross-modal interactions. Our model builds on a multimodal diffusion transformer and fuses transcript-aligned speech latent with text-conditioned environmental context via joint attention. To enhance semantic consistency, we introduce a domain-specific representation alignment objective tailored to environment-aware TTS, leveraging complementary self-supervised representations from speech and audio encoders. Experimental results show that ImmersiveTTS achieves higher naturalness, intelligibility, and audio fidelity than existing approaches across objective metrics and human listening tests.
Problem

Research questions and friction points this paper is trying to address.

text-to-speech
environment-aware
multimodal generation
speech-environment integration
acoustic disparity
Innovation

Methods, ideas, or system contributions that make the work stand out.

environment-aware TTS
multimodal diffusion transformer
representation alignment
joint attention
self-supervised learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.