UNITE-AUDIO: Joint Learning of Continuous Tokenization and Latent Flow Matching for Text-to-Audio Generation

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出Unite-Audio,通过联合学习连续音频表示和潜在流匹配来解决文本到音频生成问题,改进了传统的两阶段方法。
📝 Abstract
Text-to-audio (TTA) generation aims to synthesize realistic audio that faithfully reflects natural-language descriptions. Most TTA systems adopt a two-stage latent paradigm: an audio tokenizer is optimized for reconstruction and then frozen, after which a generative model is trained in the resulting latent space. However, reconstruction-oriented representations may be suboptimal for generation, motivating joint representation and generative learning. To this end, we introduce \textbf{Unite-Audio}, to our knowledge, is the \textbf{first} to jointly learn continuous audio representations and latent flow matching for TTA. By coupling reconstruction with self-supervised generative prediction, Unite-Audio allows the generative objective to directly shape the latent space rather than treating it as a fixed intermediate representation. We further employ Flow-GRPO post-training to improve text-conditioned generation. Experiments show competitive TTA performance with a compact latent flow model, while ablation studies confirm the benefit of jointly learning the audio representation and generative model. Audio samples are available at https://runwushi.github.io/Unite-Audio.
Problem

Research questions and friction points this paper is trying to address.

Text-to-audio
Latent space
Reconstruction
Generative model
Joint learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Joint Learning
Continuous Audio Representations
Latent Flow Matching
Text-to-Audio Generation
Flow-GRPO
🔎 Similar Papers
No similar papers found.