AudioGAR: Bridging Reconstruction and Generation in Latent Audio Generative Models

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the distribution mismatch between the training and inference phases of decoders in latent audio generation models. To bridge this gap, the authors propose an intermediate latent construction strategy that requires no source correspondence. Specifically, by perturbing encoder latents and subsequently denoising them via a diffusion model, the method establishes a smooth transition trajectory from reconstruction to generation, enabling effective decoder fine-tuning to align the divergent distributions. This approach significantly enhances end-to-end audio generation quality while requiring only 1.5% of the original training data and 0.26% of the computational cost, demonstrating highly efficient adaptation under extremely low-resource conditions.
📝 Abstract
Latent audio generative models are typically trained in two stages: an audio codec is learned first, followed by a latent generative model. This decomposition leads to a decoder train-generation mismatch: the codec decoder is trained on encoder-induced latents but deployed on generator-produced latents at inference time. Across diverse datasets and latent generative models, we observe clear reconstruction-generation gaps under both FD and FAD, showing that strong reconstruction quality does not necessarily translate into strong end-to-end generation quality. A natural remedy is to adapt the decoder on generation-produced latents, but generated latents lack correspondence with source audio and therefore cannot directly provide the paired supervision used for decoder fine-tuning. We introduce \textbf{AudioGAR}, which constructs intermediate latents by perturbing encoder latents and denoising them through the frozen latent diffusion model. These latents form a trajectory from reconstruction toward generation, with lower-noise latents retaining source correspondence and supporting paired decoder fine-tuning. We fine-tune only the codec decoder on these latents, while keeping the codec encoder and latent generative model frozen. When applied to AudioX, AudioGAR substantially improves generative performance. It requires only 1.5\% of the original training audio hours and 0.26\% of the original training cost.
Problem

Research questions and friction points this paper is trying to address.

latent audio generation
train-generation mismatch
audio codec
reconstruction-generation gap
decoder fine-tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Audio Generation
Train-Generation Mismatch
Decoder Fine-tuning
Intermediate Latents
Latent Diffusion Model
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.