How Much Audio Is Left In An Embedding? An Inversion Audit Of Audio Encoders

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the extent to which audio encoders preserve source signal details for downstream tasks. Leveraging the Stable Audio Open latent diffusion model, we perform music reconstruction from frozen embeddings of the Million Song Dataset using VGGish, CLAP, and EnCodec encoders. We systematically evaluate how supervised, contrastive, and reconstructive objectives, alongside time-frequency resolution, affect information invertibility. Our results demonstrate that finer-grained time-frequency structures significantly enhance reconstruction quality. Furthermore, even under high compression, the embeddings retain quantifiable source-specific characteristics and high-level semantic content. These findings offer new perspectives on understanding the information bottleneck inherent in audio representations.
📝 Abstract
Pretrained audio encoders are reused for downstream tasks that are often unknown when the encoder is trained, so their usefulness depends partly on which signal properties survive the pretext objective. We study this retained information through paired source reconstruction. Using a shared Stable Audio Open latent diffusion decoder, we reconstruct five-second, 44.1-kHz stereo music from frozen representations produced by supervised classifiers (VGGish, ConvNeXt), an audio-text contrastive model (CLAP), and a waveform-reconstruction model (EnCodec). These objectives impose different pressures to preserve source detail, while their exposed interfaces vary substantially in temporal and spectral resolution. Evaluating on the Million Song Dataset (MSD), we find clear differences in reconstructability across encoder families, while within encoder comparisons show improved recovery when finer temporal or spectral structure is exposed. Even compressed task oriented embeddings support reconstructions that preserve measurable source specificity and high level musical content.
Problem

Research questions and friction points this paper is trying to address.

audio encoders
embedding inversion
information retention
source reconstruction
pretrained representations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Inversion Audit
Audio Encoders
Latent Diffusion Decoder
Source Reconstruction
Representation Analysis
🔎 Similar Papers
No similar papers found.