🤖 AI Summary
This study investigates the extent to which audio encoders preserve source signal details for downstream tasks. Leveraging the Stable Audio Open latent diffusion model, we perform music reconstruction from frozen embeddings of the Million Song Dataset using VGGish, CLAP, and EnCodec encoders. We systematically evaluate how supervised, contrastive, and reconstructive objectives, alongside time-frequency resolution, affect information invertibility. Our results demonstrate that finer-grained time-frequency structures significantly enhance reconstruction quality. Furthermore, even under high compression, the embeddings retain quantifiable source-specific characteristics and high-level semantic content. These findings offer new perspectives on understanding the information bottleneck inherent in audio representations.
📝 Abstract
Pretrained audio encoders are reused for downstream tasks that are often unknown when the encoder is trained, so their usefulness depends partly on which signal properties survive the pretext objective. We study this retained information through paired source reconstruction. Using a shared Stable Audio Open latent diffusion decoder, we reconstruct five-second, 44.1-kHz stereo music from frozen representations produced by supervised classifiers (VGGish, ConvNeXt), an audio-text contrastive model (CLAP), and a waveform-reconstruction model (EnCodec). These objectives impose different pressures to preserve source detail, while their exposed interfaces vary substantially in temporal and spectral resolution. Evaluating on the Million Song Dataset (MSD), we find clear differences in reconstructability across encoder families, while within encoder comparisons show improved recovery when finer temporal or spectral structure is exposed. Even compressed task oriented embeddings support reconstructions that preserve measurable source specificity and high level musical content.