π€ AI Summary
This study addresses the inherent trade-off among reconstruction quality, semantic structure, and inference speed in audio autoencoders by proposing a lightweight variational autoencoder. The method leverages knowledge distillation to compress embeddings from a pretrained large model into a compact latent space, achieving efficient representation learning with only 105M parameters. Experimental results demonstrate that the proposed model maintains audio quality comparable to its larger counterpart while accelerating inference speed by fourfold. Furthermore, it establishes new state-of-the-art performance across nineteen semantic probing tasks. By effectively reconciling these competing objectives, this work overcomes the conventional trilemma bottleneck in audio representation learning.
π Abstract
Audio autoencoders compress waveforms into compact latent representations that serve as the interface between raw audio and downstream models. Current systems navigate a three-way trade-off between reconstruction quality, semantic structure of the latent space, and inference speed, typically favoring one or two of these at the expense of the others. This paper introduces SAGE, Semantic Audio Generative Encoder: a compact variational autoencoder, trained solely on publicly available music, that shapes its latent by distilling embeddings from a pretrained audio-text model. This 105M-parameter model runs at the inference cost of Stable Audio Open and reaches the listening-test quality of SAME-L, an autoencoder 8x larger and 4x slower, while surpassing both on objective perceptual and distributional metrics of reconstruction. Furthermore, it sets the state of the art on all nineteen probing tasks of latent semantics, in domain and out of domain. These results establish SAGE as a lightweight audio autoencoder that strikes the best balance of the three-way trade-off among those we evaluate, combining high reconstruction fidelity, state-of-the-art semantic structure, and fast inference.