SAGE: Semantic Audio Generative Encoder

πŸ“… 2026-09-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the inherent trade-off among reconstruction quality, semantic structure, and inference speed in audio autoencoders by proposing a lightweight variational autoencoder. The method leverages knowledge distillation to compress embeddings from a pretrained large model into a compact latent space, achieving efficient representation learning with only 105M parameters. Experimental results demonstrate that the proposed model maintains audio quality comparable to its larger counterpart while accelerating inference speed by fourfold. Furthermore, it establishes new state-of-the-art performance across nineteen semantic probing tasks. By effectively reconciling these competing objectives, this work overcomes the conventional trilemma bottleneck in audio representation learning.
πŸ“ Abstract
Audio autoencoders compress waveforms into compact latent representations that serve as the interface between raw audio and downstream models. Current systems navigate a three-way trade-off between reconstruction quality, semantic structure of the latent space, and inference speed, typically favoring one or two of these at the expense of the others. This paper introduces SAGE, Semantic Audio Generative Encoder: a compact variational autoencoder, trained solely on publicly available music, that shapes its latent by distilling embeddings from a pretrained audio-text model. This 105M-parameter model runs at the inference cost of Stable Audio Open and reaches the listening-test quality of SAME-L, an autoencoder 8x larger and 4x slower, while surpassing both on objective perceptual and distributional metrics of reconstruction. Furthermore, it sets the state of the art on all nineteen probing tasks of latent semantics, in domain and out of domain. These results establish SAGE as a lightweight audio autoencoder that strikes the best balance of the three-way trade-off among those we evaluate, combining high reconstruction fidelity, state-of-the-art semantic structure, and fast inference.
Problem

Research questions and friction points this paper is trying to address.

audio autoencoder
reconstruction quality
semantic structure
inference speed
trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audio Autoencoder
Variational Autoencoder
Knowledge Distillation
Latent Semantics
Audio Generation
πŸ”Ž Similar Papers
F
Francesco Brigante
Sapienza University of Rome, Italy
L
Luca Cerovaz
Sapienza University of Rome, Italy, Paradigma
D
Davide Marincione
Sapienza University of Rome, Italy
G
Giorgio Strano
Sapienza University of Rome, Italy
L
Luca Zhou
Sapienza University of Rome, Italy
Emanuele RodolΓ 
Emanuele RodolΓ 
Professor of Computer Science, Sapienza University of Rome
Machine LearningAudioGeometric Deep LearningGeometry ProcessingComputer Vision
M
Michele Mancusi
Sapienza University of Rome, Italy, Moises Systems, Inc.