Feature Encoding in VAE-based Audio Decoders: Effects of Input, Depth and Distribution

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited interpretability of internal representations in neural audio synthesis models by systematically investigating how musical attributes such as pitch and BPM are encoded within RAVE and EnCodec decoders. By integrating variational autoencoders, linear and nonlinear probing, and cross-layer clustering techniques, the analysis reveals that joint encoding across intermediate layers yields significant advantages, nonlinear probes provide substantial gains for natural audio, and cross-layer clustering outperforms single-layer analysis. Furthermore, this work quantifies the discrepancies in feature encoding intensity between synthesized and real audio. Ultimately, these findings elucidate the representational mechanisms underlying model internals, thereby enhancing interpretability and providing both a theoretical foundation and practical control strategies for achieving precise, controllable neural audio synthesis.
📝 Abstract
Neural audio synthesis models like the Realtime Audio Variational autoEncoder (RAVE) achieve impressive genera tion quality, yet how their internal representations encode musical features remains poorly understood. We present a systematic layer-wise and cross-layer cluster analysis of RAVE decoder activations across three models trained on different musical domains, tested with four stimulus types. We then evaluate architectural generalization with a general purpose EnCodec model. For RAVE, we find that synthetic stimuli are encoded well across models and audio features (pitch |\r{ho}|=0.45, 5.1x the null, BPM |\r{ho}| = 0.76, 8.6x the null). These results are reduced but still substantively apparent when using natural audio (mean across features |\r{ho}|=0.25, 2.8x the null). Natural audio sees a stronger encoding when nonlinear probes are used (mean across features R2=0.56, 18x the null, +0.152 nonlinear gain over the linear probe R2). Encoding strength varies throughout the layers of the decoder and an increased ability to joint-encode in the middle layers is seen across all audio features (\b{eta}2 all negative, p < 0.05). The general purpose EnCodec decoder also sees similar strong synthetic responses across audio features, similar nonlinear gains for natural audio joint encoding and similar depth profiles. We find the best cross-layer cluster improves the strength (r = 0.65, p = 0.006) and prevalence (r = 0.75, p = 0.001) of BPM encoding when compared against the best whole layers within the same section, with no effect for joint encoding. These findings advance the interpretability of neural audio models and inform targeted control strategies for neural synthesis.
Problem

Research questions and friction points this paper is trying to address.

Variational Autoencoder
Audio Decoders
Feature Encoding
Neural Audio Synthesis
Interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Variational Autoencoder
Feature Encoding
Interpretability
Cross-layer Cluster Analysis
Neural Audio Synthesis
🔎 Similar Papers
No similar papers found.