Towards Interpretable Framework for Neural Audio Codecs via Sparse Autoencoders: Exploration toward Age, Gender, and Accent Steering

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the unclear encoding mechanisms of speaker attributes within neural audio codecs. To this end, it pioneers the extension of sparse autoencoders (SAEs) to the waveform level, constructing multidimensional probes for attributes such as age, gender, and accent, while enabling attribute-oriented speech reconstruction and control through activation steering techniques. The findings demonstrate that SAEs can effectively capture steerable speaker attributes and induce targeted predictive shifts, thereby revealing the fundamental challenge of disentangling attribute controllability from audio fidelity. However, attribute interventions are accompanied by elevated word error rates and degraded perceptual quality, establishing a critical baseline for future research in controllable neural audio synthesis.
πŸ“ Abstract
Neural audio codecs (NACs) are widely used in speech generation and audio-language modeling, yet how they encode speaker-trait information remains poorly understood. Prior work applied sparse autoencoders (SAEs) to investigate accent information in NACs through task-level analysis. Here, we extend this analysis to the waveform level and to age, gender, and accent, using SAE steering to probe trait-related information in sparse activations. We identify trait-associated dimensions, modify their activations, and evaluate the resulting reconstructed speech. Across five NACs, steering the selected dimensions induces target-directed shifts in speaker-trait predictions. A random-dimension baseline on Mimi produces smaller shifts, supporting the relevance of the selected dimensions. However, responses vary across codecs, traits, and steering directions, and increasing steering strength does not consistently amplify the intended shifts. Steering also generally increases word error rates and lowers predicted perceptual quality. These findings suggest that SAEs capture speaker-trait information in steerable activations, while the accompanying quality degradation highlights the need to better separate trait-related information from other information.
Problem

Research questions and friction points this paper is trying to address.

Neural Audio Codecs
Speaker Traits
Interpretability
Sparse Autoencoders
Speech Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Neural Audio Codecs
Sparse Autoencoders
Speaker-Trait Steering
Interpretability
Activation Intervention
πŸ”Ž Similar Papers
No similar papers found.