Disentangling Linguistic and Paralinguistic Information with Routed Sparse Autoencoders

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the entanglement of linguistic and paralinguistic information in self-supervised speech encoders by proposing a routing-based selective disentanglement framework utilizing Top-K sparse autoencoders. Operating on frozen SPEAR/WavLM encoders, the method integrates route-specific supervision with cross-factor adversarial training to achieve corpus-agnostic information separation and feature intervention without retraining. Experimental results demonstrate that this mechanism effectively preserves route-exclusive information while suppressing irrelevant factors. Furthermore, it exhibits strong consistency and robustness across diverse encoder architectures, corpora, and intervention scenarios.
📝 Abstract
Self-supervised speech encoders contain linguistic and paralinguistic information in a shared, entangled representation space. We combine a TopK sparse autoencoder with route-specific supervision and cross-factor adversaries. Across frozen SPEAR and WavLM encoders, independent probes show factor-specific retention and suppression: linguistic information remains stronger in the linguistic route, while paralinguistic factors, including speaker identity, emotion, and prosody, are retained in the paralinguistic route and substantially reduced in the linguistic route. The route organisation learned on LibriSpeech persists on MSP-Podcast without representation-side retraining. Feature-space route interventions further transfer the swapped factor while largely preserving the information carried by the unchanged route. These results show consistent route-selective separation across encoders, corpora, independent probes, and representation-level interventions.
Problem

Research questions and friction points this paper is trying to address.

speech representation
disentanglement
linguistic information
paralinguistic information
self-supervised speech encoders
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Autoencoders
Speech Disentanglement
Adversarial Training
Cross-corpus Generalization
Feature-space Intervention
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Beimnet Bekele Guta
Department of Engineering, University of Cambridge, Cambridge, UK
Xiaoyu Yang
Xiaoyu Yang
University of Cambridge
Speech recognitionmachine learning
Guangzhi Sun
Guangzhi Sun
University of Cambridge
Speech and language technologyconversational AI
P
Philip C. Woodland
Department of Engineering, University of Cambridge, Cambridge, UK