🤖 AI Summary
This study addresses the challenge that low-frame-rate speech tokenizers struggle to simultaneously preserve linguistic information and acoustic details. To this end, we propose a dual-stream low-frame-rate speech tokenizer. Methodologically, independent learnable query compressors are designed to aggregate semantic and acoustic features separately, while a context-aware query mechanism and an autoregressive text loss are introduced to achieve stream-specific compression. Experimental results demonstrate that the proposed model achieves superior speech reconstruction quality at equivalent frame rates compared to existing methods. Furthermore, it significantly improves accuracy in downstream automatic speech recognition tasks and enhances audio quality for speech synthesis.
📝 Abstract
Neural speech codecs increasingly serve as tokenizers for speech language models (SLMs). Lowering the frame rate reduces the computational and memory costs of SLMs, but makes it difficult to preserve both linguistic information and acoustic detail. Existing approaches rely on rule-based compression: average pooling can discard linguistic information, whereas similarity-based merging uses a fixed threshold on adjacent-frame similarity and applies the resulting boundaries to the acoustic stream. We propose Q-SPT, a low-frame-rate dual-stream speech tokenizer with separate, context-aware, learnable query-based compressors specialized for semantic and acoustic representations. In particular, queries at a fixed rate independently attend to the semantic and acoustic streams as separate key-value sources, enabling stream-specific, context-aware aggregation through two separately learned compressors. In addition, an autoregressive text loss explicitly supervises the semantic compressor to preserve linguistic information. Experimental results show that Q-SPT achieves the best reconstruction among the evaluated codecs at the same frame rate. In downstream SLMs, it yields the best speech recognition accuracy and text-to-speech perceptual quality with competitive intelligibility.