Q-SPT: Learnable Query-Based Compression for Low-Frame-Rate Speech Tokenization

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that low-frame-rate speech tokenizers struggle to simultaneously preserve linguistic information and acoustic details. To this end, we propose a dual-stream low-frame-rate speech tokenizer. Methodologically, independent learnable query compressors are designed to aggregate semantic and acoustic features separately, while a context-aware query mechanism and an autoregressive text loss are introduced to achieve stream-specific compression. Experimental results demonstrate that the proposed model achieves superior speech reconstruction quality at equivalent frame rates compared to existing methods. Furthermore, it significantly improves accuracy in downstream automatic speech recognition tasks and enhances audio quality for speech synthesis.
📝 Abstract
Neural speech codecs increasingly serve as tokenizers for speech language models (SLMs). Lowering the frame rate reduces the computational and memory costs of SLMs, but makes it difficult to preserve both linguistic information and acoustic detail. Existing approaches rely on rule-based compression: average pooling can discard linguistic information, whereas similarity-based merging uses a fixed threshold on adjacent-frame similarity and applies the resulting boundaries to the acoustic stream. We propose Q-SPT, a low-frame-rate dual-stream speech tokenizer with separate, context-aware, learnable query-based compressors specialized for semantic and acoustic representations. In particular, queries at a fixed rate independently attend to the semantic and acoustic streams as separate key-value sources, enabling stream-specific, context-aware aggregation through two separately learned compressors. In addition, an autoregressive text loss explicitly supervises the semantic compressor to preserve linguistic information. Experimental results show that Q-SPT achieves the best reconstruction among the evaluated codecs at the same frame rate. In downstream SLMs, it yields the best speech recognition accuracy and text-to-speech perceptual quality with competitive intelligibility.
Problem

Research questions and friction points this paper is trying to address.

speech tokenization
low frame rate
speech language models
neural speech codecs
compression
Innovation

Methods, ideas, or system contributions that make the work stand out.

Low-frame-rate speech tokenization
Query-based compression
Dual-stream codec
Learnable compressor
Autoregressive text loss
🔎 Similar Papers
2024-07-22arXiv.orgCitations: 4
💼 Related Jobs
No related jobs found.
J
Jeeyoung Yun
Korea University, Seoul, Republic of Korea
S
Seohwan Yun
Korea University, Seoul, Republic of Korea
Sungwoong Kim
Sungwoong Kim
Associate Professor, Korea University
artificial general intelligence