SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited content fidelity in existing continuous latent autoregressive speech generation methods, which lack explicit linguistic structural supervision. To remedy this, we propose SemBridge, a novel framework that introduces discrete semantic tokens—used only during training—to directly supervise the states of the autoregressive language model. A semantic-aligned acoustic variational autoencoder (VAE) is employed to construct a structured continuous target space. Notably, the inference pipeline remains fully continuous without any architectural modifications. The proposed approach substantially reduces word and character error rates while preserving high speaker similarity and perceptual quality. Furthermore, it enables zero-shot text-to-speech synthesis and score-conditioned singing voice generation.
📝 Abstract
Continuous-latent autoregressive speech generation has emerged as a promising alternative to discrete-token modeling by avoiding quantization loss and preserving richer acoustic information. However, continuous acoustic targets do not ex- pose linguistic structure as explicit token-level prediction tar- gets. Consequently, the autoregressive language model (LM) must acquire linguistic structure indirectly through acous- tic prediction, which can compromise the content fidelity of generated speech. We propose SemBridge, a training-only semantic-token anchoring framework for continuous-latent autoregressive speech generation. SemBridge uses discrete se- mantic tokens to directly supervise autoregressive LM states and employs a Semantic-Aligned Acoustic VAE to organize the continuous target space under the same semantic refer- ence. The semantic supervision is used only during train- ing, while inference remains entirely continuous. We evalu- ate SemBridge on zero-shot text-to-speech (TTS) and score- conditioned singing voice synthesis (SVS). Across multi- ple benchmarks, SemBridge improves content accuracy, as measured by word and character error rates (WER/CER), while maintaining competitive speaker similarity and percep- tual quality. Experimental results demonstrate that explicit semantic-token supervision for autoregressive state learning is an effective and general direction for continuous speech generation. Speech samples are available.1 The model code and checkpoints will be available at https://github.com/ASLP- lab/SemBridge
Problem

Research questions and friction points this paper is trying to address.

continuous-latent
autoregressive speech generation
linguistic structure
content fidelity
semantic tokens
Innovation

Methods, ideas, or system contributions that make the work stand out.

semantic token anchoring
continuous-latent autoregressive speech generation
Semantic-Aligned Acoustic VAE
zero-shot TTS
content fidelity
🔎 Similar Papers
No similar papers found.
Hanke Xie
Hanke Xie
Northwestern Polytechnical University
Audio speech synthesis
H
Haopeng Lin
Soul AI Lab, China
J
Jiale Qian
Soul AI Lab, China
Dake Guo
Dake Guo
Northwestern Polytechnical University
Speech ProcessingSpeech Synthesis
Yuepeng Jiang
Yuepeng Jiang
Northwestern Polytechnical University
Speech ProcessingSpeech SynthesisVoice Conversion
Z
Zhichao Wang
Soul AI Lab, China
W
Wenxiao Cao
Soul AI Lab, China
J
Jingbin Hu
Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, Xi’an, China
Guobin Ma
Guobin Ma
Northwestern Polytechnical University
Wenhao Li
Wenhao Li
Marshall School of Business, University of Southern California and NBER
Asset PricingFinancial IntermediationMacroeconomics
H
Huakang Chen
Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, Xi’an, China
C
Chengyou Wang
Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, Xi’an, China
M
Ming Tao
Soul AI Lab, China
Z
Zhonghua Fu
Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, Xi’an, China
Lei Xie
Lei Xie
Northwestern Polytechnical University
speech processingspeech recognitionspeech synthesismultimediaartificial intelligence
Xinsheng Wang
Xinsheng Wang
Hong Kong University of Science and Technology (HKUST)
speech synthesissinging voice synthesisvoice conversion