Beyond Encoder Fusion: Multi-View Discrete Token Augmentation for LLM-Based ASR

📅 2026-09-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决单令牌化系统对编码器选择敏感的问题,提出多视图离散令牌增强方法,通过生成替代令牌序列增加训练多样性,提高自动语音识别性能。
📝 Abstract
Discrete speech tokens provide a compact interface between speech encoders and large language models for automatic speech recognition, but single-tokenization systems remain sensitive to the chosen encoder. We propose multi-view discrete token augmentation, a simple strategy that augments each training utterance by generating alternative token sequences from fixed SSL encoders, such as HuBERT, WavLM, and MMS-300M. These tokenizations are treated as complementary training views for a shared LLM decoder, exposing it to more diverse discrete speech representations without requiring multi-encoder inference. At test time, the model can operate with a single encoder. On LibriSpeech, the approach consistently improves all encoders over independently trained baselines, with WavLM reaching 3.30% WER on test-clean and 8.13% on test-other. Budget-matched controls show that the gains come from encoder diversity rather than data volume. ROVER over multi-view hypotheses further improves WER to 3.03% and 7.38%.
Problem

Research questions and friction points this paper is trying to address.

Discrete Speech Tokens
Encoder Sensitivity
Automatic Speech Recognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-view discrete token augmentation
self-supervised learning encoders
shared LLM decoder
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Paul Moïse Gangbadja
LIA, Avignon University, France; EDL, France
Mickael Rouvier
Mickael Rouvier
University of Avignon - LIA
Automatic Speech RecognitionSpeaker DiarizationSpeaker Verification
F
Fabrice Lefèvre
LIA, Avignon University, France