Rethinking Length-Based Training: Batch Composition and Loss Normalization in Speech Token Language Models

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过匹配比较分析了语音模型中基于长度训练的影响因素,发现短到长排序在固定批次组成和标记暴露下无独立益处。
📝 Abstract
Short-to-long training is a simple curriculum for speech models, but its gains can be difficult to interpret. In speech token language models, length-based training can change the shuffle policy, batch composition, token retention, and token weights under batch-mean loss. We disentangle these factors through matched comparisons. In the tested settings, short-to-long ordering shows no independent benefit when batch composition and token exposure are fixed. First-epoch grouping lowers perplexity for Mimi under batch-mean loss, but this gain is not observed under token-balanced loss. The cross-tokenizer results are consistent with a link between chunk-length variation and token weighting. This work provides a systematic analysis protocol for studying length-based training in variable-length speech models.
Problem

Research questions and friction points this paper is trying to address.

length-based training
batch composition
loss normalization
speech token language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

length-based training
batch composition
loss normalization
speech token language models
token weighting
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hongjin Song
Beijing Institute of Technology, Zhuhai
Runwu Shi
Runwu Shi
Institute of Science Tokyo
Signal processingIntelligent vehicle
W
Weiqiao Shan
Northeastern University
J
Jiale Luo
Sichuan University
Yujin Wang
Yujin Wang
Ph.D. Student, Tongji University
Y
Yifei Wu
Ant Group
C
Chunxiang Jin
Ant Group