🤖 AI Summary
This study addresses the ambiguous semantics of boundaries in dynamic frame-rate speech codecs and their unclear impact on reconstruction. By integrating boundary interpretation with controlled reconstruction analysis, comparative experiments are conducted using multiple encoders, including SenseVoice and Whisper, to systematically investigate the semantic meaning and reconstruction utility of boundaries across different encoder layers. The findings reveal that boundary semantics depend on low-level representations, while high-level linguistic alignment effectively mitigates pooling distortion. Under semantically rich representations, syllable-scale boundary allocation significantly outperforms fixed frame-rate baselines, whereas phoneme-level segmentation yields no notable gains. This work provides both theoretical foundations and practical guidance for optimal boundary strategies in dynamic frame-rate neural speech codecs.
📝 Abstract
Dynamic-frame-rate neural speech codecs replace a uniform frame grid with variable-duration tokens, making boundary placement part of the representation itself. Yet it is unclear what these boundaries encode and whether interpretable boundaries are also useful for neural speech reconstruction. This work combines boundary interpretation and controlled reconstruction analysis by comparing predicted boundaries with linguistic and acoustic references. We find that boundary meaning depends on the underlying speech representation. For the ASR-oriented SenseVoice and Whisper encoders, shallow layers emphasize phonetic, voicing, and acoustic transitions, whereas deeper layers shift toward syllable and subword structure. In a comparison of six dynamic-frame-rate algorithms and a uniform (fixed-frame-rate) baseline, higher-level linguistic alignment is associated with lower pooling distortion and better reconstruction from semantic tokens. Frame-rate-matched boundary tests make the distinction concrete: a syllable-derived partition improves over both Uniform and Similarity, while a phoneme-derived partition does not improve over Uniform. We infer that for semantically rich speech representations, useful codec boundaries are best understood as allocation decisions organized around the syllable scale.