Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种基于音素的分词器,通过将音节分解为声母、韵母和声调三个部分来解决传统分词方法忽视音节内部结构的问题,从而在多种语言理解任务中取得了与现有模型相当或更好的效果。
📝 Abstract
Conventional tokenizers represent text as characters or statistically derived subwords, overlooking the internal phonological structure of syllables and often requiring large vocabularies. We introduce \textbf{Phonemic Tokenizer}, a linguistically motivated tokenizer for Vietnamese and Chinese that converts each syllable into IPA and factorizes it into three phonological components: onset, rime, and tone. The three components jointly occupy one contextual position, preserving syllable-level sequence length while enabling representation sharing across phonologically related syllables. Non-phonological and unsupported units are handled through character-level fallback. This deterministic design requires no corpus-dependent vocabulary learning and yields vocabularies of only 112 entries for Chinese and 256 for Vietnamese. Intrinsic evaluation shows that the tokenizer achieves substantially higher Rényi efficiency in both languages, represents every entry in a standard Vietnamese syllable dictionary with a Fertility of exactly one, and generally produces shorter Vietnamese sequences than existing pretrained tokenizers. We further instantiate the tokenizer in \textbf{PhonemicBERT}, which combines factorized component embeddings and reconstructs complete masked syllables using three prediction heads. Under a controlled Chinese pretraining setup, PhonemicBERT-Zh is competitive with or outperforms character, subword, and SubChar alternatives across diverse language-understanding tasks. PhonemicBERT-Vi also achieves competitive or superior results to established Vietnamese and multilingual pretrained models. These results establish phonemic factorization as a compact, efficient, and interpretable alternative to atomic and statistically segmented text representations.
Problem

Research questions and friction points this paper is trying to address.

phonological structure
syllables
tokenizer
language model pretraining
vocabulary size
Innovation

Methods, ideas, or system contributions that make the work stand out.

Phonemic Tokenizer
syllable factorization
phonological components
PhonemicBERT
representation sharing
🔎 Similar Papers
2023-10-16IEEE International Conference on Acoustics, Speech, and Signal ProcessingCitations: 8
2024-06-21arXiv.orgCitations: 0
💼 Related Jobs
No related jobs found.
N
Nghia Hieu Nguyen
University of Information Technology, Vietnam National University, Ho Chi Minh city, Viet Nam
T
Thai Bao Huynh
University of Information Technology, Vietnam National University, Ho Chi Minh city, Viet Nam
B
Binh-An Dinh-Le
University of Information Technology, Vietnam National University, Ho Chi Minh city, Viet Nam
P
Phu Gia Hoang
Independent Researcher, Germany
Dat Tien Nguyen
Dat Tien Nguyen
Unknown affiliation
Information RetrievalNatural Language ProcessingMachine Learning
Kiet Van Nguyen
Kiet Van Nguyen
University of Information Technology, VNU-HCM
Data ScienceArtificial IntelligenceComputational Linguistics
Ngan Luu-Thuy Nguyen
Ngan Luu-Thuy Nguyen
University of Information Technology
Natural Language Processing