Counting and Min-Cost Encoding for Tokenization in Large Language Models

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of mainstream LLM tokenizers, including substantial sequence length variance, high inference latency, and suboptimal compression ratios. To overcome these bottlenecks, it proposes the Count Filtering (CNF) and Minimum Cost Encoding (MCE) algorithms, which optimize vocabulary construction by globally minimizing segmentation costs. Notably, this approach eliminates reliance on merge lists or probabilistic statistics, thereby offering enhanced scalability and generalizability. The proposed method significantly improves token efficiency, increasing the compression ratio for English text by 30% and boosting token efficiency by 60% with million-scale vocabularies. Furthermore, downstream task performance remains comparable to that of Byte-Pair Encoding (BPE), demonstrating that the approach achieves superior compression without compromising model effectiveness.
📝 Abstract
Mainstream large language models rely on a tokenizer to encode text into a token sequence. Different tokenizers may yield token sequences of substantially different lengths for the same text. With a fixed model architecture, shorter token sequences correspond to lower inference time. We propose a tokenizer training approach named Counting and Filtering (CNF) and a text encoding algorithm called Min-Cost Encoding (MCE). MCE defines a cost function over a text segment, and determines the best segmentation by globally minimizing the overall segmentation cost. CNF builds a raw vocabulary by directly counting valid substrings, and then constructs the final vocabulary through a filtering step based on actual token usage when segmenting the training corpus with MCE. The CNF-MCE conbination offers several advantages over BPE, including higher token efficiency, greater scalability, and lower dependency. Across six text categories and two vocabulary-size groups, CNF-MCE consistently achieves better compression than the evaluated BPE tokenizers. With a 250K vocabulary, CNF-MCE increases compression rate by 26% and 30% on English web text over the o200k_base and qwen250k tokenizers. Experiments scaling the vocabulary to 1M entries on English web text demonstrate sustained improvements over BPE, with a token efficiency improvement of over 60% and vocabulary utilization rising from 52.9% to 96.9%. The MCE algorithm does not depend on a merge list (as in BPE) or token probability (as in UnigramLM), making it applicable to a wide range of vocabularies, including those built from BPE, UnigramLM, CNF, and others. Language models trained from scratch at the 1.8B and 8B scales achieve comparable average performance to models using the BPE tokenizers across 11 benchmarks. These results demonstrate that CNF-MCE can improve token efficiency significantly while maintaining competitive downstream performance.
Problem

Research questions and friction points this paper is trying to address.

Tokenization
Large Language Models
Token Efficiency
Text Compression
Vocabulary Utilization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Tokenization
Min-Cost Encoding
Counting and Filtering
Large Language Models
Token Efficiency
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Shuming Shi
Shuming Shi
Tencent AI Lab
NLPtext understandingknowledge miningtext generationweb search
X
Xiang Zhang
Mashang Consumer Finance Co., Ltd., China
Hao Yu
Hao Yu
Hangzhou International Innovation Institute, Beihang University, Hangzhou 311115, China
deterministic networkingsemantic networkingXRcloud and edge continuum
W
Wenbo Fei
National-Mathematics Artificial Intelligence Institute in Chongqing (NMAII), China
C
Changjian Wang
Mashang Consumer Finance Co., Ltd., China
Z
Zhan Wang
National-Mathematics Artificial Intelligence Institute in Chongqing (NMAII), China
G
Guoqing Pang
National-Mathematics Artificial Intelligence Institute in Chongqing (NMAII), China
G
Guangye Yu
Mashang Consumer Finance Co., Ltd., China
Q
Quan Lu
Mashang Consumer Finance Co., Ltd., China
N
Ning Jiang
Mashang Consumer Finance Co., Ltd., China