Universal Byte-Level Encoding: UTF-8/UTF-16 Routing to Reduce Cross-Script Token-Budget Disparities

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the excessive token overhead and constrained context windows caused by UTF-8 encoding for non-English scripts in multilingual large language models. To this end, we propose a universal byte-level encoding scheme featuring a novel dual-alphabet dynamic routing mechanism that hybridizes UTF-8 and UTF-16 encodings to optimize multi-byte character processing. Crucially, this method modifies only the underlying byte representations without altering BPE merge rules, thereby achieving lossless and efficient cross-script encoding. Experimental results demonstrate that our approach substantially reduces token counts for high-premium scripts and enhances effective context utilization. Furthermore, it accelerates inference while preserving model quality, effectively mitigating cross-lingual performance disparities.
📝 Abstract
Byte-level byte-pair encoding (BBPE) tokenizers are attractive for multilingual large language models (LLMs) because they cover all Unicode text. In UTF-8-based BBPE, however, many scripts start from a higher fallback cost than English: when no learned merges can be applied, a multibyte character requires multiple byte-derived symbols. We call this worst-case pre-merge cost the encoding floor. A higher floor can increase token counts and per-request cost and shrink usable context. Changing the text encoding can reduce this gap, but a single global encoding can make already-efficient English spans more expensive in mixed-script text. We propose Universal Byte-Level Encoding (UBE), a dual-alphabet tokenizer that keeps 1-2-byte UTF-8 characters on the UTF-8 path while routing 3-4-byte UTF-8 characters through UTF-16. This lowers the encoding floor for 3-byte Basic Multilingual Plane (BMP) characters in scripts with high token premiums (token counts relative to English) without raising it for already-efficient spans in mixed-script text. UBE changes only the byte representation presented to byte-pair encoding (BPE); the merge rule remains standard, and exact decoding is preserved. UBE also composes with alternative boundary policies and morphology-based representations. In a Unicode 17 audit, UBE exactly round-trips all Unicode scalar values and all inputs in the official normalization, grapheme-break, and emoji test suites. Across intrinsic evaluations, UBE lowers dispersion in English-normalized token-count ratios, reducing cross-lingual token-budget disparity. In multilingual language model (LM) experiments, UBE matches BBPE's LM quality. In the main multilingual settings, UBE reduces token counts most for high-premium scripts and slightly lowers English token counts, yielding more usable context under fixed token budgets and faster prompt processing in content-matched benchmarks.
Problem

Research questions and friction points this paper is trying to address.

byte-level encoding
cross-script token disparity
multilingual tokenizer
encoding floor
BBPE
Innovation

Methods, ideas, or system contributions that make the work stand out.

Byte-level byte-pair encoding
Dual-alphabet tokenizer
UTF-8/UTF-16 routing
Cross-script token-budget disparity
Multilingual large language models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.