๐ค AI Summary
This study addresses the hidden costs of pre-tokenization boundary rules on text compression efficiency and their unclear coupling with model prediction quality. To quantify this compression overhead, it computes lower bounds on token counts via shortest-path algorithms and linear programming relaxations, introducing a non-negative pricing mechanism to generate independently verifiable compression certificates. Furthermore, a novel โboundary permitโ strategy is proposed to explore intermediate trade-offs. Experiments spanning BPE, regex-based boundaries, and twelve language pairs reveal that boundary constraints inflate optimal token counts by 28.3%โ36.8%. The results demonstrate that unconstrained fitting significantly improves predictive bitrates across multilingual settings, while allocating merely 10% of the vocabulary budget suffices to recover most compression gains.
๐ Abstract
Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries. We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule. Nonnegative prices on token occurrences yield a lower bound through shortest paths and vocabulary-budget selection; maximising over all prices recovers the linear programming relaxation, and an independent integer checker certifies the reported values. On English Wikipedia, boundaries increase the optimal token count by 28.3--36.8\%. Byte pair encoding lies 2.1\% above the constrained lower bound, but 10.9\% above the unrestricted bound. Compression and prediction favour different dictionaries: at 85M non-embedding parameters and matched training-token budgets, unrestricted fitting yields higher mean held-out bits per byte under a common unrestricted decoder in all 12 languages in the paired study and 11 of 12 under independent tuning and evaluation. To study intermediate boundary policies, we introduce boundary licences, which limit the vocabulary entries permitted to cross cuts and admit the same form of certificate. On separate English and Chinese fitting corpora, licensing 10\% of the vocabulary budget recovers 85.2\% and 100.0\% of the achieved token-count reduction from removing all cuts. These results quantify the compression cost of boundaries while separating it from the prediction quality of the resulting token units.