SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance

📅 2025-08-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In natural language processing, static tokenization strategies constitute a fundamental bottleneck limiting model performance. This paper proposes SupraTok, a dynamic semantic tokenization architecture. Methodologically, it (1) learns multi-word “supertokens” in a cross-boundary manner to jointly optimize compression ratio and semantic integrity, and (2) incorporates entropy-driven data filtering and a multi-stage curriculum learning scheme to improve tokenizer training. Compared to mainstream subword methods—including BPE and WordPiece—SupraTok achieves an average 31% improvement in tokenization efficiency across 38 languages. It also substantially enhances downstream task performance: GPT-2 improves by 8.4% on HellaSwag and 9.5% on MMLU. This work is the first to systematically integrate semantic awareness and dynamic learning paradigms into tokenizer design, establishing a novel pathway toward efficient and interpretable language modeling.

Technology Category

Natural Language Processing: Lexical Semantics and MorphologyMachine Learning: Unsupervised & Self-Supervised LearningSearch and Optimization: Learning to Search

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphs
📝 Abstract
Tokenization remains a fundamental yet underexplored bottleneck in natural language processing, with strategies largely static despite remarkable progress in model architectures. We present SupraTok, a novel tokenization architecture that reimagines subword segmentation through three innovations: cross-boundary pattern learning that discovers multi-word semantic units, entropy-driven data curation that optimizes training corpus quality, and multi-phase curriculum learning for stable convergence. Our approach extends Byte-Pair Encoding by learning "superword" tokens, coherent multi-word expressions that preserve semantic unity while maximizing compression efficiency. SupraTok achieves 31% improvement in English tokenization efficiency (5.91 versus 4.51 characters per token) compared to OpenAI's o200k tokenizer and 30% improvement over Google's Gemma 3 tokenizer (256k vocabulary), while maintaining competitive performance across 38 languages. When integrated with a GPT-2 scale model (124M parameters) trained on 10 billion tokens from the FineWeb-Edu dataset, SupraTok yields 8.4% improvement on HellaSWAG and 9.5% on MMLU benchmarks without architectural modifications. While these results are promising at this scale, further validation at larger model scales is needed. These findings suggest that efficient tokenization can complement architectural innovations as a path to improved language model performance.
Problem

Research questions and friction points this paper is trying to address.

Improving tokenization efficiency in language models
Enhancing semantic preservation during text segmentation
Optimizing subword tokenization for cross-linguistic performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-boundary pattern learning for multi-word units
Entropy-driven data curation optimizing corpus quality
Multi-phase curriculum learning enabling stable convergence
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Ovidius University of Constanta
A
Andrei-Valentin Tănase
Faculty of Mathematics and Computer Science, "Ovidius" University of Constanta, Constanta, Romania
Elena Pelican
Elena Pelican
Assoc. Prof., Ovidius University of Constanta
Computer VisionMachine LearningNumerical Linear Algebra