Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational bottlenecks of byte-level language models caused by excessively long sequences and the absence of explicit textual abstraction. To overcome these limitations, this work proposes a tokenizer-free architecture that leverages a Token Hyper-position training strategy alongside hash embeddings, enabling standard Transformers to model efficiently at the byte level. The findings demonstrate that additional computation can effectively substitute fixed tokenizers, allowing models to spontaneously construct local context representation mechanisms. Furthermore, the proposed approach surpasses subword-based models in performance at large scales. Notably, by exploiting the non-uniform uncertainty inherent in generated outputs, the method facilitates speculative decoding, achieving a 3.4-fold improvement in acceptance rates.
📝 Abstract
Tokenizer-free language models remove the inductive bias of fixed tokenizers by modeling text directly as bytes, but the resulting longer sequences substantially increase computation and eliminate explicit text abstractions. We ask whether this additional computation can be useful, and whether standard Transformers can learn the abstractions that tokenization provides. We study these questions on Transformers without specialized tokenization-related architectures. With token-superposition training and hash embeddings, byte Transformers consistently outperform subword Transformers as model size scales. We further find that byte Transformers build local text abstractions as external tokenizers: a set of segmentation-like positions are used to collect local context representations, and restricting up to $25\%$ of intermediate layers to these local representations preserves downstream performance. Finally, these learned structures induce highly non-uniform generation difficulty, with uncertainty concentrated near local structure boundaries; exploiting them for speculative decoding yields $3.4\times$ more accepted tokens than in subword Transformers.
Problem

Research questions and friction points this paper is trying to address.

Byte Language Models
Tokenizer-free
Emergent Abstractions
Transformers
Innovation

Methods, ideas, or system contributions that make the work stand out.

Byte Language Models
Token-superposition Training
Hash Embeddings
Emergent Abstractions
Speculative Decoding
J
Jie Wang
School of Computer Science, East China Normal University, Shanghai, China
S
Shiwei Luo
School of Computer Science, East China Normal University, Shanghai, China
Qi Zhang
Qi Zhang
Fudan University
SAGINsatellite routing
Y
Yuanbin Wu
School of Computer Science, East China Normal University, Shanghai, China