More Than Words: Compositional Tokenization for Efficient Language Models

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem that standard tokenization yields excessively long sequences, substantially increasing inference costs for language models. To mitigate this, we propose Compositional Byte Pair Encoding (CoBPE), which maps phrases into combinations of base words and modifiers. CoBPE achieves structured representations through feature fusion in the embedding space and employs a joint output decoding mechanism to replace partial token sequences. This work establishes a new paradigm for efficient language model design. Experimental results demonstrate that CoBPE reduces sequence length by 30% while improving average performance on downstream tasks by 1.2 points, effectively balancing computational efficiency with model effectiveness.
📝 Abstract
Language models process and generate text sequentially in token units, and the tokenizer determines how much text each inference step covers. Under standard tokenization, a short English phrase such as"On the table."is usually produced as four separate predictions for the preposition (On), article (the), noun (table), and punctuation (.), where each consumes a sequence position and adds inference cost. We introduce CoBPE, a compositional tokenization approach that represents such phrases as a lexical base token (table) attached with a small set of reusable surface modifiers, composed in embedding space at input and predicted jointly at output. In controlled pretraining from scratch at 780M and 1.3B scales, CoBPE shortens sequences by 30% and improves average downstream performance by 1.2 points relative to standard BPE under matched training compute. Our results suggest that part of what is now expressed through token sequences can instead be modeled through structured representations, opening a broad design space for more token-efficient and capable language models.
Problem

Research questions and friction points this paper is trying to address.

tokenization
language models
inference cost
sequence length
token efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compositional Tokenization
CoBPE
Language Models
Token Efficiency
Structured Representations
🔎 Similar Papers
No similar papers found.