The Tokenization Bottleneck: How Vocabulary Extension Improves Chemistry Representation Learning in Pretrained Language Models

📅 2025-11-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large language models (LLMs) face a “tokenization bottleneck” in chemistry: general-purpose tokenizers fragment chemical representations—such as SMILES—into semantically incoherent subwords, compromising molecular structural integrity. To address this, we propose a vocabulary expansion method that systematically incorporates chemically significant tokens—including atoms, functional groups, and common substructures—thereby unifying the discretization of natural language and molecular representations. Our approach integrates targeted chemical-text continual pretraining without architectural modifications. By enhancing the tokenizer’s chemical expressivity and refining semantic alignment via lightweight pretraining, the model achieves markedly improved understanding of molecular semantics. Evaluated across six downstream tasks—including molecular property prediction, reaction classification, and scientific literature summarization—the method yields average performance gains of 4.2–12.7%. Results demonstrate its effectiveness, generalizability across diverse chemical NLP tasks, and deployment efficiency—requiring no inference-time overhead or model reengineering.

Technology Category

Natural Language Processing: (Large) Language ModelsMachine Learning: Large Multimodal Models (LMMs)Computer Vision: Large Vision Models

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSearch and Retrieval-Augmented AI: Large language models for search
📝 Abstract
The application of large language models (LLMs) to chemistry is frequently hampered by a "tokenization bottleneck", where tokenizers tuned on general-domain text tend to fragment chemical representations such as SMILES into semantically uninformative sub-tokens. This paper introduces a principled methodology to resolve this bottleneck by unifying the representation of natural language and molecular structures within a single model. Our approach involves targeted vocabulary extension-augmenting a pretrained LLM's vocabulary with chemically salient tokens, followed by continued pretraining on chemistry-domain text to integrate this new knowledge. We provide an empirical demonstration of the effectiveness of this strategy, showing that our methodology leads to superior performance on a range of downstream chemical tasks.
Problem

Research questions and friction points this paper is trying to address.

Addressing tokenization bottleneck in chemical representation learning
Resolving SMILES fragmentation in general-domain language models
Unifying natural language and molecular structure representations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Extending vocabulary with chemically salient tokens
Unifying natural language and molecular representations
Continued pretraining on chemistry-domain text
💼 Related Jobs
No related jobs found.
P
Prathamesh Kalamkar
Thoughtworks
Ned Letcher
Ned Letcher
Thoughtworks
M
Meissane Chami
Thoughtworks
S
Sahger Lad
Thoughtworks
S
Shayan Mohanty
Thoughtworks
P
Prasanna Pendse
Thoughtworks