đ¤ AI Summary
Large language models (LLMs) face a âtokenization bottleneckâ in chemistry: general-purpose tokenizers fragment chemical representationsâsuch as SMILESâinto semantically incoherent subwords, compromising molecular structural integrity. To address this, we propose a vocabulary expansion method that systematically incorporates chemically significant tokensâincluding atoms, functional groups, and common substructuresâthereby unifying the discretization of natural language and molecular representations. Our approach integrates targeted chemical-text continual pretraining without architectural modifications. By enhancing the tokenizerâs chemical expressivity and refining semantic alignment via lightweight pretraining, the model achieves markedly improved understanding of molecular semantics. Evaluated across six downstream tasksâincluding molecular property prediction, reaction classification, and scientific literature summarizationâthe method yields average performance gains of 4.2â12.7%. Results demonstrate its effectiveness, generalizability across diverse chemical NLP tasks, and deployment efficiencyârequiring no inference-time overhead or model reengineering.
đ Abstract
The application of large language models (LLMs) to chemistry is frequently hampered by a "tokenization bottleneck", where tokenizers tuned on general-domain text tend to fragment chemical representations such as SMILES into semantically uninformative sub-tokens. This paper introduces a principled methodology to resolve this bottleneck by unifying the representation of natural language and molecular structures within a single model. Our approach involves targeted vocabulary extension-augmenting a pretrained LLM's vocabulary with chemically salient tokens, followed by continued pretraining on chemistry-domain text to integrate this new knowledge. We provide an empirical demonstration of the effectiveness of this strategy, showing that our methodology leads to superior performance on a range of downstream chemical tasks.