Score
Designs and implements subword tokenization schemes for SMILES chemical strings by learning vocabularies from SMILES corpora, segmenting molecular strings into subword pieces, and applying pre-tokenization boundary policies. Controls token granularity and tokenization behavior through vocabulary size, merge rules, and pre-tokenization settings to produce token sequences suitable for downstream modeling or analysis.
This study addresses the lack of systematic comparison among subword tokenization algorithms in SMILES-based chemical language modeling, where unvalidated default schemes are commonly adopted. Under a fixed 165-token base vocabulary, controlled experiments across diverse chemical corpora and pre-tokenization strategies compare Byte Pair Encoding (BPE) and Unigram Language Modeling (Unigram-LM). The work reveals, for the first time, that the subword vocabularies generated by these two methods exhibit minimal overlap—Jaccard coefficients ≤0.161 overall and ≤0.05 among high-frequency tokens—and display systematic differences in segmentation granularity: Unigram-LM produces 29–41% more tokens on average, while BPE acts effectively as a coarsened variant of Unigram-LM for 80–99% of molecules. These findings demonstrate that the choice of tokenization algorithm is a critical design decision rather than a negligible default setting in molecular sequence modeling.
Large language models (LLMs) face a “tokenization bottleneck” in chemistry: general-purpose tokenizers fragment chemical representations—such as SMILES—into semantically incoherent subwords, compromising molecular structural integrity. To address this, we propose a vocabulary expansion method that systematically incorporates chemically significant tokens—including atoms, functional groups, and common substructures—thereby unifying the discretization of natural language and molecular representations. Our approach integrates targeted chemical-text continual pretraining without architectural modifications. By enhancing the tokenizer’s chemical expressivity and refining semantic alignment via lightweight pretraining, the model achieves markedly improved understanding of molecular semantics. Evaluated across six downstream tasks—including molecular property prediction, reaction classification, and scientific literature summarization—the method yields average performance gains of 4.2–12.7%. Results demonstrate its effectiveness, generalizability across diverse chemical NLP tasks, and deployment efficiency—requiring no inference-time overhead or model reengineering.
Existing SMILES pretraining models rely solely on single-token supervision, neglecting substructural semantics, and are trained only on corrupted SMILES strings—leading to weak supervisory signals and train-inference mismatch. To address these limitations, we propose SMI-Editor, an edit-based pretraining paradigm that randomly perturbs molecular substructures (rather than individual atoms or bonds) and reconstructs the original valid SMILES, thereby enabling fragment-level supervision and joint modeling of chemical validity. Built upon a Transformer architecture, SMI-Editor explicitly incorporates SMILES syntactic constraints and chemical substructure priors. This work is the first to introduce edit operations into molecular language modeling. Evaluated across multiple downstream tasks, SMI-Editor achieves state-of-the-art performance—outperforming several 3D-aware representation models—and significantly enhances molecular semantic understanding and generation capabilities.
Large language models (LLMs) exhibit limited molecular structural understanding—especially when relying solely on one-dimensional textual representations like SMILES—hindering their effectiveness in chemistry. Method: We propose MolX, a lightweight multimodal extension module that jointly encodes SMILES sequences, 2D molecular graphs (via GNNs), and expert-crafted molecular fingerprints. MolX is trained via multitask contrastive learning while keeping the LLM backbone frozen. Contribution/Results: MolX establishes the first “frozen-LLM + multimodal alignment” paradigm, introducing only 0.53%–0.82% additional trainable parameters. It achieves significant improvements over baselines across four downstream tasks—including molecule-to-text translation and retrosynthetic planning—while supporting both zero-shot inference and fine-tuning deployment. This enhances cross-task generalization of LLMs in chemistry without architectural modification or full-parameter adaptation.
This work addresses the challenge of transforming general-purpose large language models (LLMs) into attribute-controllable molecular generators. We propose a lightweight adaptation paradigm that converts open-source Llama models into chemical language models (CLMs) via supervised fine-tuning (SFT) and direct preference optimization (DPO), enabling direct SMILES string generation conditioned on multidimensional physicochemical properties (e.g., logP, aqueous solubility). To our knowledge, this is the first empirical demonstration that an adapted general LLM achieves performance on multi-objective molecular generation tasks comparable to or exceeding that of domain-specific chemically pretrained models. The approach enables a paradigm shift from “chemical knowledge question-answering” to “property-directed molecular design,” significantly enhancing controllability, interpretability, and interactive exploration of chemical space.
This study addresses the limited understanding of how chemical language models (CLMs) encode chemically meaningful molecular substructures during pretraining and fine-tuning. For the first time, it systematically evaluates the substructure awareness of eight pretrained and six randomly initialized CLMs across 78 distinct substructures, employing probing techniques to analyze how SMILES sequence representations evolve across model layers, complemented by downstream task fine-tuning experiments. The findings reveal that pretraining substantially enhances models’ comprehension of high-level molecular structures, while fine-tuning selectively strengthens representations of task-relevant substructures. Notably, even randomly initialized models effectively encode cyclic structures in their initial layers. This work elucidates the dynamic mechanisms underlying substructure representation in CLMs and offers novel insights into molecular representation learning.
This work addresses the challenge that chemical pre-trained models often forget natural language semantics while learning SMILES syntax and struggle to jointly comprehend molecular structures and textual descriptions. To overcome this, the authors propose CheMatE, a model built upon the ModernBERT architecture and trained in two stages: first, masked language modeling on hundreds of billions of SMILES-annotated scientific texts, followed by Matryoshka contrastive learning and Multiple Negative Ranking Loss optimization using synthetic SMILES–text pairs to construct a shared bilingual semantic space. This approach effectively mitigates semantic forgetting and significantly enhances cross-modal understanding, achieving strong performance on both molecular property prediction and scientific language comprehension tasks, while demonstrating robust transferability and competitive generalization capabilities.
This work addresses the limitation of small language models (SLMs) in perceiving critical graph topological structures when predicting molecular properties from SMILES strings. To overcome this, the authors propose a context-augmented prompting framework that dynamically integrates, during inference, prediction prompts generated by graph neural networks (GNNs) with interpretable subgraphs, thereby enabling structure-aware zero-shot molecular property prediction for the first time. The approach synergistically combines GNNs, subgraph extraction, confidence estimation, and edge-ablation intervention analysis. Evaluated on the MUTAG and Tox21 datasets, the method achieves up to a 74% relative improvement in accuracy, demonstrating conclusively that incorporating graph-based contextual information significantly enhances the molecular understanding capabilities of small language models.
This work addresses the challenge that molecular language models face when processing SMILES strings: character-level tokenization disrupts chemically meaningful local substructures, hindering effective modeling of long-range dependencies. To overcome this limitation, the authors propose MolGram, a conditional n-gram memory module that maps recurring local string patterns to learnable embeddings without altering the standard tokenizer. These pattern embeddings are dynamically injected into the Transformer’s hidden states via a context-aware mechanism, introducing explicit local structural memory as an efficient inductive bias. Evaluated across unconditional molecular generation, forward reaction prediction, and single-step retrosynthesis tasks, MolGram consistently outperforms baseline models—achieving results comparable to or better than those of models with three times its parameter count.
This study addresses the lack of a standardized textual representation for molecules in large language models (LLMs) and the frequent oversight of how representation choice critically impacts model performance. The authors systematically evaluate nine molecular representations—including SMILES, InChI, IUPAC, and CML—across eight chemical tasks using sixteen diverse LLMs, encompassing general-purpose, reasoning-enhanced, and chemistry-specific models. Performance is assessed through generation quality (via LLM-as-a-judge), alongside mechanistic analyses such as tokenization audits, linear probing, and attention mapping. The work reveals, for the first time, a strong dependence of representation efficacy on task type: IUPAC excels in semantic and generative correctness, structured formats are better suited for structural tasks, and CML demonstrates the strongest overall performance. Based on these findings, the authors propose a task-aware representation routing strategy, challenging the prevailing “representation-agnostic” evaluation paradigm and uncovering fundamental differences in how representations are encoded mechanistically.