smiles subword tokenization

Designs and implements subword tokenization schemes for SMILES chemical strings by learning vocabularies from SMILES corpora, segmenting molecular strings into subword pieces, and applying pre-tokenization boundary policies. Controls token granularity and tokenization behavior through vocabulary size, merge rules, and pre-tokenization settings to produce token sequences suitable for downstream modeling or analysis.

smilessubwordtokenization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.48
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the lack of systematic comparison among subword tokenization algorithms in SMILES-based chemical language modeling, where unvalidated default schemes are commonly adopted. Under a fixed 165-token base vocabulary, controlled experiments across diverse chemical corpora and pre-tokenization strategies compare Byte Pair Encoding (BPE) and Unigram Language Modeling (Unigram-LM). The work reveals, for the first time, that the subword vocabularies generated by these two methods exhibit minimal overlap—Jaccard coefficients ≤0.161 overall and ≤0.05 among high-frequency tokens—and display systematic differences in segmentation granularity: Unigram-LM produces 29–41% more tokens on average, while BPE acts effectively as a coarsened variant of Unigram-LM for 80–99% of molecules. These findings demonstrate that the choice of tokenization algorithm is a critical design decision rather than a negligible default setting in molecular sequence modeling.

BPESMILESsubword segmentation

Large language models (LLMs) face a “tokenization bottleneck” in chemistry: general-purpose tokenizers fragment chemical representations—such as SMILES—into semantically incoherent subwords, compromising molecular structural integrity. To address this, we propose a vocabulary expansion method that systematically incorporates chemically significant tokens—including atoms, functional groups, and common substructures—thereby unifying the discretization of natural language and molecular representations. Our approach integrates targeted chemical-text continual pretraining without architectural modifications. By enhancing the tokenizer’s chemical expressivity and refining semantic alignment via lightweight pretraining, the model achieves markedly improved understanding of molecular semantics. Evaluated across six downstream tasks—including molecular property prediction, reaction classification, and scientific literature summarization—the method yields average performance gains of 4.2–12.7%. Results demonstrate its effectiveness, generalizability across diverse chemical NLP tasks, and deployment efficiency—requiring no inference-time overhead or model reengineering.

Addressing tokenization bottleneck in chemical representation learningResolving SMILES fragmentation in general-domain language modelsUnifying natural language and molecular structure representations

SMI-Editor: Edit-based SMILES Language Model with Fragment-level Supervision

Dec 07, 2024
KZ
Kangjie Zheng
🏛️ Peking University | Sichuan University | University of Washington

Existing SMILES pretraining models rely solely on single-token supervision, neglecting substructural semantics, and are trained only on corrupted SMILES strings—leading to weak supervisory signals and train-inference mismatch. To address these limitations, we propose SMI-Editor, an edit-based pretraining paradigm that randomly perturbs molecular substructures (rather than individual atoms or bonds) and reconstructs the original valid SMILES, thereby enabling fragment-level supervision and joint modeling of chemical validity. Built upon a Transformer architecture, SMI-Editor explicitly incorporates SMILES syntactic constraints and chemical substructure priors. This work is the first to introduce edit operations into molecular language modeling. Evaluated across multiple downstream tasks, SMI-Editor achieves state-of-the-art performance—outperforming several 3D-aware representation models—and significantly enhances molecular semantic understanding and generation capabilities.

Current models suffer from train-inference mismatch with invalid SMILES.Existing SMILES LMs lack fragment-level molecular supervision.SMI-Editor improves molecular representation via edit-based fragment reconstruction.

MolX: Enhancing Large Language Models for Molecular Learning with A Multi-Modal Extension

Jun 10, 2024
KL
Khiem Le
🏛️ University of Notre Dame | University of California, Los Angeles

Large language models (LLMs) exhibit limited molecular structural understanding—especially when relying solely on one-dimensional textual representations like SMILES—hindering their effectiveness in chemistry. Method: We propose MolX, a lightweight multimodal extension module that jointly encodes SMILES sequences, 2D molecular graphs (via GNNs), and expert-crafted molecular fingerprints. MolX is trained via multitask contrastive learning while keeping the LLM backbone frozen. Contribution/Results: MolX establishes the first “frozen-LLM + multimodal alignment” paradigm, introducing only 0.53%–0.82% additional trainable parameters. It achieves significant improvements over baselines across four downstream tasks—including molecule-to-text translation and retrosynthetic planning—while supporting both zero-shot inference and fine-tuning deployment. This enhances cross-task generalization of LLMs in chemistry without architectural modification or full-parameter adaptation.

Enhancing LLMs for molecular learning with multi-modal inputsImproving performance on molecule-related tasks with minimal parametersOvercoming SMILES limitations in molecular representation

SmileyLlama: Modifying Large Language Models for Directed Chemical Space Exploration

Sep 03, 2024
JM
Joseph M. Cavanagh
🏛️ University of California, Berkeley | University of Wisconsin–Madison | The Herbert Wertheim UF Scripps Institute for Biomedical Innovation and Technology | Lawrence Berkeley National Laboratory

This work addresses the challenge of transforming general-purpose large language models (LLMs) into attribute-controllable molecular generators. We propose a lightweight adaptation paradigm that converts open-source Llama models into chemical language models (CLMs) via supervised fine-tuning (SFT) and direct preference optimization (DPO), enabling direct SMILES string generation conditioned on multidimensional physicochemical properties (e.g., logP, aqueous solubility). To our knowledge, this is the first empirical demonstration that an adapted general LLM achieves performance on multi-objective molecular generation tasks comparable to or exceeding that of domain-specific chemically pretrained models. The approach enables a paradigm shift from “chemical knowledge question-answering” to “property-directed molecular design,” significantly enhancing controllability, interpretability, and interactive exploration of chemical space.

Generating valid novel drug-like molecules efficientlyOptimizing molecules for 3D conformation binding affinityTransforming general LLM into chemical language model

Latest Papers

What's happening recently
View more

This study addresses the limited understanding of how chemical language models (CLMs) encode chemically meaningful molecular substructures during pretraining and fine-tuning. For the first time, it systematically evaluates the substructure awareness of eight pretrained and six randomly initialized CLMs across 78 distinct substructures, employing probing techniques to analyze how SMILES sequence representations evolve across model layers, complemented by downstream task fine-tuning experiments. The findings reveal that pretraining substantially enhances models’ comprehension of high-level molecular structures, while fine-tuning selectively strengthens representations of task-relevant substructures. Notably, even randomly initialized models effectively encode cyclic structures in their initial layers. This work elucidates the dynamic mechanisms underlying substructure representation in CLMs and offers novel insights into molecular representation learning.

Chemical Language ModelsFine-tuningMolecular Substructures

This work addresses the challenge that chemical pre-trained models often forget natural language semantics while learning SMILES syntax and struggle to jointly comprehend molecular structures and textual descriptions. To overcome this, the authors propose CheMatE, a model built upon the ModernBERT architecture and trained in two stages: first, masked language modeling on hundreds of billions of SMILES-annotated scientific texts, followed by Matryoshka contrastive learning and Multiple Negative Ranking Loss optimization using synthetic SMILES–text pairs to construct a shared bilingual semantic space. This approach effectively mitigates semantic forgetting and significantly enhances cross-modal understanding, achieving strong performance on both molecular property prediction and scientific language comprehension tasks, while demonstrating robust transferability and competitive generalization capabilities.

domain adaptationnatural languageoverfitting

This work addresses the limitation of small language models (SLMs) in perceiving critical graph topological structures when predicting molecular properties from SMILES strings. To overcome this, the authors propose a context-augmented prompting framework that dynamically integrates, during inference, prediction prompts generated by graph neural networks (GNNs) with interpretable subgraphs, thereby enabling structure-aware zero-shot molecular property prediction for the first time. The approach synergistically combines GNNs, subgraph extraction, confidence estimation, and edge-ablation intervention analysis. Evaluated on the MUTAG and Tox21 datasets, the method achieves up to a 74% relative improvement in accuracy, demonstrating conclusively that incorporating graph-based contextual information significantly enhances the molecular understanding capabilities of small language models.

graph topologymolecular property predictionsmall language models

This work addresses the challenge that molecular language models face when processing SMILES strings: character-level tokenization disrupts chemically meaningful local substructures, hindering effective modeling of long-range dependencies. To overcome this limitation, the authors propose MolGram, a conditional n-gram memory module that maps recurring local string patterns to learnable embeddings without altering the standard tokenizer. These pattern embeddings are dynamically injected into the Transformer’s hidden states via a context-aware mechanism, introducing explicit local structural memory as an efficient inductive bias. Evaluated across unconditional molecular generation, forward reaction prediction, and single-step retrosynthesis tasks, MolGram consistently outperforms baseline models—achieving results comparable to or better than those of models with three times its parameter count.

chemically meaningful motifslocality gaplong-range dependencies

This study addresses the lack of a standardized textual representation for molecules in large language models (LLMs) and the frequent oversight of how representation choice critically impacts model performance. The authors systematically evaluate nine molecular representations—including SMILES, InChI, IUPAC, and CML—across eight chemical tasks using sixteen diverse LLMs, encompassing general-purpose, reasoning-enhanced, and chemistry-specific models. Performance is assessed through generation quality (via LLM-as-a-judge), alongside mechanistic analyses such as tokenization audits, linear probing, and attention mapping. The work reveals, for the first time, a strong dependence of representation efficacy on task type: IUPAC excels in semantic and generative correctness, structured formats are better suited for structural tasks, and CML demonstrates the strongest overall performance. Based on these findings, the authors propose a task-aware representation routing strategy, challenging the prevailing “representation-agnostic” evaluation paradigm and uncovering fundamental differences in how representations are encoded mechanistically.

chemical taskslarge language modelsmolecular representation

Hot Scholars

YM

Youssef Mroueh

Principal Research Scientist, IBM T.J Watson Research Center
Machine learningArtificial intelligence
PD

Payel Das

Manager and Principal Research Staff Member, AI research, IBM Watson, NY
trustworthy MLgenerative AIbio-inspired AIAI4Science
YL

Yuqiang Li

Central South University
Internal Combustion EngineCombustionEmissionsMechansim
TF

Tianfan Fu

Nanjing University
AI for DrugAI for ScienceLarge Language Model
AC

Achuth Chandrasekhar

Graduate Student, Carnegie Mellon University
Additive ManufacturingDeep Learning