smiles-graph translation

Designs and implements algorithms and tools that convert between SMILES strings and molecular graph representations, including parsing SMILES into graph structures and generating canonical or unique SMILES from graphs. Also builds canonicalization and alignment methods that map SMILES token positions to graph nodes to support consistent atom indexing, graph-to-sequence translation, and structurally grounded molecular embeddings.

smiles-graphtranslation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses a critical limitation in current molecular large language models—their inadequate adherence to the principle that “structure determines function,” resulting in weak foundational structural understanding. To overcome this, the authors propose MolBasic, a novel framework that introduces a structure-first paradigm centered on bidirectional translation between SMILES strings and molecular graphs. By aligning topological and sequential representations through a multi-level structure-aware benchmark, and integrating progressive learning with standardized chain-of-thought prompting, MolBasic enhances the model’s capacity to evolve from basic structural perception to advanced reasoning. The approach achieves substantial improvements in structural comprehension and consistently delivers robust gains across downstream tasks, including molecular property prediction and target-oriented molecular optimization.

large language modelsmolecular graphmolecular understanding

SMI-Editor: Edit-based SMILES Language Model with Fragment-level Supervision

Dec 07, 2024
KZ
Kangjie Zheng
🏛️ Peking University | Sichuan University | University of Washington

Existing SMILES pretraining models rely solely on single-token supervision, neglecting substructural semantics, and are trained only on corrupted SMILES strings—leading to weak supervisory signals and train-inference mismatch. To address these limitations, we propose SMI-Editor, an edit-based pretraining paradigm that randomly perturbs molecular substructures (rather than individual atoms or bonds) and reconstructs the original valid SMILES, thereby enabling fragment-level supervision and joint modeling of chemical validity. Built upon a Transformer architecture, SMI-Editor explicitly incorporates SMILES syntactic constraints and chemical substructure priors. This work is the first to introduce edit operations into molecular language modeling. Evaluated across multiple downstream tasks, SMI-Editor achieves state-of-the-art performance—outperforming several 3D-aware representation models—and significantly enhances molecular semantic understanding and generation capabilities.

Current models suffer from train-inference mismatch with invalid SMILES.Existing SMILES LMs lack fragment-level molecular supervision.SMI-Editor improves molecular representation via edit-based fragment reconstruction.

HIGHT: Hierarchical Graph Tokenization for Graph-Language Alignment

Jun 20, 2024
YC
Yongqiang Chen
🏛️ The Chinese University of Hong Kong | Tsinghua University | Tencent AI Lab

Existing molecular-language alignment methods predominantly rely on graph neural networks to generate flat node-level tokens, neglecting molecules’ inherent hierarchical structure—such as functional groups—leading to semantic loss and hallucination in generation. To address this, we propose a novel hierarchical graph tokenization paradigm that explicitly models three semantic levels: atoms/nodes, motifs, and the full molecular graph. We construct the first graph–language supervised fine-tuning dataset incorporating explicit hierarchical annotations. Furthermore, we design a hierarchical graph tokenizer, a multi-granularity graph encoder, and a hierarchy-aware supervised fine-tuning strategy. Our approach achieves significant performance gains across seven molecular language understanding and generation tasks. Notably, hallucination rates decrease by 40%, while functional group comprehension and chemical semantic alignment are substantially improved.

Existing methods ignore molecular hierarchical structures in tokenizationNeglecting hierarchy causes poor molecule-language alignment and hallucinationProposing hierarchical tokenization to enhance molecular perception in LLMs

Atomas: Hierarchical Alignment on Molecule-Text for Unified Molecule Understanding and Generation

Apr 23, 2024
YZ
Yikun Zhang
🏛️ Peking University | Tencent AI Lab | Tsinghua University | Hong Kong Baptist University

Existing molecular-text cross-modal methods rely on global alignment, failing to capture fine-grained correspondences between molecular substructures (e.g., atoms or functional groups) and descriptive textual phrases, while being hindered by the scarcity of localized pairwise annotations. To address this, we propose a Hierarchical Adaptive Alignment (HAA) model that jointly aligns SMILES strings and text at three granularities—atomic, functional-group, and molecular levels. We further introduce the first end-to-end understanding-generation framework integrating a multimodal Transformer encoder, hierarchical attention mechanisms, contrastive learning, and generative pretraining (molecular captioning and SMILES generation). On retrieval tasks, our method achieves an average 30.8% improvement in Recall@1; it also establishes new state-of-the-art performance on both captioning and SMILES generation tasks. Visualization analyses confirm its chemical interpretability and fidelity to domain knowledge.

Addresses limitations of global alignment in capturing fine-grained molecular-text information.Enhances molecular representation quality for drug discovery and materials science.Proposes a framework for joint learning from SMILES strings and text.

MolX: Enhancing Large Language Models for Molecular Learning with A Multi-Modal Extension

Jun 10, 2024
KL
Khiem Le
🏛️ University of Notre Dame | University of California, Los Angeles

Large language models (LLMs) exhibit limited molecular structural understanding—especially when relying solely on one-dimensional textual representations like SMILES—hindering their effectiveness in chemistry. Method: We propose MolX, a lightweight multimodal extension module that jointly encodes SMILES sequences, 2D molecular graphs (via GNNs), and expert-crafted molecular fingerprints. MolX is trained via multitask contrastive learning while keeping the LLM backbone frozen. Contribution/Results: MolX establishes the first “frozen-LLM + multimodal alignment” paradigm, introducing only 0.53%–0.82% additional trainable parameters. It achieves significant improvements over baselines across four downstream tasks—including molecule-to-text translation and retrosynthetic planning—while supporting both zero-shot inference and fine-tuning deployment. This enhances cross-task generalization of LLMs in chemistry without architectural modification or full-parameter adaptation.

Enhancing LLMs for molecular learning with multi-modal inputsImproving performance on molecule-related tasks with minimal parametersOvercoming SMILES limitations in molecular representation

Latest Papers

What's happening recently
View more

This work systematically evaluates the effectiveness of graph neural networks (GNNs) on small-molecule regression tasks and investigates the inductive biases inherent in different architectures—namely GCN, GraphSAGE, GIN, and GAT. The authors propose a hierarchical fusion strategy that integrates GNNs with molecular fingerprints, achieving a significant performance gain: the average RMSE is reduced by over 7% compared to using GNNs alone. For the first time, centered kernel alignment (CKA) is employed to analyze representation spaces, revealing that GNN-derived and fingerprint-based representations are highly independent in latent space (CKA ≤ 0.46), whereas representations from different GNN architectures exhibit strong convergence (CKA ≥ 0.88). These findings highlight both the shared learning mechanisms across GNN variants and their complementary potential when combined with traditional molecular descriptors.

Fingerprint EmbeddingsGraph Neural NetworksMolecular Regression

This study addresses the limitations of conventional molecular representations—such as SMILES and IUPAC—in large language models (LLMs), which suffer from parsing difficulties, low generation accuracy, and poor robustness with complex structures. The work presents the first systematic evaluation of how different molecular representations affect LLM performance and introduces MolJSON, a novel representation grounded in explicit molecular graph structure. Experimental results across 78,045 samples using state-of-the-art models including GPT-5 and Claude Haiku 4.5 demonstrate that MolJSON substantially outperforms traditional formats in translation, constrained generation, and shortest-path reasoning tasks. Specifically, IUPAC-to-MolJSON translation achieves 71.0% accuracy (versus 43.7% for SMILES), constrained generation reaches 95.3% (compared to 64.0% for SMILES), and path reasoning attains 98.5% accuracy—all while requiring fewer inference tokens.

IUPAC nameslarge language modelsmolecular graph

This work addresses the limitation of small language models (SLMs) in perceiving critical graph topological structures when predicting molecular properties from SMILES strings. To overcome this, the authors propose a context-augmented prompting framework that dynamically integrates, during inference, prediction prompts generated by graph neural networks (GNNs) with interpretable subgraphs, thereby enabling structure-aware zero-shot molecular property prediction for the first time. The approach synergistically combines GNNs, subgraph extraction, confidence estimation, and edge-ablation intervention analysis. Evaluated on the MUTAG and Tox21 datasets, the method achieves up to a 74% relative improvement in accuracy, demonstrating conclusively that incorporating graph-based contextual information significantly enhances the molecular understanding capabilities of small language models.

graph topologymolecular property predictionsmall language models

This study addresses the lack of a standardized textual representation for molecules in large language models (LLMs) and the frequent oversight of how representation choice critically impacts model performance. The authors systematically evaluate nine molecular representations—including SMILES, InChI, IUPAC, and CML—across eight chemical tasks using sixteen diverse LLMs, encompassing general-purpose, reasoning-enhanced, and chemistry-specific models. Performance is assessed through generation quality (via LLM-as-a-judge), alongside mechanistic analyses such as tokenization audits, linear probing, and attention mapping. The work reveals, for the first time, a strong dependence of representation efficacy on task type: IUPAC excels in semantic and generative correctness, structured formats are better suited for structural tasks, and CML demonstrates the strongest overall performance. Based on these findings, the authors propose a task-aware representation routing strategy, challenging the prevailing “representation-agnostic” evaluation paradigm and uncovering fundamental differences in how representations are encoded mechanistically.

chemical taskslarge language modelsmolecular representation

This work proposes VQ-Atom, a novel molecular representation framework that addresses the limited chemical semantics of existing notations such as SMILES, which struggle to effectively capture local structural patterns. VQ-Atom introduces, for the first time, semantic-driven discretization into molecular representation by leveraging graph neural networks to learn atomic embeddings and integrating vector quantization to construct a chemistry-aware atomic codebook. This yields discrete semantic tokens with explicit chemical meaning, enabling the formulation of a molecular language suitable for Transformer-based pretraining. Notably, the method operates without requiring 3D structural information and demonstrates significant performance gains over conventional tokenization strategies in protein–ligand interaction prediction tasks, thereby validating the efficacy and advantage of semantically informed token representations in molecular language modeling.

atom-level representationchemical substructuresmolecular representation learning

Hot Scholars

PS

Philippe Schwaller

Assistant Professor, Laboratory of Artificial Chemical Intelligence - EPFL
Deep LearningML for ChemistryReaction PredictionSynthesis Planning
ZL

Zequn Liu

Microsoft Research AI4Science, Asia
ZX

Zhiping Xiao

Postdoc at University of Washington
CSEDMML
KZ

Kangjie Zheng

Wellcome Sanger Institute
AI4ScienceNLPLarge Language Model