Score
Designs and implements algorithms and tools that convert between SMILES strings and molecular graph representations, including parsing SMILES into graph structures and generating canonical or unique SMILES from graphs. Also builds canonicalization and alignment methods that map SMILES token positions to graph nodes to support consistent atom indexing, graph-to-sequence translation, and structurally grounded molecular embeddings.
This work addresses a critical limitation in current molecular large language models—their inadequate adherence to the principle that “structure determines function,” resulting in weak foundational structural understanding. To overcome this, the authors propose MolBasic, a novel framework that introduces a structure-first paradigm centered on bidirectional translation between SMILES strings and molecular graphs. By aligning topological and sequential representations through a multi-level structure-aware benchmark, and integrating progressive learning with standardized chain-of-thought prompting, MolBasic enhances the model’s capacity to evolve from basic structural perception to advanced reasoning. The approach achieves substantial improvements in structural comprehension and consistently delivers robust gains across downstream tasks, including molecular property prediction and target-oriented molecular optimization.
Existing SMILES pretraining models rely solely on single-token supervision, neglecting substructural semantics, and are trained only on corrupted SMILES strings—leading to weak supervisory signals and train-inference mismatch. To address these limitations, we propose SMI-Editor, an edit-based pretraining paradigm that randomly perturbs molecular substructures (rather than individual atoms or bonds) and reconstructs the original valid SMILES, thereby enabling fragment-level supervision and joint modeling of chemical validity. Built upon a Transformer architecture, SMI-Editor explicitly incorporates SMILES syntactic constraints and chemical substructure priors. This work is the first to introduce edit operations into molecular language modeling. Evaluated across multiple downstream tasks, SMI-Editor achieves state-of-the-art performance—outperforming several 3D-aware representation models—and significantly enhances molecular semantic understanding and generation capabilities.
Existing molecular-language alignment methods predominantly rely on graph neural networks to generate flat node-level tokens, neglecting molecules’ inherent hierarchical structure—such as functional groups—leading to semantic loss and hallucination in generation. To address this, we propose a novel hierarchical graph tokenization paradigm that explicitly models three semantic levels: atoms/nodes, motifs, and the full molecular graph. We construct the first graph–language supervised fine-tuning dataset incorporating explicit hierarchical annotations. Furthermore, we design a hierarchical graph tokenizer, a multi-granularity graph encoder, and a hierarchy-aware supervised fine-tuning strategy. Our approach achieves significant performance gains across seven molecular language understanding and generation tasks. Notably, hallucination rates decrease by 40%, while functional group comprehension and chemical semantic alignment are substantially improved.
Existing molecular-text cross-modal methods rely on global alignment, failing to capture fine-grained correspondences between molecular substructures (e.g., atoms or functional groups) and descriptive textual phrases, while being hindered by the scarcity of localized pairwise annotations. To address this, we propose a Hierarchical Adaptive Alignment (HAA) model that jointly aligns SMILES strings and text at three granularities—atomic, functional-group, and molecular levels. We further introduce the first end-to-end understanding-generation framework integrating a multimodal Transformer encoder, hierarchical attention mechanisms, contrastive learning, and generative pretraining (molecular captioning and SMILES generation). On retrieval tasks, our method achieves an average 30.8% improvement in Recall@1; it also establishes new state-of-the-art performance on both captioning and SMILES generation tasks. Visualization analyses confirm its chemical interpretability and fidelity to domain knowledge.
Large language models (LLMs) exhibit limited molecular structural understanding—especially when relying solely on one-dimensional textual representations like SMILES—hindering their effectiveness in chemistry. Method: We propose MolX, a lightweight multimodal extension module that jointly encodes SMILES sequences, 2D molecular graphs (via GNNs), and expert-crafted molecular fingerprints. MolX is trained via multitask contrastive learning while keeping the LLM backbone frozen. Contribution/Results: MolX establishes the first “frozen-LLM + multimodal alignment” paradigm, introducing only 0.53%–0.82% additional trainable parameters. It achieves significant improvements over baselines across four downstream tasks—including molecule-to-text translation and retrosynthetic planning—while supporting both zero-shot inference and fine-tuning deployment. This enhances cross-task generalization of LLMs in chemistry without architectural modification or full-parameter adaptation.
This work systematically evaluates the effectiveness of graph neural networks (GNNs) on small-molecule regression tasks and investigates the inductive biases inherent in different architectures—namely GCN, GraphSAGE, GIN, and GAT. The authors propose a hierarchical fusion strategy that integrates GNNs with molecular fingerprints, achieving a significant performance gain: the average RMSE is reduced by over 7% compared to using GNNs alone. For the first time, centered kernel alignment (CKA) is employed to analyze representation spaces, revealing that GNN-derived and fingerprint-based representations are highly independent in latent space (CKA ≤ 0.46), whereas representations from different GNN architectures exhibit strong convergence (CKA ≥ 0.88). These findings highlight both the shared learning mechanisms across GNN variants and their complementary potential when combined with traditional molecular descriptors.
This study addresses the limitations of conventional molecular representations—such as SMILES and IUPAC—in large language models (LLMs), which suffer from parsing difficulties, low generation accuracy, and poor robustness with complex structures. The work presents the first systematic evaluation of how different molecular representations affect LLM performance and introduces MolJSON, a novel representation grounded in explicit molecular graph structure. Experimental results across 78,045 samples using state-of-the-art models including GPT-5 and Claude Haiku 4.5 demonstrate that MolJSON substantially outperforms traditional formats in translation, constrained generation, and shortest-path reasoning tasks. Specifically, IUPAC-to-MolJSON translation achieves 71.0% accuracy (versus 43.7% for SMILES), constrained generation reaches 95.3% (compared to 64.0% for SMILES), and path reasoning attains 98.5% accuracy—all while requiring fewer inference tokens.
This work addresses the limitation of small language models (SLMs) in perceiving critical graph topological structures when predicting molecular properties from SMILES strings. To overcome this, the authors propose a context-augmented prompting framework that dynamically integrates, during inference, prediction prompts generated by graph neural networks (GNNs) with interpretable subgraphs, thereby enabling structure-aware zero-shot molecular property prediction for the first time. The approach synergistically combines GNNs, subgraph extraction, confidence estimation, and edge-ablation intervention analysis. Evaluated on the MUTAG and Tox21 datasets, the method achieves up to a 74% relative improvement in accuracy, demonstrating conclusively that incorporating graph-based contextual information significantly enhances the molecular understanding capabilities of small language models.
This study addresses the lack of a standardized textual representation for molecules in large language models (LLMs) and the frequent oversight of how representation choice critically impacts model performance. The authors systematically evaluate nine molecular representations—including SMILES, InChI, IUPAC, and CML—across eight chemical tasks using sixteen diverse LLMs, encompassing general-purpose, reasoning-enhanced, and chemistry-specific models. Performance is assessed through generation quality (via LLM-as-a-judge), alongside mechanistic analyses such as tokenization audits, linear probing, and attention mapping. The work reveals, for the first time, a strong dependence of representation efficacy on task type: IUPAC excels in semantic and generative correctness, structured formats are better suited for structural tasks, and CML demonstrates the strongest overall performance. Based on these findings, the authors propose a task-aware representation routing strategy, challenging the prevailing “representation-agnostic” evaluation paradigm and uncovering fundamental differences in how representations are encoded mechanistically.
This work proposes VQ-Atom, a novel molecular representation framework that addresses the limited chemical semantics of existing notations such as SMILES, which struggle to effectively capture local structural patterns. VQ-Atom introduces, for the first time, semantic-driven discretization into molecular representation by leveraging graph neural networks to learn atomic embeddings and integrating vector quantization to construct a chemistry-aware atomic codebook. This yields discrete semantic tokens with explicit chemical meaning, enabling the formulation of a molecular language suitable for Transformer-based pretraining. Notably, the method operates without requiring 3D structural information and demonstrates significant performance gains over conventional tokenization strategies in protein–ligand interaction prediction tasks, thereby validating the efficacy and advantage of semantically informed token representations in molecular language modeling.