Bi-semantic Chemical Embedder for Joint Representation Learning of SMILES and Natural Language

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that chemical pre-trained models often forget natural language semantics while learning SMILES syntax and struggle to jointly comprehend molecular structures and textual descriptions. To overcome this, the authors propose CheMatE, a model built upon the ModernBERT architecture and trained in two stages: first, masked language modeling on hundreds of billions of SMILES-annotated scientific texts, followed by Matryoshka contrastive learning and Multiple Negative Ranking Loss optimization using synthetic SMILES–text pairs to construct a shared bilingual semantic space. This approach effectively mitigates semantic forgetting and significantly enhances cross-modal understanding, achieving strong performance on both molecular property prediction and scientific language comprehension tasks, while demonstrating robust transferability and competitive generalization capabilities.
📝 Abstract
Transformer models have revolutionized natural language processing (NLP), and text-based molecular representations like SMILES have successfully extended these architectures to chemistry. However, domain-adaptive pre-training often causes models to overfit to chemical syntax, catastrophically forgetting their foundational semantic capabilities. To address this challenge, we introduce CheMatE, a chemistry-oriented embedding model that jointly captures molecular structure and domain-specific natural language within the same representation space. Built on a ModernBERT backbone, CheMatE learns bi-semantic representations through a two-stage training procedure: continued masked language modeling (MLM) followed by a Matryoshka contrastive learning stage via Multiple Negative Ranking Loss (MNRL). First, we train the model using MLM on a novel, large-scale corpus of SMILES-annotated, long-context scientific documents that were constructed and curated from FineWeb and ChemPile (comprising 10.4B and 11.5B tokens, respectively). Subsequently, the model undergoes contrastive learning using a synthetic dataset of SMILES-text pairs algorithmically derived from our original training corpus. This design exposes the model to SMILES-enriched scientific literature, enabling bi-semantic understanding. We evaluate CheMatE across a range of downstream tasks covering molecular property prediction and scientific language understanding. Our results demonstrate that coupling our custom-curated datasets with this sequential training strategy yields robust, highly transferable representations. By effectively unifying structural and contextual signals within a single text-based framework, CheMatE achieves competitive performance across both specialized chemistry models and general-purpose language model baselines.
Problem

Research questions and friction points this paper is trying to address.

SMILES
natural language
semantic representation
overfitting
domain adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

bi-semantic representation
SMILES
contrastive learning
domain-adaptive pre-training
chemical language modeling
🔎 Similar Papers
No similar papers found.
D
David Ming Segura
1Laboratory of Artificial Chemical Intelligence (LIAC), EPFL, Lausanne, Switzerland; 2National Centre of Competence in Research (NCCR) Catalysis, EPFL, Lausanne, Switzerland
J
Jeremy Goumaz
1Laboratory of Artificial Chemical Intelligence (LIAC), EPFL, Lausanne, Switzerland
J
Joshua W. Sin
1Laboratory of Artificial Chemical Intelligence (LIAC), EPFL, Lausanne, Switzerland; 3Process Chemistry & Catalysis, Synthetic Molecules Technical Development, F. Hoffmann-La Roche AG, Basel, Switzerland
B
Bojana Ranković
1Laboratory of Artificial Chemical Intelligence (LIAC), EPFL, Lausanne, Switzerland; 2National Centre of Competence in Research (NCCR) Catalysis, EPFL, Lausanne, Switzerland
Philippe Schwaller
Philippe Schwaller
Assistant Professor, Laboratory of Artificial Chemical Intelligence - EPFL
Deep LearningML for ChemistryReaction PredictionSynthesis PlanningAccelerated Discovery