MolSC: Leveraging Substituent Contributions to Enhance Fine-grained Molecular Understanding in LLMs

📅 2026-09-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文通过构建MolSC数据集和MolSC-Bench基准,改进了分子大语言模型对细微结构-性质关系的理解,特别是局部修改如何影响分子行为。
📝 Abstract
Recent advances in natural language processing have led to molecular Large Language Models (LLMs) with strong performance across diverse chemistry tasks. However, they still struggle to capture fine-grained structure-property relationships, particularly how small, localized modifications alter a molecule's behavior. To address this limitation, we introduce MolSC, a dataset of substituent contributions, defined as property changes induced by attaching specific substituents to molecular scaffolds. Curated from manually annotated bioactivity records, MolSC spans structural-alert liability, target-specific bioactivity, and physicochemical descriptors, and contains 181K substituent-level examples for training. We further propose MolSC-Bench, a held-out evaluation benchmark of 1,541 examples disjoint from MolSC at the scaffold, substituent, and molecule levels. Our experiments show that existing molecular LLMs and strong proprietary models such as GPT-5.2 and Gemini-3-Flash show limited reliability in substituent contribution prediction. In contrast, training on MolSC substantially improves this ability and achieves strong performance across diverse downstream molecular tasks. These results highlight substituent contribution learning as a key component of fine-grained molecular understanding.
Problem

Research questions and friction points this paper is trying to address.

fine-grained structure-property relationships
molecular Large Language Models
substituent contributions
Innovation

Methods, ideas, or system contributions that make the work stand out.

MolSC
substituent contributions
fine-grained molecular understanding
MolSC-Bench