Chemistry Integrated Language Model using Hierarchical Molecular Representation for Polymer Informatics

📅 2025-12-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Polymer informatics faces dual challenges of data scarcity and inadequate molecular representation, limiting machine learning’s efficacy in property prediction and inverse design. To address these, we propose CI-LLM: a framework leveraging the HAPPY hierarchical molecular encoder to map chemical substructures into interpretable, hierarchical tokens, augmented with numerical descriptors in a De³BERTa-enhanced Transformer encoder; coupled with a GPT-based generative model for end-to-end forward prediction and inverse design. CI-LLM delivers substructure-level interpretability in forward tasks and achieves 100% backbone retention alongside multi-objective optimization for negatively correlated properties in inverse design. Experiments demonstrate a 3.5× speedup in prediction inference, R² improvements of 0.9–4.1 percentage points, and substantial advancement in few-shot polymer intelligent design.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Conceptual Inference and ReasoningSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Machine learning has transformed material discovery for inorganic compounds and small molecules, yet polymers remain largely inaccessible to these methods. While data scarcity is often cited as the primary bottleneck, we demonstrate that strategic molecular representations can overcome this limitation. We introduce CI-LLM (Chemically Informed Language Model), a framework combining HAPPY (Hierarchically Abstracted rePeat unit of PolYmer), which encodes chemical substructures as tokens, with numerical descriptors within transformer architectures. For property prediction, De$^3$BERTa, our descriptor-enriched encoder, achieves 3.5x faster inference than SMILES-based models with improved accuracy ($R^2$ score gains of 0.9-4.1 percent across four properties), while providing interpretable structure-property insights at the subgroup level. For inverse design, our GPT-based generator produces polymers with targeted properties, achieving 100 percent scaffold retention and successful multi-property optimization for negatively correlated objectives. This comprehensive framework demonstrates both forward prediction and inverse design capabilities, showcasing how strategic molecular representation advances machine learning applications in polymer science.
Problem

Research questions and friction points this paper is trying to address.

Develops a language model for polymer property prediction and design
Overcomes data scarcity with hierarchical molecular representations
Enables interpretable structure-property insights and multi-property optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical molecular representation encoding chemical substructures as tokens
Descriptor-enriched transformer encoder for faster inference and improved accuracy
GPT-based generator for inverse design with scaffold retention and multi-property optimization
💼 Related Jobs
No related jobs found.
J
Jihun Ahn
Department of Polymer Engineering, Graduate School, Chonnam National University, Gwangju, 61186, Republic of Korea
G
Gabriella Pasya Irianti
Department of Polymer Engineering, Graduate School, Chonnam National University, Gwangju, 61186, Republic of Korea
V
Vikram Thapar
Department of Polymer Engineering, Graduate School, Chonnam National University, Gwangju, 61186, Republic of Korea
S
Su-Mi Hur
School of Polymer Science and Engineering, Chonnam National University, Gwangju, 61186, Republic of Korea