Score
Designs and builds vector representations of monomers — including HELM monomers and non‑natural monomers — that explicitly encode chemical structure, substituent (R‑group) detail, and related molecular features to produce structured chemical embeddings. Analyzes and validates these embeddings for chemical fidelity and suitability as inputs to machine‑learning and generative models.
Final-layer embeddings from molecular pre-trained encoders are not necessarily optimal for ADMET property prediction; intermediate-layer representations often yield superior performance. Method: We propose an “evaluate-then-fine-tune” strategy: first, freeze all encoder layers and perform zero-shot evaluation of each layer’s embeddings to identify the optimal representation layer; then, fine-tune only that layer (or its corresponding subnetwork) for downstream tasks. Contribution/Results: This work is the first to systematically demonstrate the superiority of intermediate-layer molecular representations. We establish a strong correlation between frozen-layer zero-shot performance and fine-tuned accuracy, enabling low-cost, reliable layer selection. On 22 ADMET benchmarks, frozen intermediate-layer embeddings achieve an average improvement of 5.4% (up to 28.6%) over standard final-layer baselines; layer-specific fine-tuning yields an average gain of 8.5% (up to 40.8%), achieving state-of-the-art performance on multiple tasks.
Polymer informatics faces dual challenges of data scarcity and inadequate molecular representation, limiting machine learning’s efficacy in property prediction and inverse design. To address these, we propose CI-LLM: a framework leveraging the HAPPY hierarchical molecular encoder to map chemical substructures into interpretable, hierarchical tokens, augmented with numerical descriptors in a De³BERTa-enhanced Transformer encoder; coupled with a GPT-based generative model for end-to-end forward prediction and inverse design. CI-LLM delivers substructure-level interpretability in forward tasks and achieves 100% backbone retention alongside multi-objective optimization for negatively correlated properties in inverse design. Experiments demonstrate a 3.5× speedup in prediction inference, R² improvements of 0.9–4.1 percentage points, and substantial advancement in few-shot polymer intelligent design.
研究通过显式编码分子的Bemis-Murcko骨架并使用不同几何对比目标来指导分子嵌入学习,以改善分子性质预测。
To address the low information density of 2D graph representations and their inability to model stereoelectronic effects in molecular machine learning, this work introduces a novel molecular graph representation that explicitly incorporates quantum-chemical stereoelectronic features—such as orbital interactions. Leveraging a dual-graph neural network architecture coupled with physics-informed representation learning, our method is the first to encode stereoelectronic information directly into the message-passing process without relying on costly quantum mechanical calculations. The approach significantly improves prediction accuracy across diverse molecular properties and demonstrates cross-scale generalizability—successfully transferring from small-molecule training sets to ultra-large systems such as proteins. We release an open-source web platform (simg.cheme.cmu.edu) enabling real-time stereoelectronic analysis and interactive visualization. The framework achieves state-of-the-art predictive performance while maintaining strong interpretability grounded in physical principles.
Traditional molecular representations—such as graphs or point clouds—struggle to jointly and continuously encode surface geometry, hydrophobic cores, and dynamic conformational ensembles. To address this, we propose Molecular Neural Fields (MNF): an implicit representation of molecules as vector-valued functions parameterized by neural networks, enabling unified, continuous modeling of shape, hydrophobicity, and atomic composition. This work introduces implicit neural representations to molecular modeling for the first time, overcoming resolution and topological limitations inherent in discrete representations, and yielding compact, differentiable, resolution-agnostic 3D field representations. We employ an auto-decoder for parameterizing protein–ligand complexes and performing super-resolution reconstruction, and an auto-encoder for learning latent volumetric embeddings. Experiments demonstrate MNF’s effectiveness in molecular structure reconstruction, super-resolution recovery, and unbiased spatiotemporal conformational interpolation. MNF establishes a new paradigm for AI-driven molecular design and dynamic simulation.
This work addresses the challenges of heterogeneous modality entanglement and geometric-chemical inconsistency in 3D molecular generation, which arise from jointly modeling discrete atom types and continuous coordinates. To overcome these issues, the authors propose a vector field–based continuous representation paradigm, where a molecule is modeled as a continuous vector field in Euclidean space, with its direction implicitly encoding local structural information—eliminating the need for explicit graph construction. The vector field is parameterized by a neural field and integrated into a latent diffusion model to enable end-to-end generation, naturally decoupling structure learning from atom instantiation. Experiments on the QM9 and GEOM-Drugs benchmarks demonstrate the effectiveness of the approach, marking the first successful demonstration of vector field representations for 3D molecular generation and highlighting their feasibility and potential.
This study addresses the computationally driven inverse design of copolymers targeting specific blend ratios and desired performance properties, without requiring sequence information. The authors propose a Mixture Vector (MV) model that represents copolymer features as convex combinations of monomer descriptors weighted by their blending ratios. For the first time, this representation is embedded within a mixed-integer linear programming (MILP) framework, enabling precise and scalable inverse generation of multi-monomer copolymers. Coupled with machine learning–based property predictors trained across ten physicochemical datasets, the approach achieves test R² values exceeding 0.7 on nine datasets (with six surpassing 0.9). The method successfully accomplishes tractable inverse design for ternary copolymer systems and demonstrates robustness through external validation.
This work addresses the limitations of existing molecular generation models, where geometric representation spaces derived from pretrained encoders are non-smooth and underutilized, hindering both efficiency and quality. To overcome this, the authors propose the LENSes framework, which introduces, for the first time, a node-level representation alignment (REPA) objective during generative training. By integrating multi-level representation heads with a molecule-aware loss, LENSes optimizes 3D molecular generation based on the UniMol encoder. The approach substantially smooths the representation space and enhances semantic consistency, establishing a novel pretraining paradigm for molecular encoders. Evaluated on GEOM-DRUG, the method achieves 97.28% validity and 98.51% stability, reduces the Lipschitz constant by 4.6×, and demonstrates superior representation quality on QM9 probe tasks.
This study addresses the critical scarcity of large-scale, synthetically accessible, and chemically recyclable ring-opening polymer datasets, which has significantly hindered the development of sustainable polymeric materials. To overcome this limitation, the work introduces a chemist-inspired heuristic rule framework integrated with virtual forward synthesis (VFS), polyBART latent space exploration, and the POLYT5 sequence-to-sequence generative model, establishing a high-throughput, automated pipeline for polymer generation and validation. This approach yields a high-quality dataset comprising one million structurally novel, synthetically feasible, and chemically recyclable ring-opening polymers, offering a foundational resource for the rational design and discovery of sustainable materials.
This work proposes a unified molecular machine learning framework that overcomes the limitations of existing models, which are often confined to specific codebases and struggle to generalize across the full periodic table or diverse molecular properties. The framework supports elements 1–100, encompassing organic, inorganic, coordination, and biomolecular systems, and enables predictions at atomic, bond, molecular, and functional group levels. It natively incorporates conditional modeling of charge and spin states and uniquely integrates E(3)-equivariant networks, Transformers, and 2D graph neural networks within a single architecture, while combining both aleatoric and epistemic uncertainty quantification. Evaluated across multiple chemical benchmarks, the model matches or exceeds state-of-the-art performance, scales to datasets containing millions of molecules, and significantly lowers the barrier to entry for researchers without computational expertise.