chemical-aware monomer embeddings

Designs and builds vector representations of monomers — including HELM monomers and non‑natural monomers — that explicitly encode chemical structure, substituent (R‑group) detail, and related molecular features to produce structured chemical embeddings. Analyzes and validates these embeddings for chemical fidelity and suitability as inputs to machine‑learning and generative models.

chemical-awaremonomerembeddings

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Final-layer embeddings from molecular pre-trained encoders are not necessarily optimal for ADMET property prediction; intermediate-layer representations often yield superior performance. Method: We propose an “evaluate-then-fine-tune” strategy: first, freeze all encoder layers and perform zero-shot evaluation of each layer’s embeddings to identify the optimal representation layer; then, fine-tune only that layer (or its corresponding subnetwork) for downstream tasks. Contribution/Results: This work is the first to systematically demonstrate the superiority of intermediate-layer molecular representations. We establish a strong correlation between frozen-layer zero-shot performance and fine-tuned accuracy, enabling low-cost, reliable layer selection. On 22 ADMET benchmarks, frozen intermediate-layer embeddings achieve an average improvement of 5.4% (up to 28.6%) over standard final-layer baselines; layer-specific fine-tuning yields an average gain of 8.5% (up to 40.8%), achieving state-of-the-art performance on multiple tasks.

Improving performance by using intermediate layer embeddingsOptimizing molecular encoder layers for better ADMET predictionsReducing computational costs via evaluate-then-finetune strategy

Polymer informatics faces dual challenges of data scarcity and inadequate molecular representation, limiting machine learning’s efficacy in property prediction and inverse design. To address these, we propose CI-LLM: a framework leveraging the HAPPY hierarchical molecular encoder to map chemical substructures into interpretable, hierarchical tokens, augmented with numerical descriptors in a De³BERTa-enhanced Transformer encoder; coupled with a GPT-based generative model for end-to-end forward prediction and inverse design. CI-LLM delivers substructure-level interpretability in forward tasks and achieves 100% backbone retention alongside multi-objective optimization for negatively correlated properties in inverse design. Experiments demonstrate a 3.5× speedup in prediction inference, R² improvements of 0.9–4.1 percentage points, and substantial advancement in few-shot polymer intelligent design.

Develops a language model for polymer property prediction and designEnables interpretable structure-property insights and multi-property optimizationOvercomes data scarcity with hierarchical molecular representations

Advancing molecular machine learning representations with stereoelectronics-infused molecular graphs

Aug 08, 2024
DA
Daniil A. Boiko
🏛️ Carnegie Mellon University | Federal University of Santa Maria | Google DeepMind | University of Toronto | Vector Institute for Artificial Intelligence | Lawrence Berkeley National Laboratory

To address the low information density of 2D graph representations and their inability to model stereoelectronic effects in molecular machine learning, this work introduces a novel molecular graph representation that explicitly incorporates quantum-chemical stereoelectronic features—such as orbital interactions. Leveraging a dual-graph neural network architecture coupled with physics-informed representation learning, our method is the first to encode stereoelectronic information directly into the message-passing process without relying on costly quantum mechanical calculations. The approach significantly improves prediction accuracy across diverse molecular properties and demonstrates cross-scale generalizability—successfully transferring from small-molecule training sets to ultra-large systems such as proteins. We release an open-source web platform (simg.cheme.cmu.edu) enabling real-time stereoelectronic analysis and interactive visualization. The framework achieves state-of-the-art predictive performance while maintaining strong interpretability grounded in physical principles.

Enabling accurate extrapolation of learned representations to larger molecular systemsEnhancing molecular graphs with stereoelectronic effects for better machine learningImproving molecular property prediction via quantum-chemical-rich information infusion

Implicit Neural Representations of Molecular Vector-Valued Functions

Feb 15, 2025
JL
Jirka Lhotka
🏛️ École Polytechnique Fédérale de Lausanne | Wageningen University & Research

Traditional molecular representations—such as graphs or point clouds—struggle to jointly and continuously encode surface geometry, hydrophobic cores, and dynamic conformational ensembles. To address this, we propose Molecular Neural Fields (MNF): an implicit representation of molecules as vector-valued functions parameterized by neural networks, enabling unified, continuous modeling of shape, hydrophobicity, and atomic composition. This work introduces implicit neural representations to molecular modeling for the first time, overcoming resolution and topological limitations inherent in discrete representations, and yielding compact, differentiable, resolution-agnostic 3D field representations. We employ an auto-decoder for parameterizing protein–ligand complexes and performing super-resolution reconstruction, and an auto-encoder for learning latent volumetric embeddings. Experiments demonstrate MNF’s effectiveness in molecular structure reconstruction, super-resolution recovery, and unbiased spatiotemporal conformational interpolation. MNF establishes a new paradigm for AI-driven molecular design and dynamic simulation.

Capturing molecular features and hydrophobic cores.Enabling resolution-independent molecular interpolation.Representing molecules via neural vector fields.

Latest Papers

What's happening recently
View more

This work addresses the challenges of heterogeneous modality entanglement and geometric-chemical inconsistency in 3D molecular generation, which arise from jointly modeling discrete atom types and continuous coordinates. To overcome these issues, the authors propose a vector field–based continuous representation paradigm, where a molecule is modeled as a continuous vector field in Euclidean space, with its direction implicitly encoding local structural information—eliminating the need for explicit graph construction. The vector field is parameterized by a neural field and integrated into a latent diffusion model to enable end-to-end generation, naturally decoupling structure learning from atom instantiation. Experiments on the QM9 and GEOM-Drugs benchmarks demonstrate the effectiveness of the approach, marking the first successful demonstration of vector field representations for 3D molecular generation and highlighting their feasibility and potential.

3D molecule generationdrug discoverygenerative modeling

This study addresses the computationally driven inverse design of copolymers targeting specific blend ratios and desired performance properties, without requiring sequence information. The authors propose a Mixture Vector (MV) model that represents copolymer features as convex combinations of monomer descriptors weighted by their blending ratios. For the first time, this representation is embedded within a mixed-integer linear programming (MILP) framework, enabling precise and scalable inverse generation of multi-monomer copolymers. Coupled with machine learning–based property predictors trained across ten physicochemical datasets, the approach achieves test R² values exceeding 0.7 on nine datasets (with six surpassing 0.9). The method successfully accomplishes tractable inverse design for ternary copolymer systems and demonstrates robustness through external validation.

copolymerinverse designmixing vector

This work addresses the limitations of existing molecular generation models, where geometric representation spaces derived from pretrained encoders are non-smooth and underutilized, hindering both efficiency and quality. To overcome this, the authors propose the LENSes framework, which introduces, for the first time, a node-level representation alignment (REPA) objective during generative training. By integrating multi-level representation heads with a molecule-aware loss, LENSes optimizes 3D molecular generation based on the UniMol encoder. The approach substantially smooths the representation space and enhances semantic consistency, establishing a novel pretraining paradigm for molecular encoders. Evaluated on GEOM-DRUG, the method achieves 97.28% validity and 98.51% stability, reduces the Lipschitz constant by 4.6×, and demonstrates superior representation quality on QM9 probe tasks.

geometric representationmolecule generative modelspretrained molecular encoders

This study addresses the critical scarcity of large-scale, synthetically accessible, and chemically recyclable ring-opening polymer datasets, which has significantly hindered the development of sustainable polymeric materials. To overcome this limitation, the work introduces a chemist-inspired heuristic rule framework integrated with virtual forward synthesis (VFS), polyBART latent space exploration, and the POLYT5 sequence-to-sequence generative model, establishing a high-throughput, automated pipeline for polymer generation and validation. This approach yields a high-quality dataset comprising one million structurally novel, synthetically feasible, and chemically recyclable ring-opening polymers, offering a foundational resource for the rational design and discovery of sustainable materials.

chemically recyclable polymerspolymer datasetring-opening polymerization

This work proposes a unified molecular machine learning framework that overcomes the limitations of existing models, which are often confined to specific codebases and struggle to generalize across the full periodic table or diverse molecular properties. The framework supports elements 1–100, encompassing organic, inorganic, coordination, and biomolecular systems, and enables predictions at atomic, bond, molecular, and functional group levels. It natively incorporates conditional modeling of charge and spin states and uniquely integrates E(3)-equivariant networks, Transformers, and 2D graph neural networks within a single architecture, while combining both aleatoric and epistemic uncertainty quantification. Evaluated across multiple chemical benchmarks, the model matches or exceeds state-of-the-art performance, scales to datasets containing millions of molecules, and significantly lowers the barrier to entry for researchers without computational expertise.

diverse chemical specieselemental compositionmolecular machine learning

Hot Scholars

WN

Weili Nie

NVIDIA Research
Machine LearningDeep LearningGenerative Models
HJ

Heng Ji

Professor of Computer Science, AICE Director, ASKS Director, UIUC, Amazon Scholar
Natural Language ProcessingLarge Language Models
DZ

Duo Zhang

Twitter, Inc.
Text MiningInformation RetrievalData MiningMachine Learning
JZ

Jiayu Zhou

University of Michigan
Machine LearningAI + Health Informatics
XL

Xiao Luo

University of Wisconsin–Madison
Machine LearningLLMML for ScienceStatistical Modeling