Score
Designs, builds, and evaluates computational pipelines that convert molecular structures into engineered descriptors or learned representations and train predictive models to estimate molecular properties such as physicochemical parameters, biological activity, or toxicity. This work includes developing and selecting molecular featurizations and representation‑learning methods (fingerprints, descriptors, graph- or sequence‑based embeddings), training and calibrating supervised models, producing ranked predictions with uncertainty estimates, and validating model outputs against experimental measurements.
This work addresses the long-standing absence of a unified taxonomy and reliable benchmarking framework in molecular property prediction, particularly in light of emerging challenges posed by foundation models. The study proposes a cohesive classification scheme encompassing molecular representations, model architectures, and interdisciplinary applications, and systematically evaluates four major paradigms: quantum chemistry, descriptor-based machine learning, geometric deep learning, and foundation models through multidimensional benchmarking. The analysis exposes critical shortcomings in current benchmarks regarding stereochemical consistency, experimental heterogeneity, and reproducibility. To advance the field, the paper advocates three key directions: physics-informed learning with quantum consistency, uncertainty-calibrated foundation models for trustworthy inference, and multimodal real-world benchmark ecosystems integrating computational and experimental data—collectively paving the way toward transparent, temporally aware, and scaffold-sensitive next-generation benchmarks.
This study addresses three critical challenges in chemical process design: low accuracy in molecular property prediction, inefficient discovery of novel molecules, and the absence of molecular–process co-design. To tackle these, we propose a hybrid architecture integrating physics- and chemistry-informed graph neural networks (GNNs) with Transformer modules, enabling high-accuracy, interpretable prediction of thermodynamic and transport properties for both pure components and mixtures. Building upon this, we establish a closed-loop framework unifying molecular generation, property prediction, and process optimization—facilitating targeted exploration of chemical space and cross-scale co-design. Furthermore, we introduce the first unified benchmark spanning molecular modeling, process simulation, and industrial validation, substantially enhancing model generalizability and engineering applicability. The resulting paradigm provides a scalable, experimentally verifiable methodology for the co-development of novel functional molecules and low-carbon chemical processes.
Current evaluations of pretrained molecular embedding models lack statistical rigor and fair benchmarking, often assuming neural models inherently outperform traditional fingerprints like ECFP without empirical validation. Method: We systematically evaluate 25 pretrained molecular embedding models across 25 benchmark datasets and introduce a hierarchical Bayesian statistical testing framework to control for multiple hypothesis testing bias in a unified manner. Contribution/Results: Contrary to prevailing assumptions, most pretrained models fail to achieve statistically significant improvements over the ECFP baseline; only CLAMP demonstrates robust superiority. This challenges the implicit “neural models are necessarily better” assumption pervasive in molecular representation learning and exposes systemic issues—including methodologically unsound evaluation protocols, weak baselines, and reporting bias. Based on these findings, we propose a standardized evaluation protocol, recommend strong baselines (e.g., ECFP with optimized hyperparameters), and advocate for small-sample validation to ensure reliability. Our work establishes a methodological foundation and practical guidelines for trustworthy molecular AI research.
To address the low information density of 2D graph representations and their inability to model stereoelectronic effects in molecular machine learning, this work introduces a novel molecular graph representation that explicitly incorporates quantum-chemical stereoelectronic features—such as orbital interactions. Leveraging a dual-graph neural network architecture coupled with physics-informed representation learning, our method is the first to encode stereoelectronic information directly into the message-passing process without relying on costly quantum mechanical calculations. The approach significantly improves prediction accuracy across diverse molecular properties and demonstrates cross-scale generalizability—successfully transferring from small-molecule training sets to ultra-large systems such as proteins. We release an open-source web platform (simg.cheme.cmu.edu) enabling real-time stereoelectronic analysis and interactive visualization. The framework achieves state-of-the-art predictive performance while maintaining strong interpretability grounded in physical principles.
This work addresses the limitations of existing molecular representation methods, which predominantly rely on atomic-level information and struggle to accurately capture true physical properties. While electronic-level descriptors offer a more fundamental characterization, their high computational cost renders them impractical for large molecules. To bridge this gap, the authors propose HEDMoL, a novel model that—through knowledge transfer—efficiently injects readily available electronic-level information from small molecules into coarse-grained representations of large molecules. This approach yields electron-aware molecular embeddings without incurring additional computational overhead. Evaluated on multiple benchmark datasets containing experimentally measured physical properties, HEDMoL achieves state-of-the-art prediction accuracy, significantly outperforming current atomic-level representation methods and demonstrating both its effectiveness and strong generalization capability.
This study investigates the scaling relationship between dataset size and prediction error for machine learning models operating in highly mutable discrete combinatorial spaces—such as proteins and small molecules—where conventional continuity assumptions break down. Method: We introduce a mutation-oriented data reordering strategy and a normalized learning curve analysis framework, integrating kernel ridge regression, synthetic multi-body-theoretic data, calibration plot clustering, and resampling techniques. Contribution/Results: We discover, for the first time, a “saturation–asymptotic” two-stage learning paradigm driven by mutational complexity, accompanied by a discontinuous drop in test error at a critical dataset size—a phase-transition-like phenomenon. Systematic validation on peptide–protein binding affinity and small-molecule solvation energy prediction tasks demonstrates that mutational complexity is the dominant factor governing learning efficiency and generalization performance, substantially outperforming conventional metrics such as sequence length or chemical diversity.
This study addresses the lack of systematic comparison among molecular encoding methods in terms of both predictive performance and interpretability for drug property prediction. The authors propose a hybrid model combining multilayer perceptrons and Transformer encoders (MLP+TL) to comprehensively evaluate topological fingerprints, substructure-based fingerprints (e.g., MACCS, PubChem), and string representations across seven molecular datasets and multiple biologically relevant classification tasks. Innovatively leveraging the model’s intrinsic attention weights—without relying on external interpretation tools—the work identifies critical chemical moieties, revealing mechanistic insights such as the influence of hydroxyl groups on blood–brain barrier permeability and Salmonella mutagenicity. The model achieves average AUC scores exceeding 0.9 in toxicity, mutagenicity, and side-effect prediction tasks, demonstrating both high predictive accuracy and inherent chemical interpretability.
This study addresses the limited generalization of traditional graph-theoretic models on large-scale, chemically diverse molecular datasets. To overcome this challenge, the authors propose a lightweight, GPU-free enhancement framework that integrates topological indices with Morgan fingerprints and incorporates physicochemical properties, regularization, feature selection, and ensemble learning strategies. The approach maintains model interpretability while substantially improving predictive performance. Evaluated using Ridge/Lasso regression and gradient boosting methods across five MoleculeNet benchmark datasets, the framework increases the average coefficient of determination (R²) from 0.24 to 0.79 (p < 0.001), with training times under five minutes. Remarkably, it achieves performance comparable to or exceeding that of deep learning models, making it well-suited for resource-constrained environments.
This work proposes a unified molecular machine learning framework that overcomes the limitations of existing models, which are often confined to specific codebases and struggle to generalize across the full periodic table or diverse molecular properties. The framework supports elements 1–100, encompassing organic, inorganic, coordination, and biomolecular systems, and enables predictions at atomic, bond, molecular, and functional group levels. It natively incorporates conditional modeling of charge and spin states and uniquely integrates E(3)-equivariant networks, Transformers, and 2D graph neural networks within a single architecture, while combining both aleatoric and epistemic uncertainty quantification. Evaluated across multiple chemical benchmarks, the model matches or exceeds state-of-the-art performance, scales to datasets containing millions of molecules, and significantly lowers the barrier to entry for researchers without computational expertise.
This work addresses the high computational cost and complex data engineering requirements of conventional molecular property prediction methods, which typically rely on molecular graphs, 3D conformations, or large language models. For the first time, it systematically investigates a purely vision-based paradigm by evaluating ten visual architectures and seven pretraining strategies across ten downstream tasks using a dataset of two million molecular scaffold images. The study introduces a chemistry-informed curriculum learning strategy that dynamically orders training samples according to molecular structural complexity. Experimental results demonstrate that accurate predictions can be achieved using only a single molecular image, with the proposed approach ranking first on five out of ten benchmarks and placing within the top two on all tasks, while reducing computational costs by up to 80× compared to state-of-the-art multimodal methods.
为解决分子性质预测中标签数据不足的问题,研究提出了一种包含程序化预训练、分子预训练和下游微调的三阶段训练流程。