Score
Designs and implements molecular feature representations (descriptors) derived from molecular structure and quantum chemical calculations, and computes quantum-chemistry-derived quantities such as partial charges, orbital energies, and electron densities. Builds, validates, and selects descriptor sets for modeling tasks and analyzes descriptor quality, redundancy, and computational cost trade-offs.
This work addresses the long-standing absence of a unified taxonomy and reliable benchmarking framework in molecular property prediction, particularly in light of emerging challenges posed by foundation models. The study proposes a cohesive classification scheme encompassing molecular representations, model architectures, and interdisciplinary applications, and systematically evaluates four major paradigms: quantum chemistry, descriptor-based machine learning, geometric deep learning, and foundation models through multidimensional benchmarking. The analysis exposes critical shortcomings in current benchmarks regarding stereochemical consistency, experimental heterogeneity, and reproducibility. To advance the field, the paper advocates three key directions: physics-informed learning with quantum consistency, uncertainty-calibrated foundation models for trustworthy inference, and multimodal real-world benchmark ecosystems integrating computational and experimental data—collectively paving the way toward transparent, temporally aware, and scaffold-sensitive next-generation benchmarks.
This work addresses the limitations of existing molecular representation methods, which predominantly rely on atomic-level information and struggle to accurately capture true physical properties. While electronic-level descriptors offer a more fundamental characterization, their high computational cost renders them impractical for large molecules. To bridge this gap, the authors propose HEDMoL, a novel model that—through knowledge transfer—efficiently injects readily available electronic-level information from small molecules into coarse-grained representations of large molecules. This approach yields electron-aware molecular embeddings without incurring additional computational overhead. Evaluated on multiple benchmark datasets containing experimentally measured physical properties, HEDMoL achieves state-of-the-art prediction accuracy, significantly outperforming current atomic-level representation methods and demonstrating both its effectiveness and strong generalization capability.
Existing molecular representations—such as SMILES and SELFIES—exhibit fundamental limitations in capturing quantum structural features, 3D geometry, electron delocalization, and syntactic validity, thereby impeding accurate reaction modeling and Bayesian inference. To address this, we introduce the first strongly typed molecular representation framework based on algebraic data types (ADTs), natively embedding quantum-chemical constructs—including subshells, atomic orbitals, coordination geometries, and delocalized electrons—into the type system, thereby guaranteeing syntactic correctness by construction. Implemented in Haskell and deeply integrated with the LazyPPL probabilistic programming library under a data–type separation paradigm, our framework enables type-driven reaction algebra and Bayesian molecular inference for the first time. The open-source library ensures zero invalid molecule generation, markedly improving model composability and inference efficiency. This work establishes a type-safe foundation for molecular programming languages.
This work addresses the underutilization of molecular shape information in representation learning, which limits the accuracy of inhibition constant (Ki) prediction. To this end, we introduce the Euler Characteristic Transform (ECT) — a multiscale geometric-topological descriptor — into molecular representation for the first time. We propose a complementary fusion strategy integrating ECT with AVALON fingerprints and embed it within a shape-aware framework combining graph neural networks and regression models. Evaluated on nine Ki prediction benchmark datasets, the ECT+AVALON combination achieves state-of-the-art or second-best performance, significantly outperforming methods relying solely on chemical descriptors or graph-structural features. These results empirically validate the indispensable role of topological shape information in bioactivity prediction. Moreover, our analysis reveals strong complementarity between ECT and conventional molecular fingerprints, establishing a novel paradigm for geometric deep learning in drug discovery.
To address the low information density of 2D graph representations and their inability to model stereoelectronic effects in molecular machine learning, this work introduces a novel molecular graph representation that explicitly incorporates quantum-chemical stereoelectronic features—such as orbital interactions. Leveraging a dual-graph neural network architecture coupled with physics-informed representation learning, our method is the first to encode stereoelectronic information directly into the message-passing process without relying on costly quantum mechanical calculations. The approach significantly improves prediction accuracy across diverse molecular properties and demonstrates cross-scale generalizability—successfully transferring from small-molecule training sets to ultra-large systems such as proteins. We release an open-source web platform (simg.cheme.cmu.edu) enabling real-time stereoelectronic analysis and interactive visualization. The framework achieves state-of-the-art predictive performance while maintaining strong interpretability grounded in physical principles.
Existing neural network potentials typically rely solely on atomic numbers and Cartesian coordinates, limiting their generalizability to chemically diverse systems with varying charge states, spin multiplicities, and other electronic configurations. This work introduces a lightweight extension strategy that, while preserving the equivariant architecture of models like TensorNet, enables end-to-end learning of charge and spin embeddings—without incorporating physics-based energy terms or task-specific modules. The approach directly alleviates input degeneracy arising from identical nuclear configurations across distinct electronic states. Experiments on both a newly curated dataset and established benchmarks (e.g., ANI-1x, QM9 subsets) demonstrate significant improvements in energy and force prediction accuracy across charge and spin states, reducing mean absolute errors by 15–30%. Crucially, the method retains the original model’s computational efficiency and out-of-distribution generalization capability. This establishes a new paradigm for developing universal, robust molecular potential energy models grounded in learnable electronic structure representations.
This work addresses the lack of principled, atomistic out-of-distribution (OOD) evaluation protocols for machine learning models, which hinders reliable assessment of their generalization to atomic properties such as partial charges and multipole moments. To this end, the authors propose a leave-one-cluster-out evaluation scheme based on SOAP descriptor clustering of atomic environments and introduce QT-Net, a rotation-augmented, non-equivariant graph neural network that incorporates quantum topological atom (QTA) properties as inductive bias to predict electron populations and multipole moments for H, C, N, and O atoms. Experiments demonstrate that QT-Net exhibits strong OOD generalization on QM9 molecules, accurately reconstructs molecular dipole moments from predicted atomic multipoles, and significantly enhances performance in downstream molecular property prediction tasks.
This work addresses the limitations of conventional molecular fingerprints, which are predominantly based on two-dimensional structures and thus fail to distinguish stereoisomers and conformers, while existing three-dimensional approaches often suffer from high computational cost or heavy data dependence. The authors propose a physics-inspired 3D molecular fingerprint that represents a molecule as a fully connected 3D graph, where edge weights encode heuristic physical interactions. By performing eigendecomposition on the graph Laplacian matrix, the method yields a fixed-length descriptor invariant to both atomic permutations and E(3) transformations. This approach uniquely integrates spectral graph theory with physically motivated 3D graph representations, achieving state-of-the-art performance across multiple chemical datasets. It offers strong interpretability, geometric awareness, and low computational complexity, making it well-suited for efficient large-scale chemical space screening and applicability domain analysis.
Conventional machine-learned interatomic potentials (MLIPs) employ scalar regression to predict energies and forces, achieving high efficiency but lacking principled quantification of epistemic uncertainty. Method: This work introduces the first classification-based MLIP framework, mapping continuous energy and force targets onto histogram-binned distributional labels and jointly optimizing probabilistic outputs via cross-entropy loss. Uncertainty is intrinsically quantified through the entropy of predicted distributions, enabling direct, interpretable confidence estimation. Contribution/Results: On DFT-derived benchmark datasets, the method achieves absolute prediction errors comparable to state-of-the-art regression-based MLIPs while providing calibrated, distribution-level uncertainty estimates. It overcomes the long-standing limitation of uncertainty-unaware modeling in MLIPs and establishes a risk-aware paradigm for molecular simulation—enabling reliability assessment, active learning, and safe decision-making in computational chemistry and materials science.
This work addresses the high computational cost and complex data engineering requirements of conventional molecular property prediction methods, which typically rely on molecular graphs, 3D conformations, or large language models. For the first time, it systematically investigates a purely vision-based paradigm by evaluating ten visual architectures and seven pretraining strategies across ten downstream tasks using a dataset of two million molecular scaffold images. The study introduces a chemistry-informed curriculum learning strategy that dynamically orders training samples according to molecular structural complexity. Experimental results demonstrate that accurate predictions can be achieved using only a single molecular image, with the proposed approach ranking first on five out of ten benchmarks and placing within the top two on all tasks, while reducing computational costs by up to 80× compared to state-of-the-art multimodal methods.
This work addresses the semantic gap between one-dimensional molecular representations (e.g., SMILES/SELFIES) and three-dimensional geometric generation. To bridge this gap, we propose a novel “grammar-driven sculpting” paradigm. Methodologically, we introduce the first end-to-end injection of frozen 1D chemical foundation model knowledge—such as a SELFIES encoder—into the conditional space of a 3D diffusion model, without fine-tuning the 1D model. This is achieved via a learnable chemical knowledge query module and a cross-modal projector that precisely maps chemical priors encoded in the 1D representation onto the 3D conformational generation process. Evaluated on GEOM-DRUGS and QM9, our approach achieves state-of-the-art performance in 3D molecular generation, significantly improving geometric accuracy (reduced bond, angle, and dihedral errors), conformational diversity (↑ Coverage), and sampling stability (↑ Validity), thereby effectively closing the semantic gap between 1D syntactic representations and 3D spatial configurations.