Score
Designs and implements algorithms, models, and tools that construct molecular graph representations (atoms as nodes, bonds as edges), including generative systems that produce new molecular graphs and components that visualize those graphs. Works to ensure generated structures meet chemical validity and stereochemical constraints, provide interpretable renderings, and support downstream analysis or property-driven filtering.
This work addresses the modeling needs for molecules, proteins, reaction pathways, and industrial processes in chemical science. Method: We propose a unified graph-structured representation paradigm that abstracts multiscale chemical entities as learnable heterogeneous graphs. Integrating graph neural networks (GNNs) with domain-specific chemical priors, we develop an end-to-end graph representation learning framework supporting structure-aware embedding generation and cross-scale prediction—including molecular property estimation, reactivity assessment, and target binding affinity prediction. Contribution/Results: First, we introduce the first standardized chemical graph modeling protocol spanning atomic, molecular, protein, reaction, and process scales. Second, we design chemistry-aware edge-type encoding and subgraph-level attention mechanisms, substantially enhancing physical interpretability and generalization. Experiments demonstrate an average 9.3% improvement in prediction accuracy across 12 benchmark tasks. The framework has been deployed in real-world applications, including novel material discovery and drug candidate optimization.
To address the critical bottleneck in drug discovery—where molecular generation models neglect synthetic feasibility, hindering experimental validation—this work proposes a novel molecular generation framework projectable onto synthetically accessible chemical space. Methodologically, it introduces synthesis path expressions (SPEs) as a novel molecular representation that intrinsically encodes retrosynthetic logic, and designs a graph-based Transformer architecture for end-to-end translation from molecular graphs to SPEs. This formulation inherently guarantees synthetic feasibility of generated molecules and enables structure-preserving, synthetically constrained analog generation for initially infeasible candidates. Experiments demonstrate substantial improvements in retrosynthetic planning accuracy and successful re-mapping of multiple state-of-the-art generative model outputs—previously deemed synthetically intractable—into property-preserved, experimentally viable analogs. The approach effectively bridges the gap between de novo molecular generation and practical synthesis.
To address low chemical validity and difficulty in satisfying substructure constraints in molecular generation, this work introduces the first pure Transformer-based generative model explicitly designed for molecular *graph* structures—bypassing sequential representations (e.g., SMILES) to directly model atomic-bond topology and 3D geometry. The method integrates graph neural networks, learnable graph positional encodings, multi-task property prediction heads, and a reinforcement learning–driven graph editing mechanism, enabling geometry-aware and property-guided controllable generation. Evaluated on chemical engineering tasks such as solvent extraction, the model achieves 92.7% molecular validity across multiple benchmarks and reduces prediction error of target distribution coefficients by 31% versus VAE and GFlowNet baselines. It establishes a novel end-to-end paradigm for controllable molecular graph generation.
To address the low information density of 2D graph representations and their inability to model stereoelectronic effects in molecular machine learning, this work introduces a novel molecular graph representation that explicitly incorporates quantum-chemical stereoelectronic features—such as orbital interactions. Leveraging a dual-graph neural network architecture coupled with physics-informed representation learning, our method is the first to encode stereoelectronic information directly into the message-passing process without relying on costly quantum mechanical calculations. The approach significantly improves prediction accuracy across diverse molecular properties and demonstrates cross-scale generalizability—successfully transferring from small-molecule training sets to ultra-large systems such as proteins. We release an open-source web platform (simg.cheme.cmu.edu) enabling real-time stereoelectronic analysis and interactive visualization. The framework achieves state-of-the-art predictive performance while maintaining strong interpretability grounded in physical principles.
This work addresses the synthetic accessibility bottleneck in molecular discovery by proposing a syntax–semantics decoupled two-level program synthesis framework. At the syntax level, Markov Chain Monte Carlo (MCMC) searches over molecular skeleton grammars; at the semantics level, a policy network—trained on fixed skeletons—generates executable retrosynthetic reaction pathways. For the first time, molecular synthesis is formulated as a structured program synthesis problem, enabling user-specified resource constraints (e.g., step count, available reagents) and inherently favoring concise, high-feasibility routes. The method achieves state-of-the-art performance on synthesizable drug-like molecule generation and analogy-based optimization of non-synthesizable molecules. It provides explicit, interpretable synthesis pathways, supports automatic pathway simplification, and integrates seamlessly with autonomous synthesis platforms. This framework establishes a novel paradigm for AI-driven retrosynthetic planning, bridging symbolic reasoning with deep learning while ensuring chemical validity and practical deployability.
This study addresses the high computational cost and low efficiency of traditional discrete graph generation methods for large-scale molecular graphs by proposing a latent space flow matching framework for molecular graph generation. The method leverages a pretrained variational autoencoder (VAE) to obtain high-fidelity latent representations and, for the first time, introduces flow matching into the whole-graph latent space for generation. Furthermore, it incorporates a latent classifier to enable validity-aware guidance, ultimately decoding the representations into high-quality molecular graphs. Experimental results demonstrate that the proposed framework achieves superior Fréchet ChemNet Distance (FCD) scores and validity across multiple molecular benchmarks, significantly improving the trade-off between generation quality and computational efficiency.
How molecular generative models internally organize discrete molecular identities remains unclear, limiting the reliability of chemical space navigation. This work proposes a molecular identity pullback method that integrates multiple molecular representations, identity conventions, decoder stochasticity, and coordinate metrics to systematically reveal— for the first time—piecewise-constant regions and hierarchically refined boundary structures within mainstream generative architectures. The study demonstrates that the internal partitioning of molecular identities exhibits stable yet dynamically evolving organizational properties, underscoring the necessity of empirical characterization rather than default assumptions. These findings lay a foundational basis for developing trustworthy mechanisms for navigating chemical space.
Existing template-free, single-step retrosynthesis models suffer from slow convergence and limited generation quality and diversity due to the difficulty of explicitly modeling chemical semantics. To address this, this work proposes a Graph Representation Guidance (GRG) framework that integrates molecular representations from a pretrained encoder into a denoising diffusion Transformer. During generation, multi-granularity alignment strategies provide deep guidance, while a representation similarity–based reranking mechanism enhances both diversity and accuracy without requiring an additional verifier. Evaluated on USPTO-50k, the model achieves top-1/3/5/10 accuracies of 58.6/77.2/83.4/87.1, respectively, with diversity improved to 15.5, training epochs reduced by 35%, and inference time shortened by 30%.
This work addresses the inconsistency in predictions and explanations arising from multiple graph representations of the same molecule in molecular graph machine learning, which violates chemical identity. To resolve this, the authors propose InChIfied Invariants—strictly invariant features constructed at node, edge, and graph levels based on the International Chemical Identifier (InChI)—that inherently guarantee identical representations, predictions, and attributions for chemically equivalent graphs. Evaluation on the large-scale PubChem Substances dataset demonstrates that the method achieves consistent representations for 99.62% of chemically equivalent graph pairs, a dramatic improvement over the 0.35% consistency achieved by conventional Daylight invariants. Furthermore, it maintains competitive predictive performance on MoleculeNet benchmark tasks while significantly enhancing model interpretability and chemical plausibility.
Existing methods struggle to reveal the dynamic evolution of latent spaces in molecular graph neural networks across different layers and training stages, as well as their relationship to chemical concepts. To address this gap, this work proposes a visual analytics system that enables, for the first time, cross-layer and cross-training-state tracking of latent representations. By clustering molecular embeddings and visualizing their evolutionary trajectories through an enhanced Sankey diagram—linked interactively with representative molecules, key substructures, and domain knowledge—the system significantly enhances model interpretability. Integrating graph neural network embeddings, clustering analysis, substructure extraction, and interactive visualization techniques, the framework is validated through two case studies, demonstrating its effectiveness in helping domain scientists understand latent space dynamics, identify meaningful molecular patterns, and gain deeper insights into model decisions.