use cheminformatics tools

Designs and implements software pipelines and tooling for representing, processing, and analyzing chemical structures and reactions, including sanitizing molecular representations and computing chemical descriptors and filters. Builds and applies validation procedures and metrics to assess synthesis-route feasibility and other formal success criteria for computational chemical data and models.

usecheminformaticstools

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.18
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$196K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the fragmented and labor-intensive pipeline in scientific machine learning—from data acquisition to model deployment—by proposing an end-to-end, high-throughput, and agent-collaborative reproducible workflow platform. The platform enables full automation of data collection, processing, modeling, validation, selection, and reporting through modular components, configuration-driven mechanisms, and standardized artifacts. Key technical features include plug-and-play model architectures, deterministic data splitting, batch execution, and structured outputs. Evaluated on quantum mechanics, physicochemical property, and bioactivity prediction tasks, the system achieves state-of-the-art performance and successfully generalizes to non-molecular domains such as time series, significantly enhancing research efficiency and reproducibility.

end-to-end workflowhigh-throughput experimentationreproducible pipeline

Chemputer and Chemputation - A Universal Chemical Compound Synthesis Machine

Aug 17, 2024
LC
Leroy Cronin
🏛️ University of Glasgow

This work aims to achieve computable chemical synthesis—i.e., precise, code-driven control of reaction pathways on general-purpose reconfigurable hardware to automate the synthesis of any stable, isolable molecule while satisfying mass conservation, finite reaction time, and analytical detectability constraints. Method: We introduce the “chemputation” paradigm, modeling synthesis as graph transformations over the space Reagents × Process × Catalyst. We formally define and prove the Universal Chemical Synthesis Theorem, introduce the notion of “analytically reachable quantity” for molecules, and establish dynamic error correction as essential. Our end-to-end implementation integrates the Chemputer hardware platform, the chempiler compiler, assembly-theory–driven reachability analysis, and a real-time sensing feedback framework. Results: Experimental validation demonstrates that chemical reactions are intrinsically programmable, observable, and correctable graph operations. We identify reactor count and sensor bandwidth as critical scalability bottlenecks for chemputation.

Developing a universal machine for synthesizing any stable moleculeExpanding the definition of molecules using assembly theoryFormalizing error correction in chemical synthesis processes

Procedural Synthesis of Synthesizable Molecules

Aug 24, 2024
MS
Michael Sun
🏛️ MIT | IBM

This work addresses the synthetic accessibility bottleneck in molecular discovery by proposing a syntax–semantics decoupled two-level program synthesis framework. At the syntax level, Markov Chain Monte Carlo (MCMC) searches over molecular skeleton grammars; at the semantics level, a policy network—trained on fixed skeletons—generates executable retrosynthetic reaction pathways. For the first time, molecular synthesis is formulated as a structured program synthesis problem, enabling user-specified resource constraints (e.g., step count, available reagents) and inherently favoring concise, high-feasibility routes. The method achieves state-of-the-art performance on synthesizable drug-like molecule generation and analogy-based optimization of non-synthesizable molecules. It provides explicit, interpretable synthesis pathways, supports automatic pathway simplification, and integrates seamlessly with autonomous synthesis platforms. This framework establishes a novel paradigm for AI-driven retrosynthetic planning, bridging symbolic reasoning with deep learning while ensuring chemical validity and practical deployability.

Decoupling syntactic skeleton from synthetic tree semantics.Designing synthetically accessible molecules and analogs.Optimizing molecular descriptors and syntactic templates jointly.

To address the critical bottleneck in drug discovery—where molecular generation models neglect synthetic feasibility, hindering experimental validation—this work proposes a novel molecular generation framework projectable onto synthetically accessible chemical space. Methodologically, it introduces synthesis path expressions (SPEs) as a novel molecular representation that intrinsically encodes retrosynthetic logic, and designs a graph-based Transformer architecture for end-to-end translation from molecular graphs to SPEs. This formulation inherently guarantees synthetic feasibility of generated molecules and enables structure-preserving, synthetically constrained analog generation for initially infeasible candidates. Experiments demonstrate substantial improvements in retrosynthetic planning accuracy and successful re-mapping of multiple state-of-the-art generative model outputs—previously deemed synthetically intractable—into property-preserved, experimentally viable analogs. The approach effectively bridges the gap between de novo molecular generation and practical synthesis.

Explores synthesizable analogs for unsynthesizable drug molecules.Generates new chemical structures ensuring synthetic accessibility.Translates molecular graphs into synthetic pathway notations.

Latest Papers

What's happening recently
View more

Existing benchmarks for chemical reasoning evaluation focus solely on final answers, making it difficult to detect logical errors in intermediate reasoning steps. To address this limitation, this work proposes ChemCoTBench-V2, a novel benchmark that employs expert-designed structured templates to guide models in generating verifiable intermediate reasoning states. By integrating deterministic chemical rules, reference trajectory alignment, and oracle-verifiable state constraints, the framework enables low-cost, auditable process-level evaluation without requiring human or LLM-based adjudication. This approach is the first to support state-constraint verification and precise error localization in open-ended tasks, revealing a significant discrepancy between answer correctness and reasoning consistency across mainstream large language models. It further facilitates fine-grained model comparison and identification of the first erroneous step in reasoning trajectories.

chemical reasoningchemistry benchmarkslarge language models

This work addresses the lack of systematic evaluation for symbolic, verifiable reasoning over molecular graph structures in current chemical large language models. Existing benchmarks often suffer from label bias or information leakage, hindering precise diagnosis of model shortcomings. To bridge this gap, we propose MolecularIQ—the first evaluation framework specifically designed for symbolic reasoning on molecular graphs. By integrating molecular graph representations, symbolic logic verification, and carefully structured reasoning tasks, MolecularIQ establishes a fine-grained benchmark that effectively uncovers systematic failure modes of contemporary models across specific molecular structures and reasoning challenges. This framework provides interpretable diagnostic insights and actionable directions for developing chemical large language models with faithful structural understanding capabilities.

chemical reasoningLLM evaluationmolecular graph

Traditional retrosynthetic tools are constrained by reaction databases and struggle to devise creative synthetic routes for highly functionalized, polycyclic natural products. This work proposes SynthEx, a framework that leverages large language models to construct an intelligent agent system employing a strategy-first planning mechanism. By generating competitive synthetic strategies, integrating critical and routine steps, and incorporating self-reflection for iterative refinement, SynthEx achieves high-quality retrosynthetic planning. Notably, it produces key disconnections comparable to those devised by human experts—validated as authentic and feasible by chemists in blind evaluations. The method successfully designs highly convergent routes for over a thousand natural products and introduces SynthAtlas, an open-access database of these pathways, which has garnered recognition from domain experts.

automated synthesis designcomplex natural productsinventive chemistry

This study addresses the unclear performance bottlenecks of AI agents in drug discovery by proposing MAGI, a modular agent designed to investigate whether tool orchestration or predictive model accuracy constrains practical outcomes. Methodologically, MAGI employs an open modular architecture coordinating molecular design, optimization monitoring, and SAR analysis with self-revising strategies. It integrates a dual-pathway mechanism combining direct large language model generation with REINFORT-delegated generation, alongside a pluggable scoring service contract. Validation on retrospective pharmaceutical projects reveals that the primary bottleneck lies in the applicability domain of scoring models rather than agent orchestration capabilities. Furthermore, MAGI generates molecules approaching expert-level quality and integrates effectively into existing computational chemistry workflows, thereby clarifying the practical role of AI agents in real-world drug discovery scenarios.

AI agentsdrug discoverylead optimization

This work addresses the heavy reliance on manual effort in chemical process modeling, which is prone to catastrophic failure due to single-point errors. To overcome this limitation, the authors propose a role-adaptive collaborative framework that decomposes the modeling task into seven specialized sub-roles. By integrating natural language, process flow diagrams, and domain knowledge, the framework generates structured models through typed intermediate representations and a deterministic engineering gating mechanism, enabling automated optimization. Leveraging a fine-tuned Qwen large language model for three critical roles—visual, topological, and specification—the system is integrated with a LangGraph workflow and the IDAES/Pyomo solvers. Evaluated on 82 held-out cases from the OpenIDAES-450 dataset, the approach achieves a 91.5% model construction success rate, with F1 scores of 0.815, 0.791, and 0.782 for unit operations, material streams, and connections, respectively.

chemical process simulationdecision couplingexecutable model construction

Hot Scholars

YL

Yuqiang Li

Central South University
Internal Combustion EngineCombustionEmissionsMechansim
PS

Philippe Schwaller

Assistant Professor, Laboratory of Artificial Chemical Intelligence - EPFL
Deep LearningML for ChemistryReaction PredictionSynthesis Planning
TF

Tianfan Fu

Nanjing University
AI for DrugAI for ScienceLarge Language Model
CW

Connor W. Coley

Massachusetts Institute of Technology
machine learningdrug discoveryautomationsynthetic chemistry
TD

Tomasz Danel

Jagiellonian University, insitro
deep learningcomputer-aided drug designgenerative models