MolDesignBench: Evaluating LLM-based Agent for Scenario-grounded Molecular Design

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决现有基准未能反映真实分子设计复杂性的问题,提出MolDesignBench,通过结合隐式要求与显式约束评估LLM代理,揭示了其在隐式约束理解和不可行性检测上的不足。
📝 Abstract
Real-world molecular design remains challenging for large language model (LLM)-based agents. It requires them to interpret design contexts, satisfy multiple constraints, identify infeasible specifications, and reason over multi-step tool outputs. Existing benchmarks do not capture this complexity, focusing instead on explicit and narrow constraints, only feasible problems, and single-path solutions. To address this gap, we propose MolDesignBench, a scenario-grounded benchmark that more closely reflects real-world molecular design for evaluating tool-augmented LLM agents. MolDesignBench comprises 2K generation and optimization instances that combine implicit requirements embedded in design narratives with explicit property and functional-group constraints, including infeasible cases, and require the effective use of 17 specialized chemistry tools. Experiments across diverse frontier LLMs reveal low success rates--with the best achieving only $\sim43$\%--and frequent failures in implicit-constraint reasoning, infeasibility detection, and tool reasoning. The corresponding fine-grained failure-mode analysis identifies implicit constraint interpretation and infeasibility detection as the primary bottlenecks, establishing MolDesignBench as a rigorous testbed to guide future research on chemical agents. The benchmark, tool interface, and evaluation code are publicly available.
Problem

Research questions and friction points this paper is trying to address.

molecular design
large language model
benchmark
constraint satisfaction
infeasibility detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

MolDesignBench
scenario-grounded molecular design
tool-augmented LLM agents
implicit constraints
infeasibility detection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yongjun Jeong
Department of Artificial Intelligence, Korea University
H
Hanbum Ko
Department of Artificial Intelligence, Korea University
Y
Ye Rin Kim
Department of Artificial Intelligence, Korea University
Chanhui Lee
Chanhui Lee
GIST
Computer VisionAdversarial Attack
R
Rodrigo Hormazabal
LG AI Research
J
Jaewan Lee
LG AI Research
S
Sehui Han
LG AI Research
S
Sungbin Lim
Department of Statistics, Korea University
Sungwoong Kim
Sungwoong Kim
Associate Professor, Korea University
artificial general intelligence