MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in molecular optimization where single-round generation struggles to simultaneously balance validity, property improvement, and structural similarity. To overcome this limitation, this work proposes a bounded multi-turn proxy reinforcement learning framework. Utilizing the Qwen large language model as its backbone, the method constructs iterative "proposal-feedback-revision" trajectories. A molecular editor is trained via Group Relative Policy Optimization (GRPO) with aggregated shaped rewards, while verifier feedback guides synergistic multi-objective optimization. Evaluated on the MuMOInstruct benchmark, the proposed approach achieves state-of-the-art performance on the product metric of multi-objective property success rate and structural similarity. These results demonstrate that the multi-turn feedback mechanism substantially enhances molecular optimization performance.
📝 Abstract
Molecular optimization is inherently iterative: a candidate is proposed, evaluated against several objectives, and revised while preserving a relationship to the source molecule. Most instruction-following models instead emit one edited molecule, forcing validity, property improvement, and similarity control into a single response. We introduce MARCO, an evaluator-grounded reinforcement-learning framework that trains molecular editors on bounded proposal--feedback--revision trajectories. MARCO aggregates shaped turn rewards into an undiscounted trajectory return for group-relative policy optimization. We evaluate two consequences of this training: Same-1 tests the trained policy under a one-response budget, while Same-5 tests whether the same policy can use verifier feedback when up to five responses are available. Across the three-objective MuMOInstruct benchmark, three Qwen backbones, and seen/unseen instruction splits, SFT-initialized MARCO obtains the highest product of property success rate and similarity in every reported primary setting. Same-5 further improves the observed score under the tested budget, while four-objective and public-checkpoint experiments test transfer across constraint sets and initialization regimes.
Problem

Research questions and friction points this paper is trying to address.

Molecular Optimization
Instruction-following Models
Multi-round Iteration
Conditional Molecular Editing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-round Agentic Reinforcement Learning
Molecular Optimization
Group-Relative Policy Optimization
Trajectory Return
Verifier Feedback
💼 Related Jobs
No related jobs found.
S
Shicheng Fang
Fudan University
Yuxin Wang
Yuxin Wang
Fudan University
Zhuo Yang
Zhuo Yang
Xidian University & Shanghai AI Laboratory
Lauge Language ModelAI for Science
X
Xiaohu Xu
Shanghai Innovation Institute
J
Jiahao Lu
Fudan University
C
Chuanyuan Tan
Soochow University
T
Tong Zhu
Shanghai Innovation Institute
Y
Yining Zheng
Fudan University
X
Xipeng Qiu
Fudan University