🤖 AI Summary
This study addresses the challenge in molecular optimization where single-round generation struggles to simultaneously balance validity, property improvement, and structural similarity. To overcome this limitation, this work proposes a bounded multi-turn proxy reinforcement learning framework. Utilizing the Qwen large language model as its backbone, the method constructs iterative "proposal-feedback-revision" trajectories. A molecular editor is trained via Group Relative Policy Optimization (GRPO) with aggregated shaped rewards, while verifier feedback guides synergistic multi-objective optimization. Evaluated on the MuMOInstruct benchmark, the proposed approach achieves state-of-the-art performance on the product metric of multi-objective property success rate and structural similarity. These results demonstrate that the multi-turn feedback mechanism substantially enhances molecular optimization performance.
📝 Abstract
Molecular optimization is inherently iterative: a candidate is proposed, evaluated against several objectives, and revised while preserving a relationship to the source molecule. Most instruction-following models instead emit one edited molecule, forcing validity, property improvement, and similarity control into a single response. We introduce MARCO, an evaluator-grounded reinforcement-learning framework that trains molecular editors on bounded proposal--feedback--revision trajectories. MARCO aggregates shaped turn rewards into an undiscounted trajectory return for group-relative policy optimization. We evaluate two consequences of this training: Same-1 tests the trained policy under a one-response budget, while Same-5 tests whether the same policy can use verifier feedback when up to five responses are available. Across the three-objective MuMOInstruct benchmark, three Qwen backbones, and seen/unseen instruction splits, SFT-initialized MARCO obtains the highest product of property success rate and similarity in every reported primary setting. Same-5 further improves the observed score under the tested budget, while four-objective and public-checkpoint experiments test transfer across constraint sets and initialization regimes.