🤖 AI Summary
This study addresses the unclear performance bottlenecks of AI agents in drug discovery by proposing MAGI, a modular agent designed to investigate whether tool orchestration or predictive model accuracy constrains practical outcomes. Methodologically, MAGI employs an open modular architecture coordinating molecular design, optimization monitoring, and SAR analysis with self-revising strategies. It integrates a dual-pathway mechanism combining direct large language model generation with REINFORT-delegated generation, alongside a pluggable scoring service contract. Validation on retrospective pharmaceutical projects reveals that the primary bottleneck lies in the applicability domain of scoring models rather than agent orchestration capabilities. Furthermore, MAGI generates molecules approaching expert-level quality and integrates effectively into existing computational chemistry workflows, thereby clarifying the practical role of AI agents in real-world drug discovery scenarios.
📝 Abstract
Agentic systems increasingly coordinate molecular-design tools, but it is unclear which layer of the stack limits outcomes on real projects. We developed MAGI, an open modular agent that authors objectives, launches and monitors optimization, interprets structure--activity relationships, and revises its strategy accordingly. MAGI generates molecules either directly through the LLM or by delegating to REINVENT 4, with scoring services interchangeable behind a common contract. We tested it across nine retrospective lead-optimization campaigns from three pharmaceutical companies, replayed under fixed temporal cutoffs. Both routes produced valid structures: LLM proposals stayed closer to local chemistry and reached comparable or higher primary activity in fewer operations, whereas REINVENT explored broader chemical space. Whether a campaign met its objective depended on the predictive models, not on the generation route: attainment followed model accuracy on the chemistry proposed, dropping once that chemistry moved outside the model's applicability domain. Separately, a blinded evaluation asked whether the MAGI's output could pass as expert work: chemists were not able to discriminate agentic proposals from held-out compounds, and judged the SAR reasoning broadly plausible yet incomplete. Together, these results position MAGI as a coordination layer pluggable into existing computational chemistry workflows. The ceiling on real projects, however, remains currently set by scorer applicability rather than by tool orchestration.