🤖 AI Summary
This study addresses the challenge of training and evaluating AI for drug target discovery, given the absence of causal ground truth in real-world biobanks. To overcome this limitation, we construct a synthetic biobank with latent causal structures, programmatically generating multimodal data encompassing genomics, proteomics, and MRI to drive large language model agents through an end-to-end target discovery pipeline, from phenotypic modeling to virtual intervention. The core innovation lies in introducing a verifiable reward mechanism wherein the causal structure is disclosed to evaluators but concealed from agents, thereby enabling scalable benchmarking of scientific tasks. Experimental results demonstrate that top-performing models recover an average of 64% of causally driven proteins; however, distinguishing non-causal confounders remains a significant challenge.
📝 Abstract
Drug target discovery requires distinguishing molecules that causally drive disease from those that are merely associated with it. Training and evaluating AI agents to perform this workflow end-to-end is difficult because real world biobanks lack known causal ground truth and participant-level data is access controlled. We introduce DrugTargetWorld, a framework that procedurally generates simulated biobanks, or "worlds," with known but concealed causal structure. Each world contains genotypes, proteins, health records, outcomes, and synthetic magnetic resonance imaging (MRI) for 54,000 participants. Agents must construct a disease phenotype, identify causal driver proteins, infer the beneficial direction of modulation, and optionally conduct virtual 'wet lab' experiments. We evaluated nine agents in 540 episodes across 20 cardiovascular worlds and three experimental budgets. Opus 5 and GPT-5.6 Sol achieved the highest mean composite scores, 39.98 and 35.38 of 100, respectively, and both recovered 64% of causal drivers on average. However, no agent reliably distinguished misleading non-causal proteins, and performance remained limited by the integrative judgments required to connect phenotype construction, causal evidence, and intervention decisions. By making each world's causal structure known to the evaluator but hidden from the agent, DrugTargetWorld turns end-to-end drug target discovery into a scalable training and evaluation problem with verifiable reward.