🤖 AI Summary
This study addresses the limitations of biological interventions, which typically lack closed-loop semantic feedback and require costly paired data acquisition. To overcome these challenges, this work proposes a purely offline learning paradigm that leverages vision-language models (VLMs) to evaluate historical experimental outcomes as reward signals. By integrating offline reinforcement learning, the framework trains a mapping from natural language instructions to biological interventions, enabling effective bio-control without necessitating additional wet-lab experiments or manual verification. The primary contribution lies in introducing the first framework driven exclusively by VLM-based evaluation for intervention policy generation. Empirical results demonstrate that the proposed approach achieves 80.0% accuracy on a held-out test set, surpassing existing baselines and performing comparably to supervised learning methods.
📝 Abstract
Artificial intelligence increasingly serves as a natural-language interface to complex technical systems, letting people accomplish sophisticated tasks by describing what they want rather than specifying how to do it. Extending this interface to living systems is harder: unlike code or images, a biological intervention has no closed-form linguistic meaning, and the paired language-intervention-outcome data needed to learn such a mapping is expensive to collect, since each example requires its own wet-lab experiment. One way around this is to treat an existing archive of interventions and their already-observed outcomes as a fixed, offline dataset, and use a vision-language model to judge, without any new experiments, whether an archived outcome matches a natural-language description. But whether that judgment is reliable enough to train a language-to-intervention mapping on -- without new experiments and without human validation -- has remained unclear. Here we show that a natural-language interface for a living organism -- a xenobot, a synthetic multicellular construct with no nervous system -- can be learned entirely offline this way, using a vision-language model's own judgment as the sole training reward: an instruction is mapped to the intervention already on record as producing the described behavior. This mapping generalizes to entirely new instructions, evaluated against archive data withheld from training (80.0% held-out accuracy vs a $66.7% chance baseline, matching a network trained directly on ground-truth labels).