🤖 AI Summary
Predicting cellular responses to unseen chemical perturbations is challenging due to unknown targets and mechanisms, high-dimensional gene expression profiles, and scarce experimental data. This work proposes PerturbPFN, the first framework to integrate synthetic structural priors with structural causal models (SCMs) within a Prior-Function Network (PFN) architecture for context-aware learning. By inferring latent regulatory graphs, sparse intervention targets, and their strengths—and propagating perturbation effects through an SCM-based decoder—the method enables efficient prediction without requiring gradient updates at test time. Combining graph neural networks with a biologically inspired synthetic data simulator, PerturbPFN achieves competitive predictive performance on both real single-cell and synthetic benchmarks while accurately recovering intervention targets, strengths, and regulatory structures, thereby balancing interpretability and computational efficiency.
📝 Abstract
Predicting cellular responses to unseen chemical perturbations is challenging due to unknown targets and mechanisms, high-dimensional expression responses, and limited experimental coverage of the large small-molecule design space. We propose PerturbPFN, a PFN-style amortized model for unknown-target perturbation prediction under a hierarchical synthetic structural prior. Instead of directly regressing high-dimensional expression responses, PerturbPFN infers a latent system graph, sparse atomic intervention targets, and intervention strengths, then propagates their effects through an SCM decoder. The model is trained entirely on prior-predictive synthetic episodes generated from biologically motivated graph and expression simulators, enabling structured in-context learning without test-time gradient updates. We evaluate PerturbPFN on both real single-cell perturbation data and synthetic benchmarks, covering effect prediction, target identification, and regulatory structure discovery. Our results show that PerturbPFN offers a complementary trade-off to specialized baselines, achieving competitive perturbation prediction with low inference cost while exposing interpretable intermediate estimates of targets, strengths, and system structure.