🤖 AI Summary
This study addresses the challenge of noisy and misleading neighbor retrieval in mechanism-of-action prediction from Cell Painting data. To this end, we propose PhenoAIR, a multi-agent framework that pioneers a paradigm shift from representation matching to calibrated evidential reasoning. Methodologically, the framework employs a controller to guide refined retrieval and offline source reliability calibration, enabling reliability-aware assessment and filtering of evidence. Technically, it integrates multi-agent collaboration, an evidential memory mechanism, and high-content morphological analysis. Extensive evaluations on the JUMP benchmark demonstrate that PhenoAIR consistently outperforms conventional representation-matching approaches and large language model baselines across multiple settings, thereby validating the effectiveness of the proposed evidential reasoning paradigm for phenotypic mechanism prediction.
📝 Abstract
Cell Painting is a high-content morphological profiling assay widely used for phenotype-based biological inference, with mechanism of action (MOA) prediction as a central application. Existing approaches largely formulate Cell Painting-based inference as representation matching, assigning predictions from nearby reference perturbations in morphological feature space. However, retrieved neighbors are often noisy and partially misleading evidence due to batch effects, non-specific cytotoxicity, phenotypic convergence, and source-dependent variability. We reformulate Cell Painting-based MOA prediction as a calibrated evidence reasoning problem, where retrieved neighbors are treated as uncertain observations that must be evaluated, compared, and sometimes rejected before supporting a mechanistic conclusion. We propose PhenoAIR, a reliability-aware multi-agent framework that maintains a candidate-centric evidence memory and performs controller-guided refinement over phenotype- and mechanism-side evidence. PhenoAIR uses offline reference-set calibration to weight evidence by source reliability, phenotype stability, and mechanism-level confusion. We evaluate PhenoAIR on a benchmark constructed from JUMP Cell Painting profiles and annotations, covering controlled, realistic, and discovery-oriented open-world MOA prediction settings. PhenoAIR outperforms representation-matching and LLM-based baselines across all settings.