PertReason: A Knowledge-Grounded Benchmark and Framework for Cell-State-Conditioned Mechanistic Reasoning of Perturbation Effects

๐Ÿ“… 2026-07-21
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Existing models often produce accurate predictions in scientific reasoning tasks while relying on incorrect mechanistic explanations, particularly failing to ensure mechanistic plausibility under distributional shifts across cell types and perturbations. To address this, this work introduces PertReasonQA, a novel evaluation benchmark, and PertReasonLM, a corresponding framework that establishes the first mechanism-reasoning assessment system conditioned on cellular states. The approach integrates single-cell multi-omics perturbation data, cell-type-specific knowledge graphs, and dynamically conditioned pathway mechanisms, and incorporates a mechanismโ€“outcome alignment training paradigm. Experiments reveal that prevailing models exhibit latent failures in logical consistency, contextual awareness, and directional mechanistic reasoning. In contrast, PertReasonLM substantially enhances mechanistic faithfulness, enabling more reliable and interpretable inference of perturbation effects under complex distributional shifts.
๐Ÿ“ Abstract
Evaluating machine learning in scientific domains requires separating correct predictions from correct reasons under realistic distribution shifts. We introduce PertReason, a knowledge-grounded benchmark and framework suite for cell-state--conditioned reasoning about perturbation effects. At its core, PertReasonQA is a benchmark that tests whether models can generate mechanistically faithful explanations while remaining robust to complex shifts, such as new cells and unseen perturbations. PertReasonQA combines single-cell genetic and chemical perturbation data across multiple cellular contexts with knowledge graphs, and dynamically conditions pathways on cell-specific basal states to avoid generic memorization. Evaluations on state-of-the-art models reveal systematic gaps between predictive accuracy and mechanistic reasoning. Specifically, these models exhibit failure modes largely invisible to standard benchmarks, such as deriving correct answers through flawed logic, ignoring cellular context, and generating directionally inconsistent mechanisms. As a reference probe of the benchmark, we present PertReasonLM, a large language model trained to align outcome predictions with context-specific mechanistic reasoning. Our model targets the identified failure modes by grounding rationales in context-specific pathways and tightening agreement between outcomes and mechanisms. Together, we provide a diagnostic framework for exposing and mitigating failures in faithful reasoning in data-rich scientific systems.
Problem

Research questions and friction points this paper is trying to address.

mechanistic reasoning
perturbation effects
cell-state conditioning
distribution shift
knowledge grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

mechanistic reasoning
cell-state conditioning
knowledge-grounded benchmark
perturbation effects
context-specific pathways