🤖 AI Summary
This study addresses critical limitations in gene perturbation prediction evaluation, where absolute metrics conflate background responses, lack gene-level validation, and overlook coverage dependencies. To overcome these issues, we construct a specificity-aware, gene-resolved, and coverage-aware evaluation benchmark. Methodologically, we introduce a novel score-pairing strategy that matches predictions against same-split training means, employ gene-coordinate permutation tests to validate gene-level accuracy, and quantify the relationship between representation space coverage and performance gains. Our analysis reveals that high scores in existing methods largely stem from shared background signals rather than genuine predictive power, demonstrating that detectable improvements depend on coverage rather than data scale. Furthermore, we validate the reproducibility of target-specific gains on the RPE1 dataset.
📝 Abstract
Predicting cellular responses to genetic perturbations helps prioritize experiments in single-cell genomics, where exhaustive measurement is infeasible. While computational models increasingly predict these responses, three evaluation deficiencies obscure what their scores demonstrate. First, absolute metrics cannot separate target-specific predictions from a shared background response. Second, common metrics remain high under gene shuffling, so gene-level accuracy is never verified. Third, a score at one training size says nothing about coverage, which depends on representation-space proximity and response-constraining power. We propose AmbiModBench, a specificity-aware, gene-resolved and coverage-aware benchmark. It pairs every score with a training-mean reference fitted on the same split, screens each readout by gene-coordinate permutation, and links embedding distance to response variation. Across K562, RPE1 and Norman, strong absolute scores largely reflect shared background rather than target-specific learning. Widely used readouts track response magnitude distributions rather than the affected genes. Detectable gain follows representation-space coverage rather than training-set size. Nonetheless, on RPE1 the protocol yields a reproducible target-specific gain across five additional splits and three gene selections, which absolute scores alone cannot distinguish from shared background.