Bison: Cross-Dataset Learning for Unseen-Compound Perturbation Prediction

๐Ÿ“… 2026-09-26
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenge of predicting transcriptional responses for unseen compounds caused by cross-dataset heterogeneity. To this end, it establishes an eight-dataset evaluation benchmark and proposes a cross-dataset joint training architecture based on shared gene representations and discrete diffusion models. The core innovation lies in designing a matched-drug contrastive mechanism that transforms complementary screening into shared molecular supervision signals, enabling synergistic multi-source optimization while preserving native gene measurement characteristics. Experimental results demonstrate that the proposed method achieves the highest predictive correlation, with the drug contrastive mechanism improving correlation by 27.4% and comprehensively outperforming independently trained baselines.
๐Ÿ“ Abstract
Predicting transcriptional responses to unseen compounds is limited by fragmented chemical coverage and heterogeneous experimental platforms and gene panels. To assess molecular generalization across these settings, we build on Chem-PerturBridge to benchmark eight datasets with 16,771 compounds, withholding test compounds from every training dataset. This comparison reveals that high overall response agreement can coexist with weak prediction of drug-specific differences, despite reproducible signals across repeated measurements. To exploit complementary chemical supervision while targeting these differences, we introduce Bison: a shared gene representation connects native panels, while two discrete diffusion models compose context-dependent responses with molecular deviations learned through matched drug-contrast supervision. A single Bison model jointly trained across all eight datasets achieves the highest mean overall-response and drug-contrast Pearson correlations on the full benchmark in comparison with 11 methods trained independently per dataset. Compared with dataset-specific training of the same architecture, joint training increases mean drug-contrast correlation by 27.4\%, with gains across all eight datasets and improvements in overall response prediction. These results demonstrate how matched drug contrasts turn complementary screens into shared molecular supervision for unseen-drug response prediction while preserving native gene measurements.
Problem

Research questions and friction points this paper is trying to address.

unseen-compound perturbation prediction
cross-dataset learning
transcriptional response
chemical space fragmentation
drug-specific differences
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Dataset Learning
Discrete Diffusion Models
Shared Gene Representation
Drug-Contrast Supervision
Perturbation Prediction
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.