π€ AI Summary
This study addresses the challenges of clinical prediction from transcriptomic data, characterized by high dimensionality, strong correlations, scarce annotations, and limited performance of existing large models. To this end, we propose a 4.2M-parameter tabular foundation model. Methodologically, we design a transcriptome-aware pretraining paradigm coupled with a semi-synthetic data strategy, integrating a parameter-efficient recurrent architecture with in-context learning to enable gradient-free inference. Furthermore, we introduce a Cox residual regression-based dimensionality reduction technique that facilitates right-censored survival prediction without task-specific training. Evaluated across 80 tasks, the proposed model achieves superior overall performance while reducing parameter count by 387-fold compared to existing approaches. It supports second-level predictions on standard CPUs, thereby enabling efficient and precise clinical translation.
π Abstract
Gene expression is widely measured in biomedicine, yet clinical outcome prediction remains challenging due to high dimensionality, strong feature correlations, and limited labeled data. Large self-supervised transcriptomic foundation models often fail to outperform simple supervised baselines. Tabular foundation models offer an alternative through in-context learning, but are typically pretrained on generic synthetic data rather than transcriptomic structure. We ask whether transcriptomics-aware pretraining, rather than scale, is the missing ingredient. Towards this end, we introduce GeneICL, a 4.2M-parameter tabular foundation model combining a semi-synthetic pretraining prior built from measured bulk expression profiles with a parameter-efficient recurrent architecture. We further enable right-censored survival prediction via a training-free reduction to regression using Cox partial-likelihood residuals. We evaluate GeneICL on 80 clinical outcome-prediction tasks spanning classification, regression, and survival. Tabular foundation models consistently outperform self-supervised transcriptomic models, while GeneICL achieves the best overall rank among evaluated foundation models and tuned baselines. GeneICL does so with up to 387$\times$ fewer parameters, no gradient updates at inference, and predictions within seconds on a laptop CPU.