🤖 AI Summary
This study addresses the high cost and scarcity of expert annotation for tumor proportion score (TPS) assessment in non-small cell lung cancer (NSCLC), as well as the limitations of existing multiple instance learning (MIL) approaches in handling a large number of non-expressive (zero-class) image patches. To overcome these challenges, the authors propose an end-to-end deep MIL framework that integrates CNN-based feature embedding with a multi-class network to model patch-level characteristics. Notably, they introduce, for the first time, a zero-inflated Beta distribution to probabilistically model TPS at the whole-slide level, effectively capturing both the excess of zero-class instances and the continuous nature of TPS values. Compared to linear and ridge regression baselines, the proposed method significantly improves prediction accuracy and provides calibrated confidence estimates through the concentration parameter of the learned distribution, yielding more reliable and interpretable TPS assessments.
📝 Abstract
Accurate assessment of tumor proportion score (TPS) in non-small cell lung cancer (NSCLC) is critical for treatment planning and prognosis. Key challenges include the tedious manual work required to annotate each slide, combined with the limited number of experts certified for this task. Multiple instance learning (MIL) has proven to be an effective approach for predicting TPS scores at the slide level; however, existing methods struggle with non-expressive (zero class) images. Our approach involves two models: (1) an embedding-extraction and multiclass-classification network that captures the histopathological features of individual patches, and (2) a MIL model that aggregates these embeddings to predict zero-inflated beta (ZIBeta) parameters representing the overall TPS probability distribution for the entire slide. Using only slide-level TPS scores as labels, we demonstrate how this end-to-end framework can leverage a novel distribution-based architecture to improve prediction accuracy and explainability. ZIBeta modeling significantly outperforms baseline linear and ridge regression while capturing expected accuracy through distribution concentration.