Exploiting Spatial Structure for Transductive Few-Shot Classification of Whole-Slide Images

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high prediction noise in zero-shot whole slide image (WSI) analysis and the difficulty of existing few-shot methods in accommodating spatial structures and class imbalance. We introduce SlideTIM, the first framework to incorporate latent-constrained transductive inference memory (LC-TIM) into WSI analysis. Built upon vision-language models and transductive reasoning, SlideTIM enforces patch-level semantic consistency through spatial-latent regularization and calibrates class proportions via class distribution priors, enabling global joint optimization. Evaluated on four datasets, the proposed method achieves an 8.1 percentage point improvement in Macro-F1 over the strongest baseline and a 19.4 percentage point gain compared to zero-shot approaches, significantly outperforming existing TIM variants.
📝 Abstract
Automating the analysis of whole-slide images (WSIs), a key step in cancer diagnosis, has high clinical value, as it can reduce pathologist's workload while improving diagnosis accuracy. Recently, vision-language models have shown promising performance for patch-level classification without requiring any annotation, yet these zero-shot (ZS) predictions remain noisy on fine-grained tasks and must be further refined. A promising direction is to refine all predictions jointly, i.e., a transductive approach. However, most existing methods are not tailored to WSIs. We thus propose SlideTIM, an adaptation to WSIs of the recent transductive approach LC-TIM, which introduces a combined spatial--latent regularizer together with a prior on the patch class distribution. The former enforces spatially and semantically close patches to receive the same predictions, while the prior calibrates the predicted class proportions. Together, they address the complex spatial organization and the strong class imbalance of WSIs. Evaluated on four histology datasets, SlideTIM consistently outperforms all TIM variants, improving the macro-F1 by +8.1pp over the best competing baseline at 1 shot. Compared to the ZS, it raises the macro-F1 by +19.4pp at 1 shot. The code will be made available after submission.
Problem

Research questions and friction points this paper is trying to address.

Whole-Slide Images
Few-Shot Classification
Transductive Learning
Class Imbalance
Spatial Structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transductive Few-Shot Classification
Whole-Slide Images
Spatial-Latent Regularizer
Class Distribution Prior
Vision-Language Models
🔎 Similar Papers
2024-08-15European Conference on Computer VisionCitations: 0