From Patches to Evidence Balls: Class-Conditioned Evidence Retrieval for Few-Shot Whole Slide Image Classification

πŸ“… 2026-08-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of learning reliable global aggregation mechanisms from sparse, spatially distributed, and class-dependent diagnostic cues in whole-slide images (WSIs) under few-shot supervision, which hinders class-specific reasoning and interpretability. To this end, the authors propose EviBall, a framework that organizes local image patches into semantically and spatially coherent β€œevidence balls” and retrieves supportive evidence via language- or molecular-guided class queries, enabling class-conditional evidence representation and prediction. By reframing few-shot WSI classification as a structured evidence retrieval and class competition problem, EviBall substantially reduces reliance on unconstrained global aggregation. Experiments demonstrate that EviBall consistently outperforms existing multiple instance learning (MIL) and vision-language baselines across diverse few-shot settings on four morphological and molecular endpoint tasks, while providing spatially precise, class-specific interpretable evidence.
πŸ“ Abstract
Whole slide image (WSI) classification is an evidence-driven task, where diagnostic cues are often sparse, spatially organized, and class-dependent. Existing MIL and vision-language methods aggregate a large pool of patch features into a single global slide representation. Under few-shot supervision, limited slide-level labels make it difficult to learn a reliable aggregation mechanism that organizes sparse local cues into compact and coherent diagnostic evidence. Moreover, a shared slide representation compresses evidence supporting a candidate class and its alternatives into the same feature, limiting class-specific reasoning and interpretability. To address these issues, we propose EviBall, a class-conditioned evidence retrieval framework for few-shot WSI classification. EviBall organizes local patches into Evidence Balls through semantic-spatial assignment and center refinement, yielding compact and spatially coherent evidence units under weak supervision. It then uses task-specific class queries, including language-guided queries for morphology-oriented tasks and molecular-guided queries for molecular endpoint prediction, to retrieve supporting evidence balls and produce class-conditioned evidence representations for direct class-wise prediction. By introducing structured evidence units and task-relevant semantic guidance, EviBall reduces the reliance on learning an unconstrained global aggregation mechanism from scarce slide-level labels. It therefore reformulates few-shot WSI classification as structured evidence retrieval and competition among candidate classes. Extensive experiments across four morphology-oriented and molecular endpoint WSI tasks demonstrate that EviBall consistently outperforms conventional and vision-language MIL baselines under diverse few-shot settings, while providing spatially localized and class-specific evidence for each prediction.
Problem

Research questions and friction points this paper is trying to address.

few-shot learning
whole slide image classification
evidence retrieval
class-conditioned representation
multiple instance learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evidence Balls
Class-Conditioned Retrieval
Few-Shot WSI Classification
Semantic-Spatial Assignment
Vision-Language Guidance