🤖 AI Summary
This study addresses the challenge of nuclei segmentation in renal pathology images caused by low contrast, dense nuclear distributions, and complex morphologies. We propose a mixed-data fine-tuning framework that stratifies samples into easy, moderate, and hard difficulty levels. By constructing a hierarchical annotation system integrating human-in-the-loop pseudo-label generation with expert consensus annotations, we systematically evaluate fine-tuning strategies across multiple foundation models. Our analysis reveals that the optimal annotation combination is model-dependent. Experimental results demonstrate consistent performance improvements across all evaluated models, with LSP-DETR achieving the highest F1 score of 0.8725 and StarDist exhibiting a substantial increase to 0.8332. These findings validate the effectiveness of the proposed framework for segmenting highly challenging pathological images.
📝 Abstract
Accurate nuclei instance segmentation is essential for quantitative renal pathology, yet general-purpose models often struggle with low contrast, dense nuclei, complex morphology, and strong background staining. In this work, we extended a human-in-the-loop framework by combining 5,901 foundation-model-generated pseudo-labels from well-segmented cases (Easy), 860 newly expert-annotated unresolved challenging cases (Medium), and 198 expert-annotated consensus failure cases (Hard). These annotations, spanning different levels of segmentation difficulty, enabled the systematic evaluation of seven single-source and mixed-source fine-tuning strategies across nine cell segmentation model configurations. Fine-tuning improved all models, with Medium data included in seven of the nine best-performing strategies. LSP-DETR achieved the highest F1 score of 0.8725 with Hard-only fine-tuning, while StarDist showed the largest improvement, increasing from 0.7380 to 0.8332 with Medium-only fine-tuning. These findings show that annotations spanning multiple difficulty levels support effective model adaptation, although the optimal annotation composition remains model dependent.