🤖 AI Summary
This study addresses the high inference latency of large Vision Transformer (ViT)-based pathology segmentation models, which hinders their deployment in real-time clinical settings. We propose a lightweight nuclei instance segmentation method that distills the UNI2-h foundation model into a ConvNeXt-Tiny architecture. Our investigation reveals that multi-scale gated convolutions are ineffective within the ViT framework and demonstrates that output-level knowledge distillation alone suffices for efficient cross-architecture knowledge transfer. The resulting student model reduces parameters to 1/20th of the teacher network while retaining 98.8% of its mean Panoptic Quality (mPQ) and achieving a 21.8× inference speedup. This approach effectively balances accuracy and computational efficiency, satisfying the stringent requirements of real-time clinical analysis.
📝 Abstract
Nuclei instance segmentation is a core task in digital pathology, yet high-accuracy models rely on large vision transformer (ViT) encoders whose inference speed cannot meet real-time clinical demands. We propose a lightweight scheme that distills the UNI2-h pathology foundation model into a ConvNeXt-Tiny student (Ours-T, 34.7M parameters, 1/20 of the teacher) via output-level knowledge distillation. Ours-T achieves an mPQ of 0.519 on PanNuke (98.8% of the teacher), a zero-shot bPQ of 0.668 on MoNuSeg, and an inference speed of 634.3 img/s, requiring only 0.045 s for full-resolution 1024^2 analysis (21.8x speedup). Experiments further show that multi-scale gated convolution (MALA) yields no gain under ViT encoders, and output-level distillation alone suffices for efficient knowledge transfer.