🤖 AI Summary
To address high computational overhead, excessive GPU memory consumption, and accuracy degradation due to anatomical variability in real-time clinical prostate MRI segmentation, this paper proposes a lightweight and efficient U-Net variant. Methodologically, it introduces two key innovations: (1) a novel dynamic K-NN attention mechanism that adaptively adjusts the number of neighboring nodes based on spatial location, thereby balancing local contextual modeling capability and computational efficiency; and (2) a Cross Stage Partial (CSP) encoder that significantly reduces feature redundancy, lowering both computational cost and GPU memory usage. Evaluated on the PROMISE12 and PROSTATEx benchmarks, the model achieves a Dice coefficient of 92.6%, with a 37% speedup in inference time and a 42% reduction in GPU memory footprint. To our knowledge, this is the first approach that simultaneously maintains high segmentation accuracy and meets the stringent latency and resource constraints required for real-time deployment on clinical workstations.
📝 Abstract
Real-time deployment of prostate MRI segmentation on clinical workstations is often bottlenecked by computational load and memory footprint. Deep learning-based prostate gland segmentation approaches remain challenging due to anatomical variability. To bridge this efficiency gap while still maintaining reliable segmentation accuracy, we propose KLO-Net, a dynamic K-Nearest Neighbor attention U-Net with Cross Stage Partial, i.e., CSP, encoder for efficient prostate gland segmentation from MRI scan. Unlike the regular K-NN attention mechanism, the proposed dynamic K-NN attention mechanism allows the model to adaptively determine the number of attention connections for each spatial location within a slice. In addition, CSP blocks address the computational load to reduce memory consumption. To evaluate the model's performance, comprehensive experiments and ablation studies are conducted on two public datasets, i.e., PROMISE12 and PROSTATEx, to validate the proposed architecture. The detailed comparative analysis demonstrates the model's advantage in computational efficiency and segmentation quality.