Score
Designs and implements segmentation systems that operate on streaming anatomical image data to produce continuously updated, temporally consistent anatomical labels. This includes methods for anatomical grounding to anchor predictions to known structures, reduce temporal drift, and enable real‑time inference with online updates or adaptation during ongoing procedures.
Evaluating anatomical segmentation models in the absence of ground-truth annotations remains a critical challenge in medical AI. Method: This paper proposes the first anatomy-term-driven unsupervised collaborative evaluation framework. It harmonizes and standardizes multi-model segmentation outputs at the anatomical-structure level via the JSON-Seg schema, enabling automated, cross-model and cross-structure comparison. The framework integrates a 3D Slicer plugin and OHIF Viewer for interactive visualization and provides Summary Plot analytics. Contribution/Results: Evaluated on the NLST CT dataset across 31 anatomical structures using six leading open-source models (e.g., TotalSegmentator, MOOSE), the framework successfully identified high inter-model consistency in lung segmentation and systematic failures in vertebral and rib segmentation—demonstrating its validity and practical utility for ground-truth-free model assessment.
In whole-body medical image segmentation, segmented anatomical structures often exhibit geometric inaccuracies, compromising clinical utility. Method: This paper proposes a lightweight, plug-and-play post-processing framework—ShapeKit—that rectifies segmentation outputs during inference without model retraining. ShapeKit integrates explicit geometric constraints with morphological analysis to enforce anatomical plausibility, operating independently of the underlying segmentation model. Contribution/Results: By avoiding architectural modifications or costly retraining, ShapeKit significantly lowers deployment barriers. Evaluated on a multi-organ segmentation benchmark, it achieves an average Dice score improvement of over 8%, substantially exceeding the typical <3% gain from model-level enhancements. This demonstrates that shape-centric post-processing is both effective and practical for improving anatomical fidelity in medical image segmentation.
This paper addresses the generic anatomical localization problem in multimodal medical imaging (CT/MRI), proposing a GPS-inspired anatomical localization paradigm: mapping arbitrary image coordinates to a standardized atlas space to uniformly support downstream tasks including matching, registration, classification, and segmentation. Methodologically, we design a lightweight regression network that employs sparse input sampling and a modality-agnostic architecture for end-to-end coordinate estimation. Key contributions include: (1) the first cross-modal anatomical coordinate regression framework; (2) zero-shot cross-modal generalization without modality-specific fine-tuning; and (3) sub-millisecond single-point localization latency, enabling real-time inference on commodity CPUs. Experiments demonstrate high accuracy and strong generalization across CT and MRI datasets, establishing a foundational localization capability for multimodal medical image analysis.
Automatic pancreatic segmentation in abdominal CT suffers from low accuracy and high false-negative rates. Method: We propose a 3D full-resolution nnU-Net framework integrated with anatomical prior knowledge—specifically, multi-organ anatomical labels from TotalSegmentator, leveraging spatial constraints from neighboring organs as strong priors, and jointly trained end-to-end on the PANORAMA dataset. Contribution/Results: Our method significantly improves segmentation robustness: Dice score increases by 6.0% (p < 0.001), Hausdorff distance decreases by 36.5 mm (p < 0.001), and achieves 100% pancreatic detection with zero false negatives. This work empirically validates the critical value of anatomical priors for fine-grained single-organ segmentation, establishing a highly reliable foundation for radiomic biomarker extraction and pancreatic lesion identification.
Medical image segmentation for multiple organs suffers from prohibitively high annotation costs, while existing public datasets are limited in scale and organ coverage. To address this, we propose the first closed-loop active learning framework integrating multi-model debiased labeling, uncertainty-driven active sampling, saliency heatmap–guided interactive correction, and error typology modeling—enabling bidirectional iterative optimization between AI and human annotators. Within just three weeks, our framework enabled construction of the largest abdominal CT multi-organ segmentation dataset to date, comprising 8,448 cases, 3.2 million slices, and annotations for eight critical organs—accelerating annotation by over 2,600× compared to estimated expert-only labeling (31 years). Annotation quality matches or exceeds that of domain experts. The dataset is publicly released, establishing a high-quality benchmark resource for large-scale medical image analysis.
Existing medical report generation methods rely on expert-annotated detection modules, incurring high annotation costs and exhibiting poor cross-dataset generalizability. To address this, we propose a self-supervised anatomical consistency learning framework that requires no expert annotations. Our approach constructs a hierarchical anatomical graph structure and jointly optimizes recursive region reconstruction and region-level contrastive learning. Furthermore, we introduce a prompt-driven attention guidance mechanism to achieve precise multimodal alignment between images and text in both anatomical space and semantic space. The proposed method significantly enhances report interpretability and clinical applicability: it improves vocabulary accuracy by 10%, increases clinical validity by 25%, and outperforms state-of-the-art vision foundation models by 8% on zero-shot visual localization.
This work addresses the conflicting objectives and training instability between active learning and semi-supervised learning in medical image segmentation under extremely limited annotation budgets. To resolve this, the authors propose RegAL, a framework that unifies both objectives through topology-aware Pareto optimization. RegAL selects informative samples by jointly considering voxel-wise uncertainty and feature diversity, and leverages the same criteria to guide diffeomorphic registration-based data augmentation, thereby enabling stable training from cold start. Evaluated on BraTS 2021, dHCP, and ProstateX datasets, RegAL consistently outperforms existing methods, demonstrating particularly robust performance at very low annotation rates across key metrics including Dice score, Average Surface Distance (ASD), and Hausdorff Distance at 95% (HD95).
Existing medical image segmentation models are constrained by single-segmentation protocols or reliance on manual prompts, limiting their ability to automatically support concurrent segmentation across diverse semantic protocols—such as tissue, anatomy, and pathology—in unseen domains. Method: We propose a novel multi-protocol consistent segmentation paradigm, built upon a vision–semantics-aligned multi-head decoder architecture. It incorporates protocol-aware feature disentanglement and cross-image semantic consistency regularization, trained via a zero-shot domain generalization strategy to enable fully automatic, multi-label, semantically consistent segmentation without human intervention. Contribution/Results: Evaluated on seven hold-out datasets, our method achieves state-of-the-art performance in multi-protocol segmentation plausibility, cross-image semantic consistency, and zero-shot generalization—significantly alleviating foundational model dependencies on single-protocol constraints or strong manual prompting.
SAM faces critical challenges in medical image segmentation, including ill-defined boundaries, insufficient modeling of anatomical relationships, and lack of uncertainty quantification. To address these, we propose KG-SAM—a novel framework that first integrates an anatomical knowledge graph (KG) and an energy-driven conditional random field (CRF) into the SAM architecture. The KG encodes prior anatomical structural constraints, while the CRF jointly optimizes segmentation boundaries and models prediction uncertainty. An uncertainty-aware fusion module further enhances robustness. This design significantly improves segmentation consistency, interpretability, and clinical applicability. Extensive evaluation on multi-center datasets demonstrates state-of-the-art performance: Dice scores of 82.69% for prostate segmentation, and 78.05% and 79.68% for abdominal MRI and CT segmentation, respectively—each substantially surpassing existing methods.
This study addresses the challenges of anatomical structure recognition in minimally invasive surgery, which are hindered by scarce annotated data and the bias of existing methods toward natural scenes. To overcome these limitations, the authors introduce ATLAS-120k, a large-scale semantic segmentation dataset comprising 120,000 video frames spanning 14 distinct surgical procedures. They propose the context-aware ATLAS model, which integrates foundation model embeddings with lightweight temporal reasoning, leveraging surgical type, procedural phase, and short-term visual memory to enhance both accuracy and temporal consistency. An innovative, scalable annotation pipeline combining expert labeling, automated propagation, and surgeon validation is also developed. The proposed approach achieves real-time inference while significantly improving segmentation performance, thereby laying a foundation for surgical navigation and clinical decision-support systems.