Score
Designs and implements algorithms and pipelines to localize anatomical landmarks and to segment anatomical structures and boundaries in medical images or video. Uses these outputs to track landmarks across frames, align surface geometry to underlying anatomy, and build person‑specific anatomical models.
Precise segmentation of orthopedic anatomical landmarks—particularly fine-grained fiducial points in pelvic X-ray images—remains challenging: general foundation models like SAM lack orthopedic priors, while MedSAM targets only coarse-grained organ-level structures, failing to achieve the sub-millimeter accuracy required for clinical diagnosis. To address this, we propose YOLO-SAM, the first hybrid framework integrating YOLOv8 object detection with SAM-based segmentation. YOLOv8 generates high-confidence bounding boxes that serve as spatial prompts for SAM, enabling accurate delineation of 72 critical landmarks and complex bony structures. This two-stage paradigm synergistically combines robust localization with shape-preserving segmentation. Evaluated on pelvic X-ray data, YOLO-SAM significantly improves segmentation accuracy for small and irregularly shaped targets (mAP ↑12.6%, Dice ↑9.3%). The framework establishes an interpretable, deployable paradigm for landmark-driven quantitative orthopedic analysis.
Current medical image keypoint localization tools suffer from insufficient anatomical accuracy, limited model customizability, and poor clinical adaptability. To address these limitations, we present the first modular, extensible PyTorch-based open-source toolkit specifically designed for anatomical landmark localization. It supports both 2D/3D static and adaptive heatmap regression, integrates multi-format I/O, plug-and-play preprocessing pipelines, and flexible model architectures. Our core contribution is a medical imaging–specific framework that explicitly incorporates anatomical structural priors, enables cross-modal generalization, and ensures compatibility with clinical workflows—thereby bridging critical gaps left by generic pose estimation methods in anatomical precision, few-shot adaptation, and multi-center deployment. Experiments demonstrate significant improvements in localization accuracy and substantial reductions in algorithm development, evaluation, and transfer overhead. The toolkit has been successfully deployed across multiple clinical imaging analysis tasks.
Existing automated methods struggle to emulate the reasoning process orthodontists employ when localizing cephalometric landmarks based on anatomical structures and geometric rules. This work proposes a five-stage anatomy-guided initialization pipeline that, for the first time, encodes clinical anatomical knowledge as spatial attention priors and integrates them into an HRNet-W32 detector via confidence-weighted fusion. This inductive bias substantially enhances the model’s generalization capability. Evaluated on 1,502 multi-source cephalograms, the method achieves a mean radial error of 1.04 mm across 25 landmarks—representing a 15.4% improvement over the current state of the art—with 12 landmarks exhibiting sub-millimeter accuracy. Ablation studies confirm that the performance gain stems from improved anatomical plausibility rather than the addition of extra input channels.
This paper addresses the generic anatomical localization problem in multimodal medical imaging (CT/MRI), proposing a GPS-inspired anatomical localization paradigm: mapping arbitrary image coordinates to a standardized atlas space to uniformly support downstream tasks including matching, registration, classification, and segmentation. Methodologically, we design a lightweight regression network that employs sparse input sampling and a modality-agnostic architecture for end-to-end coordinate estimation. Key contributions include: (1) the first cross-modal anatomical coordinate regression framework; (2) zero-shot cross-modal generalization without modality-specific fine-tuning; and (3) sub-millisecond single-point localization latency, enabling real-time inference on commodity CPUs. Experiments demonstrate high accuracy and strong generalization across CT and MRI datasets, establishing a foundational localization capability for multimodal medical image analysis.
In whole-body medical image segmentation, segmented anatomical structures often exhibit geometric inaccuracies, compromising clinical utility. Method: This paper proposes a lightweight, plug-and-play post-processing framework—ShapeKit—that rectifies segmentation outputs during inference without model retraining. ShapeKit integrates explicit geometric constraints with morphological analysis to enforce anatomical plausibility, operating independently of the underlying segmentation model. Contribution/Results: By avoiding architectural modifications or costly retraining, ShapeKit significantly lowers deployment barriers. Evaluated on a multi-organ segmentation benchmark, it achieves an average Dice score improvement of over 8%, substantially exceeding the typical <3% gain from model-level enhancements. This demonstrates that shape-centric post-processing is both effective and practical for improving anatomical fidelity in medical image segmentation.
Evaluating anatomical segmentation models in the absence of ground-truth annotations remains a critical challenge in medical AI. Method: This paper proposes the first anatomy-term-driven unsupervised collaborative evaluation framework. It harmonizes and standardizes multi-model segmentation outputs at the anatomical-structure level via the JSON-Seg schema, enabling automated, cross-model and cross-structure comparison. The framework integrates a 3D Slicer plugin and OHIF Viewer for interactive visualization and provides Summary Plot analytics. Contribution/Results: Evaluated on the NLST CT dataset across 31 anatomical structures using six leading open-source models (e.g., TotalSegmentator, MOOSE), the framework successfully identified high inter-model consistency in lung segmentation and systematic failures in vertebral and rib segmentation—demonstrating its validity and practical utility for ground-truth-free model assessment.
This work addresses the limitations of conventional methods in 3D anatomical landmark detection—namely, low efficiency, poor generalization, and inadequate cross-species adaptability—by proposing LmPT, a unified framework based on a conditional point cloud Transformer. By incorporating a conditioning mechanism, LmPT flexibly accommodates diverse input types and, for the first time, enables joint modeling and transfer learning of homologous skeletal landmarks across species, such as humans and dogs. Experimental results on a newly curated dataset of human femurs and a newly annotated dataset of canine femurs demonstrate that LmPT achieves high accuracy while exhibiting strong cross-species generalization capabilities. The code and datasets are publicly released to facilitate further research.
This study addresses the high cost of manual annotation in deep learning–based anatomical landmark segmentation by proposing a method grounded in statistical shape models (SSMs). Requiring only a single manual annotation, the approach generates thousands of high-quality synthetic 3D liver annotations suitable for training deep networks to segment the liver’s anterior edge and falciform ligament. The method substantially reduces annotation effort while achieving an average Intersection over Union (IoU) of 91.4% on 500 unseen synthetic samples. Qualitative evaluation on clinical data demonstrates strong generalization capability, suggesting the framework’s potential applicability to other anatomical structures beyond the liver.
This work addresses the challenge of graph-based medical image segmentation, which typically relies on scarce point-to-point anatomical landmark annotations. To overcome this limitation, the authors propose Mask-HybridGNet, a framework that trains a fixed-topology graph model using only standard pixel-level segmentation masks. By aligning deformable object boundaries with fixed landmark nodes via Chamfer distance and incorporating edge regularization with differentiable rasterization, the method implicitly learns anatomical correspondences without requiring explicit landmark supervision. The approach enables temporal tracking, cross-sectional reconstruction, and population-level morphological analysis, while extracting stable anatomical graphs from any high-quality segmentation. Evaluated on chest X-ray, cardiac ultrasound, cardiac MRI, and fetal imaging tasks, the method achieves segmentation accuracy comparable to state-of-the-art pixel-level approaches while preserving boundary connectivity and anatomical plausibility.
This study addresses the discrepancy between automatic segmentation and expert manual delineation of deep brain structures in MRI by proposing an anatomical landmark-guided 3D segmentation method. The approach explicitly integrates anatomical landmarks into the deep segmentation pipeline for the first time, employing a global-to-local network to accurately detect 16 key landmarks. These landmarks are then combined with 3D semantic segmentation and landmark-driven post-processing to refine 12 coarse labels into 26 fine-grained structures. By incorporating local anatomical constraints through this strategy, the method substantially improves boundary precision, yielding automated segmentations that align more closely with the manual segmentation protocol defined by the Harvard–Oxford atlas.