SAMI3D-DW: Interactive Segmentation of Any 3D Medical Images
为解决3D医学图像中复杂解剖结构和病理的分割难题,提出SAMI3D-DW模型,通过大规模数据训练实现高效交互式分割。
为解决3D医学图像中复杂解剖结构和病理的分割难题,提出SAMI3D-DW模型,通过大规模数据训练实现高效交互式分割。
Existing methods for evaluating transferability in 3D medical image segmentation rely on time-consuming fine-tuning, which struggles to meet the stringent demands for boundary precision and anatomical consistency. This work proposes the first fine-tuning-free, topology-driven framework that aligns sparse features with semantic labels via minimum spanning trees (MSTs). It assesses transferability at dual scales—local boundary token separability (LBTC) and global representation topological divergence (GRTD)—and incorporates a task-adaptive gated fusion mechanism. Theoretically, we prove that the MST leakage rate constitutes a finite-sample lower bound of the Bayes error and reveal that randomly initialized decoders stabilize topological alignment. Evaluated on a large-scale benchmark encompassing 114,000 3D medical images, our method achieves state-of-the-art performance, improving the weighted Kendall metric by 0.36 on average and accelerating evaluation by 56×.
This study addresses the challenge of effectively integrating global and local features in deep learning–based classification of benign and malignant pulmonary nodules, a task further complicated by the lack of clinical validation in existing approaches. To this end, we propose DeepFAN, a Transformer-based model trained on over 10,000 pathologically confirmed cases—the largest such dataset reported to date—and rigorously evaluated through a multicenter, multi-reader clinical trial. DeepFAN synergistically combines global and local imaging features and incorporates explainability analysis to assess feature contributions. The model achieves an internal test AUC of 0.939 and a clinical trial AUC of 0.954. When used as a diagnostic aid, it improves junior radiologists’ average AUC by 10.9%, with significant gains in sensitivity, specificity, and accuracy, and elevates inter-rater agreement from fair to moderate.
This work proposes an anatomy-informed synthetic supervised pre-training framework that addresses the limitations of existing methods, which rely on generic geometric shapes and fail to capture the morphological complexity, spatial layout, and inter-organ relationships inherent in real anatomical structures, thereby lacking the global structural priors essential for medical imaging. By integrating anatomical logic—such as spatial anchors and organ topology graphs—into the synthetic data generation process, the framework leverages a lightweight repository of realistic anatomical shapes and a structure-aware sequential placement strategy to enhance physiological plausibility. Evaluated on the Vision Transformer architecture, the method outperforms the current state-of-the-art FDSL baseline by 1.74% on BTCV and surpasses SSL approaches by 1.66% on MSD, while demonstrating robust scalability with increasing synthetic data volume.
This work addresses a critical limitation in existing synthetic data approaches for medical Vision Transformers (ViTs), which often neglect the intricate tissue textures and noise inherent in real medical images, leading to inaccurate learning of anatomical boundaries. To resolve this, the authors propose a physics-inspired spatially decoupled synthesis framework that, for the first time, identifies and mitigates an optimization conflict—termed boundary aliasing—between high-frequency texture modeling and boundary delineation. By introducing a gradient masking buffer and a core texture injection mechanism, the method enables orthogonal yet synergistic learning of shape and texture. Integrating equation-driven supervised learning, boundary distance fields, and physically informed spectral texture synthesis, the approach outperforms current FDSL and real-data-pretrained self-supervised methods by 1.43% and 1.51% on the BTCV and MSD datasets, respectively, establishing an efficient, annotation-free, and scalable training paradigm for medical ViTs.
为解决3D医学图像中复杂解剖结构和病理的分割难题,提出SAMI3D-DW模型,通过大规模数据训练实现高效交互式分割。
Existing methods for evaluating transferability in 3D medical image segmentation rely on time-consuming fine-tuning, which struggles to meet the stringent demands for boundary precision and anatomical consistency. This work proposes the first fine-tuning-free, topology-driven framework that aligns sparse features with semantic labels via minimum spanning trees (MSTs). It assesses transferability at dual scales—local boundary token separability (LBTC) and global representation topological divergence (GRTD)—and incorporates a task-adaptive gated fusion mechanism. Theoretically, we prove that the MST leakage rate constitutes a finite-sample lower bound of the Bayes error and reveal that randomly initialized decoders stabilize topological alignment. Evaluated on a large-scale benchmark encompassing 114,000 3D medical images, our method achieves state-of-the-art performance, improving the weighted Kendall metric by 0.36 on average and accelerating evaluation by 56×.
This study addresses the challenge of effectively integrating global and local features in deep learning–based classification of benign and malignant pulmonary nodules, a task further complicated by the lack of clinical validation in existing approaches. To this end, we propose DeepFAN, a Transformer-based model trained on over 10,000 pathologically confirmed cases—the largest such dataset reported to date—and rigorously evaluated through a multicenter, multi-reader clinical trial. DeepFAN synergistically combines global and local imaging features and incorporates explainability analysis to assess feature contributions. The model achieves an internal test AUC of 0.939 and a clinical trial AUC of 0.954. When used as a diagnostic aid, it improves junior radiologists’ average AUC by 10.9%, with significant gains in sensitivity, specificity, and accuracy, and elevates inter-rater agreement from fair to moderate.
This work proposes an anatomy-informed synthetic supervised pre-training framework that addresses the limitations of existing methods, which rely on generic geometric shapes and fail to capture the morphological complexity, spatial layout, and inter-organ relationships inherent in real anatomical structures, thereby lacking the global structural priors essential for medical imaging. By integrating anatomical logic—such as spatial anchors and organ topology graphs—into the synthetic data generation process, the framework leverages a lightweight repository of realistic anatomical shapes and a structure-aware sequential placement strategy to enhance physiological plausibility. Evaluated on the Vision Transformer architecture, the method outperforms the current state-of-the-art FDSL baseline by 1.74% on BTCV and surpasses SSL approaches by 1.66% on MSD, while demonstrating robust scalability with increasing synthetic data volume.
This work addresses a critical limitation in existing synthetic data approaches for medical Vision Transformers (ViTs), which often neglect the intricate tissue textures and noise inherent in real medical images, leading to inaccurate learning of anatomical boundaries. To resolve this, the authors propose a physics-inspired spatially decoupled synthesis framework that, for the first time, identifies and mitigates an optimization conflict—termed boundary aliasing—between high-frequency texture modeling and boundary delineation. By introducing a gradient masking buffer and a core texture injection mechanism, the method enables orthogonal yet synergistic learning of shape and texture. Integrating equation-driven supervised learning, boundary distance fields, and physically informed spectral texture synthesis, the approach outperforms current FDSL and real-data-pretrained self-supervised methods by 1.43% and 1.51% on the BTCV and MSD datasets, respectively, establishing an efficient, annotation-free, and scalable training paradigm for medical ViTs.