Score
Designing and applying morphological image-processing filters and heuristics (e.g., dilation/erosion, opening/closing, local tolerance fills) to compute clinically aligned features, remove spurious connections, and extract reproducible patient-specific imaging signatures.
Medical image tumor segmentation suffers from substantial inter-observer annotation variability, poor model interpretability, and challenges in clinical validation. To address these issues, this work pioneers the systematic integration of spatial logic–driven model checking techniques—such as Signal Temporal Logic (STL) and its spatial extension (SLCS)—into tumor semantic recognition, establishing a formally verifiable framework for automatic/semi-automatic region-of-interest identification. Methodologically, the approach unifies image-space modeling, formal verification toolchains, and standardized robustness evaluation protocols to provide mathematically provable guarantees on segmentation outputs—particularly regarding boundary fidelity and anatomical plausibility. Compared with conventional deep learning methods, our framework significantly improves consistency in tumor boundary delineation, effectively mitigates annotation bias across multi-center datasets, and simultaneously enhances clinical interpretability and deployment reliability—without compromising segmentation accuracy.
This study addresses the challenge of discontinuous vessel responses in retinal vascular segmentation, where traditional filters such as Frangi’s often produce fragmented vessels, thereby compromising accurate extraction of vascular structures across multimodal medical images. To overcome this limitation, the authors propose LS-CF, an unsupervised post-processing method that, for the first time, integrates local sensitivity–aware connectivity constraints with a heuristic tolerance mechanism to evaluate and reconnect broken vessel segments at the pixel level—without requiring any training data, thus enabling generalization across diverse imaging modalities. Built upon Frangi responses, LS-CF combines local connectivity analysis with morphological operations to form a lightweight yet effective unsupervised filter. The method achieves state-of-the-art performance on multiple benchmark datasets—including OSIRIX, IOSTAR, DRIVE, STARE, and CHASE_DB—with particularly notable gains over existing unsupervised approaches on CHASE_DB.
To address the limited interpretability and poor clinical integration of AI systems in fundus image analysis, this study proposes a “non-diagnostic” clinical assistance paradigm, implementing a modular fundus anatomical structure analysis system. Methodologically, it integrates a Transformer-based feature encoder with classical computer vision techniques—including vessel segmentation and geometric modeling of the optic cup and optic disc—to achieve structured, multi-granularity extraction of anatomical features, without generating diagnostic conclusions—only objective, verifiable measurements. Its key contributions are: (i) the first methodology for modular CV design explicitly aligned with ophthalmic clinical reasoning; and (ii) an interpretable, interoperable AI architecture conforming to real-world clinical workflows. Multi-center validation demonstrates significant improvements over existing screening models: +23.6% consistency in feature detection, 41% faster report generation, and 89.2% clinical adoption rate.
To address radiologists’ clinical need for interpretable and verifiable style adjustments in X-ray imaging, this paper proposes an explainable style transfer method specifically designed for mammograms. Methodologically, it replaces conventional handcrafted feature extractors with a trainable local Laplacian filter; models nonlinear style mapping via a multilayer perceptron (MLP); and incorporates a learnable adaptive normalization layer to explicitly decouple style semantics from physical meaning. This framework is the first in medical image style transfer to jointly satisfy semantic interpretability and reliability verifiability. Evaluated on a mammography dataset, it achieves an SSIM of 0.94—significantly surpassing the baseline (0.82)—while generating clinically preferred images that preserve radiologically meaningful physical interpretations.
In biomedical image segmentation validation, metrics such as the Hausdorff distance suffer from implementation inconsistencies across open-source toolkits, compromising benchmark reliability, introducing biomarker bias, and posing clinical deployment risks. To address this, we systematically evaluate 11 widely used toolkits and introduce, for the first time, a reference implementation based on high-fidelity 3D surface meshes. Our framework integrates real-world clinical data and a cross-platform consistency analysis. Statistical analysis reveals significant inter-tool variation in Hausdorff distance computations (p < 0.001), with interpolation strategy, boundary handling, and sampling density identified as primary sources of discrepancy. Based on these findings, we propose a reproducible and verifiable paradigm for distance-based evaluation, accompanied by standardized computational guidelines. This work substantially enhances the reliability, comparability, and clinical translatability of segmentation assessment.
Synthetic medical images often exhibit anatomical shape artifacts introduced by generative models, compromising model generalizability and clinical trustworthiness. Method: We propose a two-stage anomaly detection framework grounded in anatomical priors. First, we construct a boundary-angle gradient distribution feature space; then, we integrate an angle-gradient feature extractor with Isolation Forest to enable unsupervised, single-image-level detection of subtle morphological anomalies. Contribution/Results: Our core innovation lies in formalizing anatomical plausibility as a quantifiable geometric distribution constraint, explicitly embedded into the anomaly detection pipeline. Evaluated on two synthetic mammography datasets, the method achieves AUCs of 0.97 and 0.91. Visually, the most anomalous regions align strongly with human-annotated artifacts, and human–machine evaluation consistency significantly outperforms baseline methods. This framework enhances the anatomical fidelity and clinical applicability of synthetic medical data.
Clinical deployment of computer-aided diagnosis (CAD) systems is hindered by their reliance on hospital IT infrastructure integration, particularly with PACS and RIS. Method: This paper proposes a “zero-integration” radiology assistance framework that employs an external camera to capture medical images displayed on clinical monitors. The captured video frames undergo image restoration, object detection, and deep learning–based analysis to enable AI-assisted diagnosis and natural language report generation—bypassing PACS/RIS interfaces entirely and supporting plug-and-play deployment. Contribution/Results: Evaluated on multimodal medical imaging data, the framework achieves classification F1-scores within <2% and report-generation key metric gaps ≤1% relative to conventional CAD systems operating on native digital images. To our knowledge, this is the first work to systematically adopt visual-input paradigms for clinical imaging analysis, offering a lightweight, broadly deployable pathway for AI in healthcare.
Current medical imaging research predominantly relies on voxel-level masks, which are inadequate for capturing the structured semantics required in radiology reports. This work systematically defines, for the first time, three core components of medical image parsing—entities, attributes, and relationships—and introduces a unified framework that simultaneously supports decision-making, reconstruction, and prediction objectives. Grounded in structured semantic modeling, the approach integrates entity recognition, attribute description, and reasoning over spatial and temporal relationships, with an explicit emphasis on output consistency and closure. Evaluation across 11 representative systems reveals that existing methods largely neglect attributes, relationships, and closure properties, thereby underscoring the necessity and feasibility of the proposed paradigm in advancing models from mere measurement toward interpretable, explanatory understanding.
Existing CT image analysis methods are often confined to single tasks and lack the capacity for joint reasoning over anatomical structures and contextual information, limiting their ability to meet clinical demands for multi-granularity analysis. This work proposes a unified autoregressive vision-language framework that dynamically activates detection and segmentation heads via task-routing tokens and incorporates a “look closer” mechanism to enable coarse-to-fine region focusing. The model jointly generates visual outputs and interpretable textual descriptions in an end-to-end manner. It is the first large vision-language model to integrate task routing with a progressive look-closer strategy, unifying detection, segmentation, and appearance-based reasoning. The authors also introduce the first multimodal CT dataset annotated with pixel-level masks, bounding boxes, and structured textual descriptions. The method achieves Dice score improvements of 1.0% on BTCV and 1.7% on MosMed+, outperforming current state-of-the-art approaches.
This work addresses the limited clinical deployment of foundation models in medical image analysis, which stems from ambiguous adaptation mechanisms and insufficient consideration of robustness, calibration, and regulatory feasibility. The authors propose a strategy-centered framework that formally defines adaptation as a post-pretraining intervention, structuring it across five dimensions: parameters, representations, objectives, data, and architecture/sequence. Through a systematic review of adaptation techniques in classification, segmentation, and detection tasks, the study evaluates trade-offs in label efficiency, domain robustness, computational cost, and auditability, while aligning these considerations with validation protocols, multi-center deployment practices, and regulatory requirements. This framework offers practical guidance for developing clinically viable, robust, and auditable foundation model systems, thereby facilitating the translation of medical AI from laboratory research to real-world clinical settings.