Score
Designs, builds, and evaluates compact, discriminative representations derived from medical imaging by extracting, selecting, or learning quantitative radiomic features or deep imaging embeddings that summarize image variation relevant to a target outcome. Work covers region selection (e.g., segmentation or saliency-based candidate regions), extraction of handcrafted or deep features, aggregation and feature selection into a signature, validation with downstream classifiers, and interpretation of discriminative biomarkers (e.g., via feature‑importance methods).
This study addresses the core challenges in radiomics—feature instability, poor reproducibility, and limited clinical translation—by systematically evaluating how methodological choices across the end-to-end pipeline (including image acquisition, preprocessing, feature engineering, modeling, and evaluation) impact model robustness and generalizability. It is the first to comprehensively uncover the interdependencies among pipeline components and underscores the critical need for rigorous validation protocols to prevent data leakage and assessment bias. By integrating feature selection, dimensionality reduction, classical machine learning, and deep learning—and further exploring emerging paradigms such as hybrid AI, multimodal fusion, and federated learning—the work identifies key determinants of reliability and highlights persistent challenges related to standardization, domain shift, and clinical deployment, offering a systematic roadmap to enhance the quality and clinical applicability of radiomics research.
This study addresses the limited interpretability of current deep learning approaches in medical image–based tumor classification, which often fail to yield biologically meaningful imaging biomarkers. To overcome this, the authors propose an integrated framework that combines deep learning–based segmentation, Grad-CAM attention guidance, and radiomics analysis. Individualized quantitative imaging biomarkers are extracted from diagnostically relevant regions using mutual information–driven adaptive thresholding. These biomarkers are then validated and interpreted through conventional machine learning classifiers enhanced with SHAP (SHapley Additive exPlanations). Evaluated across multiple public and private tumor imaging datasets, the proposed method significantly outperforms whole-tumor radiomics baselines while maintaining high classification accuracy. Crucially, it enables the discovery of reproducible, interpretable, and biologically plausible imaging biomarkers, thereby bridging the gap between data-driven deep learning and clinically actionable insights.
This study addresses the limited generalization performance of lung cancer staging models caused by feature redundancy in high-dimensional, small-sample radiomics data. To this end, the authors propose a Gradient Loss-based Recursive Feature Elimination (GL-RFE) framework, which, for the first time, incorporates the gradient sensitivity of deep neural network loss functions with respect to input features into radiomics feature selection. By recursively eliminating low-contribution features, GL-RFE effectively captures nonlinear feature interactions, thereby enhancing both model interpretability and generalization. Starting from 106 CT radiomic features extracted via PyRadiomics, GL-RFE identifies a compact subset of 15 key features, achieving 90.22% accuracy and 90.16% F1-score on the test set—significantly outperforming conventional methods.
This study addresses the challenge of balancing interpretability and classification performance in radiomics analysis of knee MRI. We propose a patient-specific radiomic feature selection mechanism, integrated with a denoising diffusion model to construct individualized healthy image baselines—generated via masked inpainting to yield pathology-agnostic references—enabling lesion-guided localization and interpretable feature discovery. Furthermore, we introduce a feature importance re-weighted logistic regression within a multi-task framework (general abnormality, ACL tear, meniscal tear) to enhance discriminative performance. Experiments demonstrate that our method matches or surpasses state-of-the-art deep learning models across all three clinical tasks, while providing clinician-interpretable, feature-level explanations and personalized diagnostic evidence. To our knowledge, this is the first radiomics framework to enhance interpretability through a generative healthy baseline paradigm.
Radiomics workflow optimization has long relied on manual trial-and-error, lacking automated and standardized methodologies. To address this, we propose the first end-to-end AutoML framework for comprehensive radiomics workflow optimization, supporting modular workflow configuration, random-search-based hyperparameter tuning, model ensembling, and multi-center imaging preprocessing and feature extraction. This work represents the first systematic integration of AutoML across the entire radiomics pipeline—from image preprocessing to predictive modeling. Evaluated on 12 clinical prediction tasks spanning 12 diseases (including Alzheimer’s disease and liposarcoma), the framework achieves AUCs of 0.45–0.87, matching expert-designed pipelines and significantly outperforming baseline methods and Bayesian optimization. All data, code, and reproducibility protocols are fully open-sourced. The framework demonstrates strong cross-disease generalizability, substantially enhancing both biomarker discovery efficiency and scientific reproducibility.
This work proposes a patient-specific compact radiomic feature selection framework that addresses the limitations of conventional radiomics, which relies on population-level predefined features and struggles to balance individualized diagnostic performance with model interpretability. The approach employs a two-stage strategy: first generating diverse candidate feature sets via random sampling, then ranking them using a learnable scoring function to identify a complementary and non-redundant feature combination tailored to each patient. By moving beyond traditional top-k marginal selection, this method pioneers combinatorial optimization for patient-specific feature retrieval. Evaluated on ACL tear detection and osteoarthritis Kellgren–Lawrence grading tasks, it outperforms competing methods and matches the performance of end-to-end deep learning models while offering clinically interpretable decisions traceable to specific anatomical regions and feature types.
This study addresses the lack of systematic, disentangled comparisons between foundation models and radiomics in lung CT analysis, particularly regarding cross-cohort robustness. Through a two-stage design, the authors evaluate diverse combinations—including feature extractors (Curia, DINOv3, Radiomics), classifiers (TabPFN, XGBoost, CatBoost), and segmentation strategies (tumor vs. whole-lung)—across five clinical tasks, using worst-case cross-cohort performance as the primary metric. The work presents the first disentangled analysis of individual component contributions, revealing that segmentation predominantly influences tumor volume and staging tasks, while the classifier dominates survival, histology, and age prediction. Task-dependent design principles are proposed, with Curia combined with tumor segmentation and CatBoost recommended as a default configuration, demonstrating superior average performance across three core tasks.
This study addresses the lack of reliable uncertainty quantification in radiomics features derived from predicted segmentation masks, which often stems from model overconfidence. To this end, the authors propose ConRad, a novel framework that, for the first time, incorporates test-time segmentation boundary uncertainty into conformal prediction. ConRad constructs adaptive prediction intervals by jointly leveraging the input image, the predicted mask, and their geometric and appearance characteristics. Extensive experiments across five 2D medical imaging datasets and 171 radiomics features demonstrate that ConRad significantly outperforms existing baseline methods in interval efficiency while maintaining coverage close to the nominal level.
研究通过比较和融合放射组学与基础表示法(2D和3D MedVAE编码),提高了肾细胞癌亚型分类的准确性,3D门控融合方法表现最佳。
This study addresses the limitations of deep learning in medical image segmentation—namely, poor interpretability, excessive parameter count, and insufficient clinical trustworthiness—by proposing RadiomicNet, a lightweight dual-stream architecture that integrates handcrafted radiomic features. The method introduces a novel Radiomic Attention Gate (RAG) to inject Gray-Level Co-occurrence Matrix (GLCM) and Local Binary Pattern (LBP) features into the skip connections of a MobileNetV2 encoder-decoder framework, alongside a radiomic consistency loss to improve prediction calibration. With only 3.27 million parameters, RadiomicNet achieves Dice scores of 0.763 and 0.854 on the BUSI and Kvasir-SEG datasets, respectively, significantly outperforming U-KAN while providing inherent interpretability and explicit quantification of key feature contributions.