🤖 AI Summary
Current machine learning research in surgical risk prediction is often hindered by methodological fragmentation, poor reproducibility, and limited clinical applicability. This study conducts a scoping review of 190 end-to-end machine learning pipelines based on electronic health records, systematically examining critical components including data preprocessing, model selection, evaluation strategies, and interpretability. It presents the first structured synthesis of the entire workflow for surgical risk stratification, uncovering systemic gaps in the use of open datasets, standardized evaluation benchmarks, and deep learning methodologies. The majority of studies rely on single-center proprietary data, and only about one-third incorporate interpretability techniques—factors that severely constrain model generalizability and clinical translation. This work establishes a methodological framework and practical guidance for developing reproducible, generalizable, and clinically viable surgical prediction models.
📝 Abstract
Postoperative adverse events, including mortality and morbidity, remain a major global burden, many of which are preventable through early identification of high-risk patients and targeted perioperative care. Accurate risk stratification is therefore essential. With the growing availability of large-scale electronic health records (EHRs), machine learning (ML) provides a data-driven approach to model complex clinical patterns. However, existing studies vary widely in design, and methodological practices remain fragmented. This scoping review characterizes ML pipelines for surgical risk stratification and outcome prediction using EHR data. We reviewed 190 studies covering the ML workflow, including data preprocessing, algorithm selection, model evaluation, and explainability. Most studies relied on single-center private datasets with limited data modalities, while the scarcity of open-access surgical datasets constrained reproducibility and generalizability. Reporting of key preprocessing steps, including missing data handling, feature selection, and class imbalance, was often incomplete. Conventional ML models and simple neural networks predominated, whereas deep learning and multimodal approaches remained uncommon. Benchmark datasets and standardized evaluation protocols were largely absent, hindering cross-study comparisons. Only about one-third of studies incorporated explainability methods. This review identifies methodological gaps limiting clinically robust postoperative ML tools and provides a structured reference to support more rigorous, reproducible, and clinically meaningful ML development for perioperative care.