A comparison of variable selection methods and predictive models for postoperative bowel surgery complications

📅 2025-07-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses key challenges in predicting postoperative gastrointestinal surgical complications—namely, high-dimensional variables, class imbalance, and poor model calibration. We systematically compare three modeling approaches—weighted logistic regression, hierarchical random forests, and naive Bayes—using 34 perioperative variables, evaluating performance via class-specific Brier scores and probability calibration. Methodologically, we introduce an integrated learning framework incorporating refined missing-data imputation and sample-weighting strategies to enhance model stability and calibration. Temporal validation demonstrates that the hierarchical random forest achieves superior performance for both “any complication” and “severe complication” prediction (optimal calibration, AUC = 0.82–0.85). These results confirm that clinically actionable, individualized risk stratification is feasible using single-center data, thereby providing a methodological foundation for dynamic perioperative monitoring and precision intervention.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationReasoning under Uncertainty: Relational Probabilistic ModelsKnowledge Representation and Reasoning: Computational Complexity of Reasoning

Application Category

Graph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods, algorithms and applications for the development of semantic models, knowledge graphs and other forms of structured data models with machine-interpretable semantics
📝 Abstract
Accurate prediction of postoperative complications can support personalized perioperative care. However, in surgical settings, data collection is often constrained, and identifying which variables to prioritize remains an open question. We analyzed 767 elective bowel surgeries performed under an Enhanced Recovery After Surgery protocol at Medisch Spectrum Twente (Netherlands) between March 2020 and December 2023. Although hundreds of variables were available, most had substantial missingness or near-constant values and were therefore excluded. After data preprocessing, 34 perioperative predictors were selected for further analysis. Surgeries from 2020 to 2022 ($n=580$) formed the development set, and 2023 cases ($n=187$) provided temporal validation. We modeled two binary endpoints: any and serious postoperative complications (Clavien Dindo $ge$ IIIa). We compared weighted logistic regression, stratified random forests, and Naive Bayes under class imbalance (serious complication rate $approx$11%; any complication rate $approx$35%). Probabilistic performance was assessed using class-specific Brier scores. We advocate reporting probabilistic risk estimates to guide monitoring based on uncertainty. Random forests yielded better calibration across outcomes. Variable selection modestly improved weighted logistic regression and Naive Bayes but had minimal effect on random forests. Despite single-center data, our findings underscore the value of careful preprocessing and ensemble methods in perioperative risk modeling.
Problem

Research questions and friction points this paper is trying to address.

Identifying key variables for predicting postoperative bowel surgery complications
Comparing predictive models under data constraints and class imbalance
Evaluating variable selection impact on model performance and calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Used weighted logistic regression for class imbalance
Applied stratified random forests for better calibration
Selected 34 perioperative predictors after preprocessing
🔎 Similar Papers
2024-04-17Annual International Conference of the IEEE Engineering in Medicine and Biology SocietyCitations: 0
Özge Şahin
Özge Şahin
Delft University of Technology
copulasstatistical learningrisk analysis
A
Annemiek Kwast
Value-Based Healthcare Group, Medisch Spectrum Twente, The Netherlands
Annemieke Witteveen
Annemieke Witteveen
Personalized eHealth Technology, University of Twente
patient level modellingcancerprediction modelsoptimizationmonitoring
T
Tina Nane
Delft Institute of Applied Mathematics, Delft University of Technology, The Netherlands