What Is Missing in Surgical Risk Stratification and Outcome Prediction: A Scoping Review of End-to-End Machine Learning Approaches
Summary
A scoping review of 190 studies characterizing end-to-end machine learning pipelines for surgical risk stratification and outcome prediction using EHR data, identifying methodological gaps in preprocessing, evaluation, and explainability.
View Cached Full Text
Cached at: 08/03/26, 07:36 AM
# What Is Missing in Surgical Risk Stratification and Outcome Prediction: A Scoping Review of End-to-End Machine Learning Approaches Source: [https://arxiv.org/abs/2607.29090](https://arxiv.org/abs/2607.29090) [View PDF](https://arxiv.org/pdf/2607.29090) > Abstract:Postoperative adverse events, including mortality and morbidity, remain a major global burden, many of which are preventable through early identification of high\-risk patients and targeted perioperative care\. Accurate risk stratification is therefore essential\. With the growing availability of large\-scale electronic health records \(EHRs\), machine learning \(ML\) provides a data\-driven approach to model complex clinical patterns\. However, existing studies vary widely in design, and methodological practices remain fragmented\. This scoping review characterizes ML pipelines for surgical risk stratification and outcome prediction using EHR data\. We reviewed 190 studies covering the ML workflow, including data preprocessing, algorithm selection, model evaluation, and explainability\. Most studies relied on single\-center private datasets with limited data modalities, while the scarcity of open\-access surgical datasets constrained reproducibility and generalizability\. Reporting of key preprocessing steps, including missing data handling, feature selection, and class imbalance, was often incomplete\. Conventional ML models and simple neural networks predominated, whereas deep learning and multimodal approaches remained uncommon\. Benchmark datasets and standardized evaluation protocols were largely absent, hindering cross\-study comparisons\. Only about one\-third of studies incorporated explainability methods\. This review identifies methodological gaps limiting clinically robust postoperative ML tools and provides a structured reference to support more rigorous, reproducible, and clinically meaningful ML development for perioperative care\. ## Submission history From: Yizhi Dong \[[view email](https://arxiv.org/show-email/b3b24258/2607.29090)\] **\[v1\]**Fri, 31 Jul 2026 07:16:50 UTC \(1,171 KB\)
Similar Articles
Comprehensive Evaluation of Machine Learning for Type 2 Diabetes Risk Prediction: Large-Scale External Validation and Fairness Analysis
This paper develops a multi-dimensional framework to evaluate discrimination, calibration, interpretability, and algorithmic fairness for machine learning-based type 2 diabetes risk prediction models, revealing significant performance degradation under real-world distribution shifts and biases by age and obesity.
A Survey of Applications of ML in Healthcare
This blog post provides a high-level survey of machine learning applications in healthcare, covering medical imaging, wearables, and molecular biology. It highlights how ML can shift medicine from curative to preventative and improve hospital workflows without replacing healthcare workers.
Leveraging Physiological Signals to Predict Exam Outcomes with Machine Learning
This study investigates machine learning models to predict exam outcomes using physiological data such as electrodermal activity, heart rate, and skin temperature, finding that both deep learning approaches and simpler models like random forests can be effective.
Evaluating Reliability in Machine Learning Models for Early Chronic Kidney Disease Prediction: A Systematic Review of Data Leakage and Predictor Stability
This systematic review evaluates methodological reliability in machine learning models for early Chronic Kidney Disease prediction, revealing that data leakage inflates reported accuracy by over 15% and that more than 80% of predictors lack stability across studies.
LLMs for Cardiovascular Risk Prediction from Structured Clinical Data
This paper presents a hybrid framework that combines structured clinical data with LLM-generated narratives for coronary artery disease prediction, achieving high fidelity in variable extraction and comparing ML models with LLM-based zero-shot and few-shot classification.