Original Article


Development and Internal Validation of an Interpretable Machine Learning Prognostic Prediction Model for Postoperative Delirium After Esophagectomy

Xiaochen Meng, Wandi Wu

Abstract

Background: Postoperative delirium (POD) is a frequent neurological complication associated with adverse outcomes after esophagectomy. Existing risk-assessment tools may provide limited individualized prognostic information. This study aimed to develop and internally validate an interpretable machine learning (ML) prognostic prediction model for POD after esophagectomy.

Methods: This single-center retrospective study included 1,089 patients who underwent esophagectomy for esophageal cancer at Shengjing Hospital of China Medical University between May 2022 and January 2025. Patients with pre-existing cognitive disorders or severe organ dysfunction were excluded. POD was assessed during the first 7 postoperative days or until discharge using the Confusion Assessment Method or the Confusion Assessment Method for the Intensive Care Unit. Before data preprocessing, patients were allocated to training and internal test cohorts using stratified random sampling at a 7:3 ratio. Missing data were handled using multiple imputation by chained equations. Candidate predictors were screened within the training cohort using Boruta and least absolute shrinkage and selection operator regression. Eight ML algorithms were optimized using grid search with stratified 5-fold cross-validation. Model discrimination, calibration, and clinical utility were evaluated in the internal test cohort, with 95% confidence intervals (CIs) estimated using 1,000 bootstrap resamples. Model interpretation was performed using SHapley Additive exPlanations (SHAP).

Results: The cohort had a median age of 70 years, and 75.9% of patients were male. POD occurred in 34.0% of patients, predominantly within the first three postoperative days. CatBoost achieved the numerically highest test-set area under the receiver operating characteristic curve (AUC) of 0.930 (95% CI: 0.900–0.959), with an accuracy of 86.5%, a recall of 0.862, a specificity of 0.866, an F1 score of 0.810, and a Brier score of 0.095. Calibration and decision curve analyses indicated satisfactory agreement and potential clinical utility. The final model included age, monocyte-to-lymphocyte ratio (MLR), C-reactive protein (CRP), preoperative polypharmacy, prognostic nutritional index (PNI), American Society of Anesthesiologists classification (ASA), surgical approach, and intraoperative blood transfusion. SHAP analysis identified MLR, age, and CRP as the most influential predictors.

Conclusion: An interpretable ML prognostic prediction model was developed and internally validated for estimating POD risk within seven days after esophagectomy. Although CatBoost demonstrated strong internal performance, prospective multicenter external validation is required before clinical implementation.

Download Citation