Machine learning models predict survival in unresectable stage III non-small cell lung cancer: Surveillance, Epidemiology, and End Results and Chinese cohort study
Original Article

Machine learning models predict survival in unresectable stage III non-small cell lung cancer: Surveillance, Epidemiology, and End Results and Chinese cohort study

Ye Zhang1,2#, Shiyu Hu2#, Jiaye Wang2,3#, Zhongjie Wu4#, Chengshui Chen5, Wenyu Chen2,5*, Liang Xie6*

1Department of Emergency Medicine, Lanxi People’s Hospital, Lanxi, China; 2Department of Respiratory Medicine, Affiliated Hospital of Jiaxing University, Jiaxing, China; 3Zhejiang Chinese Medical University, Jiaxing, China; 4Department of Cardio‑Thoracic Surgery, Affiliated Hospital of Jiaxing University, Jiaxing, China; 5Department of Respiratory Medicine, The First Affiliated Hospital of Wenzhou Medical University, Wenzhou, China; 6Department of Chronic Disease Control and Prevention, Jiaxing Center for Disease Control and Prevention, Jiaxing, China

Contributions: (I) Conception and design: Y Zhang, W Chen; (II) Administrative support: W Chen, Z Wu, C Chen; (III) Provision of study materials or patients: L Xie; (IV) Collection and assembly of data: Y Zhang, S Hu, J Wang; (V) Data analysis and interpretation: Y Zhang; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

#These authors contributed equally to this work as co-first authors.

*These authors contributed equally to this work as co-corresponding authors.

Correspondence to: Wenyu Chen, MD. Department of Respiratory Medicine, Affiliated Hospital of Jiaxing University, No. 1882, Zhonghuan South Road, Nanhu District, Jiaxing 314001, China; Department of Respiratory Medicine, The First Affiliated Hospital of Wenzhou Medical University, Wenzhou, China. Email: 00135116@zjxu.edu.cn; Liang Xie, MM. Department of Chronic Disease Control and Prevention, Jiaxing Center for Disease Control and Prevention, No. 486, Wenqiao Road, Jiaxing 314050, China. Email: xl4298@163.com.

Background: Patients with unresectable stage III non-small cell lung cancer (NSCLC) have heterogeneous survival outcomes, making accurate prognostic prediction challenging. This study aimed to develop and validate machine learning (ML) models for predicting overall survival (OS) after diagnosis in this patient population.

Methods: Data from 11,675 patients retrieved from the Surveillance, Epidemiology, and End Results (SEER) database were used for model development and internal testing. An independent external cohort (n=162) from the Affiliated Hospital of Jiaxing University was used for validation. We constructed models to predict 6-month, 1-year and 2-year OS using five ML algorithms, with model performance evaluated via the area under the receiver operating characteristic curve (AUC), accuracy and calibration. The optimal model was interpreted using SHapley Additive exPlanations (SHAP).

Results: The XGBoost model achieved the best performance for 6-month OS prediction in the test set (AUC =0.742, accuracy =0.708, sensitivity =0.746, specificity =0.637). It also showed predictive efficacy for 1-year (AUC =0.696) and 2-year (AUC =0.684) OS, with the highest discriminative ability at the 6-month time point. The model had good calibration and favorable net clinical benefit, while its performance decreased in external validation (AUC =0.647, accuracy =0.586). SHAP analysis revealed chemotherapy and radiotherapy as the most important predictive factors.

Conclusions: By integrating multiple clinical and treatment variables, the XGBoost model provides supplementary prognostic information and complements the conventional tumor-node-metastasis (TNM) staging system for 6-month, 1-year and 2-year OS prediction. This data-driven framework can assist individual risk stratification and improve the anatomical assessment of the TNM system.

Keywords: Non-small cell lung cancer (NSCLC); machine learning (ML); overall survival (OS); stage III; external validation


Submitted Apr 19, 2026. Accepted for publication Jun 18, 2026. Published online Jun 29, 2026.

doi: 10.21037/jtd-2026-1063


Highlight box

Key findings

• XGBoost achieves best 6-month overall survival prediction [area under the curve (AUC): 0.742 test, 0.647 external], and provides extra prognostic value compared with single tumor-node-metastasis (TNM) staging (AUC: 0.567, 0.542).

• SHapley Additive exPlanations (SHAP) identifies chemotherapy (0.577) and radiotherapy (0.315) as top predictors, highlighting treatment over staging.

What is known and what is new?

• TNM staging provides static anatomy without dynamic clinical variables, limiting accuracy in heterogeneous stage III non-small cell lung cancer.

• This interpretable XGBoost integrates 12 clinical factors, benchmarked vs. four machine learning (ML) models and TNM, with SHAP-based factor hierarchy. Unlike omics-dependent ML, our model uses routine Electronic Health Record (EHR) data and is externally validated in an independent Chinese cohort.

What is the implication, and what should change now?

• XGBoost complements TNM for refined risk stratification, guiding treatment intensity, follow‑up, and supportive care.

• SHAP interpretability builds up trust among clinicians, facilitating integration into clinical decision-support systems.

• Prospective multi-center validation is needed. If validated successfully, the model can be embedded into EHR for real-time prognostic estimates.


Introduction

Lung cancer (LC) is one of the most common malignancies globally, and the most common subtype is non-small cell lung cancer (NSCLC), accounting for approximately 85% of all lung cancers (1,2). Surgical resection is the preferred treatment for stage I and II LC, with a 5-year overall survival (OS) rate of over 80% postoperatively (3). However, about 30% of patients have been in stage III at the initial diagnosis, in which case the biological characteristics and growth rate of the tumor greatly vary among patients. Therefore, some patients are not suitable for surgical resection due to more extensive tumor invasion, more lymph node metastases, advanced age, and poor performance status (PS) (4-7). Unlike stage I and stage II NSCLC, for which surgical treatment yields consistently favorable clinical outcomes, and stage IV NSCLC, which is primarily managed with palliative care, unresectable stage III NSCLC represents a therapeutically challenging intermediate state. Patients in this group receive combined non-surgical regimens, while marked heterogeneity in survival persists even among individuals with identical tumor-node-metastasis (TNM) staging. Similar to advanced NSCLC, unresectable stage III NSCLC is primarily managed with multimodal therapy, including radiotherapy, chemotherapy, immunotherapy, and targeted therapy. These therapies can restrain tumor growth, but their impact on long-term survival remains unclear. Despite systemic treatment, the 5-year OS rate of these patients is statistically less than 20% (2,8). Therefore, it is a great challenge to predict the treatment outcome of these patients and to develop individualized treatment strategies aiming to improve the survival rate.

Nowadays, the TNM staging system, the most commonly used clinical tool for prognostic prediction in LC (9), contributes to tumor staging, which, however, has certain subjectivity and hysteresis, restricting the efficiency of prognostic prediction (10). To compensate for these deficiencies, establishing clinical prediction models for LC prognosis has gradually become a research hotspot recently, which, usually combining multiple predictors, has achieved dynamic prediction of prognosis (11,12). Due to the limitations of simple statistical methods and linear hypothesis, it is often difficult for traditional LC prediction models to simulate the complex and diverse development of LC; long-term training and validation are required when handling large-scale and complex clinical data; the model operation is overly dependent on specific dataset characteristics, leading to deficiencies in accuracy, efficiency, and generalization ability.

In recent years, artificial intelligence technology has been developed tremendously. Using advanced algorithms, machine learning (ML) as a branch of artificial intelligence can automatically learn and extract knowledge from mass data, greatly enhancing the accuracy and efficiency of decision-making. Therefore, ML has been increasingly applied in disease diagnosis and prognosis (13,14). ML-based prognostic prediction models can efficiently process complex datasets including radiomics, genomics, proteomics, and clinical data, and possess stable operation in the face of data diversity and uncertainty, exhibiting excellent generalization ability (15-21). Despite these strengths, the “black-box” nature of many ML models often restricts insight into the relative contributions of individual predictors, which is critical for informed clinical decision-making. Moreover, the combined predictive value of specific clinical factors within an interpretable ML framework remains underexplored. Therefore, to bridge this gap and enhance clinical applicability of ML, this study: (I) investigated the association between multiple clinical factors and OS using ML methods; (II) established and validated a prognostic model using classical ML algorithms; (III) quantitatively assessed and ranked the relative contributions of key factors, including treatment methods, within an interpretable framework, thereby advancing individual risk stratification in unresectable stage III NSCLC. We present this article in accordance with the TRIPOD reporting checklist (22) (available at https://jtd.amegroups.com/article/view/10.21037/jtd-2026-1063/rc).


Methods

Patients and selection criteria

The data were sourced from the Surveillance, Epidemiology, and End Results (SEER) Research Plus Data, 8 Registries, Nov 2021 Sub database, which covers the years 1975 to 2019. From this database, we extracted cases diagnosed between 2004 and 2015 to form our study cohort. Additional ethical approval was not required for the use of de-identified SEER data, and the SEER data use agreement was adhered to. Inclusion criteria: (I) patients pathologically diagnosed with primary NSCLC; (II) patients with stage III NSCLC; (III) patients aged ≥15 years at the time of diagnosis; (IV) patients with survival follow-up data. Exclusion criteria: (I) patients with stage I, II, or IV NSCLC; (II) patients complicated with other malignancies; (III) patients with a history of lung surgery; (IV) patients with incomplete data on survival time, metastasis, and clinical stage; (V) patients with survival time <1 month. The outcomes predicted by the model were 6-month, 1-year, and 2-year OS, i.e., patients were still alive at the 6-month, 1-year, and 2-year follow-up, respectively. The flowchart of study design and analysis is shown in Figure 1. Additionally, to verify the generalizability of the results, we conducted a retrospective study of unresectable stage III NSCLC patients hospitalized at the Affiliated Hospital of Jiaxing University between January 2015 and December 2023, with a follow-up period until July 20, 2024. Data were collected through telephone interviews or review of medical records. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. This study was approved by the Ethics Committee of Affiliated Hospital of Jiaxing University (No. 2023-KY-509). Given the retrospective nature of the study and the use of de-identified patient data, the requirement for informed consent was waived by our Ethics Committee.

Figure 1 The flowchart described the process of conducting the study and statistical analysis. NSCLC, non-small cell lung cancer; ROC, receiver operating characteristic; SEER, Surveillance, Epidemiology, and End Results.

Study variables

The following data were collected: patient characteristics (age, sex, race, marital), tumor characteristics (primary, histology, grade, laterality, summary, T, N, stage, size), treatment (radiotherapy, chemotherapy), and follow-up information (vital status, survival time). Some continuous variables (e.g., age and size) were transformed into categorical variables in the analysis. According to the age at the diagnosis, the patients were split into four groups (15–59, 60–69, 70–79, and ≥80 years). The tumor size was assigned into six groups (≤20, 21–30, 31–50, 51–70, >70 mm, and other) based on the 7th edition of American Joint Committee on Cancer (AJCC) staging system. We treated the missing values for categorical variables as “unknown”, which ensured data integrity and avoided loss of information due to missing data.

Model construction and evaluation

First, 11,675 patients were randomized 7:3 into training and test cohorts. In the training cohort, variables significantly associated with LC-specific survival were screened by the univariate Cox regression analysis. Univariate Cox regression (P<0.05) was used to select prognostic variables to reduce redundant features, and then those with a P<0.05 were further incorporated into the multivariate Cox regression model, with hazard ratios (HR) and its 95% confidence intervals (CI) calculated. The results were visualized by forest plots. During model establishment, R4.4.1 and five ML algorithms (LR, SVM, RF, GBM, and XGBoost) were adopted. All five ML models were constructed based on the training dataset, and hyperparameters were set according to empirical values from published literature. No cross-validation was performed in this study.

The predictive performance of the models was evaluated on the test set using the area under the receiver operating characteristic curve (AUC), accuracy, sensitivity, and specificity, with AUC serving as the primary metric. The model demonstrating the highest predictive performance was selected as the optimal model for this study. A calibration curve was subsequently plotted to assess the agreement between observed outcomes and predicted probabilities. In addition, decision curve analysis (DCA) was conducted to evaluate the clinical utility of the model. Finally, external validation was performed using data from an independent Chinese cohort to assess the model’s generalizability and applicability. To contextualize the performance of ML models under the current clinical standard, the prognostic prediction capability of the 7th edition of the AJCC TNM staging system as a baseline model was also evaluated. Missing values labelled as “unknown” follow the original SEER coding and were retained without manual modification to preserve data integrity. The TNM stage for each patient was used as a single predictor, and its predictive capability for 6-month, 1-year, and 2-year OS was also assessed using the AUC in the test and external validation cohorts.

Model interpretation

Model interpretability is crucial in clinical settings, as it helps clinicians understand the decision-making process, thereby building trust and facilitating practical application. To interpret the predictive models, feature importance was visualized for the XGBoost model and interpreted using SHapley Additive exPlanations (SHAP) values. Feature importance, derived from the XGBoost model’s intrinsic mechanism, quantifies the overall contribution of each variable to the predictions. Meanwhile, the SHAP framework provides a unified measure of feature impact, ensuring consistent and locally accurate explanations for each prediction. We applied SHAP to analyze the prognostic model for unresectable stage III NSCLC, enhancing its transparency to guide appropriate clinical interventions when necessary.

Statistical analysis

All statistical analyses were performed using R version 4.4.1. Continuous variables with a normal distribution were expressed as mean ± standard deviation (SD), while non-normally distributed variables were presented as median with interquartile range (IQR). Categorical variables were summarized as frequencies and percentages. The specific methodologies for variable selection, machine learning model construction, hyperparameter configuration, internal and external validation, performance evaluation (including AUC, accuracy, calibration curves, and decision curve analysis), and model interpretability via SHAP are detailed in the preceding subsections (Model construction and evaluation, and Model interpretation). A two-sided P<0.05 was considered statistically significant throughout the study.


Results

Demographic composition and baseline clinical information

A total of 11,675 (prediction models) and 162 patients with unresectable stage III NSCLC (an external validation cohort) were finally screened from SEER [2004–2015] and the Affiliated Hospital of Jiaxing University [2015–2023], respectively. The patients from SEER were randomized 7:3 into a training cohort (n=8,172) and a test cohort (n=3,503) (Table 1). There were 4,348 (53.2%) males and 3,824 (46.8%) females in the training cohort, 1,903 (54.3%) males and 1,600 (45.7%) females in the test cohort, and 133 (82.1%) males and 29 (17.9%) females in the external validation cohort. LC mostly occurred in the upper lobe: 4,407 (53.9%) patients in the training cohort, 1,781 (50.8%) patients in the test cohort, and 80 (49.4%) patients in the external validation cohort. In the external validation cohort, the median age of patients was 68 years (IQR: 62–74 years), with 77.8% of patients aged 60–79 years. A total of 71.6% (n=116) of patients underwent chemotherapy, and 40.1% (n=65) underwent radiotherapy. Complete follow-up data, including vital signs and survival time, were obtained from all patients.

Table 1

Clinicopathological characteristics of patients with lung cancer in the training cohort, test cohort and validation cohort

Characteristic Training cohort (n=8,172) Test cohort (n=3,503) Validation cohort (n=162)
Age, years
   15–59 1,480 (18.1) 668 (19.1) 26 (16.0)
   60–69 2,258 (27.6) 968 (27.6) 64 (39.5)
   70–79 2,566 (31.4) 1,082 (30.9) 62 (38.3)
   ≥80 1,868 (22.9) 785 (22.4) 10 (6.2)
Sex
   Male 4,348 (53.2) 1,903 (54.3) 133 (82.1)
   Female 3,824 (46.8) 1,600 (45.7) 29 (17.9)
Race
   Black 780 (9.5) 370 (10.6) 0 (0.0)
   White 6,333 (77.5) 2,703 (77.2) 0 (0.0)
   Other 1,059 (13.0) 430 (12.3) 162 (100.0)
Primary
   Lower 1,883 (23.0) 876 (25.0) 61 (37.7)
   Upper 4,407 (53.9) 1,781 (50.8) 80 (49.4)
   Other 1,882 (23.0) 846 (24.2) 21 (13.0)
Marital
   Single 1,085 (13.3) 495 (14.1) 3 (1.9)
   Divorced 3,071 (37.6) 1,282 (36.6) 4 (2.5)
   Married 4,016 (49.1) 1,726 (49.3) 155 (95.7)
Histologic
   Squamous 2,538 (31.1) 1,069 (30.5) 98 (60.5)
   Large 2,551 (31.2) 1,066 (30.4) 14 (8.6)
   Adenocarcinoma 3,083 (37.7) 1,368 (39.1) 50 (30.9)
Grade
   I–II 1,200 (14.7) 508 (14.5) 7 (4.3)
   III 1,804 (22.1) 790 (22.6) 15 (9.3)
   IV 438 (5.4) 207 (5.9) 3 (1.9)
   Unknown 4,730 (57.9) 1,998 (57.0) 137 (84.6)
Laterality
   Left 3,274 (40.1) 1,413 (40.3) 58 (35.8)
   Right 4,776 (58.4) 2,018 (57.6) 103 (63.6)
   Unknown 122 (1.5) 72 (2.06) 1 (0.62)
Summary
   Distant 3,974 (48.6) 1,710 (48.8) 57 (35.2)
   Regional 4,193 (51.3) 1,790 (51.1) 100 (61.7)
   Unknown 5 (0.1) 3 (0.1) 5 (3.0)
T
   T0–T1 710 (8.7) 270 (7.7) 20 (12.3)
   T2 1,885 (23.1) 809 (23.1) 45 (27.8)
   T3 605 (7.4) 236 (6.7) 39 (24.1)
   T4 4,972 (60.8) 2,188 (62.5) 58 (35.8)
N
   N0 1,591 (19.5) 691 (19.7) 12 (7.4)
   N1 490 (6.0) 218 (6.2) 9 (5.6)
   N2 4,742 (58.0) 2,031 (58.0) 106 (65.4)
   N3 1,349 (16.5) 563 (16.1) 35 (21.6)
Stage
   IIIA 2,512 (30.7) 1,033 (29.5) 80 (49.4)
   IIIB 5,660 (69.3) 2,470 (70.5) 82 (50.6)
Radiation
   No 3,634 (44.5) 1,572 (44.9) 97 (59.9)
   Yes 4,538 (55.5) 1,931 (55.1) 65 (40.1)
Chemotherapy
   No 3,101 (37.9) 1,295 (37.0) 46 (28.4)
   Yes 5,071 (62.1) 2,208 (63.0) 116 (71.6)
Size
   ≤20 mm 715 (8.8) 293 (8.36) 12 (7.4)
   21–30 mm 982 (12.0) 389 (11.1) 22 (13.6)
   31–50 mm 1,951 (23.9) 847 (24.2) 39 (24.1)
   51–70 mm 1,423 (17.4) 623 (17.8) 36 (22.2)
   >70 mm 1,370 (16.8) 628 (17.9) 15 (9.3)
   Other 1,731 (21.2) 723 (20.6) 38 (23.5)

Data are presented as n (%). N, node; T, tumor.

Independent prognostic factors in the training cohort

Univariate and multivariate Cox regression analyses identified twelve variables as significant predictors of survival (Figure 2): age [HR =1.56 (95% CI: 1.44–1.69)], sex [HR =0.83 (95% CI: 0.79–0.86)], race [HR =0.78 (95% CI: 0.70–0.86)], marital [HR =0.87 (95% CI: 0.81–0.93)], histology [HR =0.87 (95% CI: 0.82–0.92)], grade [HR =1.23 (95% CI: 1.09–1.40)], summary [HR =0.75 (95% CI: 0.70–0.79)], T [HR =1.15 (95% CI: 1.02–1.29)], stage [HR =0.84 (95% CI: 0.75–0.93)], radiotherapy [HR =0.71 (95% CI: 0.68–0.75)], chemotherapy [HR =0.60 (95% CI: 0.57–0.64)], and size [HR =1.70 (95% CI: 1.53–1.88)]. These variables were used to establish five ML models (LR, SVM, RF, GBM, and XGBoost) for subsequent validation.

Figure 2 Multivariate Cox regression analysis and Forest plot based on variables for overall survival (training cohort). CI, confidence interval; HR, hazard ratio; T, tumor.

Model performance

As revealed by the receiver operating characteristic (ROC) curves, calibration curves, and decision curves, the ML model was most effective in predicting 6-month OS. In the training cohort, the XGBoost model achieved the highest AUC (0.782), followed by RF (0.773), GBM (0.764), SVM (0.747), and LR (0.745) (Figure 3). The same result was obtained in the independent test cohort: XGBoost yielded the highest AUC (0.742), followed by GBM (0.741), LR (0.734), RF (0.732), and SVM (0.725). Detailed performance metrics for all models at 6 months are provided in Table 2. For 6-month OS prediction, XGBoost also exhibited optimal accuracy (0.714), sensitivity (0.712), and specificity (0.715) (Figure 4). Although XGBoost remained the optimal algorithm for predicting 1- and 2-year OS, its discriminative power (AUC) declined over time (training cohort: 0.754 and 0.755; Figures S1,S2). Moreover, the DCA revealed that the XGBoost model demonstrated good calibration (Figure 5) and provided net clinical benefit across a wide range of threshold probabilities (0.15–0.80) (Figure 6).

Figure 3 ROC curves and AUCs at 6 months. (A) ROC curves of different ML models in the training cohort. (B) ROC curves of different ML models in the test cohort. (C) ROC curves of different ML models in the validation cohort. AUC, area under the curve; ML, machine learning; ROC, receiver operating characteristic.

Table 2

Performance metrics of the 6-month prognostic models

Model/dataset AUC Accuracy Specificity Sensitivity
LR
   Training cohort 0.745 0.679 0.657 0.719
   Test cohort 0.734 0.706 0.748 0.630
   Validation cohort 0.660 0.512 0.457 0.864
SVM
   Training cohort 0.747 0.684 0.660 0.728
   Test cohort 0.725 0.685 0.705 0.648
   Validation cohort 0.617 0.605 0.593 0.682
RF
   Training cohort 0.773 0.727 0.762 0.662
   Test cohort 0.732 0.698 0.728 0.643
   Validation cohort 0.658 0.586 0.550 0.818
GBM
   Training cohort 0.764 0.717 0.745 0.666
   Test cohort 0.741 0.696 0.711 0.667
   Validation cohort 0.644 0.679 0.686 0.636
XGBoost
   Training cohort 0.782 0.714 0.715 0.712
   Test cohort 0.742 0.708 0.746 0.637
   Validation cohort 0.647 0.586 0.557 0.773

AUC, area under the curve.

Figure 4 Performance of predictive models at 6 months. (A) The performance of different ML models in terms of accuracy, sensitivity, specificity and precision in the training cohort. (B) The performance of different ML models in terms of accuracy, sensitivity, specificity and precision in the test cohort. (C) The performance of different ML models in terms of accuracy, sensitivity, specificity and precision in the validation cohort. ML, machine learning.
Figure 5 Calibration curves at 6 months. (A) Calibration curves of different ML models in the training cohort. (B) Calibration curves of different ML models in the test cohort. (C) Calibration curves of different ML models in the validation cohort. ML, machine learning.
Figure 6 DCA curves at 6 months. (A) DCA curves of different ML models in the training cohort. (B) DCA curves of different ML models in the test cohort. (C) DCA curves of different ML models in the validation cohort. DCA, decision curve analysis; ML, machine learning.

Given that the predictive performance was most robust for 6-month OS prediction, detailed results, including ROC curves, performance metrics, calibration curves, and DCA, are presented in the main text. Extended results for 1- and 2-year OS predictions are provided in Figures S1-S8. All figures have been rendered in high resolution to ensure clarity.

Model interpretation and prognostic relevance

We performed SHAP analysis on the XGBoost model for 6-month prediction (Figure 7). The SHAP summary plot illustrated both the magnitude and direction of each feature’s contribution to the model output. The SHAP values provided a quantitative ranking of feature importance. Chemotherapy recorded a SHAP value of 0.577, and radiotherapy recorded a SHAP value of 0.315. Tumor size was the third-ranked feature (SHAP =0.244), followed by summary (0.184), age (0.177), and sex (0.166). The mean SHAP values for T descriptor (0.093), histology (0.086) and stage (0.022) were also listed.

Figure 7 SHAP analysis summary diagram based on XGBoost. Each row represents a feature; a point represents a sample; purple represents a high feature value; and yellow represents a low feature value. A further distance from a point to the baseline SHAP value of 0 indicates a greater impact on the output. SHAP, SHapley Additive exPlanations; T, tumor.

External validation

We collected data from a cohort of patients with unresectable stage III NSCLC through telephone follow-up and review of medical records. Due to challenges in patient follow-up and resource constraints, the size of the external validation dataset was limited to 162 cases. In this independent cohort, the discriminative performance of the XGBoost model declined across all time points compared to the test set. Specifically, the AUC values were 0.647 for 6-month, 0.608 for 1-year, and 0.645 for 2-year OS prediction.

Comparative performance against TNM staging system

We compared the discriminative power of the XGBoost model with the conventional TNM staging system. As illustrated in Figure 8 and Table 3, the AUC values of the two models in different cohorts and time points are listed as follows. For 6-month OS prediction, the XGBoost model yielded AUCs of 0.782, 0.742, and 0.647 in the training, test, and validation cohorts, respectively; the corresponding AUCs of the TNM staging system were 0.569, 0.567, and 0.542. For 1-year OS prediction, the AUC values were 0.754, 0.696, and 0.608 for XGBoost, versus 0.568, 0.544, and 0.541 for the TNM staging system. For 2-year OS prediction, the AUC values were 0.755,0.684, and 0.645 for XGBoost, versus 0.563,0.550, and 0.606 for the TNM staging system. The difference in AUC between XGBoost and the TNM staging system was 0.213 in the training cohort for 6-month OS prediction.

Figure 8 Comparative performance of the XGBoost model versus the TNM staging system across multiple time horizons and cohorts. Comparison of the area under the receiver operating characteristic curve between the interpretable XGBoost model and the conventional TNM staging system for predicting overall survival. (A-C) 6-month survival prediction in the training, test, and validation cohorts. (D-F) and (G-I) show corresponding comparisons for 1-year and 2-year survival predictions, respectively. AUC, area under the curve; CI, confidence interval; TNM, tumor-node-metastasis.

Table 3

Comparison of XGBoost and TNM staging for predicting OS

Predicted time point/dataset XGBoost, AUC (95% CI) TNM, AUC (95% CI)
6 months
   Training cohort 0.782 (0.771–0.792) 0.569 (0.557–0.582)
   Test cohort 0.742 (0.725–0.760) 0.567 (0.548–0.586)
   Validation cohort 0.647 (0.526–0.769) 0.542 (0.405–0.678)
1 year
   Training cohort 0.754 (0.744–0.765) 0.568 (0.556–0.580)
   Test cohort 0.696 (0.679–0.713) 0.544 (0.525–0.563)
   Validation cohort 0.608 (0.514–0.702) 0.541 (0.446–0.635)
2 years
   Training cohort 0.755 (0.743–0.767) 0.563 (0.549–0.578)
   Test cohort 0.684 (0.664–0.703) 0.550 (0.528–0.573)
   Validation cohort 0.645 (0.560–0.730) 0.606 (0.519–0.692)

AUC, area under the curve; CI, confidence interval; OS, overall survival; TNM, tumor-node-metastasis.


Discussion

Accurate prognostic prediction is essential for personalized treatment and improved outcomes in unresectable stage III NSCLC (23). ML, which can integrate and analyze complex clinical data, has emerged as a pivotal tool for prognostic prediction in oncology (24). In this study, five ML models were established and validated using a large SEER cohort to achieve accurate OS prediction in this patient population.

Key influencing factors for LC prognosis

Twelve variables were identified as predictors, with chemotherapy and radiotherapy as the strongest predictors, highlighting the decisive role of active treatment. Radiotherapy works by delivering high-energy rays to damage tumor DNA, thereby inhibiting tumor growth and cell division, while chemotherapy disrupts tumor proliferation and dissemination by pharmacological agents. However, toxicities associated with these treatments, such as radiation pneumonitis and chemotherapy-induced nausea and vomiting, may affect patients’ quality of life and treatment adherence, thus influencing prognosis. Tumor characteristics, particularly size and summary, were also potent predictors, although the impact of size may vary by subtype (e.g., solid vs. ground-glass opacities) (25). In addition, patient-related factors, including age, sex, race, and marital status, further contributed to prognostic variation. Arnold et al. (26) reported that NSCLC patients under 50 years of age have a 59% higher incidence of actionable genetic alterations than older patients, rendering them more likely to benefit from targeted therapies. Moreover, younger patients generally exhibit better PS and fewer comorbidities, enabling them to tolerate more intensive chemoradiotherapy (27). Sex influences prognosis, and women are more prone to adenocarcinoma, frequent EGFR mutation, and potentially better responses to targeted therapies (4,28,29).

XGBoost model for survival prediction of LC

XGBoost is a gradient-boosting algorithm creating sequential decision trees to minimize errors, incorporating regularization to prevent overfitting, and efficiently handling large-scale data (30). Its efficacy has been validated in other contexts of NSCLC, such as early screening and bone metastasis prediction. Ahmed et al. (31) compared XGBoost, K-Nearest Neighbor (KNN), and SVM in early LC screening, and found that XGBoost performs 100% perfectly in the precision, recall, and F1-score. In a similar study, an early warning model for LC was constructed by XGBoost plus metabolomics, and its AUC, accuracy, and sensitivity were 0.81, 75.29%, and 74%, respectively, suggesting that XGBoost combined with other omics can also achieve an excellent predictive capacity (32). Moreover, the XGBoost model outperforms other models in both the training and validation cohorts in the survival prediction of patients with NSCLC bone metastases (33). Roughly consistent with the above studies, the powerful ability of XGBoost to process large-scale complex data was demonstrated in this study. Although careful parameter tuning is required to mitigate overfitting risks, its performance in this study underscores its utility for complex prognostic prediction (34).

Superiority over conventional TNM staging system and clinical implications

A key finding of this study is the demonstrable superiority of the interpretable XGBoost model over the established TNM staging system in predicting survival in unresectable stage III NSCLC (Table 3). The TNM staging system, a foundation for anatomical tumor classification and initial treatment planning, provides a static snapshot that does not incorporate dynamic clinical variables such as treatment modalities (chemotherapy, radiotherapy), patient demographics (age, marital status), or tumor characteristics (size, grade). This inherent limitation likely accounts for its consistently lower discriminative power (an AUC of 0.542–0.606 in the validation cohort) compared with our ML model. Our ML model synthesized these heterogeneous factors into personalized risk estimates, aligning with the multifactorial nature of prognosis in precision oncology (35). Thus, it serves not as a replacement but as a complementary, data-driven tool for refined risk stratification in stage III, potentially guiding follow-up intensity and supportive care.

Comparison with state-of-the-art prognostic models

This study adds to the growing body of literature applying ML to NSCLC prognostic prediction. For instance, Zhong et al. developed a stepwise regression model based on CT radiomics to predict survival in advanced NSCLC, achieving an AUC of 0.849 (95% CI: 0.812–0.885) in the training set (36). Despite high accuracy, such models require specialized imaging preprocessing and may not be readily applicable in all settings. Another state-of-the-art model by Liu et al. incorporated genomic markers alongside clinical data to predict the clinical outcomes of atezolizumab therapy for NSCLC (37). Their model was powerful but relied on costly sequencing techniques, restricting its broad clinical utility. In contrast, our ML model relies solely on structured clinical and treatment variables, enhancing its practicality for rapid integration into electronic health records. Thus, this study provided a pragmatic and interpretable tool specifically validated for the understudied but critical cohort of patients with unresectable stage III NSCLC.

Strengths and limitations

Unlike many prior studies, five ML models were systematically compared on a large, population-based cohort. This comparative approach identified XGBoost as the optimal model across multiple survival endpoints and reinforced the methodological rigor of our conclusions. Furthermore, the SHAP framework enhanced interpretability, transforming the model from a “black box” into a transparent tool. SHAP enabled data-driven quantification of prognostic factors, clearly demonstrating, for example, that chemotherapy exerts a significantly greater influence than TNM staging in this population.

Several limitations must be acknowledged. First, the definition of unresectable disease in this study was indirectly inferred from stage classification and surgical records in the SEER database. Since multidisciplinary assessment results for resectability are not recorded in the database, this indirect definition may introduce potential bias. Second, this retrospective registry-based study is subject to treatment selection bias and unmeasured confounding factors including PS, comorbidities, patient preference, physician decisions and healthcare accessibility. Patients who received chemoradiotherapy generally had more favorable baseline characteristics. Third, incomplete data, particularly the high proportion of patients with “unknown” tumor grade, may reduce model precision. Fourth, owing to inherent limitations of the SEER registry, several modern prognostic variables—including smoking history, PS, immunotherapy-related indicators and molecular markers—were not available for model development, which limits the direct application of our model to current treatment settings. Inclusion of these factors would substantially alter feature importance identified via SHAP analysis: immunotherapy may become a predominant prognostic factor superior to chemotherapy, and programmed death ligand-1 (PD-L1) status would also represent an important prognostic marker. Therefore, the present model is suitable only for the pre-immunotherapy era and needs further optimization before clinical adoption. Fifth, our ML model had the strongest predictive performance for 6-month OS, but the AUC declined for 1-year and 2-year OS predictions, possibly due to more complete short-term follow-up and increased censoring over time. Nevertheless, accurate short-term predictions retain the model’s clinical utility for initial risk stratification. Sixth, external validation was conducted on a relatively small single-center cohort (n=162) from China. There were significant differences between the external cohort and the large-scale SEER cohort in terms of gender distribution, race, histological type, marital status, treatment patterns and the proportion of unknown tumor grade. These discrepancies, together with the small sample size, are the main reasons for the notable drop in AUC during external validation and poor model generalizability. Large-sample, multi-center cohorts with similar population characteristics are required for subsequent validation. Finally, as the ML model relied on structured clinical data, its performance could be further enhanced in the future by integrating radiomic or genomic features within the same interpretable framework.

Future directions

To address the limitations noted above, particularly concerning generalizability and clinical relevance, future research should pursue several strategic pathways. First, data-sharing consortia with institutions across diverse geographic and healthcare settings should be established to overcome the barrier of single-center validation. The adoption of common data models, such as the Observational Medical Outcomes Partnership Common Data Model, would be crucial to the harmonization of heterogeneous retrospective datasets for large-scale validation. Second, the model’s generalizability and real-world utility must be assessed through a prospective, multi-center study with standardized data collection protocols. Concurrently, future iterations of this interpretable framework should seek to integrate the omitted variables (e.g., smoking history, molecular markers) and explore the inclusion of multi-modal data, such as radiomic features from medical imaging. These concerted efforts are important for rigorously evaluating the model’s clinical effectiveness, mitigating biases, and paving the way for its potential integration into decision-support systems before personalized management in unresectable stage III NSCLC.


Conclusions

In conclusion, this study confirmed the critical prognostic value of chemotherapy, radiotherapy, tumor size, summary, age, and sex for OS in patients with unresectable stage III NSCLC. The comparative analysis further revealed that the XGBoost algorithm can complement conventional TNM staging systems and provide richer prognostic information by effectively integrating complex clinical variables. This model has preliminary predictive value and may have potential clinical utility after large-scale multi-center validation and further optimization.


Acknowledgments

The authors thank all patients and institutions involved in this study, especially the ability to have open access to the SEER database.


Footnote

Reporting Checklist: The authors have completed the TRIPOD reporting checklist. Available at https://jtd.amegroups.com/article/view/10.21037/jtd-2026-1063/rc

Data Sharing Statement: Available at https://jtd.amegroups.com/article/view/10.21037/jtd-2026-1063/dss

Peer Review File: Available at https://jtd.amegroups.com/article/view/10.21037/jtd-2026-1063/prf

Funding: This work was supported by the Key Construction Disciplines of Provincial and Municipal Co construction of Zhejiang (No. 2023-SSGJ-002), National Oncology Clinical Key Speciality (No. 2023-GJZK-001), Scientific Technology Plan Program for Healthcare in Zhejiang Province (Nos. 2023KY329, 2023KY1196, and 2023RC098), Zhejiang Province Postdoctoral Research Project (No. ZJ2024156).

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://jtd.amegroups.com/article/view/10.21037/jtd-2026-1063/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. The study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. This study was approved by the Ethics Committee of Affiliated Hospital of Jiaxing University (No. 2023-KY-509). Given the retrospective nature of the study and the use of de-identified patient data, the requirement for informed consent was waived by our Ethics Committee.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Alborelli I, Leonards K, Rothschild SI, et al. Tumor mutational burden assessed by targeted NGS predicts clinical benefit from immune checkpoint inhibitors in non-small cell lung cancer. J Pathol 2020;250:19-29. [Crossref] [PubMed]
  2. Herbst RS, Morgensztern D, Boshoff C. The biology and management of non-small cell lung cancer. Nature 2018;553:446-54. [Crossref] [PubMed]
  3. Ohtaki Y, Aokage K, Miyoshi T, et al. Function-preserving radical surgery for early-stage non-small cell lung cancer: A review of limited resection approaches. Jpn J Clin Oncol 2026;56:385-92. [Crossref] [PubMed]
  4. Miao D, Zhao J, Han Y, et al. Management of locally advanced non-small cell lung cancer: State of the art and future directions. Cancer Commun (Lond) 2024;44:23-46. [Crossref] [PubMed]
  5. Yan Y, Sun D, Hu J, et al. Multi-omic profiling highlights factors associated with resistance to immuno-chemotherapy in non-small-cell lung cancer. Nat Genet 2025;57:126-39. [Crossref] [PubMed]
  6. Sathiyapalan A, Baloush Z, Ellis PM. Update on the Management of Stage III NSCLC: Navigating a Complex and Heterogeneous Stage of Disease. Curr Oncol 2023;30:9514-29. [Crossref] [PubMed]
  7. Ma J, Jiang J. Exploration of immunotherapy modalities in stage III unresectable non-small cell lung cancer Oncol Lett 2026;31:46. (Review). [Crossref] [PubMed]
  8. Cardenal F, Nadal E, Jové M, et al. Concurrent systemic therapy with radiotherapy for the treatment of poor-risk patients with unresectable stage III non-small-cell lung cancer: a review of the literature. Ann Oncol 2015;26:278-88. [Crossref] [PubMed]
  9. Rami-Porta R, Osarogiagbon RU, Asamura H. The TNM System Is Adequate for Making Treatment Decisions and Prognostication in Lung Cancer. J Thorac Oncol 2022;17:1255-7. [Crossref] [PubMed]
  10. Hoeijmakers F, Schreurs WH, Comans EFI, et al. The TNM System Is Not Adequate to Guide Lung Cancer Multidisciplinary Teams in Treatment Decisions in the Precision Oncology Era. J Thorac Oncol 2022;17:1250-4. [Crossref] [PubMed]
  11. Godfrey CM, Shipe ME, Welty VF, et al. The Thoracic Research Evaluation and Treatment 2.0 Model: A Lung Cancer Prediction Model for Indeterminate Nodules Referred for Specialist Evaluation. Chest 2023;164:1305-14. [Crossref] [PubMed]
  12. Guo Y, Li L, Zheng K, et al. Development and validation of a survival prediction model for patients with advanced non-small cell lung cancer based on LASSO regression. Front Immunol 2024;15:1431150. [Crossref] [PubMed]
  13. Kourou K, Exarchos KP, Papaloukas C, et al. Applied machine learning in cancer research: A systematic review for patient diagnosis, classification and prognosis. Comput Struct Biotechnol J 2021;19:5546-55. [Crossref] [PubMed]
  14. Maurya SP, Sisodia PS, Mishra R, et al. Performance of machine learning algorithms for lung cancer prediction: a comparative approach. Sci Rep 2024;14:18562. [Crossref] [PubMed]
  15. Li Y, Wu X, Yang P, et al. Machine Learning for Lung Cancer Diagnosis, Treatment, and Prognosis. Genomics Proteomics Bioinformatics 2022;20:850-66. [Crossref] [PubMed]
  16. Burki TK. Predicting lung cancer prognosis using machine learning. Lancet Oncol 2016;17:e421. [Crossref] [PubMed]
  17. Huang P, Lin CT, Li Y, et al. Prediction of lung cancer risk at follow-up screening with low-dose CT: a training and validation study of a deep learning method. Lancet Digit Health 2019;1:e353-62. [Crossref] [PubMed]
  18. Wiesweg M, Mairinger F, Reis H, et al. Machine learning-based predictors for immune checkpoint inhibitor therapy of non-small-cell lung cancer. Ann Oncol 2019;30:655-7. [Crossref] [PubMed]
  19. Yu H, Wu H, Wang W, et al. Machine Learning to Build and Validate a Model for Radiation Pneumonitis Prediction in Patients with Non-Small Cell Lung Cancer. Clin Cancer Res 2019;25:4343-50. [Crossref] [PubMed]
  20. van Amsterdam WAC, Verhoeff JJC, de Jong PA, et al. Eliminating biasing signals in lung cancer images for prognosis predictions with deep learning. NPJ Digit Med 2019;2:122. [Crossref] [PubMed]
  21. Song X, Zhu J, Tan X, et al. XGBoost-Based Feature Learning Method for Mining COVID-19 Novel Diagnostic Markers. Front Public Health 2022;10:926069. [Crossref] [PubMed]
  22. Agha RA, Mathew G, Rashid R, et al. Revised Strengthening the Reporting of Cohort, Cross-Sectional and Case-Control Studies in Surgery (STROCSS) Guideline: An Update for the Age of Artificial Intelligence. Premier Journal of Science 2025;10:100081.
  23. Eberhardt WE, De Ruysscher D, Weder W, et al. 2nd ESMO Consensus Conference in Lung Cancer: locally advanced stage III non-small-cell lung cancer. Ann Oncol 2015;26:1573-88. [Crossref] [PubMed]
  24. Varlamova EV, Butakova MA, Semyonova VV, et al. Machine Learning Meets Cancer. Cancers (Basel) 2024;16:1100. [Crossref] [PubMed]
  25. Hattori A, Matsunaga T, Takamochi K, et al. Neither Maximum Tumor Size nor Solid Component Size Is Prognostic in Part-Solid Lung Cancer: Impact of Tumor Size Should Be Applied Exclusively to Solid Lung Cancer. Ann Thorac Surg 2016;102:407-15. [Crossref] [PubMed]
  26. Arnold BN, Thomas DC, Rosen JE, et al. Lung Cancer in the Very Young: Treatment and Survival in the National Cancer Data Base. J Thorac Oncol 2016;11:1121-31. [Crossref] [PubMed]
  27. Sacher AG, Dahlberg SE, Heng J, et al. Association Between Younger Age and Targetable Genomic Alterations and Prognosis in Non-Small-Cell Lung Cancer. JAMA Oncol 2016;2:313-20. [Crossref] [PubMed]
  28. Borgeaud M, Olivier T, Bar J, et al. Personalized care for patients with EGFR-mutant nonsmall cell lung cancer: Navigating early to advanced disease management. CA Cancer J Clin 2025;75:387-409. [Crossref] [PubMed]
  29. Huh Y, Sohn YJ, Kim HR, et al. Sex differences in prognosis factors in patients with lung cancer: A nationwide retrospective cohort study in Korea. PLoS One 2024;19:e0300389. [Crossref] [PubMed]
  30. Liu J, Wu J, Liu S, et al. Predicting mortality of patients with acute kidney injury in the ICU using XGBoost model. PLoS One 2021;16:e0246306. [Crossref] [PubMed]
  31. Ahmed S, Raza B, Hussain L, et al. The Deep Learning ResNet101 and Ensemble XGBoost Algorithm with Hyperparameters Optimization Accurately Predict the Lung Cancer. Applied Artificial Intelligence 2023;37:
  32. Guan X, Du Y, Ma R, et al. Construction of the XGBoost model for early lung cancer prediction based on metabolic indices. BMC Med Inform Decis Mak 2023;23:107. [Crossref] [PubMed]
  33. Huang Z, Hu C, Chi C, et al. An Artificial Intelligence Model for Predicting 1-Year Survival of Bone Metastases in Non-Small-Cell Lung Cancer Patients Based on XGBoost Algorithm. Biomed Res Int 2020;2020:3462363. [Crossref] [PubMed]
  34. Bentéjac C, Csörgő A, Martínez-Muñoz G. A comparative analysis of gradient boosting algorithms. Artif Intell Rev 2021;54:1937-67.
  35. Santo V, Brunetti L, Pecci F, et al. Host-related Determinants of Response to Immunotherapy in Non-small Cell Lung Cancer: The Interplay of Body Composition, Metabolism, Sex and Immune Regulation. Curr Oncol Rep 2025;27:1427-47. [Crossref] [PubMed]
  36. Zhong H, Zhang HH, Wu J, et al. Integrating CT Radiomics and Clinical Information to Predict Prognosis of Advanced NSCLC Patients Receiving Chemoimmunotherapy. Curr Med Sci 2025;45:1109-22. [Crossref] [PubMed]
  37. Liu M, Xia S, Zhang X, et al. Development and validation of a blood-based genomic mutation signature to predict the clinical outcomes of atezolizumab therapy in NSCLC. Lung Cancer 2022;170:148-55. [Crossref] [PubMed]
Cite this article as: Zhang Y, Hu S, Wang J, Wu Z, Chen C, Chen W, Xie L. Machine learning models predict survival in unresectable stage III non-small cell lung cancer: Surveillance, Epidemiology, and End Results and Chinese cohort study. J Thorac Dis 2026;18(7):710. doi: 10.21037/jtd-2026-1063

Download Citation