Development and internal validation of a high-resolution computed tomography radiomics and three-dimensional deep learning diagnostic prediction model for preoperative differentiation of minimally invasive and invasive adenocarcinoma in subsolid nodules
Highlight box
Key findings
• A radiomics-deep learning (DL) fusion model achieved the best performance for differentiating minimally invasive adenocarcinoma (IAC) from IAC in subsolid nodules (SSNs), with higher area under the curve (AUC) and better calibration than single-modality models.
• The fusion model provided greater net clinical benefit across a wide range of threshold probabilities in decision curve analysis (DCA).
• Radiomics features contributed the majority of predictive power (~80%), while DL features offered complementary value that improved overall performance.
What is known and what is new?
• Radiomics and deep learning have each shown promising ability to predict invasiveness in pulmonary SSNs. Existing models often focus primarily on discrimination and are commonly developed using a single feature type.
• This study demonstrates that integrating radiomics and DL features provided additional predictive information when integrated with radiomics features.
What is the implication, and what should change now?
• The fusion model may provide preliminary risk-stratification information, but its role in guiding surgical decisions requires prospective multicenter validation.
• Model evaluation should extend beyond AUC to include calibration and clinical utility (e.g., DCA) for real-world decision-making.
• Future studies should focus on external validation and integration with clinical variables to facilitate clinical implementation.
Introduction
Subsolid nodules (SSNs), including pure ground-glass nodules and part-solid nodules, are increasingly encountered on thin-section chest computed tomography (CT) in the era of widespread CT screening and opportunistic imaging (1). Their biological spectrum is broad, ranging from transient inflammatory lesions to indolent precursor glandular lesions and early lung adenocarcinoma. SSNs represent a distinct clinical entity with unique growth kinetics, prognosis, and management considerations, indicating that risk-adapted surveillance and surgical strategies should be individualized rather than applied in a one-size-fits-all manner (2). Among SSN-associated lung adenocarcinomas, distinguishing minimally invasive adenocarcinoma (MIA) from invasive adenocarcinoma (IAC) before surgery is particularly important because this distinction directly influences surgical planning, including the selection of limited resection versus more extensive resection (2).
Pathological assessment remains the reference standard for assessing invasiveness but is not suitable for preoperative decision-making (3). High-resolution CT (HRCT) is widely used for evaluation; however, conventional assessment relies on qualitative or semi-quantitative features, such as nodule size, solid component proportion, margin characteristics, and pleural indentation (4,5). Although these features have been incorporated into predictive models, their performance is variable and subject to interobserver variability. Existing clinical-radiological models, which often include factors such as age, smoking history, tumor markers, and CT morphology, provide only moderate predictive accuracy and may not adequately capture intratumoral heterogeneity (5,6). Artificial intelligence-based methods, particularly radiomics and deep learning (DL), have shown promise for noninvasive characterization of SSNs. Radiomics extracts quantitative imaging features that reflect lesion heterogeneity and has demonstrated value in predicting invasiveness and pathological subtypes of lung adenocarcinoma (3,7,8). DL methods can automatically learn complex imaging representations and have achieved encouraging results in differentiating precursor lesions, MIA, and invasive tumors (9). However, conventional radiomics generally relies on predefined handcrafted features and lesion segmentation, which may affect reproducibility, whereas DL models can automatically learn high-level imaging representations but are often limited by their “black-box” nature and insufficient interpretability (10,11).
Recent studies suggest that radiomics and DL provide complementary information. Radiomics offers interpretable quantitative descriptors of lesion morphology and texture, while DL captures high-level imaging patterns that may be difficult to identify visually (12,13). However, most studies have evaluated these approaches separately and have focused mainly on discrimination, with limited assessment of calibration, clinical utility, and interpretability (14,15). Therefore, radiomics-DL fusion may improve diagnostic robustness and clinical applicability compared with single-modality models.
Therefore, we developed and evaluated three HRCT-based models for differentiating MIA from IAC in patients with SSNs using radiomics features, DL representations, and a fused radiomics-DL feature set. We further assessed model discrimination, calibration, and clinical utility, while quantifying the relative contributions of radiomics and DL features. Our goal was to develop an interpretable and clinically applicable fusion model for more reliable preoperative assessment of pathological invasiveness in SSNs. We present this article in accordance with the TRIPOD reporting checklist (available at https://jtd.amegroups.com/article/view/10.21037/jtd-2026-0924/rc).
Methods
Study population
This study was designed as a retrospective diagnostic prediction model development and internal validation study, aiming to develop and evaluate HRCT-based models for differentiating MIA from IAC in patients with SSNs. Consecutive patients with SSNs were collected from the picture archiving and communication system of our hospital between September 2015 and June 2025. All patients underwent preoperative HRCT targeted scanning and had pathologically confirmed lung adenocarcinoma based on surgical resection specimens. Pathological diagnosis was used as the reference standard. Patients were stratified into the MIA group and an IAC group according to pathological invasiveness. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. This study was approved by the Ethics Committee of Jiaxing Hospital of Traditional Chinese Medicine (Approval No. JX2025.06-11). The requirement for informed consent was waived due to the retrospective nature of the study and the use of anonymized data.
Histopathological diagnosis was used as the reference standard. Pathological classification of MIA and IAC was performed according to the 2021 World Health Organization (WHO) classification of thoracic tumors by two board-certified pulmonary pathologists with 10 and 15 years of experience, respectively. Both pathologists were blinded to radiomics features, DL features, and model outputs. Disagreements were resolved by consensus.
Sample size estimation
Since this was a retrospective study, the sample size was determined by the number of consecutive eligible patients during the study period. No a priori sample size calculation was performed before data collection. To assess whether the available sample was acceptable for internal model development, we performed a post hoc adequacy assessment. In the final dataset, 374 patients were included, comprising 250 IAC events and 124 MIA cases. The maximum number of nonzero predictors retained in the final ElasticNet model was 10, resulting in an events-per-variable ratio of 25. ElasticNet regularization and nested stratified five-fold cross-validation were used to reduce overfitting and evaluate model stability. Therefore, the available sample size was considered acceptable for model development and internal validation, although external validation remains necessary to confirm generalizability.
Inclusion and exclusion criteria
The inclusion criteria were as follows: (I) SSNs confirmed by HRCT targeted scanning with adequate image quality for complete three-dimensional (3D) segmentation; (II) a single pulmonary lesion with complete surgical resection; (III) HRCT performed within three months before surgery without prior antitumor treatment; and (IV) pathological confirmation of MIA or IAC.
The exclusion criteria were as follows: (I) incomplete or poor-quality HRCT images unsuitable for segmentation; (II) solid nodules; (III) multiple SSNs; and (IV) unclear pathological classification.
Missing data handling
Missing data were assessed before model development. Patients with missing or non-diagnostic HRCT images, incomplete lesion segmentation, or unclear pathological classification were excluded according to the predefined exclusion criteria. In the final included cohort, no missing data were present for pathological diagnosis, HRCT images, segmentation masks, radiomics features, or DL-derived features. Therefore, complete-case analysis was performed, and no imputation was applied.
CT image acquisition
All patients underwent chest CT using either a 64-slice or 16-slice GE scanners (GE Healthcare, USA). Scanning was performed during deep-inspiration breath-hold, and patient positioning was adjusted to optimize lesion visualization. The scanning parameters were as follows: field of view 10 cm × 10 cm, tube voltage 120 kV, tube current 300 mA, detector collimation 16 mm × 0.625 mm, pitch 0.984, slice thickness and reconstruction interval 0.625 mm, matrix 512×512, and standard reconstruction algorithm (filter A). The reconstruction field of view was adjusted according to lesion size. Coronal and sagittal images were reconstructed with slice thicknesses of 3 and 5 mm, respectively.
Lesion segmentation
SSNs were identified on axial HRCT images, and manual 3D segmentation was performed using ITK-SNAP (version 4.0). Regions of interest encompassing the entire lesion were delineated by experienced radiologists blinded to pathological results. The segmentation was performed along the visible boundary of the SSN on consecutive axial slices, including both ground-glass and solid components when present, while excluding adjacent vessels, bronchi, pleura, and unrelated lung parenchyma as far as possible. Discrepancies were resolved by consensus. To reduce potential measurement bias, all radiologists were blinded to clinical data and pathological invasiveness during segmentation.
Radiomics feature extraction
Radiomics features were extracted from the segmented 3D regions of interest (ROIs) using PyRadiomics (version 3.1.0, USA). To reduce variability across scans, all CT images were resampled to a uniform voxel spacing and intensity-normalized before feature extraction. The extracted features included first-order statistics, shape features, and texture features derived from co-occurrence matrix (GLCM), gray-level run-length matrix (GLRLM), gray-level size-zone matrix (GLSZM), and neighboring gray-tone difference matrix (NGTDM), thereby characterizing lesion intensity, morphology, and heterogeneity. All features were standardized using z-score normalization before model construction to ensure comparability (16).
DL feature extraction
DL features were extracted using a 3D ResNet-18 model pretrained with MedicalNet weights. The pretrained network was used as a fixed feature extractor, and no fine-tuning was performed in the present dataset. For each patient, the manually segmented lesion volume was cropped according to the 3D bounding box of the ROI with a 5-voxel margin and resized to 64×64×64 voxels. Voxel intensities were normalized using z-score normalization before network input. Patient-level DL features were extracted from the penultimate global average pooling layer, resulting in 512 DL-derived features for each lesion. These DL features were subsequently used alone in the DL-only model or concatenated with radiomics features in the fusion model (17).
Feature processing and selection
Candidate predictors included radiomics features and DL-derived features. For the radiomics-only model, only handcrafted radiomics features were used. For the DL-only model, only pretrained network-derived features were used. For the fusion model, radiomics and DL features were concatenated at the feature level. Before feature selection, all continuous imaging features were standardized using z-score normalization. Highly redundant features were removed using pairwise correlation analysis, and one feature from each pair with a correlation coefficient greater than 0.90 was retained. ElasticNet-regularized logistic regression was then used for final feature selection and model construction. Feature selection was performed within each training fold only, and the selected feature set was subsequently applied to the corresponding validation fold to prevent information leakage.
Model construction
Three models were developed to assess pathological invasiveness: a radiomics-only model, a DL-only model, and a fusion model integrating both feature types. All models were constructed using ElasticNet-regularized logistic regression, which combines L1 and L2 penalties for simultaneous feature selection and overfitting control. For the fusion model, radiomics and DL features were concatenated at the feature level, representing an early-fusion strategy. All features were standardized using z-score normalization before model fitting. The complete modeling workflow was implemented within a nested cross-validation framework. In each outer iteration, four folds were used for model training, and the remaining fold was used for validation. Regularization parameters were optimized through inner cross-validation within the training folds. The finalized model from each training iteration was then evaluated on the corresponding held-out validation fold. This procedure ensured that feature preprocessing, feature selection, hyperparameter tuning, and model fitting were performed without using information from the validation fold, thereby reducing the risk of information leakage.
After cross-validation-based internal validation, a final fusion model was refitted using the entire cohort with the same preprocessing, feature selection, and ElasticNet-regularized logistic regression procedures. The complete final model specification, including selected predictors, regression coefficients, intercept, standardization parameters, regularization parameters, and the final probability threshold, is provided in Table S1. Predicted probability of IAC was calculated as: , where η is the linear predictor calculated from the weighted sum of standardized selected features.
Model evaluation and validation
Since this was a single-center retrospective study, stratified five-fold cross-validation was used for internal validation. The full cohort was randomly divided into five folds, with stratification by pathological invasiveness to maintain a similar MIA/IAC distribution across folds. In each iteration, four folds were used for training and the remaining fold was used for validation. Each patient was included in the validation fold once, and final performance metrics were calculated from out-of-fold predictions. Discriminative performance was assessed using receiver operating characteristic (ROC) curves and the area under the curve (AUC). Optimal probability thresholds were determined within the training data of each fold using the Youden index. Sensitivity, specificity, and accuracy were then calculated in the corresponding validation fold and averaged across folds. To facilitate clinically meaningful interpretation, performance thresholds were predefined before model comparison. An AUC of ≥0.75 was considered acceptable discrimination for a diagnostic prediction model, whereas an AUC of ≥0.80 was considered good discrimination. Because underestimation of IAC may lead to insufficient surgical resection, sensitivity for identifying IAC was prioritized, with a sensitivity of ≥0.80 regarded as clinically desirable. In parallel, a specificity of ≥0.70 was considered desirable to reduce the risk of unnecessary overtreatment of MIA. These thresholds were used to contextualize model performance and to assess whether the models could provide supportive information for multidisciplinary decision-making regarding the extent of pulmonary resection, rather than to define an automatic indication for surgery.
In addition, 1,000 bootstrap resamples were performed based on the out-of-fold predictions to estimate 95% confidence intervals (CIs) for AUC, sensitivity, specificity, accuracy, and Brier score. Bootstrap resampling was used to quantify the uncertainty of internal validation performance, not as a substitute for external validation.
Calibration assessment
Calibration performance was evaluated using calibration curves generated from out-of-fold predictions. Agreement between predicted probabilities and observed outcomes was quantified using the Brier score, calibration intercept, and calibration slope. A lower Brier score indicates better overall prediction accuracy, while a calibration intercept close to 0 and a calibration slope close to 1 indicate better agreement between predicted and observed risks.
Decision curve analysis (DCA)
Clinical utility was assessed using DCA based on out-of-fold predictions. Net benefit was calculated across threshold probabilities from 0.10 to 0.90 and compared with treat-all and treat-none strategies. In the context of SSN management, threshold probabilities between 0.20 and 0.70 were considered clinically relevant, because this range reflects the probability interval in which the choice between limited resection and more extensive resection is most uncertain.
Model interpretability analysis
Model interpretability was assessed using ElasticNet regression coefficients for the radiomics-only and fusion models. In the radiomics-only model, the most influential features were identified according to the magnitude and direction of their coefficients. In the fusion model, the relative contributions of radiomics and DL-derived features were quantified by summing the absolute coefficients within each feature group. This coefficient-based analysis was used to evaluate the relative contribution of handcrafted radiomics descriptors and DL-derived representations to model output, thereby providing a partial assessment of feature-level interpretability.
Statistical analysis
Normality of continuous variables was assessed prior to analysis. Continuous variables with approximately normal distribution were expressed as mean ± standard deviation and compared using the independent-samples t-test. Categorical variables were compared using the Chi-squared test. All statistical analyses were performed using Python (version 3.11, USA), with a two-sided P<0.05 considered statistically significant.
Results
Baseline characteristics of patients stratified by pathological invasiveness
As shown in Table 1, a total of 374 patients were included in the final analysis, comprising 124 patients with MIA and 250 patients with IAC. No significant differences were observed between the MIA and IAC groups in baseline demographic characteristics, including age (55.82±9.73 vs. 56.91±10.42 years, P=0.26), sex (male: 57.26% vs. 58.00%, P=0.88), and smoking history (33.06% vs. 36.80%, P=0.16). In contrast, conventional CT morphological features differed significantly between groups. Compared with the MIA group, the IAC group showed higher frequencies of lobulation sign (48.80% vs. 37.10%, P=0.03), spiculation sign (40.80% vs. 29.03%, P=0.04), and pleural indentation sign (34.40% vs. 23.39%, P=0.03). The IAC group also had a larger maximum nodule diameter (18.31±4.79 vs. 12.58±3.47 mm, P<0.001) and a higher proportion of solid components (0.45±0.12 vs. 0.31±0.08, P<0.001). These findings suggest that demographic characteristics were generally balanced between pathological groups, whereas greater morphological complexity, larger lesion size, and increased solid components were associated with higher pathological invasiveness. Radiomics and DL features were not included in baseline comparisons, as they were incorporated into subsequent modeling analyses.
Table 1
| Item | MIA group (n=124) | IAC group (n=250) | χ2/t value | P value |
|---|---|---|---|---|
| Age (years) | 55.82±9.73 | 56.91±10.42 | 1.12 | 0.26 |
| Sex | 0.02 | 0.88 | ||
| Male | 71 (57.26) | 145 (58.00) | ||
| Female | 53 (42.74) | 105 (42.00) | ||
| Smoking history | 1.98 | 0.16 | ||
| Yes | 41 (33.06) | 92 (36.80) | ||
| No | 83 (66.94) | 158 (63.20) | ||
| Lobulation sign | 46 (37.10) | 122 (48.80) | 4.63 | 0.03 |
| Spiculation sign | 36 (29.03) | 102 (40.80) | 4.44 | 0.04 |
| Pleural indentation sign | 29 (23.39) | 86 (34.40) | 4.89 | 0.03 |
| Maximum diameter of nodule (mm) | 12.58±3.47 | 18.31±4.79 | 10.23 | <0.001 |
| Proportion of solid components | 0.31±0.08 | 0.45±0.12 | 10.40 | <0.001 |
Data are presented as n (%) or mean ± standard deviation. IAC, invasive adenocarcinoma; MIA, minimally invasive adenocarcinoma.
Model performance comparison for predicting pathological invasiveness
Model performance was evaluated using stratified five-fold cross-validation with out-of-fold predictions, supplemented by 1,000 bootstrap resamples to estimate 95% CIs. As shown in Figure 1 and Table 2, the three models demonstrated different levels of discriminative performance for differentiating MIA from IAC. The radiomics-only model achieved an AUC of 0.833, with a sensitivity of 0.784, specificity of 0.727, and accuracy of 0.765. The DL-only model showed improved discrimination, with an AUC of 0.887, sensitivity of 0.844, specificity of 0.727, and accuracy of 0.805. The fusion model showed the highest numerical performance among the three models, with an AUC of 0.937, sensitivity of 0.864, specificity of 0.807, and accuracy of 0.845.
Table 2
| Model | AUC (95% CI) | Sensitivity (95% CI) | Specificity (95% CI) | Accuracy (95% CI) | Brier score (95% CI) |
|---|---|---|---|---|---|
| Radiomics-only | 0.833 (0.790–0.875) | 0.784 (0.729–0.831) | 0.727 (0.641–0.797) | 0.765 (0.719–0.805) | 0.183 (0.162–0.204) |
| DL-only | 0.887 (0.851–0.920) | 0.844 (0.794–0.884) | 0.727 (0.641–0.797) | 0.805 (0.762–0.842) | 0.120 (0.102–0.140) |
| Fusion | 0.937 (0.912–0.960) | 0.864 (0.816–0.901) | 0.807 (0.728–0.866) | 0.845 (0.805–0.878) | 0.099 (0.083–0.116) |
AUC, area under the curve; CI, confidence interval; DL, deep learning.
Using optimal thresholds determined by the Youden index, the fusion model consistently showed higher sensitivity, specificity, and accuracy than the radiomics-only and DL-only models. These results suggest that integrating radiomics and DL features may provide complementary information and improve the internal validation performance for preoperative differentiation between MIA and IAC.
Calibration performance of different models
Calibration performance was assessed using calibration curves generated from out-of-fold predictions (Figure 2). The fusion model showed the lowest Brier score among the three models, followed by the DL-only model and the radiomics-only model. The Brier scores were 0.099 for the fusion model, 0.120 for the DL-only model, and 0.183 for the radiomics-only model. The fusion model also showed better calibration, with a calibration intercept of 0.02 and a calibration slope of 0.94, compared with the DL-only model, which had an intercept of 0.04 and slope of 0.88, and the radiomics-only model, which had an intercept of 0.07 and a slope of 0.71. These findings suggest that the fusion model provided more reliable probability estimates than the single-modality models.
Clinical utility assessed by DCA
The clinical utility of the three models was evaluated using DCA based on out-of-fold predictions from five-fold cross-validation (Figure 3). Across threshold probabilities from 0.20 to 0.70, the fusion model generally provided higher net benefit than the radiomics-only model, DL-only model, treat-all strategy, and treat-none strategy. This finding suggests that integrating radiomics and DL features may improve clinical decision-support value for identifying IAC. However, because DCA was based on internally validated single-center data, the observed net benefit should be interpreted as preliminary and requires confirmation in external cohorts.
Model interpretability
Model interpretability was assessed based on ElasticNet coefficients. In the radiomics-only model, the most influential features were identified according to coefficient magnitude (Figure 4A). These features mainly included texture features derived from GLCM, GLSZM, first-order intensity features, and a shape feature. Several GLCM and first-order features showed positive coefficients, suggesting an association with IAC, whereas selected GLSZM and shape features showed negative coefficients, suggesting an association with MIA. These findings indicate that texture heterogeneity, intensity distribution, and morphological characteristics may contribute to the prediction of pathological invasiveness.
For the fusion model, the relative contributions of radiomics and DL features were quantified by summing absolute coefficients within each feature group (Figure 4B). Radiomics features accounted for approximately 80% of the total coefficient weight, while DL features accounted for about 20%. This distribution suggests that handcrafted radiomics features provided the major contribution to the model output, while DL-derived features supplied additional predictive information when integrated with radiomics features. However, because DL features are latent representations extracted from the pretrained network, their direct biological or radiological interpretation remains limited. Therefore, this coefficient-based analysis should be considered a partial interpretability assessment rather than a complete explanation of model behavior.
The final fusion model retained 10 predictors, including seven radiomics features and three DL-derived features. The final linear predictor was: η =−0.684 + 0.742Z1 + 0.913Z2 + 0.621Z3 + 0.584Z4 + 0.462Z5 + 0.431Z6 − 0.392Z7 + 0.352Z8 + 0.314Z9 + 0.276Z10; , where represents the standardized value of each selected predictor. The final probability threshold selected by the Youden index was 0.61. The full names of the selected predictors, feature types, standardization parameters, and coefficients are provided in Table S1.
Discussion
In this study, we developed and compared three HRCT-based models for differentiating MIA from IAC in pathologically confirmed SSNs. All three models showed acceptable discriminatory performance, and the fusion model achieved the best overall performance in terms of AUC, calibration, and DCA. However, these findings should be interpreted as preliminary evidence of the mathematical and methodological feasibility of integrating radiomics and DL features. Although the fusion model demonstrated better calibration and higher net benefit in internal validation, its actual value in surgical decision-making requires confirmation through prospective and external validation.
The clinical importance of preoperatively differentiating MIA from IAC is well-established. MIA is associated with excellent recurrence-free survival after limited resection, whereas IAC has greater malignant potential and may require more extensive surgical management, such as lobectomy with systematic lymph node evaluation, to optimize oncologic outcomes (18). Accurate assessment of invasiveness may therefore help inform surgical planning by balancing pulmonary function preservation against oncological safety. In clinical practice, intraoperative frozen section evaluation is often used to assess invasiveness, but its accuracy may be limited by sampling error and inter-pathologist variability, supporting the need for more reliable noninvasive preoperative risk stratification (19). Nevertheless, any preoperative prediction model that may influence the extent of pulmonary resection should be interpreted cautiously. False-negative prediction of IAC could lead to insufficient resection, whereas false-positive prediction of MIA could contribute to unnecessary overtreatment. Therefore, the present model should be regarded as a potential adjunctive risk-stratification tool.
Conventional radiological evaluation based on semantic CT features, such as nodule size, consolidation component, and qualitative margin characteristics, provides clinically useful information but remains subjective and susceptible to inter-reader variability, scanner protocols and reconstruction parameters (20). In addition, the overlapping morphological appearances of MIA and IAC may limit the discriminative ability of traditional CT features, particularly in nodules with subtle or intermediate imaging characteristics (21). These limitations support the use of quantitative artificial intelligence (AI) methods that can extract high-dimensional imaging information beyond visual assessment.
Previous studies have demonstrated the value of radiomics for invasiveness assessment in SSNs (22). CT-based radiomic signatures, often combined with clinical or semantic features, have achieved AUC values of approximately 0.80–0.88 for differentiating MIA from IAC, suggesting that high-dimensional texture and shape descriptors can capture subtle imaging heterogeneity associated with invasive growth. For example, a model integrating CT radiomics with clinical variables achieved an AUC of 0.878 in distinguishing MIA from IAC in SSNs, indicating that quantitative tumor descriptors may complement conventional parameters such as nodule size and consolidation ratio (5). However, standalone radiomics models remain dependent on lesion segmentation, handcrafted feature definitions, and preprocessing settings, which may limit their reproducibility across scanners, reconstruction protocols, and institutions. In this context, the potential advantage of our fusion approach lies not merely in the observed improvement in AUC, but in its ability to integrate relatively interpretable radiomics descriptors of lesion morphology and texture with DL-derived abstract representations that may capture additional high-level imaging patterns.
Similar to radiomics, DL models have increasingly been applied to SSN characterization (23). DL frameworks based on 3D convolutional networks have shown promising performance in differentiating invasive from noninvasive pulmonary nodules, with AUCs exceeding 0.80 in multiple cohorts (24). These findings suggest that hierarchical latent features may capture complex morphological and contextual information beyond handcrafted descriptors. However, DL models are often developed independently of radiomics features and may be affected by dataset size, case heterogeneity, and limited interpretability.
More recent studies have explored the integration of radiomics and DL to leverage complementary imaging information. A retrospective study comparing multiple fusion strategies found that late fusion of radiomics and DL features achieved the highest AUC for predicting MIA versus IAC among tested models, although the differences was not statistically significance in that cohort (13). Similarly, a multicenter study using a multiple-instance learning framework that integrated radiomics and DL representations reported more stable performance across external test sets than single-modality models, with favorable calibration and decision curve results (25). These findings are consistent with our results and suggest that integrating heterogeneous feature families may enhance model robustness compared with approaches based solely on handcrafted or learned features. Compared with standalone models, the present fusion strategy may be advantageous because radiomics features provide explicit quantitative descriptors of lesion morphology, intensity, and texture, whereas DL features may encode higher-order spatial patterns that are difficult to predefine visually or mathematically. In addition, ElasticNet regularization enabled feature selection and coefficient-based contribution analysis, thereby providing partial interpretability compared with fully end-to-end DL systems.
Unlike many previous studies that focused primarily on discrimination metrics, the present study also evaluated calibration and decision curve performance, which are important for assessing whether predicted probabilities may support clinical decision-making. The relatively good calibration and higher net benefit of the fusion model across relevant threshold probabilities suggest that it may provide more informative risk stratification than single-modality models in the internal validation setting. Moreover, by quantifying the relative contribution of radiomics and DL features within an ElasticNet modeling framework, we provided interpretable evidence of complementary predictive information, addressing an important translational barrier related to model explainability (26). However, this interpretability remains incomplete. Although ElasticNet coefficients can indicate the direction and relative contribution of selected predictors, DL features extracted from the penultimate layer do not necessarily correspond to readily recognizable imaging phenotypes. Therefore, the model may still be perceived as partially opaque by clinicians, which could affect trust and adoption in preoperative workflows. Future studies should incorporate more intuitive explainability tools, prospective human-AI interaction assessment, standardized reporting of model outputs, and evaluation of how model predictions influence multidisciplinary decision-making.
From a translational perspective, the proposed model is not yet ready for direct clinical implementation. Before it can be incorporated into a thoracic surgery workflow, several requirements must be met, including external validation across institutions and scanners, assessment of segmentation robustness, prospective evaluation of model-assisted decision-making, and clarification of how model-derived risk probabilities should be integrated with radiologists’ interpretation, intraoperative frozen section findings, patient comorbidities, and surgeon judgment. Thus, the current model should be viewed as a preliminary decision-support framework rather than a validated clinical tool.
However, this study has some limitations. First, this was a retrospective single-center study with only internal validation using stratified five-fold cross-validation and 1,000 bootstrap resampling. Although this approach provided an initial estimate of model stability, it cannot replace independent temporal or geographic external validation; therefore, the generalizability of the model across different institutions, populations, scanners, and imaging protocols requires further multicenter validation. Second, semantic CT features and clinical variables were not systematically integrated into the fusion model, and future studies should determine whether incorporating these factors can improve model performance and interpretability. Third, manual 3D ROI segmentation may introduce observer-dependent variability. Although segmentation was performed by experienced radiologists blinded to pathology and discrepancies were resolved by consensus, inter- and intra-observer reproducibility were not fully assessed using intraclass correlation coefficient (ICC) analysis, which may affect radiomics feature stability. Finally, prospective studies are needed to evaluate the real-world utility of the model and its impact on surgical decision-making and clinical workflows.
Conclusions
In conclusion, this study provides preliminary proof-of-concept evidence that integrating HRCT-based radiomics and DL features is mathematically feasible and may improve the noninvasive differentiation of MIA from IAC in SSNs compared with single-modality approaches. By combining complementary imaging representations and evaluating discrimination, calibration, and decision-analytic performance, this study supports the potential value of integrative AI for early lung adenocarcinoma characterization. However, because this was a retrospective single-center study without independent external validation, the clinical utility and generalizability of the proposed fusion model remain unproven. Prospective, multicenter validation is required before the model can be considered for routine clinical use or for guiding surgical decision-making.
Acknowledgments
None.
Footnote
Reporting Checklist: The authors have completed the TRIPOD reporting checklist. Available at https://jtd.amegroups.com/article/view/10.21037/jtd-2026-0924/rc
Data Sharing Statement: Available at https://jtd.amegroups.com/article/view/10.21037/jtd-2026-0924/dss
Peer Review File: Available at https://jtd.amegroups.com/article/view/10.21037/jtd-2026-0924/prf
Funding: This study was supported by
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://jtd.amegroups.com/article/view/10.21037/jtd-2026-0924/coif). The authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. This study was approved by the Ethics Committee of Jiaxing Hospital of Traditional Chinese Medicine (Approval No. JX2025.06-11). The requirement for informed consent was waived due to the retrospective nature of the study and the use of anonymized data.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Raad RA, Garrana S, Moreira AL, et al. Imaging and Management of Subsolid Lung Nodules. Radiol Clin North Am 2025;63:517-35. [Crossref] [PubMed]
- Chen H, Kim AW, Hsin M, et al. The 2023 American Association for Thoracic Surgery (AATS) Expert Consensus Document: Management of subsolid lung nodules. J Thorac Cardiovasc Surg 2024;168:631-647.e11. [Crossref] [PubMed]
- Du H, Shen J, Chen F, et al. Integrating CT-based radiomics and deep learning for invasive prediction of ground-glass nodules in lung adenocarcinoma: a multicohort study. Insights Imaging 2025;16:271. [Crossref] [PubMed]
- Zhao Z, Yang H, Wang W. The value analysis of high-resolution thin-layer CT in the identification of early lung adenocarcinoma: An observation study. Medicine (Baltimore) 2024;103:e39608. [Crossref] [PubMed]
- Liu M, Duan R, Xu Z, et al. CT-based radiomics combined with clinical features for invasiveness prediction and pathological subtypes classification of subsolid pulmonary nodules. Eur J Radiol Open 2024;13:100584. [Crossref] [PubMed]
- Zhao Y, Ye Z, Yan Q, et al. Predicting the invasiveness of ground-glass opacity predominant lung adenocarcinoma with clinical stage Ia: a CT-based semantic and radiomics analysis. J Thorac Dis 2024;16:6713-26. [Crossref] [PubMed]
- Lin M, Li L, Hui Y, et al. Machine learning-based prediction of invasiveness in lung adenocarcinoma presenting as ground-glass nodules using radiomics and clinical CT features. BMC Cancer 2025;25:1693. [Crossref] [PubMed]
- Zuo Z, Zeng Y, Deng J, et al. Intratumoral heterogeneity score enhances invasiveness prediction in pulmonary ground-glass nodules via stacking ensemble machine learning. Insights Imaging 2025;16:209. [Crossref] [PubMed]
- Shao M, Wang J, Zhu L, et al. Deep learning for Three-Class Classification of ground-glass nodules on non-enhanced chest CT: A multicenter comparative study of CNN architectures. Eur J Radiol Open 2025;15:100690. [Crossref] [PubMed]
- Rundo L, Militello C. Image biomarkers and explainable AI: handcrafted features versus deep learned features. Eur Radiol Exp 2024;8:130. [Crossref] [PubMed]
- Demircioğlu A. Reproducibility and interpretability in radiomics: a critical assessment. Diagn Interv Radiol 2025;31:321-8. [Crossref] [PubMed]
- Tran AT, Wen J, Abou Karam G, et al. Comparing Handcrafted Radiomics Versus Latent Deep Learning Features of Admission Head CT for Hemorrhagic Stroke Outcome Prediction. BioTech (Basel) 2025;14:87. [Crossref] [PubMed]
- Sun Q, Yu L, Song Z, et al. Deep learning and radiomics fusion for predicting the invasiveness of lung adenocarcinoma within ground glass nodules. Sci Rep 2025;15:29285. [Crossref] [PubMed]
- Zheng X, He B, Hu Y, et al. Diagnostic Accuracy of Deep Learning and Radiomics in Lung Cancer Staging: A Systematic Review and Meta-Analysis. Front Public Health 2022;10:938113. [Crossref] [PubMed]
- Collins GS, Reitsma JB, Altman DG, et al. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): the TRIPOD statement. BMJ 2015;350:g7594. [Crossref] [PubMed]
- Shin HB, Sheen H, Oh JH, et al. Evaluating feature extraction reproducibility across image biomarker standardization initiative-compliant radiomics platforms using a digital phantom. J Appl Clin Med Phys 2025;26:e70110. [Crossref] [PubMed]
- Shen T, Hou R, Ye X, et al. Predicting Malignancy and Invasiveness of Pulmonary Subsolid Nodules on CT Images Using Deep Learning. Front Oncol 2021;11:700158. [Crossref] [PubMed]
- Li D, Deng C, Wang S, et al. Ten-year follow-up of lung cancer patients with resected adenocarcinoma in situ or minimally invasive adenocarcinoma: Wedge resection is curative. J Thorac Cardiovasc Surg 2022;164:1614-1622.e1. [Crossref] [PubMed]
- Mukhopadhyay S. Thoracic Frozen Section Pitfalls: Lung Adenocarcinoma Versus Selected Mimics. Arch Pathol Lab Med 2025;149:e93-e99. [Crossref] [PubMed]
- Li Y, Long Y, Zheng Y, et al. Ternary-Classification Habitat Model for Invasiveness and Grade of Lung Adenocarcinoma Presenting as a Subsolid Nodule on Low-Dose Chest CT: A Multicenter Study. AJR Am J Roentgenol 2025;225:e2533352. [Crossref] [PubMed]
- Fu CL, Yang ZB, Li P, et al. Discrimination of ground-glass nodular lung adenocarcinoma pathological subtypes via transfer learning: A multicenter study. Cancer Med 2023;12:18460-9. [Crossref] [PubMed]
- Dong N, Yan Y, Li Y, et al. CT-based habitat radiomics for preoperative differentiation of adenocarcinoma in situ/minimally invasive adenocarcinoma from invasive adenocarcinoma manifesting as ground-glass nodules: a multicenter study. Front Oncol 2025;15:1660071. [Crossref] [PubMed]
- Chassagnon G, De Margerie-Mellon C, Vakalopoulou M, et al. Artificial intelligence in lung cancer: current applications and perspectives. Jpn J Radiol 2023;41:235-44. [Crossref] [PubMed]
- Chen J, Yan W, Shi Y, et al. Radiomics and deep learning methods for predicting the growth of subsolid nodules based on CT images. Medicine (Baltimore) 2025;104:e44104. [Crossref] [PubMed]
- Liu J, Qi L, Wang Y, et al. Development of a combined radiomics and CT feature-based model for differentiating malignant from benign subcentimeter solid pulmonary nodules. Eur Radiol Exp 2024;8:8. [Crossref] [PubMed]
- Houssein EH, Gamal AM, Younis EM, et al. Explainable artificial intelligence for medical imaging systems using deep learning: a comprehensive review. Cluster Comput 2025;28:469.

