Original Article
Clinical Utility of CT-Based Deep Learning for Differentiating Pulmonary Aspergillosis from Bacterial Pneumonia: A Multicenter Retrospective Study
Abstract
Background: Early differentiation of pulmonary aspergillosis (PA) from bacterial pneumonia (BP) remains challenging in respiratory practice because of overlapping clinical and computed tomography (CT) features. Conventional microbiological and serological tests may be invasive, delayed, or insufficiently sensitive in selected patients. This study aimed to develop and externally validate CT-based deep learning models as adjunctive tools for differentiating PA from BP.
Methods: This multicenter retrospective study included patients diagnosed with PA or BP at the First Affiliated Hospital of Guangzhou Medical University and the Second Affiliated Hospital of Guangzhou Medical University. CT images from the First Affiliated Hospital of Guangzhou Medical University were used for model development and internal validation, while data from the Second Affiliated Hospital of Guangzhou Medical University were used for external validation. Two CT-based deep learning models were developed: an unannotated supervised training (UST) model using CT images without manual lesion annotation and an annotated supervised training (AST) model incorporating physician-annotated lesions. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC), sensitivity, specificity, accuracy, positive predictive value, and negative predictive value. A standardized image-only questionnaire was also used to explore clinician performance in differentiating PA from BP.
Results: A total of 387 PA cases and 311 BP cases were included. In internal validation, the AST model achieved an AUC of 0.83 (95% confidence interval [CI], 0.76–0.90), with 81.7% sensitivity, 73.2% specificity, and 78.0% accuracy, whereas the UST model achieved an AUC of 0.78 (95% CI, 0.69–0.86). In external validation, the AST model achieved an AUC of 0.81 (95% CI, 0.71–0.92), compared with 0.56 (95% CI, 0.41–0.70) for the UST model. In the standardized image-only clinician assessment, 158 valid responses were analyzed, with an average diagnostic accuracy of 68.2%. Exploratory analysis suggested that model performance varied across disease stages, highlighting the potential influence of temporal radiological changes.
Conclusions: The annotated CT-based deep learning model showed potential as an adjunctive imaging tool for differentiating PA from BP. However, the external sensitivity of 60.6% indicates that the model should not be used as a stand-alone diagnostic or rule-out tool. Prospective multicenter validation and integration of clinical and microbiological data are required before clinical implementation.

