Preprint
Article

This version is not peer-reviewed.

Evaluation of an Established Semi-Quantitative Chest CT Scoring System for Assessing the Severity of COVID-19 Pneumonia: What is Its Diagnostic Value Regarding Patient Outcomes?

Submitted:

26 June 2026

Posted:

29 June 2026

You are already at the latest version

Abstract
Background/Objectives: Investigation of whether visual and AI-based assessments of the severity of COVID-19 pneumonia using an established semi-quantitative chest CT scoring system (Pan score) correlate with laboratory parameters as well as pulmonary function, and of the score’s diagnostic value in predicting the patients’ clinical outcome. Methods: This retrospective analysis comprises patients with PCR-confirmed COVID-19, who received a chest CT scan (not more than three days prior to or after the positive PCR test) between March 21, 2020, and December 27, 2021. Each of the five lung lobes was assessed separately using a scoring system ranging from 0 (no pulmonary involvement) to 5 (> 75% pulmonary involvement) by a radiology specialist, an experienced resident physician, a medical student, and a dedicated AI-based chest CT software tool. In addition, pulmonary function and laboratory parameters, the duration of ICU stays and of any required mechanical ventilation, as well as the clinical outcome (discharge vs. death) were recorded, and their correlation with the obtained CT score was analysed. Statistical analyses comprised descriptive baseline comparisons using non-parametric tests, bivariate correlation matrices, and ROC curves to assess diagnostic accuracy. Furthermore, multivariable logistic, ordinal, and age-adjusted spline regression models were constructed to calculate odds ratios and estimate predicted probabilities for cumulative ICU and mechanical ventilation duration thresholds. Results: In total, 351 consecutive patients with confirmed COVID-19 (223 males [63.5%], 128 females [36.5%]; mean age 67.0 years) were included, all of whom underwent at least one chest CT scan. Compared with patients who were discharged, deceased patients had a significantly (p < 0.05) higher mean Pan score (11.7 ± 6.0 vs. 8.8 ± 5.0), higher rates of mechanical ventilation (56.0 vs. 32.1 %), and both a higher incidence (42.7 vs. 25.4 %) and longer duration (10.5 [6.0–20.0] vs. 6.0 [3.0–10.0] days) of ICU stays. The Pan score showed a strong and consistent association with the requirement for mechanical ventilation (r = 0.54; p < 0.01) and the duration of the ICU stay (r = 0.39; p < 0.01). Conclusions: The investigated semi-quantitative CT score is a simple, reliable tool for assessing the extent of COVID-19 pneumonia and can be evaluated both by radiologists and fully automated AI software. While its predictive value for all-cause mortality was only moderate, it showed good performance in predicting the need for and duration of mechanical ventilation, as well as intensive care requirement.
Keywords: 
;  ;  ;  ;  

1. Introduction

In March 2020, the World Health Organization (WHO) officially declared coronavirus disease 2019 (COVID-19), caused by severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), a global pandemic. By the end of February 2026, more than 779,000,000 cases of COVID-19 and approximately 7,000,000 deaths had been reported worldwide [1]. Although mortality has steadily declined over time due to vaccination programmes as well as public health and hygiene measures, COVID-19 is still considered endemic in Germany. The high level of population immunity has contributed to a reduction in severe disease courses and long-term sequelae. Nevertheless, further waves of infection are expected, which will predominantly affect elderly individuals and those with pre-existing conditions [2].
Since the onset of the pandemic, chest imaging – particularly computed tomography (CT) – has played a crucial role. It is not only essential for monitoring disease progression and identifying complications, but also for supporting the diagnosis in suspected cases, for example with polymerase chain reaction (PCR) test results that are inconsistent with a clear clinical picture including typical COVID-19 symptoms [3].
On CT imaging, COVID-19 pneumonia most commonly presents with multifocal, predominantly bilateral ground-glass opacities (GGO). These typically exhibit a patchy or geographic distribution and are mainly located in the peripheral and posterobasal regions of the lungs. Consolidations and the so-called crazy-paving pattern are observed less frequently and tend to occur in more advanced stages of the disease [4]. In addition, interlobular septal thickening, vascular enlargement and the halo sign are frequently reported parenchymal findings associated with COVID-19 pneumonia [5].
Quantitative evaluation of lung parenchymal involvement has been shown to be an important factor in predicting disease outcome [6]. One approach is the semi-quantitative CT score proposed by Pan et al., in which each lung lobe is graded according to the extent of involvement (from 0 for no involvement to 5 for > 75 % involvement). A total score of ≥ 18 – calculated as the sum of all five lobes – is associated with a significantly increased risk of a fatal outcome [7,8]. However, this scoring system considers only the extent of lung involvement and does not account for the specific type of parenchymal abnormalities.
A previous study has already shown that inter-rater agreement on the scores is high, even among readers with varying levels of experience, and that fully automated, AI-based evaluations also produce reliable results [9]. It has also been displayed that the score correlates positively with inflammatory parameters, the age of the patients, and the severity of the disease [7].
The aim of this study was to evaluate the diagnostic value of the Pan score with regard to disease severity, i. e. the need for intensive care admission and mechanical ventilation, as well as the patients’ outcomes, and to investigate the extent to which it correlates with various laboratory parameters, the pulmonary function, and the age of the patients.
Our hypothesis was that a higher score would be associated with increased inflammatory parameters, a higher likelihood of intensive care admission, a greater probability of a fatal outcome, and poorer pulmonary function – independent of the patients’ age.

2. Materials and Methods

For this retrospective study, we included 351 consecutive patients (223 males [63.5%], 128 females [36.4%]; mean age 67.0 years) with polymerase chain reaction (PCR)-confirmed COVID-19, who underwent chest CT at our hospital between March 21, 2020, and December 27, 2021. The first case of COVID-19 in our district was confirmed in March 2020.
Patients were included based on the following criteria: (a) positive PCR test for COVID-19, (b) chest CT scan performed not more than three days prior or three days after the positive PCR test, (c) age ≥ 18 years. 38 of the original 607 patients were excluded owing to motion/breathing artifacts, prior extensive lung surgery, incomplete CT data sets, or other technical difficulties that prevented AI-based evaluation. Furthermore, a further 218 patients were excluded due to incomplete clinical and laboratory data for the study period. Approval for this study was granted by the local ethics committee (Medical Association Westphalia-Lippe and University of Münster) and conducted following the principles of the Declaration of Helsinki.
CT images were obtained from the hospital’s picture archiving and communication system (PACS; CentricityTM Universal Viewer, Version 6.0, GE HealthCare Technologies Inc., Chicago, USA) and organised in a dedicated list within in the radiological information system (RIS; CentricityTM RIS-i7, Version 7.0.4.4, GE HealthCare Technologies Inc., Chicago, USA).
Laboratory parameters, the results of blood gas analyses and pulmonary function tests, as well as the need for and duration of intensive care unit (ICU) admission and mechanical ventilation (both non-invasive [NIV] and invasive), were retrieved from the hospital information system (IS-H, SAP Deutschland SE & Co. KG, Walldorf, Germany).
Chest CT scans were performed using either a 128-slice (SOMATOM Definition Flash, SIEMENS Healthineers AG, Forchheim, Germany) or a 64-slice CT scanner (SOMATOM go.Top, SIEMENS Healthineers AG, Forchheim, Germany). Images were acquired during a single breath-hold from the lung bases to the apices. Scans performed to confirm (or exclude) suspected COVID-19 pneumonia were obtained without i. v. contrast media, while patients scanned for other clinical indications, such as suspected pulmonary embolism or oncologic evaluation, received 40–80 ml of an iodine-containing contrast agent (ACCUPAQUETM-300 or -350, GE HealthCare Buchler GmbH & Co. KG, Braunschweig, Germany) i. v. All scans utilised an automatic exposure control for tube voltage and current. Images were reconstructed using a sharp (pulmonary) and a soft (mediastinal) kernel, with lung images generated at slice thicknesses of 0.8, 1, 2, 3, or 5 mm.
CT images were independently reviewed by a radiology specialist (senior physician), an experienced assistant physician, and a medical student. Readers assessed typical COVID-19 pulmonary features, including ground-glass opacities (GGO), consolidations, and crazy-paving patterns, as well as other abnormalities such as nodules or halo signs. All evaluations were performed blinded to clinical data (laboratory values, blood gas analyses, intensive medical care, etc.) of the patients and to each other’s assessments.
No formal calibration or consensus session was conducted prior to the actual image assessment. The medical student received focused training to accurately identify lung lobes, recognise relevant pathologies, and quantify their extent. The two radiologists, with 14 and 6 years of chest CT experience, were briefed only on the specific pathologies to assess (according to the Fleischner Society’s glossary of terms) and the application of the CT scoring system. As the evaluation of the CT data began in 2023, the definitions of the radiological terms are based on the 2008 version of the Fleischner glossary [10].
Based on the semi-quantitative scoring system described by Pan et al., the extent of parenchymal involvement was assessed separately for each of the five lung lobes. Each lobe was assigned a subjective visual score ranging from 0 to 5 (Table 1). The individual lobar scores were summed to yield a total score ranging from 0 to 25.
In addition to the assessments performed by the three human readers, a fully automated score was generated for each patient using a dedicated AI-based software tool (ADVANCE chest CT, contextflow GmbH, Vienna, Austria; Figure 1). Consequently, four total scores per patient were available for comparative analysis.
The software employs a deep learning-based algorithm applied to lung kernel images to automatically segment the lungs and quantify parenchymal abnormalities. No manual adjustments to the AI-derived results were performed. According to the manufacturer, the algorithm was trained on a multicentric CT dataset; however, detailed information regarding the training process and model architecture is not publicly disclosed.
All statistical analyses and visualisations were performed using R (version 4.x; R Core Team, 2024) [11]. Descriptive statistics were calculated for all variables. Continuous variables are presented as the mean ± standard deviation (SD) or median with interquartile range (IQR), depending on their distribution, while categorical variables are reported as counts and percentages. Group comparisons between survivors and non-survivors were performed using the Wilcoxon rank sum test for continuous variables and the chi-square test or Fisher’s exact test for categorical variables, as appropriate.
Receiver operating characteristics (ROC) curve analyses were conducted to evaluate the discriminatory performance of the Pan score in predicting mortality and the need for mechanical ventilation. The area under the curve (AUC), along with the corresponding confidence intervals, was calculated using the pROC package [12]. Associations between clinical variables, mortality, mechanical ventilation, and ICU length of stay were assessed using correlation analyses based on Pearson’s correlation coefficient.
Logistic regression models were used to evaluate the association between the Pan score and clinically relevant outcomes, including mechanical ventilation and prolonged ICU stay (≥ 7 days). Both univariate and multivariable models were fitted, adjusting for potential confounders including age, sex, and laboratory parameters. Results are reported as odds ratios (OR) with 95% confidence intervals (CI).
To further explore the relationship between the Pan score and outcome probabilities, generalised linear models with a binomial distribution were fitted. To account for the potentially non-linear effect of age on the outcome, age was modelled using restricted cubic splines (with three degrees of freedom) in all multivariable regression models. This approach provides greater flexibility than assuming a linear relationship between age and the outcome and reduces the risk of residual confounding due to misspecification of the functional form. Predicted probabilities with 95% CI were derived and visualised across the range of Pan scores for cumulative outcome thresholds (e. g. ICU stay ≥ 1, ≥ 2, ≥ 3, and ≥ 7 days; ventilation duration ≥ 24, ≥ 48, ≥ 72, and ≥ 168 hours).
Data handling and transformation were performed using functions from the tidyverse collection [13]. Model outputs were processed using the broom package [14]. Visualisations were generated using ggplot2 [15], and tables were created using the knitr package [16]. Additional functionality was provided by the splines2, MASS, and nnet packages [17,18,19].
A two-sided p-value of < 0.05 was considered statistically significant.

3. Results

A total of 351 patients were included in the analysis: 223 males [63.5%], 128 females [36.5%], comprising 276 survivors and 75 non-survivors, corresponding to an overall mortality rate of 21.4%. The mean age of the cohort was 67.0 ± 15.9 years, with non-survivors being older on average than survivors (74.0 ± 11.8 vs. 65.1 ± 16.4 years; p < 0.001).
The mean Pan score (averaged across all raters) was 9.4 ± 5.3 points and was significantly higher in non-survivors than in survivors (11.7 ± 6.0 vs. 8.8 ± 5.0; p < 0.001). Consistent differences were observed across individual raters as well as the AI-based assessment.
Mechanical ventilation was required in 132 of 351 patients (37.6%), while 102 patients (29.1%) were admitted to the intensive care unit (ICU). Among those treated in the ICU, the mean length of stay was 7.0 days [3.2–12.0] and was significantly longer in non-survivors compared with survivors (10.5 [6.0–20.0] vs. 6.0 [3.0–10.0] days; p < 0.003).
Laboratory parameters showed only modest differences between survivors and non-survivors (Table 2). Laboratory parameters were not available for all patients. Specifically, erythrocyte and leukocyte counts were available for only 349 patients each, CRP values for 297 patients, and D-dimer values for only 155 patients.
Pulmonary function tests, which could only be performed in selected cases due to hygiene regulations during the COVID-19 pandemic, were conducted in only 16 patients within our cohort. Consequently, the statistical validity of these functional parameters is insufficient for meaningful analyses. Therefore, pulmonary function parameters are not presented in tabular form.
Pan scores of all four readers were aggregated to derive a single mean Pan score per patient and subsequently stratified by outcome (Figure 2). Non-survivors (n = 75) exhibited significantly higher Pan scores than survivors (n = 276), with mean values of 11.7 ± 6.0 and 8.8 ± 5.0, respectively (Wilcoxon rank sum test, p < 0.001).
However, the ability to discriminate all-cause mortality was limited, as demonstrated by ROC analysis. The Pan score (overall mean) achieved an area under the curve (AUC) of 0.639, indicating only modest predictive performance for all-cause mortality (Figure 3).
Correlation analysis demonstrated that mortality was only weakly associated with the Pan score (r = 0.22, p < 0.01), age (r = 0.23, p < 0.01), and selected laboratory parameters (leukocytes: r = 0.11, p < 0.045; CRP: r = 0.16, p = 0.007). In contrast, certain variables, such as D-dimer, showed no meaningful correlation.
The Pan score was markedly higher in patients requiring mechanical ventilation (n = 132) compared with those who were not ventilated (n = 219) (Figure 4). In the correlation matrix, the strongest association was observed between the Pan score and mechanical ventilation (r = 0.54, p < 0.01), followed by ICU length of stay (r =0.39, p < 0.01). In contrast, other markers such as CRP (r =0.12, p = 0.036) and D-dimer (0.2, p = 0.013) demonstrated considerably weaker correlations with the need for ventilation (Table 3).
ROC analysis for the prediction of mechanical ventilation using a logistic model based on the Pan score demonstrated good discriminatory performance, with an AUC of 0.804. The optimal probability threshold, determined by Youden’s index, was 0.51, yielding a sensitivity of 0.58, a specificity of 0.89, and an overall accuracy of 0.77 for predicting the need for mechanical ventilation (Figure 5).
Patients with prolonged ICU treatment (≥ 7 days) exhibited markedly higher Pan scores than those with shorter or no ICU stay (Figure 6). The Pan score demonstrated good predictive performance for an ICU stay of ≥ 7 days, with an AUC of 0.803 (95% CI: 0.733–0.873). The optimal cutoff, determined by Youden’s index, was approximately twelve points (11.9), corresponding to a sensitivity of 0.71, a specificity of 0.81, and an overall accuracy of 0.79 for identifying patients with prolonged ICU stays.
In univariate logistic regression, each one-point increase in the Pan score was associated with a 31% higher odds of requiring mechanical ventilation (p < 0.001). This association remained robust after adjustment for age, sex, CRP, D-dimer, leukocytes, and erythrocytes (p < 0.001).
Similarly, for an ICU stay ≥ 7 days, the univariate odds ratio per one-point increase in Pan score was 1.27 (95% CI: 1.19–1.36; p < 0.001), and the association persisted after multivariate adjustment (adjusted OR = 1.26, 95% CI: 1.12–1.46; p < 0.001) (Table 4).
The Pan score demonstrated a strong and consistent association with both the need for mechanical ventilation and the duration of ventilatory support. Patients requiring invasive ventilation had significantly higher Pan scores than those who were not ventilated.
In age-adjusted logistic regression analyses, the Pan score emerged as an independent predictor of mechanical ventilation across all examined duration thresholds. For ventilation lasting ≥ 24 hours, each one-point increase in Pan score was associated with a 33% increase in the odds of requiring ventilation (OR = 1.33, 95% CI: 1.24–1.44; p < 0.001). Comparable effect sizes were observed for longer ventilation durations (≥ 48 hours: OR = 1.31, 95% CI: 1.22–1.42; ≥ 72 hours: OR = 1.31, 95% CI: 1.22–1.42; ≥ 7 days: OR = 1.30, 95% CI: 1.20–1.42; all p < 0.001).
Discriminatory performance was consistently high across all thresholds, with areas under the curve ranging from 0.853 to 0.857. The highest performance was observed for ventilation ≥ 24 hours (AUC 0.857, 95% CI: 0.799–0.916), with only minimal variation across longer ventilation durations.
In an ordinal regression model treating ventilation duration as an ordered outcome (no ventilation [n = 219], < 24 hours [n = 79], ≥ 24 hours [n = 6], ≥ 48 hours [n = 5], ≥ 72 hours [n = 13], ≥ 7 days [n = 29], the Pan score showed a significant positive trend (β = 0.276, p < 0.001), indicating a stepwise increase in ventilation duration with rising Pan scores (Figure 7).
Overall, higher Pan scores at baseline CT were associated not only with a greater likelihood of subsequent mechanical ventilation but also with progressively longer durations of ventilatory support.
Predicted probabilities of ICU admission across all duration thresholds (≥ 1, ≥ 2, ≥ 3, and ≥ 7 days) increased monotonically with the baseline Pan score. Adjusted for the cohort’s median age of 69 years, a Pan score of 15 corresponded to a predicted probability of > 50% for an ICU stay of ≥ 7 days. As the duration thresholds increased, the probability curves showed a consistent rightward shift (Figure 8).
To further characterise these findings, a bivariate correlation analysis was performed between the mean Pan score across all readers and selected clinical, laboratory, and outcome parameters (Figure 9). Significant positive monotonic correlations were observed between the mean Pan score and markers of systemic inflammation and tissue injury, markedly LDH (ρ = 0.41, p < 0.001) and CRP (ρ = 0.29, p < 0.001). In addition, higher Pan scores were associated with increased healthcare resource utilisation, as reflected by longer ICU stays (ρ = 0.43, p < 0.001) and prolonged duration of mechanical ventilation (ρ = 0.43, p < 0.001). In contrast, significant negative correlations were found with respiratory parameters, including oxygen saturation (ρ = -0.24, p < 0.001) and arterial pO2 (ρ = -0.13, p = 0.002). Moreover, higher Pan scores were associated with reduced overall survival (death after PCR: ρ = -0.37, p = 0.001). No significant correlations were identified for D-dimer levels, leukocyte count, or estimated glomerular filtration rate (eGFR).

4. Discussion

The aim of this study was to investigate the extent to which the semi-quantitative CT scoring system proposed by Pan et al. [8] for assessing the severity of COVID-19 pneumonia can predict the clinical course of affected patients. This involved assessing the correlation of the score with various laboratory parameters and the patients’ pulmonary function, as well as evaluating its ability to predict the need for mechanical ventilation, admission to intensive care, and survival.
The Pan score was determined independently by three different readers with varying levels of radiological experience, as well as by an AI-based software application, as described previously [9]. For the present study, the mean score of all four readers was used for comparison with other parameters in each case. Compared with other studies investigating the outcome of COVID-19 patients or the correlation of SARS-CoV-2-associated changes on chest CT with various clinical parameters [6,7,20,21,22,23], the patient cohort examined in our study (n = 351) was considerably larger.
The overall mortality rate in our cohort was 21.4%, which is consistent with the rates reported in other studies, ranging from 15.4 to 33.1% [6,7,22,23]. The COVID-positive patients who died were, on average, significantly older than those who were discharged from hospital alive. In addition, the deceased patients showed significantly higher Pan scores (11.7 ± 6.0 vs. 8.8 ± 5.0). These findings are consistent with the observations of Li et al., who demonstrated that the mortality rate in COVID-19 infections increases with higher CT scores [23]. Although they used an earlier CT scoring system to semi-quantitatively assess the extent of pulmonary involvement, it employs the same threshold values as the Pan score and, in terms of its assessment – based on separate evaluation of different lung segments – is also comparable to it [24].
Of the laboratory parameters examined, only erythrocytes (4.4 vs. 3.9 /pl) and CRP (4.6 vs. 7.3 mg/dL) showed significant differences between survivors and non-survivors; leucocytes and D-dimer did not differ significantly.
In our study, the predictive power of the Pan score with regard to all-cause mortality was only moderately pronounced. This contradicts the findings of Francone et al., who observed that the risk of death significantly increased at a CT score of ≥ 18, and to the observations of Szabó et al., who demonstrated that death was significantly associated with higher CT scores [6,7].
By contrast, the Pan score showed good discriminatory ability in predicting the need for mechanical ventilation, and furthermore, the score was strongly correlated with the duration of mechanical ventilation. The proportion of patients requiring mechanical ventilation was significantly higher among non-survivors than among survivors (56.0 vs. 32.1%), and patients who required invasive ventilation exhibited significantly higher Pan scores. Each additional point in the Pan score was associated with a 31% increase in the odds of requiring mechanical ventilation, and with regard to ventilation duration, the score showed a significant positive trend, with a stepwise increase in the duration of mechanical ventilation as scores rose. For ventilation lasting ≥ 24 hours, each additional point in the Pan score corresponded to a 33% increase in the odds of requiring ventilation.
The Pan score also showed good predictive performance in assessing the need for intensive care admission in patients with COVID-19 pneumonia, especially for ICU stays of ≥ 7 days. Both the rate of ICU admissions (42.7 vs. 25.4%) and the length of ICU stays (10.5 vs. 6.0 days) were significantly higher among non-survivors than among survivors. At a threshold of 11.9 points, the Pan score demonstrated an accuracy of 0.79 in identifying patients requiring prolonged intensive care. After adjustment for the cohort’s median age of 69 years, a Pan score of 15 was associated with a predicted probability of > 50% or an ICU stay of ≥ 7 days.
Among the laboratory parameters examined, only weak overall correlations were observed with the endpoints of mortality, mechanical ventilation, and intensive care admission: D-dimer showed a weak correlation with the need for mechanical ventilation (r = 0.20), while CRP and leucocyte count were likewise weakly correlated with patient mortality (r = 0.16 and 0.11, respectively). Pulmonary function parameters showed no meaningful correlation with the above-mentioned endpoints. Furthermore, pulmonary function tests were performed in only a small number of patients in our cohort, precluding any reliable statistical analysis. Patient age also demonstrated only a weak correlation with mortality (r = 0.23).
The main limitation of our study is its retrospective design. As a result, we were confronted, on the one hand, with different CT protocols, as patients with a confirmed COVID-19 infection predominantly underwent non-contrast chest CT to assess the extent of pulmonary involvement [25,26], whereas others received contrast-enhanced CT scans (e. g. for oncological staging), where COVID-19 was not the primary clinical concern. We had already confirmed in a previous study the hypothesis that patients undergoing non-contrast CT for symptomatic COVID-19 infection exhibit higher Pan scores than, for example, patients in oncological follow-up [9].
On the other hand, we had to deal with the fact that the laboratory parameters of interest were partially incompletely documented, and pulmonary function testing was performed in only a small number of patients. The latter is explained by the fact that, during the COVID-19 pandemic, pulmonary function testing was generally recommended only in emergency situations for hygiene reasons [27].
In recent years, the number of COVID-19 infections has declined markedly due to established vaccination programmes and the population immunity acquired through prior infection, meaning that COVID-19-related pulmonary changes are now observed only rarely. Nevertheless, further waves of transmission are still considered possible, and severe disease courses may continue to occur, particularly among older individuals and those with certain pre-existing conditions. Furthermore, antigenic shifts remain possible, which could temporarily lead to increased viral circulation and a higher disease burden once again [2].
Therefore, the assessment of semi-quantitative CT scores and their predictive value remains relevant, as these scoring systems may be adapted in the future for evaluating other pulmonary infectious diseases. Moreover, the observation that higher Pan scores are associated with increased healthcare resource utilisation may also be relevant when further evaluating therapeutic strategies and their effectiveness.

5. Conclusions

The semi-quantitative CT score investigated is a simple and reliable diagnostic tool for assessing the extent of COVID-19 pneumonia, which can be determined both subjectively by radiologists with varying levels of experience and fully automatically by AI-based software. The score showed only moderate predictive performance for all-cause mortality but performed well in predicting the need for and duration of mechanical ventilation, as well as intensive care requirement in affected patients.

Author Contributions

Conceptualization, A.J.H. and M.E.; methodology, E.N. and A.J.H.; software, E.N.; validation, A.J.H., J.P.A. and M.E.; formal analysis, E.N.; investigation, A.M., S.T.S. and A.J.H.; resources, H.V. and M.E.; data curation, J.P.A. and A.J.H.; writing—original draft preparation, A.J.H.; writing—review and editing, E.N., H.V., J.P.A. and M.E.; visualization, E.N.; supervision, A.J.H.; project administration, A.J.H.; funding acquisition, A.J.H. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded as part of the Female Clinician Scientist Fellowship (FCSF) of the Medical School and University Medical Center OWL, Bielefeld University, Germany.

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki and approved by the Ethics Committee of the Medical Association Westphalia-Lippe and the University of Münster (2022-861-f-S; January 10, 2023).

Data Availability Statement

The data sets (chest CT images) presented in this article are not readily available due to privacy and ethical restrictions.

Acknowledgments

We would like to thank Saskia Polnau and Lubana Al Haj Hossen for their support with the inclusion of patients and the collection of clinical patient data. We would also like to thank Antonia Dresselhaus for processing and collating the collected data and for her support with the statistical analysis. Finally, we would like to thank Mia Ilic from contextflow for the AI-based analysis of the CT images and the technical support.

Conflicts of Interest

The authors declare no conflicts of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

Abbreviations

The following abbreviations are used in this manuscript:
AI Artificial intelligence
AUC Area under the curve
CI Confidence interval
COVID-19 Coronavirus disease 2019
CRP C-reactive protein
CT Computed tomography
GFR Glomerular filtration rate
GGO Ground-glass opacity
Hb Haemoglobin
ICU Intensive care unit
IQR Interquartile range
i. v. Intravenous
K Potassium
LDH Lactate dehydrogenase
LLL Left lower lobe
LUL Left upper lobe
ML Middle lobe
Na Sodium
NIV Non-invasive ventilation
OR Odds ratio
PACS Picture archiving and communication system
PCR Polymerase chain reaction
PCT Procalcitonin
pH Potential of hydrogen
pO2 Partial pressure of oxygen
PTT Partial thromboplastin time
RIS Radiological information system
RLL Right lower lobe
ROC Receiver operating characteristic
RUL Right upper lobe
SD Standard deviation
SARS-CoV-2 Severe acute respiratory syndrome coronavirus 2
WHO World Health Organization

References

  1. WHO COVID-19 dashboard. Available online: https://data.who.int/dashboards/covid19/cases?n=c (accessed on 19 March 2026).
  2. Robert Koch-Institut: RKI-Ratgeber COVID-19. Epid Bull. 2024, 22, 3–14.
  3. Vogel-Claussen, J.; Ley-Zaporozhan, J.; Agarwal, P. Recommendations of the thoracic imaging section of the German Radiological Society of clinical application of chest imaging and structured CT reporting in the COVID-19 pandemic. Rofo 2020, 192, 633–640. [Google Scholar] [CrossRef] [PubMed]
  4. Ye, Z.; Zhang, Y.; Wang, Y. Chest CT manifestations of new coronavirus disease 2019 (COVID-19): a pictorial review. Eur. Radiol. 2020, 30, 4381–4389. [Google Scholar] [CrossRef] [PubMed]
  5. Li, J.; Yan, R.; Zhai, Y. Chest CT findings in patients with coronavirus disease 2019 (COVID-19): a comprehensive review. Diagn. Interv. Radiol. 2021, 27, 621–632. [Google Scholar] [CrossRef] [PubMed]
  6. Szabó, M.; Kardos, Z.; Kostyál, L. The importance of chest CT severity score and lung CT patterns in risk assessment in COVID-19-associated pneumonia: a comparative study. Front. Med. 2023, 10, 1125530. [Google Scholar] [CrossRef]
  7. Francone, M.; Iafrate, F.; Masci, G.M. Chest CT score in COVID-19 patients: correlation with disease severity and short-term prognosis. Eur. Radiol. 2020, 30, 6808–6817. [Google Scholar] [CrossRef] [PubMed]
  8. Pan, F.; Ye, T.; Sun, P. Time course of lung changes at chest CT during recovery from coronavirus disease 2019 (COVID-19). Radiology 2020, 295, 715–721. [Google Scholar] [CrossRef] [PubMed]
  9. Neumann, E.; Movlilishvili, A.; Scherfeld, S.T. Visual and AI-based assessment of COVID-19 pneumonia: practicability and reproducibility of an established semi-quantitative chest CT scoring system. Diagnostics 1987, 2025, 15. [Google Scholar]
  10. Hansell, D.M.; Bankier, A.A.; MacMahon, H. Fleischner Society: glossary of terms for thoracic imaging. Radiology 2008, 246, 697–722. [Google Scholar] [CrossRef] [PubMed]
  11. R Core Team, The R Foundation. Available online: https://www.r-project.org (accessed on 4 May 2026).
  12. Robin, X.; Turck, N.; Hainard, A. pROC: an open-source package for R and S+ to analyse and compare ROC curves. BMC Bioinform. 2011, 12, 77. [Google Scholar]
  13. Wickham, H.; Averick, M.; Bryan, J. Welcome to the tidyverse. J. Open Source Softw. 2019, 4, 1686. [Google Scholar] [CrossRef]
  14. Robinson, D.; Hayes, A.; Couch, S. broom: convert statistical objects into tidy tibbles. Available online: https://cran.r-project.org/web/packages/broom/index.html (accessed on 4 May 2026).
  15. Wickham, H. ggplot2: elegant graphics for data analysis. Available online: https://ggplot2.tidyverse.org (accessed on 4 May 2026).
  16. Xie, Y. knitr: a general-purpose package for dynamic report generation in R. Available online: https://yihui.org/knitr (accessed on 4 May 2026).
  17. Wang, W.; Yan, J. splines2: regression spline functions and classes. Available online: https://cran.r-project.org/web/packages/splines2/index.html (accessed on 4 May 2026).
  18. Venables, W.N.; Ripley, B.D. Modern applied statistics with S, 4th edition; Springer: New York.
  19. Ripley, B.; Venables, W. nnet: feed-forward neural networks and multinomial log-linear models. Available online: https://cran.r-project.org/web/packages/nnet/index.html (accessed on 4 May 2026).
  20. Li, K.; Fang, Y.; Li, W. CT image visual quantitative evaluation and clinical classification of coronavirus disease (COVID-19). Eur. Radiol. 2020, 30, 4407–4416. [Google Scholar] [CrossRef] [PubMed]
  21. Li, Q.Y.; An, Z.Y.; Pan, Z.H. Severe/critical COVID-19 early warning system based on machine learning algorithms using novel imaging scores. World J. Clin. Cases 2023, 11, 2716–2728. [Google Scholar] [CrossRef] [PubMed]
  22. Pan, F.; Zheng, C.; Ye, T. Different computed tomography patterns of coronavirus disease 2019 (COVID-19) between survivors and non-survivors. Sci. Rep. 2020, 10, 11336. [Google Scholar] [CrossRef] [PubMed]
  23. Li, L.; Yang, L.; Gui, S. Association of clinical and radiographic findings with the outcomes of 93 patients with COVID-19 in Wuhan, China. Theranostics 2020, 10, 6113–6121. [Google Scholar] [CrossRef] [PubMed]
  24. Ooi, G.C.; Khong, P.L.; Müller, N.L. Severe acute respiratory syndrome: temporal lung changes at thin-section CT in 30 patients. Radiology 2004, 230, 836–844. [Google Scholar] [CrossRef] [PubMed]
  25. Rodrigues, J.C.L.; Hare, S.S.; Edey, A. An update on COVID-19 for the radiologist – a British Society of Thoracic Imaging statement. Clin. Radiol. 2020, 75, 323–325. [Google Scholar] [PubMed]
  26. Kwee, T.C.; Kwee, R.M. Chest CT in COVID-19: what the radiologist needs to know. Radiographics 2020, 40, 1848–1865. [Google Scholar] [CrossRef] [PubMed]
  27. Ochmann, U.; Nowak, D.; Criée, C. Recommendations for performance of lung function in times of SARS-CoV-2 pandemic. Pneumologie 2020, 74, 582–584. [Google Scholar] [PubMed]
Figure 1. Fully automated evaluation of a chest CT scan by an AI-based software tool (ADVANCE chest CT, contextflow GmbH, Vienna, Austria). (a) CT image without any software overlays, (b) automated segmentation of the lung lobes (light green, right upper lobe; turquoise, middle lobe; light blue, right lower lobe; blue, lingula; dark blue, left lower lobe), (c) AI-generated detection of circular consolidations in the right lower lobe, marked in yellow, as well as ground-glass opacities (GGO) in both lower lobes and in the middle lobe, marked in orange.
Figure 1. Fully automated evaluation of a chest CT scan by an AI-based software tool (ADVANCE chest CT, contextflow GmbH, Vienna, Austria). (a) CT image without any software overlays, (b) automated segmentation of the lung lobes (light green, right upper lobe; turquoise, middle lobe; light blue, right lower lobe; blue, lingula; dark blue, left lower lobe), (c) AI-generated detection of circular consolidations in the right lower lobe, marked in yellow, as well as ground-glass opacities (GGO) in both lower lobes and in the middle lobe, marked in orange.
Preprints 220361 g001
Figure 2. Distribution of Pan scores by outcome across all raters. Violin and box plots demonstrate consistently higher Pan scores in non-survivors (n = 75) compared with survivors (n = 276) across all rater groups (radiology specialist, assistant physician, medical student, and AI-based software tool). The greatest separation between groups was observed for the AI-based assessment.
Figure 2. Distribution of Pan scores by outcome across all raters. Violin and box plots demonstrate consistently higher Pan scores in non-survivors (n = 75) compared with survivors (n = 276) across all rater groups (radiology specialist, assistant physician, medical student, and AI-based software tool). The greatest separation between groups was observed for the AI-based assessment.
Preprints 220361 g002
Figure 3. ROC curves for mortality prediction based on Pan scores are shown. Overall, discriminatory performance was limited, with area under the curve (AUC) values ranging from low to moderate. The highest AUC reached 0.639, indicating only modest ability to differentiate between survivors and non-survivors.
Figure 3. ROC curves for mortality prediction based on Pan scores are shown. Overall, discriminatory performance was limited, with area under the curve (AUC) values ranging from low to moderate. The highest AUC reached 0.639, indicating only modest ability to differentiate between survivors and non-survivors.
Preprints 220361 g003
Figure 4. Distribution of Pan scores stratified by mechanical ventilation status. Patients requiring mechanical ventilation (n = 132) showed substantially higher Pan scores compared with non-ventilated patients (n = 219; Wilcoxon rank sum test, p < 0.001).
Figure 4. Distribution of Pan scores stratified by mechanical ventilation status. Patients requiring mechanical ventilation (n = 132) showed substantially higher Pan scores compared with non-ventilated patients (n = 219; Wilcoxon rank sum test, p < 0.001).
Preprints 220361 g004
Figure 5. ROC curve for the prediction of mechanical ventilation based on the Pan score.
Figure 5. ROC curve for the prediction of mechanical ventilation based on the Pan score.
Preprints 220361 g005
Figure 6. Distribution of Pan scores according to ICU length of stay (< 7 vs. ≥ 7 days). Patients with prolonged ICU treatment (≥ 7 days) exhibited markedly higher Pan scores than those with shorter ICU stays, in line with the findings from the logistic regression analysis.
Figure 6. Distribution of Pan scores according to ICU length of stay (< 7 vs. ≥ 7 days). Patients with prolonged ICU treatment (≥ 7 days) exhibited markedly higher Pan scores than those with shorter ICU stays, in line with the findings from the logistic regression analysis.
Preprints 220361 g006
Figure 7. Predicted probabilities of mechanical ventilation across different duration thresholds as a function of the Pan score. The plot shows estimated probabilities for ventilation > 0, ≥ 24, ≥ 48, ≥ 72, and ≥ 168 hours (7 days). Predictions were derived from separate logistic regression models including Pan score and age, with age fixed at the cohort median (69 years). Shaded areas indicate 95% confidence intervals. As the thresholds are cumulative (e. g., ≥ 168 hours is a subset of ≥ 72 hours), the curves are not additive.
Figure 7. Predicted probabilities of mechanical ventilation across different duration thresholds as a function of the Pan score. The plot shows estimated probabilities for ventilation > 0, ≥ 24, ≥ 48, ≥ 72, and ≥ 168 hours (7 days). Predictions were derived from separate logistic regression models including Pan score and age, with age fixed at the cohort median (69 years). Shaded areas indicate 95% confidence intervals. As the thresholds are cumulative (e. g., ≥ 168 hours is a subset of ≥ 72 hours), the curves are not additive.
Preprints 220361 g007
Figure 8. Predicted probabilities of ICU length of stay across different duration thresholds as a function of the Pan score. The figure depicts estimated probabilities of an ICU stay of 0, ≥ 1, ≥ 2, ≥ 3, and ≥ 7 days. Predictions were derived from separate logistic regression models including Pan score and age, with age fixed at the cohort median (69 years). Shaded bands indicate 95% confidence intervals. As the thresholds are cumulative (e. g., ≥ 7 days is a subset of ≥ 3 days), the curves are not additive. additive.
Figure 8. Predicted probabilities of ICU length of stay across different duration thresholds as a function of the Pan score. The figure depicts estimated probabilities of an ICU stay of 0, ≥ 1, ≥ 2, ≥ 3, and ≥ 7 days. Predictions were derived from separate logistic regression models including Pan score and age, with age fixed at the cohort median (69 years). Shaded bands indicate 95% confidence intervals. As the thresholds are cumulative (e. g., ≥ 7 days is a subset of ≥ 3 days), the curves are not additive. additive.
Preprints 220361 g008
Figure 9. Spearman rank correlation coefficients (ρ) between clinical parameters and the mean Pan score. Bars represent the strength and direction of the monotonic association, ordered from the strongest positive (dark green) to the strongest negative correlation (dark red). Statistical significance is indicated as *** p < 0.001, ** p < 0.01, * p < 0.05, and ns = not significant. Strong positive correlations were observed for LDH, ICU length of stay, duration of mechanical ventilation, and CRP, whereas respiratory parameters (oxygen saturation and pO2) and overall survival showed the strongest inverse correlations. LDH, lactate dehydrogenase; ICU, intensive care unit; CRP, C-reactive protein; PCT, procalcitonin; PTT, partial thromboplastin time; GFR, glomerular filtration rate; K, potassium; pH, potential of hydrogen; Na, sodium; Hb, haemoglobin; pO2, partial pressure of oxygen; PCR, polymerase chain reaction.
Figure 9. Spearman rank correlation coefficients (ρ) between clinical parameters and the mean Pan score. Bars represent the strength and direction of the monotonic association, ordered from the strongest positive (dark green) to the strongest negative correlation (dark red). Statistical significance is indicated as *** p < 0.001, ** p < 0.01, * p < 0.05, and ns = not significant. Strong positive correlations were observed for LDH, ICU length of stay, duration of mechanical ventilation, and CRP, whereas respiratory parameters (oxygen saturation and pO2) and overall survival showed the strongest inverse correlations. LDH, lactate dehydrogenase; ICU, intensive care unit; CRP, C-reactive protein; PCT, procalcitonin; PTT, partial thromboplastin time; GFR, glomerular filtration rate; K, potassium; pH, potential of hydrogen; Na, sodium; Hb, haemoglobin; pO2, partial pressure of oxygen; PCR, polymerase chain reaction.
Preprints 220361 g009
Table 1. A semi-quantitative chest CT scoring system, as described by Pan et al. [8], to assess the proportion of lung parenchyma affected by pathological changes in COVID-19 pneumonia.
Table 1. A semi-quantitative chest CT scoring system, as described by Pan et al. [8], to assess the proportion of lung parenchyma affected by pathological changes in COVID-19 pneumonia.
Score Extent of Involvement of Each Lung Lobe
0 no involvement
1 <5%
2 5–25%
3 26–49%
4 50–75%
5 >75%
Table 2. Values of the different variables in the overall cohort, the cohort of survivors, and the cohort of non-survivors. Data are presented as mean ± standard deviation or as median with interquartile range (IQR; in square brackets). Corresponding p-values are also reported. AI, artificial intelligence; ICU, intensive care unit; CRP, C-reactive protein.
Table 2. Values of the different variables in the overall cohort, the cohort of survivors, and the cohort of non-survivors. Data are presented as mean ± standard deviation or as median with interquartile range (IQR; in square brackets). Corresponding p-values are also reported. AI, artificial intelligence; ICU, intensive care unit; CRP, C-reactive protein.
Variable Overall Cohort
(n = 351)
Survivors
(n = 276)
Non-Survivors
(n = 75)
p-Value
Age [years] 67.0 ± 15.9 65.1 ± 16.4 74.0 ± 11.8 <0.001
Pan score
 total
 radiology specialist
 assistant physician
 medical student
 AI

9.4 ± 5.3
8.1 ± 5.7
11.2 ± 6.1
8.0 ± 6.4
10.3 ± 4.9

8.8 ± 5.0
7.4 ± 5.2
10.8 ± 5.9
7.2 ± 5.8
9.7 ± 4.6

11.7 ± 6.0
10.7 ± 6.6
12.5 ± 6.5
11.0 ± 7.7
12.5 ± 5.4

< 0.001
< 0.001
0.062
< 0.001
< 0.001
ICU
 rate [%]
length of stay [days]

29.1
7.0 [3.2–12.0]

25.4
6.0 [3.0–10.0]

42.7
10.5 [6.0–20.0]

0.002
0.003
Mechanical ventilation rate [%] 37.6 32.1 56.0 < 0.001
Laboratory parameters
 Erythrocytes [/pl]
 Leukocytes [/nl]
 CRP [mg/dL]
 D-dimer [mg/L]

4.4 [3.7–4.9]
7.5 [5.6–10.4]
5.0 [2.2–11.4]
1.5 [0.8–3.5]

4.4 [3.9–4.9]
7.5 [5.7–10.1]
4.6 [2.1–10.2]
1.5 [0.8–3.5]

3.9 [3.4–4.7]
7.9 [5.2–10.9]
7.3 [2.7–15.4]
2.3 [0.9–3.4]

0.001
0.558
0.016
0.554
Table 3. Correlations between clinical variables and mortality, mechanical ventilation, ICU length of stay. The Pan score showed the strongest associations with mechanical ventilation (r = 0.54), ICU stay (r = 0.39), and mortality (r = 0.22). Age was primarily correlated with mortality (r = 0.23). CRP and leukocytes demonstrated weaker associations, whereas D-dimer showed no meaningful correlation with mortality or ICU stay. ICU, intensive care unit; CRP, C-reactive protein.
Table 3. Correlations between clinical variables and mortality, mechanical ventilation, ICU length of stay. The Pan score showed the strongest associations with mechanical ventilation (r = 0.54), ICU stay (r = 0.39), and mortality (r = 0.22). Age was primarily correlated with mortality (r = 0.23). CRP and leukocytes demonstrated weaker associations, whereas D-dimer showed no meaningful correlation with mortality or ICU stay. ICU, intensive care unit; CRP, C-reactive protein.
Variable Death Ventilation ICU Stay
Correlation p-Value Correlation p-Value Correlation p-Value
Age 0.23 < 0.01 0.026 0.63 -0.02 0.72
Biological sex 0.11 0.047 -0.05 0.38 0.02 0.70
Pan score [total mean] 0.22 < 0.01 0.54 < 0.01 0.39 < 0.01
Laboratory
parameters
 Erythrocytes
 Leukocytes
 CRP
 D-dimer

-0.17
0.107
0.16
0.06

0.001
0.045
0.007
0.43

-0.068
0.078
0.12
0.2

0.2
0.15
0.036
0.013

-0.15
-0.01
0.1
0

< 0.01
0.8
0.1
1
Table 4. Associations between the Pan score and clinical outcomes (logistic regression). Each one-point increase in the Pan score was associated with higher odds of mechanical ventilation as well as prolonged ICU stay. All associations were statistically significant (p < 0.001).
Table 4. Associations between the Pan score and clinical outcomes (logistic regression). Each one-point increase in the Pan score was associated with higher odds of mechanical ventilation as well as prolonged ICU stay. All associations were statistically significant (p < 0.001).
Outcome Odds Ratio (CI) p-Value
Mechanical ventilation
 univariate
 adjusted

1.31 (1.23–1.39)
1.26 (1.15–1.41)

< 0.001
< 0.001
ICU stay ≥ 7 days
 univariate
 adjusted

1.27 (1.19–1.36)
1.26 (1.12–1.46)

< 0.001
< 0.001
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.