Preprint
Article

This version is not peer-reviewed.

Deep Temporal-Sepsis: A Calibrated Transformer Framework for Early ICU Sepsis Prediction with Grad-CAM Temporal Phenotyping

A peer-reviewed version of this preprint was published in:
BioMedInformatics 2026, 6(4), 40. https://doi.org/10.3390/biomedinformatics6040040

Submitted:

12 June 2026

Posted:

12 June 2026

You are already at the latest version

Abstract
Background/Objectives: Sepsis is responsible for approximately 270,000 deaths annually in the United States. Conventional scoring systems, such as SOFA and qSOFA, are largely reactive and do not effectively leverage longitudinal ICU data for early prediction. This study aimed to develop a deep learning framework capable of predicting sepsis onset up to 6 hours before Sepsis-3 criteria are met, while also providing clinically interpretable temporal explanations. Methods: The PhysioNet/CinC 2019 Challenge dataset, comprising 1,552,210 patient-hours from 40,336 ICU patients, was utilized. A Temporal Transformer Encoder (TTE) was trained using 12-hour look-back windows with 92 engineered features. Severe class imbalance (2.6% positive rate) was addressed through weighted random sampling and focal loss. Five-fold patient-level cross-validation was employed to prevent temporal leakage. Platt scaling was applied for probability calibration. Grad-CAM was adapted for temporal explainability, while SHAP was used for feature-level attribution. BiLSTM-Attention and XGBoost models served as baseline comparators. Results: The TTE model achieved a cross-validated AUROC of 0.8320±0.0032 and an AUPRC of 0.1505±0.0148, significantly outperforming BiLSTM Attention (AUROC: 0.7859) and XGBoost (AUROC: 0.7731; DeLong p < 0.0001). Platt scaling reduced the Expected Calibration Error from 0.3154 to 0.0017. The median alert lead time was 46.5 hours (IQR: 21–84 h), with 95.3% of septic patients receiving alerts at least 3 hours before onset. Grad-CAM analysis identified timesteps t − 10 and t −9 as the most predictive. However, high-severity patients (SOFA proxy ≥ 3)demonstrated substantially reduced performance (AUROC: 0.257). Conclusions: The proposed TTE framework demonstrates strong and well-calibrated early sepsis prediction with substantial clinical lead time. The concentration of predictive signals 10–11 hours prior to alert generation supports the feasibility of continuous automated ICU monitoring from admission onward. Reduced performance in high-severity patients highlights the need for severity-stratified modelling in future research.
Keywords: 
;  ;  ;  ;  ;  ;  ;  

1. Introduction

1.1. Background and Clinical Motivation

Sepsis is one of the most critical health crises faced by humanity in the twenty-first century. Sepsis is “a life-threatening organ dysfunction caused by a dysregulated host response to infection” as proposed by Sepsis-3 [1]. An estimated 48.9 million people worldwide are affected by sepsis every year, resulting in 11 million deaths annually [2]. Patients with sepsis make up only 10–15% of all ICU admissions but consume a disproportionate amount of resources [3].
The primary difficulty in treating sepsis is time: each delayed hour of appropriate antibiotic treatment increases mortality risk by 7% [4]. However, it is only after considerable organ damage has occurred that doctors are able to identify sepsis through SOFA or qSOFA criteria, making it impossible to be proactive instead of reactive.
Currently, Electronic Health Records (EHR) used by ICUs contain a high volume of data, including hourly updates, irregularly time-spaced laboratory results, and free-text nursing notes. These represent rich data sources for developing predictive models but at the same time pose technical challenges for developing early warning systems using machine learning.

1.2. Limitations of Existing Methods

Existing clinical scoring systems such as SOFA [1], qSOFA [5], NEWS2 [6], and MEWS [7] are point-in-time systems that do not leverage the longitudinal nature of ICU data. They are unable to capture the progression of deterioration, are subject to noise in individual measurements, and are not probabilistic in nature. Rule-based systems like SIRS have been found to have low sensitivity and specificity in prospective studies [8].
Machine learning systems have advantages over rule-based systems but also have their own limitations. In logistic regression and random forest systems such as those proposed by Horng et al. [9] and Churpek et al. [10], each measurement is used independently without considering the temporal nature of the data. The trend of increasing lactate levels, for example, is more informative than any single reading.
Transformer architectures [13], originally developed for natural language processing, offer principled handling of long-range temporal dependencies via self-attention, processing all time-steps simultaneously with no gradient degradation. Their application to clinical time-series, particularly regarding interpretability, remains underexplored.

1.3. Related Work

Prior machine learning-based sepsis prediction methods used logistic regression on engineered EHR features [9]. Churpek et al. demonstrated gradient boosted machines outperforming conventional scoring on large retrospective cohorts [10]. Shashikumar et al. presented the 1D-CNN model AISE reporting AUROC 0.831 on MIMIC [14].
Recurrent architectures were popularised by Lipton et al., who applied LSTM to clinical time series with missing-value imputation, reporting AUROC 0.82 for ICU mortality [11]. GRU-D modelled missing values via a learnable decay [15]; Moor et al. applied it for early sepsis, reporting AUROC 0.848 on PhysioNet [16]. The PhysioNet/CinC 2019 Challenge [17] established ensemble methods as strong baselines (winning utility score: 0.441). Transformer-based EHR models include BEHRT [19], SAnD [20], and Med-BERT [21]. SHAP [24] provides model-agnostic attribution; Grad-CAM [25] has been applied to ECG classification [26] but rarely to Transformers on multivariate clinical sequences.
Key unaddressed gaps: (1) most studies use single train–test splits rather than patient-level cross-validation; (2) calibration is rarely reported; (3) temporal attribution identifying which hours drive an alert is seldom analysed; (4) alert lead-time distributions are rarely characterised. This work directly addresses all four.

1.4. Contributions

  • A Temporal Transformer Encoder (TTE) achieving AUROC 0.8320 ± 0.0032 under 5-fold patient-level cross-validation on PhysioNet 2019—5.9% above XGBoost and 4.6% above BiLSTM.
  • A novel temporal Grad-CAM method for sequence Transformers, revealing prediction importance concentrated at hours t 10 and t 9 , identifying the 12-hour clinical prodrome.
  • The first reported median alert lead time of 46.5 h on PhysioNet 2019, with 95.3% of septic patients alerted ≥3 h before clinical onset.
  • A complete post-hoc calibration pipeline reducing ECE from 0.3154 to 0.0017 via Platt scaling.
  • Subgroup analysis revealing dramatically reduced performance for high-severity patients (SOFA proxy ≥3, AUROC = 0.257), motivating stratified modelling.

2. Materials and Methods

2.1. Data Source and Population

This work uses the PhysioNet/CinC 2019 Sepsis Early Detection Challenge dataset [17] (physionet.org), containing hourly data from 40,336 patients (20,336 from Training Set A; 20,000 from Training Set B) admitted to ICUs at two US academic medical centres (Beth Israel Deaconess Medical Center and Emory University Hospital), totalling 1,552,210 patient-hour data points stored as pipe-separated values.
Sepsis labels follow Sepsis-3 criteria: suspected infection (antibiotics administered and body fluids cultured) plus a ≥2-point rise in SOFA score. The overall hourly sepsis rate is 1.80%; 7.3% of patients (2,932 of 40,336) develop sepsis during their ICU stay. ICU stay lengths are right-skewed with a median of 40 hours and a long tail extending beyond 200 hours.
Figure 1. Data characterisation: ICU stay distribution (left), sepsis onset hour distribution (centre), and patient-level class imbalance (right). Sepsis onset peaks within the first 5 hours of admission; the dataset exhibits severe class imbalance with 37,404 non-septic versus 2,932 septic patients.
Figure 1. Data characterisation: ICU stay distribution (left), sepsis onset hour distribution (centre), and patient-level class imbalance (right). Sepsis onset peaks within the first 5 hours of admission; the dataset exhibits severe class imbalance with 37,404 non-septic versus 2,932 septic patients.
Preprints 218225 g001

2.2. Feature Space

The dataset contains 40 raw clinical variables across four categories. After feature engineering, the final feature matrix contains 92 variables as detailed in Table 1.

2.3. Label Construction

We define a prospective 6-hour early label: a patient-hour is labelled positive if the Sepsis-3 threshold is crossed within the next 6 hours, targeting the clinical objective of alerting early enough for intervention. The resulting label prevalence is 2.56%.

2.4. Preprocessing Pipeline

The raw data poses several challenges that require systematic preprocessing. The vital sign data has missing values due to various reasons including disconnected sensors, patient transportation, and delayed recording. Laboratory data also has high sparsity; variables such as Bilirubin_direct and Fibrinogen have more than 98% missing values.
The following preprocessing steps are applied:
  • Step 1: Missingness Encoding. A binary feature is created for each laboratory feature to account for the clinical importance of an unrecorded test result.
  • Step 2: Forward Filling for Vital Signs. Vital sign data is forward-filled for each patient stay, reflecting clinical practice in which the last value serves as a reasonable estimate of the current state.
  • Step 3: Median Imputation. Remaining missing values are imputed using the median of each feature, computed from the training set and applied consistently to validation and test sets.
  • Step 4: Standardisation. All features are normalised using z-score standardisation:
    x = x μ σ
    where μ and σ represent the mean and standard deviation computed from the training data.
  • Step 5: NaN Clamping. After scaling, any NaN or Inf values arising from zero-variance cases are replaced with 0.
  • Step 6: Feature Engineering. Four composite features are derived: Shock Index (HR/SBP), Pulse Pressure (SBP−DBP), Temperature Deviation ( | Temp 37 C | ), and a Simplified SOFA Score computed using binary thresholds for creatinine, bilirubin, platelets, MAP, respiratory rate, and WBC. Additionally, temporal features are generated using a 3-hour rolling window (mean, standard deviation, and first-order difference) for seven key variables: HR, MAP, Resp, Temp, O 2 Sat, Lactate, and WBC.

2.5. Data Splitting and Imbalance Handling

All data partitioning is performed at the patient level to prevent temporal leakage. Patients are randomly assigned to training, validation, and test sets in the ratio 70%, 15%, and 15%, respectively, yielding 28,235/6,050/6,051 patients and 1,085,255/230,301/236,654 patient-hours per split.
Class imbalance is addressed through weighted random sampling during training, where each sequence is assigned a weight inversely proportional to its class frequency. Focal Loss [27] is also used as the loss function:
FL ( p t ) = α ( 1 p t ) γ log p t
with α = 0.25 and γ = 2.0 , as recommended in the original paper.

2.6. Proposed Architecture: Temporal Transformer Encoder (TTE)

2.6.1. Input Representation

Given a patient at hour t, we construct a 12-hour look-back sequence
X = [ x t 11 , x t 10 , , x t ]
where each x i R 92 is a 92-dimensional feature vector. The sequence window SEQ _ LEN = 12 is selected to capture the acute prodromal phase of sepsis while remaining computationally tractable. The target label y t indicates whether sepsis onset occurs within the subsequent 6 hours.

2.6.2. Model Architecture

The TTE comprises four main components: (a) input projection, (b) positional encoding, (c) Transformer encoder layers, and (d) a classification head.
Input Projection.
A linear layer maps the input features into a d model = 128 dimensional embedding space:
H 0 = X W proj + b proj , W proj R 92 × 128
Positional Encoding.
Sinusoidal positional encodings are added to retain temporal order:
PE ( pos , 2 i ) = sin pos 10000 2 i / d model , PE ( pos , 2 i + 1 ) = cos pos 10000 2 i / d model
Transformer Encoder.
Three encoder layers are stacked, each with 8 attention heads, feed-forward dimension 256, dropout 0.2, and a pre-norm architecture (LayerNorm applied before each sub-layer for improved training stability). The multi-head self-attention mechanism is:
Attention ( Q , K , V ) = softmax Q K d k V
where Q, K, and V are linear projections of the hidden representations. With 8 heads and d model = 128 , each head operates in dimension d k = 16 .
Classification Head.
Adaptive average pooling is applied over the temporal dimension, followed by LayerNorm, a fully connected layer with 64 units and GELU activation, dropout regularisation, and a final output neuron:
y ^ = σ W 2 GELU W 1 AvgPool ( H L ) + b 1 + b 2
where σ ( · ) denotes the sigmoid activation. The complete model contains 417,921 trainable parameters, deliberately compact to mitigate overfitting given the 2.6% positive rate.

2.7. Training Protocol

The model is trained using the AdamW optimiser [28] with a learning rate of 5 × 10 5 and weight decay of 1 × 10 4 . A cosine annealing schedule is applied after a 5-epoch linear warm-up phase. Early stopping with patience 15 on validation AUROC is used to avoid overfitting, and gradient clipping with maximum norm 1.0 prevents exploding gradients. Hyperparameter configuration is summarised in Table 2.

2.8. Baseline Models

BiLSTM with Temporal Attention: A 2-layer bidirectional LSTM with hidden size 128, followed by a temporal attention layer where a weighted sum of hidden states is computed, representing the standard deep learning approach for clinical time series.
XGBoost: A gradient-boosted decision ensemble operating on flat feature vectors, using GPU-accelerated histogram-based feature construction, scale_pos_weight for class imbalance, and early stopping on validation AUPRC.

2.9. Grad-CAM for Temporal Explainability

We adapt Grad-CAM [25] to Transformer sequence models by hooking the final encoder layer output. For sequence input X of shape ( 1 , T , d model ) , the Grad-CAM score at timestep t is:
L GradCAM ( t ) = ReLU k α k A k ( t ) ,
where A k ( t ) is the k-th channel of the final encoder activation at time t, and
α k = mean t y ^ A k ( t ) .
ReLU retains only positive contributions; the resulting T-dimensional vector is min-max normalised to [ 0 , 1 ] for visualisation. Input-gradient saliency S ( t , f ) = | x ( t , f ) · y ^ / x ( t , f ) | produces a ( T × F ) feature-by-time importance matrix.

3. Results

3.1. Overall Performance

Evaluation results on the held-out test set are displayed in Table 3. Figure 2 shows the full evaluation dashboard including training AUROC curves, ROC curves, precision-recall curves, confusion matrix, XGBoost feature importance, and focal loss curves.

3.2. Calibration Analysis

Calibration—the match between predicted probabilities and empirical event rates—is necessary for clinical deployment. The raw TTE suffers from severe overconfidence (ECE = 0.3154), a known consequence of focal loss training [29]. Platt scaling reduces the ECE to 0.0017, making TTE one of the best-calibrated clinical AI models reported. The intermediate ECE of 0.1910 for XGBoost aligns with the known calibration behaviour of tree ensembles.
Figure 3. Calibration curves (left) and Expected Calibration Error (right) for TTE raw, TTE after Platt scaling, and XGBoost. Platt scaling reduces TTE ECE from 0.3154 to 0.0017, achieving near-perfect calibration.
Figure 3. Calibration curves (left) and Expected Calibration Error (right) for TTE raw, TTE after Platt scaling, and XGBoost. Platt scaling reduces TTE ECE from 0.3154 to 0.0017, achieving near-perfect calibration.
Preprints 218225 g003

3.3. Alert Lead-Time Analysis

Figure 4 presents the alert lead-time distribution and cumulative early-alert coverage across true-positive septic patients in the test set.
The median alert lead time of 46.5 h (IQR: 21–84 h) substantially exceeds the 3–6 h target typically specified in clinical AI deployment guidelines. The cumulative coverage curve shows 95% of septic patients receive an alert ≥3 h before onset, 93% at ≥6 h, and 84% at ≥12 h, compatible with overnight ICU monitoring scenarios.

3.4. Subgroup Analysis

Table 4 and Figure 5 present subgroup AUROCs stratified by age, sex, and severity of illness.
The dramatic fall in AUROC for high-severity patients (SOFA proxy ≥3, AUROC = 0.257) is clinically important: patients with pre-existing high SOFA scores have an abnormal sepsis phenotype in which the deterioration signal is confounded by pre-existing organ dysfunction. The model should therefore be deployed primarily for patients not in critical illness at admission, with stratified models for high SOFA patients investigated in future work.

3.5. SHAP Feature Importance

Figure 6 and Figure 7 present SHAP beeswarm plots for the XGBoost model and the Transformer-aligned XGBoost model, respectively.
SHAP analysis reveals that ICULOS (ICU length of stay) is the most important feature, clinically intuitive since sepsis risk accumulates over the ICU stay. Oxygen saturation and its rate of change are also strong predictors, reflecting respiratory compromise as a hallmark of sepsis. The sofa_proxy feature has high positive SHAP values, validating alignment with the Sepsis-3 definition. Temperature variation, HR/MAP ratio, and lactate variation indicate haemodynamic and thermoregulatory instability consistent with SIRS.

3.6. Grad-CAM Temporal Analysis

Figure 8 and Figure 9 present individual and population-level Grad-CAM analyses, respectively.
The most clinically relevant finding is that the model’s highest Grad-CAM scores are concentrated at timesteps t 10 and t 9 —approximately 10–11 hours before the prediction time—rather than at the most recent observations. This counter-intuitive result reflects the fact that subtle physiological changes preceding sepsis begin long before acute deterioration becomes apparent via conventional scoring.
For Patient 78 (septic, pred = 0.68), the saliency heatmap implicates WBC at the final timestep as the dominant feature, likely reflecting sustained leukocytosis. For Patient 283 (septic, pred = 0.39), a brief high-importance window at t 10 to t 9 centres on temperature and HR_mean3, consistent with febrile tachycardia in early infection.
At the population level (Figure 9), the mean Grad-CAM profile shows a strong peak at t 10 and t 9 , gradually decreasing towards t 0 . This gradient suggests that early-phase physiological changes are more discriminative than noisy measurements immediately preceding the alert, reinforcing the clinical case for continuous automated monitoring from the time of ICU admission.

4. Discussion

4.1. Ablation Study

Table 5 presents a systematic ablation study on the validation set, isolating the contribution of each architectural and training decision.
The most impactful components are the WeightedRandomSampler (−0.064) and focal loss (−0.051), confirming the dominance of class imbalance handling. The 12-hour sequence length validates the hypothesis that the entire prodromal window is necessary (−0.028 for SEQ_LEN=6). Three encoder layers outperform one by −0.032; adding a fourth yields no significant improvement. Rolling statistical features contribute −0.016, indicating that feature engineering still plays a meaningful role alongside end-to-end representation learning.

4.2. Theoretical Justification

4.2.1. Why Self-Attention Outperforms Recurrence for Clinical Time Series

The fundamental limitation of LSTM and GRU for clinical time series is sequential processing: information at time step t must pass through all intermediate hidden states to influence t + k , leading to information loss via vanishing gradients for long-range dependencies such as rising creatinine levels 8 hours before acute haemodynamic instability.
Self-attention circumvents this by computing direct pairwise relationships between all timesteps simultaneously:
A ( t i , t j ) = softmax q i k j d k ,
with no decay over distance—uniquely suited to sepsis, where the clinically relevant signal may manifest hours before the most recent observation.

4.2.2. Optimisation Landscape Under Severe Imbalance

With only 2.6% positive examples, binary cross-entropy gravitates towards the trivial zero-probability solution for the negative class. The focal loss modulation term ( 1 p t ) γ down-weights easy examples: at γ = 2 , an example with p t = 0.9 receives 100 × less gradient weight than one with p t = 0.1 . The ablation confirms this empirically: BCE achieves AUPRC 0.0891 versus 0.1686 for focal loss, an improvement of 89%.

5. Conclusions

This work introduces a comprehensive Temporal Transformer Encoder architecture for early sepsis detection among ICU patients, validated on the largest publicly available benchmark dataset for this task. Our method achieves a five-fold cross-validated AUROC of 0.8320 ± 0.0032 , a median clinical alert lead time of 46.5 hours, and 95.3% coverage at the 3-hour threshold, demonstrating strong performance with full methodological rigour including cross-validation, calibration, and statistical significance testing.
The Grad-CAM temporal analysis—the primary methodological contribution of this work—reveals that the Transformer learns to recognise sepsis through a temporally distributed pattern centred at t 10 to t 9 hours within the 12-hour observation window. This early physiological signal precedes clinically evident deterioration as defined by conventional scoring systems, underlining the importance of continuous automated physiological monitoring from the time of ICU admission rather than episodic assessment.
A key limitation identified through subgroup analysis is poor performance for high-severity patients (SOFA proxy ≥3, AUROC = 0.257). Future work will investigate stratified modelling with specialised models for high-SOFA patients, integration of unstructured clinical text via Clinical BERT embeddings, and extension to multi-outcome prediction (septic shock, in-hospital mortality) via multi-task learning in a human-in-the-loop deployment environment.

Author Contributions

Md Shahnawaj: Conceptualization, Methodology, Software, Validation, Formal Analysis, Investigation, Data Curation, Writing—Original Draft Preparation, Visualization; Hamim Islam Hellol: Methodology, Validation, Writing—Review & Editing; Mohammad Hasibul Hasan: Validation, Writing—Review & Editing; Roise Uddin: Investigation, Writing—Review & Editing; Novera Mahjabin Hossain: Data Curation, Writing—Review & Editing; Sumaia Benta Arif: Writing—Review & Editing; Shamim Akhtar: Conceptualization, Resources, Supervision; Md Nafis Azad Nobel: Writing—Review & Editing. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Ethical review and approval were waived for this study as the PhysioNet 2019 dataset consists entirely of de-identified retrospective data released under an approved data-use agreement by the Beth Israel Deaconess Medical Center and Emory University Hospital. No direct patient involvement occurred.

Data Availability Statement

The PhysioNet/CinC 2019 Sepsis Early Detection Challenge dataset is publicly available at https://physionet.org/content/challenge-2019/1.0.0/ under a credentialled data-use agreement. Code and trained model weights will be released upon acceptance.

Acknowledgments

The authors are grateful to the PhysioNet/CinC 2019 Challenge organisers, as well as Beth Israel Deaconess Medical Center and Emory University Hospital, for provision of the clinical data. The PhysioNet Sepsis Early Detection Challenge dataset was accessed via the PhysioNet credentialled health data repository [17]; all applicable data usage agreements and IRB exemption provisions for de-identified retrospective data analysis were followed. Computational experiments were carried out on the Kaggle cloud computing platform using NVIDIA Tesla T4 GPU resources provided free of charge via the Kaggle Research Accelerator programme.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AUROC Area Under the Receiver Operating Characteristic Curve
AUPRC Area Under the Precision-Recall Curve
ECE Expected Calibration Error
EHR Electronic Health Record
GRU Gated Recurrent Unit
ICU Intensive Care Unit
IQR Interquartile Range
LSTM Long Short-Term Memory
MAP Mean Arterial Pressure
MEWS Modified Early Warning Score
NEWS2 National Early Warning Score 2
qSOFA Quick Sequential Organ Failure Assessment
SHAP SHapley Additive exPlanations
SIRS Systemic Inflammatory Response Syndrome
SOFA Sequential Organ Failure Assessment
TTE Temporal Transformer Encoder

References

  1. Singer, M.; Deutschman, C. S.; Seymour, C. W.; Shankar-Hari, M.; Annane, D.; Bauer, M.; Angus, D. C. The Third International Consensus Definitions for Sepsis and Septic Shock (Sepsis-3). JAMA 2016, 315(8), 801–810. [Google Scholar] [CrossRef] [PubMed]
  2. Rudd, K. E.; Johnson, S. C.; Agesa, K. M.; Shackelford, K. A.; Tsoi, D.; Kievlan, D. R.; Naghavi, M. Global, regional, and national sepsis incidence and mortality, 1990–2017: analysis for the Global Burden of Disease Study. Lancet 2020, 395(10219), 200–211. [Google Scholar] [CrossRef] [PubMed]
  3. Fleischmann-Struzek, C.; Goldfarb, D. M.; Schlattmann, P.; Schlapbach, L. J.; Reinhart, K.; Kissoon, N. The global burden of paediatric and neonatal sepsis. Lancet Respir. Med. 2018, 6(3), 223–230. [Google Scholar] [CrossRef] [PubMed]
  4. Kumar, A.; Roberts, D.; Wood, K. E.; Light, B.; Parrillo, J. E.; Sharma, S.; Cheang, M. Duration of hypotension before initiation of effective antimicrobial therapy is the critical determinant of survival in human septic shock. Crit. Care Med. 2006, 34(6), 1589–1596. [Google Scholar] [CrossRef] [PubMed]
  5. Seymour, C. W.; Liu, V. X.; Iwashyna, T. J.; Brunkhorst, F. M.; Rea, T. D.; Scherag, A.; Angus, D. C. Assessment of Clinical Criteria for Sepsis. JAMA 2016, 315(8), 762–774. [Google Scholar] [CrossRef] [PubMed]
  6. Royal College of Physicians. (2017). National Early Warning Score (NEWS) 2. RCP. RC.
  7. Subbe, C. P.; Kruger, M.; Rutherford, P.; Gemmel, L. Validation of a modified Early Warning Score in medical admissions. QJM 2001, 94(10), 521–526. [Google Scholar] [CrossRef] [PubMed]
  8. Bone, R. C.; Balk, R. A.; Cerra, F. B.; Dellinger, R. P.; Fein, A. M.; Knaus, W. A.; Sibbald, W. J. Definitions for sepsis and organ failure and guidelines for the use of innovative therapies in sepsis. Chest 1992, 101(6), 1644–1655. [Google Scholar] [CrossRef] [PubMed]
  9. Horng, S.; Sontag, D. A.; Halpern, Y.; Jernite, Y.; Shapiro, N. I.; Nathanson, L. A. Creating an automated trigger for sepsis clinical decision support at emergency department triage using machine learning. PLoS ONE 2017, 12(4), e0174708. [Google Scholar] [CrossRef] [PubMed]
  10. Churpek, M. M.; Yuen, T. C.; Winslow, C.; Meltzer, D. O.; Kattan, M. W.; Edelson, D. P. Multicenter comparison of machine learning methods and conventional regression for predicting clinical deterioration on the wards. Crit. Care Med. 2016, 44(2), 368–374. [Google Scholar] [CrossRef] [PubMed]
  11. Lipton, Z. C.; Kale, D. C.; Elkan, C.; Wetzel, R. (2016). Learning to diagnose with LSTM recurrent neural networks. In Proceedings of the International Conference on Learning Representations (ICLR).
  12. Cho, K.; van Merrienboer, B.; Gulcehre, C.; Bahdanau, D.; Bougares, F.; Schwenk, H.; Bengio, Y. (2014). Learning phrase representations using RNN encoder-decoder for statistical machine translation. In Proceedings of EMNLP. 2014.
  13. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. (NeurIPS) 2017, 30. [Google Scholar]
  14. Shashikumar, S. P.; Josef, C. S.; Sharma, A.; Nemati, S. DeepAISE—an interpretable and recurrent neural survival model for early prediction of sepsis. Artif. Intell. Med. 2021, 113, 102036. [Google Scholar] [CrossRef] [PubMed]
  15. Che, Z.; Purushotham, S.; Cho, K.; Sontag, D.; Liu, Y. Recurrent neural networks for multivariate time series with missing values. Sci. Rep. 2018, 8(1), 6085. [Google Scholar] [CrossRef] [PubMed]
  16. Moor, M.; Horn, M.; Rieck, B.; Roqueiro, D.; Borgwardt, K. (2021). Early warning in the emergency room. In Proceedings of the ACM Conference on Health, Inference, and Learning (CHIL).
  17. Reyna, M. A.; Josef, C. S.; Jeter, R.; Shashikumar, S. P.; Westover, M. B.; Nemati, S.; Clifford, G. D. Early prediction of sepsis from clinical data: The PhysioNet/Computing in Cardiology Challenge 2019. Crit. Care Med. 2020, 48(2), 210–217. [Google Scholar] [CrossRef] [PubMed]
  18. Reyna, M. A.; Lehman, L.-W. H.; Clifford, G. D. Rethinking the early sepsis definition. Am. J. Respir. Crit. Care Med. 2021, 204(5), 487–489. [Google Scholar]
  19. Li, Y.; Rao, S.; Solares, J. R. A.; Hassaine, A.; Ramakrishnan, R.; Canoy, D.; Nadarajah, R. BEHRT: Transformer for electronic health records. Sci. Rep. 2020, 10, 7155. [Google Scholar] [CrossRef] [PubMed]
  20. Shukla, S. N.; Marlin, B. M. (2019). Interpolation-prediction networks for irregularly sampled time series. In Proceedings of ICLR.
  21. Rasmy, L.; Xiang, Y.; Xie, Z.; Tao, C.; Zhi, D. Med-BERT: Pretrained contextualized embeddings on large-scale structured EHR for disease prediction. npj Digit. Med. 2021, 4, 86. [Google Scholar] [CrossRef] [PubMed]
  22. Wu, H.; Hu, T.; Liu, Y.; Zhou, H.; Wang, J.; Long, M. (2023). TimesNet: Temporal 2D-variation modeling for general time series analysis. In Proceedings of ICLR.
  23. Nie, Y.; Nguyen, N. H.; Sinthong, P.; Kalagnanam, J. (2023). A time series is worth 64 words. In Proceedings of ICLR.
  24. Lundberg, S. M.; Lee, S.-I. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems (NeurIPS); 2017; p. 30. [Google Scholar]
  25. Selvaraju, R. R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. (2017). Grad-CAM: Visual explanations from deep networks via gradient-based localization. In Proceedings of ICCV.
  26. Strodthoff, N.; Wagner, P.; Schaeffter, T.; Samek, W. Deep learning for ECG analysis. IEEE J. Biomed. Health Inform. 2021, 25(5), 1519–1528. [Google Scholar] [CrossRef] [PubMed]
  27. Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. (2017). Focal loss for dense object detection. In Proceedings of ICCV.
  28. Loshchilov, I.; Hutter, F. (2019). Decoupled weight decay regularization. In Proceedings of ICLR.
  29. Guo, C.; Pleiss, G.; Sun, Y.; Weinberger, K. Q. (2017). On calibration of modern neural networks. In Proceedings of ICML.
  30. DeLong, E. R.; DeLong, D. M.; Clarke-Pearson, D. L. Comparing the areas under two or more correlated receiver operating characteristic curves. Biometrics 1988, 44(3), 837–845. [Google Scholar] [CrossRef] [PubMed]
Figure 2. Full evaluation dashboard: validation AUROC curves during training (top-left), ROC curves (top-centre), precision-recall curves (top-right), confusion matrix at clinical threshold (bottom-left), XGBoost feature importance (bottom-centre), and focal loss curves (bottom-right).
Figure 2. Full evaluation dashboard: validation AUROC curves during training (top-left), ROC curves (top-centre), precision-recall curves (top-right), confusion matrix at clinical threshold (bottom-left), XGBoost feature importance (bottom-centre), and focal loss curves (bottom-right).
Preprints 218225 g002
Figure 4. Alert lead-time distribution (left) and cumulative early alert coverage as a function of minimum required lead time (right). The model achieves a median lead time of 46.5 h with 95.3% of septic patients alerted at least 3 h before onset.
Figure 4. Alert lead-time distribution (left) and cumulative early alert coverage as a function of minimum required lead time (right). The model achieves a median lead time of 46.5 h with 95.3% of septic patients alerted at least 3 h before onset.
Preprints 218225 g004
Figure 5. Subgroup AUROC stratified by age group (blue), severity of illness (orange), and sex (purple). Dashed line indicates overall AUROC (0.826). High-severity patients (SOFA proxy ≥3) show substantially reduced performance.
Figure 5. Subgroup AUROC stratified by age group (blue), severity of illness (orange), and sex (purple). Dashed line indicates overall AUROC (0.826). High-severity patients (SOFA proxy ≥3) show substantially reduced performance.
Preprints 218225 g005
Figure 6. SHAP beeswarm plot showing feature impact on XGBoost sepsis predictions. Each point represents one patient-hour; colour encodes feature value (red = high, blue = low). Positive SHAP values indicate increased sepsis risk.
Figure 6. SHAP beeswarm plot showing feature impact on XGBoost sepsis predictions. Each point represents one patient-hour; colour encodes feature value (red = high, blue = low). Positive SHAP values indicate increased sepsis risk.
Preprints 218225 g006
Figure 7. SHAP feature impact for the Transformer-aligned XGBoost model, showing the top 15 features by mean absolute SHAP value.
Figure 7. SHAP feature impact for the Transformer-aligned XGBoost model, showing the top 15 features by mean absolute SHAP value.
Preprints 218225 g007
Figure 8. Grad-CAM analysis for three representative patients: two septic (Patients 78 and 283) and one non-septic (Patient 10). Each row shows: (left) vital signs with Grad-CAM intensity as background shading, (centre) temporal importance bar chart, (right) feature-by-timestep saliency heatmap.
Figure 8. Grad-CAM analysis for three representative patients: two septic (Patients 78 and 283) and one non-septic (Patient 10). Each row shows: (left) vital signs with Grad-CAM intensity as background shading, (centre) temporal importance bar chart, (right) feature-by-timestep saliency heatmap.
Preprints 218225 g008
Figure 9. Population-level Grad-CAM analysis across all septic test patients. Left: mean temporal importance with standard deviation band. Right: per-patient CAM heatmap showing consistent emphasis on early timesteps t 10 and t 9 .
Figure 9. Population-level Grad-CAM analysis across all septic test patients. Left: mean temporal importance with standard deviation band. Right: per-patient CAM heatmap showing consistent emphasis on early timesteps t 10 and t 9 .
Preprints 218225 g009
Table 1. Feature categories and engineering summary.
Table 1. Feature categories and engineering summary.
Category N Examples Engineering
Vital signs 8 HR, MAP, Temp, O2Sat Forward-fill; 3-hr rolling mean/std/diff
Laboratory values 26 Lactate, WBC, Creatinine Median imputation; missingness flags
Static/demographics 6 Age, Gender, ICULOS As-is; ICULOS as temporal anchor
Derived/engineered 5 shock_index, sofa_proxy HR/SBP ratio; simplified SOFA
Missingness flags 26 Lactate_miss, PTT_miss Binary lab-presence indicators
Rolling statistics 21 HR_mean3, MAP_std3 3-hr window per vital/lab
Total 92 After all engineering steps
Table 2. Hyperparameter configuration.
Table 2. Hyperparameter configuration.
Hyperparameter Value Justification
d model 128 Capacity vs. overfitting balance
Attention heads 8 Standard for d model = 128
Encoder layers 3 Minimal gain observed at 4
Sequence length 12 12-hour clinical prodrome
Batch size 512 Saturates T4 GPU memory
Learning rate 5 × 10 5 Stable convergence with warmup
Focal loss α 0.25 Lin et al. [27]
Focal loss γ 2.0 Lin et al. [27]
Dropout 0.2 Regularisation
Early-stop patience 15 epochs Prevents premature termination
Optimiser AdamW Weight decay regularisation
Table 3. Model comparison on held-out test set.
Table 3. Model comparison on held-out test set.
Model AUROC AUPRC S@90 DeLong p
TTE (proposed) 0.8264 0.1686 0.5125
BiLSTM + Attn 0.7859 0.1098 0.4647 <0.0001
XGBoost 0.7731 0.1278 0.4711 <0.0001
CV 5-fold (TTE) 0.8320 ± 0.0032 0.1505 ± 0.0148 0.5461 ± 0.0127
S@90 = Sensitivity at 90% Specificity. DeLong test vs. TTE.
Table 4. Subgroup analysis results.
Table 4. Subgroup analysis results.
Subgroup AUROC Δ Clinical Interpretation
Female 0.840 +0.014 Consistent with overall
>65 years 0.834 +0.008 Elderly well represented
<45 years 0.824 −0.002 Near-average performance
Male 0.818 −0.008 Slight underperformance
45–65 years 0.820 −0.006 Moderate performance
Low severity (0) 0.805 −0.021 Subtle presentation
Medium severity (1–2) 0.795 −0.031 Boundary cases; hardest
High severity (3+) 0.257 −0.569 Critical failure at high SOFA
Table 5. Ablation study results on validation set.
Table 5. Ablation study results on validation set.
Configuration AUROC AUPRC Δ AUROC
Full TTE (proposed) 0.8264 0.1686
w/o positional encoding 0.8071 0.1542 −0.019
w/o pre-norm (post-norm) 0.8183 0.1601 −0.008
1 encoder layer (vs. 3) 0.7942 0.1388 −0.032
SEQ_LEN = 6 (vs. 12) 0.7980 0.1441 −0.028
BCE loss (vs. focal) 0.7753 0.0891 −0.051
No WeightedSampler 0.7621 0.0713 −0.064
No rolling features 0.8102 0.1503 −0.016
No missingness flags 0.8211 0.1649 −0.005
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings