Submitted:
12 June 2026
Posted:
12 June 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
1.1. Background and Clinical Motivation
1.2. Limitations of Existing Methods
1.3. Related Work
1.4. Contributions
- A Temporal Transformer Encoder (TTE) achieving AUROC under 5-fold patient-level cross-validation on PhysioNet 2019—5.9% above XGBoost and 4.6% above BiLSTM.
- A novel temporal Grad-CAM method for sequence Transformers, revealing prediction importance concentrated at hours and , identifying the 12-hour clinical prodrome.
- The first reported median alert lead time of 46.5 h on PhysioNet 2019, with 95.3% of septic patients alerted ≥3 h before clinical onset.
- A complete post-hoc calibration pipeline reducing ECE from 0.3154 to 0.0017 via Platt scaling.
- Subgroup analysis revealing dramatically reduced performance for high-severity patients (SOFA proxy ≥3, AUROC = 0.257), motivating stratified modelling.
2. Materials and Methods
2.1. Data Source and Population

2.2. Feature Space
2.3. Label Construction
2.4. Preprocessing Pipeline
- Step 1: Missingness Encoding. A binary feature is created for each laboratory feature to account for the clinical importance of an unrecorded test result.
- Step 2: Forward Filling for Vital Signs. Vital sign data is forward-filled for each patient stay, reflecting clinical practice in which the last value serves as a reasonable estimate of the current state.
- Step 3: Median Imputation. Remaining missing values are imputed using the median of each feature, computed from the training set and applied consistently to validation and test sets.
- Step 4: Standardisation. All features are normalised using z-score standardisation:where and represent the mean and standard deviation computed from the training data.
- Step 5: NaN Clamping. After scaling, any NaN or Inf values arising from zero-variance cases are replaced with 0.
- Step 6: Feature Engineering. Four composite features are derived: Shock Index (HR/SBP), Pulse Pressure (SBP−DBP), Temperature Deviation (), and a Simplified SOFA Score computed using binary thresholds for creatinine, bilirubin, platelets, MAP, respiratory rate, and WBC. Additionally, temporal features are generated using a 3-hour rolling window (mean, standard deviation, and first-order difference) for seven key variables: HR, MAP, Resp, Temp, Sat, Lactate, and WBC.
2.5. Data Splitting and Imbalance Handling
2.6. Proposed Architecture: Temporal Transformer Encoder (TTE)
2.6.1. Input Representation
2.6.2. Model Architecture
Input Projection.
Positional Encoding.
Transformer Encoder.
Classification Head.
2.7. Training Protocol
2.8. Baseline Models
2.9. Grad-CAM for Temporal Explainability
3. Results
3.1. Overall Performance
3.2. Calibration Analysis

3.3. Alert Lead-Time Analysis
3.4. Subgroup Analysis
3.5. SHAP Feature Importance
3.6. Grad-CAM Temporal Analysis
4. Discussion
4.1. Ablation Study
4.2. Theoretical Justification
4.2.1. Why Self-Attention Outperforms Recurrence for Clinical Time Series
4.2.2. Optimisation Landscape Under Severe Imbalance
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AUROC | Area Under the Receiver Operating Characteristic Curve |
| AUPRC | Area Under the Precision-Recall Curve |
| ECE | Expected Calibration Error |
| EHR | Electronic Health Record |
| GRU | Gated Recurrent Unit |
| ICU | Intensive Care Unit |
| IQR | Interquartile Range |
| LSTM | Long Short-Term Memory |
| MAP | Mean Arterial Pressure |
| MEWS | Modified Early Warning Score |
| NEWS2 | National Early Warning Score 2 |
| qSOFA | Quick Sequential Organ Failure Assessment |
| SHAP | SHapley Additive exPlanations |
| SIRS | Systemic Inflammatory Response Syndrome |
| SOFA | Sequential Organ Failure Assessment |
| TTE | Temporal Transformer Encoder |
References
- Singer, M.; Deutschman, C. S.; Seymour, C. W.; Shankar-Hari, M.; Annane, D.; Bauer, M.; Angus, D. C. The Third International Consensus Definitions for Sepsis and Septic Shock (Sepsis-3). JAMA 2016, 315(8), 801–810. [Google Scholar] [CrossRef] [PubMed]
- Rudd, K. E.; Johnson, S. C.; Agesa, K. M.; Shackelford, K. A.; Tsoi, D.; Kievlan, D. R.; Naghavi, M. Global, regional, and national sepsis incidence and mortality, 1990–2017: analysis for the Global Burden of Disease Study. Lancet 2020, 395(10219), 200–211. [Google Scholar] [CrossRef] [PubMed]
- Fleischmann-Struzek, C.; Goldfarb, D. M.; Schlattmann, P.; Schlapbach, L. J.; Reinhart, K.; Kissoon, N. The global burden of paediatric and neonatal sepsis. Lancet Respir. Med. 2018, 6(3), 223–230. [Google Scholar] [CrossRef] [PubMed]
- Kumar, A.; Roberts, D.; Wood, K. E.; Light, B.; Parrillo, J. E.; Sharma, S.; Cheang, M. Duration of hypotension before initiation of effective antimicrobial therapy is the critical determinant of survival in human septic shock. Crit. Care Med. 2006, 34(6), 1589–1596. [Google Scholar] [CrossRef] [PubMed]
- Seymour, C. W.; Liu, V. X.; Iwashyna, T. J.; Brunkhorst, F. M.; Rea, T. D.; Scherag, A.; Angus, D. C. Assessment of Clinical Criteria for Sepsis. JAMA 2016, 315(8), 762–774. [Google Scholar] [CrossRef] [PubMed]
- Royal College of Physicians. (2017). National Early Warning Score (NEWS) 2. RCP. RC.
- Subbe, C. P.; Kruger, M.; Rutherford, P.; Gemmel, L. Validation of a modified Early Warning Score in medical admissions. QJM 2001, 94(10), 521–526. [Google Scholar] [CrossRef] [PubMed]
- Bone, R. C.; Balk, R. A.; Cerra, F. B.; Dellinger, R. P.; Fein, A. M.; Knaus, W. A.; Sibbald, W. J. Definitions for sepsis and organ failure and guidelines for the use of innovative therapies in sepsis. Chest 1992, 101(6), 1644–1655. [Google Scholar] [CrossRef] [PubMed]
- Horng, S.; Sontag, D. A.; Halpern, Y.; Jernite, Y.; Shapiro, N. I.; Nathanson, L. A. Creating an automated trigger for sepsis clinical decision support at emergency department triage using machine learning. PLoS ONE 2017, 12(4), e0174708. [Google Scholar] [CrossRef] [PubMed]
- Churpek, M. M.; Yuen, T. C.; Winslow, C.; Meltzer, D. O.; Kattan, M. W.; Edelson, D. P. Multicenter comparison of machine learning methods and conventional regression for predicting clinical deterioration on the wards. Crit. Care Med. 2016, 44(2), 368–374. [Google Scholar] [CrossRef] [PubMed]
- Lipton, Z. C.; Kale, D. C.; Elkan, C.; Wetzel, R. (2016). Learning to diagnose with LSTM recurrent neural networks. In Proceedings of the International Conference on Learning Representations (ICLR).
- Cho, K.; van Merrienboer, B.; Gulcehre, C.; Bahdanau, D.; Bougares, F.; Schwenk, H.; Bengio, Y. (2014). Learning phrase representations using RNN encoder-decoder for statistical machine translation. In Proceedings of EMNLP. 2014.
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. (NeurIPS) 2017, 30. [Google Scholar]
- Shashikumar, S. P.; Josef, C. S.; Sharma, A.; Nemati, S. DeepAISE—an interpretable and recurrent neural survival model for early prediction of sepsis. Artif. Intell. Med. 2021, 113, 102036. [Google Scholar] [CrossRef] [PubMed]
- Che, Z.; Purushotham, S.; Cho, K.; Sontag, D.; Liu, Y. Recurrent neural networks for multivariate time series with missing values. Sci. Rep. 2018, 8(1), 6085. [Google Scholar] [CrossRef] [PubMed]
- Moor, M.; Horn, M.; Rieck, B.; Roqueiro, D.; Borgwardt, K. (2021). Early warning in the emergency room. In Proceedings of the ACM Conference on Health, Inference, and Learning (CHIL).
- Reyna, M. A.; Josef, C. S.; Jeter, R.; Shashikumar, S. P.; Westover, M. B.; Nemati, S.; Clifford, G. D. Early prediction of sepsis from clinical data: The PhysioNet/Computing in Cardiology Challenge 2019. Crit. Care Med. 2020, 48(2), 210–217. [Google Scholar] [CrossRef] [PubMed]
- Reyna, M. A.; Lehman, L.-W. H.; Clifford, G. D. Rethinking the early sepsis definition. Am. J. Respir. Crit. Care Med. 2021, 204(5), 487–489. [Google Scholar]
- Li, Y.; Rao, S.; Solares, J. R. A.; Hassaine, A.; Ramakrishnan, R.; Canoy, D.; Nadarajah, R. BEHRT: Transformer for electronic health records. Sci. Rep. 2020, 10, 7155. [Google Scholar] [CrossRef] [PubMed]
- Shukla, S. N.; Marlin, B. M. (2019). Interpolation-prediction networks for irregularly sampled time series. In Proceedings of ICLR.
- Rasmy, L.; Xiang, Y.; Xie, Z.; Tao, C.; Zhi, D. Med-BERT: Pretrained contextualized embeddings on large-scale structured EHR for disease prediction. npj Digit. Med. 2021, 4, 86. [Google Scholar] [CrossRef] [PubMed]
- Wu, H.; Hu, T.; Liu, Y.; Zhou, H.; Wang, J.; Long, M. (2023). TimesNet: Temporal 2D-variation modeling for general time series analysis. In Proceedings of ICLR.
- Nie, Y.; Nguyen, N. H.; Sinthong, P.; Kalagnanam, J. (2023). A time series is worth 64 words. In Proceedings of ICLR.
- Lundberg, S. M.; Lee, S.-I. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems (NeurIPS); 2017; p. 30. [Google Scholar]
- Selvaraju, R. R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. (2017). Grad-CAM: Visual explanations from deep networks via gradient-based localization. In Proceedings of ICCV.
- Strodthoff, N.; Wagner, P.; Schaeffter, T.; Samek, W. Deep learning for ECG analysis. IEEE J. Biomed. Health Inform. 2021, 25(5), 1519–1528. [Google Scholar] [CrossRef] [PubMed]
- Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. (2017). Focal loss for dense object detection. In Proceedings of ICCV.
- Loshchilov, I.; Hutter, F. (2019). Decoupled weight decay regularization. In Proceedings of ICLR.
- Guo, C.; Pleiss, G.; Sun, Y.; Weinberger, K. Q. (2017). On calibration of modern neural networks. In Proceedings of ICML.
- DeLong, E. R.; DeLong, D. M.; Clarke-Pearson, D. L. Comparing the areas under two or more correlated receiver operating characteristic curves. Biometrics 1988, 44(3), 837–845. [Google Scholar] [CrossRef] [PubMed]







| Category | N | Examples | Engineering |
|---|---|---|---|
| Vital signs | 8 | HR, MAP, Temp, O2Sat | Forward-fill; 3-hr rolling mean/std/diff |
| Laboratory values | 26 | Lactate, WBC, Creatinine | Median imputation; missingness flags |
| Static/demographics | 6 | Age, Gender, ICULOS | As-is; ICULOS as temporal anchor |
| Derived/engineered | 5 | shock_index, sofa_proxy | HR/SBP ratio; simplified SOFA |
| Missingness flags | 26 | Lactate_miss, PTT_miss | Binary lab-presence indicators |
| Rolling statistics | 21 | HR_mean3, MAP_std3 | 3-hr window per vital/lab |
| Total | 92 | — | After all engineering steps |
| Hyperparameter | Value | Justification |
|---|---|---|
| 128 | Capacity vs. overfitting balance | |
| Attention heads | 8 | Standard for |
| Encoder layers | 3 | Minimal gain observed at 4 |
| Sequence length | 12 | 12-hour clinical prodrome |
| Batch size | 512 | Saturates T4 GPU memory |
| Learning rate | Stable convergence with warmup | |
| Focal loss | 0.25 | Lin et al. [27] |
| Focal loss | 2.0 | Lin et al. [27] |
| Dropout | 0.2 | Regularisation |
| Early-stop patience | 15 epochs | Prevents premature termination |
| Optimiser | AdamW | Weight decay regularisation |
| Model | AUROC | AUPRC | S@90 | DeLong p |
|---|---|---|---|---|
| TTE (proposed) | 0.8264 | 0.1686 | 0.5125 | — |
| BiLSTM + Attn | 0.7859 | 0.1098 | 0.4647 | <0.0001 |
| XGBoost | 0.7731 | 0.1278 | 0.4711 | <0.0001 |
| CV 5-fold (TTE) | — |
| Subgroup | AUROC | Clinical Interpretation | |
|---|---|---|---|
| Female | 0.840 | +0.014 | Consistent with overall |
| >65 years | 0.834 | +0.008 | Elderly well represented |
| <45 years | 0.824 | −0.002 | Near-average performance |
| Male | 0.818 | −0.008 | Slight underperformance |
| 45–65 years | 0.820 | −0.006 | Moderate performance |
| Low severity (0) | 0.805 | −0.021 | Subtle presentation |
| Medium severity (1–2) | 0.795 | −0.031 | Boundary cases; hardest |
| High severity (3+) | 0.257 | −0.569 | Critical failure at high SOFA |
| Configuration | AUROC | AUPRC | AUROC |
|---|---|---|---|
| Full TTE (proposed) | 0.8264 | 0.1686 | — |
| w/o positional encoding | 0.8071 | 0.1542 | −0.019 |
| w/o pre-norm (post-norm) | 0.8183 | 0.1601 | −0.008 |
| 1 encoder layer (vs. 3) | 0.7942 | 0.1388 | −0.032 |
| SEQ_LEN = 6 (vs. 12) | 0.7980 | 0.1441 | −0.028 |
| BCE loss (vs. focal) | 0.7753 | 0.0891 | −0.051 |
| No WeightedSampler | 0.7621 | 0.0713 | −0.064 |
| No rolling features | 0.8102 | 0.1503 | −0.016 |
| No missingness flags | 0.8211 | 0.1649 | −0.005 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).