Preprint
Review

This version is not peer-reviewed.

Artificial Intelligence for Early Prediction and Diagnosis of Neonatal Sepsis: Current Evidence, Challenges, and Future Directions

Submitted:

14 July 2026

Posted:

15 July 2026

You are already at the latest version

Abstract
Neonatal sepsis remains a major cause of morbidity and mortality worldwide, while timely diagnosis continues to be challenging because of nonspecific clinical manifestations and limitations of conventional diagnostic methods. Recent advances in artificial intelligence (AI) have created new opportunities for the early prediction and diagnosis of neonatal sepsis through the analysis of large and complex clinical datasets. This structured narrative review summarizes current evidence regarding AI-based approaches for neonatal sepsis prediction and diagnosis. A literature search of PubMed, Scopus, and Google Scholar identified studies evaluating machine learning, deep learning, and advanced predictive analytics using clinical, laboratory, physiological, electronic health record, and multi-omics data. Current evidence demonstrates that AI models, particularly ensemble learning, gradient boosting, and deep learning approaches, can achieve strong predictive performance and identify infants at increased risk of sepsis hours before conventional clinical recognition. Continuous physiological monitoring and multimodal data integration appear particularly promising for real-time prediction. However, important challenges remain, including limited external validation, small and heterogeneous datasets, concerns regarding interpretability, and unresolved ethical and regulatory issues. Future progress will depend on multicenter collaboration, explainable AI frameworks, federated learning, and multimodal predictive models. Artificial intelligence has the potential to become a valuable decision-support tool that enhances early sepsis recognition and supports precision neonatal care.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

Neonatal sepsis remains a major cause of neonatal morbidity and mortality worldwide and continues to represent a significant challenge for neonatal intensive care units (NICUs) [1,2]. Despite substantial advances in perinatal care, antimicrobial therapy, and infection prevention strategies, sepsis accounts for a considerable proportion of preventable neonatal deaths, particularly among preterm and very-low-birth-weight infants [2,3]. The condition is commonly classified as early-onset sepsis (EOS), occurring within the first 72 hours of life, and late-onset sepsis (LOS), which develops later and is frequently associated with nosocomial or environmental exposures [1]. Beyond acute mortality, neonatal sepsis is associated with adverse long-term outcomes, including neurodevelopmental impairment, growth restriction, and chronic health complications [3]. Furthermore, the growing threat of antimicrobial resistance has increased the complexity of infection management and reinforced the need for earlier and more accurate diagnostic approaches [1,3].
Timely diagnosis of neonatal sepsis remains particularly challenging because clinical manifestations are often subtle, nonspecific, and highly variable [3,4]. Common signs, including respiratory distress, apnea, feeding intolerance, lethargy, temperature instability, and cardiovascular dysfunction, may occur in both infectious and noninfectious neonatal conditions, limiting their diagnostic value [1,3]. Although blood culture remains the reference standard for confirming infection, its diagnostic performance is constrained by low bacterial loads, inadequate sample volumes, prior antibiotic exposure, and prolonged turnaround times [3]. Similarly, widely used biomarkers such as C-reactive protein (CRP), procalcitonin (PCT), and interleukin-6 (IL-6) provide useful adjunctive information but lack sufficient accuracy when used in isolation [4,5]. Consequently, empirical antibiotic therapy is frequently initiated in neonates with suspected sepsis, contributing to antibiotic overuse, disruption of the developing microbiome, increased healthcare costs, and the emergence of antimicrobial resistance [3,6]. These limitations highlight the urgent need for innovative approaches capable of improving risk stratification and supporting earlier diagnosis [4,5].
Recent advances in artificial intelligence (AI) have generated considerable interest as potential solutions to these longstanding diagnostic challenges and have transformed multiple areas of healthcare by enabling the analysis of large, heterogeneous, and high-dimensional clinical datasets that support prediction, diagnosis, and clinical decision-making [7,8,9,10,11]. The widespread adoption of electronic health records, bedside monitoring systems, and high-throughput laboratory technologies has resulted in the generation of large volumes of complex clinical data that frequently exceed the analytical capabilities of conventional statistical approaches [7]. Machine learning (ML) and deep learning (DL) techniques offer the ability to identify hidden patterns, integrate heterogeneous data sources, and support data-driven clinical decision-making across a broad spectrum of medical specialties, including critical care, precision medicine, and public health [8,12,13].
Over the past decade, an increasing number of studies have investigated AI-based approaches using demographic characteristics, laboratory biomarkers, physiological signals, electronic health records, and multi-omics data, reporting encouraging results across diverse clinical settings [5,14]. More recently, multimodal AI approaches capable of integrating structured clinical information with physiological monitoring, laboratory findings, and molecular datasets have emerged as a promising strategy for improving predictive performance and supporting more comprehensive clinical decision-making [15]. This review critically examines the current evidence regarding AI applications for the early prediction and diagnosis of neonatal sepsis, with particular emphasis on data sources, machine learning methodologies, predictive performance, explainability, implementation challenges, and future directions for clinical integration.
A structured narrative review was conducted to evaluate current evidence regarding the application of artificial intelligence for the early prediction and diagnosis of neonatal sepsis. Relevant literature was identified through searches of PubMed, Scopus, and Google Scholar conducted between May and June 2026. The search strategy combined terms related to neonatal sepsis, artificial intelligence, machine learning, deep learning, predictive analytics, electronic health records, physiological monitoring, biomarkers, and multi-omics data. Additional studies were identified through manual screening of reference lists.
Articles published in English between 2006 and 2026 were considered eligible. Original studies and review articles investigating AI-based approaches for neonatal sepsis prediction or diagnosis were included, whereas duplicate publications, conference abstracts, and studies not directly relevant to neonatal sepsis were excluded. A total of 268 records were initially identified, of which 53 studies were considered relevant and included in the final narrative synthesis. Given the substantial methodological heterogeneity of published AI studies, including differences in study design, data sources, predictive algorithms, outcome definitions, validation strategies, and performance metrics, a structured narrative review was considered the most appropriate approach to provide a comprehensive and clinically meaningful synthesis of the available evidence. Therefore, no formal quantitative risk-of-bias assessment was performed.

2. Artificial Intelligence in Neonatal Healthcare

Artificial intelligence (AI) refers to the ability of computer systems to perform tasks that traditionally require human intelligence, including pattern recognition, prediction, decision-making, and problem-solving. In healthcare, its rapid expansion has been driven by the increasing availability of electronic health records (EHRs), digital imaging, physiological monitoring systems, and high-throughput molecular technologies [7,8,9,10,16]. AI algorithms can analyze large and complex datasets to identify patterns and interactions, supporting predictive analytics and precision medicine across numerous medical specialties, including neonatology [7,12].
Within AI, machine learning (ML) algorithms learn from data to generate predictions without being explicitly programmed for every possible scenario [5]. Most clinical prediction models use supervised learning, with commonly applied algorithms including logistic regression, decision trees, random forests, support vector machines (SVMs), gradient boosting, and extreme gradient boosting (XGBoost) [5,13]. These approaches have demonstrated broad utility in disease prediction, risk stratification, and clinical decision support, including the early detection and monitoring of sepsis [17].
Deep learning (DL), a subset of ML based on multilayer artificial neural networks, can automatically learn hierarchical representations from complex datasets [7]. Beyond image analysis and biomedical signal interpretation, DL has evolved toward multimodal architectures integrating structured clinical data, imaging, physiological signals, laboratory results, and free-text documentation [15]. In neonatal care, architectures such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and long short-term memory (LSTM) networks have been increasingly applied to physiological monitoring and EHR data [5,7].
Modern NICUs provide a particularly data-rich environment, generating continuous physiological measurements alongside laboratory, microbiological, therapeutic, maternal, and perinatal information [18]. Integrating these heterogeneous data streams offers substantial opportunities for AI-driven prediction [7,14], while multimodal learning and privacy-preserving approaches such as federated learning may facilitate collaborative model development across institutions [19]. Consequently, AI provides a promising framework for transforming complex neonatal data into clinically actionable information and supporting precision neonatal care [14].
Despite this potential, most AI models for neonatal sepsis remain in the development or validation stage, with data quality, limited external validation, interpretability, algorithmic bias, and regulatory oversight restricting widespread clinical implementation [20,21]. Emerging foundation models and generative AI may create new opportunities for clinical decision support, but their application in neonatal medicine will require rigorous prospective validation, ethical oversight, and transparent governance [22,23]. Addressing these challenges is essential for the safe, effective, and equitable integration of AI into routine neonatal practice [13,16].

3. Data Sources for AI-Based Neonatal Sepsis Prediction

3.1. Clinical and Demographic Data

Clinical and demographic variables constitute the foundation of many artificial intelligence (AI) models for neonatal sepsis prediction because they are routinely collected and readily available within electronic health records [11]. Commonly used predictors include gestational age, birth weight, sex, mode of delivery, Apgar scores, maternal characteristics, and perinatal risk factors such as prolonged rupture of membranes, maternal fever, chorioamnionitis, intrapartum antibiotic exposure, and Group B Streptococcus colonization [1,3]. These variables support early risk stratification and have been incorporated into predictive models and neonatal sepsis calculators to reduce unnecessary antibiotic use [5,24,25]. More recently, an externally validated machine learning model based on routinely available clinical variables demonstrated the potential to identify neonates at low risk of early-onset sepsis and reduce unnecessary antibiotic exposure [26]. However, clinical and demographic variables alone rarely provide sufficient predictive accuracy and are generally most effective when integrated with laboratory, physiological, and molecular data within multimodal AI models [5,15].

3.2. Laboratory Biomarkers

Laboratory biomarkers remain integral components of many AI-based neonatal sepsis prediction models. Frequently used biomarkers include C-reactive protein (CRP), procalcitonin (PCT), white blood cell count (WBC), immature-to-total neutrophil ratio, platelet count, and interleukin-6 (IL-6) [3]. Although none provides sufficient diagnostic accuracy independently, combining multiple biomarkers within machine learning frameworks may improve predictive performance by capturing complex inflammatory and hematological interactions [4,5].
Several predictive models have incorporated laboratory variables as key input features. Goldberg et al. reported improved prediction of late-onset neonatal sepsis using clinical assessment, neutrophil-to-lymphocyte ratio, and CRP [27], while broader evidence has identified CRP and WBC count among the most frequently reported predictors in machine learning studies of neonatal sepsis [28]. Multiple studies have similarly identified hematological indices and inflammatory biomarkers among the most influential features in AI-based risk assessment [5].
Recent advances have expanded beyond conventional biomarkers toward high-dimensional molecular datasets. An et al. identified a predictive gene-expression signature capable of detecting neonatal sepsis before clinical presentation [29], while Mithal et al. demonstrated the potential of cord blood proteomics combined with machine learning for early-onset sepsis prediction [30]. Integrating molecular biomarkers with clinical, laboratory, physiological, and EHR data through multimodal AI frameworks may further improve predictive performance and support precision medicine [15]. However, further validation is required before routine clinical implementation.

3.3. Electronic Health Records

The widespread adoption of electronic health records (EHRs) has transformed predictive analytics in neonatal medicine by providing access to large volumes of longitudinal clinical data, including demographic characteristics, maternal and perinatal history, laboratory findings, medication records, microbiological results, clinical observations, and therapeutic interventions [7]. Consequently, EHRs have become one of the most frequently used data sources for AI-based neonatal sepsis prediction and a cornerstone of contemporary data-driven healthcare [5,9,11].
One of the earliest and most influential studies in this field was conducted by Mani et al., who developed machine learning models for the early detection of late-onset neonatal sepsis using routinely collected NICU data. Multiple algorithms, including support vector machines, random forests, logistic regression, decision trees, and Bayesian classifiers, were evaluated, with several outperforming physician-guided antibiotic initiation [31]. These findings highlighted the potential of routinely collected EHR data to support earlier identification of high-risk infants [5].
Further evidence was provided by Masino et al., who developed predictive models using EHR data collected at least four hours before clinical diagnosis. Their analysis incorporated demographic information, laboratory findings, vital signs, and clinical observations and demonstrated AUC values exceeding 0.80 across several machine learning models [14]. Predictive performance improved when both culture-positive and clinically diagnosed cases were included, suggesting that AI may capture clinically meaningful disease patterns even without microbiological confirmation [5,14]. Other neonatal studies have similarly demonstrated the potential of AI models integrating routinely collected clinical and laboratory data for sepsis prediction in NICU populations [32], these findings support the growing value of heterogeneous clinical data within multimodal AI frameworks [15].
Despite their promise, EHR-based models face challenges related to missing data, inconsistent documentation, institutional variability, and the predominance of retrospective single-center datasets, which limit external validation and generalizability [13]. Privacy-preserving approaches such as federated learning may facilitate multicenter model development without transferring sensitive patient-level data [19]. Nevertheless, EHR-derived systems remain among the most practical approaches for clinical implementation because they rely on routinely collected information [7], while recent studies increasingly explore their integration into clinical decision-support systems for neonatal sepsis and other complications [33].

3.4. Physiological Monitoring and Vital Signs

Continuous physiological monitoring represents one of the most promising data sources for early neonatal sepsis detection. Modern NICUs routinely collect high-frequency physiological data, including heart rate, respiratory rate, oxygen saturation, blood pressure, and temperature [18,34]. These signals provide a real-time representation of neonatal physiology and may reveal subtle alterations preceding overt clinical manifestations. Unlike intermittent laboratory testing, continuous monitoring is particularly suited to early warning systems capable of detecting deterioration before clinical recognition [18].
Among physiological signals, heart rate variability (HRV) has received particular attention as a potential marker of neonatal sepsis. Alterations in autonomic regulation during systemic infection may produce measurable changes in heart rate dynamics before traditional clinical signs become evident. Gomez et al. demonstrated that HRV-derived parameters could differentiate neonates with sepsis from healthy controls using machine learning algorithms [35]. AdaBoost achieved the highest predictive performance, with an AUC approaching 0.94, supporting the potential utility of noninvasive physiological monitoring for early sepsis detection [5,35].
Several investigators have expanded this concept by integrating multiple physiological signals. Song et al. demonstrated that routinely collected vital signs, including blood pressure, oxygen saturation, and body temperature, could predict late-onset neonatal sepsis up to 48 hours before diagnosis [36]. Garstman et al. further demonstrated the potential of machine learning for early LOS detection in extremely preterm infants, with a random forest model achieving an AUROC of 0.973 [37]. Similarly, Honoré et al. analyzed high-frequency monitoring data from 325 neonates and achieved an AUC of approximately 0.82, predicting sepsis up to 24 hours before clinical suspicion [18]. The risk score increased substantially before diagnosis, highlighting the potential for earlier intervention [5,18]. Integrating continuous physiological signals with EHR, laboratory, and clinical data within multimodal AI frameworks may further improve real-time neonatal risk assessment [15].
Deep learning has further enhanced physiological signal analysis by enabling automated extraction of complex temporal features [34]. Kallonen et al. reported that a model based on electrocardiographic and respiratory impedance signals could predict late-onset sepsis approximately 44 hours before clinical suspicion while maintaining high sensitivity and acceptable specificity [34]. Future decision-support systems may leverage foundation models to integrate heterogeneous physiological and clinical data [22], further strengthening the role of continuous monitoring in real-time AI-based neonatal sepsis prediction [5].

3.5. Multi-Omics and Emerging Data Sources

Although most currently available AI models rely on clinical, laboratory, and physiological data, advances in high-throughput molecular technologies have created new opportunities for precision neonatal medicine. Multi-omics approaches integrate genomics, transcriptomics, proteomics, metabolomics, and other molecular platforms to provide a more comprehensive understanding of disease biology [38]. By simultaneously assessing thousands of biological variables, these technologies may identify molecular signatures associated with infection before clinical symptoms emerge. Their increasing availability has stimulated interest in AI techniques capable of analyzing high-dimensional biological information and integrating molecular and clinical data within unified predictive frameworks [15,29,30,38,39].
Transcriptomic analyses have shown particular promise for early sepsis detection by characterizing host immune responses during the earliest stages of infection. An et al. developed a machine learning model based on gene-expression data and identified a four-gene signature capable of distinguishing neonates who subsequently developed sepsis from those who remained infection-free [29]. Its strong discriminatory performance before overt clinical presentation suggests that transcriptomic biomarkers may provide predictive information beyond conventional laboratory markers [29].
Proteomics has also demonstrated potential in neonatal sepsis prediction. Mithal et al. applied machine learning to cord blood proteomic profiles and identified candidate biomarkers associated with early-onset sepsis [30]. Proteomic signatures combined with logistic regression and random forests achieved clinically meaningful predictive performance while providing insight into underlying disease mechanisms. Because cord blood is available at birth, this approach may enable identification of vulnerable infants before clinical manifestations emerge [30].
Beyond transcriptomics and proteomics, metabolomics, microbiome profiling, and integrated multi-omics frameworks are increasingly being explored for neonatal sepsis prediction [38,39]. By combining multiple layers of biological information, these approaches may capture complex interactions among host immunity, microbial colonization, metabolism, and environmental exposures [38,39]. Advances in multimodal AI and foundation models may further facilitate integration of heterogeneous biological, physiological, and clinical datasets [22,40]. However, cost, standardization, limited sample sizes, and insufficient external validation continue to restrict routine clinical application [7,16,39].
Collectively, these emerging data sources illustrate the transition toward data-driven precision medicine and precision public health [12]. Future predictive models may increasingly integrate clinical, physiological, laboratory, and molecular information within unified multimodal frameworks for individualized neonatal sepsis risk assessment [15]. Figure 1 summarizes the major data sources currently used in AI-based neonatal sepsis prediction and their integration into machine learning and deep learning frameworks to support early clinical decision-making.

4. Machine Learning Models for Neonatal Sepsis Prediction

4.1. Logistic Regression Models

Logistic regression remains a widely used benchmark for clinical prediction modeling because of its simplicity, computational efficiency, and interpretability [11,13]. Despite advances in machine learning, it continues to provide robust performance in structured clinical datasets [5]. In neonatal sepsis prediction, Mani et al. reported competitive discriminatory performance using logistic regression alongside more complex algorithms [31], while Masino et al. demonstrated useful predictive accuracy using EHR-derived data collected several hours before clinical diagnosis [14]. Mithal et al. further applied logistic regression to cord blood proteomic data for early-onset sepsis prediction [30]. However, its limited ability to capture complex nonlinear interactions has encouraged increasing use of more sophisticated machine learning approaches [13].

4.2. Decision Trees and Random Forests

Decision trees classify observations through hierarchical decision rules and can accommodate nonlinear relationships and interactions among variables, although individual trees are prone to overfitting and limited external generalizability [13]. Random forests address these limitations by combining multiple decision trees, thereby improving predictive stability and reducing variance.
Several studies have identified random forests among the best-performing algorithms for neonatal sepsis prediction [5]. Mithal et al. successfully applied random forests to cord blood proteomic data for early-onset sepsis prediction [30], while Garstman et al. reported strong predictive performance for early LOS detection in extremely preterm infants, with a random forest model achieving an AUROC of 0.973 [37]. A recent scoping review identified random forests among the highest-performing approaches across neonatal sepsis machine learning studies [28]. More recently, an externally validated model for early-onset neonatal sepsis demonstrated strong sensitivity using a random forest algorithm with feature selection, supporting its potential utility in reducing unnecessary antibiotic exposure [26].
Random forests can also rank predictor importance, supporting model interpretability and explainable AI approaches [41]. However, individual-level interpretation remains challenging, particularly with large ensembles [13].

4.3. Support Vector Machines

Support vector machines (SVMs) are supervised learning algorithms capable of modeling complex nonlinear relationships through kernel-based approaches, making them suitable for biomedical datasets with intricate interactions among variables [5]. Mani et al. demonstrated strong predictive performance using SVMs in longitudinal NICU data for late-onset sepsis prediction [31], while Gomez et al. reported excellent discrimination between septic and non-septic neonates using heart rate variability data [35].
Despite their predictive capabilities, SVMs are generally less interpretable than logistic regression, highlighting the importance of explainable AI approaches for improving transparency and clinician trust [41]. Their computational complexity may also increase substantially with large datasets. Nevertheless, SVMs remain important benchmark algorithms in neonatal sepsis prediction [13].

4.4. Gradient Boosting and Extreme Gradient Boosting (XGBoost)

Gradient boosting algorithms are among the most effective machine learning techniques for clinical prediction, particularly in structured healthcare datasets. By sequentially combining weak learners, typically decision trees, with each model correcting errors made by its predecessors, these methods can capture complex nonlinear relationships while maintaining strong predictive performance [13]. Consequently, gradient boosting has been increasingly applied in biomedical research, including critical care, sepsis prediction, and outcome forecasting [13,42].
Extreme Gradient Boosting (XGBoost) has gained particular prominence because of its computational efficiency, robustness to missing data, and ability to handle high-dimensional clinical datasets [13]. Yang et al. developed an XGBoost-based framework for continuous prediction of late-onset sepsis in preterm infants using 102 demographic and physiological features [43]. The model achieved the highest performance among the evaluated algorithms, with an AUC of 0.875 six hours before clinical suspicion, outperforming logistic regression, support vector machines, k-nearest neighbors, heart-rate-characteristics-based models, and deep learning architectures [43]. Combining heart rate, respiratory rate, and oxygen saturation further improved discrimination, while the model identified 96.1% of infants who subsequently developed late-onset sepsis before clinical recognition [5,43].
An additional strength of XGBoost is its compatibility with explainable AI techniques. Yang et al. applied SHapley Additive exPlanations (SHAP) to quantify individual feature contributions and identify physiological variables associated with sepsis risk [43]. This combination of predictive performance and interpretability makes XGBoost a particularly promising approach for neonatal sepsis prediction and future clinical decision-support systems [13,43].

4.5. Ensemble Learning Approaches

Ensemble learning combines predictions from multiple models to improve accuracy, robustness, and generalizability. Methods such as random forests and gradient boosting are themselves ensemble techniques, integrating multiple decision trees to reduce variance and capture complex nonlinear interactions [13]. These properties are particularly relevant to neonatal sepsis, where heterogeneous datasets may include demographic characteristics, laboratory biomarkers, electronic health records, and continuous physiological monitoring data [5,13].
Several neonatal sepsis studies have reported strong performance with ensemble approaches [13,39,42], while a neonatal-specific scoping review identified random forests and neural networks among the highest-performing machine learning methods [28]. Nevertheless, increasing model complexity may reduce transparency and complicate clinical implementation [16]. Future research should therefore balance predictive performance with interpretability, external validation, and real-world applicability [16].

4.6. Comparative Overview of Artificial Intelligence Algorithms

No single machine learning or deep learning algorithm consistently outperforms all others across neonatal sepsis datasets and clinical settings [5,9,13,27,42]. Model performance depends on data quality, dataset size, predictor selection, outcome definitions, class imbalance, and validation methodology [5,13]. Accordingly, algorithm selection should balance predictive performance with interpretability, computational complexity, data availability, and clinical applicability [9,11,16].
Logistic regression remains attractive because of its simplicity and transparency, whereas random forests and XGBoost can better capture complex nonlinear interactions and often achieve strong predictive performance [5,13,42]. XGBoost additionally offers computational efficiency and compatibility with explainability techniques such as SHAP [43,44]. Support vector machines perform well in relatively small, high-dimensional datasets but require careful hyperparameter optimization and are generally less interpretable [5,13,41]. Deep learning architectures, including convolutional neural networks and long short-term memory networks, are particularly suited to continuous physiological monitoring, temporal data analysis, and multimodal integration [15], although they typically require larger datasets and rigorous external validation [34,43].
Across published studies, many models have reported AUC values exceeding 0.80 and, in some cases, approaching or surpassing 0.90 [5]. However, substantial heterogeneity in datasets, outcome definitions, and evaluation methods limits direct comparison and underscores the need for standardized validation frameworks [13]. Table 1 summarizes the principal characteristics, strengths, limitations, and potential clinical applications of the machine learning and deep learning approaches discussed in this review.
Overall, no single algorithm is universally superior for neonatal sepsis prediction, and model selection should be guided by the intended clinical application, data availability, interpretability, and validation requirements [21]. While simpler models such as logistic regression offer greater transparency, ensemble and deep learning approaches may achieve higher predictive performance when sufficiently large and diverse datasets are available. Future studies should prioritize standardized reporting and validation frameworks [45], prospective multicenter datasets, and collaborative approaches such as federated learning to improve model robustness, external validity, and clinical translation [19].

5. Deep Learning and Advanced Predictive Analytics

5.1. Artificial Neural Networks

Artificial neural networks (ANNs) form the foundation of modern deep learning and learn complex relationships directly from data through interconnected computational layers. Unlike traditional machine learning algorithms that often require manual feature engineering, ANNs can automatically identify relevant patterns within large and heterogeneous datasets [7]. This capability is particularly relevant to neonatal sepsis, which reflects complex interactions among physiological, immunological, and environmental factors [16].
In neonatal medicine, neural networks have been explored for mortality prediction, respiratory outcome assessment, and sepsis detection [5,46], including the prediction of in-hospital mortality among neonates with clinically suspected sepsis [46]. Their ability to integrate clinical variables, laboratory biomarkers, physiological signals, and molecular information offers substantial potential for neonatal risk stratification [5]. However, their performance remains highly dependent on data quality, sample size, and rigorous validation [13].

5.2. Convolutional Neural Networks

Convolutional neural networks (CNNs), originally developed for image analysis, have increasingly been applied to biomedical signal processing and physiological monitoring. Through automated feature extraction, CNNs can identify complex patterns that may not be readily captured by conventional analytical approaches [7], making them particularly relevant to data-rich NICU environments.
Hu et al. developed a CNN-based framework for early detection of late-onset neonatal sepsis using continuously monitored vital signs [47]. Physiological signals were transformed into image-like representations and analyzed using CNNs, demonstrating the feasibility of deep learning-based prediction from noninvasive monitoring data and its potential application in continuous sepsis surveillance [47].
More recently, Kallonen et al. reported a prospective multicenter deep learning study using noninvasive biosignals for early detection of late-onset neonatal sepsis [34]. Their multimodal CNN-based model achieved an AUC of approximately 0.81 in external validation and relied exclusively on continuously acquired biosignals, supporting the feasibility of noninvasive and automated risk assessment within routine NICU workflows [34].
Collectively, these findings suggest that CNNs may extract clinically relevant information from complex physiological signals and contribute to future neonatal sepsis prediction systems, particularly when integrated with complementary clinical and laboratory data [34,47].

5.3. Recurrent Neural Networks and Long Short-Term Memory Networks

Recurrent neural networks (RNNs) are designed to analyze sequential data and capture temporal relationships among observations, making them particularly relevant to neonatal sepsis, where physiological deterioration develops progressively over time [7]. By incorporating information from previous observations, RNNs can model dynamic physiological processes and identify evolving patterns associated with impending infection.
Long short-term memory (LSTM) networks were developed to overcome the limitations of traditional RNNs in learning long-range temporal dependencies and have become widely used for healthcare time-series analysis. Yang et al. evaluated an attention-based LSTM model for continuous prediction of late-onset neonatal sepsis using physiological monitoring data [43]. Although the model demonstrated competitive performance, XGBoost achieved superior discrimination within the same dataset, illustrating that more complex deep learning architectures do not necessarily outperform well-designed machine learning models, particularly when dataset sizes are relatively modest [43].
Kallonen et al. further demonstrated the value of temporal deep learning using electrocardiographic and respiratory impedance signals, predicting late-onset sepsis approximately 44 hours before clinical suspicion [34]. These findings indicate that continuous physiological monitoring contains substantial predictive information that can be leveraged through advanced neural network architectures [5,34].

5.4. Real-Time Predictive Analytics and Early Warning Systems

Real-time predictive analytics represents one of the most promising applications of AI in neonatal care. Unlike conventional diagnostic approaches based on intermittent clinical assessment, AI-driven early warning systems can continuously process streaming physiological data and dynamically update sepsis risk as new information becomes available [18]. Such systems may support clinicians by identifying early signs of deterioration [8], while recent work has increasingly focused on translating predictive algorithms into clinical decision-support systems for late-onset sepsis and other critical neonatal conditions [33].
Predictive monitoring has already demonstrated clinical value in neonatal care [8,16]. Earlier heart rate characteristic studies showed that abnormal physiological patterns may precede clinical recognition of sepsis [8,35], while more recent AI approaches integrate multiple physiological signals, demographic characteristics, and clinical variables [18,34,43]. Studies by Honoré et al., Kallonen et al., and Yang et al. suggest that continuous monitoring may enable clinically meaningful prediction hours before conventional recognition [12,34,43]. Similar advances across the broader sepsis field further highlight the potential of AI for early detection, real-time monitoring, and personalized clinical decision support [17].
Despite encouraging results, prospective multicenter validation, integration into clinical workflows, management of false-positive alerts, and demonstration of improved patient outcomes remain essential for successful clinical translation [13,16]. Addressing these challenges will determine whether AI-based early warning systems can provide meaningful clinical benefit without increasing alert fatigue or complicating neonatal decision-making.

6. Current Evidence and Performance of AI Models

The application of artificial intelligence to neonatal sepsis prediction has expanded substantially over the past decade, driven by advances in computational methodologies and the increasing availability of neonatal clinical data. Recent systematic and scoping reviews have identified a growing number of machine learning and deep learning studies using demographic characteristics, laboratory biomarkers, physiological monitoring signals, electronic health records, and multi-omics data for early sepsis detection [4,5,28]. Collectively, these studies suggest that AI-based approaches may improve risk stratification and facilitate earlier recognition of neonatal sepsis [4,5,8,9,13].
Models integrating multiple data sources generally outperform individual clinical variables or isolated biomarkers [5], while recent studies have increasingly focused on identifying the most informative predictors and improving model explainability [41,44]. Across published studies, predictive performance frequently exceeds an AUC of 0.80 and, in some cases, approaches or surpasses 0.90 [5,13]. Random forests, support vector machines, gradient boosting, and ensemble methods have demonstrated strong performance, while deep learning architectures increasingly enable the analysis of high-dimensional, continuously generated, and multimodal clinical data [13,15].
Several influential studies illustrate the potential of AI-driven neonatal sepsis prediction. Gomez et al. achieved an AUC approaching 0.94 using heart rate variability-derived features [35], while Masino et al. demonstrated that EHR-based machine learning models could identify neonates at increased risk of sepsis several hours before clinical diagnosis [14]. Yang et al. subsequently reported an AUC of 0.875 using an XGBoost-based framework for late-onset sepsis prediction in preterm infants [43]. More recently, an externally validated random forest model demonstrated high sensitivity for early-onset neonatal sepsis and potential utility in reducing unnecessary antibiotic exposure [26]. Collectively, these findings indicate that clinically meaningful predictive information can be extracted from routinely collected neonatal data [14,24,43].
Continuous physiological monitoring has produced some of the most clinically compelling results. Song et al. predicted late-onset sepsis up to 48 hours before diagnosis using routinely monitored vital signs [36], Honoré et al. reported successful prediction approximately 24 hours before clinical suspicion [18], and Kallonen et al. achieved prediction horizons of approximately 44 hours using electrocardiographic and respiratory impedance signals [34]. Together, these studies suggest that subtle physiological disturbances may precede overt clinical manifestations and provide opportunities for earlier intervention.
Molecular and multi-omics approaches have also shown promise, although the evidence remains limited. An et al. identified a four-gene transcriptomic signature capable of distinguishing neonates who subsequently developed sepsis from those who remained infection-free [29], while Mithal et al. demonstrated the potential of cord blood proteomic biomarkers combined with machine learning for early-onset sepsis prediction [30]. These findings support the potential role of molecular profiling in earlier and more individualized sepsis prediction within future multimodal AI systems [6,15].
Despite encouraging findings, substantial heterogeneity in study populations, sepsis definitions, predictor variables, outcomes, and validation methodologies limits direct comparison across studies [4,13,48]. Most models remain based on retrospective, often single-center datasets, restricting generalizability, while algorithmic bias, limited transparency, and reproducibility remain important barriers to implementation [20,21]. Standardized evaluation and reporting frameworks [45], together with prospective multicenter validation, remain essential before widespread clinical implementation [5,13]. Table 2 summarizes representative AI studies in neonatal sepsis prediction, highlighting the diversity of data sources, analytical methods, and reported performance.

7. Explainability, Challenges, and Clinical Implementation

7.1. Explainable Artificial Intelligence

As artificial intelligence systems become increasingly sophisticated, transparency and interpretability have emerged as major barriers to clinical adoption [16,20,21]. Many high-performing machine learning and deep learning models function as “black boxes,” generating predictions without clearly explaining the factors underlying their outputs. This limitation is particularly important in high-stakes scenarios such as neonatal sepsis, where decisions regarding antibiotic administration, diagnostic investigations, and intensive monitoring may have significant clinical consequences [43]. Consequently, explainability has become a key requirement for trustworthy clinical AI systems [16,49].
Explainable artificial intelligence (XAI) seeks to provide insight into how predictive models generate their outputs. Among the most widely used approaches are SHapley Additive exPlanations (SHAP) and Local Interpretable Model-Agnostic Explanations (LIME) [50,51]. SHAP quantifies the contribution of individual variables to specific predictions, whereas LIME generates simplified local explanations of model behavior [43]. These approaches enable assessment of whether model predictions are clinically plausible and biologically meaningful, while recent studies have also focused on identifying the most influential predictors of neonatal sepsis [41,44].
Recent neonatal sepsis studies have increasingly incorporated explainability into model development. Yang et al. applied SHAP analysis to identify variables contributing most strongly to sepsis prediction, demonstrating the importance of clinically relevant physiological and laboratory features [43]. Such approaches may enhance clinician confidence, facilitate model validation, and provide insight into disease pathophysiology. However, simplified explanations may not fully capture the complexity of advanced models, and future research must balance predictive performance with transparency and clinical interpretability [16,21,49].

7.2. Methodological Challenges and Limitations

Despite encouraging results, important methodological limitations continue to restrict the clinical applicability of AI-based neonatal sepsis prediction models [13,51,52]. Similar challenges have been identified across the broader sepsis-AI literature, including small sample sizes, dataset heterogeneity, limited external validation, and difficulties in translating predictive models into clinical practice [42]. Because neonatal sepsis is a relatively infrequent outcome within individual institutions, datasets often contain limited numbers of confirmed cases, increasing the risk of overfitting and reducing generalizability [5,13,20]. This limitation is particularly relevant to deep learning approaches, which typically require large training datasets.
Most available models have also been developed retrospectively at single centers, despite substantial differences among NICUs in patient populations, pathogen epidemiology, clinical practices, antimicrobial stewardship, and monitoring technologies [4,13]. Consequently, models demonstrating strong performance in their development cohort may experience substantial reductions in predictive accuracy when applied across different hospitals, healthcare systems, or geographic regions [16,42,48]. Robust external validation across diverse populations is therefore essential to establish generalizability, ensure equitable performance, and support safe clinical implementation [21,45,48,53]. Recent external validation of a machine learning model for early-onset neonatal sepsis in a distinct hospital population illustrates an important step toward addressing this evidence gap [26].
Future prospective multicenter studies involving geographically diverse neonatal populations will be essential to determine whether AI models can be reliably translated into routine neonatal care [16,44,48]. Collaborative approaches such as federated learning may further facilitate multicenter model development while preserving patient privacy and reducing barriers to interinstitutional data sharing [19].
Data quality and outcome heterogeneity represent additional challenges. Missing values, inconsistent documentation, measurement errors, and variability in data collection procedures may substantially influence model performance [5]. Furthermore, varying definitions of neonatal sepsis—including culture-confirmed infection, clinically suspected sepsis, and composite outcomes—complicate comparisons across studies and the identification of optimal predictive approaches [5,48]. Standardized definitions, reporting frameworks, and validation procedures should therefore be prioritized [5,45,48].
Reporting practices also remain inconsistent, with limited adherence to standardized frameworks such as the TRIPOD+AI Statement [45] and risk-of-bias assessment tools such as PROBAST [54]. Beyond methodological rigor, future implementation must consider ethical governance, transparency, and responsible deployment of AI systems in neonatal care [23]. Standardized reporting, prospective multicenter validation, and rigorous evaluation of clinical utility will be essential before widespread implementation of AI-based neonatal sepsis prediction systems can be achieved [5,13,48].

7.3. Ethical, Regulatory, and Clinical Considerations

The implementation of AI-driven decision-support systems in neonatal care raises important ethical, regulatory, and governance considerations [16,23,55]. Because predictive algorithms may influence critical treatment decisions in highly vulnerable patients, reliability, safety, and accountability are essential requirements [16,21]. Algorithmic bias remains a particular concern, as models trained on specific populations may perform less accurately in underrepresented groups and potentially contribute to healthcare disparities [16,20,56].
Regulatory oversight of clinical AI continues to evolve, particularly because machine learning models may be updated as new data become available. Demonstrating safety, reproducibility, clinical effectiveness, and real-world benefit therefore requires evaluation beyond conventional performance metrics [7]. Healthcare institutions must also address data privacy, cybersecurity, informed consent, governance, and infrastructure readiness when implementing AI technologies [7,23,55].
Clinician acceptance is another key determinant of successful implementation [16,51]. Even highly accurate systems may fail if perceived as unreliable, difficult to interpret, or disruptive to established workflows. Effective EHR integration, minimization of alert fatigue, clinically meaningful explanations, and seamless incorporation into clinical practice are therefore essential [43,49,55]. Ultimately, AI should complement rather than replace clinical judgment, enabling clinicians to combine algorithmic insights with professional expertise and individual patient context [16].
Beyond discrimination metrics such as the area under the receiver operating characteristic curve (AUC), successful implementation requires evaluation of calibration, clinical utility, and real-world impact [48,55]. Decision curve analysis can assess net clinical benefit across different thresholds, while prospective impact studies are needed to demonstrate improvements in diagnostic accuracy, antibiotic stewardship, patient outcomes, and resource utilization [16,48,55,56]. Future research should therefore move beyond predictive accuracy alone and prioritize clinical utility, cost-effectiveness, and integration into routine neonatal care using standardized reporting and evaluation frameworks such as TRIPOD+AI [41,48,55].

8. Future Directions

The future of artificial intelligence in neonatal sepsis prediction will likely be driven by advances in multimodal data integration, large-scale collaborative learning, foundation models, and real-time clinical decision-support systems [15,19,22]. Although most current models rely on individual data sources, emerging evidence suggests that integrating clinical, physiological, laboratory, and molecular information within unified predictive frameworks may improve predictive performance and provide a more comprehensive representation of neonatal health status [7,15,16]. Such multimodal approaches could enable individualized risk assessment and support the development of more generalizable AI systems capable of adapting across diverse clinical settings [39,40].
As AI moves toward routine clinical implementation, effective integration into neonatal intensive care workflows will become increasingly important. Beyond individual patient care, AI may also contribute to precision public health by supporting earlier diagnosis, targeted antimicrobial stewardship, and optimized allocation of healthcare resources [12]. Figure 2 illustrates a proposed clinical workflow for AI-assisted neonatal sepsis prediction, highlighting the integration of multimodal data, risk stratification, and clinical decision support.
Another promising area is federated learning and collaborative multicenter model development [13,19,52]. Because most neonatal sepsis algorithms have been trained on relatively small, single-center datasets, concerns regarding generalizability remain substantial [13]. Federated learning enables institutions to collaboratively train machine learning models without directly sharing patient-level data, facilitating access to larger and more diverse datasets while preserving privacy and regulatory compliance [19]. As neonatal research networks expand, this approach may support the development of more robust and externally validated AI systems [13,19].
Rapid advances in computational medicine are also expected to accelerate the development of next-generation predictive systems [55,57]. Foundation models and generative AI trained on large-scale healthcare datasets may enable transfer learning across multiple neonatal conditions, reducing dependence on large disease-specific cohorts and potentially improving performance in data-limited settings [22]. These models represent an emerging paradigm for generalist medical AI, with potential applications across diverse clinical prediction tasks [40,58]. Similarly, digital twin approaches may eventually enable individualized computational representations of neonatal patients, supporting dynamic simulation of disease trajectories and estimation of sepsis risk [59]. Although these technologies remain largely experimental, their development and implementation will require appropriate ethical oversight and governance [23]. Collectively, these advances illustrate the convergence of AI, systems biology, and precision medicine and may ultimately support more personalized neonatal care [15,55,58,59].
The increasing availability of transcriptomic, proteomic, metabolomic, and microbiome-derived data may further advance neonatal sepsis prediction [38,39]. Combined with machine learning and computational biology, these technologies could support precision medicine approaches capable of identifying vulnerable infants before clinical deterioration [16,38,39]. Ultimately, future AI systems may evolve from passive prediction tools into dynamic clinical decision-support platforms that integrate multimodal data, continuously update individualized risk estimates, and assist clinicians in optimizing diagnostic and therapeutic strategies [15,16,55,57]. Despite persistent methodological, ethical, regulatory, and implementation challenges, advances in AI and multimodal data integration may increasingly support earlier detection, risk stratification, and individualized management of neonatal sepsis, while contributing to broader precision medicine and precision public health initiatives [12,55,57].

9. Conclusions

Artificial intelligence is rapidly transforming neonatal sepsis prediction by enabling the integration of clinical, laboratory, physiological, and molecular data into individualized risk assessment models. Current evidence suggests that machine learning and deep learning approaches may facilitate earlier sepsis recognition, support timely clinical intervention, and advance precision neonatal care. However, important challenges remain, including limited external validation, dataset heterogeneity, model interpretability, and ethical and regulatory considerations. Future progress will depend on rigorous multicenter validation, standardized reporting, explainable and trustworthy AI, and seamless integration into clinical workflows. Rather than replacing clinical expertise, AI should serve as a decision-support tool that complements clinician judgment and supports evidence-based care. With continued technological advances and collaborative research, AI has the potential to shift neonatal sepsis management from reactive diagnosis toward proactive, personalized, and data-driven care.

Author Contributions

Conceptualization, A.I.N. and V.G.; methodology, A.I.N. and V.G.; validation, A.I.N. and V.G.; formal analysis, A.I.N.; investigation, A.I.N.; writing—original draft preparation, A.I.N.; writing—review and editing, A.I.N., N.D., N.C., M.B., N.G.P., G.C.G., S.G. and V.G.; supervision, V.G. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Ethics approval was not required for this review study.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Shane, A.L.; Sánchez, P.J.; Stoll, B.J. Neonatal sepsis. Lancet 2017, 390, 1770–1780. [Google Scholar] [CrossRef] [PubMed]
  2. Fleischmann-Struzek, C.; Goldfarb, D.M.; Schlattmann, P.; Schlapbach, L.J.; Reinhart, K.; Kissoon, N. The global burden of paediatric and neonatal sepsis: A systematic review. Lancet Respir. Med. 2018, 6, 223–230. [Google Scholar] [CrossRef] [PubMed]
  3. De Rose, D.U.; Ronchetti, M.P.; Martini, L.; Rechichi, J.; Iannetta, M.; Dotta, A.; Auriti, C. Diagnosis and management of neonatal bacterial sepsis: Current challenges and future perspectives. Trop. Med. Infect. Dis. 2024, 9, 199. [Google Scholar] [CrossRef] [PubMed]
  4. Sahu, P.; Stanly, E.A.R.; Lewis, L.E.S.; Prabhu, K.; Rao, M.; Kunhikatta, V. Prediction modelling in the early detection of neonatal sepsis. World J. Pediatr. 2022, 18, 160–175. [Google Scholar] [CrossRef] [PubMed]
  5. Rao, K.V.K.L.N.; Dadabada, P.K.; Jaipuria, S. A systematic literature review of predictive analytics methods for early diagnosis of neonatal sepsis. Discov. Public Health 2024, 21, 96. [Google Scholar] [CrossRef]
  6. Stocker, M.; Fillistorf, L.; Carra, G.; Giannoni, E. Early detection of neonatal sepsis and reduction of overall antibiotic exposure: Towards precision medicine. Arch. Pediatr. 2024, 31, 480–483. [Google Scholar] [CrossRef] [PubMed]
  7. McAdams, R.M.; Kaur, R.; Sun, Y.; Bindra, H.; Cho, S.J.; Singh, H. Predicting clinical outcomes using artificial intelligence and machine learning in neonatal intensive care units: A systematic review. J. Perinatol. 2022, 42, 1561–1575. [Google Scholar] [CrossRef] [PubMed]
  8. Sullivan, B.A.; Kausch, S.L.; Fairchild, K.D. Artificial and human intelligence for early identification of neonatal sepsis. Pediatr. Res. 2023, 93, 350–356. [Google Scholar] [CrossRef] [PubMed]
  9. Topol, E.J. High-performance medicine: The convergence of human and artificial intelligence. Nat. Med. 2019, 25, 44–56. [Google Scholar] [CrossRef] [PubMed]
  10. Esteva, A.; Robicquet, A.; Ramsundar, B.; Kuleshov, V.; DePristo, M.; Chou, K.; Cui, C.; Corrado, G.; Thrun, S.; Dean, J. A guide to deep learning in healthcare. Nat. Med. 2019, 25, 24–29. [Google Scholar] [CrossRef] [PubMed]
  11. Rajkomar, A.; Dean, J.; Kohane, I. Machine learning in medicine. N. Engl. J. Med. 2019, 380, 1347–1358. [Google Scholar] [CrossRef] [PubMed]
  12. Panteli, D.; Adib, K.; Buttigieg, S.; Goiana-da-Silva, F.; Ladewig, K.; Azzopardi-Muscat, N.; Figueras, J.; Novillo-Ortiz, D. Artificial intelligence in public health: Promises, challenges, and an agenda for policy makers and public health institutions. Lancet Public Health 2025, 10, e428–e432. [Google Scholar] [CrossRef] [PubMed]
  13. Kainth, D.; Prakash, S.; Sankar, M.J. Diagnostic Performance of Machine Learning-based Models in Neonatal Sepsis: A Systematic Review. Pediatr. Infect. Dis. J. 2024. [Google Scholar] [CrossRef] [PubMed]
  14. Masino, A.J.; Harris, M.C.; Forsyth, D.; Ostapenko, S.; Srinivasan, L.; Bonafide, C.P.; Balamuth, F.; Schmatz, M.; Grundmeier, R.W. Machine learning models for early sepsis recognition in the neonatal intensive care unit using readily available electronic health record data. PLoS ONE 2019, 14, e0212665. [Google Scholar] [CrossRef] [PubMed]
  15. Schouten, D.; Nicoletti, G.; Dille, B.; Chia, C.; Vendittelli, P.; Schuurmans, M.; Litjens, G.; Khalili, N. Navigating the landscape of multimodal AI in medicine: A scoping review on technical challenges and clinical applications. Med. Image Anal. 2025, 105, 103621. [Google Scholar] [CrossRef] [PubMed]
  16. Sullivan, B.A.; Beam, K.; Vesoulis, Z.A.; Aziz, K.B.; Husain, A.N.; Knake, L.A.; Moreira, A.; Hooven, T.A.; Weiss, E.M.; Carr, N.R.; et al. Transforming neonatal care with artificial intelligence: Challenges, ethical consideration, and opportunities. J. Perinatol. 2024, 44, 1–11. [Google Scholar] [CrossRef] [PubMed]
  17. Li, F.; Wang, S.; Gao, Z.; Qing, M.; Pan, S.; Liu, Y.; Hu, C. Harnessing artificial intelligence in sepsis care: Advances in early detection, personalized treatment, and real-time monitoring. Front. Med. 2025, 11, 1510792. [Google Scholar] [CrossRef] [PubMed]
  18. Honoré, A.; Forsberg, D.; Adolphson, K.; Chatterjee, S.; Jost, K.; Herlenius, E. Vital sign-based detection of sepsis in neonates using machine learning. Acta Paediatr. 2023, 112, 686–696. [Google Scholar] [CrossRef] [PubMed]
  19. Miloudi, A.; Laouid, A.; Bounceur, A.; Kara, M.; Bouhamed, M.M.; Hammoudeh, M.; Kraidia, I. Federated learning in healthcare: Recent progress and challenges. Comput. Electr. Eng. 2026, 131, 110924. [Google Scholar] [CrossRef]
  20. Mittermaier, M.; Raza, M.M.; Kvedar, J.C. Bias in AI-based models for medical applications: Challenges and mitigation strategies. npj Digit. Med. 2023, 6, 113. [Google Scholar] [CrossRef] [PubMed]
  21. Kowald, D.; Scher, S.; Pammer-Schindler, V.; Müllner, P.; Waxnegger, K.; Demelius, L.; Toller, M.; Šimić, I.; Kopeinik, S.; et al. Establishing and evaluating trustworthy AI: Overview and research challenges. Front. Big Data 2024, 7, 1467222. [Google Scholar] [CrossRef] [PubMed]
  22. Sitek, A.; Bates, D.W. Beyond language: Generative artificial intelligence as a general computing model for medicine. Lancet Digit. Health 2026, 8, 101011. [Google Scholar] [CrossRef] [PubMed]
  23. Ranisch, R.; Haltaufderheide, J. Foundation models in medicine are a social experiment: Time for an ethical framework. npj Digit. Med. 2025, 8, 525. [Google Scholar] [CrossRef] [PubMed]
  24. Seyhanlı, D.; Yıldırım, T.G.; Kalkanlı, O.H.; Soysal, B.; Alkan Özdemir, S.; Devrim, I.; Çalkavur, Ş. Prediction model for early diagnosis of late-onset sepsis in preterm newborns. J. Neonatal Perinat. Med. 2024, 17, 661–671. [Google Scholar] [CrossRef] [PubMed]
  25. Kuzniewicz, M.W.; Puopolo, K.M.; Fischer, A.; Walsh, E.M.; Li, S.; Newman, T.B.; Kipnis, P.; Escobar, G.J. A quantitative, risk-based approach to the management of neonatal early-onset sepsis. JAMA Pediatr. 2017, 171, 365–371. [Google Scholar] [CrossRef] [PubMed]
  26. Kainth, D.; Gupta, A.; Singh, P.; Prakash, S.; Thukral, A.; Deorari, A.; et al. A machine learning model for prediction of early-onset neonatal sepsis in low-income and middle-income countries: Development and validation study. BMJ Paediatr. Open 2026, 10, e003561. [Google Scholar] [CrossRef] [PubMed]
  27. Goldberg, O.; Amitai, N.; Chodick, G.; Bromiker, R.; Scheuerman, O.; Ben-Zvi, H.; Klinger, G. Can we improve early identification of neonatal late-onset sepsis? A validated prediction model. J. Perinatol. 2020, 40, 1315–1322. [Google Scholar] [CrossRef] [PubMed]
  28. O’Sullivan, C.; Tsai, D.H.T.; Wu, I.C.Y.; et al. Machine learning applications on neonatal sepsis treatment: A scoping review. BMC Infect. Dis. 2023, 23, 441. [Google Scholar] [CrossRef] [PubMed]
  29. An, A.Y.; Acton, E.; Idoko, O.T.; Shannon, C.P.; Blimkie, T.M.; Falsafi, R.; Wariri, O.; Imam, A.; Dibbasey, T.; Bennike, T.B.; et al. Predictive gene expression signature diagnoses neonatal sepsis before clinical presentation. EBioMedicine 2024, 110, 105411. [Google Scholar] [CrossRef] [PubMed]
  30. Mithal, L.B.; Becker, M.E.; Ling-Hu, T.; Goo, Y.A.; Otero, S.; Kremer, A.; Pandey, S.; Lancki, N.; Li, Y.; Luo, Y.; et al. Cord blood proteomics identifies biomarkers of early-onset neonatal sepsis. JCI Insight 2025, 10, e193826. [Google Scholar] [CrossRef] [PubMed]
  31. Mani, S.; Ozdas, A.; Aliferis, C.; Varol, H.A.; Chen, Q.; Carnevale, R.; Chen, Y.; Romano-Keeler, J.; Nian, H.; Weitkamp, J.-H. Medical decision support using machine learning for early detection of late-onset neonatal sepsis. J. Am. Med. Inform. Assoc. 2014, 21, 326–336. [Google Scholar] [CrossRef] [PubMed]
  32. Iqbal, F.; Chandra, P.; Lewis, L.E.S.; Acharya, D.; Purkayastha, J.; Shenoy, P.A.; Patil, A.K. Application of artificial intelligence to predict the sepsis in neonates admitted in neonatal intensive care unit. J. Neonatal Nurs. 2023. [Google Scholar] [CrossRef]
  33. Meeus, M.; Beirnaert, C.; Mahieu, L.; Laukens, K.; Meysman, P.; Mulder, A.; Van Laere, D. Clinical decision support for improved neonatal care: The development of a machine learning model for the prediction of late-onset sepsis and necrotizing enterocolitis. J. Pediatr. 2024, 266, 113869. [Google Scholar] [CrossRef] [PubMed]
  34. Kallonen, A.; Juutinen, M.; Värri, A.; Carrault, G.; Pladys, P.; Beuchée, A. Early detection of late-onset neonatal sepsis from noninvasive biosignals using deep learning: A multicenter prospective development and validation study. Int. J. Med. Inform. 2024, 184, 105366. [Google Scholar] [CrossRef] [PubMed]
  35. Gomez, R.; Garcia, N.; Collantes, G.; Ponce, F.; Redon, P. Development of a non-invasive procedure to early detect neonatal sepsis using HRV monitoring and machine learning algorithms. In Proceedings of the 2019 IEEE 32nd International Symposium on Computer-Based Medical Systems (CBMS), Cordoba, Spain, 5–7 June 2019; pp. 132–137. [Google Scholar] [CrossRef]
  36. Song, W.; Jung, S.Y.; Baek, H.; Choi, C.W.; Jung, Y.H.; Yoo, S. A predictive model based on machine learning for the early detection of late-onset neonatal sepsis: Development and observational study. JMIR Med. Inform. 2020, 8, e15965. [Google Scholar] [CrossRef] [PubMed]
  37. Garstman, A.G.; Rodriguez Rivero, C.; Onland, W. Early Detection of Late Onset Sepsis in Extremely Preterm Infants Using Machine Learning: Towards an Early Warning System. Appl. Sci. 2023, 13, 9049. [Google Scholar] [CrossRef]
  38. Hasin, Y.; Seldin, M.; Lusis, A. Multi-omics approaches to disease. Genome Biol. 2017, 18, 83. [Google Scholar] [CrossRef] [PubMed]
  39. Duci, M.; Verlato, G.; Moschino, L.; Uccheddu, F.; Fascetti-Leon, F. Advances in artificial intelligence and machine learning for precision medicine in necrotizing enterocolitis and neonatal sepsis: A state-of-the-art review. Children 2025, 12, 498. [Google Scholar] [CrossRef] [PubMed]
  40. Vishwanath, K.; Alyakin, A.; Ghosh, M.; Hage, A.; Neifert, S.N.; Orillac, C.; Mandelberg, N.J.; Khan, H.A.; Lee, J.V.; Yao, J.J.; Small, W.R.; Varma, A.; Hewitt, D.B.; Aphinyanaphongs, Y.; Alber, D.A.; Oermann, E.K. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. Nat. Med. 2026. [Google Scholar] [CrossRef] [PubMed]
  41. Sadeghi, Z.; Alizadehsani, R.; Cifci, M.A.; Kausar, S.; Rehman, R.; Mahanta, P.; Bora, P.K.; Almasri, A.; Alkhawaldeh, R.S.; Hussain, S.; Alatas, B.; Shoeibi, A.; Moosaei, H.; Hladík, M.; Nahavandi, S.; Pardalos, P.M. A review of explainable artificial intelligence in healthcare. Comput. Electr. Eng. 2024, 118, 109370. [Google Scholar] [CrossRef]
  42. Tądel, K.; Dudek, A.; Bil-Lula, I. AI algorithms for modeling the risk, progression, and treatment of sepsis, including early-onset sepsis—A systematic review. J. Clin. Med. 2024, 13, 5959. [Google Scholar] [CrossRef] [PubMed]
  43. Yang, M.; Peng, Z.; van Pul, C.; Andriessen, P.; Dong, K.; Silvertand, D.; Li, J.; Liu, C.; Long, X. Continuous prediction and clinical alarm management of late-onset sepsis in preterm infants using vital signs from a patient monitor. Comput. Methods Programs Biomed. 2024, 255, 108335. [Google Scholar] [CrossRef] [PubMed]
  44. de Morais, F.L.; da Silva Canejo, S.P.; de Mello, M.E.F.; da Silva, R.C.L.; da Silva Barros, M.H.L.F.; da Silva Rocha, E.; Mendes, K.M.; Brandão Neto, W.; Endo, P.T. On the usage of artificial intelligence for identifying main attributes and predicting neonatal sepsis. Sci. Rep. 2026. [Google Scholar] [CrossRef] [PubMed]
  45. Collins, G.S.; Moons, K.G.M.; Dhiman, P.; Riley, R.D.; Beam, A.L.; Van Calster, B.; Ghassemi, M.; Liu, X.; Reitsma, J.B.; van Smeden, M.; et al. TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024, 385, e078378. [Google Scholar] [CrossRef] [PubMed]
  46. Hsu, J.-F.; Chang, Y.-F.; Cheng, H.-J.; Yang, C.; Lin, C.-Y.; Chu, S.-M.; Huang, H.-R.; Chiang, M.-C.; Wang, H.-C.; Tsai, M.-H. Machine learning approaches to predict in-hospital mortality among neonates with clinically suspected sepsis in the neonatal intensive care unit. J. Pers. Med. 2021, 11, 695. [Google Scholar] [CrossRef] [PubMed]
  47. Hu, Y.; Lee, V.C.S.; Tan, K. An application of convolutional neural networks for the early detection of late-onset neonatal sepsis. In Proceedings of the 2019 International Joint Conference on Neural Networks (IJCNN), Budapest, Hungary, 14–19 July 2019; IEEE: Piscataway, NJ, USA, 2019; pp. 5661–5668. [Google Scholar] [CrossRef]
  48. Husain, A.; Knake, L.; Sullivan, B.; Barry, J.; Beam, K.; Holmes, E.; Hooven, T.; McAdams, R.; Moreira, A.; Shalish, W.; et al. AI models in clinical neonatology: A review of modeling approaches and a consensus proposal for standardized reporting of model performance. Pediatr. Res. 2025, 98, 412–422. [Google Scholar] [CrossRef] [PubMed]
  49. van Leersum, C.M.; Maathuis, C. Human centred explainable AI decision-making in healthcare. J. Responsible Technol. 2025, 21, 100108. [Google Scholar] [CrossRef]
  50. Lundberg, S.M.; Lee, S.-I. A unified approach to interpreting model predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS 2017), Long Beach, CA, USA, 4–9 December 2017; pp. 4768–4777. [Google Scholar]
  51. Ribeiro, M.T.; Singh, S.; Guestrin, C. “Why Should I Trust You?”: Explaining the Predictions of Any Classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD ‘16); Association for Computing Machinery: New York, NY, USA, 2016; pp. 1135–1144. [Google Scholar] [CrossRef]
  52. Kainth, D.; Agarwal, R. Artificial intelligence in neonatal sepsis: Scope, challenges, and potential solutions! Semin. Fetal Neonatal Med. 2026, 31, 101687. [Google Scholar] [CrossRef] [PubMed]
  53. Tądel, K.; Dudek, A.; Bil-Lula, I. AI Algorithms for Modeling the Risk, Progression, and Treatment of Sepsis, Including Early-Onset Sepsis—A Systematic Review. J. Clin. Med. 2024, 13, 5959. [Google Scholar] [CrossRef] [PubMed]
  54. Wolff, R.F.; Moons, K.G.M.; Riley, R.D.; Whiting, P.F.; Westwood, M.; Collins, G.S.; Reitsma, J.B.; Kleijnen, J.; Mallett, S.; PROBAST Group. PROBAST: A Tool to Assess the Risk of Bias and Applicability of Prediction Model Studies. Ann. Intern. Med. 2019, 170, 51–58. [Google Scholar] [CrossRef] [PubMed]
  55. Barrett, R.; Lawler, B.; Liu, S.; Park, W.Y.; Davoodi, M.; Martin, B.; Kalyanam, S.M.; Makker, K.; Kuiper, J.R.; Aziz, K.B. Transforming neonatal care through informatics: A review of artificial intelligence, data, and implementation considerations. Semin. Perinatol. 2025, 49, 152144. [Google Scholar] [CrossRef] [PubMed]
  56. Vickers, A.J.; Elkin, E.B. Decision curve analysis: A novel method for evaluating prediction models. Med. Decis. Mak. 2006, 26, 565–574. [Google Scholar] [CrossRef] [PubMed]
  57. El Arab, R.A.; Al Moosa, O.A.; Albahrani, Z.; Alkhalil, I.; Somerville, J.; Abuadas, F. Integrating artificial intelligence into perinatal care pathways: A scoping review of reviews of applications, outcomes, and equity. Nurs. Rep. 2025, 15, 281. [Google Scholar] [CrossRef] [PubMed]
  58. Moor, M.; Banerjee, O.; Hossein Abad, Z.S.; Krumholz, H.M.; Leskovec, J.; Topol, E.J.; Rajpurkar, P. Foundation models for generalist medical artificial intelligence. Nature 2023, 616, 259–265. [Google Scholar] [CrossRef] [PubMed]
  59. Corral-Acero, J.; Margara, F.; Marciniak, M.; Rodero, C.; Loncaric, F.; Feng, Y.; Gilbert, A.; Fernandes, J.F.; Bukhari, H.A.; Wajdan, A.; et al. The “Digital Twin” to enable the vision of precision cardiology. Eur. Heart J. 2020, 41, 4556–4564. [Google Scholar] [CrossRef] [PubMed]
Figure 1. Conceptual framework of artificial intelligence–based neonatal sepsis prediction. Clinical and demographic characteristics, laboratory biomarkers, electronic health records, physiological monitoring signals, and multi-omics datasets serve as input sources for machine learning and deep learning algorithms. These models generate individualized sepsis risk estimates that may support earlier diagnosis, antibiotic stewardship, and precision neonatal care. Abbreviations: AI, artificial intelligence; CRP, C-reactive protein; DL, deep learning; EHRs, electronic health records; IL-6, interleukin-6; ML, machine learning; PCT, procalcitonin.
Figure 1. Conceptual framework of artificial intelligence–based neonatal sepsis prediction. Clinical and demographic characteristics, laboratory biomarkers, electronic health records, physiological monitoring signals, and multi-omics datasets serve as input sources for machine learning and deep learning algorithms. These models generate individualized sepsis risk estimates that may support earlier diagnosis, antibiotic stewardship, and precision neonatal care. Abbreviations: AI, artificial intelligence; CRP, C-reactive protein; DL, deep learning; EHRs, electronic health records; IL-6, interleukin-6; ML, machine learning; PCT, procalcitonin.
Preprints 223189 g001
Figure 2. Proposed clinical workflow for artificial intelligence-assisted neonatal sepsis prediction in the neonatal intensive care unit. Clinical, laboratory, physiological, and electronic health record data are integrated into artificial intelligence models to generate individualized sepsis risk estimates that support early clinical assessment, timely therapeutic interventions, and antibiotic stewardship, ultimately contributing to improved neonatal outcomes.
Figure 2. Proposed clinical workflow for artificial intelligence-assisted neonatal sepsis prediction in the neonatal intensive care unit. Clinical, laboratory, physiological, and electronic health record data are integrated into artificial intelligence models to generate individualized sepsis risk estimates that support early clinical assessment, timely therapeutic interventions, and antibiotic stewardship, ultimately contributing to improved neonatal outcomes.
Preprints 223189 g002
Table 1. Comparative characteristics of machine learning and deep learning algorithms for neonatal sepsis prediction.
Table 1. Comparative characteristics of machine learning and deep learning algorithms for neonatal sepsis prediction.
Algorithm Advantages Limitations Best clinical application
Logistic Regression High interpretability Linear relationships Clinical decision support
Random Forest Robust, handles nonlinear data Moderate interpretability Structured EHR datasets
SVM High accuracy in small datasets Computationally demanding Small datasets
XGBoost Highest performance Hyperparameter tuning Multimodal prediction
CNN Physiological signal analysis Large datasets needed Continuous monitoring
LSTM Temporal prediction Complex training Real-time monitoring
Abbreviations: AI, artificial intelligence; CNN, convolutional neural network; DL, deep learning; LSTM, long short-term memory; ML, machine learning; SHAP, SHapley Additive exPlanations; SVM, support vector machine; XGBoost, Extreme Gradient Boosting.
Table 2. Representative studies evaluating artificial intelligence models for neonatal sepsis prediction.
Table 2. Representative studies evaluating artificial intelligence models for neonatal sepsis prediction.
Study Data Source AI/ML Method Prediction Horizon Performance
Masino et al., 2019 [14] EHR, laboratory data, and vital signs Multiple ML algorithms ≥4 h before diagnosis AUC >0.80
Honoré et al., 2023 [18] Continuous physiological monitoring Machine learning Up to 24 h before clinical suspicion AUC ≈0.82
Kainth et al., 2026 [26] Perinatal and neonatal clinical variables Random forest with Boruta feature selection EOS prediction within the first 72 h of life Sensitivity 90.3%, specificity 40.6%; external validation: sensitivity 92.3%, NPV 95.7%
An et al., 2024 [29] Transcriptomic data Machine learning Before clinical presentation Four-gene predictive signature
Mithal et al., 2025 [30] Cord blood proteomics Logistic regression, random forest At birth for EOS risk prediction Clinically meaningful predictive performance
Mani et al., 2014 [31] EHR and clinical data Logistic regression, SVM, random forest, decision tree Before clinical diagnosis ML models outperformed physician-guided antibiotic initiation
Kallonen et al., 2024 [34] ECG and respiratory impedance signals Deep learning/CNN ~44 h before clinical suspicion AUC ≈0.81 in external validation
Gomez et al., 2019 [35] Heart rate variability AdaBoost, SVM Not reported AUC ≈0.94
Song et al., 2020 [36] Vital sign monitoring Machine learning Up to 48 h before diagnosis Early prediction of LOS
Garstman et al., 2023 [37] Continuous physiological monitoring data Random forest and other ML algorithms Early detection of LOS AUROC 0.973 for random forest
Yang et al., 2024 [43] Physiological monitoring and demographic data XGBoost, LSTM, and other ML models 6 h before clinical suspicion AUC 0.875; XGBoost was the best-performing model
Abbreviations: AI, artificial intelligence; AUC, area under the receiver operating characteristic curve; CNN, convolutional neural network; ECG, electrocardiography; EHR, electronic health record; EOS, early-onset sepsis; LSTM, long short-term memory; ML, machine learning; SVM, support vector machine; XGBoost, Extreme Gradient Boosting.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings