Preprint
Article

This version is not peer-reviewed.

Longitudinal Clinical and Minimally Invasive Readouts Enable Early Assessment of Disease Trajectory in Preclinical Infection Models

  † These authors contributed equally to this work.

Submitted:

15 August 2026

Posted:

18 August 2026

You are already at the latest version

Abstract

Murine infection models are essential for investigating host-pathogen interactions, disease progression, vaccine-induced immune responses and protection. Early identification of unfavorable disease trajectories is important for animal monitoring and collection of high-quality biological samples. However, individual clinical readouts may provide limited information. Longitudinal clinical and minimally invasive readouts (cumulative weight loss, day-to-day weight change, clinical score, and blood bacterial load) were integrated using Machine Learning approaches to assess short-term disease trajectory and fatal outcome within the subsequent three days. Two Salmonella Typhimurium infection studies were analyzed, with Study 1 used for model development and Study 2 for external validation. Four classifiers (Random Forest, Support Vector Machine, XGBoost, and penalized logistic regression) and three ensemble strategies (majority-voting, weighted-voting, and stacked meta-learner) were evaluated. Classification performance was generally good, with the stacked ensemble and SVM showing the best balance across metrics. Application to Study 2 showed that the combined readouts remained informative in an independent experiment. These findings support the feasibility of integrating longitudinal clinical and minimally invasive readouts for early assessment of disease trajectory. This framework may help identify informative sampling windows and support planning of experimental procedures, including sample collection and humane endpoints, with potential applicability to other preclinical infection models.

Keywords: 
;  ;  ;  ;  ;  ;  

1. Introduction

Preclinical murine infection models remain essential for investigating host-pathogen interactions, elucidating mechanisms of disease progression, and testing potential prophylactic or therapeutic interventions. They provide a controlled framework to capture the complexity of systemic infection and to identify quantitative markers of severity that can guide both mechanistic studies and welfare-oriented decision making. Invasive infections caused by non-typhoidal Salmonella (iNTS) represent a major global health concern, particularly affecting young children, older adults, and immunocompromised individuals, and have been extensively studied using murine infection models, which have been instrumental in elucidating disease mechanisms and host responses1,2,3,4. In preclinical experimental settings, ensuring the protection of animal well-being in accordance with FELASA (Federation of European Laboratory Animal Science Associations) guidelines is a primary ethical requirement and the basis for evidence-based severity assessment. Early identification of disease worsening is therefore essential both for monitoring the transition toward critical disease stages and for supporting experimental planning and the appropriate definition of humane endpoints.
Rigorous severity classification supports reproducibility, standardization, and the generation of biologically meaningful data from animal studies5–7. In this framework, approaches capable of providing an early indication of worsening disease trajectory may contribute to more efficient experimental designs by supporting the identification of informative sampling windows and endpoint planning. Such strategies align with the internationally recognized principles of the 3Rs (replacement, reduction, and refinement) which guide the ethical use of animals in research8,9. In particular, improved longitudinal monitoring may support more informed decisions, facilitate the selection of informative sampling time points, maximize the biologically meaningful information obtained from each experimental cohort, and contribute to timely endpoint planning.
Routine monitoring in experimental infection models commonly relies on complementary clinical and microbiological readouts, including body-weight changes, clinical severity scores and circulating bacterial burden7,10–15. However, although individually informative, these parameters may not always provide a sufficiently clear indication of short-term disease trajectory when considered in isolation7,10–15. Integrating these longitudinal measurements may therefore provide a more informative representation of disease progression than reliance on individual readouts alone.
Machine Learning (ML) approaches are increasingly being adopted in biomedical research as tools to capture complex, multivariate patterns associated with clinical outcomes. In human cohorts, ML models have been successfully applied to predict mortality risk in large populations16,17, to identify immune correlates of protection following vaccination18,19, and to forecast adverse outcomes associated with infectious diseases20. Emerging evidence from veterinary clinical settings shows that ML models can aid decision-making across multiple indications, including risk prediction from routine clinical data, early detection of infectious diseases, and diagnostic case classification21–25, suggesting that ML can also support decision-making in clinical veterinary practice.
By contrast, the application of ML models to preclinical in vivo research, and particularly to experimental infection models in laboratory animals, remains at an early stage. Yet this context represents a crucial intermediate step between mechanistic studies and translational applications. Preclinical models, such as murine infection systems, generate longitudinal datasets integrating multiple clinical and biological measurements that may be well-suited for ML analysis but are often underutilized beyond descriptive or statistical evaluations. Bridging this gap may help maximize the scientific information obtained from animal studies, while supporting more informed experimental planning, longitudinal disease monitoring and improved ethical management of laboratory animals.
In the present study, longitudinal clinical and minimally invasive readouts were integrated using complementary ML approaches in a murine model of intragastric infection with S. Typhimurium strain D23580, a well-established preclinical model of iNTS disease. Cumulative weight loss, day-to-day weight change, clinical score and bacteraemia were evaluated in combination to assess whether their integration could provide an early indication of short-term disease trajectory and fatal outcome within the subsequent three days. The analytical strategy comprised four steps: (i) identifying informative patterns through exploratory analysis; (ii) training and testing multiple supervised algorithms in Study 1; (iii) integrating base models into three more ensemble approaches; and (iv) external validation of the finalized models in Study 2, an independently conducted infection experiment.

2. Materials and Methods

2.1. Animal Infection Model

C57BL/6 mice were purchased from Charles River Italy (Lecco, Italy) and maintained under pathogen-free conditions in the animal facility of the Laboratory of Molecular Microbiology and Biotechnology (LA.M.M.B.), Department of Medical Biotechnologies at the University of Siena. The animals were housed in groups of up to six animals in cages (Tecniplast, Buguggiate, Italy) at a temperature of 20-24 °C, humidity rate of 55 (+10%), with food and water ad libitum, under conditions complying with the “3R” principles (replacement, refinement and reduction) according to the national guidelines (Italian D. Lgs n. 26 of 2014), in conformity with the European 2010/63/EU Directive on the protection of animals used for scientific purposes. Animal experiments were approved by the Ministry of Health with authorization no. 371/2019-PR.
Animals were inoculated by the intragastric route with an infection dose of 10^8 CFU of the invasive clinical isolate from Malawi S. Typhimurium D2358026 using a gavage needle (22 GA), as previously described27, in a volume of 100 μl/mouse in Luria-Bertani broth solution (Triptone –10 g/L, Oxoid, Basingstoke, UK; Sodium chloride – 10 g/L, Carlo Erba, Milan, Italy; Yeast extract – 5 g/L, Oxoid) without any previous antibiotic pretreatment. Upon infection, animals were monitored daily for body weight, clinical score, and survival. Blood bacterial burden was measured at defined timepoints by plating appropriate dilutions onto Salmonella-Shigella agar plates (Oxoid).
Both experiments (Study 1 and Study 2) were conducted as longitudinal survival studies, with predefined observation periods of 34 days for Study 1 and 27 days for Study 2. Animals were monitored throughout disease progression until spontaneous death, euthanasia upon reaching predefined humane endpoint criteria, or completion of the scheduled observation period. Completion of the observation period was not considered a fatal event. Animals reaching a clinical score of 5 were euthanized for animal-welfare reasons and not as part of a scheduled experimental sampling procedure. Euthanasia was performed by intraperitoneal administration of ketamine (Lobotor, 100 mg/kg; ACME S.r.l., Italy) and xylazine (Rompun, 8 mg/kg; Bayer S.p.A., Germany).

2.2. Experimental Model and Data Collection

Two separate and independently conducted survival experiments, hereafter referred to as Study 1 and Study 2, were analysed (Figure 1). Study 1 was used for model construction, including training and internal testing, whereas Study 2 served as a separate experimental cohort reserved exclusively for external validation of the finalized models.
Study 1 consisted of three cages housing six animals each (n = 18), while Study 2 consisted of two cages housing eight animals each (n = 16). In both experiments, animals were daily monitored, and the following key physiological and clinical parameters were collected for each subject: (i) cumulative weight loss, defined as the cumulative percentage weight loss from baseline; (ii) daily weight variation, defined as the percentage change in weight relative to the previous day (day-to-day dynamics); (iii) clinical score, defined as the measure used to assess the severity of disease in the infected animals, determined using a scale ranging from 0 (normal) to 5 (severe clinical condition requiring euthanasia according to the predefined humane endpoint criteria), based on the observation of different markers, such as fur aspect, behaviour, activity, posture and eyelids. Specifically, the severity of the disease was classified with: 0 absence of symptoms; 1, animal with piloerection and reduced motility; 2, animal with curved back and semi-closed eyes; 3, animal that takes more than 5 seconds to turn from supine position to prone position when placed on the back; 4, animal that cannot turn in a prone position from the supine position; 5, severe clinical condition requiring euthanasia according to the predefined humane endpoint criteria; (iv) bacteraemia: circulating bacterial load expressed as Colony Forming Units (CFU)/mL of blood. Blood sampling was performed on alternating days across the cages, such that each animal was bled at longer intervals. This alternating design ensured equivalent temporal coverage across the cohort while adhering to guidelines aimed at minimizing blood volume loss and procedural stress.
For each experiment, observations were excluded if any of the four predictor variables (cumulative weight loss, daily weight change, clinical score, or blood CFU) were missing or if sufficient subsequent follow-up was not available to determine the three-day outcome, as detailed below.
For analytical purposes, a fatal outcome was defined as either spontaneous death or euthanasia required upon reaching the predefined humane endpoint. At each eligible observation time point, the binary outcome “fatal outcome within 3 days” was assigned a value of 1 when either event occurred within the subsequent three days and 0 when the animal remained alive for more than three days after that observation. To ensure complete outcome ascertainment, only observations for which the subsequent three-day period was fully covered by the experimental follow-up were considered eligible. Accordingly, observations from animals remaining alive at the end of the study were excluded when the corresponding three-day outcome window extended beyond the predefined observation period (day 34 for Study 1 and day 27 for Study 2).

2.3. Pre-Processing and Hyperparameter Tuning

The Study 1 dataset was split into training (70%) and test (30%) sets using stratification by outcome. All predictor variables were standardized using z-score transformation, with mean and standard deviation estimated exclusively from the training data within each cross-validation fold.
Four supervised learning algorithms were trained within an R 4.5.0 environment following tidymodels v1.3.0 workflow: (i) Random Forests (RF; ranger v0.17.0), a tree-based model that aggregates the results of multiple decision trees to reduce overfitting; (ii) Support Vector Machines with radial basis function kernel (SVM; kernlab v0.9-33), a classifier that maps data into a higher-dimensional space to find the best boundary between classes; (iii) gradient boosting decision trees (XGBoost; xgboost v1.7.11.1), a tree-boosting algorithm that builds sequential models to correct the errors of the previous ones; (iv) penalized logistic regression with elastic net regularization (glmnet v4.1-9), a linear model used to predict binary outcomes, which includes regularization to avoid overfitting
For each algorithm, hyperparameters controlling model complexity (i.e., number of variables sampled at each split and minimum node size for Random Forest; cost and RBF sigma for SVM; learning rate, tree depth, minimum node size number of variables sampled at each split and loss reduction for XGBoost; penalty and mixing ratio for logistic regression) were tuned through a systematic grid search.
Model selection relied on stratified 5-fold cross-validation within the training set: data were partitioned into five equal folds, with four used for training and one for validation in rotation, ensuring balanced outcome proportions across folds. This procedure was repeated for all hyperparameter combinations (tune grid = 20), and average performance metrics were computed across folds. Optimization targeted multiple standard metrics including overall accuracy, sensitivity – defined as the ability to correctly classify true positives (i.e., mice with a fatal outcome) – and specificity – defined as the ability to correctly identify true negatives (i.e., surviving mice) –, with specificity prioritized to minimize false-positive classifications of animals surviving beyond the subsequent three days.
Importantly, the 30% test set was kept fully independent throughout model development and was used only for final performance evaluation, providing an unbiased estimate of generalization.

2.4. Ensemble Models Building

In addition to base models, three ensemble models, which combine the prediction of the four base models to improve robustness, were built: (i) a majority vote classifier, in which the final prediction in the class is predicted by at least three out of four models; (ii) a weighted vote, in which predictions are weighted by the specificity of each base model, giving more influence to those with higher specificity values; and (iii) a stacked ensemble, in which predictions from the base models are used as input for a second model (a penalized logistic regression), which learns how to optimally combine them. This meta-model was tuned to balance the contribution of each base model while preventing overfitting through regularization.

3. Results

In the present study, four longitudinal clinical and minimally invasive readouts – cumulative weight loss, daily weight variation, clinical score, and bacteraemia – were integrated using multiple ML approaches to assess short-term disease trajectory following Salmonella Typhimurium infection. These parameters, selected for their established relevance in murine infection models, were evaluated across two independently conducted experimental studies.

3.1. Study Cohort

Two independently conducted infection studies, Study 1 and Study 2, were analysed (Figure 1 and Table 1). The first dataset (Study 1) was used for model construction, encompassing both training and internal testing phases. Out of 18 infected mice, 17 met the predefined inclusion criteria, yielding 82 longitudinal observations across the four clinical or minimally invasive readouts variables considered (cumulative weight loss, daily weight variation, clinical score, and bacteraemia).
The second dataset (Study 2) consisted of a separate experimental cohort and was reserved exclusively for external validation. It comprised 16 infected mice, of which 8 satisfied inclusion criteria, resulting in 23 eligible observations for model evaluation.This design ensured a clear separation between the development and validation cohorts and allowed performance metrics to be assessed on an independent dataset.

3.2. Exploratory Data Analysis

To examine how cumulative weight loss, daily weight variation, clinical score, and bacteraemia related to short-term disease trajectory and fatal outcome, preliminary exploratory analyses were performed.
The time-course plot in Figure 2A showed that animals experiencing a fatal outcome within the subsequent three days exhibited more pronounced cumulative weight loss, steeper day-to-day declines, a progressive worsening of the clinical score, and, when available, rising bacteraemia compared with animals surviving beyond this interval. However, none of these variables alone displayed a clear threshold capable of consistently distinguishing animals approaching a short-term fatal outcome.
For instance, regarding the clinical score (Figure 2A, top-right panel), some mice that progressed from a score of 0 to 1 subsequently experienced a fatal outcome within three days, whereas others showing a similar early trajectory recovered and returned toward baseline. Similarly, no clear cutoff was observed for either cumulative weight loss (Figure 2A, bottom-left panel) or daily weight variation (Figure 2A, bottom-right panel): some animals recovered despite losing more than 10% of their body weight, while others reached a fatal outcome despite smaller losses. A similar pattern was observed in the heatmap (Figure 2B), which summarizes the values of the four variables at the last available observation preceding the fatal outcome. Indeed, while a few mice displayed clear increases in all four variables shortly before death, for most animals experiencing a fatal outcome within the subsequent three days, such a sharp distinction was not consistently observable at the level of individual variables. In contrast, the Principal Component Analysis (PCA; Figure 2C), capable of capturing the linear joint effect of all four variables, revealed a clearer separation between observations associated with survival beyond three days and those preceding a fatal outcome.
In practice, animals nearing death exhibited a joint signature of increased cumulative weight loss, more pronounced day-to-day decline, higher clinical scores, and, where measured, elevated blood bacterial burden, whereas animals surviving beyond three days remained within milder, less coordinated fluctuations across the same variables. Interestingly, while individual variables alone were not always sufficient to distinguish observations preceding a short-term fatal outcome, their combined evaluation revealed a multivariate pattern associated with worsening disease trajectory. This observation prompted further investigation of the combined information provided by the four readouts through the application of various ML algorithms.

3.3. Training on the Study 1 Dataset

To investigate the combined discriminatory information provided by the four variables considered (bacteraemia, clinical score, cumulative weight loss and daily weight variation), four ML classifiers, including Random Forest, SVM, XGBoost and penalized logistic regression, were implemented.
These four ML models were trained and tested on the Study 1 dataset (Figure 1 and Table 1). During training, all models underwent hyperparameter tuning through grid search combined with 5-fold cross-validation on the training data. This procedure systematically tested multiple parameter combinations to optimize predictive performance while minimizing overfitting, to finally select the best ones based on specificity, defined as the ability to correctly identify observations associated with survival beyond the subsequent three days. The cross-validated metrics obtained during this tuning phase were visualized (Figure 3): each point corresponded to a parameter configuration, with mean performance and associated variability shown across folds. All models achieved good performances in terms of accuracy (0.83 - 1) and sensitivity (0.9 - 1). However, specificity showed greater variability, particularly for SVM and XGBoost, which exhibited values ranging from 0 to 0.72 across different folds, suggesting reduced reliability in identifying true survivors under certain parameter settings.
The final set of hyperparameters selected as those yielding highest specificity for each model (Supplementary Table S1) were then used to refit the models on the full training Study 1 dataset before evaluation on the independent test partition.

3.4. Independent Testing of Ensemble and ML Models on the Study 1 Testing Dataset

Based on the promising performance observed during training, all four individual ML classifiers were further evaluated on the independent test set of Study 1 using their respective optimized hyperparameters.
In addition to the base classifiers, three ensemble strategies, including a majority voting, a weighted voting (with weights defined as model-specificity values) and a stacked meta-learner ensemble model, were implemented to assess whether combining predictions could improve overall classification performance.
The meta-learner ensemble model selected a penalty of 0.144 and retained contributions from one SVM, one XGBoost, and 20 Random Forest models (Supplementary Figure S1). Models with higher stacking coefficients contributed more strongly to the final decision, while redundant or less informative members were effectively down weighted.
All seven models – four base classifiers and three ensemble approaches – achieved high performance on the test set in terms of accuracy (ranging from 0.92 to 0.96) and sensitivity (1.00 for all models). However, specificity remained more variable across models, ranging from 0.50 to 0.75 (Table 2; Supplementary Figure S2). Notably, the stacked meta-learner and the SVM classifier reached the highest specificity (0.75), indicating better identification of observations associated with survival beyond three days.

3.5. External Independent Validation on the Study 2 Dataset

When tested on the external and independent Study 2 dataset (as described in Figure 1), all models retained generally high sensitivity (0.89–1), indicating that most observations preceding a fatal outcome were correctly classified as positive.
The stacked ensemble achieved the best overall performance, with a sensitivity of 1.00, specificity of 0.75, and an overall accuracy of 0.96 (Table 3; Supplementary Figure S2). This indicated its strong ability to correctly identify observations preceding a fatal outcome within the subsequent three days (purple dots in Figure 2A), while reducing false negative classifications. Observations from mice surviving beyond the three-day window (green dots in Figure 2A) were also generally identified as low risk, although the frequency of false-positive classifications varied across models.
These results indicate that the stacked ensemble model maintained good classification performance across both internal testing and external validation datasets. Importantly, the model relies solely on longitudinal clinical and minimally invasive readouts, supporting its potential application in preclinical infection studies to assist disease monitoring, optimize sampling timepoints and better define humane endpoints while minimizing unnecessary interventions.

4. Discussion

In this study, four longitudinal clinical and minimally invasive readouts – cumulative weight loss, daily weight variation, clinical score, and bacteraemia – were integrated using multiple ML approaches to assess short-term disease trajectory in a murine model of S. Typhimurium infection. These parameters were selected for their well-established relevance in murine bacterial infection models and for the complementary information they provide on disease progression. Their combined evaluation supported the identification of multivariate patterns associated with a fatal outcome within the subsequent three days.
Among the considered variables, cumulative weight loss remains one of the most widely adopted markers of disease severity, reflecting both metabolic alterations and the progression of systemic illness10. Its sensitivity makes it a valuable marker of health status, but reliance on rigid thresholds poses clear limitations. As highlighted by Brochut et al.7, stringent cut-offs can lead to premature euthanasia of animals otherwise capable of recovery, while in other cases deterioration may occur before the threshold is reached. These observations underline both the importance and the limitations of weight loss as a single criterion, emphasizing the need for integrative approaches that combine it with additional physiological and clinical parameters.
Unlike cumulative weight loss, which reflects the overall progression of illness and decreased wellbeing, day-to-day weight variation provides a more dynamic perspective, capturing short-term trajectories that may precede clinical collapse. In murine sepsis models, Osuchowski et al.11 demonstrated that rapid and pronounced weight declines closely tracked the terminal phase of disease in non-survivors, highlighting the capacity of daily changes to capture short-term deterioration that might remain less evident when relying exclusively on cumulative measures.
Beyond weight-related parameters, clinical scoring systems offer a complementary, multidimensional view of animal wellbeing. By integrating non-invasive observations of fur condition, posture, mobility, behaviour, and responsiveness, they provide a structured evaluation of disease severity. Several studies have demonstrated their predictive value, showing that composite scores anticipate mortality with high accuracy in murine sepsis models and can serve as more humane and reproducible endpoints than death itself7,12,13. In neonatal mice infection models, structured scoring has even outperformed conventional biomarkers such as circulating cytokines in predicting both survival and time-to-death14. Simplified 0–5 clinical scoring systems have been shown to improve reproducibility in experimental settings15. In this study, a 0–5 scale from normal status (0) to severe clinical condition requiring euthanasia according to the predefined humane endpoint criteria (5) was applied, consistent with FELASA guidelines for standardized severity assessment and humane end-point definition and grimace scale28,29.
Finally, blood bacterial load represents direct microbiological evidence of systemic infection and is strongly associated with outcome. Elevated bacterial counts in peripheral blood are consistently linked to increased mortality in murine models of invasive bacterial disease. In preclinical infection models, quantitative blood cultures have been correlated with fatal outcomes and heightened inflammatory responses in non-survivors30–32. These data confirm the pathobiological and prognostic relevance of bacterial dissemination as a marker of disease severity, providing a microbiological readout that complements host-based indicators such as weight metrics and clinical scores.
While these findings highlight the individual value of each parameter, the exploratory analyses also showed that no single readout consistently distinguished observations preceding a short-term fatal outcome. Rather, worsening disease was characterized by the joint behaviour of multiple clinical and microbiological measurements. ML approaches provide a flexible means of integrating such complementary variables in both linear and non-linear forms, allowing multivariate relationships to be evaluated without reliance on a single predefined threshold33. Building on this rationale, four supervised classifiers (RF, SVM, XGBoost, and penalized logistic regression) as well as three ensemble strategies (majority voting, weighted voting, and a stacked model) were implemented. Models were developed and internally evaluated on a discovery study (Study 1) and subsequently validated on an independent study (Study 2). The persistence of useful classification information in Study 2 provides further support for the feasibility of combining these routinely collected readouts across separate experimental settings.
Translating these findings into practice, early identification of animals approaching a fatal outcome within the subsequent three days may support more informed experimental decision-making. This may facilitate the selection of biologically informative sampling windows and the scheduling of experimental procedures before animals reach advanced disease stages. Integration of longitudinal clinical and minimally invasive readouts may support timely endpoint planning, contributing to the refinement of experimental procedures in accordance with the 3Rs principles. In this context, ML-based integration may further assist the earlier recognition of critical disease stages, helping align ethical considerations with scientific rigor.
Nonetheless, some limitations should be acknowledged. The model was developed and validated across two studies (Study 1 for model construction and Study 2 for external validation), representing a relatively limited sample size. Furthermore, analyses were restricted to a single host background (C57BL/6) and one bacterial strain (S. Typhimurium D23580). Broader validation across different genetic backgrounds, infectious agents and doses, and experimental settings will be required to consolidate transferability. Looking ahead, incorporation of other immunological features, such as plasma cytokine profiles, may further enhance predictive performance and provide additional insight into host–pathogen dynamics.
In conclusion, the present work supports the feasibility of integrating routinely collected longitudinal clinical and minimally invasive readouts to obtain an early indication of short-term disease trajectory in a murine model of intragastric infection with S. Typhimurium. Cumulative weight loss, day-to-day weight variation, clinical score and bacteraemia provided complementary information that became more informative when evaluated jointly rather than in isolation. Among all tested models, the stacked ensemble meta-learner achieved the highest performance during both the model construction phase and validation on an external cohort. The analytical framework proposed should therefore be considered a proof-of-concept for integrating routinely available measurements to support longitudinal disease monitoring and the identification of informative experimental windows. By supporting early recognition of critical disease progression, this approach may support more refined planning of experimental procedures and sampling, while also enabling timely implementation of humane endpoints. These elements contribute directly to improving adherence to the internationally recognized 3Rs principles which promote ethically responsible animal research8,9. In particular, this work could contribute to the reduction principle by increasing the informational value obtained from each individual animal and potentially improving the efficiency of experimental designs. As highlighted by Richter34, innovation in study design is a key strategy to reduce animal use without compromising scientific rigor. In this perspective, integrative computational approaches may represent a valuable contribution to more refined and ethically conscious experimental planning.

Supplementary Materials

The following supporting information can be downloaded at: Preprints.org, Table S1: Tuned hyperparameters for each classifier. Final hyperparameter values selected after grid search and 5-fold cross-validation for each model. The table reports the computational engine, the hyperparameters tuned, the final selected values, and the total number of parameters optimized; Figure S1: Stacked model interpretation plots. Stacking coefficients assigned by the penalized logistic regression meta-model for the ensemble members. Higher coefficients indicate stronger contribution to the final model. The model selected a penalty of 0.144, retaining one SVM (beige), one XGBoost (red), and 20 Random Forest models (blue). This configuration highlights the predominance of Random Forest members while preserving complementary signals from non-linear classifiers, ensuring both parsimony and balanced performance; Figure S2: Comparative performance of base and ensemble models on the test and validation datasets. Barplots report accuracy, sensitivity, and specificity for the four base classifiers, namely Random Forest (RF), Support Vector Machine with RBF kernel (SVM), XGBoost (XGB), and penalized logistic regression (Logit), and the three ensemble strategies: majority vote (Major.), weighted vote (Weight.) and stacked (Stacked). Performances are shown separately for the internal test set (light purple) and the external validation set (dark purple). Accuracy and sensitivity were consistently high across all models in both datasets, while specificity displayed greater variability. Notably, the stacked ensemble and SVM maintained the best balance between sensitivity and specificity, with the stacked model showing maintaining good overall performance across datasets..

Author Contributions

Conceptualization: GM, VM, EP; Data curation: VM, FF, GM; Formal analysis: GM, VM; Methodology: GM, VM, FF, EP; Project Administration: EP, DM, GM, VM, FF; Resources: DM; Software: GM, VM; Supervision: EP, DM, FF, GM; Writing – original draft: VM, GM, EP, FF, DM. All authors read and approved the final manuscript.

Funding

This research was supported by the Department of Medical Biotechnologies of the University of Siena (D.M.).

Institutional Review Board Statement

Animal experiments were approved by the Ministry of Health with authorization no. 371/2019-PR.

Data Availability Statement

Data and code will be made available upon reasonable request to the corresponding author.

Conflicts of interest

The Authors declare no competing interests.

References

  1. MacLennan, C. A.; Levine, M. M. Invasive nontyphoidal Salmonella disease in Africa: current status. Expert Rev. Anti Infect. Ther. 2013, 11, 443–446. [Google Scholar] [CrossRef] [PubMed]
  2. Uche, I. V.; MacLennan, C. A.; Saul, A. A Systematic Review of the Incidence, Risk Factors and Case Fatality Rates of Invasive Nontyphoidal Salmonella (iNTS) Disease in Africa (1966 to 2014). PLoS Negl. Trop. Dis. 2017, 11, e0005118. [Google Scholar] [CrossRef] [PubMed]
  3. Santos, R. L.; et al. Animal models of Salmonella infections: enteritis versus typhoid fever. Microbes Infect. 2001, 3, 1335–1344. [Google Scholar] [CrossRef] [PubMed]
  4. Tsolis, R. M.; Xavier, M. N.; Santos, R. L.; Bäumler, A. J. How To Become a Top Model: Impact of Animal Experimentation on Human Salmonella Disease Research ▿. Infect. Immun. 2011, 79, 1806–1814. [Google Scholar] [CrossRef] [PubMed]
  5. Keubler, L. M.; et al. Where are we heading? Challenges in evidence-based severity assessment. Lab Anim. 2020, 54, 50–62. [Google Scholar] [CrossRef] [PubMed]
  6. Peppermüller, P. P.; et al. Grimace scale assessment during Citrobacter rodentium inflammation and colitis development in laboratory mice. Front Vet. Sci. 2023, 10, 1173446. [Google Scholar] [CrossRef] [PubMed]
  7. Brochut, M.; et al. Using weight loss to predict outcome and define a humane endpoint in preclinical sepsis studies. Sci. Rep. 2024, 14, 21150. [Google Scholar] [CrossRef] [PubMed]
  8. Prescott, M. J.; Lidster, K. Improving quality of science through better animal welfare: the NC3Rs strategy. Lab Anim. (NY) 2017, 46, 152–156. [Google Scholar] [CrossRef] [PubMed]
  9. The Principles of Humane Experimental Technique. Med. J. Aust. 1960, 1, 500–500. [CrossRef]
  10. Talbot, S. R.; et al. Defining body-weight reduction as a humane endpoint: a critical appraisal. Lab Anim. 2020, 54, 99–110. [Google Scholar] [CrossRef] [PubMed]
  11. Osuchowski, M. F.; Welch, K.; Yang, H.; Siddiqui, J.; Remick, D. G. Chronic sepsis mortality characterized by an individualized inflammatory response. J. Immunol. 2007, 179, 623–630. [Google Scholar] [CrossRef] [PubMed]
  12. Shrum, B.; et al. A robust scoring system to evaluate sepsis severity in an animal model. BMC Res. Notes 2014, 7, 233. [Google Scholar] [CrossRef] [PubMed]
  13. Mai, S. H. C.; et al. Body temperature and mouse scoring systems as surrogate markers of death in cecal ligation and puncture sepsis. Intensive Care Med. Exp. 2018, 6, 20. [Google Scholar] [CrossRef] [PubMed]
  14. Fehlhaber, B.; et al. A sensitive scoring system for the longitudinal clinical evaluation and prediction of lethal disease outcomes in newborn mice. Sci. Rep. 2019, 9, 5919. [Google Scholar] [CrossRef] [PubMed]
  15. Sulzbacher, M. M.; et al. Adapted Murine Sepsis Score: Improving the Research in Experimental Sepsis Mouse Model. BioMed Res. Int. 2022, 2022, 5700853. [Google Scholar] [CrossRef] [PubMed]
  16. Qiu, W.; et al. Interpretable machine learning prediction of all-cause mortality. Commun. Med. (Lond) 2022, 2, 125. [Google Scholar] [CrossRef] [PubMed]
  17. Yu, Q.; et al. Predicting all-cause mortality and premature death using interpretable machine learning among a middle-aged and elderly Chinese population. Heliyon 2024, 10, e36878. [Google Scholar] [CrossRef] [PubMed]
  18. Montesi, G.; et al. Machine Learning Approaches to Dissect Hybrid and Vaccine-Induced Immunity dataset. Zenodo 2025. [Google Scholar] [CrossRef]
  19. Montesi, G.; et al. Predicting humoral responses to primary and booster SARS-CoV-2 mRNA vaccination in people living with HIV: a machine learning approach. J. Transl. Med. 2024, 22, 432. [Google Scholar] [CrossRef] [PubMed]
  20. Baker, M. R.; Buyrukoğlu, S.; Buyrukoğlu, G.; Moreira, J.; Topalcengiz, Z. Efficacy of machine learning models for the prediction of death occurrence and counts associated with foodborne illnesses and hospitalizations in the United States. Microb. Risk Anal. 2025, 30, 100351. [Google Scholar] [CrossRef]
  21. Franzo, G.; et al. Comparison and validation of different models and variable selection methods for predicting survival after canine parvovirus infection. Vet. Rec. 2020, 187, e76. [Google Scholar] [CrossRef] [PubMed]
  22. Reagan, K. L.; et al. Use of machine-learning algorithms to aid in the early detection of leptospirosis in dogs. J. Vet. Diagn. Invest 2022, 34, 612–621. [Google Scholar] [CrossRef] [PubMed]
  23. Ferreira, T. S.; et al. Diagnostic Classification of Cases of Canine Leishmaniasis Using Machine Learning. Sensors 2022, 22, 3128. [Google Scholar] [CrossRef] [PubMed]
  24. Kim, Y.; et al. Machine learning-based risk prediction model for canine myxomatous mitral valve disease using electronic health record data. Front Vet. Sci. 2023, 10, 1189157. [Google Scholar] [CrossRef] [PubMed]
  25. Sanaei, N.; Zamani-Ahmadmahmudi, M.; Nassiri, S. M. Development of machine learning models to predict clinical outcome and recovery time in dogs with parvovirus enteritis. Front Vet. Sci. 2025, 12, 1555714. [Google Scholar] [CrossRef] [PubMed]
  26. MacLennan, C. A.; et al. The neglected role of antibody in protection against bacteremia caused by nontyphoidal strains of Salmonella in African children. J. Clin. Invest 2008, 118, 1553–1562. [Google Scholar] [CrossRef] [PubMed]
  27. Walker, G. T.; Gerner, R. R.; Nuccio, S.-P.; Raffatellu, M. Murine Models of Salmonella Infection. Curr. Protoc. 2023, 3, e824. [Google Scholar] [CrossRef] [PubMed]
  28. Langford, D. J.; et al. Coding of facial expressions of pain in the laboratory mouse. Nat. Methods 2010, 7, 447–449. [Google Scholar] [CrossRef] [PubMed]
  29. Smith, D.; et al. Classification and reporting of severity experienced by animals used in scientific procedures: FELASA/ECLAM/ESLAV Working Group report. Lab Anim. 2018, 52, 5–57. [Google Scholar] [CrossRef] [PubMed]
  30. van den Berg, S.; et al. Distinctive cytokines as biomarkers predicting fatal outcome of severe Staphylococcus aureus bacteremia in mice. PLoS ONE 2013, 8, e59107. [Google Scholar] [CrossRef] [PubMed]
  31. Wu, D.; Zhou, S.; Hu, S.; Liu, B. Inflammatory responses and histopathological changes in a mouse model of Staphylococcus aureus-induced bloodstream infections. J. Infect. Dev. Ctries. 2017, 11, 294–305. [Google Scholar] [CrossRef] [PubMed]
  32. Wills, B. M.; et al. Identification of Virulence Factors Involved in a Murine Model of Severe Achromobacter xylosoxidans Infection. Infect. Immun. 2023, 91, e0003723. [Google Scholar] [CrossRef] [PubMed]
  33. Wiens, J.; Shenoy, E. Machine Learning for Healthcare: On the Verge of a Major Shift in Healthcare Epidemiology. Clin. Infect. Dis. An. Off. Publ. Infect. Dis. Soc. Am. 2018, 66. [Google Scholar]
  34. Richter, S. H. Challenging current scientific practice: how a shift in research methodology could reduce animal use. Lab Anim. (NY) 2024, 53, 9–12. [Google Scholar] [CrossRef] [PubMed]
Figure 1. Experimental design and data analysis workflow. C57BL/6 mice, divided into two independent cohorts (Study 1 and Study 2) were intragastrically infected with S. Typhimurium. The predefined observation periods were 34 days for Study 1 and 27 days for Study 2. Completion of these periods was not considered a fatal event. Daily monitoring included weight and clinical score assessment in both cohorts. In Study 1, blood samples for CFU determination were collected on alternating days across cages to minimize stress and blood volume loss, while ensuring equivalent temporal coverage. Data from Study 1 were used for model development (70% employed for ML model training, 30% used for internal ML model testing), whereas Study 2 served as an independent validation cohort. After preprocessing, four classifiers (Random Forest, RF; Support Vector Machine, SVM; XGBoost; and penalized logistic regression) were trained, and ensemble strategies (majority voting, weighted voting, stacked ensemble) were implemented. Model performance was evaluated in terms of accuracy, sensitivity, and specificity. Only observations with complete predictor data and sufficient subsequent follow-up to determine the three-day outcome were included in the ML analysis.
Figure 1. Experimental design and data analysis workflow. C57BL/6 mice, divided into two independent cohorts (Study 1 and Study 2) were intragastrically infected with S. Typhimurium. The predefined observation periods were 34 days for Study 1 and 27 days for Study 2. Completion of these periods was not considered a fatal event. Daily monitoring included weight and clinical score assessment in both cohorts. In Study 1, blood samples for CFU determination were collected on alternating days across cages to minimize stress and blood volume loss, while ensuring equivalent temporal coverage. Data from Study 1 were used for model development (70% employed for ML model training, 30% used for internal ML model testing), whereas Study 2 served as an independent validation cohort. After preprocessing, four classifiers (Random Forest, RF; Support Vector Machine, SVM; XGBoost; and penalized logistic regression) were trained, and ensemble strategies (majority voting, weighted voting, stacked ensemble) were implemented. Model performance was evaluated in terms of accuracy, sensitivity, and specificity. Only observations with complete predictor data and sufficient subsequent follow-up to determine the three-day outcome were included in the ML analysis.
Preprints 228499 g001
Figure 2. Exploratory analysis of clinical and minimally invasive readouts associated with short-term disease trajectory. (A) Longitudinal trajectories of the four measured variables—bacteraemia (CFU/mL, log scale), clinical score (ordinal 0–5 scale), cumulative weight loss (% from baseline), and daily weight variation (% change from previous day)—across the full observation period. Each dot represents a daily measurement from a single mouse. Observations preceding a fatal outcome within the subsequent three days are shown in purple, whereas observations followed by survival beyond this interval are shown in green. Time is indicated on the x-axis (days post-infection). (B) Heatmap of standardized variable values (z-scores) at the critical time point (i.e., last available observation preceding the fatal outcome or an equivalent matched time point for controls). Rows represent individual animals, and columns represent the four variables. Colour intensity reflects the degree of deviation from the variable mean: yellow tones indicate values above the cohort mean (worsening), while green tones indicate below-average values (more stable). Worsening patterns are apparent in observations preceding a fatal outcome. (C) Principal Component Analysis (PCA) of the same z-scored variables at the critical time point. The first two principal components are plotted, capturing 78.9% (PC1) and 12.2% (PC2) of the variance, respectively. Dots represent individual animals; separation between observations preceding a fatal outcome within three days (purple) and those associated with survival beyond this interval (green) is visible along PC1, indicating multivariate structure among the predictors. .
Figure 2. Exploratory analysis of clinical and minimally invasive readouts associated with short-term disease trajectory. (A) Longitudinal trajectories of the four measured variables—bacteraemia (CFU/mL, log scale), clinical score (ordinal 0–5 scale), cumulative weight loss (% from baseline), and daily weight variation (% change from previous day)—across the full observation period. Each dot represents a daily measurement from a single mouse. Observations preceding a fatal outcome within the subsequent three days are shown in purple, whereas observations followed by survival beyond this interval are shown in green. Time is indicated on the x-axis (days post-infection). (B) Heatmap of standardized variable values (z-scores) at the critical time point (i.e., last available observation preceding the fatal outcome or an equivalent matched time point for controls). Rows represent individual animals, and columns represent the four variables. Colour intensity reflects the degree of deviation from the variable mean: yellow tones indicate values above the cohort mean (worsening), while green tones indicate below-average values (more stable). Worsening patterns are apparent in observations preceding a fatal outcome. (C) Principal Component Analysis (PCA) of the same z-scored variables at the critical time point. The first two principal components are plotted, capturing 78.9% (PC1) and 12.2% (PC2) of the variance, respectively. Dots represent individual animals; separation between observations preceding a fatal outcome within three days (purple) and those associated with survival beyond this interval (green) is visible along PC1, indicating multivariate structure among the predictors. .
Preprints 228499 g002
Figure 3. Cross-validated performance during hyperparameter tuning. Accuracy, sensitivity, and specificity values obtained during 5-fold cross-validation for each hyperparameter configuration tested across four classifiers: Random Forest (RF, dark blue), Support Vector Machine with RBF kernel (SVM, light orange), XGBoost (XGB, dark red), and penalized logistic regression (Logit, light green). Each dot represents the mean performance of a specific hyperparameter combination across folds, and the vertical bars indicate the associated standard error.
Figure 3. Cross-validated performance during hyperparameter tuning. Accuracy, sensitivity, and specificity values obtained during 5-fold cross-validation for each hyperparameter configuration tested across four classifiers: Random Forest (RF, dark blue), Support Vector Machine with RBF kernel (SVM, light orange), XGBoost (XGB, dark red), and penalized logistic regression (Logit, light green). Each dot represents the mean performance of a specific hyperparameter combination across folds, and the vertical bars indicate the associated standard error.
Preprints 228499 g003
Table 1. Study cohorts. The Study 1 dataset was employed for model development while the Study 2 dataset was employed for external validation. *Observations = total number of time-points contributing to the ML analysis.
Table 1. Study cohorts. The Study 1 dataset was employed for model development while the Study 2 dataset was employed for external validation. *Observations = total number of time-points contributing to the ML analysis.
Dataset Total number of mice Mice included after exclusion criteria Variables observed Total observations (after exclusion criteria) * Purpose
Study 1 18 17 4 (cumulative weight loss, daily weight variation, clinical score and bacteraemia) 272 (82) Training + internal testing
Study 2 16 8 4 (cumulative weight loss, daily weight variation, clinical score and bacteraemia) 90 (23) External validation
Table 2. Comparative performance of base and ensemble models on the test dataset. Comparative accuracy, sensitivity, and specificity for the four base classifiers, namely Random Forest, SVM, XGBoost, and penalized logistic regression (Logit), and the three ensemble strategies (majority vote, weighted vote, stacked). For each model, the three-performance metrics are also shown as horizontal bar plots: fuchsia for accuracy, light purple for sensitivity, and dark purple for specificity.
Table 2. Comparative performance of base and ensemble models on the test dataset. Comparative accuracy, sensitivity, and specificity for the four base classifiers, namely Random Forest, SVM, XGBoost, and penalized logistic regression (Logit), and the three ensemble strategies (majority vote, weighted vote, stacked). For each model, the three-performance metrics are also shown as horizontal bar plots: fuchsia for accuracy, light purple for sensitivity, and dark purple for specificity.
Preprints 228499 i001
Table 3. Comparative performance of base and ensemble models on the validation dataset. Comparative accuracy, sensitivity, and specificity for the four base classifiers, namely Random Forest, SVM, XGBoost, and penalized logistic regression (Logit), and the three ensemble strategies (majority vote, weighted vote, stacked). For each model, the three-performance metrics are also shown as horizontal bar plots: fuchsia for accuracy, light purple for sensitivity, and dark purple for specificity.
Table 3. Comparative performance of base and ensemble models on the validation dataset. Comparative accuracy, sensitivity, and specificity for the four base classifiers, namely Random Forest, SVM, XGBoost, and penalized logistic regression (Logit), and the three ensemble strategies (majority vote, weighted vote, stacked). For each model, the three-performance metrics are also shown as horizontal bar plots: fuchsia for accuracy, light purple for sensitivity, and dark purple for specificity.
Preprints 228499 i002
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.