Submitted:
01 September 2026
Posted:
02 September 2026
You are already at the latest version
Abstract
Salmonella enterica comprises genetically and ecologically diverse lineages associated with different hosts, food-production chains, environmental reservoirs, and transmission routes. Although whole-genome sequencing provides extensive information for epidemiological source attribution, ecological classification is complicated by incomplete and inconsistent metadata and the absence of reliable ground-truth datasets. We developed a semi-supervised, genotype-guided hierarchical machine-learning approach for predicting probable ecological associations of S. enterica isolates from genome sequences. In total, 2,406 complete genomes of S. enterica retrieved from NCBI were compared using the Panaroo pangenome analysis tool, and the resulting matrix of single-nucleotide polymorphisms was used for progressive dimensionality reduction, scoring, filtering, and training of a multilevel framework of node-associated decision-making models that associate query sequences with specific habitats. The resulting classification model predicts combinations of habitats in which genomes with similar patterns of diagnostic polymorphisms have been discovered, which may reflect possible dissemination pathways of the pathogen. Testing of the program on a test dataset comprising genomes from NCBI with known sources of isolation showed 62% exact matches. Functional analysis of genes carrying informative mutations revealed associations with the glyoxylate shunt and TCA cycle, carbohydrate and amino-acid utilization, anaerobic respiration, redox metabolism, ion and metal homeostasis, motility, transcriptional regulation, and type III secretion, suggesting that the selected genomic features capture biologically meaningful mechanisms of ecological adaptation. The resulting Salmonella_classifier tool provides a freely available framework for inferring ecological profiles of newly sequenced S. enterica isolates and demonstrates an approach for constructing interpretable genomic classifiers when reliable ecological ground truth is unavailable. The program is available from a public GitHub repository (https://github.com/SeqWord/Classifiers).
Keywords:
Salmonella enterica
; genotyping
; classification
; machine learning
1. Introduction
Salmonella enterica is the causative agent of salmonellosis, a zoonotic infection transmitted through contaminated food or water [1]. Disease severity ranges from mild gastroenteritis to invasive systemic infection that can cause multiple-organ failure. The global emergence of antibiotic-resistant strains complicates treatment and poses a serious public health challenge, particularly in developing countries.
This species comprises numerous serovars, many of which were previously classified as separate species [2]. These diverse pathogenic lineages differ significantly in disease manifestation, host specificity, transmission routes, and environmental persistence. Thus, S. enterica serovars Typhi and Paratyphi spread primarily between humans through food or water contaminated with human feces. They commonly cause systemic enteric fever rather than gastroenteritis [3,4,5]. Sources of infection are often asymptomatic individuals who chronically carry the pathogen. By contrast, S. enterica serovar Typhimurium infects numerous animal species and is transmitted to humans through meat, particularly pork and poultry. It can also spread through contaminated water, fresh produce, contact with animals, and the broader farm environment [6,7].
Serovars Enteritidis, Heidelberg, Infantis, Kentucky, and Hadar are also strongly associated with eggs and poultry meat [8,9], whereas serovars Dublin, Derby, and Choleraesuis are more frequently associated with cattle, pork, and industrial meat-production chains [10,11]. Strains circulating in livestock and poultry production systems frequently develop antibiotic resistance because of the inappropriate use of antibiotics [12,13,14,15]. Serovars Montevideo, Senftenberg, Agona, Oranienburg and Typhimurium are notable for their persistence in low-moisture foods and food-processing environments, including nuts, seeds, spices, cereals, chocolate, and animal feed [16,17]. Their principal survival strategy is a strong tolerance of desiccation, which enables prolonged persistence without active growth [18,19].
Serovars Newport, Saintpaul, Javiana, and Braenderup have caused outbreaks associated with tomatoes, leafy vegetables, melons, sprouts, and other fresh produce [20,21]. Contamination may occur through irrigation water, manure, soil, wildlife, or handling [22]. Some of these strains can attach to plant surfaces, colonize the rhizosphere, or enter internal plant tissues [23].
Many apparently sporadic infections cannot be attributed to a specific source, commonly because the contaminated food or relevant environmental exposure is no longer available when the illness is investigated. Therefore, serovar identity may indicate probable reservoirs and survival strategies but does not, by itself, establish the source of an individual infection. For example, salmonellosis outbreak in Brazil caused by S. enterica serovar Newport was transmitted to humans from parrots [24].
Advances in next-generation sequencing (NGS) technologies have provided more detailed insights into the genetic diversity and plasticity of different S. enterica lineages [25]. Multilocus sequence typing (MLST) has revealed numerous sequence types (STs) with distinct geographical distributions [26,27]. For example, S. enterica serovar Typhimurium ST313 has been identified as a major cause of invasive non-typhoidal salmonellosis, particularly among young children and immunocompromised individuals in sub-Saharan Africa [28]. Another study reported that S. enterica serovar Enteritidis ST11 was the most prevalent sequence type among hospital isolates in Kazakhstan [29].
Although serotyping and genotyping may provide valuable information about possible sources of infection and support epidemiological surveillance, ecological classification has not been a primary objective of most such studies. We hypothesized that adaptation to particular survival and dissemination strategies creates characteristic patterns of genetic polymorphisms in bacterial genomes, producing distinguishable allelic profiles even among strains that evolved from different lineages. The availability of numerous whole-genome sequences in public databases, together with metadata describing isolation sources and phenotypic characteristics, enables the application of machine-learning methods and neural-network models to the ecological classification of isolates [30,31,32].
However, the application of these computational approaches is hindered by inconsistent standards for recording metadata in public databases, erroneous records resulting from contamination or inaccurate reporting at the time of sampling, and the subjective assignment of bacterial strains to ecological categories. These limitations impede the construction of reliable ground-truth datasets required for supervised neural-network training. Nevertheless, previous studies have demonstrated the feasibility of attributing Salmonella isolates to probable animal hosts or zoonotic sources using machine-learning methods, including neural networks [30,31]. Despite their promising results, these studies did not provide maintained, ready-to-use software that would enable non-specialist users to classify newly sequenced Salmonella isolates. In the present study, we implemented a semi-supervised, genotype-guided hierarchical learning network to predict the probable habitats and ecovariants of S. enterica isolates and to identify the major genes involved in the ecological adaptation of these pathogens.
2. Materials and Methods
In total, 5,893 complete genomes of S. enterica strains, together with all associated metadata records, were retrieved from the NCBI database.
Core genome alignment and a matrix of single-nucleotide polymorphisms were generated using Panaroo v.1.6.0 [33].
The Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN) algorithm, implemented in the Python module hdbscan v.0.8.44, was used for unsupervised clustering.
The chi-square test implemented in the scipy.stats module v.1.17.1 was used for feature scoring. To control for multiple testing, p-values were corrected using the Benjamini–Hochberg false discovery rate (FDR) procedure implemented in statsmodels v.0.14.6. Custom scripts were developed in Python 3.14.4. An automatic model-selection script was used to evaluate eight supervised machine-learning algorithms: logistic regression (LR), support vector classification (SVC), Random Forest (RF), decision tree (DT), multilayer perceptron (MLP), k-nearest neighbours (KNN), categorical Naive Bayes (NBayes), and XGBoost. Candidate models were implemented using the scikit-learn Python package v.1.8.0, with XGBoost provided by the xgboost package v.3.4.1. Categorical SNP states were one-hot encoded for all models except NBayes, which used ordinal encoding; MLP inputs were additionally standardized using sklearn.preprocessing.StandardScaler. Candidate models were ranked by mean cross-validation balanced accuracy, and the highest-scoring algorithm was selected. Model parameters were heuristically adjusted according to sample size, feature dimensionality, number of classes, data type, and class imbalance; class weighting was applied when the largest-to-smallest class ratio was ≥2. When stratified cross-validation was not possible, NBayes was used as a conservative fallback for nodes with < 50 samples and RF for larger nodes.
The OpenAI ChatGPT service was used to assist with custom Python scripting and the design of the architecture of the model PKL files.
3. Results
This section may be divided by subheadings. It should provide a concise and precise description of the experimental results, their interpretation, and the experimental conclusions that can be drawn.
3.1. Machine-Learning Algorithm
3.1.1. Genomic Data Collection and Ontology-Based Grouping
The ontology-based classification scheme shown in Table 1 was developed to assign S. enterica strains to two-level ecological categories.
The resulting two-level hierarchy was used as input for the ontology-guided supervision of the ML workflow by assigning composite labels to genome sequences. These labels do not represent genomic ground truth but rather prior biological hypotheses inferred from metadata according to predefined rules. Of the 5,893 S. enterica genomes retrieved from NCBI, 2,406 were assigned to two-level ecological categories through keyword searches. A genome could be assigned to multiple categories when its metadata contained keywords associated with different categories. The resulting dataset comprised 2,408 records, including only a few genomes assigned to more than one category.
3.1.2. Genomic Feature Selection for Machine Learning
The aim of this step was to construct a matrix of single-nucleotide polymorphisms (SNPs) suitable for distinguishing among S. enterica habitats. The selected features needed to be sufficiently diverse to ensure their broad applicability. A feature-discovery cohort was created by randomly selecting five reference strains from each composite category. Because Table 1 defined 20 unique two-level composite categories, 100 strains were selected for the feature-discovery cohort.
The selected genomes were analyzed using Panaroo, a pangenome analysis tool that generates a core-genome SNP matrix for a given set of genomes. SNPs with a minor allele frequency below 5% were removed to exclude variants specific to individual genomes and potential sequencing errors. SNP positions containing non-canonical nucleotide symbols were also removed. In total, 172,397 SNPs were retained.
In parallel, an SNP annotation file was created containing the following fields: a unique SNP ID; the SNP position in the core-genome alignment generated by Panaroo; its position and DNA strand in the reference genome; the observed allelic states (e.g., A/T); the locus tag of the corresponding reference-genome gene, if the SNP was located within a coding sequence; the protein-coding gene annotation, where applicable; the resulting amino acid substitution for nonsynonymous mutations (e.g., N131K); and an 81-bp context sequence containing the variable nucleotide at the central position, flanked by 40-bp upstream and downstream DNA sequences.
Genome assembly GCA_053410235.1 was selected as the reference genome. There was no specific biological rationale for selecting this particular assembly beyond its high level of completeness. Because the reference genome was used exclusively for SNP annotation, its selection did not influence any subsequent ML steps.
An additional file was generated containing the most common allelic state at each selected SNP position, as determined from the core-genome alignment of all genomes. Using this information, the SNP dataset was converted from a nucleotide matrix (A/T/G/C) into a binary matrix (0/1), in which 0 represented the most common allelic state and 1 represented the minor allelic state or states. This conversion resulted in little information loss because only 5,111 positions contained more than two allelic states (approximately 3% of all selected positions). At the same time, representing allelic states numerically facilitated subsequent pattern comparisons.
The number of potentially diagnostic SNPs was substantially reduced during subsequent processing steps. At each reduction step, the SNP annotation and allelic-state datasets were updated to keep their records synchronized. During the first reduction step, the annotation data were used to remove SNPs located in non-coding regions and synonymous substitutions within coding sequences. This step reduced the SNP matrix to 35,838 positions.
3.1.3. Unsupervised Clustering and Further Dimensionality Reduction
It can be assumed that S. enterica strains sharing the same habitat do not necessarily belong to phylogenetically related lineages. Nevertheless, adaptation to similar environmental conditions may result in similar patterns of adaptive genetic changes, while lineage-specific SNP patterns are also retained. These lineage- and environment-specific signals must be distinguished because they may produce conflicting patterns. To address this issue, unsupervised clustering was applied to the matrix of allelic states at 35,838 SNP positions in the feature-discovery cohort of genomes.
For unsupervised clustering, the Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN) algorithm was selected because it is well suited to high-dimensional binary datasets for which the number of clusters is not predefined. Its ability to identify noise is particularly relevant because metadata-derived ecological categories are not necessarily expected to correspond closely to genomic distances. The HDBSCAN parameters were configured as follows: Jaccard distance was used because it is appropriate for binary SNP data; the minimum cluster size was set to 10 genomes; and the min_samples parameter was set to 5. The cluster identifiers generated by HDBSCAN were added as the first level of the composite labels, as illustrated below. Vertical bars separate the cluster identifier, hierarchical ecological categories, and genome accession:
1|Environment|Environmental swabs and sponges|GCA_053410235.1
2|Environment|Environmental swabs and sponges|GCA_053411835.1
5|Environment|Environmental swabs and sponges|GCA_053410975.1
In this example, three genomes identified from their metadata as environmental isolates belonging to the same second-level category, “Environmental swabs and sponges,” were assigned to three different genomic clusters (1, 2, and 5), because they represented distinct genetic lineages.
The next step assessed whether SNP allele distributions were associated with the proposed three-level hierarchical classification of S. enterica genomes. For each SNP, a contingency table was constructed from its observed allelic states and the genome groups defined at each hierarchical level. A chi-square test of independence was then performed separately for every SNP at every hierarchy level. To account for multiple testing, the resulting p-values were adjusted using the Benjamini-Hochberg procedure. The script calculated two types of adjusted p-values: a level-specific FDR-adjusted p-value across all SNPs tested at a particular hierarchy level and a global FDR-adjusted p-value across all valid SNP-by-level tests. The global correction was used for feature selection because a SNP was retained if it was significantly associated with the grouping at least at one hierarchical level. SNPs with a global FDR-adjusted p-value below 0.05 were classified as potentially diagnostic features. This analysis retained 21,290 SNPs associated with at least one level of the proposed hierarchical classification.
A small number of significant SNPs would indicate limited genomic support for the proposed classification. Conversely, a large number of retained SNPs indicates a strong statistical association between allelic distributions and at least one hierarchy level.
Further dimensionality reduction was performed by controlling for redundancy among allelic-state patterns and for the genomic proximity of selected SNPs. Although neural-network models do not strictly require independent input features, highly correlated SNPs provide little additional information, may disproportionately weight particular genomic signals, and increase model complexity.
First, SNPs were grouped into genomic neighborhoods according to their coordinates in the reference genome. SNPs were assigned to the same neighborhood when consecutive SNP positions were separated by no more than 80 bp. Because each SNP is subsequently identified using an 81-bp BLASTN query sequence containing 40-bp upstream and downstream flanking regions, closely located SNPs may have overlapping query sequences and may interfere with reliable allele calling. Therefore, only the statistically highest-ranked SNP was retained from each genomic neighborhood.
Redundancy among the remaining SNPs was then assessed by comparing their allelic-state patterns across all genomes in the feature-discovery cohort. Pairwise dissimilarity was calculated as the proportion of genomes with mismatching allelic states, equivalent to the normalized Hamming distance. The SNPs were grouped using complete-linkage hierarchical clustering with a maximum mismatch threshold of 10%. Consequently, every pair of SNP patterns within a redundancy cluster was at least 90% identical. Up to five SNPs were retained from each cluster, prioritizing those with the lowest global FDR-adjusted p-values; level-specific FDR-adjusted and raw p-values were used as successive tie-breakers. Retaining several representatives from each redundancy cluster provides robustness when genomic regions containing otherwise informative SNPs are absent or cannot be reliably analyzed in newly sequenced genomes.
These filtering steps reduced the feature set to 183 SNPs, providing a manageable input matrix for subsequent ML analyses. However, the final classifier determines the allelic state of each SNP through a separate BLASTN search, making the analysis of all 183 loci time demanding. The remaining SNPs were therefore ranked according to their statistical significance, and the 100 highest-ranked SNPs were retained for model development.
3.1.4. SNP Calling and Machine Learning
Allelic states at the 100 selected diagnostic SNP positions were determined for all 2,406 S. enterica genomes assigned to ecological categories by aligning SNP-context sequences against the genome assemblies using BLASTN. An additional hierarchy level representing genomic lineage was generated by unsupervised clustering with HDBSCAN, as described above. The resulting matrix was divided into non-overlapping training and test datasets. The test dataset comprised 554 genomes randomly selected to represent the different ecological categories, whereas the remaining genomes were assigned to the training dataset.
A separate supervised classification model was trained at each internal node of the hierarchical classification tree to distinguish among its immediate child nodes. Candidate models were compared using mean balanced accuracy obtained through stratified cross-validation, with a maximum of five folds. The number of folds was reduced when required by the size of the smallest child class. Balanced accuracy was used to reduce bias caused by unequal representation of ecological categories. The highest-scoring algorithm was selected independently at each eligible node and was subsequently trained using all training genomes assigned to that node. Supervised models were trained only at splitting nodes containing at least 10 genomes and two or more child classes. Nodes containing fewer than 10 genomes were classified using allele-profile similarity rather than a fitted supervised model. For nodes with only one possible child, that child was assigned directly without model fitting. Allele-frequency profiles were calculated and stored for every node and each of its immediate children. These profiles provide a fallback when supervised-model training fails or when a prediction is considered unreliable. In particular, a model prediction is rejected when its normalized Shannon entropy exceeds 0.80, indicating an uncertain probability distribution across the candidate child classes; classification then proceeds according to similarity between the query SNP profile and the stored child profiles. The same profile-based procedure may be used when missing SNP calls prevent reliable model-based prediction.
The resulting hierarchical ensemble of supervised machine-learning models was packaged in a single model.pkl file. In addition to the fitted node-specific models, the file contained the hierarchical network structure; SNP-context sequences used for BLASTN searches; common allelic states; SNP annotations; allele-frequency profiles for all nodes and their immediate children; model-selection parameters and scores; training-set accuracy values; the reference-genome identifier; the original training matrix; and descriptive metadata. This self-contained model package was used to classify genomes in the test dataset and can also be applied to newly sequenced S. enterica genomes.
Evaluation using the independent test dataset resulted in correct habitat classification for 52% of the 554 genomes. The initial model was subsequently applied to all 2,406 ecologically categorized genomes to identify recurrent patterns of cross-identification. For each genome, all successful terminal predictions were collected from the classification reports. Within each original HDBSCAN cluster, genomes sharing the same complete set of terminal predictions were assigned to the same cross-group. Genomes without a successful terminal prediction were assigned to a separate cross-group with the signature “Not identified.” Cross-group identifiers were numbered independently within each HDBSCAN cluster and inserted between the top-level HDBSCAN clusters and the two ontology-derived ecological levels creating a four-level hierarchy:
HDBSCAN cluster | cross_group | first-level category | second-level category
The hierarchical ensemble was then retrained using the training dataset and the revised four-level labels. Incorporating the cross-group level increased the proportion of genomes receiving correct terminal habitat predictions from 52% to 62%. The architecture of the final four-level hierarchical classification network stored in the model.pkl file is shown in Figure 1. The figure indicates the supervised machine-learning algorithm selected at each internal node and its training accuracy. This value represents the proportion of genomes assigned to the correct immediate child node when the fitted model was applied to the same data used for its training. An accuracy of 1.0 indicates complete separation of the child classes in the training dataset, whereas lower values indicate that their SNP profiles overlap and are more difficult to distinguish. Nodes without a fitted model were classified using allele-profile similarity and therefore have no model-training accuracy value. Because these values were calculated from the training data, they describe model fit rather than predictive performance on independent genomes and may overestimate generalization accuracy.
3.2. Salmonella Enterica Habitat Classifier
3.2.1. Program Overview
The trained model was incorporated into a software framework developed in Python 3. The resulting program, Salmonella_classifier, is available from the SeqWord Classifiers repository (https://github.com/SeqWord/Classifiers).
The program does not require a formal installation and can be run after been downloaded and the required Python dependencies have been installed. The program directory contains the main script, classifier.py, together with the ‘input’, ‘output’, ‘model’, and supporting program directories. Genome assemblies in FASTA or GenBank format may contain either a single chromosome sequence or multiple contigs. GZ-archives of these files can also be used as an input. Input files should be placed in a project subfolder newly created by the user in folder ‘input’. Classification can then be initiated using the following command:
python3 classifier.py new_project_folder
The program saves the classification results as text files in folder ./output/new_project_folder created by the program.
Each query genome is listed in the output file, followed by an indented list of predicted ecological categories ranked according to their classification scores. A query genome may reach multiple terminal nodes, producing a composite profile of predicted ecological associations. If none of the possible terminal categories satisfies the classification criteria, the query is reported as “Not identified”.
Users can adjust several additional command-line parameters that control the sensitivity and specificity of predictions, as described on the project website.
3.2.2. Validation of the Program Using Newly Sequenced Genomes
We applied the program to several recently sequenced S. enterica genomes described in our previous publication [29]. These strains were isolated from clinical samples collected in Kazakhstan, and most belonged to MLST ST11. The ST11 isolates were genetically homogeneous; however, one strain, 19S, exhibited a distinct genome-wide adenine-methylation pattern. We hypothesized that this alternative methylation pattern might be associated with adaptation to a different habitat or life-cycle strategy. The classifier results were consistent with this hypothesis. 19S was assigned a combination of probable ecological categories that differed from those predicted for the other ST11 isolates, which were grouped together by the classifier as likely waste water and clinical isolates. Conversely, the program suggested an association of the strain 19S livestock and food chain environment. The predictions for 19S and a representative ST11 strain, 1S, are compared in Table 2.
These results indicate that the classifier distinguished 19S from the otherwise genetically homogeneous ST11 group. Although this observation does not establish a causal relationship between methylation and ecological adaptation, it provides a basis for further investigation of the distinctive biological characteristics of the isolate 19S
As a specificity control, sequences from the closely related species Escherichia coli, as well as datasets containing mixtures of Salmonella and Escherichia contigs, were submitted to the classifier. In both cases, the program returned “Not identified,” indicating that these sequences did not produce a supported terminal classification within the trained S. enterica hierarchy.
3.3. Functional Analysis of Genes Bearing Marker Mutations
Assuming that the patterns of habitat-specific mutations are formed due to adaptive changes to specific environments and pathogen survival strategies, an analysis of functions of genes bearing these mutations can tell great deal about molecular mechanisms of adaptation. The analysis of the distribution of 183 top-selected marker SNPs representing non-synonymous mutations in protein-coding genes showed non-random scattering of these polymorphisms across housekeeping proteins summarized in Table 3.
4. Discussion
4.1. Classification Model Design
This study presents a semi-supervised algorithm for constructing a hierarchical framework of trained models to predict the probable ecological habitats of S. enterica isolates from their SNP genotypes. The workflow begins with human-interpretable ecological concepts, which are translated into provisional two-level labels. The genomic data are then used to determine whether the proposed categories are genetically distinguishable. The initial two-level hierarchy therefore provides ontology-guided weak supervision for genome classification. An additional level generated through unsupervised HDBSCAN clustering accounts for genomic relatedness and helps stratify lineage-associated variation before ecological associations are evaluated.
The resulting three-level hierarchy should not be regarded as ground truth. Rather, it represents a prior biological hypothesis inferred from metadata according to expert-defined rules and supplemented with genomic clustering [34] The proposed workflow circumvents the requirement of many supervised machine-learning methods for a fully validated ground-truth dataset at the beginning of the analysis. Instead, it evaluates the initial hypothesis by testing whether SNP allele distributions are significantly associated with the proposed hierarchical categories. The number of SNPs passing the global FDR threshold, as well as their level-specific FDR-adjusted and unadjusted p-values, depends on the extent to which the proposed hypothesis corresponds to reproducible genomic patterns. An arbitrary or biologically uninformative grouping would be expected to yield none or few significant features at the ecological hierarchy levels, whereas enrichment of significant associations provides evidence that the proposed categories capture reproducible genomic structure.
After establishing that the proposed hierarchy was associated with a substantial set of genomic features, we performed further dimensionality reduction. Principal component analysis (PCA) is commonly used to reduce a large set of correlated features to a small number of composite variables [31]. However, the principal component eigenvalues cannot usually be attributed directly to individual polymorphisms, genes, or other readily interpretable biological characteristics. PCA was therefore unsuitable for the aims of this study, which included identifying candidate molecular mechanisms underlying the adaptation and dissemination of Salmonella pathogens, in addition to assigning isolates to probable ecological niches.
Instead, a series of feature-filtering procedures was used to select informative SNPs while controlling for redundant allele patterns and the genomic proximity of polymorphic sites. Allelic states at the resulting subset of informative SNP positions were then determined in the complete collection of categorized genomes. This targeted approach substantially reduced the computational cost of feature extraction while preserving the biological interpretability of individual SNPs and their associated genes.
The resulting matrix of allelic states for the training genomes was used to construct a hierarchical ensemble of supervised classifiers, with a separate model trained at each eligible internal node (Figure 1). Different machine-learning algorithms were evaluated independently at each node, and the best-performing algorithm was selected to distinguish among its immediate child nodes. Nodes containing insufficient training data, nodes for which model fitting was unsuccessful, and uncertain predictions were handled using stored allele-frequency profiles. Efficacy of this approach was demonstrated previously in the model trained for identification of patterns of antibiotic resistance of Mycobacterium tuberculosis isolates [35].
Cross-validation and subsequent application of the initial model revealed recurrent cross-identification among some ecological categories. These patterns were used to introduce an additional cross-group level between the top-level genomic clusters and the two ontology-derived ecological levels. Similar approaches of model improving were discussed previously [36]. This post-training hierarchy refinement aimed at latent class discovery resulted in four-level hierarchy with an improved correspondence between the predicted and metadata-derived terminal categories.
The resulting program, which is freely available from a public GitHub repository, returns a ranked set of ecological categories assigned to each query genome, as illustrated in Table 2. These categories should not be interpreted merely as mutually exclusive alternatives. Collectively, they constitute a predicted ecological profile that may reflect multiple habitats, reservoirs, transmission routes, or clinical sources associated with genetically similar isolates. For example, the output for strain 1S in Table 2 associates its SNP profile with categories representing sediment and waste, livestock, food-chain products, and several types of clinical samples. These predictions suggest possible ecological and epidemiological associations; however, do not, by themselves, demonstrate a specific transmission pathway or clinical outcome. In Figure 1, the terminal ecological categories are represented by circles of different colors. The combination of terminal categories reached by a query genome describes a predicted survival and dissemination strategy associated with specific genomic profiles.
4.2. Molecular Mechanisms of Adaptation of S. enterica to Versatile Ecological Niches
This study presents a semi-supervised algorithm for constructing a hierarchical framework of trained models to predict the probable ecological habitats of S. enterica isolates from their SNP genotypes. The workflow begins with human-interpretable ecological concepts, which are translated into provisional two-level labels. The genomic data are then used to determine whether the proposed categories are genetically distinguishable.
The analysis of functional genes bearing mutations associated with ecological adaptation provides important insights into the strategies and molecular mechanisms exploited by these pathogens (Table 3). Of particular interest are the genes:
iclR transcriptional regulator: glyoxylate shunt ← aceA ⇄ icd → aceK → TCA
which form a tightly connected regulatory and metabolic module controlling the balance between the TCA cycle and the glyoxylate shunt, particularly when bacteria utilize acetate or fatty acids as major carbon sources. The glyoxylate shunt bypasses the two CO2-generating steps of the TCA cycle, thereby conserving carbon skeletons that can be channeled into anabolic pathways through gluconeogenesis [37]. Microorganisms with an inactive or absent glyoxylate shunt may utilize lipids for energy production but have a limited capacity to convert lipid-derived acetyl-CoA into biomass. This metabolic capability may therefore contribute to the preferential colonization and persistence of pathogens in particular host organs and tissues [38]. For example, the ability of Mycobacterium tuberculosis and Pseudomonas aeruginosa to persist in lipid-rich host environments, including the lungs, has been associated with the activity of the glyoxylate shunt and utilization of host-derived fatty acids [39,40]. Other host compartments, such as lymphoid tissues and the brain, can also provide lipid-rich nutritional niches, whereas the intestinal tract and bloodstream generally offer greater access to carbohydrates and other readily metabolizable carbon sources. Environmental habitats such as wastewater and sediments may likewise be relatively depleted in readily available carbohydrates while containing diverse organic acids and lipid-derived compounds. Adaptation to these contrasting nutritional environments requires precise regulation of catabolic and anabolic pathways, particularly the balance between glycolysis, the TCA cycle, the glyoxylate shunt, and gluconeogenesis. Strain-specific mutations in proteins associated with anaerobic respiration, including the NarX nitrate/nitrite sensor, the NarK nitrate/nitrite transporter, and nitrate reductases, which catalyses a key reaction in nitrate respiration, contribute to the versatility of Salmonella adaptation to diverse environments [41,42].
There is a broad collection of carbohydrate-related proteins bearing strain-specific mutations in different S. enterica isolates. These proteins are involved in the metabolism of mannose, sorbitol/glucitol, chitobiose/N-acetylglucosamine, fucose, and β-glucosides, as well as in central carbohydrate metabolism. This finding suggests that a substantial fraction of strain-specific diversification is associated with nutritional niche specialization rather than merely with modifications of core glycolysis. Salmonella pathogens disseminated via contaminated plants can actively degrade plant-derived polysaccharides including cellulose, providing substrates for energy production and potentially facilitating access to internal plant tissues [23]. On the other hand, mutations in fucK and nagK/chbC are particularly interesting in Salmonella, because fucose and amino sugars are abundant in warm-blooded host-associated environments, where they occur in host glycans and mucus [43,44].
The type III secretion system is one of the major virulence factors of Salmonella, used by the pathogen to inject effector molecules directly into the cytoplasm of host cells, including those of warm-blooded animals and plants. The large number of type III secretion system components bearing environment-associated mutations (Table 3) highlights the importance of this secretion system in the adaptation of the pathogen to different ecological niches [45].
Several other adaptation mechanisms can be inferred from the list of genes bearing strain-specific mutations in Table 3, including those encoding enzymes involved in amino acid utilization and nitrogen metabolism, aromatic compound utilization, the synthesis and salvage of important vitamins and cofactors, as well as global regulators of gene transcription.
4.3. Prospects for Improving the Accuracy of the Classifier
The performance of the current version of the classifier can be improved through reciprocal analysis and refinement of the initial grouping of genomes into two-level ecological categories. Although the accuracy values assigned to individual nodes of the classification framework (Figure 1), which describe model fit, should not be interpreted as general estimates of overall system performance, they can identify parts of the framework that require improvement. Such improvements may be achieved by adding more genomes from the relevant ecological categories to the training dataset, restructuring the ecological grouping of these genomes, or searching for additional latent layers within the classification framework.
An additional source of information for model training is the matrix of accessory genes present in subsets of genomes. This matrix is automatically generated by Panaroo, a program designed for pangenomic analysis, and is therefore readily available. It uses a binary 0/1 encoding similar to that employed in the current version of the classifier for alternative SNP alleles. Thus, the SNP and accessory-gene matrices can, in principle, be combined. The applicability of accessory-gene presence/absence matrices to bacterial classification has been demonstrated in previous studies (Lupolova et al., 2019). However, we decided against using this matrix in the current study and reserved its incorporation for future updates for several reasons.
The major concern is the difficulty in distinguishing the true absence of an accessory gene from a false-negative result caused by incompleteness of the query genome or by a high level of sequence divergence within the sequence context used for BLASTN alignment. The current classifier, which is based on patterns of SNP alleles, treats BLASTN alignment failures as missing data rather than as a specific allelic state. In contrast, when accessory-gene presence/absence is encoded as a binary variable, an alignment failure could be incorrectly interpreted as gene absence. Nevertheless, we do not exclude the possibility that careful selection of conserved, reliably detectable, and ecologically informative accessory genes could improve the accuracy of future versions of the classifier.
5. Conclusions
This study demonstrates that ecological information encoded in heterogeneous and potentially incomplete genome metadata can be combined with genomic variation to construct an interpretable hierarchical framework for predicting the probable habitats and survival strategies of S. enterica. Rather than treating metadata-derived ecological categories as ground truth, the proposed semi-supervised approach uses them as an initial biological hypothesis and subsequently evaluates and refines this hypothesis according to patterns present in the genomic data. Unsupervised clustering, statistical selection of habitat-associated SNPs, redundancy reduction, and node-specific supervised classification together provide a framework capable of accommodating both phylogenetic structure and ecological diversification.
An important outcome of the study is the biological interpretability of the selected genomic markers. Genes bearing habitat-associated nonsynonymous mutations participate in interconnected processes including carbon and nitrogen utilization, the balance between the TCA cycle and glyoxylate shunt, anaerobic respiration, redox metabolism, vitamin and cofactor metabolism, environmental sensing, motility, stress responses, transcriptional regulation, and type III secretion. Their functional diversity is consistent with adaptation of S. enterica to markedly different nutritional, physicochemical, environmental, and host-associated niches.
The developed Salmonella_classifier translates this approach into a freely available tool applicable to newly sequenced genomes. Its architecture also provides opportunities for iterative improvement as additional genomes and better ecological metadata become available. More broadly, the approach illustrates how weak ecological supervision combined with unsupervised genomic structure and interpretable molecular markers can be used to investigate microbial ecological adaptation when a reliable ground-truth classification cannot be established in advance.
Author Contributions
software, D.T.Y and O.N.R.; validation, D.T.Y, A.A.A. and O.N.R.; supervision, A.A.A. and O.N.R.; funding acquisition, A.A.A.; writing—review and editing, D.T.T., A.A.A. and O.N.R.
Funding
This study was supported by the grant project IRN AP23490794 (2024–2026), “Investigation of the Distribution of Salmonellosis and Enteroviral Infection Pathogens Using PCR-Based Diagnostic Assays Developed to Improve Epidemiological Surveillance.”
Data Availability Statement
The classifier program is available from the public GitHub repository at https://github.com/SeqWord/Classifiers.
Conflicts of Interest
The authors declare no conflicts of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| DT | Decision tree |
| KNN | k-nearest neighbours |
| LR | Logistic regression |
| ML | Machine Learning |
| MLP | Multilayer perceptron |
| RF | Random Forest |
| SVC | Support vector classification |
References
- Ranjan, A.; Chandna, M.; Stevens, N.J.; Kandil, J.; Dinh, B.; Kuhn, M.; Mian, N.; Tran, B.; Hamid, A.; Kim, P.; Desin, T.S. Salmonella Infections: Global Trends and Emerging Challenges. Microorganisms 2026, 14(4), 816. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
- Ryan, M.P.; O'Dwyer, J.; Adley, C.C. Evaluation of the Complex Nomenclature of the Clinically and Veterinary Significant Pathogen Salmonella. BioMed Res. Int. 2017, 2017, 3782182. [Google Scholar] [CrossRef]
- Parry, C.M.; Hien, T.T.; Dougan, G.; White, N.J.; Farrar, J.J. Typhoid fever. N Engl. J. Med. 2002, 347(22), 1770–82. [Google Scholar] [CrossRef]
- Meiring, J.E.; Khanam, F.; Basnyat, B.; Charles, R.C.; Crump, J.A.; Debellut, F.; Holt, K.E.; Kariuki, S.; Mugisha, E.; Neuzil, K.M.; Parry, C.M.; Pitzer, V.E.; Pollard, A.J.; Qadri, F.; Gordon, M.A. Typhoid fever. Nat. Rev. Dis. Primers 2023, 9(1), 71. [Google Scholar] [CrossRef]
- Kuehn, R.; Rahden, P.; Hussain, H.S.; Karkey, A.; Qamar, F.N.; Rupali, P.; Parry, C.M. Enteric (typhoid and paratyphoid) fever. Lancet 2025, 406(10509), 1283–1294. [Google Scholar] [CrossRef]
- Sun, H.; Wan, Y.; Du, P.; Bai, L. The Epidemiology of Monophasic Salmonella Typhimurium. Foodborne Pathog. Dis. 2020, 17(2), 87–97. [Google Scholar] [CrossRef]
- Simon, S.; Lamparter, M.C.; Pietsch, M.; Borowiak, M.; Fruth, A.; Rabsch, W.; Fischer, J. The zoonotic agent Salmonella. Zoonoses Infect. Affect. Hum. Anim. 2023, 295–327. [Google Scholar] [CrossRef]
- Gantois, I.; Ducatelle, R.; Pasmans, F.; Haesebrouck, F.; Gast, R.; Humphrey, T.J.; Van Immerseel, F. Mechanisms of egg contamination by Salmonella Enteritidis. FEMS Microbiol. Rev. 2009, 33(4), 718–38. [Google Scholar] [CrossRef]
- Foley, S.L.; Nayak, R.; Hanning, I.B.; Johnson, T.J.; Han, J.; Ricke, S.C. Population dynamics of Salmonella enterica serotypes in commercial egg and poultry production. Appl. Environ. Microbiol. 2011, 77(13), 4273–9. [Google Scholar] [CrossRef]
- Chiu, C.H.; Tang, P.; Chu, C.; Hu, S.; Bao, Q.; Yu, J.; Chou, Y.Y.; Wang, H.S.; Lee, Y.S. The genome sequence of Salmonella enterica serovar Choleraesuis, a highly invasive and resistant zoonotic pathogen. Nucleic Acids Res. 2005, 33(5), 1690–1698. [Google Scholar] [CrossRef]
- Bonardi, S. Salmonella in the pork production chain and its impact on human health in the European Union. Epidemiol. Infect. 2017, 145(8), 1513–1526. [Google Scholar] [CrossRef]
- Van Hoorebeke, S.; Van Immerseel, F.; Haesebrouck, F.; Ducatelle, R.; Dewulf, J. The influence of the housing system on Salmonella infections in laying hens: a review. Zoonoses Public Health 2011, 58(5), 304–311. [Google Scholar] [CrossRef]
- Ferrari, R.G.; Rosario, D.K.; Cunha-Neto, A.; Mano, S.B.; Figueiredo, E.E.; Conte-Junior, C.A. Worldwide epidemiology of Salmonella serovars in animal-based foods: a meta-analysis. Appl. Environ. Microbiol. 2019, 85(14), e00591-19. [Google Scholar] [CrossRef]
- Kenney, S.M.; M'ikanatha, N.M.; Ganda, E. Genomic evolution of Salmonella Dublin in cattle and humans in the United States. Appl. Environ. Microbiol. 2025, 91(9), e0068925. [Google Scholar] [CrossRef]
- Wang, Y.; Xu, X.; Jia, S.; Qu, M.; Pei, Y.; Qiu, S.; Zhang, J.; Liu, Y.; Ma, S.; Lyu, N.; Hu, Y.; Li, J.; Zhang, E.; Wan, B.; Zhu, B.; Gao, G.F. A global atlas and drivers of antimicrobial resistance in Salmonella during 1900-2023. Nat. Commun. 2025, 16(1), 4611. [Google Scholar] [CrossRef]
- Liu, F.; Duan, P.; Xiao, H.; Zhang, H.; Guo, H.; Zhang, R.; Jiang, S. Investigation into the occurrence and molecular characteristics of Salmonella from food animals in Shandong, China. Poult. Sci. 2025, 104(10), 105628. [Google Scholar] [CrossRef]
- Podolak, R.; Enache, E.; Stone, W.; Black, D.G.; Elliott, P.H. Sources and risk factors for contamination, survival, persistence, and heat resistance of Salmonella in low-moisture foods. J. Food Prot. 2010, 73(10), 1919–36. [Google Scholar] [CrossRef]
- Gruzdev, N.; Pinto, R.; Sela, S. Effect of desiccation on tolerance of Salmonella enterica to multiple stresses. Appl. Environ. Microbiol. 2011, 77(5), 1667–1673. [Google Scholar] [CrossRef]
- Finn, S.; Condell, O.; McClure, P.; Amézquita, A.; Fanning, S. Mechanisms of survival, responses and sources of Salmonella in low-moisture environments. Front Microbiol. 2013, 4, 331. [Google Scholar] [CrossRef]
- Mody, R.K.; Greene, S.A.; Gaul, L.; Sever, A.; Pichette, S.; Zambrana, I.; Dang, T.; Gass, A.; Wood, R.; Herman, K.; Cantwell, L.B.; Falkenhorst, G.; Wannemuehler, K.; Hoekstra, R.M.; McCullum, I.; Cone, A.; Franklin, L.; Austin, J.; Delea, K.; Behravesh, C.B.; Sodha, S.V.; Yee, J.C.; Emanuel, B.; Al-Khaldi, S.F.; Jefferson, V.; Williams, I.T.; Griffin, P.M.; Swerdlow, D.L. National outbreak of Salmonella serotype saintpaul infections: importance of Texas restaurant investigations in implicating jalapeño peppers. PLoS ONE 2011, 6(2), e16579. [Google Scholar] [CrossRef]
- Olaimat, A.N.; Holley, R.A. Factors influencing the microbial safety of fresh produce: a review. Food Microbiol. 2012, 32(1), 1–19. [Google Scholar] [CrossRef] [PubMed]
- Hintz, L.D.; Boyer, R.R.; Ponder, M.A.; Williams, R.C.; Rideout, S.L. Recovery of Salmonella enterica Newport introduced through irrigation water from tomato (Lycopersicum esculentum) fruit, roots, stems, and leaves. HortScience 2010, 45(4), 675–678. [Google Scholar] [CrossRef]
- Karmakar, K.; Nath, U.; Nataraja, K.N.; Chakravortty, D. Root mediated uptake of Salmonella is different from phyto-pathogen and associated with the colonization of edible organs. BMC Plant Biol. 2018, 18(1), 344. [Google Scholar] [CrossRef]
- Saidenberg, A.B.S.; Stegger, M.; Semmler, T.; Rocha, V.G.P.; Cunha, M.P.V.; Souza, V.A.F.; Cristina Menão, M.; Milanelo, L.; Petri, B.S.S.; Knöbl, T. Salmonella Newport outbreak in Brazilian parrots: confiscated birds from the illegal pet trade as possible zoonotic sources. Environ. Microbiol. Rep. 2021, 13(5), 702–707. [Google Scholar] [CrossRef]
- Alikhan, N.F.; Zhou, Z.; Sergeant, M.J.; Achtman, M. A genomic overview of the population structure of Salmonella. PLoS Genet. 2018, 14(4), e1007261. [Google Scholar] [CrossRef]
- Achtman, M.; Wain, J.; Weill, F.X.; Nair, S.; Zhou, Z.; Sangal, V.; Krauland, M.G.; Hale, J.L.; Harbottle, H.; Uesbeck, A.; Dougan, G.; Harrison, L.H.; Brisse, S. S. enterica MLST Study Group. Multilocus sequence typing as a replacement for serotyping in Salmonella enterica. PLoS Pathog. 2012, 8(6), e1002776. [Google Scholar] [CrossRef]
- Pijnacker, R.; van den Beld, M.; van der Zwaluw, K.; Verbruggen, A.; Coipan, C.; Segura, A.H.; Mughini-Gras, L.; Franz, E.; Bosch, T. Comparing Multiple Locus Variable-Number Tandem Repeat Analyses with Whole-Genome Sequencing as Typing Method for Salmonella Enteritidis Surveillance in The Netherlands, January 2019 to March 2020. Microbiol. Spectr. 2022, 10(5), e0137522. [Google Scholar] [CrossRef]
- Okoro, C.K.; Barquist, L.; Connor, T.R.; Harris, S.R.; Clare, S.; Stevens, M.P.; Arends, M.J.; Hale, C.; Kane, L.; Pickard, D.J.; Hill, J.; Harcourt, K.; Parkhill, J.; Dougan, G.; Kingsley, R.A. Signatures of adaptation in human invasive Salmonella Typhimurium ST313 populations from sub-Saharan Africa. PLoS Negl. Trop. Dis. 2015, 9(3), e0003611. [Google Scholar] [CrossRef]
- Yessimseit, D.T.; Rysbekova, A.K.; Zhumadilova, Z.B.; Abdeliyev, B.Z.; Tukhanova, N.B.; Kassenova, A.K.; Abdrakhmanova, A.K.; Mereke, A.; Agzam, S.D.; Nurpeisova, A.S.; Nissanova, R. Genetic and epigenetic diversity of Salmonella enterica isolates from Kazakhstan from clinical and veterinary sources. bioRxiv 2026, 2026–06. [Google Scholar] [CrossRef]
- Zhang, S.; Li, S.; Gu, W.; den Bakker, H.; Boxrud, D.; Taylor, A.; Roe, C.; Driebe, E.; Engelthaler, D.M.; Allard, M.; Brown, E.; McDermott, P.; Zhao, S.; Bruce, B.B.; Trees, E.; Fields, P.I.; Deng, X. Zoonotic Source Attribution of Salmonella enterica Serotype Typhimurium Using Genomic Surveillance Data, United States. Emerg. Infect. Dis. 2019, 25(1), 82–91. [Google Scholar] [CrossRef]
- Lupolova, N.; Lycett, S.J.; Gally, D.L. A guide to machine learning for bacterial host attribution using genome sequence data. Microb. Genom. 2019, 5(12), e000317. [Google Scholar] [CrossRef]
- Pascoe, B.; Futcher, G.; Pensar, J.; Bayliss, S.C.; Mourkas, E.; Calland, J.K.; Hitchings, M.D.; Joseph, L.A.; Lane, C.G.; Greenlee, T.; Arning, N.; Wilson, D.J.; Jolley, K.A.; Corander, J.; Maiden, M.C.J.; Parker, C.T.; Cooper, K.K.; Rose, E.B.; Hiett, K.; Bruce, B.B.; Sheppard, S.K. Machine learning to attribute the source of Campylobacter infections in the United States: A retrospective analysis of national surveillance data. J. Infect. 2024, 89(5), 106265. [Google Scholar] [CrossRef]
- Tonkin-Hill, G.; MacAlasdair, N.; Ruis, C.; Weimann, A.; Horesh, G.; Lees, J.A.; Gladstone, R.A.; Lo, S.; Beaudoin, C.; Floto, R.A.; Frost, S.D.W.; Corander, J.; Bentley, S.D.; Parkhill, J. Producing polished prokaryotic pangenomes with the Panaroo pipeline. Genome Biol. 2020, 21(1), 180. [Google Scholar] [CrossRef]
- Ratner, A.; De Sa, C.; Wu, S.; Selsam, D.; Ré, C. Data Programming: Creating Large Training Sets, Quickly. Adv. Neural Inf. Process Syst. 2016, 3567–3575. [Google Scholar]
- Muzondiwa, D.; Mutshembele, A.; Pierneef, R.E.; Reva, O.N. Resistance Sniffer: An online tool for prediction of drug resistance patterns of Mycobacterium tuberculosis isolates using next generation sequencing data. Int. J. Med. Microbiol. 2020, 310(2), 151399. [Google Scholar] [CrossRef]
- Silva-Palacios, D.; Ferri, C.; Ramírez-Quintana, M.J. Improving performance of multiclass classification by inducing class hierarchies. Procedia Comput. Sci. 2017, 108, 1692–1701. [Google Scholar] [CrossRef]
- Fang, F.C.; Libby, S.J.; Castor, M.E.; Fung, A.M. Isocitrate lyase (AceA) is required for Salmonella persistence but not for acute lethal infection in mice. Infect. Immun. 2005, 73(4), 2547–2549. [Google Scholar] [CrossRef]
- Lorenz, M.C.; Fink, G.R. Life and death in a macrophage: role of the glyoxylate cycle in virulence. Eukaryot. Cell 2002, 1(5), 657–662. [Google Scholar] [CrossRef]
- McKinney, J.D.; Höner zu Bentrup, K.; Muñoz-Elías, E.J.; Miczak, A.; Chen, B.; Chan, W.T.; Swenson, D.; Sacchettini, J.C.; Jacobs, W.R., Jr.; Russell, D.G. Persistence of Mycobacterium tuberculosis in macrophages and mice requires the glyoxylate shunt enzyme isocitrate lyase. Nature 2000, 406(6797), 735–738. [Google Scholar] [CrossRef]
- Hagins, J.M.; Scoffield, J.A.; Suh, S.J.; Silo-Suh, L. Influence of RpoN on isocitrate lyase activity in Pseudomonas aeruginosa. Microbiology 2010, 156 Pt 4, 1201–1210. [Google Scholar] [CrossRef]
- Rowley, G.; Hensen, D.; Felgate, H.; Arkenberg, A.; Appia-Ayme, C.; Prior, K.; Harrington, C.; Field, S.J.; Butt, J.N.; Baggs, E.; Richardson, D.J. Resolving the contributions of the membrane-bound and periplasmic nitrate reductase systems to nitric oxide and nitrous oxide production in Salmonella enterica serovar Typhimurium. Biochem J. 2012, 441(2), 755–762. [Google Scholar] [CrossRef]
- Li, W.; Li, L.; Yan, X.; Wu, P.; Zhang, T.; Fan, Y.; Ma, S.; Wang, X.; Jiang, L. Nitrate Utilization Promotes Systemic Infection of Salmonella Typhimurium in Mice. Int. J. Mol. Sci. 2022, 23(13), 7220. [Google Scholar] [CrossRef]
- Kalai Chelvam, K.; Yap, K.P.; Chai, L.C.; Thong, K.L. Variable Responses to Carbon Utilization between Planktonic and Biofilm Cells of a Human Carrier Strain of Salmonella enterica Serovar Typhi. PLoS ONE 2015, 10(5), e0126207. [Google Scholar] [CrossRef]
- Wheeler, K.M.; Gold, M.A.; Stevens, C.A.; Tedin, K.; Wood, A.M.; Uzun, D.; Cárcamo-Oyarce, G.; Turner, B.S.; Fulde, M.; Song, J.; Kramer, J.R.; Ribbeck, K. Mucus-derived glycans are inhibitory signals for Salmonella Typhimurium SPI-1-mediated invasion. Cell Rep. 2025, 44(10), 116304. [Google Scholar] [CrossRef]
- Dos Santos, A.M.P.; Ferrari, R.G.; Conte-Junior, C.A. Type three secretion system in Salmonella Typhimurium: the key to infection. Genes Genom. 2020, 42(5), 495–506. [Google Scholar] [CrossRef]
Figure 1.
A. hitecture of the final four-level hierarchical classification network for predicting the ecological associations of S. enterica isolates. Squares represent internal decision nodes, and their colors indicate the supervised machine-learning algorithm selected at each node: k-nearest neighbors (KNN), logistic regression (LR), multilayer perceptron (MLP), categorical Naive Bayes (NBayes), or Random Forest (RF). Black squares denote nodes without a fitted supervised model. Classification at these nodes was based on allele-profile similarity or direct assignment when only one child node was available. Numerical values indicate training accuracy, defined as the proportion of training genomes assigned to the correct immediate child node. These values describe model fit and should not be interpreted as estimates of performance on independent genomes. Colored circles represent terminal ecological categories.
Figure 1.
A. hitecture of the final four-level hierarchical classification network for predicting the ecological associations of S. enterica isolates. Squares represent internal decision nodes, and their colors indicate the supervised machine-learning algorithm selected at each node: k-nearest neighbors (KNN), logistic regression (LR), multilayer perceptron (MLP), categorical Naive Bayes (NBayes), or Random Forest (RF). Black squares denote nodes without a fitted supervised model. Classification at these nodes was based on allele-profile similarity or direct assignment when only one child node was available. Numerical values indicate training accuracy, defined as the proportion of training genomes assigned to the correct immediate child node. These values describe model fit and should not be interpreted as estimates of performance on independent genomes. Colored circles represent terminal ecological categories.

Table 1.
Ontology-based grouping of S. enterica strains into two-level ecological categories by associated metadata keywords.
Table 1.
Ontology-based grouping of S. enterica strains into two-level ecological categories by associated metadata keywords.
| First level category | Second level category | Associated metadata keywords |
|---|---|---|
| Human clinical | Blood and tissues | blood; blood culture; cerebrospinal fluid; synovial fluid; peritoneal effusion; pericardic fluid; pleura; biopsy; human tissue; knee tissue; subcutaneous tissue |
| Gastrointestinal and urogenital | stool; stools; feces specimen; fecal swab; rectal swab; perianal screening culture; urine | |
| Wounds and abscesses | abscess; pus; wound; wound swab; draining tract; aspirate | |
| Livestock and Other animals | Cattle | animal-cattle-beef cow; animal-cattle-dairy cow; animal-cattle-heifer; animal-cattle-steer; animal-cattle-bull; animal-cattle-beef cow (cecal); animal-cattle-dairy cow (cecal); animal-cattle-heifer (cecal); animal-cattle-steer (cecal); bos taurus; bovine; fecal (bos taurus) |
| Swine | animal-swine-market swine (cecal); animal-swine-sow (cecal); swine; pig; pig carcass; pig feces; rectal swab from pig | |
| Chicken and poultry birds | animal-chicken-heavy fowl; animal-chicken-young chicken (cecal); animal-turkey-young turkey (cecal); animal-turkey-turkey carcass sponge; chick; chicken; broiler farm; poultry; poultry internal organs; poultry processing plant; poultry shed; duck; duck embryo egg; turkey; hatchery; lairage; intestinal contents of broiler; environmental swab of broiler flock | |
| Companion animals | dog; canine; cat; domestic shorthair cat- abscess in the cat lung | |
| Wildlife and other animals | bird nest; cull; foal; frilled lizard; lamb; pigeon; wild boar; wildlife | |
| Animal organs and body sites | anal swab; bone swab; brain swab; cecum contents; cloacal swab; colon; gi tract; gut; heart; heart surface; homogenized tissue; intestinal contents; intestinal tract; intestine; joint swab; kidney; liver; liver & gallbladder; lung; lymph node; pericardic fluid; peritoneal effusion; pleura; rectum; spleen; spleen surface; stomach contents; tonsils; urachus swab | |
| Food chain | General food and retail | food; food sample; miscellaneous foods; meat; meat retail; retail chicken meat; food: retail chicken meat; raw intact chicken; nonintact chicken; rte product; rtc boneless skinless chicken breast with rib meat; feed |
| Beef products | beef; comminuted beef; ground beef; product-raw-intact-beef; milk; milk: raw bovine | |
| Pork products | pork; pork chop; pork meat; pork (cecal content); product-raw-ground, comminuted or otherwise nonintact-pork; product-raw-intact-pork | |
| Chicken products | chicken - young chicken carcass rinse (post-chill); chicken breast; chicken breasts; chicken carcass; chicken chicken gizzard; chicken chicken liver; chicken drumsticks; chicken from wet market; chicken gizzard; chicken gizzard (giblets); chicken gizzards; chicken gut; chicken heart; chicken hearts (giblets); chicken liver; chicken livers; chicken meat; chicken thighs; comminuted chicken; ground chicken; mechanically separated chicken; organic drumett; product-rinse-chicken; spicey chicken breast; product-eggs-dried-yellow egg products; raw eggs; retail shell egg | |
| Turkey and duck products | comminuted turkey; duck meat; turkey fresh mst; turkey ground turkey | |
| Seafood and aquatic foods | aquatic products; cooked mussels; fresh water fish; frozen layang scad whole round; frozen raw shrimp; frozen shrimp; hairtail fish; rohu fish; shrimp; silver pomfret fish; tilapia | |
| Plant-derived foods | baking flour; basil; black pepper; cold dish; coriander; cucumber; fermented soybean; holy basil; jujube; leafy greens; moringa; oat; peanuts; pistachio; pistachio kernel; pistachio paste; raw almond; sesame seeds; sweet basil; tahini; tahini sesame seed paste | |
| Mixed, processed, and specialty foods | apple sausage; fragrance product; greens and vegetable blend dietary supplement with moringa powder; insect based food (crickets); jelly; kangaroo tail dog chew; kratom; nadiyadi mix; pet food dog; sausage; spice mix for chicken; spices | |
| Environment | Environmental swabs and sponges | carcass; drag (farm); drag swab; environmental swab; environmental swab sponge; swab |
| Water and wastewater | filter from dairy farm; irrigation water; reservoir; river water; sewage; wastewater; wastewater influent; water | |
| Sediment and waste | faces; faecal samples; feces; floor litter; livestock manure; sediment |
Table 2.
Results of classification of S. enterica strains by program Salmonella_classifier.
| Query sequence | Classifier predictions |
|---|---|
| 19S_chr.gbk | Livestock and Other animals | Chicken and poultry birds, 0.9666; Food chain | Pork products, 0.9666; Food chain | Chicken products, 0.9666 ; Food chain | General food and retail, 0.9666; Environment | Environmental swabs and sponges, 0.9666; |
| 1S_chr.gbk | Environment | Sediment and waste, 0.9520; Human clinical | Blood and tissues, 0.9520; Human clinical | Gastrointestinal and urogenital, 0.9520; Livestock and Other animals | Chicken and poultry birds, 0.9520; Food chain | General food and retail, 0.9520; Food chain | Chicken products, 0.9520; |
Table 3.
Major housekeeping, central metabolism and virulence-associated genes affected by habitat-specific adaptive mutations in S. enterica genomes.
Table 3.
Major housekeeping, central metabolism and virulence-associated genes affected by habitat-specific adaptive mutations in S. enterica genomes.
| Functional/metabolic theme | Representative proteins | Possible ecological significance |
|---|---|---|
| Glyoxylate cycle and TCA | aceA, aceK, icd, iclR | Controls partitioning of isocitrate between the TCA cycle and glyoxylate bypass; important when growing on acetate/fatty acids and during nutrient limitation. |
| Proline / glutamate metabolism | putA, putP, gdhA, aspC, astC, astE, ansA | Nitrogen utilization, amino-acid catabolism, osmotic adaptation and alternative carbon/nitrogen sources. |
| Carbohydrate uptake and utilization | pfkB, ppsA, manX, srlA, chbC, nagK, amyA, 6-phospho-beta-glucosidase, D-hexose-6-P mutarotase, gutQ, gutM, kdgA | Carbohydrate utilization and metabolism. |
| Fucose metabolism | fucK | Utilization of host/mucosal carbohydrates; potentially relevant to intestinal ecology |
| Fatty-acid / acetyl-CoA metabolism | fadI, acetyl-CoA C-acetyltransferase, fabD, cfa | Membrane composition, β-oxidation and adaptation to environmental/stress conditions |
| Anaerobic respiration / nitrate respiration | narH, nitrate reductase α, narK, narX | Survival in oxygen imitated environments. |
| Hydrogen metabolism | hyaB, hyaD, hypE | Hydrogen utilization, production and anaerobic energy metabolism |
| Electron transfer / redox metabolism | rsxG, cytochrome b, NAD(P)H nitroreductase, nemA, oxidoreductases | Redox balancing and adaptation to alternative electron donors/acceptors |
| Pentose phosphate pathway | glucose-6-phosphate dehydrogenase | NADPH production, oxidative-stress resistance and anabolic metabolism |
| NAD metabolism | pncA, pncB, pncC, nadE | NAD biosynthesis and salvage |
| Purine metabolism | purT, add | Purine synthesis and salvage |
| Vitamin/cofactor metabolism | thiK, riboflavin synthase, btuC | Thiamine, riboflavin and vitamin B12 acquisition and metabolism |
| Aromatic amino-acid biosynthesis | aroH | Entry into the shikimate pathway, aromatic amino-acid biosynthesis |
| 4-hydroxyphenylacetate degradation | hpaG, hpaE, hpaX | Utilization of aromatic compounds as carbon sources |
| Sulfur / Fe-S metabolism | sufA, sufE, tusE, DsrE/DsrF/TusD-family protein | Fe-S cluster maintenance and sulfur transfer, oxidative/metal stress response |
| Selenium metabolism | selD | Selenoprotein biosynthesis and redox metabolism |
| Metal homeostasis | zinT, copper-resistance protein, heavy-metal sensor kinase | Adaptation to metal availability and toxicity |
| Ion/pH/osmotic homeostasis | chaA, chaB, mechanosensitive channel | Osmotic, ionic and pH adaptation |
| Type III secretion | ssaJ, ssaI, ssaU, ssaE, sseB, sseC, pipB, sopB, sipB, spaM, sigE | Virulence |
| Chemotaxis and flagellar motility | tar, cheB, flhA, flhE, flgA, flgE, fliI | Response to environmental stimuli |
| Top-level transcriptional regulators | PhoQ, BarA, NarX, HprR, SlyA, IclR, GutM, MerR-family regulators, heavy-metal sensor kinase, RelA | Stress response and general metabolism regulation |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.