Preprint
Article

This version is not peer-reviewed.

Serralysin Marker States Across Pseudomonas: The AprX SVMSY Pentapeptide Is Shared with Pseudomonas aeruginosa AprA and Tracks Lineage Rather than Proteolytic Phenotype

Submitted:

21 September 2026

Posted:

22 September 2026

You are already at the latest version

Abstract
Psychrotrophic Pseudomonas spp. contribute to proteolytic spoilage of milk through heat-stable extracellular proteases, principally the serralysin-family metalloprotease AprX, and there is sustained interest in sequence features that could serve as molecular markers of spoilage risk. This study reports a reference-anchored catalogue of Pseudomonas serralysins and uses it to evaluate the diagnostic behaviour of a recently proposed five-residue AprX marker, SVMSY. Three cohorts were analysed. In 433 annotated RefSeq Pseudomonas type-material genomes, 231 full-length serralysin candidates were identified and classified by phylogenetic proximity to curated AprX and Pseudomonas aeruginosa AprA references (162 AprX-proximal, 69 AprA-proximal). Zinc-anchored catalytic-region mapping reproduced all 231 five-residue assignments obtained by global alignment. Exact SVMSY occurred in 90/162 (55.6%) AprX-proximal but 65/69 (94.2%) AprA-proximal proteins (odds ratio 13.0, 95% confidence interval 4.52-37.38, Fisher p = 8.1 x 10-10), so the pentapeptide is neither universal among AprX-proximal serralysins nor specific to them. Superposition of all six experimentally determined serralysin structures showed why: the region is the metzincin Met-turn, occupying residues 212-216 with the methionine sulfur 6.15 +/- 0.05 A from the catalytic zinc in every structure, and across 2,611 serralysins positions 3-5 were invariant (Met 99.96%, Ser 100%, Tyr 100%) while only position 2 varied. The proposed marker is therefore a fixed element of the catalytic scaffold carrying one variable position, and that position is fully buried (relative solvent accessibility 0.0% in all six structures) 9.4-9.7 A from the zinc. An independent strain-level cohort of 2,611 RefSeq Pseudomonas serralysins, not restricted to one genome per species, replicated this contrast closely (53.1% versus 94.7%; odds ratio 15.71, 95% confidence interval 9.43-26.16, p = 1.6 x 10-52) and quantified its consequences: of 187 species carrying AprX-proximal serralysins, an exact-SVMSY probe would fully detect 102, partially miss 40 and entirely miss 45. Forty-two species carried more than one marker state, and assembly-level resolution through Identical Protein Groups confirmed strain-level SVMSY/SLMSY variation within Pseudomonas mandelii, Pseudomonas poae and Pseudomonas koreensis, identifying closely related strains suitable for a controlled test of the marker. In an 87-genome phenotype-linked cohort, the marker was lineage-locked, was not associated with day-7 proteolysis in unadjusted analysis (mean difference -0.52, p = 0.343) or under phylogenetic regularization (beta = 0.397, p = 0.397), and a within-species residue scan yielded no false-discovery-rate-significant AprX positions; sequence distance to the AprX reference was the stronger correlate of phenotype (Spearman rho = -0.640, p = 5.4 x 10-9) but also attenuated under phylogenetic regularization (p = 0.095). Homology-based reconstruction of the aprX-lipA2 region recovered PrtA-like and PrtB-like proteins that annotation-string searches missed entirely, and showed a large phenotype difference between local-PrtAB and distal-PrtAB architectures that was nonetheless inseparable from species background. Together these results define where a pentapeptide-level AprX assay will and will not work, and identify the strain panels needed to separate marker effects from lineage.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

Psychrotrophic Pseudomonas species are a persistent challenge in dairy processing because refrigerated raw-milk storage selects for psychrotolerant pseudomonads that secrete extracellular proteases and lipases, many of which remain active after heat processing (De Jonghe et al., 2011; Marchand et al., 2009; Zhang et al., 2019). AprX is of particular importance because residual proteolytic activity can destabilize casein micelles during storage of pasteurized or ultra-high-temperature (UHT) milk, contributing to sedimentation, age gelation, bitterness and shortened shelf life (Bagliniere et al., 2013; Zhang et al., 2019; Sinha and Kelly, 2026). AprX is a serralysin-family zinc metalloprotease, commonly encoded within the aprX-lipA2 region together with an AprI-family inhibitor, the AprDEF type-I secretion system, variable autotransporter-like proteins and lipase genes (Woods et al., 2001; Maier et al., 2020).
The presence of Pseudomonas, or even of an aprX-like gene, does not translate into a predictable spoilage phenotype. Marked strain-to-strain differences are reported in caseinolytic activity, UHT-milk destabilization and AprX-associated spoilage potential (Marchand et al., 2009; Bagliniere et al., 2012; Maier et al., 2020; Aguilera-Toro et al., 2023). This has motivated a search for genomic features that distinguish high-risk strains before extensive enzyme production or visible product failure, and molecular assays directed at aprX are an active area of method development in which specificity is the central requirement (Sinha and Kelly, 2026).
Designing such an assay requires knowing how the target sequence is distributed across the genus. Serralysin nomenclature in Pseudomonas is historically inconsistent, particularly the AprX/AprA distinction, and the same catalytic-domain sequence may occur in proteins of quite different provenance. A marker that is shared with the alkaline protease AprA of the opportunistic pathogen Pseudomonas aeruginosa, for example, carries a different interpretive risk from one confined to dairy-associated AprX. Comprehensive, reference-anchored descriptions of serralysin sequence space in Pseudomonas are not currently available, and this gap is what the present catalogue is intended to fill.
A further consideration is statistical. Bacterial populations are strongly structured and extensive linkage disequilibrium can make many variants track the same clonal lineage, so a residue may appear associated with a phenotype in pooled data even when it has little independent explanatory value; the same applies to gene-cluster organization when a particular architecture occurs predominantly within one clade (Earle et al., 2016; Lees et al., 2016). Pooled association tests therefore overstate the generality of a candidate biomarker unless lineage structure and the number of independent evolutionary transitions are examined explicitly.
A recent study integrating whole-genome analysis, cold-storage phenotyping and molecular docking proposed a five-residue AprX sequence, SVMSY, as a feature consistently present in high-spoilage strains, suggested that a valine-to-leucine substitution at the second position could alter substrate-pocket interactions, and put the sequence forward as a theoretical target for molecular detection of high-risk isolates (Liang et al., 2026). That hypothesis provides a well-defined test case for the catalogue reported here: a pentapeptide intended for general spoilage-risk detection must retain predictive value across separately sampled strain collections and must carry information that is not simply a proxy for phylogenetic lineage.
This study therefore had three objectives. The first was to build a reference-anchored catalogue of full-length serralysins across the annotated Pseudomonas type-material genomes, with independent verification of catalytic-region positional homology. The second was to determine, from that catalogue and from an independent strain-level cohort, how an exact-pentapeptide assay would behave across the genus, species by species. The third was to test, in a separately published phenotype-linked cohort, whether marker state or locus architecture carries information about proteolytic phenotype that is separable from lineage. The objective was not to test whether a SVMSY-to-SLMSY substitution can influence AprX biochemistry in a defined protein background, which remains an open and reasonable mechanistic question.

2. Materials and Methods

2.1. Study Design and Cohorts

Three cohorts were analysed. The type-material cohort was assembled from currently annotated RefSeq Pseudomonas genomes flagged by NCBI as type material, selecting one assembly per organism name and excluding atypical assemblies and metagenome-assembled genomes, yielding 433 genomes. Because this collection is taxonomically ascertained rather than environmentally sampled, it was used to characterize sequence diversity and taxonomic distribution, not the environmental prevalence of AprX or the frequency of milk-spoilage phenotypes. One consequence of the one-genome-per-species design is that the cohort cannot, by construction, contain two strains of the same species, so it cannot be used to assess within-species marker variation.
The strain-level cohort was assembled for that purpose and is independent of the type-material cohort in its sampling design. All RefSeq Pseudomonas protein records annotated as serralysin-family metalloproteases, alkaline metalloproteinases, AprX, AprA or M10-family metallopeptidases were retrieved through NCBI E-utilities (3,372 records), without restriction on the number of genomes per species. Because RefSeq WP_ accessions are non-redundant across strains, protein-record counts are not strain counts; strain and assembly attribution was recovered separately through NCBI Identical Protein Groups (Sayers et al., 2022) as described in Section 2.6.
The phenotype-linked cohort comprised 87 Pseudomonas genome assemblies associated with BioProject PRJNA612806 and the milk-proteolysis study of Maier et al. (2020). RefSeq assemblies were used where duplicate GenBank/RefSeq representations of the same isolate were present. Published categorical proteolysis information was extracted from the original study figure by automated row matching and colour classification, followed by visual quality control of the extracted assignments; this constraint and its implications are stated in Section 4.6. Only successfully resolved observations were retained and no missing phenotype values were imputed. Day-7 phenotype categories were mapped to an ordinal score of 0-4 (non-proteolytic, weak, moderate, strong, very strong).

2.2. Identification of Serralysin-Family Proteins

The AprX protein of Pseudomonas fluorescens CY091 (AAC38255.2) was used as the primary homology query because CY091 is a well-characterized AprX reference for which extracellular protease activity, heat stability and the corresponding aprX sequence were established experimentally (Liao and McCallus, 1998). Proteomes of the phenotype-linked and type-material cohorts were screened with BLASTP using BLAST+ (Camacho et al., 2009). Candidates were evaluated for query coverage, protein length and the serralysin/metzincin catalytic zinc-binding motif HEXXHXXGXXH (Bode et al., 1993).
Conservative full-length criteria were applied to the type-material cohort: query coverage at least 90%, E-value at most 1e-50, protein length 430-550 amino acids and presence of the catalytic motif. The coverage threshold excludes partial proteins and the length window encompasses the well-characterized approximately 50-kDa/477-residue AprX homologues while excluding substantially larger or truncated M12 metallopeptidases (Liao and McCallus, 1998; Zhang et al., 2009; Bode et al., 1993). These criteria identified 231 candidates. Genomes without a protein-level hit were additionally checked by TBLASTN so that annotation failure alone would not be recorded as absence. The same length and catalytic-motif criteria were applied to the strain-level cohort, retaining 2,611 of 3,372 retrieved records. Throughout, AprX-reference-like and AprX-reference-proximal are used as descriptive sequence and phylogenetic terms rather than as gene-name assignments, because serralysin nomenclature in Pseudomonas is historically inconsistent.

2.3. Reference Panel and Serralysin Phylogeny

A curated serralysin reference panel contained AprX sequences from strains CY091, A506, TSS and B52 together with Pseudomonas aeruginosa AprA references from PAO1 and PA14. Candidate and reference proteins were aligned with MAFFT v7 (Katoh and Standley, 2013) and maximum-likelihood trees inferred with IQ-TREE 2 (Minh et al., 2020) using ModelFinder for substitution-model selection (Kalyaanamoorthy et al., 2017) with 1,000 ultrafast bootstrap and 1,000 SH-like approximate likelihood-ratio replicates. Type-material candidates were classified as AprX-reference-proximal or Pseudomonas aeruginosa-AprA-reference-proximal by minimum patristic distance to the two reference groups. For the strain-level cohort, where the number of sequences made per-cohort tree inference unnecessary for this purpose, the same two-group assignment was made by highest global pairwise identity (BLOSUM62, gap open -11, gap extend -1) to the nearest member of each reference group. Reference-proximity labels are comparative descriptors and are not definitive functional gene annotations.

2.4. Core-Genome Phylogeny of the Phenotype-Linked Cohort

To model bacterial lineage independently of aprX itself, a core-protein phylogeny was constructed for the 87-genome cohort. Marker homologues were identified with HMMER (Eddy, 2011) using the published UBCG bacterial marker hidden Markov model set as a marker source (Na et al., 2018) within a custom alignment and concatenation workflow, individually aligned with MAFFT, filtered for completeness and concatenated. Seventy-four retained core markers produced a concatenated amino-acid alignment of 17,789 positions, and the maximum-likelihood core tree was inferred with IQ-TREE 2. All 87 genomes were represented in the final tree.

2.5. Mapping of the Five-Residue Marker

The five-residue region corresponding to the proposed SVMSY marker was mapped using an exact-SVMSY anchor sequence (GCF_012985445.1 | WP_170037105.1). Two independent procedures were used. In the first, candidates were aligned globally and the marker position transferred through the alignment column. In the second, positional homology was established directly from protein structure-defining sequence: the conserved M10 zinc-binding motif HEXXHXXGXXH was located in the unaligned sequence, the surrounding catalytic region extracted, independently aligned to the corresponding anchor region, and the anchor five-residue position transferred through that local alignment. In the anchor, CY091 AprX and PAO1 AprA the marker begins exactly 36 residues after the first histidine of the zinc motif. Agreement between the two procedures was recorded before any downstream frequency or association analysis.

2.6. Strain-Level Cohort and Within-Species Analysis

Species attribution for the strain-level cohort was taken from the RefSeq protein definition line. Records labelled MULTISPECIES (a protein observed in more than one named species) and records attributed only to Pseudomonas sp. were excluded from per-species tabulation and are reported separately. For six species of dairy relevance or of demonstrated marker variability, protein records were resolved to genome assemblies through NCBI Identical Protein Groups (Sayers et al., 2022), allowing three distinct situations to be separated: species in which all assemblies carry the same marker state; species in which different assemblies carry different states, which constitutes genuine strain-level variation; and individual assemblies carrying two serralysin paralogues with different states, for which a genome-level marker state is not well defined.
Per-species outcomes for an assay targeting the exact SVMSY pentapeptide were classified as fully detected (all AprX-proximal serralysins of that species carry SVMSY), partially missed (some do) or entirely missed (none do). The classification describes the sequence-level coverage of an exact-match target only; the sensitivity and specificity of any realized assay also depend on probe chemistry, target region and amplification design.

2.7. Sequence-Phenotype Analysis

Associations with day-7 phenotype in the phenotype-linked cohort were examined using pooled non-parametric comparisons and a high-versus-low dichotomization. To identify associations not driven solely by between-species differences, residue scans were repeated after restricting to alignment positions varying within at least one species, using species-centred and permutation logic, with false-discovery-rate correction across tested sites. A phylogenetically regularized generalized least-squares model with a Pagel-type lambda estimated by likelihood, using a normalized phylogenetic correlation matrix derived from the core-genome tree, was fitted for the ordinal day-7 score. Because the phenotype is ordinal rather than continuous, this model is reported as a sensitivity analysis rather than as the sole inferential test. Spearman correlation was used for the relationship between phenotype and patristic distance to the nearest AprX reference.
Marker frequencies in the type-material and strain-level cohorts were compared between reference-proximity classes by Fisher exact test, with an approximate 95% confidence interval for the odds ratio calculated on the log-odds scale. The minimum number of marker-state changes on the serralysin phylogeny was estimated by Fitch parsimony as a lower bound on evolutionary transitions.

2.8. Homology-Based Reconstruction of the aprX-lipA2 Region

Operon components were reconstructed by protein homology rather than by annotation string, because annotation vocabularies do not reliably label the variable autotransporter-like proteins of this region. Nine reference proteins corresponding to the aprX-lipA2 system (GenBank AGL85002.1-AGL85010.1), consistent with the operon framework analysed by Maier et al. (2020), were screened against all proteomes. High-confidence calls required at least 60% query coverage, at least 25% amino-acid identity, E-value at most 1e-20 and a reciprocal best match to the corresponding reference within the nine-protein panel. Gene coordinates were taken from GFF annotations and mapped relative to the serralysin locus.
AprI is short, and generic E-value thresholds calibrated on much larger proteins under-detect it; a gene-specific rule was therefore applied for AprI (query coverage at least 70%, identity at least 25%, E-value at most 1e-10, and either reciprocal-reference support or an AprI/Inh-family annotation). Local components were defined as occurring on the same assembly contig and within 50 kb of the serralysin anchor. Type-1-like architecture was defined as a local aprX/aprI/aprDEF region with local PrtA-like, PrtB-like and LipA2 homologues; type-8-like architecture allowed PrtA-like and PrtB-like homologues to occur distally, including on another contig, while aprX/aprI/aprDEF/LipA2 remained local. These labels describe reconstructed relative architecture rather than exact reproductions of every published operon subtype.

2.9. Structural Superposition of the Catalytic Domain

To establish where the five-residue region lies in the folded enzyme, all six experimentally determined serralysin structures available in the Protein Data Bank were analysed: the Pseudomonas aeruginosa alkaline protease structures 1KAP (Baumann et al., 1993) and 1AKL (Miyatake et al., 1995), and the Serratia serralysin structures 1SAT (Baumann, 1994), 1SMP (Baumann et al., 1995), 1SRP (Hamada et al., 1996) and 1AF0. Chains were superposed onto 1KAP by combinatorial extension as implemented in Biopython (Cock et al., 2009), which uses backbone geometry rather than a sequence alignment, so that positional correspondence is not imposed by the same alignment logic used elsewhere in this study. In each structure the catalytic zinc-binding motif HEXXHXXGXXH was located in the chain sequence and the five-residue region identified at the offset used throughout.
Three distances were measured in the common reference frame: from the catalytic zinc ion to the Met-turn methionine S-delta atom, to the position-2 C-beta atom, and to the most distal atom of the position-2 side chain. Relative solvent accessibility of the position-2 residue was computed by the Shrake-Rupley algorithm as implemented in Biopython and expressed as a percentage of that residue type maximum accessible surface area. Local backbone conservation was quantified as the root-mean-square deviation of the five C-alpha atoms of the region after global superposition, for comparison with the global superposition root-mean-square deviation of the same pair of chains.

2.10. Reproducibility

All sequence retrieval, candidate identification, alignment, phylogenetic analysis, marker mapping, homology screening, statistical analysis and figure generation were scripted in Python and shell workflows, with intermediate tables retained at each stage. Supplementary Tables S1-S14 report cohort composition, the principal statistical analyses, the external reference-proximity results, the homology-based operon reconstruction, the per-species probe outcomes, the assembly-level within-species resolution, the per-structure geometry of the five-residue region and its position-wise residue composition. They are supplied as a single spreadsheet workbook, AprX_supplementary_tables_S1-S14.xlsx, with a contents sheet, and the staged scripts that generated them are in the repository given under Data and code availability.

3. Results

3.1. A Reference-Anchored Serralysin Catalogue Across 433 Pseudomonas Type Genomes

Of 433 annotated RefSeq Pseudomonas type-material genomes, 231 contained a full-length serralysin candidate satisfying the coverage, length, catalytic-motif and E-value criteria. The candidate phylogeny, containing these 231 proteins and six curated references (237 tips), separated candidates by relative proximity to the AprX and Pseudomonas aeruginosa AprA reference groups, placing 162 as AprX-reference-proximal and 69 as AprA-reference-proximal (Figure 1).
All 231 candidates yielded an unambiguous five-residue marker state, and the independent zinc-anchored catalytic-region mapping reproduced every assignment obtained by global alignment: each mapped to five contiguous residues and no binary SVMSY classification changed. The external calls are therefore not an artifact of transferring a particular set of global alignment columns from a single anchor.

3.2. The SVMSY Pentapeptide Is Neither Universal Among AprX-proximal Serralysins nor Specific to Them

Across the 231 candidates, SVMSY was the most common state (155), followed by SLMSY (60), SIMSY (14) and TVMSY (2). The distribution differed sharply by reference-proximity class: SVMSY occurred in 90/162 (55.6%) AprX-proximal candidates but in 65/69 (94.2%) AprA-proximal candidates (Figure 2A). The Fisher exact comparison gave an odds ratio of 13.0 for SVMSY in the AprA-proximal relative to the AprX-proximal class (approximate 95% confidence interval 4.52-37.38; two-sided p = 8.1 x 10-10). The four-state marker required at least 20 changes on the serralysin phylogeny, indicating repeated evolution or replacement of the motif across the sequence space sampled.
This is the central sequence-level result of the catalogue. An assay targeted only to the exact pentapeptide would be expected to generate false negatives for AprX-proximal lineages that do not carry SVMSY and false positives for non-AprX serralysins that do, with the balance depending on assay design and on the taxonomic composition of the sample.

3.3. Homology-Based Reconstruction Recovers Locus Components That Annotation Strings Miss

Among the 231 candidates, AprD, AprE and AprF homologues occurred within 50 kb of the serralysin locus in 214 (92.6%), 216 (93.5%) and 215 (93.1%) genomes respectively, and AprI in 186 (80.5%). The variable autotransporter-like and lipase components were less frequently local: PrtA-like 103 (44.6%), PrtB-like 117 (50.6%), LipA1 59 (25.5%) and LipA2 158 (68.4%) (Figure 2B). The phenotype-linked serralysin-positive cohort showed similar conservation of AprDEF with a higher frequency of local LipA2 (66/76, 86.8%).
These values differ substantially from an annotation-string screen of the same 231 genomes, which recovered no PrtA/PrtB-like proteins at all and over-counted AprI by matching generic protease-inhibitor annotations (Supplementary Table S4, retained for comparison and superseded by Supplementary Tables S5 and S6). The difference is methodological rather than biological: annotation vocabularies for autotransporter-like proteins are inconsistent, and homology screening recovers proteins that string matching does not. All locus-context values reported in this study derive from the homology-based reconstruction.
Figure 2. Marker distribution and locus context across the 231 full-length serralysin candidates of the type-material cohort. (A) Distribution of SVMSY, SLMSY, SIMSY and TVMSY in AprX-reference-proximal and Pseudomonas aeruginosa-AprA-reference-proximal candidates. (B) Frequency of homology-supported AprI, AprD, AprE, AprF, PrtA-like, PrtB-like, LipA1 and LipA2 homologues within 50 kb of the serralysin anchor. Annotation-string equivalents of panel B are given in Supplementary Table S4 and are superseded by these values.
Figure 2. Marker distribution and locus context across the 231 full-length serralysin candidates of the type-material cohort. (A) Distribution of SVMSY, SLMSY, SIMSY and TVMSY in AprX-reference-proximal and Pseudomonas aeruginosa-AprA-reference-proximal candidates. (B) Frequency of homology-supported AprI, AprD, AprE, AprF, PrtA-like, PrtB-like, LipA1 and LipA2 homologues within 50 kb of the serralysin anchor. Annotation-string equivalents of panel B are given in Supplementary Table S4 and are superseded by these values.
Preprints 234484 g002

3.4. The Five-Residue Region Is the Metzincin Met-Turn and Occupies One Position Beneath the Catalytic Zinc

The sequence results raise a structural question: does the five-residue region occupy the same physical location in the serralysin fold, and where is it relative to the catalytic zinc? All six experimentally determined serralysin structures in the Protein Data Bank were therefore superposed. Two are Pseudomonas aeruginosa alkaline protease (1KAP, 1AKL; identical sequences) and four are Serratia serralysins (1SAT, 1SMP, 1SRP, 1AF0); the two groups share 57-59% amino-acid identity, and their five-residue states are SVMSY and SLMSY respectively, so the comparison spans both of the common marker states.
The region is the metzincin Met-turn. In every structure it occupies residues 212-216, with the invariant methionine at position 214 and the three zinc-coordinating histidines at 176, 180 and 186, in identical numbering across both groups. After superposition the zinc ions coincided to within 0.41-0.53 A, and the methionine S-delta atom lay 6.10-6.22 A from the catalytic zinc in all six structures (mean 6.15 A, standard deviation 0.053 A). The five C-alpha atoms of the region deviated from their 1KAP positions by only 0.37-0.60 A, less than the global superposition deviation of the same chains (1.19-2.68 A), so the region is more strongly conserved in position than the catalytic domain that carries it (Figure 3A, 3B; Table 2). The Met-turn is a defining feature of the metzincin clan, in which the methionine packs beneath the three zinc-binding histidines and provides a hydrophobic base for the catalytic site (Bode et al., 1993); the five-residue region proposed as a spoilage marker is that structural element.
This explains the sequence pattern. Across the 2,611 serralysins of the strain-level cohort, position 3 was methionine in 2,610 (99.96%), position 4 serine in 2,611 (100%) and position 5 tyrosine in 2,611 (100%), and position 1 was serine in 2,579 (98.8%). Only position 2 varied appreciably: valine in 1,519 (58.2%), leucine in 962 (36.8%), isoleucine in 107 (4.1%) and seven other residues in the remainder (Figure 3C). Four of the five residues of the proposed marker are therefore fixed by the requirements of the catalytic scaffold and carry no discriminating information at all; the SVMSY-versus-SLMSY distinction is a single conservative substitution at position 2 of the Met-turn.
Where that position sits also bears on mechanism. The position-2 side chain is directed away from the catalytic cleft and into the protein interior: its relative solvent accessibility was 0.0% in all six structures, its C-beta atom lay 9.35-9.66 A from the zinc, and its most distal side-chain atom 9.06-9.15 A away in the valine-bearing structures and 10.77-10.94 A in the leucine-bearing ones, the additional methylene extending further from the zinc rather than toward it. In these structures the residue makes no contact with solvent and is too distant to participate directly in substrate binding at the zinc; any catalytic consequence of a valine-to-leucine substitution would have to act indirectly, through packing beneath the active site. That is a testable prediction, and it is a different claim from direct modification of a substrate pocket.

3.5. An Independent Strain-Level Cohort Replicates the Specificity Failure and Resolves It Species by Species

Because the type-material cohort contains one genome per species, it cannot show how a marker behaves within a species or across the many sequenced strains of the species that matter in dairy contexts. An independent strain-level cohort was therefore assembled from RefSeq without a per-species restriction. Of 3,372 retrieved Pseudomonas serralysin-family protein records, 2,611 satisfied the length and catalytic-motif criteria; 2,311 were AprX-proximal and 300 AprA-proximal. Marker states were assigned by the zinc-anchored local-alignment procedure, which resolved all 2,611 records to a five-residue state; the simpler fixed-offset transfer agreed with it for 96.4% of records, the discrepancies arising where an insertion or deletion between the zinc motif and the marker shifts the reading frame of a fixed offset.
The specificity contrast replicated closely in this independently sampled cohort: SVMSY occurred in 1,226/2,311 (53.1%) AprX-proximal and 284/300 (94.7%) AprA-proximal serralysins (odds ratio 15.71, approximate 95% confidence interval 9.43-26.16; two-sided Fisher p = 1.6 x 10-52; Table 3). That two cohorts assembled under different sampling designs give AprA-proximal SVMSY frequencies of 94.2% and 94.7% indicates the near-fixation of this pentapeptide in the Pseudomonas aeruginosa AprA group is a stable property of the sequence space rather than an artifact of type-material ascertainment.
Of the 2,611 records, 1,275 carried an unambiguous binomial species attribution; 623 were labelled MULTISPECIES, that is, the identical protein is found in more than one named species, and 713 were attributed only to Pseudomonas sp. Among the species-resolved records, 187 species carried at least one AprX-proximal serralysin. An exact-SVMSY probe would fully detect the AprX-proximal serralysins of 102 of these species, partially miss those of 40, and entirely miss those of 45 (Figure 4A; Supplementary Table S9 and Table S12). Species that would be entirely missed include Pseudomonas iridis, Pseudomonas simiae, Pseudomonas monsensis and Pseudomonas veronii, while Pseudomonas lundensis, Pseudomonas fragi, Pseudomonas viridiflava, Pseudomonas gingeri and Pseudomonas atacamensis were uniformly SVMSY. Pseudomonas fluorescens, the species most often invoked in dairy proteolysis, carried four different states across 94 protein records, with only 42.6% SVMSY.

3.6. Within-Species Marker Variation Identifies Strains for a Controlled Test

Forty-two species carried more than one marker state among their AprX-proximal serralysin records, so the within-species contrast that the phenotype-linked and type-material cohorts cannot provide does exist in the wider RefSeq collection. Because RefSeq WP_ accessions are non-redundant across strains, protein-record counts alone cannot distinguish strain-level allelic variation from paralogy within a single genome, and six species were therefore resolved to assemblies through Identical Protein Groups (Figure 4B; Supplementary Table S10).
Of 102 resolved assemblies, three species showed genuine strain-level variation. Pseudomonas mandelii comprised 10 SLMSY and 2 SVMSY assemblies, Pseudomonas poae 4 SVMSY and 3 SLMSY, and Pseudomonas koreensis 4 SVMSY and 15 SLMSY assemblies together with 7 assemblies carrying two serralysin paralogues of different state. Pseudomonas lundensis (12/12), Pseudomonas fragi (36/36) and Pseudomonas proteolytica (9/9 resolved assemblies) were uniformly SVMSY; a single recent Pseudomonas proteolytica record carrying SLMSY had no assembly attribution in the Identical Protein Groups output and is therefore not counted at assembly level. The seven Pseudomonas koreensis assemblies carrying two paralogues of different state are a distinct problem for assay design, because for those genomes a single genome-level marker state is not defined at all.
These three species provide the design that the present data cannot supply: closely related strains of a single species that differ at the marker position, in which the residue can be varied while lineage is held approximately constant.

3.7. In the Phenotype-Linked Cohort the Marker Is Lineage-Locked and Unassociated with Proteolysis

The 87-genome phenotype-linked cohort contained 74 AprX-reference-like serralysin candidates, two Pseudomonas aeruginosa-AprA-like serralysins, ten genomes with no convincing AprX-like homologue under the applied criteria, and one low-identity M12-family metallopeptidase outlier. Among the 74 AprX-reference-like proteins, 57 carried SVMSY, 16 SLMSY and one SIMSY. These states were not randomly distributed across the core-genome tree but were largely lineage-locked: no species in this cohort contained more than one of the major marker states among the strains available (Figure 5).
In pooled day-7 data, exact SVMSY was not associated with a higher proteolytic score. Among 67 strains with both a resolved marker and a day-7 phenotype, the mean difference for SVMSY versus non-SVMSY was -0.52 score units and the Mann-Whitney comparison was not significant (p = 0.343). Under dichotomization, SVMSY showed no enrichment among high-activity strains (odds ratio 0.32; p = 0.088). The lambda-regularized phylogenetic sensitivity analysis gave beta = 0.397 (standard error 0.466; p = 0.397), again providing no evidence of an independent positive association.

3.8. Naive Residue Associations Collapse After Lineage-Aware Filtering, and Distance to the AprX Reference Is the Stronger Correlate

A pooled alignment-wide residue scan initially returned 161 of 191 tested AprX positions as false-discovery-rate-significant. This was not biologically credible, because most variants were perfectly or near-perfectly linked to species background. Restricting the analysis to sites varying within at least one species left only 11 testable positions, none significant after correction. The proposed five-residue marker could not be tested within species at all, because the cohort contained no within-species marker-state variation; Section 3.6 identifies the species in which such a test is possible.
Two further associations in this cohort are reported for completeness. Genomes recovered from raw milk were more likely to carry an AprX-reference-like serralysin than those from other sources (odds ratio 15.64; p = 4.9 x 10-5). More importantly, patristic distance from each protein to the nearest AprX reference was strongly and negatively correlated with day-7 proteolysis (Spearman rho = -0.640; p = 5.4 x 10-9), a considerably stronger association than anything obtained for the pentapeptide. Under phylogenetic regularization, however, this correlate also attenuated (beta = -2.56, standard error 1.51; p = 0.095). The contrast is instructive rather than contradictory: a continuous, whole-protein measure of similarity to a characterized AprX carries more phenotype information in pooled data than a single five-residue state, and both are substantially confounded with lineage.

3.9. Operon Architecture Differences Are Large but Inseparable from Lineage

Direct reconstruction of the phenotype-linked cohort architectures identified 39 type-1-like local-PrtAB genomes, 10 type-8-like distal-PrtAB genomes, 16 type-2-like genomes lacking PrtAB under the operational definition and 11 other or incomplete configurations. The type-8-like reconstruction captured a characteristic Pseudomonas lundensis pattern in which strong PrtA-like and PrtB-like homologues were present but located on a different assembly contig from the local serralysin/LipA2 region.
Among strains with day-7 phenotype data, 33 type-1-like strains had a mean score of 3.64 (median 4.0) and 10 type-8-like strains a mean of 0.80 (median 0.0), a difference of 2.84 score units with a highly significant unadjusted Mann-Whitney comparison (p = 7.6 x 10-7). A lambda-PGLS sensitivity model returned an almost identical coefficient (beta = 2.84, standard error 0.29; p = 4.3 x 10-12; fitted lambda approximately zero). These are reported as sensitivity analyses rather than as evidence of an independent operon effect, because all 10 type-8-like strains were Pseudomonas lundensis whereas type-1-like strains were distributed across multiple taxa, and the binary trait required only one minimum Fitch change across the 49 classified tips (Figure 6). Architecture and lineage are effectively non-identifiable in this cohort. The very small fitted lambda should not be read as showing that phylogeny is irrelevant: residual phylogenetic covariance in the outcome and identifiability of the predictor are different statistical questions, and a predictor that is nearly collinear with lineage cannot be separated from it by a model that only adjusts the error structure.
Table 1. Summary of the principal analyses and results.
Table 1. Summary of the principal analyses and results.
Analysis Cohort Key count Principal result
Serralysin catalogue 433 type-material genomes 231 full-length candidates 162 AprX-reference-proximal; 69 AprA-reference-proximal; all 231 marker states reproduced by independent zinc-anchored
mapping
Marker distribution 231 candidates 155 SVMSY; 60
SLMSY; 14
SIMSY; 2 TVMSY
SVMSY in 55.6% of AprX-proximal but 94.2% of AprA-proximal candidates (OR 13.0, 95% CI
4.52-37.38, p = 8.1 x 10-10)
Independent replication 2,611 strain-level RefSeq serralysins 2,311 AprX-
proximal; 300
AprA-proximal
SVMSY in 53.1% versus 94.7% (OR 15.71, 95%
CI 9.43-26.16, p = 1.6 x 10-52)
Species-level probe coverage 187 species with AprX-proximal
serralysins
102 / 40 / 45 Fully detected / partially missed / entirely missed by an exact-SVMSY target
Structural position 6 serralysin crystal structures residues 212-216 Region is the metzincin Met-turn; Met sulfur 6.15
+/- 0.05 A from the catalytic zinc; local backbone RMSD 0.37-0.60 A versus 1.19-2.68 A globally
Positional constraint 2,611 serralysins 4 of 5 positions invariant Met 99.96%, Ser 100%, Tyr 100%, Ser 98.8%;
only position 2 varies (V 58.2% / L 36.8% / I 4.1%), and it is buried (0.0% relative accessibility)
9.4-9.7 A from the zinc
Within-species variation 42 species; 102 assemblies resolved
for six species
3 species with strain-level variation P. mandelii (10 SLMSY vs 2 SVMSY), P. poae (4 vs 3), P. koreensis (15 vs 4, plus 7 assemblies with
two paralogues)
Marker versus phenotype 67 phenotyped strains mean difference - 0.52 Not significant unadjusted (p = 0.343) or under phylogenetic regularization (beta = 0.397, p = 0.397); no within-species variation available to
test
Residue scan 191 AprX positions 161 FDR-significant
pooled; 0 of 11 within-species
Pooled sequence-phenotype association dominated by shared ancestry
Alternative correlate 67 phenotyped strains Spearman rho = - 0.640 Distance to nearest AprX reference correlated with phenotype (p = 5.4 x 10-9) but attenuated
under regularization (p = 0.095)
Reconstructed architecture 76 serralysin-positive genomes 39 type-1-like; 10
type-8-like; 16 type-
2-like; 11 other
Type-8-like class confined to P. lundensis; one minimum Fitch transition; phenotype difference
not separable from lineage

4. Discussion

4.1. What the Catalogue Shows About the Marker

The clearest result is a specificity failure that is quantified twice, in two independently sampled cohorts. The SVMSY pentapeptide is close to fixed in Pseudomonas aeruginosa AprA-proximal serralysins (94.2% in the type-material cohort, 94.7% in the strain-level cohort) while being present in only about half of AprX-proximal serralysins (55.6% and 53.1%). A pentapeptide-level target therefore does not distinguish the dairy-relevant protease from the alkaline protease of an opportunistic human pathogen, and it does not cover a large fraction of the AprX-proximal sequence space it is intended to detect. Earlier aprX-directed molecular screening already encountered substantial sequence heterogeneity among milk-associated Pseudomonas, and current reviews of AprX detection identify specificity as the central methodological requirement (Marchand et al., 2009; Sinha and Kelly, 2026); the present catalogue supplies the sequence-level detail behind that requirement.
The per-species resolution is the practically usable part. Of 187 species carrying AprX-proximal serralysins, 45 would be entirely missed by an exact-match probe and 40 partially missed. For a dairy laboratory the relevant entries are specific: Pseudomonas lundensis and Pseudomonas fragi are uniformly SVMSY and would be detected reliably, whereas Pseudomonas fluorescens carries four states with only 42.6% SVMSY and would be detected erratically. This is not an argument that the pentapeptide is uninformative; it is a statement of where an exact-match assay built on it will and will not work, which is the information an assay developer needs before deployment. Section 4.2 explains why a pentapeptide target carries so little discriminating information in the first place.

4.2. The Marker Is a Catalytic-Scaffold Element with One Variable Position

The structural superposition changes what the proposed marker is. The five-residue region is the Met-turn of the metzincin clan, the short segment whose methionine packs beneath the three zinc-coordinating histidines and forms the hydrophobic floor of the catalytic site (Bode et al., 1993). It occupies residues 212-216 in both Pseudomonas aeruginosa alkaline protease and Serratia serralysin, proteins sharing 57-59% identity, and its methionine sulfur lies 6.15 A from the catalytic zinc with a standard deviation of 0.053 A across six independently determined structures. Its backbone is held more tightly in place than the surrounding catalytic domain. Four of its five residues are correspondingly invariant in sequence across 2,611 serralysins.
A pentapeptide of which four positions cannot vary is not a five-residue marker; it is a one-residue marker with four positions of scaffold attached. Read that way, several otherwise puzzling observations become expected. The near-fixation of SVMSY in the Pseudomonas aeruginosa AprA group is fixation of valine at a single buried position within an otherwise obligatory motif. The recurrence of the same state in distant lineages, requiring at least 20 changes on the serralysin phylogeny, is what a single conservative substitution at a constrained position would produce. And an assay targeted at the full pentapeptide inherits no specificity from the four invariant positions, because every serralysin carries them, including the serralysins of species the assay is not intended to detect.
The position of the variable residue also constrains mechanistic interpretation. In all six structures the position-2 side chain is fully buried, with a relative solvent accessibility of 0.0%, and its C-beta atom lies 9.4-9.7 A from the catalytic zinc, with the additional methylene of leucine extending further from the zinc rather than toward it. On this geometry the residue does not contact solvent or substrate at the active site, so a valine-to-leucine substitution cannot alter a substrate pocket by direct contact; an effect on catalysis would have to be transmitted indirectly through packing beneath the zinc site. This is a narrower and more testable hypothesis than a direct substrate-pocket interaction, and it is compatible with a real but modest biochemical effect. It should be noted that no experimental AprX structure exists, so these measurements come from serralysin homologues spanning both marker states rather than from AprX itself, and that structures determined without bound substrate cannot exclude local rearrangement on substrate binding.

4.3. Why the Phenotype Association Does Not Survive Lineage Adjustment

The phenotype-linked cohort illustrates the general problem with pooled bacterial genotype-phenotype analysis. A naive residue scan suggested widespread association between AprX sequence and proteolytic phenotype, and almost all of that signal disappeared once the analysis was restricted to residues that actually vary within species. Because clonal structure creates extensive genome-wide linkage, a residue can tag a lineage whose members share many other genetic and regulatory properties (Earle et al., 2016; Lees et al., 2016). Without repeated independent transitions or within-lineage variation, the residue and the lineage cannot be separated statistically.
The same caution applies to locus architecture, which produced a larger phenotypic contrast than the pentapeptide and is consistent with earlier reports linking aprX-lipA2 organization with proteolytic activity (Maier et al., 2020; Aguilera-Toro et al., 2023). In this cohort, however, the architecture classes were confined to single species and the binary trait required one minimum evolutionary transition, so the contrast cannot be attributed to architecture independently of background.
It is worth stating what this does not establish. None of these analyses contradicts the possibility that a SVMSY-to-SLMSY substitution alters substrate recognition within a defined protein background; a mechanistic effect at a residue and the utility of that residue as a population-level marker are separate claims with separate evidence requirements (Liang et al., 2026). The present results bear only on the second.

4.4. Defining the Experiment That Would Settle It

Section 3.6 converts this from a caution into a design. Pseudomonas mandelii, Pseudomonas poae and Pseudomonas koreensis each contain sequenced strains that differ at the marker position within a single species, which is the configuration required to vary the residue while approximately holding lineage constant. A phenotyping panel drawn from these strains, assayed for caseinolytic activity and UHT-milk destabilization, would provide a within-species test that no cohort analysed here could support. The stronger design remains engineered isogenic backgrounds differing only at the SVMSY/SLMSY position, with recombinant AprX proteins compared for kinetic activity against caseins and defined peptides followed by milk-system validation; the natural strain panel is the cheaper first step and is available now. The role of PrtA/PrtB localization would likewise be better tested by genetic manipulation within one strain background than by comparing naturally occurring species-level operon classes.

4.5. Methodological Implications for Locus Reconstruction

Annotation-string searches recovered no PrtA/PrtB-like proteins among the 231 candidates, whereas homology screening recovered them in roughly half. Several short
AprI/Inh-family proteins were biologically plausible and annotation-supported but failed a generic E-value threshold calibrated on much larger proteins. In Pseudomonas lundensis, strong PrtA-like and PrtB-like homologues were frequently located on a different contig from the local serralysin/LipA2 region. Gene-specific thresholds and explicit genomic-coordinate logic are therefore necessary if annotation conventions and assembly fragmentation are not to be recorded as biological absence. This point generalizes beyond the aprX-lipA2 system to any comparative analysis of variable accessory loci.

4.6. Implications for Spoilage-Risk Prediction

These results suggest that molecular prediction of proteolytic spoilage risk will be more robust when it combines taxonomic background, whole-locus architecture, aprX allele or alignment group, regulatory context and, where possible, a direct measure of expression or enzyme activity, in a framework that explicitly controls population structure rather than treating each polymorphism as independent (Earle et al., 2016; Lees et al., 2016). The observation that distance to a characterized AprX reference outperformed the pentapeptide in pooled data, while also attenuating under phylogenetic regularization, points in the same direction. For industrial use a hierarchical workflow is the practical form: identify the Pseudomonas lineage first, then apply lineage-specific markers or operon signatures that have been validated within that lineage.

4.7. Limitations

The phenotype cohort was not generated prospectively for this analysis, and its categorical proteolysis values were recovered from a published figure rather than from an original machine-readable phenotype table; residual digitization error cannot be excluded, and the phenotype-linked results should be read with that constraint in mind. The day-7 score is ordinal, so continuous PGLS was used only as a sensitivity analysis. The type-material collection is taxonomically broad but is not a representative survey of dairy environments and cannot be used to estimate environmental prevalence. The strain-level RefSeq cohort is larger and not restricted to one genome per species, but it is ascertainment-biased toward taxa that have been sequenced intensively for clinical or agricultural reasons, so its species-level proportions describe sequence space rather than dairy prevalence. Reference-proximity classification depends on the available curated AprX and AprA reference set and is not a nomenclatural solution. Reference-proximity assignment in the strain-level cohort used pairwise identity rather than tree-based patristic distance, which is a coarser criterion, although the close agreement between the two cohorts’ class-specific frequencies suggests it did not distort the comparison. Assembly fragmentation can separate genuinely linked genes across contigs. The structural analysis rests on six crystal structures of serralysin homologues rather than on AprX itself, none with bound substrate, and solvent accessibility and interatomic distances were computed on single static models. Finally, the per-species probe outcomes describe exact-match coverage of a pentapeptide at the protein level and are not a substitute for empirical validation of any realized assay.

5. Conclusions

A reference-anchored catalogue of Pseudomonas serralysins shows that the proposed AprX five-residue marker SVMSY is close to fixed in Pseudomonas aeruginosa AprA-proximal proteins while covering only about half of AprX-proximal proteins, a specificity failure reproduced independently in a 231-candidate type-material cohort and a 2,611-protein strain-level cohort. Species by species, an exact-match probe would entirely miss 45 of 187 species carrying AprX-proximal serralysins. Structurally, the region is the metzincin Met-turn: it occupies residues 212-216 with its methionine sulfur 6.15 A from the catalytic zinc in every available serralysin structure, four of its five positions are invariant across 2,611 serralysins, and its one variable position is fully buried 9.4-9.7 A from the zinc, so the proposed marker is a fixed element of the catalytic scaffold distinguished only by a single conservative substitution. In a separately published phenotype-linked cohort the marker was lineage-locked and showed no independent association with day-7 proteolysis, and although reconstructed operon architecture produced a large phenotype difference, that difference was inseparable from species background. Forty-two species nonetheless carry more than one marker state, and three of them contain sequenced strains differing at the marker position within a single species, defining a tractable panel for the controlled test this question requires. Marker-based spoilage-risk prediction should be validated across separately sampled lineages, and reported with the species-level coverage of the intended target.

Supplementary Materials

The following supporting information can be downloaded at the website of this paper posted on Preprints.org.

Author Contributions

A.S. is the sole author and was responsible for conceptualization, methodology, software, formal analysis, investigation, data curation, visualization, and writing of the original draft and its revision.

Funding

This research received no specific grant from any funding agency in the public, commercial or not-for-profit sectors.

Data Availability Statement

All data analysed in this study are publicly available. Genome assemblies for the phenotype-linked cohort are associated with BioProject PRJNA612806; type-material and strain-level genome and protein records were retrieved from NCBI RefSeq. Supplementary Tables S1-S14 accompany this preprint as a single spreadsheet workbook with a contents sheet. The structural superposition of 1KAP and 1SAT used in Section 3.4 is provided as a coordinate file. The analysis code is available at https://github.com/aritrasinha26/aprx-lineage-analysis and is archived at https://doi.org/10.5281/zenodo.22832350 (release v1.0.0), which is the version used to produce the results reported here.

Conflicts of Interest

The author declares no competing interests.

References

  1. Aguilera-Toro, M., Kragh, M. L., Thomasen, A. V., Piccini, V., Rauh, V., Xiao, Y., Wiking, L., Poulsen, N. A., Hansen, L. T., & Larsen, L. B. (2023). Proteolytic activity and heat resistance of the protease AprX from Pseudomonas in relation to genotypic characteristics. International Journal of Food Microbiology, 391-393, 110147. [CrossRef]
  2. Bagliniere, F., Tanguy, G., Jardin, J., Mateos, A., Briard-Bion, V., Rousseau, F., Robert, B., Beaucher, E., Humbert, G., Dary, A., Gaillard, J. L., Amiel, C., & Gaucheron, F. (2012). Quantitative and qualitative variability of the caseinolytic potential of different strains of Pseudomonas fluorescens: implications for the stability of casein micelles of UHT milks during their storage. Food Chemistry, 135(4), 2593-2603. [CrossRef]
  3. Bagliniere, F., Mateos, A., Tanguy, G., Jardin, J., Briard-Bion, V., Rousseau, F., Robert, B., Beaucher, E., Gaillard, J. L., Amiel, C., Humbert, G., Dary, A., & Gaucheron, F. (2013). Proteolysis of ultra high temperature-treated casein micelles by AprX enzyme from Pseudomonas fluorescens F induces their destabilisation. International Dairy Journal, 31(2), 55-61. [CrossRef]
  4. Baumann, U. (1994). Crystal structure of the 50 kDa metallo protease from Serratia marcescens. Journal of Molecular Biology, 242(3), 244-251. PDB 1SAT.
  5. Baumann, U., Wu, S., Flaherty, K. M., & McKay, D. B. (1993). Three-dimensional structure of the alkaline protease of Pseudomonas aeruginosa: a two-domain protein with a calcium binding parallel beta roll motif. EMBO Journal, 12(9), 3357-3364. PDB 1KAP.
  6. Baumann, U., Bauer, M., Letoffe, S., Delepelaire, P., & Wandersman, C. (1995). Crystal structure of a complex between Serratia marcescens metallo-protease and an inhibitor from Erwinia chrysanthemi. Journal of Molecular Biology, 248(3), 653-661. PDB 1SMP.
  7. Bode, W., Gomis-Ruth, F. X., & Stockler, W. (1993). Astacins, serralysins, snake venom and matrix metalloproteinases exhibit identical zinc-binding environments (HEXXHXXGXXH and Met-turn) and topologies and should be grouped into a common family, the metzincins. FEBS Letters, 331(1-2), 134-140. [CrossRef]
  8. Camacho, C., Coulouris, G., Avagyan, V., Ma, N., Papadopoulos, J., Bealer, K., & Madden, T. L. (2009). BLAST+: architecture and applications. BMC Bioinformatics, 10, 421. [CrossRef]
  9. Cock, P. J. A., Antao, T., Chang, J. T., et al. (2009). Biopython: freely available Python tools for computational molecular biology and bioinformatics. Bioinformatics, 25(11), 1422-1423. [CrossRef]
  10. De Jonghe, V., Coorevits, A., Van Hoorde, K., Messens, W., Van Landschoot, A., De Vos, P., & Heyndrickx, M. (2011). Influence of storage conditions on the growth of Pseudomonas species in refrigerated raw milk. Applied and Environmental Microbiology, 77(2), 460-470. [CrossRef]
  11. Earle, S. G., Wu, C.-H., Charlesworth, J., et al. (2016). Identifying lineage effects when controlling for population structure improves power in bacterial association studies. Nature Microbiology, 1, 16041. [CrossRef]
  12. Eddy, S. R. (2011). Accelerated profile HMM searches. PLoS Computational Biology, 7(10), e1002195. [CrossRef]
  13. Hamada, K., Hata, Y., Katsuya, Y., Hiramatsu, H., Fujiwara, T., & Katsube, Y. (1996). Crystal structure of Serratia protease, a zinc-dependent proteinase from Serratia sp. E-15, containing a beta-sheet coil motif at 2.0 A resolution. Journal of Biochemistry, 119(5), 844-851. PDB 1SRP.
  14. Kalyaanamoorthy, S., Minh, B. Q., Wong, T. K. F., von Haeseler, A., & Jermiin, L. S. (2017). ModelFinder: fast model selection for accurate phylogenetic estimates. Nature Methods, 14, 587-589. [CrossRef]
  15. Katoh, K., & Standley, D. M. (2013). MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Molecular Biology and Evolution, 30(4), 772-780. [CrossRef]
  16. Lees, J. A., Vehkala, M., Valimaki, N., et al. (2016). Sequence element enrichment analysis to determine the genetic basis of bacterial phenotypes. Nature Communications, 7, 12797. [CrossRef]
  17. Liang, L., Wang, P., Zhao, X., Wang, Z., Xu, B., Ji, Q., Chen, Y., & Wu, D. (2026). Whole-genome analysis and molecular docking reveal how nucleotide variations in the aprX gene drive the differential spoilage potential of Pseudomonas in milk. Food Microbiology, 141, 105262. [CrossRef]
  18. Liao, C.-H., & McCallus, D. E. (1998). Biochemical and genetic characterization of an extracellular protease from Pseudomonas fluorescens CY091. Applied and Environmental Microbiology, 64(3), 914-921. [CrossRef]
  19. Maier, C., Huptas, C., von Neubeck, M., Scherer, S., Wenning, M., & Lucking, G. (2020). Genetic organization of the aprX-lipA2 operon affects the proteolytic potential of Pseudomonas species in milk. Frontiers in Microbiology, 11, 1190. [CrossRef]
  20. Marchand, S., Vandriesche, G., Coorevits, A., Coudijzer, K., De Jonghe, V., Dewettinck, K., De Vos, P., Devreese, B., Heyndrickx, M., & De Block, J. (2009). Heterogeneity of heat-resistant proteases from milk Pseudomonas species. International Journal of Food Microbiology, 133(1-2), 68-77. [CrossRef]
  21. Minh, B. Q., Schmidt, H. A., Chernomor, O., Schrempf, D., Woodhams, M. D., von Haeseler, A., & Lanfear, R. (2020). IQ-TREE 2: new models and efficient methods for phylogenetic inference in the genomic era. Molecular Biology and Evolution, 37(5), 1530-1534. [CrossRef]
  22. Miyatake, H., Hata, Y., Fujii, T., Hamada, K., Morihara, K., & Katsube, Y. (1995). Crystal structure of the unliganded alkaline protease from Pseudomonas aeruginosa IFO3080 and its conformational changes on ligand binding. Journal of Biochemistry, 118(3), 474-479. PDB 1AKL.
  23. Na, S.-I., Kim, Y. O., Yoon, S.-H., Ha, S.-M., Baek, I., & Chun, J. (2018). UBCG: up-to-date bacterial core gene set and pipeline for phylogenomic tree reconstruction. Journal of Microbiology, 56(4), 280-285. [CrossRef]
  24. Sayers, E. W., Bolton, E. E., Brister, J. R., et al. (2022). Database resources of the National Center for Biotechnology Information. Nucleic Acids Research, 50(D1), D20-D26. [CrossRef]
  25. Sinha, A., & Kelly, A. L. (2026). Techniques to detect and quantify the bacterial metalloprotease AprX in bovine milk: a review. Comprehensive Reviews in Food Science and Food Safety, 25(2), e70415. [CrossRef]
  26. Woods, R. G., Burger, M., Beven, C.-A., & Beacham, I. R. (2001). The aprX-lipA operon of Pseudomonas fluorescens B52: a molecular analysis of metalloprotease and lipase production. Microbiology, 147(2), 345-354. [CrossRef]
  27. Zhang, C., Bijl, E., Svensson, B., & Hettinga, K. (2019). The extracellular protease AprX from Pseudomonas and its spoilage potential for UHT milk: a review. Comprehensive Reviews in Food Science and Food Safety, 18(4), 834-852. [CrossRef]
  28. Zhang, W.-W., Hu, Y.-H., Wang, H.-L., & Sun, L. (2009). Identification and characterization of a virulence-associated protease from a pathogenic Pseudomonas fluorescens strain. Veterinary Microbiology, 139(1-2), 183-188. [CrossRef]
Figure 1. Reference-aware serralysin phylogeny for the type-material cohort, with the maximum-likelihood topology drawn as a cladogram. The tree contains 231 candidate serralysins and six curated AprX/AprA references, the latter marked with stars and labelled. The inner side track gives the nearest-reference class and the outer track the aligned five-residue marker state; both are keyed at right. Reference-proximity classes were assigned from patristic distances to the curated AprX and Pseudomonas aeruginosa AprA references and are summarized in Figure 2A.
Figure 1. Reference-aware serralysin phylogeny for the type-material cohort, with the maximum-likelihood topology drawn as a cladogram. The tree contains 231 candidate serralysins and six curated AprX/AprA references, the latter marked with stars and labelled. The inner side track gives the nearest-reference class and the outer track the aligned five-residue marker state; both are keyed at right. Reference-proximity classes were assigned from patristic distances to the curated AprX and Pseudomonas aeruginosa AprA references and are summarized in Figure 2A.
Preprints 234484 g001
Figure 3. The five-residue region is the metzincin Met-turn. (A) Superposition of the zinc-binding motif and Met-turn of all six experimentally determined serralysin structures, coloured by five-residue state; the catalytic zinc of 1KAP is shown, with the Met-turn backbone drawn heavy and the Met-turn sulfur and the position-2 side-chain tip marked. Chains were superposed by combinatorial extension on backbone geometry, not by sequence alignment. (B) Distances from the catalytic zinc to the Met-turn sulfur, the position-2 C-beta atom and the most distal position-2 side-chain atom, for each structure. (C) Residue composition at each of the five positions across the 2,611 serralysins of the strain-level cohort; residues occupying less than 5% of a position are drawn without a label. The interactive superposition of 1KAP and 1SAT is provided as an accompanying structure file.
Figure 3. The five-residue region is the metzincin Met-turn. (A) Superposition of the zinc-binding motif and Met-turn of all six experimentally determined serralysin structures, coloured by five-residue state; the catalytic zinc of 1KAP is shown, with the Met-turn backbone drawn heavy and the Met-turn sulfur and the position-2 side-chain tip marked. Chains were superposed by combinatorial extension on backbone geometry, not by sequence alignment. (B) Distances from the catalytic zinc to the Met-turn sulfur, the position-2 C-beta atom and the most distal position-2 side-chain atom, for each structure. (C) Residue composition at each of the five positions across the 2,611 serralysins of the strain-level cohort; residues occupying less than 5% of a position are drawn without a label. The interactive superposition of 1KAP and 1SAT is provided as an accompanying structure file.
Preprints 234484 g003
Figure 4. Species-level behaviour of an exact-SVMSY probe in the independent strain-level RefSeq cohort. (A) Percentage of AprX-proximal serralysins carrying exact SVMSY, for the 20 species with at least 10 AprX-proximal records; the number of protein records is given at right and bars at zero are drawn as stubs. (B) Assembly-level marker states for six focus species resolved through NCBI Identical Protein Groups; ‘both’ denotes an assembly carrying two serralysin paralogues with different states. Protein records are non-redundant across strains, which is why panel B is resolved to assemblies rather than to protein counts.
Figure 4. Species-level behaviour of an exact-SVMSY probe in the independent strain-level RefSeq cohort. (A) Percentage of AprX-proximal serralysins carrying exact SVMSY, for the 20 species with at least 10 AprX-proximal records; the number of protein records is given at right and bars at zero are drawn as stubs. (B) Assembly-level marker states for six focus species resolved through NCBI Identical Protein Groups; ‘both’ denotes an assembly carrying two serralysin paralogues with different states. Protein records are non-redundant across strains, which is why panel B is resolved to assemblies rather than to protein counts.
Preprints 234484 g004
Figure 5. Core-genome phylogeny of the 87-genome phenotype-linked Pseudomonas cohort. Three tracks are shown and keyed at right: the aligned five-residue AprX-region state (squares), the day-7 categorical proteolysis phenotype (circles) and the homology-reconstructed operon class (triangles), where Type1-like_local_prtAB and Type8-like_distal_prtAB correspond to the type-1-like and type-8-like architectures analysed in Section 3.9 and Figure 6. Tips without a mark on a track had no resolved value for it. The reconstructed operon classes carried on the third track are analysed in Figure 6.
Figure 5. Core-genome phylogeny of the 87-genome phenotype-linked Pseudomonas cohort. Three tracks are shown and keyed at right: the aligned five-residue AprX-region state (squares), the day-7 categorical proteolysis phenotype (circles) and the homology-reconstructed operon class (triangles), where Type1-like_local_prtAB and Type8-like_distal_prtAB correspond to the type-1-like and type-8-like architectures analysed in Section 3.9 and Figure 6. Tips without a mark on a track had no resolved value for it. The reconstructed operon classes carried on the third track are analysed in Figure 6.
Preprints 234484 g005
Figure 6. Phenotype and lineage structure of reconstructed operon classes in the phenotype-linked cohort. (A) Day-7 phenotype for the homology-reconstructed type-1-like local-PrtAB and type-8-like distal-PrtAB classes; the comparison is descriptive because operon class is strongly lineage-restricted. (B) Species distribution of reconstructed operon classes, showing that all type-8-like genomes were Pseudomonas lundensis and that the binary class required only one minimum Fitch transition.
Figure 6. Phenotype and lineage structure of reconstructed operon classes in the phenotype-linked cohort. (A) Day-7 phenotype for the homology-reconstructed type-1-like local-PrtAB and type-8-like distal-PrtAB classes; the comparison is descriptive because operon class is strongly lineage-restricted. (B) Species distribution of reconstructed operon classes, showing that all type-8-like genomes were Pseudomonas lundensis and that the binary class required only one minimum Fitch transition.
Preprints 234484 g006
Table 2. Position of the five-residue region relative to the catalytic zinc in all six experimentally determined serralysin structures, after superposition onto 1KAP by combinatorial extension. The region occupies residues 212-216 in every structure, with the Met-turn methionine at 214 and the zinc ligands at His176, His180 and His186. Relative accessibility is expressed as a percentage of the maximum accessible surface area for that residue type.
Table 2. Position of the five-residue region relative to the catalytic zinc in all six experimentally determined serralysin structures, after superposition onto 1KAP by combinatorial extension. The region occupies residues 212-216 in every structure, with the Met-turn methionine at 214 and the zinc ligands at His176, His180 and His186. Relative accessibility is expressed as a percentage of the maximum accessible surface area for that residue type.
Structure Five-residue state Met-turn Sδ to Zn²⁺ (Å) Position-
2 Cβ to Zn²⁺ (Å)
Position-2 side-chain tip to Zn²⁺ (Å) Position-2 relative accessibility (%) Five-residue Cα RMSD
to 1KAP
(Å)
Global RMSD
to1KAP (Å)
1AF0 —
Serratia marcescens serralysin (inhibitor
complex)
SLMSY 6.22 9.59 10.94 0.0 0.52 2.29
1SAT —
Serratia marcescens
serralysin
SLMSY 6.10 9.56 10.82 0.0 0.59 2.53
1SMP —
Serratia marcescens serralysin (inhibitor
complex)
SLMSY 6.11 9.66 10.83 0.0 0.60 2.51
1SRP —
Serratia sp. E-15 serralysin
SLMSY 6.14 9.52 10.77 0.0 0.57 2.68
1AKL —
Pseudomonas aeruginosa alkaline protease
(AprA)
SVMSY 6.22 9.38 9.06 0.0 0.37 1.19
1KAP —
Pseudomonas aeruginosa alkaline protease (AprA)
SVMSY 6.14 9.35 9.15 0.0 0.00 0.00
Table 3. Replication of the AprX/AprA five-residue contrast in two cohorts assembled under different sampling designs. The type-material cohort takes one assembly per organism name; the strain-level cohort applies no per-species restriction.
Table 3. Replication of the AprX/AprA five-residue contrast in two cohorts assembled under different sampling designs. The type-material cohort takes one assembly per organism name; the strain-level cohort applies no per-species restriction.
Quantity Type-material cohort (this
study, 433 genomes)
Strain-level RefSeq cohort (this
study, independent)
Serralysin candidates (n) 231 2611
One genome per species yes no
AprX-reference-proximal
(n)
162 2311
SVMSY among AprX-
proximal (%)
55.6 53.1
AprA-reference-proximal
(n)
69 300
SVMSY among AprA-
proximal (%)
94.2 94.7
Fisher odds ratio 13.00 15.71
Approximate 95%
confidence interval
4.52–37.38 9.43–26.16
Fisher two-sided p 8.1 × 10⁻¹⁰ 1.6 × 10⁻⁵²
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.