Preprint
Article

This version is not peer-reviewed.

Comparative Genomic Analysis of Clonal Variants of a Spirabiliibacterium Lineage Recovered from a Chicken in the United States

Submitted:

03 August 2026

Posted:

04 August 2026

You are already at the latest version

Abstract
The family Pasteurellaceae includes pathogenic and opportunistic bacteria affecting poultry, yet taxonomic resolution remains challenging due to phenotypic heterogeneity, limited genomic representation, and the poor performance of routine diagnostic systems for uncommon taxa. At the Clemson Veterinary Diagnostic Center, a Gram-negative bacterium was isolated from a backyard hen. The isolate produced two distinct colony morphotypes with identical biochemical profiles. Both were misidentified as Sphingomonas paucimobilis by VITEK® 2, while MALDI-TOF MS failed to identify them. Whole-genome sequencing using Oxford Nanopore long reads with PacBio HiFi polishing generated near-complete chromosome-level assemblies. Comparative genomic analyses demonstrated that morphotypes CVDC-smooth and CVDC-rough represent clonal variants of a distinct Spirabiliibacterium lineage sharing 93% average nucleotide identity (ANI) and 52% digital DNA–DNA hybridization (dDDH) with the closest relative, Spirabiliibacterium mucosae. Copy-number variation within a tandemly duplicated ~20-kb genomic region likely contributes to differences in colony morphology. Pan-genome analysis revealed an open genome with lineage-specific gene content and few virulence-associated genes. This study identifies a genomically distinct lineage within the genus Spirabiliibacterium, reports the first isolation and chromosome-level genome assemblies of a Spirabiliibacterium lineage from a chicken in the United States, and highlights the limitations of routine diagnostic methods for identifying uncommon bacterial taxa.
Keywords: 
;  ;  ;  ;  

1. Introduction

The family Pasteurellaceae comprises Gram-negative, non-motile, non-spore-forming coccobacilli or short rods that are often pleomorphic and possess relatively small genomes [1,2]. These fastidious bacteria grow under aerobic or facultatively anaerobic conditions at approximately 37 °C and commonly colonize the upper respiratory and reproductive tracts of birds and other animals as commensals, although some species act as opportunistic pathogens under conditions of stress, immunosuppression, or coinfection [1,3,4]. In poultry, infections caused by members of this family range from acute septicemia to chronic respiratory or reproductive disease and may occasionally involve joints or other tissues [1,4,5]. Among avian Pasteurellaceae, Pasteurella multocida and Gallibacterium anatis are well-recognized pathogens that encode virulence factors contributing to colonization and disease [5,6,7]. Although many members exhibit host-associated ecological preferences, their pathogenic potential and evolutionary relationships remain incompletely understood [1,8,9].
Traditional diagnostic methods, such as biochemical testing and use of VITEK® 2, provide limited resolution for closely related taxa, including members of the Pasteurellaceae [10,11,12]. Although MALDI-TOF mass spectrometry enables rapid bacterial identification in many diagnostic settings, it may fail to reliably classify rare, divergent, or poorly represented species in reference databases [13,14]. Phenotypic heterogeneity further complicates identification, particularly when colony morphology varies, despite similar metabolic profiles; such variation may arise from underlying genomic differences not detected by routine biochemical testing [15,16]. Consequently, bacterial classification increasingly relies on molecular approaches, including 16S rRNA gene sequencing and whole-genome sequencing (WGS) [1,10,17]. The value of WGS for resolving the identity of uncommon bacterial isolates that cannot be reliably identified by routine diagnostic methods has also been demonstrated in other bacterial genera [18]. Long-read sequencing technologies, including Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio), enable complete genome assembly, resolution of repetitive regions, and robust comparative genomic and phylogenetic analyses [19,20,21]. Despite these advances, extensive genomic diversity and highly open pangenomes continue to complicate taxonomic resolution within the family Pasteurellaceae, resulting in ongoing revisions [1,17,22]. This taxonomic uncertainty is exemplified by several Bisgaard taxa that could not be reliably classified based on their phenotypic characteristics and have been reported from a variety of avian species and rodents, with or without associated clinical disease [1,3,10].
Bisgaard taxon 14, initially characterized by its phenotypic and biochemical characteristics, has been associated with respiratory disease and fowl cholera–like disease in poultry [3,9]. Over time, numerous Pasteurella-like isolates have been assigned to this taxon, although some have remained unidentified using conventional diagnostic methods [11]. Bisgaard taxon 14 was subsequently reclassified as Spirabiliibacterium mucosae (S. mucosae) based on phylogenetic, genomic, and phenotypic analysis [10]. Publicly available genomic resources for S. mucosae remain limited, including a fragmented multi-contig assembly for the type strain 20609/3ᵀ (GCF_014884965.1, NCBI RefSeq) and Illumina short-read data for the TN_CUL_2021 (SRR19976462; NCBI Sequence Read Archive) [10,17]. Genomic analyses of these isolates identified limited virulence and antimicrobial resistance (AMR) genes, suggesting their potential relevance to avian health [17]. However, the limited availability of complete genome assemblies continues to hinder comparative genomic analyses, taxonomic resolution, and the accurate identification of uncommon clinical and veterinary isolates.
In 2022, an eight-year-old Dominique hen from a small backyard flock in South Carolina was submitted to the Clemson Veterinary Diagnostic Center (CVDC) for evaluation of chronic neurologic disease. Necropsy and histopathology revealed multifactorial lesions, including granulomatous pneumonia, but routine diagnostic testing failed to identify a definitive etiologic agent. During bacterial culture, a Gram-negative isolate recovered from the liver could not be reliably identified using the VITEK® 2 biochemical system and remained unidentified by MALDI-TOF mass spectrometry. Extended incubation revealed two distinct colony morphotypes from the same isolate, designated CVDC-smooth and CVDC-rough. To resolve their taxonomic identity, whole-genome sequencing with Oxford Nanopore Technologies, followed by PacBio HiFi polishing, generated near-complete chromosome-level assemblies for both morphotypes. Comparative genomic analyses indicated that the isolates represent clonal variants of a genomically distinct Spirabiliibacterium lineage closely related to S. mucosae with average nucleotide identity (ANI) and digital DNA–DNA hybridization (dDDH) values below commonly accepted species-level thresholds [23,24]. They differ in the copy number of a ~20-kb tandem repeat region.
The objectives of this study were to identify an unclassified bacterial isolate recovered from a chicken, characterize its genomic features and compare it with related members of the family Pasteurellaceae, and provide high-quality genomic references for the understudied genus Spirabiliibacterium. To our knowledge, this work presents the first chromosome-level genome assemblies for the genus Spirabiliibacterium, documents the first isolation of a Spirabiliibacterium lineage from a chicken in the United States, and expands the genomic resources available for future comparative, evolutionary, and diagnostic investigations of this poorly characterized lineage.

2. Materials and Methods

2.1. Clinical History and Routine Diagnostic Testing

In 2022, an 8-year-old Dominique hen from a backyard flock of six birds was submitted to CVDC, Columbia, South Carolina, USA, with a reported 1-year history of progressive partial paralysis, weight loss, and wing lesions. Following euthanasia, a complete necropsy was performed, including gross and histopathologic examination. Routine diagnostic methods were used to investigate potential infectious etiologies. Uncut, formalin-fixed, paraffin-embedded (FFPE) tissue sections were submitted to the Michigan State University Veterinary Diagnostic Laboratory for immunohistochemistry (IHC) to rule out Marek’s disease virus (MDV), avian leukosis virus (ALV), and reticuloendotheliosis virus (REV). Swabs were collected from multiple tissues, including the liver, and from lesions identified at necropsy, and cultured on 5% sheep blood agar (Hardy Diagnostics, CA, USA) at 37 °C with 5% CO₂. Antimicrobial susceptibility testing was performed using Kirby-Bauer disk diffusion with β-lactams, aminoglycosides, tetracyclines, trimethoprim–sulfamethoxazole, macrolides, rifamycin, and sulfonamides (BD Diagnostics, Sparks, MD, USA).
Bacterial isolates were initially characterized by Gram staining, basic biochemical assays, the VITEK® 2 system (bioMérieux, NC, USA), and MALDI-TOF MS (Bruker, MA, USA); however, one Gram-negative bacterial isolate recovered from the liver could not be identified by these methods. To further characterize this bacterial isolate, incubation was extended to 72 h, resulting in the appearance of two distinct colony morphotypes (CVDC-smooth and CVDC-rough) on blood agar. Additionally, the two morphotypes were cultured on MacConkey and chocolate agar to evaluate their growth on alternative media. These morphotypes were subcultured separately and selected for whole-genome sequencing.

2.2. DNA Extraction and Sequencing

Genomic DNA was extracted from CVDC-smooth and CVDC-rough isolates using the Monarch® Genomic DNA Purification Kit (New England Biolabs, USA). Sequencing libraries were prepared using the ONT SQK-LSK114 ligation kit (Oxford Nanopore Technologies, Oxford, UK), and sequenced on a GridION (Oxford Nanopore Technologies, Oxford, UK) platform with FLO-MIN114 flow cells (one flow cell per isolate) for approximately 72 h following the manufacturer’s instructions. In parallel, both isolates were sequenced using PacBio HiFi technology at the University of Delaware DNA Sequencing and Genotyping Center by submitting bacterial isolates on 5% sheep blood agar slants.

2.3. Basecalling, Genome Assembly and Annotation, Taxonomic Identification

Oxford Nanopore (ONT) POD5 files were base-called using Dorado in super-accuracy mode. The resulting BAM files were converted to FASTQ using SAMtools [25,26], filtered with NanoFilt (minimum Q score 9), [27] and quality-checked with SeqKit [28]. No filters were applied for PacBio HiFi reads. Initial genome assembly attempts using Flye and the Nextflow EPI2ME wf-bacterial de novo assembly workflow failed to generate genome assemblies, likely due to the large ONT read dataset. Therefore, ONT and PacBio HiFi reads were assembled using AutoCycler, [29], a de novo genome assembly pipeline incorporating read subsampling and multiple assemblers. The ONT assemblies of both isolates were initially polished by self-alignment to ONT reads using the Nextflow EPI2ME wf-bacterial genome workflow in reference and isolate modes. Final polishing was performed with NextPolish, [30] using PacBio HiFi reads, which provide higher per-base accuracy than ONT long reads, and enable more reliable detection of small insertions and deletions and variant fractions. Isolate purity for both CVDC-smooth and CVDC-rough was assessed using SAMtools with Nextflow EPI2ME-generated self-alignments. Genome circularization was confirmed by visual inspection of assembly graphs in Bandage (reference), whereas assembly completeness and uniform read coverage were evaluated by mapping sequencing reads and inspecting coverage profiles in the Integrative Genomics Viewer (IGV) (reference) prior to downstream analyses.
Assembly base-level accuracy and completeness were assessed with Merqury (k-mer size 21) [31]. Genome size and GC content were calculated using QUAST [32]. Genome completeness was assessed using BUSCO with the Pasteurellales_odb12 dataset [33]. Genome annotation was performed with Prokka [34]. Taxonomic identification was performed using web-based Basic Local Alignment Search Tool for nucleotides (BLASTn (core_nt), [35], PubMLST, [36], GTDB-Tk v2.3.2 (via KBase), [37,38], and Kraken2 [39]. Genome assemblies of CVDC-smooth and CVDC-rough were also aligned to the genomes of S. mucosae 20609/3T and TN_CUL_2021 using Minimap2, and alignment statistics were calculated with SAMtools to evaluate genomic similarity and species-level relatedness [25,26,40].

2.4. Phylogenetics, Average Nucleotide Identity (ANI), and Genome-To-Genome Distance Analysis

Phylogenetic analysis was performed on 28 genomes, including 22 representative Pasteurellaceae genomes, 4 Spirabiliibacterium genomes from NCBI, CVDC-smooth and CVDC-rough isolates and S. mucosae strain TN_CUL_2021 (kindly provided by the authors) [17]. Multiple sequence alignments were generated using REALPHY, [41], with Spirabiliibacterium genomes as references. Phylogenetic trees were constructed with IQ-TREE, [42] and visualized using a custom Python Script. Accession numbers for all genomes were provided in the Supplementary Materials (Table S1).
Average nucleotide identity (ANI) between CVDC-smooth, CVDC-rough, S. mucosae type strain 20609/3ᵀ, and TN_CUL_2021 was calculated using FastANI [43]. Digital DNA–DNA hybridization (dDDH) values were calculated using the Genome-to-Genome Distance Calculator (GGDC) with BLAST+ local alignment (Formula 2) [44]. Commonly accepted species-level genomic thresholds of ≥95% ANI and ≥70% dDDH were used to assess genomic relatedness [23,24].

2.5. SNP Distance Analysis

Pairwise single-nucleotide polymorphism (SNP) distances were calculated using SNP-dists v0.8.2 [45]. Whole-genome alignments were generated with REALPHY, [41], using the CVDC-smooth, CVDC-rough, and publicly available Spirabiliibacterium genomes used as reference sequences. The resulting merged alignment was processed with SNP-dists using default parameters to produce a pairwise SNP distance matrix. Reported SNP distances represent single-nucleotide substitutions identified within genomic regions shared across all aligned genomes.

2.6. Pan-Genome, Virulence, and AMR Gene Analysis

Pan-genome analysis of 28 genomes, including CVDC-smooth and CVDC-rough, four Spirabiliibacterium genomes, and 22 Pasteurellaceae genomes, was performed using Roary v3.13.0 with a 95% BLASTp identity threshold [46]. The resulting gene presence–absence matrix was analyzed with Scoary2 to identify lineage-associated genes [47]. Default parameters were used unless otherwise specified. Core, soft-core, shell, and cloud gene categories were defined based on Roary output.
Virulence genes in the 28 genomes were screened against the Virulence Factor Database (VFDB) [48] and visualized using a custom Python script. AMR genes in CVDC-smooth and CVDC-rough were screened using ResFinder and the Comprehensive Antibiotic Resistance Database (CARD) via ABRicate [49,50]. A minimum threshold of ≥70% sequence identity and ≥80% coverage was applied for both virulence and AMR gene detection.

2.7. Variant Calling and Tandem Repeat Analysis

Genomic differences between the CVDC-smooth and CVDC-rough isolates were assessed using the BV-BRC Variation Analysis pipeline, which integrates BWA-MEM for read mapping, FreeBayes for variant calling, and SnpEff for functional annotation [51]. For this analysis, reads from CVDC-rough were aligned to the CVDC-smooth genome. PacBio HiFi reads were processed through the BV-BRC pipeline, while Oxford Nanopore Technologies (ONT) reads were analyzed using the Nextflow EPI2ME wf-bacterial-genome workflow. Variant profiles from both platforms were highly consistent; therefore, PacBio HiFi–derived variants were used for downstream analyses due to their higher per-base accuracy. Structural variants and tandem-repeat copy-number differences between CVDC-smooth and CVDC-rough were assessed by visualizing Minimap2 alignments in IGV and examining repeat junctions, read lengths, read coverage, and mapping percentage [40,52]. The repeat structures were further evaluated by performing BLASTn comparisons between CVDC-smooth and CVDC-rough, S. mucosae type strain 20609/3T (GCF_014884965.1) and TN_CUL_2021 [35].

2.8. Computational Analysis and Data Summary

All data analyses were conducted on the Palmetto 2 high-performance computing cluster (Clemson University, USA) using the SLURM scheduler, except for web-based tools [53].
Genome assemblies and raw sequencing reads generated in this study have been deposited in the National Center for Biotechnology Information (NCBI) under BioProject accessions PRJNA1447504 and PRJNA1448032. The isolates CVDC-smooth and CVDC-rough are associated with BioSample accessions SAMN56998712 and SAMN56998713, respectively. 16S rRNA gene sequences are also deposited in NCBI GenBank, and associated accession numbers are PZ459834 and PZ459835. All software tools, versions, and parameters used in this study are provided in Supplementary Table S2.

3. Results

3.1. Clinical Findings

The bird was alive but weak on presentation, responsive to external stimuli, and in lateral recumbency. Necropsy of the hen revealed multifocal granulomatous pneumonia associated with inhaled dust, along with locally invasive follicular cysts on the wings, consistent with avian keratoacanthoma. The examination also revealed an overgrown, penetrating spur and peripheral nerve infiltrates, and avian keratoacanthoma, which may have contributed to clinical lameness, inability to move, and subsequent health deterioration.
IHC testing for MDV, ALV, and REV was negative, supporting the absence of known viral lymphoproliferative disease. Aerobic culture and antimicrobial susceptibility testing of a right-wing abscess yielded Escherichia coli (E. coli), Pseudomonas aeruginosa (P. aeruginosa), and Proteus mirabilis (P. mirabilis). Multidrug resistance was observed in the P. aeruginosa and P. mirabilis isolates, whereas the E. coli isolate was fully susceptible to all antimicrobials tested. A Gram-negative organism was additionally isolated from the liver. The colonies produced by this isolate were greyish on culture plates and appeared pleomorphic under microscopy. The bacterial colonies were pinpoint after overnight incubation, but enlarged substantially after 72 h at 37 °C with 5% CO₂ on 5% sheep blood agar. Growth occurred on Chocolate agar but not on MacConkey agar. Two distinct colony morphotypes (CVDC-smooth and CVDC-rough) were observed: a uniformly grey “smooth” variant (CVDC-smooth) and a “rough” variant with grey centers and translucent margins (CVDC-rough).
Initial identification using the VITEK® 2 Gram-negative (GN) card classified both isolates as Sphingomonas paucimobilis with 90% probability; however, MALDI-TOF MS did not provide any identification. VITEK® 2 GN card profiles were identical between the two morphotypes, despite their distinct colony morphology (Table 1). Both isolates were positive for D-glucose, sucrose, γ-glutamyl transferase, phosphatase, and arylamidases and negative for D-maltose, sugar alcohols, H₂S production, urease, citrate and malonate utilization, and decarboxylases, lysine and ornithine decarboxylase activities. Both morphotypes were catalase-positive, oxidase-negative, and indole production-negative. These results demonstrated that routine phenotypic identification methods were insufficient for reliably identifying the isolate, prompting subsequent genomic characterization by whole-genome sequencing.

3.2. Read Quality and Genome Assembly and Annotation

SeqKit analysis of ONT and PacBio HiFi reads showed that N50 values ranged from 11,598 bp (PacBio HiFi) to 21,299 bp (ONT), with a mean GC content of ~49% (Table 2). AutoCycler assembly produced one large circular contig (CVDC-smooth: 2.18 Mb with ONT reads and 2.14 Mb with PacBio reads; CVDC-rough: 2.14 Mb with both ONT and PacBio reads) and a small contig of 2,428 bp for each isolate. Assembly polishing with NextPolish improved base-level accuracy, as indicated by increased Merqury QV scores from 52.2 to 63.3 for CVDC-rough and from 57.1 to ≥63 for CVDC-smooth, with a slight increase in completeness (96.18–96.20% and 97.36–97.37%, respectively). BUSCO completeness scores were high (CVDC-smooth: 97.3%; CVDC-rough: 97.4%), and QUAST confirmed genome sizes and GC content, with both morphotypes exhibiting a GC content of 49.15%. Self-alignment of ONT assemblies to raw reads using the EPI2ME wf-bacterial-genome workflow yielded 100% primary read mapping for both isolates, consistent with high assembly accuracy and isolate purity, with no detectable contamination. Genome annotation with Prokka identified 2,187 coding sequences (CDSs) in CVDC-smooth and 2,140 in CVDC-rough. Both genomes contained 19 rRNA genes, 59 tRNAs, and one tmRNA. The additional 47 CDSs in CVDC-smooth correspond to genes duplicated within the tandemly repeated ~20-kb genomic region. After accounting for overlapping CDSs at the repeat boundaries, the increase in CDS number was fully explained by the additional repeat copies rather than the acquisition of novel genes. Overall, the combined ONT/PacBio workflow produced high-quality, near-complete chromosome-level genome assemblies suitable for downstream comparative genomic analyses.

3.3. Taxonomic Identification

Whole-genome analysis using Nextflow Epi2ME wf-bacterial genome classified both CVDC-smooth and CVDC-rough as S. mucosae, with 96.4% ANI and 15% reference genome coverage. BLASTn alignment of the S. mucosae type strain 20609/3T to CVDC-smooth and CVDC-rough assemblies revealed the largest contiguous alignment of ~74 kb with 94.7% identity, covering ~3.4% of the reference genome. Read mapping with Minimap2 showed 91.4–97.6% of primary reads aligned to S. mucosae, but the taxonomic assignments varied across multiple databases, reflecting the limited representation of Spirabiliibacterium in public repositories. BLASTn searches against core_nt identified CVDC-smooth and CVDC-rough as Avibacterium paragallinarum (9% coverage, 83% identity), whereas PubMLST classified them as S. mucosae (53% identity). Kraken2 failed to assign taxonomy, and GTDB-Tk placed both genomes within genus Spirabiliibacterium, with S. mucosae as the closest reference (ANI 93.2%, RED 0.9966). A small 2,428 bp contig in each isolate showed similarity to a Neisseria gonorrhoeae plasmid (CP145022.1; 75.8% identity) based on a BLASTn search.

3.4. Phylogenetic and Comparative Genomic Analyses

Phylogenetic analysis of 28 genomes revealed a well-supported monophyletic Spirabiliibacterium clade, clearly separated from Avibacterium gallinarum (bootstrap 95–100; Figure 1). The CVDC-smooth and CVDC-rough isolates clustered closely with S. mucosae, forming a distinct subclade within the Spirabiliibacterium lineage. Pairwise ANI analysis showed 100% identity between CVDC-smooth and CVDC-rough but approximately 93% identity relative to two S. mucosae strains, below the commonly accepted 95% ANI species-level threshold. Digital DNA–DNA hybridization (dDDH) values were 52%, with only a 24.96% probability that the true dDDH value exceeds the 70% species-level threshold, further supporting genomic divergence from the currently described S. mucosae strains. Additionally, SNP distance analysis revealed no SNP differences between the CVDC-smooth and CVDC-rough (Table 3), indicating they are genomically identical within shared alignment regions. In contrast, comparisons between the CVDC isolates and the available Spirabiliibacterium reference genomes identified more than 4,000 SNP differences (Table 3). Together, these findings are consistent with the ANI and dDDH analyses, confirming that CVDC-smooth and CVDC-rough are clonal variants while distinguishing them from currently available Spirabiliibacterium reference genomes.

3.5. Pan-Genome Structure, Virulence and AMR Profiles

Roary analysis of the 28-genome dataset revealed an extremely open pangenome comprising 32,617 genes, including 16 core genes and 1 soft-core gene. Scoary2 identified 312 genes conserved across all members of the genus Spirabiliibacterium (Table S3). Lineage-specific comparisons showed 377 genes unique to S. mucosae (Table S4), 732 genes shared among CVDC-smooth, CVDC-rough, and S. mucosae genomes (Table S5), and 543 genes associated with CVDC-smooth and CVDC-rough isolates (Table S6). Virulence gene screening using VFDB (Figure 2) identified three conserved genes (rfaD, lpxA, lpxC) across all 28 genomes. The CVDC-smooth and CVDC-rough and the two S. mucosae strains shared six virulence-associated genes (rfaD, lpxA, lpxC, katA, ctrC, and ctrD). The heat shock protein gene htpB was present in all the analyzed genomes except those of the S. mucosae strains. The AMR gene screening with ResFinder and ABRicate detected no acquired resistance genes in either CVDC isolate, whereas CARD identified limited intrinsic determinants, including penicillin-binding protein variants, elfamycin markers, and SMR-family efflux pumps. Phenotypic antimicrobial susceptibility testing confirmed susceptibility to β-lactams, aminoglycosides, tetracyclines, macrolides, rifamycin, and sulfonamides. Streptomycin resistance was observed despite the absence of known resistance genes. Overall, the CVDC isolates exhibited a limited repertoire of virulence-associated and antimicrobial resistance genes, consistent with their close genomic relationship to S. mucosae.

3.6. Variant Calling and Tandem Repeat Analysis of CVDC-Smooth and CVDC-Rough

Variant calling of ONT and PacBio HiFi reads from CVDC-rough generated identical profiles; consequently, PacBio HiFi-derived variants were reported. No single-nucleotide substitutions were detected between CVDC-smooth and CVDC-rough. Six localized insertion–deletion variants were identified (Table S7), comprising five intergenic variants and one coding insertion within the nucleoside transporter gene nupC. An additional frameshift insertion occurred within a hypothetical protein. All variants exhibited high variant fractions (0.78–0.99), indicating reliable, fixed mutations.
Despite the absence of core-genome SNPs, a 40-kb size difference was observed between CVDC-smooth and CVDC-rough genomes. Structural variation analysis using Minimap2 and IGV identified a prominent tandem duplication in CVDC-smooth, comprising three consecutive ~20-kb repeat units spanning positions 183,749–244,952 bp (Figure 3.A). This configuration was supported by long ONT reads spanning the entire repeat array, uniform read depth across the duplicated region, and uninterrupted coverage at the flanking boundaries. In contrast, the corresponding region in CVDC-rough lacked long-read support and was represented only by short-read alignments, which showed uneven coverage, zero mapping quality, and gaps at the repeat boundaries, consistent with a single-copy configuration. Accordingly, the corresponding genomic region was present as a single copy in CVDC-rough (positions 334,915–355,320 bp) (Figure 3.B). Although PacBio HiFi reads (~11 kb) were insufficient to span the full repeat array in CVDC-smooth, alignments demonstrated continuous coverage across the flanking regions. In CVDC-rough, HiFi alignments contained gaps across the repeat-region junctions, further supporting its single-copy status. BLASTn comparisons confirmed that the repeat expansion was unique to CVDC-smooth and absent from CVDC-rough and two strains of S. mucosae, all of which harbored a single-copy version of this region. The ~20-kb repeat region encodes genes involved in metabolism, transport, and stress response, including representative genes such as frdA–D, pykA, modA–E, btuD, pspB–C, the Sap A-D and SapF transporter system, and the membrane-associated enzyme uppP. Collectively, these findings indicate that copy-number variation within the tandemly duplicated ~20-kb region represents the primary genomic difference between the CVDC-smooth and CVDC-rough morphotypes and likely underlies their distinct colony phenotypes.

4. Discussion

This study describes the isolation and characterization of two phenotypically distinct but genomically closely related bacterial isolates recovered from the liver of a backyard hen with multifactorial disease. The bird had concurrent disease conditions, including granulomatous pneumonia, keratoacanthoma, overgrown penetrating spur and peripheral nerve infiltrates, which may have compromised host defenses and facilitated opportunistic invasion. Members of the family Pasteurellaceae include both primary pathogens and opportunistic colonizers in avian hosts, particularly in immunocompromised or stressed individuals [1,8,9]. In this context, recovery of an unclassified Gram-negative organism from a normally sterile site warranted further investigation.
Phenotypically, CVDC-smooth and CVDC-rough displayed features consistent with S. mucosae, including growth on blood and chocolate agar but not MacConkey agar, pleomorphism, and a catalase-positive, oxidase-negative, indole-negative profile [17]. Despite distinct colony morphologies, basic biochemical testing and VITEK 2 profiles were indistinguishable between CVDC-smooth and CVDC-rough, indicating that the observed phenotypic variation is not captured by standard metabolic assays. Both basic and Vitek2 card biochemical profiles were broadly comparable to those of S. mucosae, with minor lineage-specific variations. In particular, β-glucosidase activity, coumarate utilization, O/129 resistance, and Ellman reactions were positive in S. mucosae but negative in CVDC-smooth and CVDC-rough isolates (Table S8) [10,17]. A notable finding was the strong concordance between observed phenotypes and genome-derived metabolic predictions. Catalase positivity, oxidase negativity, and indole negativity were consistent with the presence of katA [54] and the absence of cytochrome c oxidase components [55] and tnaA [56], respectively, supporting the reliability of genome-based metabolic inference for poorly represented taxa and strengthening confidence in the taxonomic placement of both CVDC-smooth and CVDC-rough within Spirabiliibacterium [54,55,56]. Despite this overall agreement, VITEK® 2 misidentified the organism, and MALDI-TOF MS failed to assign its taxonomy, highlighting the limitations of database-dependent systems for uncommon taxa, consistent with previous reports of S. mucosae misidentified as Burkholderia mallei [12,13,17]. Similar diagnostic challenges have been reported for other uncommon bacterial taxa, where routine phenotypic methods were unable to provide reliable identification and whole-genome sequencing was ultimately required to resolve taxonomic identity [18]. In the present study, whole-genome sequencing (WGS) provided high-resolution taxonomic placement and further demonstrates its value for identifying unusual bacterial isolates in diagnostic settings [17,23,24]. These findings also emphasize the importance of expanding MALDI-TOF MS and automated biochemical reference databases to improve the identification of uncommon and underrepresented veterinary bacterial taxa.
High-quality hybrid assemblies generated using ONT and PacBio HiFi reads yielded near-complete genomes suitable for comparative analysis. Database-driven taxonomic tools produced inconsistent classifications, reflecting the limited representation of Spirabiliibacterium in public repositories. However, genome-based methods consistently placed both CVDC-smooth and CVDC-rough within this genus. Average nucleotide identity (~93%) and digital DNA–DNA hybridization (~52%) [23,24], values relative to S. mucosae fell below accepted species-delineation thresholds, indicating that both CVDC-smooth and CVDC-rough represent a distinct lineage within Spirabiliibacterium. Phylogenetic analysis further supported placement of the isolates within the genus Spirabiliibacterium [10,17]. Previously reported Spirabiliibacterium isolates were recovered from pigeons and ducks, whereas the isolate described in this study was obtained from the liver of a backyard chicken. This finding expands the known avian host range of the genus and, to our knowledge, represents the first report of a Spirabiliibacterium lineage recovered from a chicken in the United States. Consistent with the ANI and dDDH analyses, SNP distance analysis identified 4,193 SNP differences between CVDC-smooth and CVDC-rough and S. mucosae genomes, supporting their recognition as a genomically distinct lineage within the genus Spirabiliibacterium [17,57]. In contrast, SNP distance analysis showed no differences between the smooth and rough morphotypes, indicating that both represent the same lineage rather than separate strains [17,57]. Although genomic analyses indicate that the CVDC-smooth and CVDC-rough isolates are distinct from the currently described S. mucosae strains, the objective of this study was not to formally describe a new species but rather to characterize a genomically distinct Spirabiliibacterium lineage recovered from a chicken. To facilitate future taxonomic, comparative genomic, and phenotypic investigations, deposition of a representative isolate in a public culture collection is currently in progress.
Pan-genome analysis indicates that Spirabiliibacterium possesses a highly open genome structure, consistent with the genomic plasticity described for Pasteurellaceae [17,22]. The presence of conserved genes across all genomes supports shared functional traits and confirms the placement of CVDC-smooth and CVDC-rough within this genus. Lineage-specific gene patterns distinguished CVDC-smooth and CVDC-rough from S. mucosae, supporting them as a separate lineage. Genes shared exclusively between CVDC-smooth and CVDC-rough further confirm that these morphotypes represent the same genomic lineage despite phenotypic differences [17].
Comparative analysis of 28 Pasteurellaceae genomes showed that Spirabiliibacterium species harbor fewer virulence-associated genes than other genera within the family. Both CVDC-smooth and CVDC-rough carried seven virulence-associated genes commonly found in Pasteurellaceae [58]. Specifically, lpxA and lpxC participate in lipid A biosynthesis, a core component of lipopolysaccharide [59]. The gene rfaD contributes to the formation of the LPS core oligosaccharide [60]. The capsular polysaccharide transport genes ctrC and ctrD [61,62], support cell surface integrity, and htpB [63], encodes a heat shock chaperone involved in protein folding under stress conditions. Database-dependent variation was observed for the luxS, a gene involved in quorum sensing and biofilm formation [17,64]. VFDB detected the luxS gene only in S. mucosae strains,37 whereas genome annotation pipelines identified it in CVDC-smooth and CVDC-rough [17]. This discrepancy likely reflects differences in database coverage and annotation sensitivity rather than true biological absence. AMR determinants were sparse in CVDC-smooth and CVDC-rough, consistent with their phenotypically susceptible profiles. As previously reported for the S. mucosae type strain 20609/3ᵀ, the tetracycline resistance gene tet(B) was not detected in the CVDC genomes, aligning with their observed tetracycline susceptibility.[17] In contrast, another study showed that the tet (B) gene has been reported in the S. mucosae strain TN_CUL_2021[17]. This distribution suggests that tetracycline resistance within this broader clade is not uniformly maintained. Notably, phenotypic streptomycin resistance in the absence of known resistance genes suggests a potential intrinsic or uncharacterized mechanism [65].
Despite limited variation in virulence and AMR profiles, CVDC-smooth and CVDC-rough were highly similar at the nucleotide level, with only a small number of INDELs detected in the CVDC-rough relative to the CVDC-smooth. The major genomic difference was a three-copy expansion of an ~20-kb region in CVDC-smooth, whereas CVDC-rough and the two strains of S. mucosae retained a single copy. Long-read mapping and uniform coverage across the duplicated region confirmed that this expansion represents a genuine genomic feature rather than an assembly artifact. The duplicated locus encodes genes involved in metabolism, membrane integrity, and stress response, and increased gene dosage provides a plausible mechanism for variation in colony morphology between the two morphotypes [66]. Despite differences in genome size, the absence of core genome SNPs, the presence of only a few small indels, and 100% ANI indicate that the morphotypes represent clonal variants rather than distinct strains. Further studies are needed to determine the functional impact of this tandem duplication. In addition, investigations into the prevalence, host range, and clinical significance of this Spirabiliibacterium lineage in poultry and other avian species will improve our understanding of its epidemiology and biological significance.
The clinical significance of this organism remains uncertain. Concurrent disease processes may have facilitated opportunistic invasion, and isolation from the liver suggests potential for systemic dissemination under compromised conditions, similar to other members of the Pasteurellaceae [1,67]. However, the single-case nature of this study precludes conclusions about pathogenicity. This study expands the known diversity of Spirabiliibacterium and identifies a genomically distinct lineage associated with poultry. Structural genomic variation, rather than SNP-level divergence, appears to underlie phenotypic heterogeneity within this lineage. Collectively, these findings highlight limitations of conventional diagnostic methods and support the use of whole-genome sequencing for accurate identification and characterization of unusual bacterial isolates in veterinary diagnostics.

Supplementary Materials

The following supporting information can be downloaded at the website of this paper posted on Preprints.org. Table S1: NCBI Genome accession numbers used for phylogenetic and comparative genomic analyses; Table S2: Bioinformatics tools, versions, and parameters used in this study; Table S3: Scoary2 identified 312 core genes conserved across all six Spirabiliibacterium genomes; Table S4: Species-specific genes in Spirabiliibacterium mucosae identified by Scoary2 pangenome analysis; Table S5: Shared gene content among CVDC-smooth, CVD-rough, and Spirabiliibacterium mucosae based on Scoary2 pangenome analysis; Table S6: Genes uniquely associated with CVDC-smooth and CVDC-rough isolates within Pasteurellaceae identified by Scoary2 pangenome analysis; Table S7: High-confidence insertion–deletion variants between CVDC-smooth and CVDC-rough isolates; Table S8: Biochemical characteristics of CVDC-smooth and CVDC-rough compared with published Spirabiliibacterium mucosae. A custom Python script for phylogenetic tree visualization and heatmap generation will be provided upon request.

Author Contributions

Conceptualization, L.S.; methodology, L.S., D.D., T.K., and J.R.P.; software, L.S. and J.R.P.; validation, L.S. and J.R.P.; formal analysis, L.S. and J.R.P.; investigation, L.S., J.R.P., D.D., and T.K.; resources, D.D. and T.K.; data curation, L.S.; writing original draft preparation, L.S. and J.R.P.; review and editing, J.R.P., R.K., T.K., R.H., and M.B.; visualization, L.S. and J.R.P.; project administration, L.S. All authors have read and agreed to the published version of the manuscript.. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Ethical review and approval were waived because samples were obtained during routine diagnostic investigations and no animals were used specifically for research purposes.

Data Availability Statement

Genome assemblies and raw sequencing reads generated in this study have been deposited in the National Center for Biotechnology Information (NCBI) under BioProject accessions PRJNA1447504 and PRJNA1448032. The isolates CVDC-smooth and CVDC-rough are associated with BioSample accessions SAMN56998712 and SAMN56998713, respectively. Both isolate 16S rRNA gene sequences are also deposited in NCBI GenBank, and associated accession numbers are PZ459834 and PZ459835. All software tools, versions, and parameters used in this study are provided in Supplementary Table S2. Deposition of a representative isolate in a public culture collection is currently in progress, and the accession number will be provided once available. Additional data supporting the findings of this study are available from the corresponding author upon reasonable request.

Acknowledgments

The authors thank Austin Compton and Christie Dix of Oxford Nanopore Technologies Inc. for their valuable guidance with bioinformatic data analysis. The authors also thank Frances Pearsal, Clemson Veterinary Diagnostic Center, Livestock Poultry Health, Clemson University, for preparing the uncut slides for Immunohistochemistry.

Conflicts of Interest

The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.

References

  1. Christensen, H.; Kuhnert, P.; Nørskov-Lauritsen, N.; Planet, P.J.; Bisgaard, M. The family Pasteurellaceae. In The Prokaryotes: Gammaproteobacteria; Rosenberg, E., Ed.; Springer: London, 2014; pp. 535–564. [Google Scholar]
  2. Plata Barril, B. Francisella, Brucella and Pasteurella. In Encyclopedia of Infection and Immunity; Rezaei, N., Ed.; Elsevier, 2022; pp. 673–684. [Google Scholar]
  3. Bisgaard, M.; Mutters, R. A new facultatively anaerobic Gram-negative fermentative rod obtained from different pathological lesions in poultry and tentatively designated taxon 14. Int. J. Syst. Evol. Microbiol. 1986. [Google Scholar] [CrossRef] [PubMed]
  4. Frey, J. The role of RTX toxins in host specificity of animal pathogenic Pasteurellaceae. Vet. Microbiol. 2011, 153, 51–58. [Google Scholar] [CrossRef] [PubMed]
  5. Geda, A.M. Fowl cholera in chickens: current trends in diagnosis and phenotypic drug resistance in Gondar City, Ethiopia. Vet. Med. Int. 2024, 2024, 6613019. [Google Scholar] [CrossRef] [PubMed]
  6. Harper, M.; Boyce, J.D.; Adler, B. The key surface components of Pasteurella multocida: capsule and lipopolysaccharide. Curr. Top. Microbiol. Immunol. 2012, 361, 39–51. [Google Scholar] [CrossRef] [PubMed]
  7. Persson, G.; Bojesen, A.M. Bacterial determinants of importance in the virulence of Gallibacterium anatis in poultry. Vet. Res. 2015, 46, 57. [Google Scholar] [CrossRef] [PubMed]
  8. Christensen, J.P.; Bisgaard, M. Avian pasteurellosis: taxonomy of the organisms involved and aspects of pathogenesis. Avian Pathol. 1997, 26, 461–483. [Google Scholar] [CrossRef] [PubMed]
  9. Bisgaard, M. Ecology and significance of Pasteurellaceae in animals. Zentralbl Bakteriol. 1993, 279, 7–26. [Google Scholar] [CrossRef] [PubMed]
  10. Bisgaard, M.; Christensen, H. Classification of Bisgaard’s taxa 14 and 32 and a taxon from kestrels demonstrating satellitic growth and proposal of Spirabiliibacterium gen. nov., including three species: S. mucosae sp. nov., S. pneumoniae sp. nov., and S. falconis sp. nov. Int. J. Syst. Evol. Microbiol. 2019, 71. [Google Scholar] [CrossRef] [PubMed]
  11. Bisgaard, M.; Günther, R.; Christensen, H. Further investigations on the involvement of Taxon 14 in upper respiratory tract infections and blepharoconjunctivitis in turkeys. In Proceedings of the 4th International Symposium on Turkey Production; Hafez, M.H., Ed.; Berlin, 2007; pp. 305–316. [Google Scholar]
  12. Zong, Z.; Wang, X.; Deng, Y.; Zhou, T. Misidentification of Burkholderia pseudomallei as Burkholderia cepacia by the VITEK 2 system. J. Med. Microbiol. 2012, 61, 1483–1484. [Google Scholar] [CrossRef] [PubMed]
  13. Podin, Y. Reliability of automated biochemical identification of Burkholderia pseudomallei is regionally dependent. J. Clin. Microbiol. 2013, 51, 3076–3078. [Google Scholar] [CrossRef] [PubMed]
  14. Tsuchida, S.; Umemura, H.; Nakayama, T. Current Status of Matrix-Assisted Laser Desorption/Ionization-Time-of-Flight Mass Spectrometry (MALDI-TOF MS) in Clinical Diagnostic Microbiology. Molecules 2020, 25. [Google Scholar] [CrossRef] [PubMed]
  15. Ackermann, M. A functional perspective on phenotypic heterogeneity in microorganisms. Nat. Rev. Microbiol. 2015, 13, 497–508. [Google Scholar] [CrossRef] [PubMed]
  16. Kovacs, A.T. Colony morphotype diversification as a signature of bacterial evolution. Microlife 2023, 4, uqad041. [Google Scholar] [CrossRef] [PubMed]
  17. Karthik, K.; Anbazhagan, S.; Chitra, M.A.; Ramya, R.; Sridhar, R.; Raj, G.D. Foremost report of the whole genome of Spirabiliibacterium mucosae from India and comparative genomics of the novel genus Spirabiliibacterium. Gene 2023, 867, 147359. [Google Scholar] [CrossRef] [PubMed]
  18. Kweon, O.J.; Lim, Y.K.; Kim, H.R.; Kim, T.H.; Ha, S.M.; Lee, M.K. Isolation of a novel species in the genus Cupriavidus from a patient with sepsis using whole genome sequencing. PLoS ONE 2020, 15, e0232850. [Google Scholar] [CrossRef] [PubMed]
  19. Amarasinghe, S.L.; Su, S.; Dong, X.; Zappia, L.; Ritchie, M.E.; Gouil, Q. Opportunities and challenges in long-read sequencing data analysis. Genome Biol. 2020, 21, 30. [Google Scholar] [CrossRef] [PubMed]
  20. Jain, M.; Olsen, H.E.; Paten, B.; Akeson, M. The Oxford Nanopore MinION: delivery of nanopore sequencing to the genomics community. Genome Biol. 2016, 17, 239. [Google Scholar] [CrossRef] [PubMed]
  21. Rhoads, A.; Au, K.F. PacBio Sequencing and Its Applications. Genom. Proteom. Bioinform. 2015, 13, 278–289. [Google Scholar] [CrossRef] [PubMed]
  22. De Luca, E. Comparative genomics analyses support the reclassification of Bisgaard taxon 40 as Mergibacter gen. nov. Front Microbiol. 2021, 12, 667356. [Google Scholar] [CrossRef] [PubMed]
  23. Chun, J.; Oren, A.; Ventosa, A.; Christensen, H.; Arahal, D.R.; da Costa, M.S.; Rooney, A.P.; Yi, H.; Xu, X.W.; De Meyer, S.; et al. Proposed minimal standards for the use of genome data for the taxonomy of prokaryotes. Int. J. Syst. Evol. Microbiol. 2018, 68, 461–466. [Google Scholar] [CrossRef] [PubMed]
  24. Riesco, R.; Trujillo, M.E. Update on the proposed minimal standards for the use of genome data for the taxonomy of prokaryotes. Int. J. Syst. Evol. Microbiol. 2024, 74. [Google Scholar] [CrossRef] [PubMed]
  25. Danecek, P.; Bonfield, J.K.; Liddle, J.; Marshall, J.; Ohan, V.; Pollard, M.O.; Whitwham, A.; Keane, T.; McCarthy, S.A.; Davies, R.M.; et al. Twelve years of SAMtools and BCFtools. GigaScience 2021, 10, giab008. [Google Scholar] [CrossRef] [PubMed]
  26. Li, H.; Handsaker, B.; Wysoker, A.; Fennell, T.; Ruan, J.; Homer, N.; Marth, G.; Abecasis, G.; Durbin, R.; Genome Project Data Processing, S. The Sequence Alignment/Map format and SAMtools. Bioinformatics 2009, 25, 2078–2079. [Google Scholar] [CrossRef] [PubMed]
  27. De Coster, W.; D’Hert, S.; Schultz, D.T.; Cruts, M.; Van Broeckhoven, C. NanoPack: visualizing and processing long-read sequencing data. Bioinformatics 2018, 34, 2666–2669. [Google Scholar] [CrossRef] [PubMed]
  28. Shen, W.; Le, S.; Li, Y.; Hu, F. SeqKit: A Cross-Platform and Ultrafast Toolkit for FASTA/Q File Manipulation. PLoS ONE 2016, 11, e0163962. [Google Scholar] [CrossRef] [PubMed]
  29. Wick, R.; Howden, B.; Stinear, T. Autocycler: long-read consensus assembly for bacterial genomes. Bioinformatics 2025, btaf474. [Google Scholar] [PubMed]
  30. Hu, J.; Fan, J.; Sun, Z.; Liu, S. NextPolish: a fast and efficient genome polishing tool for long-read assembly. Bioinformatics 2020, 36, 2253–2255. [Google Scholar] [CrossRef] [PubMed]
  31. Rhie, A.; Walenz, B.P.; Koren, S.; Phillippy, A.M. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. Genome Biol. 2020, 21, 245. [Google Scholar] [CrossRef] [PubMed]
  32. Gurevich, A.; et al. QUAST: quality assessment tool for genome assemblies. Bioinformatics 2013, 29, 1072–1075. [Google Scholar] [CrossRef] [PubMed]
  33. Simão, F.A.; Waterhouse, R.M.; Ioannidis, P.; Kriventseva, E.V.; Zdobnov, E.M. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 2015, 31, 3210–3212. [Google Scholar] [CrossRef] [PubMed]
  34. Seemann, T. Prokka: rapid prokaryotic genome annotation. Bioinformatics 2014, 30, 2068–2069. [Google Scholar] [CrossRef] [PubMed]
  35. Altschul, S.F.; et al. Basic local alignment search tool. J. Mol. Biol. 1990, 215, 403–410. [Google Scholar] [CrossRef] [PubMed]
  36. Jolley, K.A.; Bray, J.E.; Maiden, M.C.J. Open-access bacterial population genomics: BIGSdb software, the PubMLST.org website and their applications. Wellcome Open Res. 2018, 3, 124. [Google Scholar] [CrossRef] [PubMed]
  37. Chaumeil, P.A.; Mussig, A.J.; Hugenholtz, P.; Parks, D.H. GTDB-Tk: a toolkit to classify genomes with the Genome Taxonomy Database. Bioinformatics 2019, 36, 1925–1927. [Google Scholar] [CrossRef] [PubMed]
  38. Arkin, A.P.; Cottingham, R.W.; Henry, C.S.; Harris, N.L.; Stevens, R.L.; Maslov, S.; Dehal, P.; Ware, D.; Perez, F.; Canon, S.; et al. KBase: The United States Department of Energy Systems Biology Knowledgebase. Nat. Biotechnol. 2018, 36, 566–569. [Google Scholar] [CrossRef] [PubMed]
  39. Wood, D.E.; Lu, J.; Langmead, B. Improved metagenomic analysis with Kraken 2. Genome Biol. 2019, 20, 257. [Google Scholar] [CrossRef] [PubMed]
  40. Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 2018, 34, 3094–3100. [Google Scholar] [CrossRef] [PubMed]
  41. Bertels, F.; Silander, O.K.; Pachkov, M.; Rainey, P.B.; van Nimwegen, E. Automated Reconstruction of Whole-Genome Phylogenies from Short-Sequence Reads. Mol. Biol. Evol. 2014, 31, 1077–1088. [Google Scholar] [CrossRef] [PubMed]
  42. Nguyen, L.-T.; et al. IQ-TREE: a fast and effective stochastic algorithm for estimating maximum-likelihood phylogenies. Mol. Biol. Evol. 2015, 32, 268–274. [Google Scholar] [CrossRef] [PubMed]
  43. Jain, C.; Rodriguez-R, L.M.; Phillippy, A.M.; et al. High throughput ANI analysis of 90K prokaryotic genomes reveals clear species boundaries. Nat. Commun. 2018, 9, 5114. [Google Scholar] [CrossRef] [PubMed]
  44. Meier-Kolthoff, J.P.; Auch, A.F.; Klenk, H.P.; Goker, M. Genome sequence-based species delimitation with confidence intervals and improved distance functions. BMC Bioinform. 2013, 14, 60. [Google Scholar] [CrossRef] [PubMed]
  45. Seemann, T. snp-dists. [PubMed]
  46. Page, A.J.; Cummins, C.A.; Hunt, M.; et al. Roary: rapid large-scale prokaryote pan genome analysis. Bioinformatics 2015, 31, 3691–3693. [Google Scholar] [CrossRef] [PubMed]
  47. Roder, T.; Pimentel, G.; Fuchsmann, P.; Stern, M.T.; von Ah, U.; Vergeres, G.; Peischl, S.; Brynildsrud, O.; Bruggmann, R.; Bar, C. Scoary2: rapid association of phenotypic multi-omics data with microbial pan-genomes. Genome Biol. 2024, 25, 93. [Google Scholar] [CrossRef] [PubMed]
  48. Liu, B.; Zheng, D.; Jin, Q.; Chen, L.; Yang, J. VFDB 2019: a comparative pathogenomic platform with an interactive web interface. Nucleic Acids Res. 2019, 47, D687–D692. [Google Scholar] [CrossRef] [PubMed]
  49. Zankari, E.; Hasman, H.; Cosentino, S.; Vestergaard, M.; Rasmussen, S.; Lund, O.; Aarestrup, F.M.; Larsen, M.V. Identification of acquired antimicrobial resistance genes. J. Antimicrob. Chemother. 2012, 67, 2640–2644. [Google Scholar] [CrossRef] [PubMed]
  50. Alcock, B.P.; Raphenya, A.R.; Lau, T.T.Y.; Tsang, K.K.; Bouchard, M.; Edalatmand, A.; Huynh, W.; Nguyen, A.V.; Cheng, A.A.; Liu, S.; et al. CARD 2020: antibiotic resistome surveillance with the comprehensive antibiotic resistance database. Nucleic Acids Res. 2020, 48, D517–D525. [Google Scholar] [CrossRef] [PubMed]
  51. Olson, R.D.; Assaf, R.; Brettin, T.; Conrad, N.; Cucinell, C.; Davis, J.J.; Dempsey, D.M.; Dickerman, A.; Dietrich, E.M.; Kenyon, R.W.; et al. Introducing the Bacterial and Viral Bioinformatics Resource Center (BV-BRC): a resource combining PATRIC, IRD and ViPR. Nucleic Acids Res. 2023, 51, D678–D689. [Google Scholar] [CrossRef] [PubMed]
  52. Robinson, J.T.; Thorvaldsdottir, H.; Winckler, W.; Guttman, M.; Lander, E.S.; Getz, G.; Mesirov, J.P. Integrative genomics viewer. Nat. Biotechnol. 2011, 29, 24–26. [Google Scholar] [CrossRef] [PubMed]
  53. Antao, A.; Burton, J.D.; Dawson, D.; Gemmill, J.; Gerstener, Z.; Godfrey, B.; Groel, S.; Jordan, Z.; Ligon, B.; Smith, D.; et al. Modernizing Clemson University's Palmetto Cluster: Lessons Learned from 17 Years of HPC Administration. In Proceedings of the Practice and Experience in Advanced Research Computing 2024: Human Powered Computing, Providence, RI, USA, 2024; p. Article 14. [Google Scholar]
  54. Barriere, C.; Bruckner, R.; Centeno, D.; Talon, R. Characterisation of the katA gene encoding a catalase and evidence for at least a second catalase activity in Staphylococcus xylosus, bacteria used in food fermentation. FEMS Microbiol. Lett. 2002, 216, 277–283. [Google Scholar] [CrossRef] [PubMed]
  55. Jurtshuk, P., Jr.; McQuitty, D.N. Use of a quantitative oxidase test for characterizing oxidative metabolism in bacteria. Appl. Env. Microbiol. 1976, 31, 668–679. [Google Scholar] [CrossRef] [PubMed]
  56. Rezwan, F.; Lan, R.; Reeves, P.R. Molecular basis of the indole-negative reaction in Shigella strains: extensive damages to the tna operon by insertion sequences. J. Bacteriol. 2004, 186, 7460–7465. [Google Scholar] [CrossRef] [PubMed]
  57. Pightling, A.W.; Pettengill, J.B.; Luo, Y.; Baugher, J.D.; Rand, H.; Strain, E. Interpreting Whole-Genome Sequence Analyses of Foodborne Bacteria for Regulatory Applications and Outbreak Investigations. Front Microbiol. 2018, 9, 1482. [Google Scholar] [CrossRef] [PubMed]
  58. Peng, Z.; Liang, W.; Liu, W.; Wu, B.; Tang, B.; Tan, C.; Zhou, R.; Chen, H. Genomic characterization of Pasteurella multocida HB01, a serotype A bovine isolate from China. Gene 2016, 581, 85–93. [Google Scholar] [CrossRef] [PubMed]
  59. Peng, Z.; Wang, X.; Zhou, R.; Chen, H.; Wilson, B.A.; Wu, B. Pasteurella multocida: Genotypes and Genomics. Microbiol. Mol. Biol. Rev. 2019, 83. [Google Scholar] [CrossRef] [PubMed]
  60. Nichols, W.A.; Gibson, B.W.; Melaugh, W.; Lee, N.G.; Sunshine, M.; Apicella, M.A. Identification of the ADP-L-glycero-D-manno-heptose-6-epimerase (rfaD) and heptosyltransferase II (rfaF) biosynthesis genes from nontypeable Haemophilus influenzae 2019. Infect. Immun. 1997, 65, 1377–1386. [Google Scholar] [CrossRef] [PubMed]
  61. Cress, B.F.; Englaender, J.A.; He, W.; Kasper, D.; Linhardt, R.J.; Koffas, M.A. Masquerading microbial pathogens: capsular polysaccharides mimic host-tissue molecules. FEMS Microbiol. Rev. 2014, 38, 660–697. [Google Scholar] [CrossRef] [PubMed]
  62. Cao, X.; Gu, L.; Gao, Z.; Fan, W.; Zhang, Q.; Sheng, J.; Zhang, Y.; Sun, Y. Pathogenicity and Genomic Characteristics Analysis of Pasteurella multocida Serotype A Isolated from Argali Hybrid Sheep. Microorganisms 2024, 12. [Google Scholar] [CrossRef] [PubMed]
  63. Garduno, R.A.; Chong, A.; Nasrallah, G.K.; Allan, D.S. The Legionella pneumophila Chaperonin - An Unusual Multifunctional Protein in Unusual Locations. Front Microbiol. 2011, 2, 122. [Google Scholar] [CrossRef] [PubMed]
  64. He, Z.; Liang, J.; Tang, Z.; Ma, R.; Peng, H.; Huang, Z. Role of the luxS gene in initial biofilm formation by Streptococcus mutans. J. Mol. Microbiol. Biotechnol. 2015, 25, 60–68. [Google Scholar] [CrossRef] [PubMed]
  65. Michael, G.B.; Bossé, J.T.; Schwarz, S. Antimicrobial resistance in Pasteurellaceae of veterinary origin. Microbiol. Spectr. 2018, 6. [Google Scholar] [CrossRef] [PubMed]
  66. Waters, E.V.; Cameron, S.K.; Langridge, G.C.; Preston, A. Bacterial genome structural variation: prevalence, mechanisms, and consequences. Trends Microbiol. 2025, 33, 875–886. [Google Scholar] [CrossRef] [PubMed]
  67. Abd El-Ghany, W.A.; Algammal, A.M.; Hetta, H.F.; Elbestawy, A.R. Gallibacterium anatis infection in poultry: a comprehensive review. Trop. Anim. Health Prod. 2023, 55, 383. [Google Scholar] [CrossRef] [PubMed]
Figure 1. Maximum-likelihood phylogenetic tree of 28 Pasteurellaceae genomes. The tree illustrates the phylogenetic placement of CVDC-smooth and CVDC-rough (highlighted in orange), which form a distinct subclade within the Spirabiliibacterium lineage. These isolates share a common internal node with the Spirabiliibacterium mucosae subclade (shown in green), indicating a shared evolutionary history. Together with S. pneumoniae and S. falconis, these taxa form a well-supported monophyletic Spirabiliibacterium clade distinct from the Avibacterium gallinarum outgroup. Node values represent bootstrap support (%) based on 1,000 replicates. Genome labels indicate strain designation, year of isolation, country of origin, and host source.
Figure 1. Maximum-likelihood phylogenetic tree of 28 Pasteurellaceae genomes. The tree illustrates the phylogenetic placement of CVDC-smooth and CVDC-rough (highlighted in orange), which form a distinct subclade within the Spirabiliibacterium lineage. These isolates share a common internal node with the Spirabiliibacterium mucosae subclade (shown in green), indicating a shared evolutionary history. Together with S. pneumoniae and S. falconis, these taxa form a well-supported monophyletic Spirabiliibacterium clade distinct from the Avibacterium gallinarum outgroup. Node values represent bootstrap support (%) based on 1,000 replicates. Genome labels indicate strain designation, year of isolation, country of origin, and host source.
Preprints 226611 g001
Figure 2. Heatmap generated using a custom Python Script showing the distribution of virulence-associated genes identified using ABRicate (VFDB) across 28 Pasteurellaceae genomes. Spirabiliibacterium mucosae strains and the isolates CVDC-smooth and CVDC-rough share six conserved genes (rfaD, lpxA, lpxC, katA, ctrC, and ctrD). The gene htpB is present in all genomes except S. mucosae, whereas three genes (rfaD, lpxA, and lpxC) are conserved across all 28 Pasteurellaceae genomes. The x-axis represents virulence-associated genes, and the y-axis lists the analyzed genomes. Heatmap intensity reflects gene coverage (%).
Figure 2. Heatmap generated using a custom Python Script showing the distribution of virulence-associated genes identified using ABRicate (VFDB) across 28 Pasteurellaceae genomes. Spirabiliibacterium mucosae strains and the isolates CVDC-smooth and CVDC-rough share six conserved genes (rfaD, lpxA, lpxC, katA, ctrC, and ctrD). The gene htpB is present in all genomes except S. mucosae, whereas three genes (rfaD, lpxA, and lpxC) are conserved across all 28 Pasteurellaceae genomes. The x-axis represents virulence-associated genes, and the y-axis lists the analyzed genomes. Heatmap intensity reflects gene coverage (%).
Preprints 226611 g002
Figure 3. Structural variation analysis of CVDC-smooth and CVDC-rough genomes. a. CVDC-smooth shows a tandem duplication of three ~20-kb repeat units supported by long ONT reads, uniform coverage, and continuous flanking alignment. b. CVDC-rough contains a single copy of this region, with uneven coverage, gaps at repeat boundaries, and the lack of long-read support, consistent with a single-copy configuration.
Figure 3. Structural variation analysis of CVDC-smooth and CVDC-rough genomes. a. CVDC-smooth shows a tandem duplication of three ~20-kb repeat units supported by long ONT reads, uniform coverage, and continuous flanking alignment. b. CVDC-rough contains a single copy of this region, with uneven coverage, gaps at repeat boundaries, and the lack of long-read support, consistent with a single-copy configuration.
Preprints 226611 g003
Table 1. VITEK® 2 gram-negative (GN) card biochemical profiles of CVDC-smooth and CVDC-rough isolates.
Table 1. VITEK® 2 gram-negative (GN) card biochemical profiles of CVDC-smooth and CVDC-rough isolates.
Well Test Mnemonic CVDC-smooth CVDC-rough
2 Ala-Phe-Pro-ARYLAMIDASE APPA - -
3 ADONITOL ADO - -
4 L-Pyrrolydonyl-ARYLAMIDASE PyrA - -
5 L-ARABITOL IARL - -
7 D-CELLOBIOSE dCEL - -
9 BETA-GALACTOSIDASE BGAL - -
10 H2S PRODUCTION H2S - -
11 BETA-N-ACETYL-GLUCOSAMINIDASE BNAG - -
12 Glutamyl Arylamidase pNA AGLTp - -
13 D-GLUCOSE dGLU + +
14 GAMMA-GLUTAMYL-TRANSFERASE GGT + +
15 FERMENTATION/GLUCOSE OFF - -
17 BETA-GLUCOSIDASE BGLU - -
18 D-MALTOSE dMAL - -
19 D-MANNITOL dMAN - -
20 D-MANNOSE dMNE - -
21 BETA-XYLOSIDASE BXYL - -
22 BETA-Alanine arylamidase pNA BAIap - -
23 L-Proline ARYLAMIDASE ProA + +
26 LIPASE LIP - -
27 PALATINOSE PLE - -
29 Tyrosine ARYLAMIDASE TyrA + +
31 UREASE URE - -
32 D-SORBITOL dSOR - -
33 SACCHAROSE/SUCROSE SAC + +
34 D-TAGATOSE dTAG - -
35 D-TREHALOSE dTRE - -
36 CITRATE (SODIUM) CIT - -
37 MALONATE MNT - -
39 5-KETO-D-GLUCONATE 5KG - -
40 L-LACTATE alkalinization ILATk - -
41 ALPHA-GLUCOSIDASE AGLU - -
42 SUCCINATE alkalinization SUCT - -
43 Beta-N-ACETYL-GALACTOSAMINIDASE NAGA - -
44 ALPHA-GALACTOSIDASE AGAL - -
45 PHOSPHATASE PHOS + +
46 Glycine ARYLAMIDASE GlyA + +
47 ORNITHINE DECARBOXYLASE ODC - -
48 LYSINE DECARBOXYLASE LDC - -
53 L-HISTIDINE assimilation IHISa - -
56 COUMARATE CMT - -
57 BETA-GLUCURONIDASE BGUR - -
58 O/129 RESISTANCE (comp. vibrio.) O129R - -
59 Glu-Gly-Arg-ARYLAMIDASE GGAA - -
61 L-MALATE assimilation IMLTa (-) -
62 ELLMAN ELLM - -
64 L-LACTATE assimilation ILATa - -
Table 2. Seqkit-derived quality metrics for sequencing reads from CVDC-smooth and CVDC-rough.
Table 2. Seqkit-derived quality metrics for sequencing reads from CVDC-smooth and CVDC-rough.

Parameter
CVDC-rough CVDC-rough CVDC-smooth CVDC-smooth
PacBio ONT PacBio ONT
Format FASTQ FASTQ FASTQ FASTQ
Type DNA DNA DNA DNA
Num_seqs 448,313 3,860,174 488,711 2,629,437
Sum_len 5263368463 34387493888 6101321108 20028319930
Min_len 270 5 148 5
Avg_len 11740.4 8908.3 12484.5 7617
Max_len 44430 1121731 46772 808968
Q1 10225 1462 10518 958
Q2 11216 4863 11817 3977
Q3 12807 12084 13826 9332
Sum_gap 0 0 0 0
N50 11598 18585 12375 16396
N50_num 10924 49631 12071 52989
Q20(%) 98.15 91.52 98.02 89.95
Q30(%) 95.67 85.13 95.34 82.37
AvgQual 26.3 19.41 26.03 18.78
GC(%) 49.09 49.1 49.11 49.08
Table 3. Pairwise single-nucleotide polymorphism (SNP) distances between CVDC isolates and Spirabiliibacterium genomes.
Table 3. Pairwise single-nucleotide polymorphism (SNP) distances between CVDC isolates and Spirabiliibacterium genomes.
Sample S. mucosae TN_CUL_
2021
S. falconis strain NCTC S. mucosae strain 20609/3 CVDC-rough CVDC-smooth S. pneumoniae strain HPA106
S. mucosae TN_CUL_2021 0 10432 749 4254 4254 7864
S. falconis strain NCTC 10432 0 10254 11574 11574 11802
S. mucosae strain 20609/3 749 10254 0 4193 4193 7723
CVDC-rough 4254 11574 4193 0 0 9020
CVDC-smooth 4254 11574 4193 0 0 9019
S. pneumoniae strain HPA106 7864 11802 7723 9020 9019 0
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings