Preprint
Review

This version is not peer-reviewed.

Advances in Decoding Bacterial N-Terminal Proteoforms: Technologies, Challenges, and Functional Insights

Submitted:

04 June 2026

Posted:

05 June 2026

You are already at the latest version

Abstract

The bacterial proteome is a highly dynamic landscape rather than a static reflection of the genome. Recent research revealed that proteome complexity extends far beyond canonical gene annotation, with N-terminal (Nt-)proteoforms emerging as an important underexplored additional regulatory layer. These molecular variants originate from a single genetic locus through alternative translation initiation at internal or external in-frame start sites, thereby generating N-terminal heterogeneity that can influence protein stability, subcellular localization, interaction networks, and the stoichiometric assembly of multiprotein complexes. While recent advances in riboproteogenomics, N-terminomics, and computational annotation strategies have enabled proteoform mapping at single-amino acid resolution, rapid high-throughput discovery currently outpaces downstream functional characterization. This review discusses the technological advances driving Nt-proteoform discovery, including emerging ribosome profiling and proteogenomic approaches, and further evaluates strategies for the functional characterization of Nt-proteoform. Particular emphasis is placed on the transition from conventional plasmid-based heterologous expression systems toward precise genome-engineering approaches that enable selective manipulation of alternative translation initiation events within their native genomic context. Such targeted strategies are essential to bridge the gap between Nt-proteoform identification and functional understanding, ultimately uncovering how individual bacterial genomic loci can encode proteoforms with distinct and potentially polarized roles in bacterial physiology and pathogenesis.

Keywords: 
;  ;  ;  ;  ;  

1. N-Terminal Proteoforms Expand Bacterial Proteome Complexity

The central dogma of molecular biology, originally proposed by Francis Crick, describes a unidirectional flow of genetic information from DNA through RNA to functional proteins [1]. However, the discovery of alternative splicing by Sharp and Roberts demonstrated that eukaryotic gene sequences are often non-contiguous, thereby challenging the presumed linear relationship between genes and their protein products. More recently, the identification of defense-associated reverse transcriptases (DRTs) capable of synthesizing alternating poly(GT/AC) double-stranded DNA (dsDNA) revealed a protein-templated mechanism for sequence-specific DNA synthesis, further expanding the conceptual boundaries of information flow in biology [2].
In bacteria, proteome complexity is substantially amplified by alternative translation initiation events and a broad repertoire of co- and post-translational modifications, collectively generating a diverse array of proteoforms (Figure 1). Proteoforms encompass all distinct protein variants derived from a single genetic locus [3].
Among bacterial protein modifications, one of the most prevalent is the co-translational removal of the N-terminal formyl group from the initiator methionine (iMet) by peptide deformylase (PDF), which associates with the ribosomal exit tunnel during translation elongation [4,5,6]. This processing event affects more than 94% of proteins in Escherichia coli (E. coli) [7] and exposes a free N-terminal amine group that permits further N-terminal maturation. Subsequently, methionine aminopeptidase (MetAP) may remove the iMet depending on the identity of the penultimate acid residue, a process estimated to affect more than half of mature bacterial N-termini [7,8,9].
In addition, protein N-termini may undergo (partial) Nt-acetylation mediated by N-terminal acetyltransferases (NATs), a process that occurs post-translationally in bacteria [10]. NATs can irreversibly transfer an acetyl group from acetyl-coenzyme A (acetyl-CoA) to the α-amino group of either the deformylated initiator methionine or the newly exposed N-terminal residue following iMet excision. Although bacterial N-terminal acetylation is generally observed at substoichiometric levels, up to 10-15% of bacterial proteins are estimated to contain detectable N-terminal acetylation [11,12,13,14,15]. Other protein modifications, including phosphorylation, lipidation, glycosylation, and pupylation, further contribute to bacterial proteome diversification (Figure 1A). Proteoform diversity is additionally expanded through coordinated proteolytic processing events, most notably the cleavage of signal peptides by dedicated signal peptidases [16]. Prior to removal, these N-terminal sequences serve essential regulatory functions by directing subcellular localization and mediating interactions with molecular chaperones that maintain proteins in a translocation-competent state. Subsequent site-specific proteolytic processing ultimately yields the mature and localized proteoform (Figure 1B) [17].
Among these protein variants, distinct molecular forms originating from the same gene locus represent a major source of bacterial protein diversity. Beyond the previously discussed N-terminal maturation and processing events, Nt-proteoforms may additionally arise through the selection of alternative translation initiation sites (aTISs) distinct from the database-annotated or canonical translation initiation site (dbTISs) (Figure 1C). Translation from internal or external in-frame and out-of-frame start sites significantly expands the bacterial proteome and can yield Nt-proteoforms that differ in their N-terminal composition and, consequently, their biological properties [18,19]. These alternative initiation events generate functionally distinct proteins that contribute to a more complex and versatile ‘alternative’ proteome than previously recognized.
Different protein variants originating from a single gene locus can arise through a combination of transcriptional and translational mechanisms. At the transcriptional level, the use of alternative transcription start sites (TSSs) can directly influence the selection and usage of aTISs. For instance, TSSs located within annotated coding regions may generate shorter transcripts that alter the set of translation initiation sites accessible to the ribosome. Accordingly, the observation that some experimentally mapped TSS reside precisely between dbTISs and aTISs further illustrates how transcriptional boundaries can shape alternative translation initiation patterns [19].
Alongside transcription, multiple translational features further influence TIS selection. Due to bacterial transcription-translation coupling, local mRNA secondary structures or riboswitches may promote or inhibit translation initiation at specific TISs by masking or exposing ribosome binding sites (RBSs) [20,21,22,23,24,25,26]. In addition, the presence and strength of the Shine-Dalgarno (SD) sequences, together with their spacing relative to the initiating codon, strongly affect translation initiation efficiency. However, the presence of an SD sequence followed by an AUG triplet does not necessarily guarantee translation initiation [27], and because some leaderless mRNAs can be efficiently translated in the absence of an SD sequence [28], TIS selection remains highly dynamic and context-dependent.
Although AUG is the predominant start codon in E. coli, accounting for approximately 83% of annotated genes, near-cognate initiating codons such as GUG (14%) and UUG (3%) are also frequently utilized [29,30]. Similar patterns of start codon frequency have been observed across the bacterial domain [31]. Beyond these three major initiation codons, several additional non-AUG codons have been reported to support translation initiation [32]. Importantly, our previous work demonstrated that near-cognate start codons account for approximately one-quarter of newly identified Nt-proteoforms in Salmonella enterica serovar Typhimurium (S. Typhimurium) [33].
Collectively, these mechanisms demonstrate that the bacterial translatome is not merely a static reflection of genome annotation, but rather a highly dynamic and adaptable landscape. In particular, the widespread occurrence of alternative translation initiation suggests that bacterial coding capacity extends substantially beyond canonical gene models. By generating multiple Nt-proteoforms from single genetic loci, bacteria can expand their functional repertoire without increasing genome size. Consequently, N-terminal heterogeneity constitutes an important and still underexplored regulatory layer that can influence protein stability, subcellular localization, interaction networks, enzymatic activity, and multiprotein complex assembly, among other biological processes.

2. Biological Relevance of Nt-Proteoforms

Heterogeneity at the N-terminus of bacterial proteins represents an important regulatory layer that may influence subcellular localization, multiprotein complex assembly, protein stability, and ultimately protein function (Figure 2).
N-terminal variation fundamentally dictates spatial organization, as illustrated by the architecture of cyanobacterial β-carboxysomes. Within these proteinaceous microcompartments, ribulose-1,5-biphosphate carboxylase/oxygenase (Rubisco), the key enzyme of the Calvin cycle, is sequestered to optimize CO2 fixation efficiency [34]. In Synechococcus sp. strain PCC 7942, an internal GUG216 initiation codon within ccmM generates two distinct isoforms, forming an N-terminal proteoform pair consisting of the full-length CcmM58 (58 kDa) and the N-terminally truncated CcmM35 (35 kDa). This 23 kDa difference fundamentally dictates their cellular functions. CcmM58, retaining its complete N-terminal region, is localized to the inner shell of the carboxysome, where it recruits the carboxysomal carbonic anhydrase CcaA and interconnects outer shell components [35,36,37]. Conversely, the truncated CcmM35 proteoform is confined to the carboxysome lumen and functions exclusively in organizing Rubisco into a paracrystalline matrix. This example highlights how N-terminal diversity can physically segregate structural and scaffolding functions within a single microcompartment to ensure efficient CO2 fixation.
Beyond mediating spatial organization within intracellular bacterial microcompartments, Nt-proteoforms can additionally control protein targeting and secretion in bacterial virulence systems. In bacterial pathogens such as S. Typhimurium, more than 40 type III effector (T3E) proteins are secreted through the type III secretion system (T3SS) to manipulate host cell biology and promote bacterial survival, replication, and dissemination [38,39,40]. The effector SseL exists as two proteoforms: a long isoform (SseLL, 38 kDa) and a shorter isoform lacking 23 N-terminal residues (SseLS, 29 kDa) [33]. The N-terminal extension unique to SseLL contains a T3E secretion signal, thereby enabling translocation into the host cell environment. In contrast, SseLS remains confined to the bacterial cytosol and may therefore fulfill distinct intracellular functions. In this context, it is important to note that T3Es were long assumed to remain inactive within the bacterium and only adopt their functional conformations following host-cell translocation, owing to chaperone-mediated maintenance of a secretion-competent unfolded state. However, recent studies have demonstrated that certain T3Es can exert important intrabacterial activities involved in the regulation of bacterial physiology [41,42]. This heteromorphic phenomenon is further corroborated by CyaA’ (calmodulin-dependent adenylate cyclase) translocation assays demonstrating that the SseL leader sequence constitutes the minimal determinant required for host-cell entry [43]. Accordingly, several additional candidate Nt-proteoform pairs, including SsaQ, YdgA, YfhG, MgtC, OmpX, and PagC, are predicted to exhibit distinct subcellular localization among their constituting Nt-proteoform members in Salmonella [33].
Beyond influencing effector localization and host-cell targeting, Nt-proteoforms can additionally modulate the assembly and stoichiometry of larger multiprotein machineries such as the T3SS sorting platform. The T3SS injectisome itself is composed of a cytosolic sorting complex, multiple membrane-embedded structures, a hollow needle spanning the bacterial membranes, and a translocon complex located at the tip (extensively reviewed in [44] and [45]). Central to this machinery is the cytosolic sorting platform, a multiprotein complex positioned at the base of the injectisome that coordinates effector sorting and unfolding prior to secretion. Previous studies demonstrated that SpaOS (11 kDa) forms homodimers that subsequently associate with a single SpaOL (34 kDa) subunit to generate a heterotrimeric complex with a strict 2:1 stoichiometry (2 SpaOS: 1 SpaOL). Within this assembly, SpaOL acts as a central interaction hub that promotes SP pod formation [46,47]. Although deletion of SpaOS partially impairs T3SS functionality, secretion activity is not completely abolished, suggesting an ancillary yet functionally important role in T3SS-mediated protein translocation [47,48]. Consistent with this observation, deletion of the homologues Spa33S (12 kDa) (proteoform in Shigella) resulted in residual, albeit detectable, T3SS activity. Conversely, conventional in vitro studies in both Yersinia [49,50] and Shigella [51] indicate that the short SpaOS isoform is indispensable for proper T3SS function.
In addition to controlling protein localization and multiprotein complex assembly, Nt-proteoforms may also directly modulate biochemical activity and signalling behaviour, as exemplified by the chemotaxis regulator CheA. This histidine kinase functions as a central regulator of bacterial chemotaxis. Upon binding of chemical stimuli to chemoreceptors, CheA undergoes autophosphorylation, thereby initiating a signal transduction cascade that governs flagellar rotation, receptor adaptation, and methylation [52,53]. In E. coli, an internal AUG98 initiation site generates the shorter CheAS (60 kDa) isoform alongside the full-length CheAL (71 kDa) isoform. The N-terminal difference between these variants dictates both their complex assembly and biochemical activities. CheAL forms both homodimers and heterodimers with CheAS and functions exclusively as a kinase responsible for phosphorylating the response regulator CheY [54]. In contrast, CheAS lacks the critical His48 residue and instead fulfils distinct regulatory functions. Specifically, CheAS localizes as small foci at the cell poles and participates in both phosphorylating and dephosphorylating complexes [55,56]. Furthermore, CheAS interacts with and enhances the activity of CheZ-mediated dephosphorylation of phospho-CheY, while CheAL lacks this capability [56,57]. This example illustrates how N-terminal heterogeneity enables a single genetic locus to generate proteins with functionally polarized roles.
Besides influencing localization, multiprotein complex assembly, and biochemical specialization, Nt-proteoforms may also profoundly affect protein stability. Importantly, not all Nt-proteoforms differ by extensive N-terminal truncations or extensions, as many identified proteoform architectures vary by only a few amino acid residues [19,33]. Nevertheless, even these subtle N-terminal differences may drastically influence protein stability through the N-degron pathway (formerly the N-end rule pathway), in which the identity of the N-terminal residue acts as a primary determinant of proteolytic fate [58,59,60]. In bacteria, this pathway is primarily mediated by the ATP-dependent ClpAP and its adaptor ClpS, a 12-kDa Leu/N-recognin that captures substrates and delivers them to the protease machinery [61]. ClpS preferentially recognizes primary destabilizing residues such as Leu, Phe, Trp, and Tyr, whereas proteins carrying stabilizing N-terminal residues (Met, Ser, Thr, Ala, Val, and Gly) generally exhibit prolonged half-lives. Although bacterial N-degron pathways firmly establish the importance of N-terminal identity in determining protein stability, endogenous Nt-proteoform architectures differing only minimally at their N-termini have rarely been directly linked to divergent degradation kinetics. However, analogous functional consequences of subtle N-terminal variation are well established in eukaryotes, where human Nt-proteoforms differing by only a few residues can display markedly distinct stability profiles [62].

3. Mapping the Bacterial Translatome: Genome Annotation, Ribosome Profiling, and N-Terminomics

The increasing prevalence of Nt-proteoforms within the bacterial translatome reveals a regulatory layer that is frequently overlooked by traditional genome annotation approaches. Accurate delineation of this proteomic landscape increasingly relies on integrated frameworks combining computational prediction with high-resolution experimental data (Figure 3). Three complementary pillars currently underpin this strategy: computational modeling (Figure 3C), ranging from classical ab initio gene prediction algorithms to modern genomic language models (gLMs); ribosome profiling (Figure 3A), which identifies actively translated regions and TISs; and N-terminomics (Figure 3B), which provides direct mass spectrometry (MS)-based evidence of protein start sites. Together, these complementary approaches are shifting bacterial genome annotation from static gene predictions toward dynamic translatome- and proteome-aware annotation strategies [63].

3.1. Classical Genome Annotation Approaches and Their Limitations

Historically, the existence of Nt-proteoforms became apparent through observations that single gene loci could produce multiple protein variants differing at their N-termini. However, accurately defining the genomic boundaries and translation initiation events underlying this heterogeneity remained technically challenging. At the turn of the millennium, the first generation of microbial genome annotation tools emerged, most notably GLIMMER [64] and GeneMark [65]. These algorithms used probabilistic sequence models to identify coding regions within microbial genomes and represented major advances in automated gene prediction. However, early annotation tools were often computationally intensive and displayed limited accuracy in delineating complex genomic architectures, particularly TISs.
Methodological development subsequently focused on improving the precision of TIS identification, most notably through the development of the PROkaryotic Dynamic programming Gene-finding Algorithm (Prodigal) [66]. Despite their widespread use, substantial discrepancies persist among annotation platforms. Comparative analysis of GLIMMER, GeneMark, and Prodigal revealed only 70% consensus between predictions, highlighting the inherent algorithmic biases and limitations within individual tools [67]. To meet increasing high-throughput demands, automated annotation pipelines such as the TIGR Annotation Engine [68] and the Bacterial Annotation System (BASys) [69] were subsequently developed to streamline these workflows. This evolution ultimately culminated in fully automated annotation services for archaeal and bacterial genomes, including Rapid Annotations using Subsystems Technology (RAST) [70], which was later integrated into Bacterial and Viral Bioinformatics Resource Center (BV-BRC) [71]. Despite the comparatively simple gene architecture of prokaryotes relative to eukaryotes, modern annotation tools still struggle to resolve internal genomic complexity. Most annotation pipelines fail to detect multiple internal TISs (iTISs) within a single coding sequence (CDS). Widely used annotation frameworks such as the NCBI Prokaryotic Genome Annotation Pipeline (PGAP) rely heavily on homology-based methodologies, integrating Hidden Markov Models (HMMs), BlastRules, and curated Conserved Domain Database (CDD) architectures [72,73]. Currently, approximately 79% of the RefSeq proteins are annotated based on matches to a curated protein family model (PFM) [74]. Importantly, these tools frequently prioritize the longest possible ORF to maximize coding information, although this does not necessarily reflect the true bacterial translatome [63].
Although current ab initio gene prediction algorithms robustly identify canonical ORFs, their sensitivity declines substantially for non-canonical genomic elements. First, many annotation tools exclude sequences shorter than 150-300 nucleotides, thereby systematically overlooking small open reading frames (sORFs) [75,76,77]. Increasing evidence demonstrates that these small proteins participate in critical biological processes, including protein recruitment [78,79], protein stability [80,81], and multiprotein complex assembly [82,83,84], among others. Second, current annotation frameworks generally fail to identify aTISs that give rise to Nt-proteoform. Homology-driven pipelines such as PGAP frequently enforce previously annotated start sites across related genomes, thereby masking biologically relevant variation in translation initiation [73]. Finally, annotation tools prioritizing computational speed, including RAST and Prokka, often perform poorly when annotating specialized or highly divergent genes lacking close homologs.

3.2. Machine Learning and Genomics Language Models

Recent advances in machine learning have substantially reshaped genome annotation strategies. In particular, gLMs such as DNABERT and Evo 2 interpret DNA as structured sequential information, enabling the extraction of both local sequence motifs and long-range contextual relationships [85,86,87]. Compared with traditional alignment- or homology-based approaches, these models offer improved adaptability for resolving complex bacterial gene architectures, offering a more adaptive alternative to traditional genome annotation.
Various deep learning-based frameworks have been proposed for gene prediction, including conventional neural networks [88]. More recently, GeneLM, a DNABERT-based architecture, was applied to prokaryotic gene prediction using a two-stage approach involving initial CDS identification followed by precise TIS delineation across diverse bacterial species [89]. Benchmarking analyses demonstrated that GeneLM consistently outperformed classical tools like Prodigal, GeneMark, and Glimmer, as well as specialized deep learning approaches including TITER [90] and DeepGSR [91], both of which exhibit comparatively lower TIS prediction performance.
Despite these advances, the performance of gLMs remains highly dependent on the quality and diversity of their training datasets. Nevertheless, gLMs are expected to provide increasingly accurate frameworks for identifying novel gene structures and refining gene boundaries in both well-characterized and previously unannotated genomes.

3.3. Ribosome Profiling-Based Translation Initiation Site Mapping

The limitations of purely computational annotation strategies underscore the importance of experimental validation and evidence-based annotation approaches. Proteogenomic (re-)annotation efforts have therefore been widely applied to identify novel genes and to correct missing or inaccurate annotations [92,93,94,95]. In particular, the rapid evolution of riboproteogenomics has largely been driven by the advent of ribosome profiling (Ribo-seq). General Ribo-seq approaches infer translation by capturing ribosome footprints across the transcriptome, thereby substantially improving ORF detection and Nt-proteoform discovery. One of the first bacterial genome annotation approaches integrating Ribo-seq-derived translational evidence was RibosomE Profiling Assisted (Re-)AnnotaTION (REPARATION), which enabled de novo ORF delineation in prokaryotic genomes [96]. Consequently, extensive re-annotation efforts have been performed in so-called well-annotated bacteria, including E. coli, S. Typhimurium, and Bacillus subtilis [75,96,97].
To move beyond general ORF delineation and improve TIS resolution, several specialized Ribo-seq approaches have been developed to selectively enrich initiating ribosomes. One of the earliest approaches, tetracycline-inhibited ribosome profiling (TetRP), was introduced as a relatively simple and comprehensive strategy for bacterial TIS annotation [98]. Tetracycline blocks aminoacyl-tRNA entry into the ribosomal A-site and is therefore classically regarded as an elongation inhibitor. To further improve TIS resolution, Meydan et al. subsequently developed retapamulin-assisted ribosome profiling (Ribo-RET), which selectively enriches initiating ribosomes through retapamulin-mediated stalling. Using this approach, internal start codons were identified in more than one hundred E. coli genes [19]. Application of Ribo-RET in S. Typhimurium further revealed 26 candidate Nt-proteoform pairs following manual curation [33], demonstrating that in both E. coli and S. Typhimurium, the majority of identified aTISs cluster near the beginning of the CDSs.
Despite these methodological advances, precise TIS delineation remains highly challenging due to several technical and biological limitations. One major limitation stems from incomplete elongation inhibition during antibiotic treatment. Because tetracycline can interact with ribosomes throughout multiple elongation cycles, TetRP cannot reliably discriminate between elongating and initiating ribosomes, thereby limiting its sensitivity for detecting aTISs [98]. Retapamulin-based approaches face a related constraint, as signals originating from elongating ribosomes are not completely eliminated. This residual background complicates confident identification of internal start codons, particularly within highly expressed ORFs [99]. Furthermore, Ribo-RET peaks at putative initiation sites do not inherently demonstrate productive, full-length translation in the absence of antibiotic treatment, meaning that some detected events may represent abortive or non-functional initiation [100].
In addition to antibiotic-related artifacts, ribosome profiling approaches also suffer from intrinsic resolution constraints because translation initiation is inferred indirectly through ribosome footprints. Since translating ribosomes physically protect RNA fragments of approximately 22-35 nucleotides, overlapping footprints complicate precise discrimination between closely positioned TISs [101].
Moreover, the reliance of ribosome profiling on living cells can complicate the distinction between direct and indirect translational effects. Translation-targeting antibiotics and their associated uptake mechanisms may inadvertently induce generalized stress responses and widespread off-target transcriptional and translational changes that obscure native cellular physiology [102,103].
To minimize such confounding in vivo artifacts, in vitro Ribo-seq (INRI-seq) was developed as a cell-free alternative that enables translation analysis on customizable synthetic transcriptomes [104]. Although INRI-seq successfully confirmed the annotated TIS for approximately 70% of E. coli genes and validated 51 out of the 64 aTISs previously identified by in vivo Ribo-RET [19], the method remains constrained by its current library design, which captures only the first 50 codons of a CDS. Consequently, alternative TIS detection is inherently restricted to the first ~150 nucleotides of a gene, thereby limiting the identification of more downstream internal initiation events.

3.4. N-terminomics Approaches for Nt-Proteoform Discovery

Together with ribosome profiling, N-terminomics has become an important tool for high-resolution identification of prokaryotic translation initiation sites (TISs). In conventional bottom-up proteomics workflows, alternative Nt-proteoforms are often difficult to distinguish from their canonical counterparts because most internal tryptic peptides are shared. Proteoform-specific information is mainly confined to the altered N-terminal peptide and, in some cases, to the absence of peptides from the truncated N-terminal region, making dedicated N-terminal proteomics (N-terminomics) strategies essential for direct TIS identification [105,106].
Techniques such as Combined FRActional Diagonal Chromatography (COFRADIC) [107,108], Terminal Amine Isotopic Labeling (TAILS) [109], and LysN Amino Terminal Enrichment (LATE) [110] employ selective chemical derivatization and chromatographic enrichment to isolate N-terminal peptides prior to LC-MS/MS analysis. By reducing sample complexity and enriching proteoform-specific N-terminal peptides, these approaches facilitate direct detection of both canonical and alternative TISs. Comprehensive overviews of bacterial N-terminomics strategies have previously been provided by Berry et al. [106]. Recently, a deformylation-assisted N-terminomics workflow termed TRAnslation Initiation SPOTTER (TRAINSPOTTER) was developed [111]. By exploiting the presence of the nascent N-formyl group—optionally stabilized through peptide deformylase (PDF) inhibition—TRAINSPOTTER enables proteome-wide detection of nascent N-termini and thereby provides direct molecular evidence for translation initiation at single-amino acid resolution. Nevertheless, the efficacy of MS-based TIS delineation remains highly dependent on both the completeness of the underlying protein sequence databases and the applied database search strategy. The composition of the proteomic search space fundamentally defines the boundaries of novel proteoform discovery. Historically, unannotated TISs, including aTISs and initiation events originating from non-canonical start codons, frequently escaped MS-based detection because they were absent from standard reference databases. However, this limitation is not absolute and becomes particularly problematic when stringent enzymatic cleavage constraints are imposed during database searching or when non-AUG initiation events are excluded from the databases.
To circumvent these limitations, the TRAINSPOTTER workflow interrogates MS datasets against a customized six-frame translation (6-FT) database combined with semi-specific search parameters that account for iMet processing rules. This relatively unconstrained search strategy enables identification of internal and alternative N-termini, thereby providing direct physical evidence for productive translation initiation from both canonical and non-canonical start codons.
Despite these advances, current N-terminomics workflows remain limited in their ability to capture TIS-indicative N-termini of proteins undergoing signal peptide cleavage. In such cases, the original N-terminus is removed during protein translocation, effectively erasing the molecular evidence of the initial TIS and complicating comprehensive mapping of the bacterial N-terminal proteome landscape.

4. Methodological Strategies for Decoupling Bacterial N-Terminal Proteoforms

While the growing number of identified aTISs reveals a highly modular bacterial proteome, the capacity for high-throughput discovery currently exceeds downstream functional validation. Consequently, substantial gaps remain in our understanding of the precise physiological roles of individual Nt-proteoforms.

4.1. Classical Heterologous Complementation Approaches

Initial functional characterization and decoupling of bacterial Nt-proteoforms—defined as the selective expression and analysis of individual proteoforms—primarily relied on deletion of the endogenous gene followed by heterologous complementation. Within this framework, individual proteoforms are typically expressed from (inducible) plasmids to evaluate their ability to complement a mutant phenotype. This strategy provided some of the earliest functional insights into the isoforms of Initiation Factor 2 (IF2), encoded by infB. Expression of IF2α and IF2β from pEV1 and pBAD24 vectors enabled characterization of their distinct ribosomal binding properties and respective contributions to replication restart [112,113]. Similarly, plasmid-based complementation experiments demonstrated that a 46-amino acid shorter proteoform of the major penicillin-binding protein PBP-1B, encoded by mrcB, could restore thermosensitive growth defects in E. coli [114]. Another classical example is represented by the ClpB proteoform pair, originally identified due to pronounced molecular weight differences and subsequently decoupled using pBS or pGB2-based expression systems [115,116,117].
Despite their widespread utility in gain-of-function studies, heterologous expression systems frequently fail to recapitulate the native biological context. Plasmid-based protein overexpression often drives protein concentrations far beyond physiological levels, potentially inducing molecular crowding, aberrant subcellular localization, inclusion body (IB) formation, or artificial protein-protein interactions (PPIs). For example, overexpression of periplasmic proteins may saturate Sec-translocon capacity, thereby impairing the secretion of endogenous proteins and promoting cytoplasmic protein aggregation [118]. Similarly, excessive production of aggregation-prone proteins can promote artificial non-physiological PPIs or sequester essential factors into toxic aggregates [119]. Such artifacts are particularly disruptive for multiprotein complexes and signaling pathways, where strict stoichiometric balance is essential for proper function. Furthermore, proteoform expression is often tightly coupled to specific growth phases or environmental stimuli [115,120,121,122,123,124], whereas inducible plasmid systems frequently bypass the native transcriptional and post-transcriptional regulatory architecture. Finally, episomal maintenance often requires continuous antibiotic selection, imposing an additional metabolic burden that may fundamentally alter cellular physiology [125].

4.2. Endogenous Chromosomal Manipulation Strategies

To overcome the limitations associated with heterologous overexpression systems, increasing emphasis has been placed on endogenous genome engineering strategies. By manipulating genes directly within their native chromosomal context, proteoform expression remains under endogenous promoter and regulatory control. This approach can be combined with small high-sensitivity peptide tags, such as HiBiT or 3 x FLAG, to facilitate proteoform detection at physiological expression levels while avoiding disruptive overexpression [126,127].
The S. Typhimurium protein SpaO represents a landmark example of Nt-proteoform decoupling via endogenous genome manipulation. The spaO gene contains an internal GTG203 initiation codon to yield two distinct proteoforms: the full-length SpaOL and the truncated SpaOS proteoforms. To dissect their individual functions, Lara-Tejero et al. employed R6K-based allelic exchange to steer selective Nt-proteoform expression [47,128]. This suicide plasmid-based allelic exchange strategy enables precise chromosomal mutagenesis through homologous recombination. Following counterselection using markers such as sacB, a second recombination event removes the plasmid backbone and leaves behind the desired scarless mutation. Specifically, this approach was used to mutate the internal GTG203 start codon to the non-initiating GCG codon, successfully eliminating SpaOS production and enabling selective functional interrogation of SpaOL. Additionally, loss-of-function (∆spaO mutant) phenotypes were complemented by expressing a wild-type spaO copy in trans from the pSB4545 plasmid.

4.3. (Multiplex) Recombineering Approaches

More recent developments have implemented multiplexed recombineering approaches to precisely manipulate Nt-proteoform expression while simultaneously enabling high-sensitivity detection through integration of peptide tags such as the luminescent HiBiT tag [33]. This 11-amino acid peptide (VSGWRLFKKIS) associates with high affinity to the complementary Large BiT (LgBiT) subunit to form an active luciferase, thereby enabling highly sensitive bioluminescent quantification. A cornerstone of modern bacterial genome engineering is the λ-Red recombineering system, which mediates homologous recombination through the phage-derived Gam, Bet, and Exo proteins. This platform enables precise genomic engineering using either single-stranded DNA (ssDNA) oligonucleotides or dsDNA repair templates [129].
The use of ssDNA oligonucleotides, commonly referred to as oligo-mediated allelic replacement (OMAR) [130], is particularly well-suited for Nt-proteoform decoupling because it permits seamless conversion of individual TIS codons into non-initiating codons without introducing selection markers that could perturb operon polarity. Maintaining operon integrity is biologically important in prokaryotes, where genes are often tightly clustered and co-transcribed. Residual recombination scars or selection markers may disrupt important transcriptional and translational regulatory sequences, including promoters, terminators, and RBSs [131,132,133]. Consequently, OMAR is particularly well suited for TIS mutagenesis while preserving the native physiological context of the targeted genetic locus.
Nevertheless, careful oligonucleotide design remains essential when modifying start codons. Near-cognate codons such as ATA should generally be avoided because they may still permit low-level “leaky” translation initiation [32,134]. Instead, codons lacking initiation potential are preferred, ideally while preserving the encoded amino acid or introducing only conservative amino acid substitutions where possible. For example, CTT can serve as a reliable non-initiating alternative in certain leucine-compatible contexts [32]. Furthermore, because closely spaced TISs may overlap with SD-like sequences or other cis-regulatory elements, even single nucleotide substitutions may inadvertently disrupt translational coupling or generate unintended polar effects across polycistronic operons [135,136,137]. Predictive tools such as the Genetic Systems Calculator [138] therefore provide valuable support for evaluating the translational consequences of engineered mutations.
Beyond design constraints, execution of precise tag-free TIS modifications poses substantial screening challenges. Because OMAR typically generates scarless and non-selectable mutations, the identification of correctly edited clones becomes experimentally challenging. Reported baseline efficiencies for OMAR vary considerably, typically ranging from 6% to 20% depending on the genetic background, mutation type, and oligonucleotide modifications [130]. To overcome these low baseline frequencies in the absence of direct selection, co-selection MAGE (Cos-MAGE) can be implemented. By co-targeting a nearby selectable chromosomal marker with an additional ssDNA oligo, CoS-MAGE enriches for subpopulations containing both mutations, thereby substantially increasing the allelic replacement frequency (ARF) of the unselectable TIS mutation [139].
To identify successful mutants within these enriched populations, alternative screening approaches such as allele-specific colony PCR (ASC-PCR) are required [140]. While OMAR efficiencies have historically been estimated using selectable chromosomal markers that generate easily screenable phenotypes on selective media—thereby enabling ARF calculation as a percentage of the surviving population [141]—our previous work demonstrated that ASC-PCR enables determination of actual TIS conversion frequencies [142]. Broader optimization strategies for recombineering efficiency have been comprehensively reviewed elsewhere [130].
Overall, multiplexed recombineering approaches have previously been shown to effectively steer Nt-proteoform expression at the genomic level while simultaneously allowing sensitive detection of endogenous expression [33]. Decoupling Nt-proteoform expression is particularly relevant for validation purposes when multiple TISs are located in close proximity within the bacterial genome. Furthermore, this approach enables characterization of conditionally expressed proteoforms, as exemplified by the S. Typhimurium FruK and E. coli RNA polymerase Sigma S (σS) proteoform pairs [33,143].

4.4. CRISPR-Based Nt-Proteoform Engineering

Since the development of clustered regularly interspaced short palindromic repeats (CRISPR)-based genome engineering, CRISPR/Cas systems have become widely implemented for bacterial genome manipulation [144,145,146,147]. These approaches rely on guide RNA (gRNA)-directed Cas-mediated cleavage at target protospacer sequences adjacent to protospacer-adjacent motifs (PAMs), thereby enabling counterselection of nonedited wild-type cells. Importantly, CRISPR-based strategies can eliminate the need for selectable markers, classical counterselection systems, or dedicated recombineering platforms. Many early CRISPR/Cas-based bacterial editing strategies required the introduction of additional mutations within the protospacer or PAM sequence to prevent repeated Cas-mediated cleavage following successful genome editing. While initial CRISPR-Cas9-mediated oligonucleotide-directed mutagenesis approaches achieved relatively high editing efficiencies for two-to three-base mutations, introduction of single-nucleotide mutations proved substantially more challenging [148,149]. Such single-base edits were often obtained at frequencies below 3%, largely due to the mismatch tolerance of the CRISPR/Cas system [150]. However, subsequent optimization of sgRNA design strategies dramatically improved single-nucleotide editing efficiencies to between 36% and 95% [150,151].
More recently, the RECKLEEN platform combines λ-Red recombineering with CRISPR-Cas9-mediated counterselection, achieving editing efficiencies approaching 100% in Klebsiella [152]. Nevertheless, CRISPR-mediated genome engineering still faces several important limitations, including off-target mutagenesis and the occurrence of ‘escaper’ colonies containing wild-type clones or unintended genomic alterations [153]. In addition, excessive Cas9 expression may exert cytotoxic effects, thereby limiting the broader applicability of certain CRISPR-based editing systems [153]. Despite these limitations, increasingly precise CRISPR-based genome engineering technologies hold potential for Nt-proteoform research by enabling selective manipulation of endogenous proteoform expression within native chromosomal contexts.

5. Functional Characterization of Nt-Proteoforms

The rapid expansion of computational TIS prediction, ribosome profiling, and N-terminomics datasets has substantially increased the number of candidate bacterial Nt-proteoforms. Nevertheless, functional validation and mechanistic characterization of these proteoforms remain comparatively limited. Recent advances in endogenous genome engineering now enable the selective decoupling of individual proteoforms within their native chromosomal context, thereby facilitating direct investigation of their physiological relevance. Such strategies provide opportunities to study proteoform-specific effects of cellular physiology, subcellular localization, interaction networks, stability, and conditional regulation.
Recently, multiplexed recombineering toolkits have been extended toward genomic integration of promiscuous biotin ligases (PBLs), thereby enabling proximity-dependent biotinylation approaches such as BioID [142]. These systems exploit engineered ligases, including BioID/BirA* [154], BioID2 [155], TurboID, and miniTurboID [156] to map local protein interaction environments, although their application in prokaryotes remains limited [157,158]. Resolving the specific proxeomes of decoupled Nt-proteoforms may provide important insights into how N-terminal variation influences protein interactions, complex assembly, and functional specialization (Figure 4) [159,160].
Nt-proteoform variation may additionally influence protein stability. In eukaryotes, altered N-termini have been shown to substantially affect proteoform half-life [62]. Protein stability is commonly investigated using pulse-chase approaches such as pulse SILAC (pSILAC) combined with COFRADIC, which enables temporal monitoring of proteoform degradation dynamics [108,111,161,162,163,164]. Although initially developed for eukaryotic cells, recent methodological adaptations have enabled highly efficient SILAC labeling in bacteria, thereby expanding the applicability of these approaches to prokaryotic proteome dynamics [165]. Interestingly, a recent E. coli study suggested that N-terminal identity may not always represent the dominant determinant of protein stability in vivo [166]. However, unstable proteoforms may evade detection entirely due to rapid turnover, complicating the interpretation of such datasets.
Collectively, the combination of precise endogenous genome engineering with complementary proteomic and interaction-mapping technologies provides an emerging framework for systematic functional characterization of bacterial Nt-proteoforms.

6. Discussion

Although Nt-proteoforms have now been identified across phylogenetically diverse bacterial species, their overall prevalence and evolutionary distribution across bacterial phyla remain largely unresolved. Current evidence is still largely limited to a relatively small number of experimentally investigated lineages, spanning taxonomically distant groups such as Streptomyces and Cyanobacteria [35,167]. However, recent riboproteogenomic advances are increasingly enabling more systematic and large-scale identification of bacterial Nt-proteoforms. By revealing a far more dynamic and heterogeneous proteomic landscape than previously appreciated, these methods have fundamentally transformed our understanding of the bacterial translatome. In particular, the widespread occurrence of aTIS events demonstrates that single bacterial loci can generate multiple Nt-proteoforms with potentially distinct biological properties. However, the rapid expansion of high-throughput discovery datasets is currently outpacing downstream functional characterization, leaving the physiological relevance of many candidate Nt-proteoforms unresolved.
A central challenge in the field is distinguishing functional Nt-proteoforms from pervasive or non-productive translational noise. Ribosome profiling approaches such as Ribo-RET provide highly sensitive snapshots of ribosome occupancy and initiation-site selection, but fundamentally measure ribosome-protected mRNA fragments, rather than stable protein products. Consequently, not every detected initiation event necessarily yields a biologically relevant proteoform. Nevertheless, our previous work demonstrated that aggregate Ribo-seq and Ribo-RET signals correlate strongly with experimentally determined protein abundances. This correlation indicates that ribosome occupancy metrics can serve as valuable quantitative proxies for translation output [33], and that signal intensity, reproducibility, positional conservation, sequence context, and integration with complementary proteomic evidence serve as critical parameters for prioritizing candidate functional Nt-proteoforms.
Accordingly, orthogonal peptide-level validation remains essential. N-terminomics approaches such as TRAINSPOTTER provide direct molecular evidence for productive translation initiation by experimentally identifying proteoform-specific N-termini [111]. The integration of quantitative ribosome profiling with high-confidence N-terminal proteomics, therefore represents a particularly powerful framework for distinguishing genuine proteoforms from pervasive translational noise and for delineating bacterial translation initiation events at single-amino-acid resolution.
From an evolutionary perspective, alternative translation initiation likely represents an efficient strategy for expanding bacterial functional capacity without increasing genome size. Bacterial genomes are highly compact and coding dense, with approximately 90% of genomic DNA dedicated to coding regions [168]. Within this constrained genomic architecture, the generation of multiple Nt-proteoforms from a single transcript provides a mechanism to diversify protein functionality while minimizing the genetic burden associated with gene duplication. In this regard, Nt-proteoform generation may represent an efficient strategy for expanding proteomic and regulatory complexity in streamlined prokaryotic systems.
Beyond coding efficiency, alternative translation initiation may additionally provide important regulatory and kinetic advantages. Because transcription and translation are tightly coupled in bacteria, modulation of translation initiation can enable highly rapid proteomic adaptation without requiring de novo transcriptional programs [169]. Changes in local mRNA structure, ribosome binding site accessibility, or environmental cues may rapidly shift TIS usage and thereby alter proteoform production under changing growth conditions. Such mechanisms could facilitate fast post-transcriptional remodelling of the bacterial proteome during environmental, thermal, nutritional, or host-associated stress responses. Consistent with this concept, several characterized proteoform-encoding genes, including if2, clpB, safA, and σS, display condition-dependent expression or specialized stress-associated functionality [122,170,171].
Alternative translation initiation may also contribute to stoichiometric regulation of multi-protein complexes. In the T3SS sorting platform of S. Typhimurium, defined ratios between long and short SsaQ and SpaO proteoforms are required for proper injectisome assembly and functionality [47,48,172]. Encoding multiple structural variants within a single transcriptional unit may therefore provide a robust mechanism to coordinate relative subunit abundance while minimizing the need for additional transcriptional regulatory layers.
In addition, Nt-proteoform diversity may influence protein stability and conditional protein turnover. Although bacterial N-degron pathways clearly establish the importance of N-terminal identity in proteolytic targeting, endogenous Nt-proteoforms differing only minimally at their N-termini have thus far rarely been systematically linked to distinct degradation kinetics. It nevertheless remains highly plausible that certain alternative proteoforms evolved as transient or condition-specific molecular variants with specialized accessory functions and tightly regulated half-lives.
Ultimately, systematic functional characterization of bacterial Nt-proteoforms will require a transition away from artifact-prone plasmid overexpression systems toward precise endogenous genome engineering approaches. The use of advanced genome engineering technologies now enables selective manipulation of individual proteoforms at endogenous levels. Combined with complementary approaches, including proximity labeling, quantitative proteomics, and proteoform stability profiling, these tools may facilitate bridging the gap between high-throughput Nt-proteoform discovery and deep mechanistic biological understanding. Collectively, continued integration of riboproteogenomics, N-terminomics, and endogenous genome engineering is expected to substantially refine our understanding of bacterial proteome organization and regulatory complexity. Resolving how Nt-proteoforms contribute to bacterial physiology, adaptation, and pathogenesis will likely represent an important frontier in future bacterial proteome biology.

Author Contributions

Conceptualization, V.S. and P.V.D.; writing-original draft preparation, V.S.; writing—review and editing V.S. and P.V.D.; visualization, V.S.; supervision, P.V.D.; project administration, P.V.D.; funding acquisition, P.V.D. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Research Foundation – Flanders (FWO-Vlaanderen; project no. G088726N), awarded to P.V.D., and the Special Research Fund (BOF) of Ghent University (reference no. BOF25/CDV/006).

Institutional Review Board Statement

Not applicable.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
6-FT Six-frame translation
ARF Allelic replacement frequency
ASC-PCR Allele-specific colony PCR
BASys Bacterial Annotation system
BV-BRC Bacterial and Viral Bioinformatics Resource Center
CDD Conserved Domain Database
CDS Coding sequence
COFRADIC Combined FRActional Diagonal Chromatography
Cos-MAGE Co-selection MAGE
CRISPR Clustered Regularly Interspaced Short Palindromic Repeats
CyaA Calmodulin-dependent adenylate cyclase
dbTIS Database-annotated translation initiation site
DRTs Defense-associated reverse transcriptases
dsDNA Double stranded DNA
E. coli Escherichia coli
gLM Genomic language model
gRNA guide RNA
HMM Hidden Markov Model
IB Inclusion Body
IF2 Initiation Factor 2
iMet Initiator methionine
INRI-seq in vitro Ribo-seq
iTIS Internal translation initiation site
kDa Kilodalton
LATE LysN Amino Terminal Enrichment
LC-MS/MS Liquid Chromatography-Tandem Mass Spectrometry
LgBiT Large BiT
MAGE Multiplex Automated Genome Engineering
MetAP Methionine aminopeptidase
mRNA Messenger RNA
MS Mass spectrometry
NAT N-terminal acetyl transferase
N-terminomics N-terminal proteomics
Nt-proteoforms N-terminal proteoforms
OMAR Oligo-mediated allelic replacement
ORF Open reading frame
PAM Protospacer-adjacent motifs
PBL Promiscuous biotin ligase
PDF Peptide deformylase
PFM Protein family model
PGAP Prokaryotic Genome Annotation Pipeline
PPIs Protein-protein interactions
Prodigal Prokaryotic Dynamic programming Gene-finding Algorithm
pSILAC Pulse Stable isotope labeling by amino acids in cell culture
RAST Rapid Annotations using subsystems Technology
RECKLEEN Recombineering/CRISPR-based KLebsiella Engineering for Efficient Nucleotide editing
REPARATION RibosomE Profiling Assisted (Re-)AnnotaTION
Ribo-RET Retapamulin-assisted ribosome profiling
Ribo-seq Ribosome profiling
Rubisco Ribulose-1,5-biphosphate carboxylase/oxygenase
S. Typhimurium Salmonella enterica serovar Typhimurium
SD Shine-Dalgarno
sgRNA Single guide RNA
sORF Small open reading frame
SP Signal peptidase
SP Sorting platform
T3E Type III secretion system effector
T3SS Type III secretion system
TAILS Terminal Amine Isotopic Labeling
TetRP Tetracycline-inhibited ribosome profiling
TRAINSPOTTER TRAnslation Initiation SPOTTER
TSS Transcription start site

References

  1. Crick, F. Central Dogma of Molecular Biology. Nature 1970, 227, 561–563. [Google Scholar] [CrossRef]
  2. Deng, P.; Lee, H.; Armijo, C.; Wang, H.; Gao, A. Protein-templated synthesis of dinucleotide repeat DNA by an antiphage reverse transcriptase. Science 2026, 0, eaed1656. [Google Scholar] [CrossRef]
  3. Smith, L.M.; Kelleher, N.L.; Linial, M.; Goodlett, D.; Langridge-Smith, P.; Ah Goo, Y.; Safford, G.; Bonilla*, L.; Kruppa, G.; Zubarev, R.; et al. Proteoform: a single term describing protein complexity. Nat. Methods 2013, 10, 186–187. [Google Scholar] [CrossRef]
  4. Bingel-Erlenmeyer, R.; Kohler, R.; Kramer, G.; Sandikci, A.; Antolić, S.; Maier, T.; Schaffitzel, C.; Wiedmann, B.; Bukau, B.; Ban, N. A peptide deformylase–ribosome complex reveals mechanism of nascent chain processing. Nature 2008, 452, 108–111. [Google Scholar] [CrossRef]
  5. Giglione, C.; Fieulaine, S.; Meinnel, T. N-terminal protein modifications: Bringing back into play the ribosome. Biochimie 2015, 114, 134–146. [Google Scholar] [CrossRef] [PubMed]
  6. Meinnel, T.; Mechulam, Y.; Blanquet, S. Methionine as translation start signal: A review of the enzymes of the pathway in Escherichia coli. Biochimie 1993, 75, 1061–1075. [Google Scholar] [CrossRef]
  7. Bienvenut, W.V.; Giglione, C.; Meinnel, T. Proteome-wide analysis of the amino terminal status of Escherichia coli proteins at the steady-state and upon deformylation inhibition. PROTEOMICS 2015, 15, 2503–2518. [Google Scholar] [CrossRef]
  8. Solbiati, J.; Chapman-Smith, A.; Miller, J.L.; Miller, C.G.; Cronan, J.E. Processing of the N termini of nascent polypeptide chains requires deformylation prior to methionine removal Edited by M. Gottesman. J. Mol. Biol. 1999, 290, 607–614. [Google Scholar] [CrossRef]
  9. Frottin, F.; Martinez, A.; Peynot, P.; Mitra, S.; Holz, R.C.; Giglione, C.; Meinnel, T. The Proteomics of N-terminal Methionine Cleavage. Mol. Cell. Proteom. 2006, 5, 2336–2349. [Google Scholar] [CrossRef] [PubMed]
  10. VanDrisse, C.M.; Escalante-Semerena, J.C. Protein Acetylation in Bacteria. Annu. Rev. Microbiol. 2019, 73, 111–132. [Google Scholar] [CrossRef] [PubMed]
  11. Thompson, C.R.; Champion, M.M.; Champion, P.A. Quantitative N-Terminal Footprinting of Pathogenic Mycobacteria Reveals Differential Protein Acetylation. J. Proteome Res. 2018, 17, 3246–3258. [Google Scholar] [CrossRef]
  12. Christensen, D.G.; Baumgartner, J.T.; Xie, X.; Jew, K.M.; Basisty, N.; Schilling, B.; Kuhn, M.L.; Wolfe, A.J. Mechanisms, Detection, and Relevance of Protein Acetylation in Prokaryotes. mBio 2019, 10. [Google Scholar] [CrossRef] [PubMed]
  13. Bonissone, S.; Gupta, N.; Romine, M.; Bradshaw, R.A.; Pevzner, P.A. N-terminal Protein Processing: A Comparative Proteogenomic Analysis*. Mol. Cell. Proteom. 2013, 12, 14–28. [Google Scholar] [CrossRef]
  14. Ouidir, T.; Jarnier, F.; Cosette, P.; Jouenne, T.; Hardouin, J. Characterization of N-terminal protein modifications in Pseudomonas aeruginosa PA14. J. Proteom. 2015, 114, 214–225. [Google Scholar] [CrossRef]
  15. Jones, J.D.; O’Connor, C.D. Protein acetylation in prokaryotes. PROTEOMICS 2011, 11, 3012–3022. [Google Scholar] [CrossRef]
  16. Hegde, R.S.; Bernstein, H.D. The surprising complexity of signal sequences. Trends Biochem. Sci. 2006, 31, 563–571. [Google Scholar] [CrossRef]
  17. Kaushik, S.; He, H.; Dalbey, R.E. Bacterial Signal Peptides- Navigating the Journey of Proteins. Front. Physiol. 2022, 13–2022. [Google Scholar] [CrossRef]
  18. Huvet, M.; Stumpf, M.P.H. Overlapping genes: a window on gene evolvability. BMC Genom. 2014, 15, 721. [Google Scholar] [CrossRef] [PubMed]
  19. Meydan, S.; Marks, J.; Klepacki, D.; Sharma, V.; Baranov, P.V.; Firth, A.E.; Margus, T.; Kefi, A.; Vázquez-Laslop, N.; Mankin, A.S. Retapamulin-Assisted Ribosome Profiling Reveals the Alternative Bacterial Proteome. Mol. Cell 2019, 74, 481–493.e486. [Google Scholar] [CrossRef]
  20. Altuvia, S.; Zhang, A.; Argaman, L.; Tiwari, A.; Storz, G. The Escherichia coli OxyS regulatory RNA represses fhlA translation by blocking ribosome binding. EMBO J. 1998, 17, 6069–6075. [Google Scholar] [CrossRef] [PubMed]
  21. Majdalani, N.; Cunning, C.; Sledjeski, D.; Elliott, T.; Gottesman, S. DsrA RNA regulates translation of RpoS message by an anti-antisense mechanism, independent of its action as an antisilencer of transcription. Proc. Natl. Acad. Sci. 1998, 95, 12462–12467. [Google Scholar] [CrossRef]
  22. Storz, G.; Opdyke, J.A.; Zhang, A. Controlling mRNA stability and translation with small, noncoding RNAs. Curr. Opin. Microbiol. 2004, 7, 140–144. [Google Scholar] [CrossRef]
  23. Winkler, W.; Nahvi, A.; Breaker, R.R. Thiamine derivatives bind messenger RNAs directly to regulate bacterial gene expression. Nature 2002, 419, 952–956. [Google Scholar] [CrossRef]
  24. Mandal, M.; Breaker, R.R. Gene regulation by riboswitches. Nat. Rev. Mol. Cell Biol. 2004, 5, 451–463. [Google Scholar] [CrossRef] [PubMed]
  25. Moine, H.; Romby, P.; Springer, M.; Grunberg-Manago, M.; Ebel, J.-P.; Ehresmann, B.; Ehresmann, C. Escherichia coli threonyl-tRNA synthetase and tRNAThr modulate the binding of the ribosome to the translational initiation site of the ThrS mRNA. J. Mol. Biol. 1990, 216, 299–310. [Google Scholar] [CrossRef]
  26. Babitzke, P.; Baker, C.S.; Romeo, T. Regulation of Translation Initiation by RNA Binding Proteins. Annu. Rev. Microbiol. 2009, 63, 27–44. [Google Scholar] [CrossRef]
  27. Perham, R.N. Structural aspects of biomolecular recognition and self-assembly. Biosens. Bioelectron. 1994, 9, 753–760. [Google Scholar] [CrossRef]
  28. Scharff, L.B.; Childs, L.; Walther, D.; Bock, R. Local Absence of Secondary Structure Permits Translation of mRNAs that Lack Ribosome-Binding Sites. PLoS Genet. 2011, 7, e1002155. [Google Scholar] [CrossRef]
  29. Nirenberg, M.; Leder, P. RNA Codewords and Protein Synthesis. Science 1964, 145, 1399–1407. [Google Scholar] [CrossRef] [PubMed]
  30. Blattner, F.R.; Plunkett, G.; Bloch, C.A.; Perna, N.T.; Burland, V.; Riley, M.; Collado-Vides, J.; Glasner, J.D.; Rode, C.K.; Mayhew, G.F.; et al. The Complete Genome Sequence of Escherichia coli K-12. Science 1997, 277, 1453–1462. [Google Scholar] [CrossRef]
  31. Villegas, A.; Kropinski, A.M. An analysis of initiation codon utilization in the Domain Bacteria – concerns about the quality of bacterial genome annotation. Microbiology 2008, 154, 2559–2661. [Google Scholar] [CrossRef]
  32. Hecht, A.; Glasgow, J.; Jaschke, P.R.; Bawazer, L.A.; Munson, M.S.; Cochran, J.R.; Endy, D.; Salit, M. Measurements of translation initiation from all 64 codons in E. coli. Nucleic Acids Res. 2017, 45, 3615–3626. [Google Scholar] [CrossRef] [PubMed]
  33. Fijalkowski, I.; Snauwaert, V.; Damme, P.V. Proteins à la carte: riboproteogenomic exploration of bacterial N-terminal proteoform expression. mBio 2024, 15, e00333–00324. [Google Scholar] [CrossRef]
  34. Cannon, G.C.; Bradburne, C.E.; Aldrich, H.C.; Baker, S.H.; Heinhorst, S.; Shively, J.M. Microcompartments in Prokaryotes: Carboxysomes and Related Polyhedra. Appl. Environ. Microbiol. 2001, 67, 5351–5361. [Google Scholar] [CrossRef] [PubMed]
  35. Long, B.M.; Tucker, L.; Badger, M.R.; Price, G.D. Functional Cyanobacterial β-Carboxysomes Have an Absolute Requirement for Both Long and Short Forms of the. Plant Physiol. 2010, 153, 285–293. [Google Scholar] [CrossRef]
  36. Long, B.M.; Badger, M.R.; Whitney, S.M.; Price, G.D. Analysis of Carboxysomes from Synechococcus PCC7942 Reveals Multiple Rubisco Complexes with Carboxysomal Proteins CcmM and CcaA. J. Biol. Chem. 2007, 282, 29323–29335. [Google Scholar] [CrossRef]
  37. Cot, S.S.-W.; So, A.K.-C.; Espie, G.S. A Multiprotein Bicarbonate Dehydration Complex Essential to Carboxysome Function in Cyanobacteria. J. Bacteriol. 2008, 190, 936–945. [Google Scholar] [CrossRef]
  38. Van Damme, P.; Jonckheere, V.; Simoens, L. Systematic real-time profiling of Salmonella type III effector translocation provides quantitative resolution of the T3SS-1/T3SS-2 secretion dichotomy; 2026. [Google Scholar]
  39. Worley, M.J. Salmonella Type III Secretion System Effectors. Int. J. Mol. Sci. 2025, 26, 2611. [Google Scholar] [CrossRef]
  40. Pillay, T.D.; Hettiarachchi, S.U.; Gan, J.; Diaz-Del-Olmo, I.; Yu, X.-J.; Muench, J.H.; Thurston, T.L.M.; Pearson, J.S. Speaking the host language: how Salmonella effector proteins manipulate the host. Microbiology 2023, 169. [Google Scholar] [CrossRef] [PubMed]
  41. El Qaidi, S.; Scott, N.E.; Hays, M.P.; Geisbrecht, B.V.; Watkins, S.; Hardwidge, P.R. An intra-bacterial activity for a T3SS effector. Sci. Rep. 2020, 10, 1073. [Google Scholar] [CrossRef] [PubMed]
  42. Hasan, M.K.; El Qaidi, S.; Hardwidge, P.R. The T3SS Effector Protease NleC Is Active within Citrobacter rodentium. Pathogens 2021, 10, 589. [Google Scholar] [CrossRef]
  43. Niemann, G.S.; Brown, R.N.; Mushamiri, I.T.; Nguyen, N.T.; Taiwo, R.; Stufkens, A.; Smith, R.D.; Adkins, J.N.; McDermott, J.E.; Heffron, F. RNA Type III Secretion Signals That Require Hfq. J. Bacteriol. 2013, 195, 2119–2125. [Google Scholar] [CrossRef]
  44. Büttner, D. Protein export according to schedule: architecture, assembly, and regulation of type III secretion systems from plant- and animal-pathogenic bacteria. Microbiol. Mol. Biol. Rev. 2012, 76, 262–310. [Google Scholar] [CrossRef] [PubMed]
  45. Abrusci, P.; McDowell, M.A.; Lea, S.M.; Johnson, S. Building a secreting nanomachine: a structural overview of the T3SS. Curr. Opin. Struct. Biol. 2014, 25, 111–117. [Google Scholar] [CrossRef] [PubMed]
  46. Morita-Ishihara, T.; Ogawa, M.; Sagara, H.; Yoshida, M.; Katayama, E.; Sasakawa, C. Shigella Spa33 Is an Essential C-ring Component of Type III Secretion Machinery. J. Biol. Chem. 2006, 281, 599–607. [Google Scholar] [CrossRef]
  47. Lara-Tejero, M.; Qin, Z.; Hu, B.; Butan, C.; Liu, J.; Galán, J.E. Role of SpaO in the assembly of the sorting platform of a Salmonella type III secretion system. PLoS Pathog. 2019, 15, e1007565. [Google Scholar] [CrossRef]
  48. Soto, J.E.; Wang, T.; Galán, J.E.; Lara-Tejero, M. Interplay between SpaO variants shapes the architecture of the Salmonella type III secretion sorting platform. mBio 2026, 17, e00155–00126. [Google Scholar] [CrossRef]
  49. Bzymek, K.P.; Hamaoka, B.Y.; Ghosh, P. Two Translation Products of Yersinia yscQ Assemble To Form a Complex Essential to Type III Secretion. Biochemistry 2012, 51, 1669–1677. [Google Scholar] [CrossRef] [PubMed]
  50. Diepold, A.; Kudryashev, M.; Delalez, N.J.; Berry, R.M.; Armitage, J.P. Composition, Formation, and Regulation of the Cytosolic C-ring, a Dynamic Component of the Type III Secretion Injectisome. PLoS Biol. 2015, 13, e1002039. [Google Scholar] [CrossRef]
  51. McDowell, M.A.; Marcoux, J.; McVicker, G.; Johnson, S.; Fong, Y.H.; Stevens, R.; Bowman, L.A.H.; Degiacomi, M.T.; Yan, J.; Wise, A.; et al. Characterisation of Shigella Spa33 and Thermotoga FliM/N reveals a new model for C-ring assembly in T3SS. Mol. Microbiol. 2016, 99, 749–766. [Google Scholar] [CrossRef]
  52. Wadhams, G.H.; Armitage, J.P. Making sense of it all: bacterial chemotaxis. Nat. Rev. Mol. Cell Biol. 2004, 5, 1024–1037. [Google Scholar] [CrossRef]
  53. Levit, M.N.; Stock, J.B. Receptor Methylation Controls the Magnitude of Stimulus-Response Coupling in Bacterial Chemotaxis. J. Biol. Chem. 2002, 277, 36760–36765. [Google Scholar] [CrossRef]
  54. Hess, J.F.; Oosawa, K.; Kaplan, N.; Simon, M.I. Phosphorylation of three proteins in the signaling pathway of bacterial chemotaxis. Cell 1988, 53, 79–87. [Google Scholar] [CrossRef]
  55. Wolfe, A.J.; Stewart, R.C. The short form of the CheA protein restores kinase activity and chemotactic ability to kinase-deficient mutants. Proc. Natl. Acad. Sci. 1993, 90, 1518–1522. [Google Scholar] [CrossRef]
  56. Wang, H.; Matsumura, P. Phosphorylating and dephosphorylating protein complexes in bacterial chemotaxis. J. Bacteriol. 1997, 179, 287–289. [Google Scholar] [CrossRef] [PubMed]
  57. O’Conno, C.; Matsumura, P. The Accessibility of Cys-120 in CheAS Is Important for the Binding of CheZ and Enhancement of CheZ Phosphatase Activity. Biochemistry 2004, 43, 6909–6916. [Google Scholar] [CrossRef]
  58. Tobias, J.W.; Shrader, T.E.; Rocap, G.; Varshavsky, A. The N-End Rule in Bacteria. Science 1991, 254, 1374–1377. [Google Scholar] [CrossRef] [PubMed]
  59. Dougan, D.A.; Truscott, K.N.; Zeth, K. The bacterial N-end rule pathway: expect the unexpected. Mol. Microbiol. 2010, 76, 545–558. [Google Scholar] [CrossRef] [PubMed]
  60. Heo, A.J.; Kim, S.B.; Kwon, Y.T.; Ji, C.H. The N-degron pathway: From basic science to therapeutic applications. Biochim. Et. Biophys. Acta (BBA) -Gene Regul. Mech. 2023, 1866, 194934. [Google Scholar] [CrossRef]
  61. Varshavsky, A. N-degron and C-degron pathways of protein degradation. Proc. Natl. Acad. Sci. 2019, 116, 358–366. [Google Scholar] [CrossRef]
  62. Gawron, D.; Ndah, E.; Gevaert, K.; Van Damme, P. Positional proteomics reveals differences in N-terminal proteoform stability. Mol. Syst. Biol. 2016, 12, MSB156662. [Google Scholar] [CrossRef] [PubMed]
  63. Fijalkowska, D.; Fijalkowski, I.; Willems, P.; Van Damme, P. Bacterial riboproteogenomics: the era of N-terminal proteoform existence revealed. FEMS Microbiol. Rev. 2020, 44, 418–431. [Google Scholar] [CrossRef]
  64. Salzberg, S.L.; Delcher, A.L.; Kasif, S.; White, O. Microbial gene identification using interpolated Markov models. Nucleic Acids Res. 1998, 26, 544–548. [Google Scholar] [CrossRef]
  65. Borodovsky, M.; Rudd, K.E.; Koonin, E.V. Intrinsic and extrinsic approaches for detecting genes in a bacterial genome. Nucleic Acids Res. 1994, 22, 4756–4767. [Google Scholar] [CrossRef] [PubMed]
  66. Hyatt, D.; Chen, G.-L.; LoCascio, P.F.; Land, M.L.; Larimer, F.W.; Hauser, L.J. Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinform. 2010, 11, 119. [Google Scholar] [CrossRef]
  67. Tripp, H.J.; Sutton, G.; White, O.; Wortman, J.; Pati, A.; Mikhailova, N.; Ovchinnikova, G.; Payne, S.H.; Kyrpides, N.C.; Ivanova, N. Toward a standard in structural genome annotation for prokaryotes. Stand. Genom. Sci. 2015, 10, 45. [Google Scholar] [CrossRef] [PubMed]
  68. Poole, F.L.; Gerwe, B.A.; Hopkins, R.C.; Schut, G.J.; Weinberg, M.V.; Jenney, F.E.; Adams, M.W.W. Defining Genes in the Genome of the Hyperthermophilic Archaeon Pyrococcus furiosus: Implications for All Microbial Genomes. J. Bacteriol. 2005, 187, 7325–7332. [Google Scholar] [CrossRef]
  69. Van Domselaar, G.H.; Stothard, P.; Shrivastava, S.; Cruz, J.A.; Guo, A.; Dong, X.; Lu, P.; Szafron, D.; Greiner, R.; Wishart, D.S. BASys: a web server for automated bacterial genome annotation. Nucleic Acids Res. 2005, 33, W455–W459. [Google Scholar] [CrossRef]
  70. Aziz, R.K.; Bartels, D.; Best, A.A.; DeJongh, M.; Disz, T.; Edwards, R.A.; Formsma, K.; Gerdes, S.; Glass, E.M.; Kubal, M.; et al. The RAST Server: Rapid Annotations using Subsystems Technology. BMC Genom. 2008, 9, 75. [Google Scholar] [CrossRef]
  71. Olson, R.D.; Assaf, R.; Brettin, T.; Conrad, N.; Cucinell, C.; Davis, James J.; Dempsey, Donald M.; Dickerman, A.; Dietrich, Emily M.; Kenyon, Ronald W.; et al. Introducing the Bacterial and Viral Bioinformatics Resource Center (BV-BRC): a resource combining PATRIC, IRD and ViPR. Nucleic Acids Res. 2022, 51, D678–D689. [Google Scholar] [CrossRef]
  72. Haft, D.H.; Selengut, J.D.; Richter, R.A.; Harkins, D.; Basu, M.K.; Beck, E. TIGRFAMs and Genome Properties in 2013. Nucleic Acids Res. 2012, 41, D387–D395. [Google Scholar] [CrossRef]
  73. Tatusova, T.; DiCuccio, M.; Badretdin, A.; Chetvernin, V.; Nawrocki, E.P.; Zaslavsky, L.; Lomsadze, A.; Pruitt, K.D.; Borodovsky, M.; Ostell, J. NCBI prokaryotic genome annotation pipeline. Nucleic Acids Res. 2016, 44, 6614–6624. [Google Scholar] [CrossRef]
  74. Li, W.; O’Neill, K.R.; Haft, D.H.; DiCuccio, M.; Chetvernin, V.; Badretdin, A.; Coulouris, G.; Chitsaz, F.; Derbyshire, Myra K.; Durkin, A.S.; et al. RefSeq: expanding the Prokaryotic Genome Annotation Pipeline reach with protein family model curation. Nucleic Acids Res. 2020, 49, D1020–D1028. [Google Scholar] [CrossRef]
  75. Baek, J.; Lee, J.; Yoon, K.; Lee, H. Identification of Unannotated Small Genes in Salmonella. G3 Genes|Genomes|Genetics 2017, 7, 983–989. [Google Scholar] [CrossRef]
  76. Fijalkowski, I.; Willems, P.; Jonckheere, V.; Simoens, L.; Van Damme, P. Hidden in plain sight: challenges in proteomics detection of small ORF-encoded polypeptides. microLife 2022, 3. [Google Scholar] [CrossRef]
  77. Simoens, L.; Fijalkowski, I.; Van Damme, P. Exposing the small protein load of bacterial life. FEMS Microbiol. Rev. 2023, 47. [Google Scholar] [CrossRef]
  78. Kim, E.Y.; Tyndall, E.R.; Huang, K.C.; Tian, F.; Ramamurthi, K.S. Dash-and-Recruit Mechanism Drives Membrane Curvature Recognition by the Small Bacterial Protein SpoVM. Cell Syst. 2017, 5, 518–526.e513. [Google Scholar] [CrossRef]
  79. Raina, M.; Aoyama, J.J.; Bhatt, S.; Paul, B.J.; Zhang, A.; Updegrove, T.B.; Miranda-Ríos, J.; Storz, G. Dual-function AzuCR RNA modulates carbon metabolism. Proc. Natl. Acad. Sci. 2022, 119, e2117930119. [Google Scholar] [CrossRef]
  80. Yeom, J.; Shao, Y.; Groisman, E.A. Small proteins regulate Salmonella survival inside macrophages by controlling degradation of a magnesium transporter. Proc. Natl. Acad. Sci. 2020, 117, 20235–20243. [Google Scholar] [CrossRef]
  81. Ebmeier, S.E.; Tan, I.S.; Clapham, K.R.; Ramamurthi, K.S. Small proteins link coat and cortex assembly during sporulation in Bacillus subtilis. Mol. Microbiol. 2012, 84, 682–696. [Google Scholar] [CrossRef]
  82. Gaßel, M.; Möllenkamp, T.; Puppe, W.; Altendorf, K. The KdpF Subunit Is Part of the K+-translocating Kdp Complex of Escherichia coli and Is Responsible for Stabilization of the Complex in Vitro. J. Biol. Chem. 1999, 274, 37901–37907. [Google Scholar] [CrossRef]
  83. Chen, H.; Luo, Q.; Yin, J.; Gao, T.; Gao, H. Evidence for the requirement of CydX in function but not assembly of the cytochrome bd oxidase in Shewanella oneidensis. Biochim. Et. Biophys. Acta (BBA) -General. Subj. 2015, 1850, 318–328. [Google Scholar] [CrossRef]
  84. Sun, Y.-H.; de Jong, M.F.; den Hartigh, A.B.; Roux, C.M.; Rolan, H.G.; Tsolis, R.M. The small protein CydX is required for function of cytochrome bd oxidase in Brucella abortus. Front. Cell. Infect. Microbiol. 2012, 2–2012. [Google Scholar] [CrossRef]
  85. Benegas, G.; Ye, C.; Albors, C.; Li, J.C.; Song, Y.S. Genomic language models: opportunities and challenges. Trends Genet. 2025, 41, 286–302. [Google Scholar] [CrossRef]
  86. Gorenstein, L.; Konen, E.; Green, M.; Klang, E. Bidirectional Encoder Representations from Transformers in Radiology: A Systematic Review of Natural Language Processing Applications. J. Am. Coll. Radiol. 2024, 21, 914–941. [Google Scholar] [CrossRef]
  87. Brixi, G.; Durrant, M.G.; Ku, J.; Naghipourfar, M.; Poli, M.; Sun, G.; Brockman, G.; Chang, D.; Fanton, A.; Gonzalez, G.A.; et al. Genome modelling and design across all domains of life with Evo 2. Nature 2026, 652, 1349–1361. [Google Scholar] [CrossRef]
  88. Al-Ajlan, A.; El Allali, A. CNN-MGP: Convolutional Neural Networks for Metagenomics Gene Prediction. Interdiscip. Sci. Comput. Life Sci. 2019, 11, 628–635. [Google Scholar] [CrossRef]
  89. Akotenou, G.; El Allali, A. Genomic language models (gLMs) decode bacterial genomes for improved gene prediction and translation initiation site identification. Brief. Bioinform. 2025, 26. [Google Scholar] [CrossRef]
  90. Zhang, S.; Hu, H.; Jiang, T.; Zhang, L.; Zeng, J. TITER: predicting translation initiation sites by deep learning. Bioinformatics 2017, 33, i234–i242. [Google Scholar] [CrossRef]
  91. Kalkatawi, M.; Magana-Mora, A.; Jankovic, B.; Bajic, V.B. DeepGSR: an optimized deep-learning structure for the recognition of genomic signals and regions. Bioinformatics 2019, 35, 1125–1132. [Google Scholar] [CrossRef]
  92. Kumar, D.; Yadav, A.K.; Kadimi, P.K.; Nagaraj, S.H.; Grimmond, S.M.; Dash, D. Proteogenomic Analysis of Bradyrhizobium japonicum USDA110 Using Genosuite, an Automated Multi-algorithmic Pipeline*. Mol. Cell. Proteom. 2013, 12, 3388–3397. [Google Scholar] [CrossRef]
  93. Potgieter, M.G.; Nakedi, K.C.; Ambler, J.M.; Nel, A.J.M.; Garnett, S.; Soares, N.C.; Mulder, N.; Blackburn, J.M. Proteogenomic Analysis of Mycobacterium smegmatis Using High Resolution Mass Spectrometry. Front. Microbiol. 2016, 7–2016. [Google Scholar] [CrossRef]
  94. Čuklina, J.; Hahn, J.; Imakaev, M.; Omasits, U.; Förstner, K.U.; Ljubimov, N.; Goebel, M.; Pessi, G.; Fischer, H.-M.; Ahrens, C.H.; et al. Genome-wide transcription start site mapping of Bradyrhizobium japonicum grown free-living or in symbiosis – a rich resource to identify new transcripts, proteins and to study gene regulation. BMC Genom. 2016, 17, 302. [Google Scholar] [CrossRef]
  95. Abendroth, U.; Adlung, N.; Otto, A.; Grüneisen, B.; Becher, D.; Bonas, U. Identification of new protein-coding genes with a potential role in the virulence of the plant pathogen Xanthomonas euvesicatoria. BMC Genom. 2017, 18, 625. [Google Scholar] [CrossRef]
  96. Ndah, E.; Jonckheere, V.; Giess, A.; Valen, E.; Menschaert, G.; Van Damme, P. REPARATION: ribosome profiling assisted (re-)annotation of bacterial genomes. Nucleic Acids Res. 2017, 45, e168–e168. [Google Scholar] [CrossRef]
  97. Giess, A.; Jonckheere, V.; Ndah, E.; Chyżyńska, K.; Van Damme, P.; Valen, E. Ribosome signatures aid bacterial translation initiation site identification. BMC Biol. 2017, 15, 76. [Google Scholar] [CrossRef]
  98. Nakahigashi, K.; Takai, Y.; Kimura, M.; Abe, N.; Nakayashiki, T.; Shiwa, Y.; Yoshikawa, H.; Wanner, B.L.; Ishihama, Y.; Mori, H. Comprehensive identification of translation start sites by tetracycline-inhibited ribosome profiling. DNA Res. 2016, 23, 193–201. [Google Scholar] [CrossRef]
  99. Eisenberg, A.R.; Higdon, A.L.; Hollerer, I.; Fields, A.P.; Jungreis, I.; Diamond, P.D.; Kellis, M.; Jovanovic, M.; Brar, G.A. Translation Initiation Site Profiling Reveals Widespread Synthesis of Non-AUG-Initiated Protein Isoforms in Yeast. Cell Syst. 2020, 11, 145–160.e145. [Google Scholar] [CrossRef]
  100. Stringer, A.; Smith, C.; Mangano, K.; Wade, J.T. Identification of Novel Translated Small Open Reading Frames in Escherichia coli Using Complementary Ribosome Profiling Approaches. J. Bacteriol. 2022, 204, e00352–00321. [Google Scholar] [CrossRef]
  101. Limbu, M.S.; Xiong, T.; Wang, S. A review of Ribosome profiling and tools used in Ribo-seq data analysis. Comput. Struct. Biotechnol. J. 2024, 23, 1912–1918. [Google Scholar] [CrossRef]
  102. Kole, R.; Krainer, A.R.; Altman, S. RNA therapeutics: beyond RNA interference and antisense oligonucleotides. Nat. Rev. Drug Discov. 2012, 11, 125–140. [Google Scholar] [CrossRef]
  103. Pifer, R.; Greenberg, D.E. Antisense antibacterial compounds. Transl. Res. 2020, 223, 89–106. [Google Scholar] [CrossRef] [PubMed]
  104. Hör, J.; Jung, J.; Ðurica-Mitić, S.; Barquist, L.; Vogel, J. INRI-seq enables global cell-free analysis of translation initiation and off-target effects of antisense inhibitors. Nucleic Acids Res. 2022, 50, e128–e128. [Google Scholar] [CrossRef]
  105. Koch, A.; Gawron, D.; Steyaert, S.; Ndah, E.; Crappé, J.; De Keulenaer, S.; De Meester, E.; Ma, M.; Shen, B.; Gevaert, K.; et al. A proteogenomics approach integrating proteomics and ribosome profiling increases the efficiency of protein identification and enables the discovery of alternative translation start sites. PROTEOMICS 2014, 14, 2688–2698. [Google Scholar] [CrossRef]
  106. Berry, I.J.; Steele, J.R.; Padula, M.P.; Djordjevic, S.P. The application of terminomics for the identification of protein start sites and proteoforms in bacteria. PROTEOMICS 2016, 16, 257–272. [Google Scholar] [CrossRef]
  107. Gevaert, K.; Goethals, M.; Martens, L.; Van Damme, J.; Staes, A.; Thomas, G.R.; Vandekerckhove, J. Exploring proteomes and analyzing protein processing by mass spectrometric identification of sorted N-terminal peptides. Nat. Biotechnol. 2003, 21, 566–569. [Google Scholar] [CrossRef]
  108. Staes, A.; Van Damme, P.; Helsens, K.; Demol, H.; Vandekerckhove, J.; Gevaert, K. Improved recovery of proteome-informative, protein N-terminal peptides by combined fractional diagonal chromatography (COFRADIC). PROTEOMICS 2008, 8, 1362–1370. [Google Scholar] [CrossRef]
  109. Kleifeld, O.; Doucet, A.; auf dem Keller, U.; Prudova, A.; Schilling, O.; Kainthan, R.K.; Starr, A.E.; Foster, L.J.; Kizhakkedathu, J.N.; Overall, C.M. Isotopic labeling of terminal amines in complex samples identifies protein N-termini and protease cleavage products. Nat. Biotechnol. 2010, 28, 281–288. [Google Scholar] [CrossRef] [PubMed]
  110. Hanna, R.; Rozenberg, A.; Saied, L.; Ben-Yosef, D.; Lavy, T.; Kleifeld, O. In-Depth Characterization of Apoptosis N-Terminome Reveals a Link Between Caspase-3 Cleavage and Posttranslational N-Terminal Acetylation. Mol. Cell. Proteom. 2023, 22, 100584. [Google Scholar] [CrossRef]
  111. Van Damme, P. TRAINSPOTTER: Profiling Nascent Protein N-Termini Indicative of Bacterial Translation Initiation via Deformylation-Assisted N-Terminomics. Nucleic Acids Res. Accepted Author Manuscript. 2026. [Google Scholar] [CrossRef] [PubMed]
  112. Caserta, E.; Tomšic, J.; Spurio, R.; La Teana, A.; Pon, C.L.; Gualerzi, C.O. Translation Initiation Factor IF2 Interacts with the 30 S Ribosomal Subunit via Two Separate Binding Sites. J. Mol. Biol. 2006, 362, 787–799. [Google Scholar] [CrossRef] [PubMed]
  113. North, S.H.; Kirtland, S.E.; Nakai, H. Translation factor IF2 at the interface of transposition and replication by the PriA-PriC pathway. Mol. Microbiol. 2007, 66, 1566–1578. [Google Scholar] [CrossRef] [PubMed]
  114. Kato, J.-i.; Suzuki, H.; Hirota, Y. Overlapping of the coding regions for α and γ components of penicillin-binding protein 1 b in Escherichia coli. Mol. General. Genet. MGG 1984, 196, 449–457. [Google Scholar] [CrossRef]
  115. Park, S.K.; Kim, K.I.; Woo, K.M.; Seol, J.H.; Tanaka, K.; Ichihara, A.; Ha, D.B.; Chung, C.H. Site-directed mutagenesis of the dual translational initiation sites of the clpB gene of Escherichia coli and characterization of its gene products. J. Biol. Chem. 1993, 268, 20170–20174. [Google Scholar] [CrossRef]
  116. Beinker, P.; Schlee, S.; Groemping, Y.; Seidel, R.; Reinstein, J. The N Terminus of ClpB from Thermus thermophilus Is Not Essential for the Chaperone Activity. J. Biol. Chem. 2002, 277, 47160–47166. [Google Scholar] [CrossRef]
  117. Nagy, M.; Guenther, I.; Akoyev, V.; Barnett, M.E.; Zavodszky, M.I.; Kedzierska-Mieszkowska, S.; Zolkiewski, M. Synergistic Cooperation between Two ClpB Isoforms in Aggregate Reactivation. J. Mol. Biol. 2010, 396, 697–707. [Google Scholar] [CrossRef]
  118. Schlegel, S.; Rujas, E.; Ytterberg, A.J.; Zubarev, R.A.; Luirink, J.; de Gier, J.-W. Optimizing heterologous protein production in the periplasm of E. coli by regulating gene expression levels. Microb. Cell Fact. 2013, 12, 24. [Google Scholar] [CrossRef]
  119. Bhattacharyya, S.; Bershtein, S.; Yan, J.; Argun, T.; Gilson, A.I.; Trauger, S.A.; Shakhnovich, E.I. Transient protein-protein interactions perturb E. coli metabolome and cause gene dosage toxicity. eLife 2016, 5, e20309. [Google Scholar] [CrossRef]
  120. Omairi-Nasser, A.; de Gracia, A.G.; Ajlani, G. A larger transcript is required for the synthesis of the smaller isoform of ferredoxin:NADP oxidoreductase. Mol. Microbiol. 2011, 81, 1178–1189. [Google Scholar] [CrossRef] [PubMed]
  121. Thomas, J.-C.; Ughy, B.; Lagoutte, B.; Ajlani, G. A second isoform of the ferredoxin:NADP oxidoreductase generated by an in-frame initiation of translation. Proc. Natl. Acad. Sci. 2006, 103, 18368–18373. [Google Scholar] [CrossRef]
  122. Giuliodori, A.M.; Brandi, A.; Gualerzi, C.O.; Pon, C.L. Preferential translation of cold-shock mRNAs during cold adaptation. Rna 2004, 10, 265–276. [Google Scholar] [CrossRef] [PubMed]
  123. Squires, C.L.; Pedersen, S.; Ross, B.M.; Squires, C. ClpB is the Escherichia coli heat shock protein F84.1. J. Bacteriol. 1991, 173, 4254–4262. [Google Scholar] [CrossRef]
  124. Yoshida, A.; Tomita, T.; Kuzuyama, T.; Nishiyama, M. Mechanism of Concerted Inhibition of α2β2-type Hetero-oligomeric Aspartate Kinase from Corynebacterium glutamicum. J. Biol. Chem. 2010, 285, 27477–27486. [Google Scholar] [CrossRef]
  125. Wein, T.; Hülter, N.F.; Mizrahi, I.; Dagan, T. Emergence of plasmid stability under non-selective conditions maintains antibiotic resistance. Nat. Commun. 2019, 10, 2595. [Google Scholar] [CrossRef]
  126. Schwinn, M.K.; Machleidt, T.; Zimmerman, K.; Eggers, C.T.; Dixon, A.S.; Hurst, R.; Hall, M.P.; Encell, L.P.; Binkowski, B.F.; Wood, K.V. CRISPR-Mediated Tagging of Endogenous Proteins with a Luminescent Peptide. ACS Chem. Biol. 2018, 13, 467–474. [Google Scholar] [CrossRef]
  127. Einhauer, A.; Jungbauer, A. The FLAG™ peptide, a versatile fusion tag for the purification of recombinant proteins. J. Biochem. Biophys. Methods 2001, 49, 455–465. [Google Scholar] [CrossRef]
  128. Penfold, R.J.; Pemberton, J.M. An improved suicide vector for construction of chromosomal insertion mutations in bacteria. Gene 1992, 118, 145–146. [Google Scholar] [CrossRef]
  129. Fels, U.; Gevaert, K.; Van Damme, P. Bacterial Genetic Engineering by Means of Recombineering for Reverse Genetics. Front. Microbiol. 2020, 11–2020. [Google Scholar] [CrossRef] [PubMed]
  130. Wang, H.H.; Xu, G.; Vonner, A.J.; Church, G. Modified bases enable high-efficiency oligonucleotide-mediated allelic replacement via mismatch repair evasion. Nucleic Acids Res. 2011, 39, 7336–7347. [Google Scholar] [CrossRef] [PubMed]
  131. Brandis, G.; Cao, S.; Hughes, D. Operon Concatenation Is an Ancient Feature That Restricts the Potential to Rearrange Bacterial Chromosomes. Mol. Biol. Evol. 2019, 36, 1990–2000. [Google Scholar] [CrossRef]
  132. Gao, G.; Le, D.; Huang, L.; Lu, H.; Narumi, I.; Hua, Y. Internal promoter characterization and expression of the Deinococcus radiodurans pprI-folP gene cluster. FEMS Microbiol. Lett. 2006, 257, 195–201. [Google Scholar] [CrossRef]
  133. Knöppel, A.; Näsvall, J.; Andersson, D.I. Compensating the Fitness Costs of Synonymous Mutations. Mol. Biol. Evol. 2016, 33, 1461–1477. [Google Scholar] [CrossRef]
  134. K. E. Köpke, A.; A. Leggatt, P. Initiation of translation at an AUA codon for an archaebacterial protein gene expressed in E.coli. Nucleic Acids Res. 1991, 19, 5169–5172. [Google Scholar] [CrossRef]
  135. Johnson, Z.I.; Chisholm, S.W. Properties of overlapping genes are conserved across microbial genomes. Genome Res. 2004, 14, 2268–2272. [Google Scholar] [CrossRef]
  136. Huber, M.; Vogel, N.; Borst, A.; Pfeiffer, F.; Karamycheva, S.; Wolf, Y.I.; Koonin, E.V.; Soppa, J. Unidirectional gene pairs in archaea and bacteria require overlaps or very short intergenic distances for translational coupling via termination-reinitiation and often encode subunits of heteromeric complexes. Front. Microbiol. 2023, 14–2023. [Google Scholar] [CrossRef]
  137. Brown, K.M.; Wade, J.T. Translational coupling of neighboring genes in prokaryotes. J. Bacteriol. 2025, 207, e00255–00225. [Google Scholar] [CrossRef] [PubMed]
  138. Salis, H.M. Chapter two - The Ribosome Binding Site Calculator. In Methods in Enzymology; Voigt, C., Ed.; Academic Press, 2011; Volume 498, pp. 19–42. [Google Scholar]
  139. Wang, H.H.; Kim, H.; Cong, L.; Jeong, J.; Bang, D.; Church, G.M. Genome-scale promoter engineering by coselection MAGE. Nat. Methods 2012, 9, 591–593. [Google Scholar] [CrossRef] [PubMed]
  140. Kwok, P.-Y. Methods for Genotyping Single Nucleotide Polymorphisms. Annu. Rev. Genom. Hum. Genet. 2001, 2, 235–258. [Google Scholar] [CrossRef] [PubMed]
  141. Ellis, H.M.; Yu, D.; DiTizio, T.; Court, D.L. High efficiency mutagenesis, repair, and engineering of chromosomal DNA using single-stranded oligonucleotides. Proc. Natl. Acad. Sci. 2001, 98, 6742–6746. [Google Scholar] [CrossRef]
  142. Snauwaert, V.; Jonckheere, V.; Van Damme, P. Unraveling N-Terminal Proteoform Interactomes via Multiplexed Recombineering in Salmonella. In Proximity-Dependent Protein Biotinylation: Methods and Protocols; Van Damme, P., Ed.; Springer: New York, NY, USA, 2025; pp. 81–102. [Google Scholar]
  143. Subbarayan, P.R.; Sarkar, M. A stop codon-dependent internal secondary translation initiation region in Escherichia coli rpoS. Rna 2004, 10, 1359–1365. [Google Scholar] [CrossRef]
  144. Jiang, W.; Bikard, D.; Cox, D.; Zhang, F.; Marraffini, L.A. RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nat. Biotechnol. 2013, 31, 233–239. [Google Scholar] [CrossRef]
  145. Li, Y.; Lin, Z.; Huang, C.; Zhang, Y.; Wang, Z.; Tang, Y.-j.; Chen, T.; Zhao, X. Metabolic engineering of Escherichia coli using CRISPR–Cas9 meditated genome editing. Metab. Eng. 2015, 31, 13–21. [Google Scholar] [CrossRef]
  146. Huang, C.; Ding, T.; Wang, J.; Wang, X.; Guo, L.; Wang, J.; Zhu, L.; Bi, C.; Zhang, X.; Ma, X.; et al. CRISPR-Cas9-assisted native end-joining editing offers a simple strategy for efficient genetic engineering in Escherichia coli. Appl. Microbiol. Biotechnol. 2019, 103, 8497–8509. [Google Scholar] [CrossRef]
  147. Huang, C.; Guo, L.; Wang, J.; Wang, N.; Huo, Y.-X. Efficient long fragment editing technique enables large-scale and scarless bacterial genome engineering. Appl. Microbiol. Biotechnol. 2020, 104, 7943–7956. [Google Scholar] [CrossRef]
  148. Reisch, C.R.; Prather, K.L.J. The no-SCAR (Scarless Cas9 Assisted Recombineering) system for genome editing in Escherichia coli. Sci. Rep. 2015, 5, 15096. [Google Scholar] [CrossRef]
  149. Ronda, C.; Pedersen, L.E.; Sommer, M.O.A.; Nielsen, A.T. CRMAGE: CRISPR Optimized MAGE Recombineering. Sci. Rep. 2016, 6, 19452. [Google Scholar] [CrossRef]
  150. Lee, H.J.; Kim, H.J.; Lee, S.J. CRISPR-Cas9-mediated pinpoint microbial genome editing aided by target-mismatched sgRNAs. Genome Res. 2020, 30, 768–775. [Google Scholar] [CrossRef]
  151. Lim, S.R.; Lee, H.J.; Kim, H.J.; Lee, S.J. Multiplex Single-Nucleotide Microbial Genome Editing Achieved by CRISPR-Cas9 Using 5′-End-Truncated sgRNAs. ACS Synth. Biol. 2023, 12, 2203–2207. [Google Scholar] [CrossRef]
  152. Elsayed, E.M.; Stukenberg, D.; Meier, D.; Schmeck, B.; Becker, A. RECKLEEN is a lambda Red/CRISPR-Cas9 based single plasmid platform for enhanced genome editing in Klebsiella pneumoniae. Commun. Biol. 2025, 8, 1509. [Google Scholar] [CrossRef]
  153. Vento, J.M.; Crook, N.; Beisel, C.L. Barriers to genome editing with CRISPR in bacteria. J. Ind. Microbiol. Biotechnol. 2019, 46, 1327–1341. [Google Scholar] [CrossRef]
  154. Roux, K.J.; Kim, D.I.; Raida, M.; Burke, B. A promiscuous biotin ligase fusion protein identifies proximal and interacting proteins in mammalian cells. J. Cell Biol. 2012, 196, 801–810. [Google Scholar] [CrossRef]
  155. Kim, D.I.; Jensen, S.C.; Noble, K.A.; KC, B.; Roux, K.H.; Motamedchaboki, K.; Roux, K.J.; Zheng, Y. An improved smaller biotin ligase for BioID proximity labeling. Mol. Biol. Cell 2016, 27, 1188–1196. [Google Scholar] [CrossRef]
  156. Branon, T.C.; Bosch, J.A.; Sanchez, A.D.; Udeshi, N.D.; Svinkina, T.; Carr, S.A.; Feldman, J.L.; Perrimon, N.; Ting, A.Y. Efficient proximity labeling in living cells and organisms with TurboID. Nat. Biotechnol. 2018, 36, 880–887. [Google Scholar] [CrossRef]
  157. Herfurth, M.; Müller, F.; Søgaard-Andersen, L.; Glatter, T. A miniTurbo-based proximity labeling protocol to identify conditional protein interactomes in vivo in Myxococcus xanthus. STAR Protoc. 2023, 4, 102657. [Google Scholar] [CrossRef] [PubMed]
  158. Remy, O.; Santin, Y.G.; Jonckheere, V.; Tesseur, C.; Kaljević, J.; Damme, P.V.; Laloux, G. Distinct dynamics and proximity networks of hub proteins at the prey-invading cell pole in a predatory bacterium. J. Bacteriol. 2024, 206, e00014–00024. [Google Scholar] [CrossRef]
  159. Jonckheere, V.; Van Damme, P. N-Terminal Acetyltransferase Naa40p Whereabouts Put into N-Terminal Proteoform Perspective. Int. J. Mol. Sci. 2021, 22, 3690. [Google Scholar] [CrossRef] [PubMed]
  160. Bogaert, A.; Fijalkowska, D.; Staes, A.; Van de Steene, T.; Vuylsteke, M.; Stadler, C.; Eyckerman, S.; Spirohn, K.; Hao, T.; Calderwood, M.A.; et al. N-terminal proteoforms may engage in different protein complexes. Life Sci. Alliance 2023, 6, e202301972. [Google Scholar] [CrossRef]
  161. Schwanhäusser, B.; Busse, D.; Li, N.; Dittmar, G.; Schuchhardt, J.; Wolf, J.; Chen, W.; Selbach, M. Global quantification of mammalian gene expression control. Nature 2011, 473, 337–342. [Google Scholar] [CrossRef]
  162. Jayapal, K.P.; Sui, S.; Philp, R.J.; Kok, Y.-J.; Yap, M.G.S.; Griffin, T.J.; Hu, W.-S. Multitagging Proteomic Strategy to Estimate Protein Turnover Rates in Dynamic Systems. J. Proteome Res. 2010, 9, 2087–2097. [Google Scholar] [CrossRef] [PubMed]
  163. Fierro-Monti, I.; Racle, J.; Hernandez, C.; Waridel, P.; Hatzimanikatis, V.; Quadroni, M. A Novel Pulse-Chase SILAC Strategy Measures Changes in Protein Decay and Synthesis Rates Induced by Perturbation of Proteostasis with an Hsp90 Inhibitor. PLoS ONE 2013, 8, e80423. [Google Scholar] [CrossRef]
  164. Boisvert, F.-M.; Ahmad, Y.; Gierliński, M.; Charrière, F.; Lamont, D.; Scott, M.; Barton, G.; Lamond, A.I. A Quantitative Spatial Proteomics Analysis of Proteome Turnover in Human Cells. Mol. Cell. Proteom. 2012, 11. [Google Scholar] [CrossRef]
  165. Han, J.; Yi, S.; Zhao, X.; Zheng, Y.; Yang, D.; Du, G.; Yang, X.-Y.; He, Q.-Y.; Sun, X. Improved SILAC method for double labeling of bacterial proteome. J. Proteom. 2019, 194, 89–98. [Google Scholar] [CrossRef]
  166. Gupta, M.; Johnson, A.N.T.; Cruz, E.R.; Costa, E.J.; Guest, R.L.; Li, S.H.-J.; Hart, E.M.; Nguyen, T.; Stadlmeier, M.; Bratton, B.P.; et al. Global protein turnover quantification in Escherichia coli reveals cytoplasmic recycling under nitrogen limitation. Nat. Commun. 2024, 15, 5890. [Google Scholar] [CrossRef]
  167. Xue, Y.; Sherman, D.H. Alternative modular polyketide synthase expression controls macrolactone structure. Nature 2000, 403, 571–575. [Google Scholar] [CrossRef] [PubMed]
  168. Bohlin, J.; Pettersson, J.H.-O. Evolution of Genomic Base Composition: From Single Cell Microbes to Multicellular Animals. Comput. Struct. Biotechnol. J. 2019, 17, 362–370. [Google Scholar] [CrossRef] [PubMed]
  169. Pan, T.; Sosnick, T. RNA FOLDING DURING TRANSCRIPTION. Annu. Rev. Biophys. 2006, 35, 161–175. [Google Scholar] [CrossRef] [PubMed]
  170. Chow, I.T.; Baneyx, F. Coordinated synthesis of the two ClpB isoforms improves the ability of Escherichia coli to survive thermal stress. FEBS Lett. 2005, 579, 4235–4241. [Google Scholar] [CrossRef]
  171. Ozin, A.J.; Costa, T.; Henriques, A.O.; Moran, C.P. Alternative Translation Initiation Produces a Short Form of a Spore Coat Protein in Bacillus subtilis. J. Bacteriol. 2001, 183, 2032–2040. [Google Scholar] [CrossRef]
  172. Yu, X.-J.; Liu, M.; Matthews, S.; Holden, D.W. Tandem Translation Generates a Chaperone for the Salmonella Type III Secretion System Protein SsaQ*. J. Biol. Chem. 2011, 286, 36098–36107. [Google Scholar] [CrossRef]
Figure 1. Mechanisms contributing to bacterial N-terminal (Nt-) proteoform diversity. Distinct Nt-proteoforms can arise through multiple co- and post-translational mechanisms that collectively expand bacterial proteome complexity. (A) N-terminal processing and protein modifications: Nascent bacterial proteins initially carry a formylmethionine residue. The Nt-formyl group (f) can be removed co-translationally by peptide deformylase (PDF), after which the initiator methionine (M) may either be retained or cleaved by methionine aminopeptidase (MetAP). Mature proteins may subsequently undergo additional post-translational modifications (PTMs), including phosphorylation, pupylation (addition of a prokaryotic ubiquitin-like protein), glycosylation, and other modifications, thereby further contributing to N-terminal heterogeneity. (B) Proteolytic processing: Cleavage of N-terminal signal peptides (SPs) by dedicated signal peptidases (SPases) generates mature processed proteoforms with altered N-termini following protein translocation. (C) Alternative translation initiation: Multiple proteoforms can originate from a single gene through the use of alternative translation initiation sites (aTISs) within the same transcript. Translation may initiate from the canonical AUG start codon or from near-cognate initiation codons such as GUG and UUG. Promoter architecture, including the −10 and −35 promoter elements, as well as the Shine–Dalgarno (SD) sequence involved in ribosome recruitment, is indicated. Utilization of distinct TISs results in Nt-proteoforms differing in N-terminal composition and length. Figure created using BioRender.
Figure 1. Mechanisms contributing to bacterial N-terminal (Nt-) proteoform diversity. Distinct Nt-proteoforms can arise through multiple co- and post-translational mechanisms that collectively expand bacterial proteome complexity. (A) N-terminal processing and protein modifications: Nascent bacterial proteins initially carry a formylmethionine residue. The Nt-formyl group (f) can be removed co-translationally by peptide deformylase (PDF), after which the initiator methionine (M) may either be retained or cleaved by methionine aminopeptidase (MetAP). Mature proteins may subsequently undergo additional post-translational modifications (PTMs), including phosphorylation, pupylation (addition of a prokaryotic ubiquitin-like protein), glycosylation, and other modifications, thereby further contributing to N-terminal heterogeneity. (B) Proteolytic processing: Cleavage of N-terminal signal peptides (SPs) by dedicated signal peptidases (SPases) generates mature processed proteoforms with altered N-termini following protein translocation. (C) Alternative translation initiation: Multiple proteoforms can originate from a single gene through the use of alternative translation initiation sites (aTISs) within the same transcript. Translation may initiate from the canonical AUG start codon or from near-cognate initiation codons such as GUG and UUG. Promoter architecture, including the −10 and −35 promoter elements, as well as the Shine–Dalgarno (SD) sequence involved in ribosome recruitment, is indicated. Utilization of distinct TISs results in Nt-proteoforms differing in N-terminal composition and length. Figure created using BioRender.
Preprints 217048 g001
Figure 2. Biological relevance and functional implications of bacterial N-terminal (Nt-)proteoforms. The bacterial translatome constitutes a highly dynamic landscape in which N-terminal heterogeneity introduces multiple regulatory layers that influence protein fate and function. Protein stability: Variations at the N-terminus can modulate protein half-life through the bacterial N-degron pathway, in which specific N-terminal residues determine recognition by the ClpS-ClpAP proteolytic machinery. Functional polarization: Distinct N-termini can alter biochemical activity and signaling behaviour. The long CheA Nt-proteoform (CheAL) contains an N-terminal autophosphorylation site that is absent in the shorter CheAS proteoform. Consequently, CheAL primarily functions as a kinase that phosphorylates CheY, whereas CheAS interacts with CheZ to promote CheY dephosphorylation. Subcellular localization: N-terminal extensions can encode transport or secretion signals, such as N-terminal type III effector (T3E) secretion signals, thereby enabling selective translocation of specific proteoforms (e.g., SseLL) into the host cell, while truncated proteoforms lacking these signals (e.g., SseLS) remain confined to the bacterial cytoplasm. Multiprotein complex assembly: Nt-proteoform diversity can further regulate the stoichiometric assembly and functionality of higher-order molecular complexes, exemplified by the T3SS sorting platform (SP), where defined ratios between long and short isoforms (e.g., one SpaOL and two SpaOS subunits) are required for proper injectisome assembly and activity. Figure created using BioRender.
Figure 2. Biological relevance and functional implications of bacterial N-terminal (Nt-)proteoforms. The bacterial translatome constitutes a highly dynamic landscape in which N-terminal heterogeneity introduces multiple regulatory layers that influence protein fate and function. Protein stability: Variations at the N-terminus can modulate protein half-life through the bacterial N-degron pathway, in which specific N-terminal residues determine recognition by the ClpS-ClpAP proteolytic machinery. Functional polarization: Distinct N-termini can alter biochemical activity and signaling behaviour. The long CheA Nt-proteoform (CheAL) contains an N-terminal autophosphorylation site that is absent in the shorter CheAS proteoform. Consequently, CheAL primarily functions as a kinase that phosphorylates CheY, whereas CheAS interacts with CheZ to promote CheY dephosphorylation. Subcellular localization: N-terminal extensions can encode transport or secretion signals, such as N-terminal type III effector (T3E) secretion signals, thereby enabling selective translocation of specific proteoforms (e.g., SseLL) into the host cell, while truncated proteoforms lacking these signals (e.g., SseLS) remain confined to the bacterial cytoplasm. Multiprotein complex assembly: Nt-proteoform diversity can further regulate the stoichiometric assembly and functionality of higher-order molecular complexes, exemplified by the T3SS sorting platform (SP), where defined ratios between long and short isoforms (e.g., one SpaOL and two SpaOS subunits) are required for proper injectisome assembly and activity. Figure created using BioRender.
Preprints 217048 g002
Figure 3. Integrated riboproteogenomic framework for high-resolution translation initiation site (TIS) mapping and bacterial genome annotation. Complementary experimental and computational approaches collectively enable the identification and validation of annotated and alternative translation initiation sites (aTISs) within bacterial genomes. The central circular overview integrates genome-wide TIS mapping results supported by ribosome profiling, N-terminomics, and computational prediction approaches. Annotated TISs, experimentally supported aTISs, and computationally predicted TISs are represented alongside the corresponding open reading frame (ORF) architecture. (A) Ribosome profiling. Conventional Ribo-seq captures actively translated genomic elements through ribosome-protected mRNA fragments. Specialized approaches such as retapamulin-assisted ribosome (Ribo-RET) enrich initiating ribosomes through antibiotic-mediated stalling at start codons, thereby improving the detection of canonical and aTISs. (B) N-terminomics. N-terminal proteomics workflows combine protein extraction, proteolytic digestion, selective enrichment of N-terminal peptides, and LC-MS/MS analysis to identify proteoform-specific N-termini. The presence of formylated methionine (fM) serves as a direct proxy for an actual initiator methionine. Consequently, detection of fM-containing N-terminal peptides provides direct molecular evidence for translation initiation events and enables TIS delineation at single-amino acid resolution. (C) Computational TIS prediction. Computational approaches ranging from classical ab initio gene finders (e.g., GeneMark, Prodigal, and GLIMMER) to modern genomic language models (gLMs) are used to predict coding sequences (CDSs) and TISs. Figure created using BioRender.
Figure 3. Integrated riboproteogenomic framework for high-resolution translation initiation site (TIS) mapping and bacterial genome annotation. Complementary experimental and computational approaches collectively enable the identification and validation of annotated and alternative translation initiation sites (aTISs) within bacterial genomes. The central circular overview integrates genome-wide TIS mapping results supported by ribosome profiling, N-terminomics, and computational prediction approaches. Annotated TISs, experimentally supported aTISs, and computationally predicted TISs are represented alongside the corresponding open reading frame (ORF) architecture. (A) Ribosome profiling. Conventional Ribo-seq captures actively translated genomic elements through ribosome-protected mRNA fragments. Specialized approaches such as retapamulin-assisted ribosome (Ribo-RET) enrich initiating ribosomes through antibiotic-mediated stalling at start codons, thereby improving the detection of canonical and aTISs. (B) N-terminomics. N-terminal proteomics workflows combine protein extraction, proteolytic digestion, selective enrichment of N-terminal peptides, and LC-MS/MS analysis to identify proteoform-specific N-termini. The presence of formylated methionine (fM) serves as a direct proxy for an actual initiator methionine. Consequently, detection of fM-containing N-terminal peptides provides direct molecular evidence for translation initiation events and enables TIS delineation at single-amino acid resolution. (C) Computational TIS prediction. Computational approaches ranging from classical ab initio gene finders (e.g., GeneMark, Prodigal, and GLIMMER) to modern genomic language models (gLMs) are used to predict coding sequences (CDSs) and TISs. Figure created using BioRender.
Preprints 217048 g003
Figure 4. Multiplexed recombineering enables proteoform-specific proximity interactome (”proxeome”) mapping. Schematic representation of a bacterial genomic locus comprising genes X, Y and Z. Gene Y encodes two Nt-proteoforms generated through alternative translation initiation. The long proteoform (YLong) is translated from the database-annotated translation initiation site (dbTIS; blue flag), whereas the short proteoform (YShort) originates from an alternative translation initiation site (aTIS; green flag). Multiplexed recombineering is used to selectively inactivate either TIS while simultaneously integrating a C-terminal promiscuous biotin ligase (PBL) fusion. This genome engineering strategy generates strains selectively expressing either YLong or YShort -PBL under endogenous regulatory control. Subsequent proximity-dependent biotin identification (BioID) approaches enable mapping of the specific proxeomes associated with each Nt-proteoform. Solid lines denote direct or stable protein interactions, whereas dashed lines represent indirect, transient, or shared interaction networks. Figure created using BioRender.
Figure 4. Multiplexed recombineering enables proteoform-specific proximity interactome (”proxeome”) mapping. Schematic representation of a bacterial genomic locus comprising genes X, Y and Z. Gene Y encodes two Nt-proteoforms generated through alternative translation initiation. The long proteoform (YLong) is translated from the database-annotated translation initiation site (dbTIS; blue flag), whereas the short proteoform (YShort) originates from an alternative translation initiation site (aTIS; green flag). Multiplexed recombineering is used to selectively inactivate either TIS while simultaneously integrating a C-terminal promiscuous biotin ligase (PBL) fusion. This genome engineering strategy generates strains selectively expressing either YLong or YShort -PBL under endogenous regulatory control. Subsequent proximity-dependent biotin identification (BioID) approaches enable mapping of the specific proxeomes associated with each Nt-proteoform. Solid lines denote direct or stable protein interactions, whereas dashed lines represent indirect, transient, or shared interaction networks. Figure created using BioRender.
Preprints 217048 g004
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings