Preprint
Article

This version is not peer-reviewed.

Structure-Guided Discovery Reveals Recurrent Bioactive Peptide Architectures Across Coleoptera

Submitted:

20 July 2026

Posted:

22 July 2026

You are already at the latest version

Abstract
Bioactive peptides are an important source of therapeutic molecules and molecular scaffolds involved in defense, signaling, and immune regulation. Despite the extraordinary diversity of Coleoptera, the structural landscape of beetle-derived bioactive peptides remains largely unexplored, limiting our understanding of their evolutionary diversity and biotechnological potential. Here, we performed a large-scale structural survey of predicted toxin-like peptide scaffolds across publicly available Coleoptera transcriptomes by integrating transcriptome mining, peptide maturation prediction, physicochemical characterization, AlphaFold 3 structural modeling, structural similarity analyses, and interpretable machine learning. We identified 291 candidate peptides, of which 155 contained canonical signal peptides and 273 produced mature peptides within the expected size range of known bioactive peptides. Structural analyses revealed that, despite extensive sequence diversity, many candidates converged toward a comparatively restricted repertoire of compact cysteine-rich architectures, suggesting that structural conservation exceeds primary sequence conservation during peptide diversification. Comparative structural analyses further identified recurrent protein architectures shared across multiple beetle lineages, while machine learning prioritization integrated structural and biochemical descriptors to identify high-confidence candidates for future functional characterization. These analyses establish the first structural atlas of predicted toxin-like peptides across Coleoptera and demonstrate that structure-guided transcriptome mining provides a powerful framework for uncovering evolutionarily conserved bioactive peptide scaffolds that would remain largely undetected using sequence-based approaches alone. Beyond expanding our understanding of peptide evolution in beetles, this resource is a foundation for future structural, functional, and biotechnological exploration of bioactive peptides in underexplored animal groups.
Keywords: 
;  ;  ;  ;  

Introduction

Bioactive peptides constitute one of the most diverse classes of biologically active molecules in nature, participating in a wide range of physiological processes including defense, signaling, immune regulation, enzyme inhibition, and interspecific interactions. Their remarkable structural diversity, combined with high specificity and favorable pharmacological properties, has made these molecules valuable sources of therapeutic agents, agricultural bioproducts, and molecular probes [1,2,3]. Among these molecules, small secreted cysteine-rich peptides are particularly notable because compact disulfide-stabilized architectures frequently preserve structural integrity while supporting extensive functional diversification [4,5,6] .
The discovery of bioactive peptides has focused predominantly on a relatively small number of well-characterized venomous organisms, including snakes, spiders, cone snails, scorpions, and marine invertebrates [6,7]. However, accumulating evidence suggests that structurally related peptide scaffolds extend far beyond these traditional groups and may perform diverse biological functions unrelated to venom, including antimicrobial activity, immune modulation, and cellular signaling [2,8,9]. Consequently, taxonomically diverse yet poorly explored groups may harbor a vast and largely unrecognized diversity of bioactive peptide scaffolds.
Coleoptera represents an outstanding example of this knowledge gap. As the largest animal order, comprising nearly one-quarter of all described animal species, beetles occupy virtually every terrestrial ecosystem and exhibit extraordinary ecological and evolutionary diversification. Despite this remarkable diversity, comparatively little is known about the repertoire, structural organization, and evolutionary relationships of their small secreted peptides. Most available studies have focused on individual antimicrobial peptides or isolated toxin-like molecules, leaving the broader structural diversity of beetle-derived bioactive peptides largely unexplored [3,8]. This disparity raises an important question: does the apparent scarcity of beetle bioactive peptides reflect genuine biological rarity, or simply limited exploration?
Recent advances in artificial intelligence and structural bioinformatics have fundamentally transformed protein discovery. High-accuracy structure prediction methods, particularly AlphaFold 3, together with large-scale structural comparison tools such as Foldseek, now enable protein relationships to be explored beyond primary sequence conservation [10,11]. Increasingly, protein structure is recognized as a more conserved and functionally informative descriptor than sequence alone, allowing the identification of evolutionarily related protein architectures that would otherwise remain undetected [12,13]. These developments have shifted peptide discovery from predominantly sequence-based approaches toward integrative structure-guided strategies capable of exploring previously inaccessible regions of protein structural space.
Here, we present the first large-scale structural survey of predicted toxin-like peptide scaffolds across Coleoptera. By integrating transcriptome mining, peptide maturation prediction, physicochemical characterization, AlphaFold 3 structural modeling, structural similarity analyses, and interpretable machine learning, we generated a comprehensive structural atlas of bioactive peptide candidates spanning diverse beetle lineages. Rather than focusing solely on peptide identification, our objective was to investigate how structural information organizes peptide diversity across Coleoptera and whether recurrent protein architectures emerge despite extensive sequence divergence. Together, our results demonstrate that beetle transcriptomes harbor a previously underappreciated repertoire of structurally conserved bioactive peptide scaffolds and establish a scalable framework for structure-guided peptide discovery in underexplored animal groups.

Materials and Methods

To systematically identify and prioritize toxin-like peptides (TLPs) in Coleoptera, we developed an integrative computational pipeline combining sequence-based filtering, structural characterization, and machine learning-based classification. The workflow is summarized below.

2.1. Dataset Construction and Candidate Peptide Identification

We assembled/obtained a dataset comprising 88 publicly available transcriptomes spanning multiple beetle lineages (Table S1). Transcriptome assemblies were primarily obtained from the NCBI Transcriptome Shotgun Assembly (TSA) database, representing 70 unique TSA projects, with additional publicly available transcriptomic resources incorporated when required to maximize taxonomic coverage (Table S1). Raw RNA-seq datasets representing multiple Coleoptera species were assembled de novo using Trinity v2.15.2 [14]. The resulting transcriptomes were translated into putative open reading frames (ORFs) using TransDecoder v5.7.1, generating a comprehensive protein sequence dataset. To minimize redundancy and reduce sampling bias in downstream analyses, protein sequences were clustered using CD-HIT v4.8.1 (-c 0.9) [15], and representative sequences from each cluster were retained to construct a non-redundant dataset. Special attention was given to preserving short ORFs typically excluded during conventional annotation pipelines. Candidate secretory proteins were identified using SignalP 5.0 [16], and only sequences containing a predicted N-terminal signal peptide were retained, consistent with the canonical secretory pathway described for extracellular bioactive peptides [17]. Predicted signal peptide cleavage sites were subsequently used to delimit mature peptide regions.
Predicted precursor sequences were further examined for features associated with peptide maturation, including propeptide regions and proteolytic cleavage motifs [18]. Mature peptide sequences were computationally extracted whenever possible, allowing downstream analyses to focus on biologically active regions rather than full-length precursor proteins. Candidate peptides were subsequently filtered according to physicochemical characteristics commonly associated with structurally constrained bioactive peptides, including peptide length, net charge, cysteine content and spacing, amphipathicity, hydrophobic moment (μH), amino acid composition, and other sequence-derived descriptors [19]. Rather than excluding atypical sequences, these criteria were used to enrich the dataset while preserving structural and taxonomic diversity, allowing the inclusion of both canonical and potentially novel toxin-like peptide scaffolds.
All candidate sequences were screened against curated profile hidden Markov models using HMMER v3.4 [20] to identify conserved domains associated with well-characterized peptide families, including Kunitz inhibitors, Kazal domains, defensins, and cystine-knot peptides. Because short extracellular peptides frequently evolve rapidly and exhibit limited sequence conservation, both complete and partial HMM matches were considered during annotation. Importantly, sequences lacking significant HMM matches were retained for subsequent analyses, ensuring that highly divergent or previously undescribed peptide scaffolds were not excluded from the discovery pipeline.

2.2. Structure Prediction and Comparative Structural Analyses

Three-dimensional structural models were generated for all candidate mature peptides using AlphaFold 3 [11]. Model confidence was evaluated using the predicted Local Distance Difference Test (pLDDT), and only models with acceptable confidence scores were considered for downstream structural analyses. The resulting structures constituted the basis for all subsequent comparative analyses, enabling structural characterization independently of primary sequence similarity.
To comprehensively describe the structural properties of each peptide, multiple complementary descriptors were extracted from the predicted models. Secondary structure assignments were obtained using DSSP 4 [21], whereas solvent-accessible surface area (SASA) was calculated using FreeSASA v2.2.1 [22]. Additional structural descriptors, including radius of gyration, structural compactness, residue contact patterns, and other geometric measurements, were calculated with MDTraj v1.10.3 [23].
Structural relationships among candidate peptides were investigated through pairwise structural comparisons performed using Foldseek [10], enabling the identification of structurally similar proteins even in the absence of detectable sequence homology. Foldseek similarity scores were subsequently used to identify recurrent structural architectures shared among peptides from distinct Coleoptera lineages. Unlike conventional sequence-based similarity searches, this approach enables the detection of evolutionarily conserved protein folds that remain structurally similar despite extensive primary sequence divergence. The structural descriptors extracted from AlphaFold 3 models together with Foldseek similarity metrics were integrated into downstream analyses to characterize the structural landscape of the identified peptide repertoire.

2.3. Structure-Guided Machine Learning Prioritization

We developed an interpretable machine learning framework to prioritize candidates for downstream analyses integrating sequence-derived, physicochemical, and structural descriptors into a unified predictive model. Unlike conventional sequence-based peptide discovery pipelines, our strategy incorporated information extracted from AlphaFold 3 models and comparative structural analyses, allowing structurally similar peptides to be identified even in the presence of extensive sequence divergence. Because experimentally validated negative examples of toxin-like peptides are largely unavailable, candidate prioritization was formulated as a Positive-Unlabeled (PU) learning problem. This strategy enables robust classification when only confirmed positive examples and unlabeled sequences are available, avoiding biases associated with artificially generated negative datasets. The training dataset consisted of experimentally characterized toxin-like peptides representing diverse peptide families, whereas all remaining candidates were treated as unlabeled observations.
Each candidate was represented by a comprehensive feature set combining sequence-derived descriptors, physicochemical properties, structural metrics extracted from predicted three-dimensional models, and Foldseek similarity scores. These features included peptide length, molecular weight, amino acid composition, net charge, isoelectric point, hydrophobicity, aliphatic index, cysteine content, secondary-structure composition, solvent-accessible surface area, radius of gyration, structural compactness, pLDDT confidence scores, and comparative structural similarity metrics. This integrated representation enabled simultaneous evaluation of biochemical and structural characteristics during candidate prioritization.
Classification models were implemented in Python using scikit-learn [24]. Hyperparameter optimization and model selection were performed using repeated cross-validation, and predictive performance was evaluated through Receiver Operating Characteristic (ROC) and Precision–Recall (PR) curves. Because peptide discovery represents a highly imbalanced classification problem, PR curves were considered particularly informative for assessing classifier performance [25], while complementary probabilistic performance measures followed the recommendations of Gajda and Chlebus [26].
Model predictions were examined using SHapley Additive exPlanations (SHAP), which quantify the contribution of individual features to each prediction and improve biological interpretability [27,28]. The resulting prioritization scores were not interpreted as definitive functional annotations but rather as confidence estimates integrating multiple independent sources of evidence. These scores were subsequently combined with structural clustering analyses to guide the selection of representative candidates for the structural atlas and future experimental validation.
Taxonomic information associated with each transcriptome was standardized using the taxize package [29] in R [30]. Candidate peptides were subsequently grouped according to their taxonomic classification to investigate the distribution of predicted bioactive peptide scaffolds across Coleoptera lineages and to identify peptide families exhibiting lineage-specific enrichment. Associations between peptide occurrence and taxonomic groups were evaluated using Fisher's Exact Test [31]. To account for multiple hypothesis testing, P-values were adjusted using the Benjamini–Hochberg false discovery rate (FDR) procedure [32], and corrected values below 0.05 were considered statistically significant.
All statistical analyses and data manipulation were performed in R [30], primarily using the dplyr [33] and tidyr [34] packages. Summary statistics were calculated for sequence-derived, physicochemical, structural, and machine learning-derived descriptors, and graphical visualizations were generated to support comparative analyses presented throughout the study.

2.4. Structural Atlas Construction

The final structural atlas was generated by integrating sequence-derived descriptors, physicochemical properties, predicted three-dimensional structures, comparative structural analyses, and machine learning prioritization into a unified dataset. Each candidate peptide was represented by its mature amino acid sequence, structural model, physicochemical profile, structural descriptors, Foldseek similarity metrics, and prioritization score, providing a comprehensive overview of its predicted structural and biochemical characteristics.
Peptides were organized according to structural similarity to facilitate comparative analyses. Foldseek-derived structural comparisons enabled the identification of recurrent protein architectures shared among candidate peptides, allowing structurally related molecules to be grouped even when sequence conservation was minimal. Representative structures from each structural group were selected to summarize the diversity of the identified peptide repertoire while preserving the major architectural features observed across the dataset. Machine learning prioritization scores were subsequently integrated with structural information to support the selection of high-confidence candidates for downstream analyses.

Results

3.1. Large-Scale Discovery of Toxin-like Peptide Candidates Across Coleoptera

The analyses identified 291 short peptide candidates exhibiting characteristics compatible with toxin-like molecules, providing one of the first large-scale overviews of putative toxin-like peptides across the order (Table 1). These candidates were recovered from multiple beetle lineages, indicating that short secreted peptides are broadly distributed throughout Coleoptera rather than being restricted to a small number of taxa.
Among the identified candidates, 155 peptides (53.3%) were predicted to contain N-terminal signal peptides, supporting their potential secretion through the classical secretory pathway (Table 1). Because extracellular secretion represents a fundamental requirement for most peptide toxins and other bioactive peptides [17], this result substantially enriches the candidate dataset for molecules with potential biological activity.
To assess their maturation potential, precursor sequences were computationally processed using predicted signal peptide cleavage sites together with canonical proteolytic processing rules. Following maturation inference, 273 candidates (93.8%) generated mature peptides within the expected size range of 8-80 amino acids, consistent with the size distribution commonly reported for bioactive toxin-like peptides (Table 1) [8]. Only 18 sequences (6.2%) produced mature peptides outside this interval and were therefore considered lower-confidence candidates. Proteolytic processing signals were also frequently observed. A total of 127 candidates (43.6%) contained canonical dibasic cleavage motifs (e.g., KR, RR, and KK) downstream of the signal peptide, supporting precursor architectures compatible with regulated peptide maturation [1,35]. The distributions of mature peptide length, cleavage motifs and cysteine content (Figure 1) reveal considerable diversity while remaining within the structural and biochemical boundaries typically associated with secreted bioactive peptides.

3.2. Physicochemical Landscape of Coleoptera Toxin-like Peptides

We analyzed the physicochemical properties of their mature peptide sequences, including peptide length, molecular weight, net charge, isoelectric point, hydrophobicity, aliphatic index, and cysteine content (Figure 2; Table S1). Together, these descriptors provide a comprehensive overview of the biochemical landscape occupied by the identified candidates and offer important insights into their potential structural and functional properties.
The predicted mature peptides exhibited substantial variability in sequence length and molecular weight, reflecting the broad diversity expected among naturally occurring bioactive peptides while remaining within the size range commonly reported for toxin-like molecules (Figure 2A; Table S1). Similarly, predicted net charge spanned both negatively and positively charged peptides (Figure 2B), indicating considerable electrostatic diversity that may influence molecular recognition, receptor binding, and interactions with biological membranes. Such variation is characteristic of bioactive peptide families that have evolved to recognize distinct molecular targets while maintaining compact extracellular architectures [8].
Additional physicochemical descriptors further supported the diversity of the identified candidates. Predicted isoelectric point, hydrophobicity, and aliphatic index varied substantially among peptides (Table S1), suggesting differences in folding behavior, molecular stability, and interaction potential with proteins or membrane-associated targets. Despite this variability, the overall physicochemical distributions remained consistent with those reported for structurally constrained secreted peptides, which encompass a wide range of biological activities beyond classical toxin function, including antimicrobial defense, immune modulation, signaling, and enzyme inhibition [1,2].
One of the most distinctive characteristics of the dataset was the abundance of cysteine-rich peptides (Figure 2C). Many candidates contained multiple cysteine residues arranged in patterns compatible with disulfide bond formation, a defining feature of numerous extracellular bioactive peptide families. Disulfide bonds play a central role in stabilizing compact three-dimensional conformations, increasing resistance to proteolytic degradation while preserving structural integrity and molecular specificity [4,5,6]. The widespread occurrence of cysteine-rich architectures therefore suggests that structural stabilization has been a major determinant of peptide evolution within the identified candidates. These physicochemical analyses indicate that the predicted peptides occupy a biochemical space highly consistent with structurally constrained bioactive peptides. Although the candidates display considerable diversity in their primary sequence composition and physicochemical properties, they also share conserved molecular characteristics commonly associated with secreted extracellular peptides. This combination of diversity and conserved biochemical signatures suggests that common structural frameworks may underlie these molecules despite extensive sequence variation, providing the rationale for the structural analyses presented in the following section.

3.3. Structural Landscape of Coleoptera Toxin-like Peptides

To investigate the structural diversity of the identified candidates, three-dimensional models were generated using AlphaFold 3 and subsequently characterized through multiple complementary structural descriptors, including predicted Local Distance Difference Test (pLDDT), secondary structure composition, radius of gyration (Rg), solvent-accessible surface area (SASA), and disulfide bond content (Figure 3; Table S2).
The predicted structures exhibited high confidence, with most candidates presenting mean pLDDT values compatible with reliable backbone prediction (Figure 3; Table S2). As expected for short secreted peptides, regions displaying lower confidence were predominantly restricted to terminal residues and exposed loop segments, whereas the structural cores generally exhibited consistently high confidence scores. These observations indicate that the predicted models provide a suitable framework for comparative structural analyses. Despite considerable sequence diversity, the candidate peptides shared several common structural characteristics. Most models adopted compact globular conformations composed of short α-helices, β-strands, and connecting loops arranged in diverse combinations (Figure 3). Radius of gyration and solvent-accessible surface area varied according to peptide length, although the majority of candidates remained within the expected range for small extracellular peptides.
A prominent feature of the dataset was the high frequency of cysteine-rich structures. In many candidates, cysteine residues occupied conserved spatial positions compatible with the formation of stabilizing disulfide bonds, generating compact architectures characteristic of structurally constrained peptide families. Such structural stabilization is widely recognized as a defining characteristic of numerous toxin-like peptides, contributing to both conformational stability and resistance to proteolytic degradation [5].
Individual structures displayed substantial topological variation, however, the overall structural landscape revealed a relatively limited repertoire of recurrent architectural solutions. This observation suggests that diverse primary sequences frequently converge toward compact three-dimensional organizations that preserve structural stability while potentially accommodating distinct biological functions. These results establish the first large-scale structural overview of predicted toxin-like peptides across Coleoptera and is the foundation for subsequent analyses of structural similarity and candidate prioritization.
Interestingly, structural diversity appeared considerably lower than the observed sequence diversity. While mature peptide sequences varied substantially in amino acid composition, length, and physicochemical properties, many candidates converged toward a relatively limited repertoire of compact three-dimensional organizations. This pattern suggests that structural constraints may play a stronger role than primary sequence conservation in shaping toxin-like peptide evolution, preserving stable molecular architectures despite extensive sequence divergence. These observations motivated a comprehensive structural similarity analysis to determine whether the predicted peptides could be organized into recurrent structural families across Coleoptera.

3.4. Structural Convergence Reveals Recurrent Peptide Architectures

Unlike sequence-based comparisons, structural similarity analysis enables the identification of conserved protein architectures even among peptides sharing little primary sequence identity. Structural clustering revealed that the 291 predicted peptides could be organized into a limited number of recurrent structural groups, despite exhibiting substantial sequence variability. Many candidates repeatedly adopted similar three-dimensional organizations characterized by compact folds stabilized by short secondary-structure elements and disulfide-rich cores. These results indicate that structurally constrained peptide architectures recur across multiple beetle lineages, suggesting the existence of common structural solutions underlying toxin-like peptide evolution. The largest Foldseek-derived clusters contained recurrent architectures represented by dozens of candidates, with the predominant clusters displaying consistent Pfam annotations and generally high structural confidence (Figure S1). The highest-scoring PU-learning representative from each cluster was retained to facilitate structural interpretation and downstream visualization.
Several structural clusters included peptides originating from phylogenetically distant Coleoptera families, indicating that similar protein architectures are not restricted to closely related taxa. This observation is consistent with structural convergence, whereby distinct amino acid sequences independently adopt comparable three-dimensional organizations capable of maintaining molecular stability and potentially supporting similar biological functions. In addition, some structural groups were represented by only a few candidates, highlighting the presence of lineage-specific architectures that may correspond to novel structural scaffolds. These recurrent and unique folds reveal that the structural landscape of Coleoptera toxin-like peptides is characterized by both conservation and innovation, expanding the currently recognized diversity of small extracellular peptide architectures.
Structural similarity analyses demonstrate that the remarkable sequence diversity observed among the candidate peptides collapses into a comparatively restricted structural landscape. This reduction in structural complexity suggests that protein fold conservation represents a major organizing principle of toxin-like peptide diversity in Coleoptera and establishes the first structural atlas of these molecules across the order.

3.5. Structural Features Enable the Prioritization of High-Confidence Toxin-like Peptides

Following structural characterization, we integrated sequence-derived, physicochemical, and structural descriptors into a PU learning framework to prioritize candidates with the highest probability of representing functional toxin-like peptides (Figure 4). Unlike conventional binary classifiers, PU-learning is particularly suitable for biological discovery problems where experimentally validated positive examples are available, whereas true negative examples remain largely unknown.
Model performance indicated that the integration of sequence-derived, physicochemical, and structural descriptors supported effective candidate discrimination. Structural variables, including AlphaFold-derived confidence scores and three-dimensional descriptors, contributed to candidate prioritization, indicating that protein architecture provides complementary information beyond primary sequence composition. Feature importance analysis using SHAP revealed that structural characteristics consistently ranked among the most informative predictors driving candidate prioritization (Figure 4). Rather than relying on a single descriptor, the model integrated multiple independent sources of evidence, including secretion signals, peptide maturation, physicochemical properties, and structural features, allowing robust discrimination of high-confidence candidates.
The resulting prioritization identified a subset of top-ranked peptides that simultaneously exhibited canonical secretion signals, mature peptide lengths consistent with known toxin families, structurally stable AlphaFold models, and physicochemical properties compatible with extracellular bioactive peptides. These candidates represent the most promising molecules for future experimental validation and functional characterization. The prioritization framework demonstrates that combining structural information with interpretable machine learning presents an effective strategy for reducing the large search space generated by transcriptome mining while preserving molecular diversity. Rather than replacing biological interpretation, the machine learning model functions as a complementary decision-support tool that systematically integrates diverse lines of evidence into a unified confidence ranking.
We compiled the 20 highest-confidence candidates integrating secretion prediction, physicochemical properties, structural confidence, structural clustering, and Positive-Unlabeled learning scores (Table 2). Most top-ranked peptides displayed high AlphaFold 3 confidence (mean pLDDT > 85), multiple cysteine residues compatible with disulfide-rich architectures, and were consistently recovered across independent ensemble runs. Several candidates belonged to recurrent structural clusters shared among phylogenetically distant Coleoptera species, whereas others represented unique structural architectures, highlighting both conserved and potentially lineage-specific peptide scaffolds for future biochemical characterization.
Representative structures from the major Foldseek-derived clusters were visually compared (Figure 5). Despite originating from taxonomically distant beetle species, several candidates shared similar three-dimensional architectures, whereas others represented distinct structural scaffolds, illustrating the structural diversity recovered by the prioritization pipeline.

3.6. A Structural Atlas of Coleoptera Toxin-like Peptides

The integration of transcriptome mining, physicochemical characterization, structural prediction, structural similarity analyses, and machine learning prioritization resulted in the first large-scale structural atlas of predicted toxin-like peptides across Coleoptera (Figure 5; Table S3). Rather than representing an isolated collection of peptide sequences, the resulting dataset is a comprehensive resource linking sequence, biochemical properties, three-dimensional structure, structural similarity, and prioritization scores for each candidate.
The atlas encompasses peptides distributed across multiple beetle lineages and captures a broad spectrum of structural organizations, ranging from highly conserved compact architectures to lineage-specific structural variants. While recurrent structural folds dominated the dataset, several candidates occupied unique regions of the structural landscape, representing potentially novel peptide scaffolds that currently lack close structural counterparts. These molecules constitute promising targets for future structural, biochemical, and pharmacological investigations.
The prioritization framework further facilitates exploration of this structural diversity by identifying candidates that simultaneously exhibit canonical secretion signals, appropriate maturation patterns, structurally reliable models, and favorable physicochemical characteristics. Rather than limiting future studies to a small number of highly similar peptides, the atlas preserves the observed structural diversity while providing a rational strategy for selecting representative candidates from distinct structural groups. This resource establishes a foundation for future studies investigating the evolution, molecular function, and biotechnological potential of beetle-derived toxin-like peptides. By integrating complementary computational approaches into a unified structural framework, the atlas considerably expands the currently available repertoire of predicted peptide structures from Coleoptera.

Discussion

4.1. Coleoptera Harbor an Overlooked Structural Diversity of Toxin-like Peptides

To our knowledge, this is the first large-scale structural survey of predicted toxin-like peptides across Coleoptera, revealing that beetle transcriptomes harbor a considerably broader repertoire of secreted and structurally constrained peptides than previously recognized. Although beetles comprise nearly one-quarter of all described animal species, their peptide diversity has received remarkably limited attention compared with classical venomous organisms such as snakes, spiders, cone snails, scorpions, and marine cnidarians [5,36]. Consequently, the structural diversity of beetle-derived bioactive peptides has remained largely unexplored.
Our analyses indicate that this apparent lack of diversity likely reflects a historical sampling bias rather than a genuine biological absence of toxin-like peptides. By integrating transcriptome mining with structural prediction and comparative analyses, we identified hundreds of candidates displaying the hallmark characteristics of secreted bioactive peptides, including canonical signal peptides, conserved maturation features, cysteine-rich architectures, and compact three-dimensional structures. Collectively, these observations suggest that beetles represent an overlooked reservoir of structurally diverse peptide scaffolds with considerable evolutionary and biotechnological relevance.
Importantly, our findings should not be interpreted as evidence that all identified candidates function as toxins. Instead, the observed structural characteristics indicate that these peptides occupy a region of protein structural space commonly associated with extracellular bioactive molecules. This distinction is particularly important because similar structural frameworks may support diverse biological functions, including antimicrobial activity, immune regulation, signaling, enzyme inhibition, and defense. Thus, the structural atlas presented here substantially expands the repertoire of candidate peptide scaffolds available for future functional characterization while highlighting the value of transcriptome-scale structural exploration in taxonomically underrepresented animal groups.

4.2. Structural Convergence Shapes the Diversity of Coleoptera Bioactive Peptide Scaffolds

One of the most striking observations of this study is that the extensive sequence diversity identified across Coleoptera collapses into a comparatively restricted repertoire of three-dimensional protein architectures. Although the predicted mature peptides differed substantially in amino acid composition, length, and physicochemical properties, structural analyses consistently revealed the recurrence of compact folds stabilized by similar secondary-structure arrangements and cysteine-rich cores. This apparent reduction in structural complexity suggests that protein architecture is considerably more conserved than primary sequence during the diversification of these peptides. Similar observations have increasingly emerged from large-scale structural analyses, where protein folds remain conserved despite extensive sequence divergence, reinforcing the concept that structural space is considerably more constrained than sequence space [11,13].
Structural convergence has long been recognized as a recurrent feature of peptide evolution. In numerous venomous organisms, peptides with little or no detectable sequence similarity frequently adopt nearly identical three-dimensional folds while targeting similar molecular processes [5,7]. More recently, comparative genomic and structural studies have shown that similar molecular architectures repeatedly emerge across evolution through independent recruitment, co-option, and functional diversification of ancestral protein scaffolds [9,37,38]. Such convergence reflects strong evolutionary constraints imposed by protein folding, structural stability, and functional requirements, allowing diverse amino acid sequences to preserve common molecular architectures while acquiring distinct biological functions. Our results suggest that a similar phenomenon may also occur across Coleoptera, despite these insects not being traditionally regarded as major sources of peptide toxins.
The structural gallery (Figure 5) further illustrates that multiple high-confidence candidates converge into recurrent three-dimensional architectures despite their taxonomic diversity, supporting the hypothesis that structural conservation may exceed sequence conservation among Coleoptera toxin-like peptides. This observation reinforces the value of integrating structure-based analyses with machine learning for large-scale peptide discovery.
The predominance of cysteine-rich candidates further supports this interpretation. Disulfide bonds represent one of the most efficient evolutionary solutions for stabilizing small extracellular peptides, increasing resistance to proteolytic degradation while maintaining well-defined conformations compatible with receptor recognition and molecular specificity [4,5]. Recent advances in peptide engineering further demonstrate that conserved cysteine motifs and oxidative folding pathways constitute fundamental determinants of structural stability and molecular recognition, often preserving protein architecture even after extensive sequence diversification [6]. Likewise, computational and structural studies have increasingly highlighted that disulfide-rich peptide scaffolds provide highly adaptable frameworks capable of supporting multiple biological activities while maintaining remarkable structural robustness [12].
Structural convergence should not be interpreted as evidence of functional equivalence. Similar protein folds can support a wide variety of biological activities, including antimicrobial defense, immune modulation, signaling, enzyme inhibition, and toxin activity. Indeed, recent studies have demonstrated extensive functional plasticity among structurally related peptide families, with similar scaffolds being repeatedly recruited for distinct physiological and pathological processes [2,3]. Rather than predicting a single biological role, the conserved architectures identified here likely represent versatile bioactive peptide scaffolds that have been repeatedly co-opted for distinct physiological functions throughout beetle evolution.
Taken together, these results support the view that the diversity of Coleoptera bioactive peptides is organized around a relatively limited number of recurrent protein architectures. This observation shifts the emphasis from sequence conservation toward structural conservation as a more informative framework for exploring peptide evolution. More broadly, our results reinforce the growing importance of structure-guided approaches for peptide discovery, in which advances in protein structure prediction and comparative structural bioinformatics enable the identification of evolutionarily conserved scaffolds that would remain undetected using sequence-based analyses alone [11,12,13].

4.3. Structural Information Enhances Computational Discovery and Prioritization of Bioactive Peptides

A major advantage of the framework presented here lies in the integration of structural information into the computational discovery process. Traditional peptide mining strategies rely predominantly on sequence similarity, conserved motifs, or physicochemical descriptors, which often fail to detect highly divergent peptide families sharing common structural architectures. By incorporating AlphaFold-derived models, structural descriptors, and Foldseek-based comparisons (Figure 3 and Figure 4), our approach substantially expands the search space beyond primary sequence conservation, allowing the identification of candidates that would likely remain undetected using conventional homology-based strategies. Recent advances in structure prediction have similarly demonstrated that structural information is becoming an essential component of protein annotation and functional inference, particularly for rapidly evolving protein families [11,12].
The explainable machine learning framework further demonstrated that structural descriptors contributed substantially to candidate prioritization. Rather than depending on a single predictive feature, the model integrated complementary sources of evidence, including secretion signals, precursor maturation, physicochemical properties, and three-dimensional structural characteristics. Feature attribution analyses based on SHAP confirmed that several structural variables ranked among the most influential predictors driving model decisions (Figure 4), highlighting that structural information captures biologically meaningful patterns not readily identifiable from sequence alone [26,27]. This observation agrees with recent studies showing that combining structural modeling with interpretable machine learning considerably improves the discovery of bioactive peptides compared with sequence-only approaches [12,39].
An additional strength of our strategy is the use of PU learning. Unlike conventional supervised classification, PU-learning is specifically designed for discovery problems in which experimentally validated positive examples are available but reliable negative datasets are largely absent. This characteristic makes the approach particularly suitable for peptide discovery, where the absence of functional annotation cannot be interpreted as evidence of inactivity. Consequently, the prioritization framework shows probabilistic estimates of confidence rather than definitive functional assignments, reducing the risk of introducing biases associated with artificially generated negative datasets [25]. Moreover, because peptide discovery typically involves highly imbalanced datasets, the combined interpretation of ROC and Precision-Recall curves (Figure 54 displays a more informative assessment of classifier performance than either metric alone [24].
The computational framework should not be viewed as a substitute for experimental validation but rather as an efficient strategy for reducing an otherwise overwhelming search space. The structural atlas generated in this study (Table S3) contains hundreds of candidates displaying diverse architectures and biochemical characteristics, making exhaustive functional characterization impractical. By integrating structural prediction, structural similarity, and explainable machine learning into a unified prioritization workflow, our approach enables the rational selection of representative candidates spanning distinct structural groups, thereby maximizing structural diversity while minimizing experimental effort. These findings reinforce the growing paradigm that protein structure provides a more robust framework than sequence similarity alone for guiding the discovery of novel bioactive peptides. As high-accuracy structural prediction and comparative structural bioinformatics continue to advance, integrative approaches such as the one presented here are expected to play an increasingly central role in large-scale peptide discovery and functional annotation across underexplored taxa [11,12,13].

4.4. Biological Significance and Biotechnological Opportunities of the Structural Atlas

Beyond expanding our understanding of peptide diversity in Coleoptera, the structural atlas generated in this study (Table S3) is a valuable resource for exploring the evolutionary and functional diversity of small secreted peptides. Historically, the discovery of bioactive peptides has focused on a relatively limited number of well-characterized venomous organisms, whereas the enormous diversity of beetles has remained largely unexplored [8]. Our results demonstrate that transcriptome-wide structural mining can substantially broaden the repertoire of candidate peptide scaffolds available for future investigation.
An important implication of our findings is that structurally related peptides should not necessarily be expected to perform identical biological functions. Numerous studies have shown that highly similar peptide scaffolds may be repeatedly recruited during evolution for distinct physiological roles, including antimicrobial defense, immune modulation, hormone signaling, enzyme inhibition, and toxin activity [2,3,9]. Consequently, the structural conservation observed across many of our candidates likely reflects the preservation of robust molecular frameworks rather than strict conservation of biological function. This functional plasticity considerably expands the range of potential applications associated with the identified peptide families.
From a biotechnological perspective, compact disulfide-rich peptide scaffolds represent particularly attractive molecular templates due to their exceptional structural stability, resistance to proteolytic degradation, and capacity to tolerate sequence diversification while maintaining a conserved three-dimensional framework [4,5,6]. These properties have stimulated increasing interest in their application as starting points for therapeutic peptides, molecular recognition molecules, enzyme inhibitors, antimicrobial agents, and environmentally friendly bioinsecticides [3,12]. The diversity of structural architectures identified in the present study therefore considerably enlarges the repertoire of candidate scaffolds that may be explored for future biotechnology-driven optimization.
Equally important, the atlas provides an objective framework for selecting representative candidates for experimental validation. Rather than prioritizing peptides solely according to sequence similarity, the integration of structural clustering, AlphaFold 3 confidence estimates, Foldseek comparisons, and explainable machine learning (Figure 3, Figure 4 and Figure 5) enables the rational selection of structurally diverse representatives occupying distinct regions of the identified structural landscape. Such an approach maximizes the probability of capturing novel molecular functions while minimizing redundancy during downstream biochemical characterization. These results demonstrate that large-scale structure-guided transcriptome mining represents a powerful strategy for accelerating peptide discovery in taxonomically underexplored organisms

4.5. Structural Bioinformatics Expands Functional Annotation Beyond Sequence Similarity

The rapid development of highly accurate protein structure prediction methods has substantially expanded the possibilities of computational biology. Until recently, large-scale structural characterization represented a major bottleneck for protein annotation, particularly in non-model organisms lacking experimentally determined structures. The emergence of AlphaFold 3 and related approaches has shifted this paradigm by enabling routine prediction of thousands of high-confidence protein structures, making the extraction of biological knowledge from these structural datasets the new computational challenge [11]. In this context, structural comparison algorithms such as Foldseek have become increasingly valuable because they enable the organization of large structural datasets according to three-dimensional similarity rather than primary sequence conservation.
Our workflow illustrates the advantages of integrating transcriptome mining, structure prediction, structural comparison, and interpretable machine learning into a unified framework for peptide discovery. We incorporated structural descriptors and Foldseek-derived relationships throughout the prioritization process, allowing candidate peptides to be evaluated within a broader structural context. This strategy enabled structurally related peptides originating from phylogenetically distant beetle lineages to be grouped into recurrent protein architectures, while simultaneously preserving lineage-specific structural variants represented in the structural atlas (Figure 6; Figure S1). These observations reinforce that three-dimensional protein architecture has biologically meaningful information that complements conventional sequence-derived descriptors.
An important consequence of this integrative strategy is that structural information becomes useful not only for describing individual peptide models but also for organizing large repertoires of predicted proteins into biologically interpretable structural families. Because short extracellular peptides frequently evolve rapidly while maintaining conserved structural scaffolds, analyses based exclusively on sequence similarity may underestimate functional relationships among highly divergent candidates. By combining AlphaFold-derived models with Foldseek structural clustering, our framework identifies recurrent molecular architectures that remain detectable despite extensive sequence diversification, facilitating the recognition of representative scaffolds suitable for downstream experimental characterization.
Beyond the discovery of toxin-like peptides in Coleoptera, the computational strategy presented here provides a general framework for large-scale functional exploration of transcriptomic datasets from non-model organisms. As public repositories continue to expand and accurate structure prediction becomes increasingly accessible, workflows integrating transcriptomics, structural prediction, structural comparison, and explainable artificial intelligence are expected to play an increasingly important role in protein annotation, candidate prioritization, and the identification of previously unrecognized protein families. In this regard, the structural atlas generated in this study represents not only a catalog of predicted Coleoptera peptide structures but also a proof of concept demonstrating how structural bioinformatics can substantially enhance large-scale peptide discovery and functional hypothesis generation.

4.6. Limitations and Future Perspectives

Although the present study substantially expands the structural landscape of predicted bioactive peptides across Coleoptera, several limitations should be acknowledged. First, the identified candidates were inferred from transcriptomic data and computational analyses rather than experimentally validated at the protein level. Consequently, peptide expression, post-translational processing, structural stability under physiological conditions, and biological activities remain to be confirmed through complementary proteomic, biochemical, and functional assays. Nevertheless, recent advances in transcriptome-guided peptide discovery and high-accuracy structural prediction have demonstrated that integrative computational approaches display an efficient strategy for prioritizing candidates for downstream validation while substantially reducing experimental search space [11,12].
A second limitation concerns functional annotation. Similar protein architectures do not necessarily imply identical biological functions, and structurally related peptides may participate in diverse physiological processes, including immune defense, signaling, enzyme inhibition, and toxin activity [2,3,9]. Accordingly, the structural atlas presented here should be interpreted as a resource for hypothesis generation rather than definitive functional annotation. Future integration with transcriptomic expression profiles, comparative proteomics, receptor-binding studies, molecular dynamics simulations, and experimental functional assays will provide a more comprehensive understanding of the biological roles of these peptide families.
Despite these limitations, the framework developed in this study establishes a scalable strategy for large-scale exploration of peptide diversity in non-model organisms. As structural databases continue to expand and artificial intelligence-based prediction methods become increasingly accurate, comparative structural analyses will likely become a central component of peptide discovery pipelines. Extending this framework to additional insect orders and other underexplored metazoan lineages will enable increasingly comprehensive structural atlases, facilitating comparative evolutionary analyses and accelerating the identification of novel bioactive peptide scaffolds with potential applications in biotechnology and drug discovery [11,12,13].

Conclusions

Here, we present the first large-scale structural atlas of predicted toxin-like peptide scaffolds across Coleoptera, revealing that beetle transcriptomes harbor a substantially broader repertoire of structurally constrained bioactive peptides than previously recognized. Despite extensive sequence divergence, many candidates converged toward a comparatively limited set of recurrent protein architectures, suggesting that structural conservation represents a major organizing principle of peptide diversity in this insect order. These findings expand current knowledge of peptide evolution and demonstrate that protein structure offers a more informative framework than primary sequence alone for exploring hidden bioactive peptide diversity.
By integrating transcriptome mining, peptide maturation prediction, structural modeling, comparative structural analyses, and interpretable machine learning, we establish a scalable framework for structure-guided peptide discovery in underexplored taxa. Beyond providing a comprehensive catalog of high-confidence candidates, the structural atlas generated here represents a valuable resource for future evolutionary, structural, biochemical, and biotechnological investigations. Collectively, our results reinforce the growing importance of structure-based approaches for uncovering novel bioactive peptide scaffolds and provide a foundation for accelerating peptide discovery across the vast unexplored diversity of animal life.

Supplementary Materials

The following supporting information can be downloaded at the website of this paper posted on Preprints.org, Figure S1: Summary of the major Foldseek-derived structural clusters identified among Coleoptera toxin-like peptide candidates; Table S1: Metadata for the Coleoptera transcriptomes analyzed in this study; Table S2: Top 100 toxin-like peptide candidates prioritized by the ensemble PU-learning pipeline; Table S3: Structural atlas and associated molecular descriptors of the identified candidates.

Funding

This study was supported by the São Paulo Research Foundation (FAPESP) (grant no. 2023/05589-4.

CRediT Author Statement

TCG: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Writing - original draft. JAT: Data curation, Investigation, Methodology, Software, Validation, Writing - review & editing. DTA: Conceptualization, Funding acquisition, Project administration, Resources, Supervision, Methodology, Visualization, Writing - review & editing.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

The processed datasets, predicted structural models, analysis scripts, structural-clustering results, and source data supporting this study are available through the GitHub repository (https://github.com/BBMDO/coleotoxmin) and the associated Zenodo record (DOI 10.5281/zenodo.21386392). Raw transcriptomic datasets were obtained from public repositories, and their accession numbers are provided in Table S1.

Conflicts of Interest

The authors declare no conflicts of interest.

Declaration of generative AI use

During the preparation of this work, the authors used ChatGPT (OpenAI) to assist with language editing, improvement of text clarity, and manuscript organization. All scientific content, interpretation of the literature, critical revision, and final approval of the manuscript were performed by the authors, who take full responsibility for the content of this publication.

References

  1. Besharati, M.; Lackner, M. Bioactive peptides: A review. EuroBiotech J. 2023, 7, 176–188. [Google Scholar] [CrossRef]
  2. Tang, Y.; Zhang, Y.; Zhang, D.; Liu, Y.; Nussinov, R.; Zheng, J. Exploring pathological link between antimicrobial and amyloid peptides. Chem. Soc. Rev. 2024, 53, 8713–8763. [Google Scholar] [CrossRef] [PubMed]
  3. Gao, B.; Yang, N.; Teng, D.; Hao, Y.; Wang, J.; Mao, R. Marine Antimicrobial Peptides: Advances in Discovery, Multifunctional Mechanisms, and Therapeutic Translation Challenges. Mar. Drugs 2025, 23, 463. [Google Scholar] [CrossRef] [PubMed]
  4. Craik, D. J.; Daly, N. L.; Bond, T.; Waine, C. Plant cyclotides: A unique family of cyclic and knotted proteins that defines the cyclic cystine knot structural motif. J. Mol. Biol. 1999, 294, 1327–1336. [Google Scholar] [CrossRef] [PubMed]
  5. King, G. F.; Hardy, M. C. Spider-Venom Peptides: Structure, Pharmacology, and Potential for Control of Insect Pests. Annu. Rev. Entomol. 2013, 58, 475–496. [Google Scholar] [CrossRef] [PubMed]
  6. Wu, C. Motif-directed oxidative folding to design and discover multicyclic peptides for protein recognition. Acc. Chem. Res. 2025, 58, 1620–1631. [Google Scholar] [CrossRef] [PubMed]
  7. Kini, M.; Doley, R. Excitement ahead: Structure, function and mechanism of snake venom phospholipase A2 enzymes. Toxicon 2003, 42, 827–840. [Google Scholar] [CrossRef] [PubMed]
  8. Linial, M.; Rappoport, N.; Ofer, D. Overlooked short toxin-like proteins: a shortcut to drug design. Toxins 2017, 9, 350. [Google Scholar] [CrossRef] [PubMed]
  9. Raoelijaona, F.; Szczepaniak, J.; Schahl, A.; Bray, J. E.; Zhou, J. C.; Baker, L.; Seiradake, E. Ancestral neuronal receptors are bacterial accessory toxins. Nat. Commun. 2026. [Google Scholar] [CrossRef] [PubMed]
  10. van Kempen, M.; Kim, S. S.; Tumescheit, C.; Mirdita, M.; Lee, J.; Gilchrist, C. L.; Steinegger, M. Fast and accurate protein structure search with Foldseek. Nat. Biotechnol. 2024, 42, 243–246. [Google Scholar]
  11. Abramson, J.; Adler, J.; Dunger, J.; Evans, R.; Green, T.; Pritzel, A.; Jumper, J. M. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 2024, 630, 493–500. [Google Scholar] [CrossRef] [PubMed]
  12. Agoni, C.; Fernández-Díaz, R.; Timmons, P. B.; Adelfio, A.; Gómez, H.; Shields, D. C. Molecular modelling in bioactive peptide discovery and characterisation. Biomolecules 2025, 15, 524. [Google Scholar] [CrossRef] [PubMed]
  13. Gonçalves, T. C.; Cortez, E. H. T.; Amaral, D. T. A Structure-Informed Atlas of Venom-Derived Peptides Reveals the Organization of Chemical Space. Mol. Inform. 2026, 45, e70039. [Google Scholar] [CrossRef] [PubMed]
  14. Grabherr, M. G.; Haas, B. J.; Yassour, M.; Levin, J. Z.; Thompson, D. A.; Amit, I.; Regev, A. Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nat. Biotechnol. 2011, 29, 644–652. [Google Scholar] [CrossRef] [PubMed]
  15. Li, W.; Godzik, A. Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics 2006, 22, 1658–1659. [Google Scholar] [CrossRef] [PubMed]
  16. Almagro Armenteros, J. J.; Tsirigos, K. D.; Sønderby, C. K.; Petersen, T. N.; Winther, O.; Brunak, S.; Nielsen, H. SignalP 5.0 improves signal peptide predictions using deep neural networks. Nat. Biotechnol. 2019, 37, 420–423. [Google Scholar] [CrossRef] [PubMed]
  17. Hegde, R. S.; Bernstein, H. D. The surprising complexity of signal sequences. Trends Biochem. Sci. 2006, 31, 563–571. [Google Scholar] [CrossRef] [PubMed]
  18. Boon, L.; Ugarte-Berzal, E.; Vandooren, J.; Opdenakker, G. Protease propeptide structures, mechanisms of activation, and functions. Crit. Rev. Biochem. Mol. Biol. 2020, 55, 111–165. [Google Scholar] [CrossRef] [PubMed]
  19. Löwik, D. W.; van Hest, J. C. Peptide based amphiphiles. Chem. Soc. Rev. 2004, 33, 234–245. [Google Scholar] [CrossRef] [PubMed]
  20. Eddy, S. R. Accelerated profile HMM searches. PLoS Comput. Biol. 2011, 7, e1002195. [Google Scholar] [CrossRef] [PubMed]
  21. Hekkelman, M. L.; Salmoral, D. Á.; Perrakis, A.; Joosten, R. P. DSSP 4: FAIR annotation of protein secondary structure. Protein Sci. 2025, 34, e70208. [Google Scholar] [CrossRef] [PubMed]
  22. Mitternacht, S. FreeSASA: An open source C library for solvent accessible surface area calculations. F1000Research 2016, 5, 189. [Google Scholar] [CrossRef] [PubMed]
  23. McGibbon, R. T.; Beauchamp, K. A.; Harrigan, M. P.; Klein, C.; Swails, J. M.; Hernández, C. X.; Pande, V. S. MDTraj: a modern open library for the analysis of molecular dynamics trajectories. Biophys. J. 2015, 109, 1528–1532. [Google Scholar] [CrossRef] [PubMed]
  24. Kramer, O. Scikit-learn. In Machine learning for evolution strategies; Springer International Publishing: Cham, 2016; pp. 45–53. [Google Scholar]
  25. Saito, T.; Rehmsmeier, M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS ONE 2015, 10, e0118432. [Google Scholar] [CrossRef] [PubMed]
  26. Gajda, S.; Chlebus, M. A probability-based models ranking approach: an alternative method of machine-learning model performance assessment. Sensors 2022, 22, 6361. [Google Scholar] [CrossRef] [PubMed]
  27. Lundberg, S. M.; Lee, S. I. A unified approach to interpreting model predictions. Adv. Neural Inf. Process. Syst. 2017, 30. [Google Scholar]
  28. Lundberg, S. M.; Erion, G.; Chen, H.; DeGrave, A.; Prutkin, J. M.; Nair, B.; Lee, S. I. From local explanations to global understanding with explainable AI for trees. Nat. Mach. Intell. 2020, 2, 56–67. [Google Scholar] [CrossRef] [PubMed]
  29. Chamberlain, S. A.; Szöcs, E. taxize: taxonomic search and retrieval in R. F1000Research 2013, 2, 191. [Google Scholar] [CrossRef] [PubMed]
  30. R Core Team. R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria; 2025. Available online: <u>https://www.R-project.org/.
  31. Fisher, R. A. On the interpretation of χ 2 from contingency tables, and the calculation of P. J. R. Stat. Soc. 1922, 85, 87–94. [Google Scholar] [CrossRef]
  32. Benjamini, Y.; Hochberg, Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J. R. Stat. Soc. Ser. B (Methodological) 1995, 57, 289–300. [Google Scholar] [CrossRef]
  33. Yarberry, W. Dplyr. CRAN recipes: DPLYR, stringr, lubridate, and regex in R; Apress: Berkeley, CA, 2021; pp. 1–58. [Google Scholar]
  34. Wickham, H.; Wickham, M. H. Package ‘tidyr’. Easily Tidy Data With’spread’and’gather () ’Functions 2017, 10. [Google Scholar]
  35. Coassolo, L.; Wiggenhorn, A.; Svensson, K. J. Understanding peptide hormones: from precursor proteins to bioactive molecules. Trends Biochem. Sci. 2025. [Google Scholar] [CrossRef] [PubMed]
  36. Fry, B. G.; Roelants, K.; Champagne, D. E.; Scheib, H.; Tyndall, J. D. A.; King, G. F.; Nevalainen, T. J.; Norman, J. A.; Lewis, R. J.; Norton, R. S.; Renjifo, C.; De La Vega, R. C. R. The Toxicogenomic Multiverse: Convergent Recruitment of Proteins Into Animal Venoms. Annu. Review Genom. Hum. Genet.> 2009, 10, 483–511. [Google Scholar] [CrossRef]
  37. Van Thiel, J.; Khan, M. A.; Wouters, R. M.; Harris, R. J.; Casewell, N. R.; Fry, B. G.; Richardson, M. K. Convergent evolution of toxin resistance in animals. Biol. Rev. 2022, 97, 1823–1843. [Google Scholar] [CrossRef] [PubMed]
  38. de Oliveira, J. L.; Roman-Ramos, H. Coevolution Between Three-Finger Toxins and Target Receptors. Receptors 2026, 5, 7. [Google Scholar] [CrossRef]
  39. Beltrán, J. F.; Herrera-Belén, L.; Parraguez-Contreras, F.; Farías, J. G.; Machuca-Sepúlveda, J.; Short, S. MultiToxPred 1.0: a novel comprehensive tool for predicting 27 classes of protein toxins using an ensemble machine learning approach. BMC Bioinform. 2024, 25, 148. [Google Scholar] [CrossRef]
Figure 1. Distributions of predicted maturation features across candidate peptides, including (A) mature peptide length, (B) dibasic cleavage motifs, and (C) cysteine content.
Figure 1. Distributions of predicted maturation features across candidate peptides, including (A) mature peptide length, (B) dibasic cleavage motifs, and (C) cysteine content.
Preprints 224180 g001
Figure 2. Violin plots summarizing the distribution of key sequence-derived features across candidate peptides identified in Coleoptera transcriptomes. (A) Peptide length distribution. (B) Net charge distribution at pH 7. (C) Cysteine count distribution. Each panel includes the median, interquartile range, and summary statistics.
Figure 2. Violin plots summarizing the distribution of key sequence-derived features across candidate peptides identified in Coleoptera transcriptomes. (A) Peptide length distribution. (B) Net charge distribution at pH 7. (C) Cysteine count distribution. Each panel includes the median, interquartile range, and summary statistics.
Preprints 224180 g002
Figure 3. Structural clustering and reproducibility of candidate peptides. (A) Relationship between cluster size and mean structural confidence (pLDDT), showing that recurrent folds are supported by consistently high-confidence predictions. (B) Pairwise Jaccard overlap of top-ranked candidates across independent runs, indicating reproducible identification of a core subset of peptides.
Figure 3. Structural clustering and reproducibility of candidate peptides. (A) Relationship between cluster size and mean structural confidence (pLDDT), showing that recurrent folds are supported by consistently high-confidence predictions. (B) Pairwise Jaccard overlap of top-ranked candidates across independent runs, indicating reproducible identification of a core subset of peptides.
Preprints 224180 g003
Figure 4. Performance of the PU-learning framework showing (A) model evaluation metrics, (B) ROC and Precision–Recall curves, (C) SHAP feature importance, and (D) prioritization of high-confidence toxin-like peptide candidates.
Figure 4. Performance of the PU-learning framework showing (A) model evaluation metrics, (B) ROC and Precision–Recall curves, (C) SHAP feature importance, and (D) prioritization of high-confidence toxin-like peptide candidates.
Preprints 224180 g004aPreprints 224180 g004b
Figure 5. Representative AlphaFold-predicted structures from the major Foldseek structural clusters. Each panel represents one of the 12 largest Foldseek structural clusters and shows the representative peptide selected as the candidate with the highest ensemble PU score within that cluster. Structures are displayed as ribbon diagrams colored by per-residue pLDDT confidence scores, ranging from low (red) to high (blue). Panels A-L correspond to clusters 1-12, respectively.
Figure 5. Representative AlphaFold-predicted structures from the major Foldseek structural clusters. Each panel represents one of the 12 largest Foldseek structural clusters and shows the representative peptide selected as the candidate with the highest ensemble PU score within that cluster. Structures are displayed as ribbon diagrams colored by per-residue pLDDT confidence scores, ranging from low (red) to high (blue). Panels A-L correspond to clusters 1-12, respectively.
Preprints 224180 g005
Figure 6. Stability of top-ranked candidates across independent runs. (A) Number of seeds in which each candidate appears among the top 20, highlighting highly recurrent peptides. (B) Distribution of candidate recurrence across seeds, showing the presence of a stable subset consistently identified by the model.
Figure 6. Stability of top-ranked candidates across independent runs. (A) Number of seeds in which each candidate appears among the top 20, highlighting highly recurrent peptides. (B) Distribution of candidate recurrence across seeds, showing the presence of a stable subset consistently identified by the model.
Preprints 224180 g006
Table 1. Summary of candidate discovery and maturation filtering in Coleoptera transcriptomes.
Table 1. Summary of candidate discovery and maturation filtering in Coleoptera transcriptomes.
Step N (%)
Initial candidates 291 100.0
Predicted secreted (Signal peptide) 155 53.3
Valid mature peptide length (8-80 aa) 273 93.8
Out-of-range mature peptides 18 6.2
Candidates with dibasic cleavage motifs 127 43.6
Table 2. Top 20 toxin-like peptide candidates prioritized by the ensemble PU learning framework. Candidates are ranked according to their ensemble PU score and represent the highest-confidence predictions obtained across five independent training seeds. For each peptide, the table summarizes the species of origin, amino acid sequence, peptide length, net charge at pH 7, cysteine content, predicted number of disulfide bonds, Foldseek-derived structural cluster, mean AlphaFold pLDDT confidence score, ensemble PU score, and ranking stability across independent model runs.
Table 2. Top 20 toxin-like peptide candidates prioritized by the ensemble PU learning framework. Candidates are ranked according to their ensemble PU score and represent the highest-confidence predictions obtained across five independent training seeds. For each peptide, the table summarizes the species of origin, amino acid sequence, peptide length, net charge at pH 7, cysteine content, predicted number of disulfide bonds, Foldseek-derived structural cluster, mean AlphaFold pLDDT confidence score, ensemble PU score, and ranking stability across independent model runs.
Rank Candidate Species Sequence Length (aa) Net charge (pH 7) Cysteines (n) Disulfide bonds (n) Fold cluster Mean pLDDT PU score Seed stability
1 diphucephala_sp_2 Diphucephala sp MKFIYFLLVLVILIVSTVVAPPPCGENEERKTCGPACHPTCANPNVSTVSCPKPCISGCFCKADFLTNSKGKCVPKNQCS 80 4.1 10 5 geotrupes_spiniger_2 91.08 9.98 5/5
2 eophileurus_sp_1 Eophileurus sp MLCAGLAKGGKDACQGDSGGPMVTTGKLTGVISWGIGCAREGYPGIYTQVSKFRNWIKQNSNI 63 4.0 3 1 dastarcus_helophoroides_2 94.48 9.90 5/5
3 aromia_moschata_1 Aromia moschata MSNVECRKTGYGYKITDNMLCAGYHDGKKDSCQGDSGGPLHVVNGSVHQVVGIVSWGEGCAQANFPGVYTRVNRYISWIKSNTRDACYC 89 2.3 6 2 dastarcus_helophoroides_2 93.15 9.90 5/5
4 platycerus_caraboides_1 Platycerus caraboides MFCAGWKSGIADTCAGDSGGGLMCPVNRTLISTAYAVQGITSFGDGCGRKNKYGIYTKVNNYLKWIQNTIEKYS 74 4.0 4 1 dastarcus_helophoroides_2 93.62 9.90 4/5
5 platycerus_caraboides_2 Platycerus caraboides MKVRLNLFQNSRCDKAYKGQYFPNGLPKTMMCVGELAGGKDTCQGDSGGPITITKSDDPCVFYTVGITSFGKACAAENTPAVYTRVSEFVSWIEKTVW 98 2.0 5 2 dastarcus_helophoroides_2 93.13 9.89 3/5
6 onthophagus_curvicornis_3 Onthophagus curvicornis MICAKDLTGGRKDTCQGDSGSGFVVDGKLFGITSWGIGCGRLNKAGVYTKVSLYREWIKLYTKV 64 5.0 3 1 dastarcus_helophoroides_2 91.56 9.88 5/5
7 microcara_testacea_2 Microcara testacea MVCAGDIKEAKDTCQGDSGGPIVVTDKKNQCLFHVIGVTSFGKGCGIKLPAIYTRVSSFVPWIESVVWP 69 1.1 4 1 dastarcus_helophoroides_2 91.46 9.86 4/5
8 onthophagus_curvicornis_1 Onthophagus curvicornis MKLFYLVVIFAALLASVFAAQTCGPNEEYRTCGSACEPTCAQKNPRACAFNCIPGCYCKSGYLHDSKGQCVKEKDC 76 2.1 10 5 geotrupes_spiniger_2 87.9 9.86 4/5
9 dichotomius_satanas_3 Dichotomius satanas MKYRASRITNYMLCAGRGSHDSCQGDSGGPLIINNGERYEIVGIVSWGVGCGRPGYPGVYTRIAKYISWLKYNLEDACLC 80 3.1 5 1 dastarcus_helophoroides_2 93.05 9.85 4/5
10 sisyphus_schaefferi_1 Sisyphus schaefferi MCAGESAGGKDACQGDSGGPLVAGGKLRGIVSWGYGCARPAYPGVYASVSNLRSYITQVAGI 62 2.0 3 1 dastarcus_helophoroides_2 94.05 9.83 2/5
11 hydroscapha_redfordi_2 Hydroscapha redfordi MMCAGQDNRDSCSGDSGGPLMINNGRWVQVGVVSWGIGCGKGQYPGVYSRVTSFLSWIVKNLK 63 3.0 3 1 dastarcus_helophoroides_2 94.39 9.83 3/5
12 ceutorhynchus_napi_2 Ceutorhynchus napi MMCAGYKNGGRDSCQGDSGGPLMLQKTGRWFLIGIVSAGYSCAQAGQPGIYHRVAHTVDWITRAIGS 67 3.2 3 1 dastarcus_helophoroides_2 94.25 9.83 3/5
13 tenebrio_molitor_1 Tenebrio molitor MMCAGYKNGGRDSCQGDSGGPLMLQKQGRWFLIGIVSAGYSCAQPGQPGIYHRVAHTVDWITRAIGV 67 3.2 3 1 dastarcus_helophoroides_2 94.32 9.81 2/5
14 pleurophorus_caesus_1 Pleurophorus caesus MLCAGEANKDSCSGDSGGPLMITNQQGRYVQAGVVSWGIGCGKGQYPGVYSRVESFLPWINKNLKD 66 1.0 3 1 dastarcus_helophoroides_2 93.94 9.80 3/5
15 neoserica_sp_2 Neoserica sp MICAGFTAGGRDACQGDSGGPLVVGNTLVGIVSWGHGCAKPNFPGVYACVGNLRNWIRTNSGV 63 2.1 4 1 dastarcus_helophoroides_2 93.91 9.80 2/5
16 dichotomius_satanas_5 Dichotomius satanas MICAGFAAGGRDACQGDSGGPLAVGNTLVGVVSWGRGCARPNFPGVYACTGNLRNWISSATGI 63 2.0 4 1 dastarcus_helophoroides_2 94.5 9.80 2/5
17 dichotomius_satanas_1 Dichotomius satanas MKLFYLVVIFAALLASVFAAQTCGPNEEYRTCGSACEPTCAQKNPRACAFNCIPGCYCKSGYLHDSKGQCVKEKDC 76 2.1 10 5 geotrupes_spiniger_2 87.45 9.79 4/5
18 alurnus_ornatus_2 Alurnus ornatus MFWVSAANTWCSKYRAESLKLRILTGYHSPGQFRVQGPLSNSDDFASDFGCPVGSKMNPEKKCRIW 66 4.1 3 0 elateroides_flabellicornis_2 94.19 9.79 2/5
19 micromalthus_debilis_1 Micromalthus debilis MIPVVSIDECKRAYANFKTTTIDQRVICAGYAKGGKDACQGDSGGPLMWGKTQASSTSLTYYLIGVVSYGFRCAEEGYPGVYSRVTQFVDWIQKNLN 97 2.0 4 2 dastarcus_helophoroides_2 94.1 9.79 3/5
20 spercheus_emarginatus_3 Spercheus emarginatus MFLRAGRHEFIPDIFLCAGHEGGGRDSCQGDSGGPLQVKGKDGRYFLAGIISWGIGCAEANLPGVCTRISKFVPWILKNVK 81 3.2 4 1 dastarcus_helophoroides_2 92.92 9.78 2/5
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings