Submitted:
14 September 2026
Posted:
15 September 2026
You are already at the latest version
Abstract
Background/Objectives: The global crisis of antimicrobial resistance and limited discov-ery of new antimicrobial classes have intensified the search for alternative sources of bio-active natural products. Environmental bacteria harbor largely unexplored biosynthetic diversity, much of which is encoded in biosynthetic gene clusters (BGCs). This review examines current strategies for genome-guided BGC bioprospecting, emphasizing antimi-crobial and environmental applications, and the transition from computational prediction to experimental validation. Methods: We review advances in bacterial genome mining, protein structure prediction, molecular docking, and experimental approaches for evalu-ating antimicrobial and bioremediation potential. Emphasis is placed on representative environmental bacterial genera and the integration of genomic, structural, metabolomic, microbiological, and environmental approaches. Results: Genome mining has revealed substantial biosynthetic potential across diverse environmental bacteria. Comparative ge-nomics and BGC similarity analyses support dereplication and prioritization, while AI-based approaches can expand the search toward atypical and poorly characterized bi-osynthetic systems. Protein structure prediction and molecular docking provide mecha-nistic hypotheses but cannot independently demonstrate metabolite production or biolog-ical activity. Experimental approaches, including One Strain-Many Compounds, heterol-ogous expression, metabolomics, functional genetics, antimicrobial assays, and environ-mentally relevant microcosms, are therefore essential for linking genomic potential to me-tabolite production and biological function. Conclusions: BGC-driven bioprospecting pro-vides a powerful framework for expanding microbial natural product discovery beyond conventional screening. Its greatest potential lies in multi-omics approaches and experi-mental validation. Future progress will depend on improving the prioritization and acti-vation of cryptic BGCs and establishing stronger gene-to-metabolite-to-function relation-ships, ultimately supporting the development of novel antimicrobials and environmental-ly relevant biotechnological applications.
Keywords:
biosynthetic gene clusters
; genome mining
; natural products
; environmental bacteria
; antimicrobial discovery
; bioremediation
; artificial intelligence
; molecular docking
1. Introduction
The global crisis of antimicrobial resistance (AMR) is a complex phenomenon driven by a range of biological factors, human practices, and socioeconomic conditions that threaten modern medicine. It is estimated that AMR contributes to about 9% of all global deaths. Data from 2021 show that 4.71 million deaths were associated with AMR, with 1.14 million deaths directly related to resistance. The excessive and inappropriate use of antibiotics in both human healthcare and agriculture is among the main causes of this crisis, creating selective pressure that favors the survival of resistant bacteria. It is estimated that about 73% of all antimicrobials used globally are destined for the livestock sector, creating reservoirs of resistance genes that can be transmitted to humans. Socioeconomic and infrastructure factors—such as lack of access to safe drinking water and basic sanitation, and low vaccination rates—contribute to the burden of AMR, as they facilitate the spread of infections, which in turn require the use of antibiotics. In low- and middle-income countries, the scarcity of basic diagnostic tools leads to widespread empirical use of antibiotics. Furthermore, human mobility facilitates the global spread of resistant strains [1,2].
During the so-called “Golden Age” of antibiotics, spanning the 1930s through the 1960s, there was a boom in the discovery of most classes of antibiotics that still sustain antimicrobial therapy currently, including aminoglycosides, glycopeptides, macrolides, and tetracyclines, among others. Since the 1980s, there has been stagnation in the discovery or development of new classes of antimicrobials, leading to excessive reliance on existing antibiotics. This situation stems from various economic and financial factors. According to Brussow (2024), the cost of launching a new antibiotic is estimated at US$ 1.4 billion per registered drug, with a low return on investment. For Ho et al. (2025), the key factors behind this stagnation are the absence of new structural classes; “false” innovation, which consists solely of structural modifications to existing classes, such as carbapenems, aminoglycosides, and macrolides; the problem of cross-resistance, since bacteria often already possess resistance genes that are effective against structurally modified antibiotics. Stagnation is particularly critical in the fight against Gram-negative bacteria, due to their permeability barrier, which makes the discovery of drugs that can penetrate and remain inside their cells a huge technological challenge. Meanwhile, pathogens such as Acinetobacter baumannii and Klebsiella pneumoniae are developing resistance to next-generation antibiotics, such as carbapenems and colistin [1,2,3].
The discovery of antibiotics during the “Golden Age” relied primarily on the cultivation of microorganisms, followed by the extraction of their secondary metabolites and biological screening to identify bioactive molecules. The most successful methodology of that era was the Waksman platform, named after its pioneer, Selman Waksman. The process involved collecting soil samples to isolate actinomycetes, which were then cultured on Petri dishes inoculated with a test pathogen. If the microorganism produced a natural antimicrobial compound, a “zone of growth inhibition” was observed around it. This method led to the discovery of several classes of antibiotics, such as beta-lactams, aminoglycosides, chloramphenicol, macrolides, tetracyclines, and glycopeptides. Starting in the 1960s, screening based on the Waksman platform began to yield increasingly lower returns. This was largely due to the frequent rediscovery of already known compounds, leading to the realization that this source of microorganisms had been extensively exploited. As a result, the pharmaceutical industry redirected its efforts toward the chemical modification of existing natural products, with the aim of increasing their activity, improving their pharmacological properties, or broadening their spectrum of action, giving rise, for example, to semisynthetic derivatives of penicillin, such as ampicillin. At the same time, investments in the development of synthetic molecules intensified, culminating in the introduction of important classes of antimicrobials, such as metronidazole and quinolones, as well as antituberculosis drugs, such as isoniazid [3,4].
Advances in genomic sequencing have revolutionized the search for new antimicrobials by revealing a much greater biosynthetic potential than had previously been inferred from approaches based exclusively on the cultivation of microorganisms. The growing availability of complete genomes has allowed the search for new compounds to move beyond reliance on traditional phenotypic screening and incorporate genomic mining strategies capable of identifying genes and biosynthetic pathways with the potential to produce secondary metabolites. This process has evolved in two major stages: an initial phase, focused on identifying genes and molecular targets from the first sequenced genomes, and a more recent phase, driven by the integration of bioinformatics, metagenomics, and artificial intelligence (AI) tools, which has significantly expanded the capacity to explore the chemical diversity encoded in microbial genomes [4].
The start of bacterial genome sequencing in 1995 marked a new phase in the discovery of antimicrobials, known as the Genomic Era or the era of target-based drug discovery, which lasted approximately until 2004. The expectation was that comparing the genomes of pathogens with those of their hosts would allow the identification of genes essential for bacterial survival that are absent in human cells, revealing new therapeutic targets and previously unknown mechanisms of action. Driven by this perspective, several pharmaceutical companies and research centers have invested in high-throughput screening programs aimed at identifying molecules capable of inhibiting these essential proteins. Although this approach has significantly expanded our understanding of bacterial biology and revealed hundreds of potential targets, it has not led to the introduction of new classes of antibiotics developed exclusively through this strategy. In many cases, the identified compounds exhibited low in vivo efficacy, pharmacokinetic limitations, or difficulty penetrating the complex structure of the bacterial cell envelope, especially in Gram-negative bacteria, and were often eliminated by efflux systems [3,4].
The rapid decline in sequencing costs and the development of bioinformatics tools marked, beginning in the mid-2000s, the transition to the Post-Genomic Era. During this period, interest shifted from focusing solely on the identification of essential genes to exploring the vast biosynthetic potential encoded in microbial genomes. At the same time, metagenomics significantly broadened this horizon by enabling direct access to the genetic material of microbial communities, including microorganisms not yet cultured in the laboratory, revealing an enormous diversity of previously unknown biosynthetic gene clusters (BGCs). More recently, the integration of genomic mining with metagenomic, transcriptomic, and artificial intelligence (AI) approaches has accelerated the identification and prioritization of promising BGCs, establishing these strategies as one of the main pillars of the discovery of new bioactive metabolites [3,4,5].
Advances in genomic approaches have broadened environmental bioprospecting by revealing genetic determinants and metabolic traits associated with contaminant transformation, removal, or immobilization. In environmental bacteria, this potential may arise from catabolic pathways that act directly on xenobiotics or from specialized metabolites that modify contaminant bioavailability, mobility, or toxicity [6,7,8,9]. This distinction is essential: genes involved in degradation, detoxification, or resistance do not necessarily constitute BGCs, which are defined as sets of genes involved in the biosynthesis of specialized metabolites. Nevertheless, catabolic pathways and biosynthetic systems can contribute in complementary ways to the same environmental phenotype, and their interpretation is strengthened when genomic data are integrated with phenotypic screening and other omics approaches [6,9,10]. The herbicide 2,4-dichlorophenoxyacetic acid (2,4-D) illustrates this progression from phenotype to genomic resolution. Bacteria recovered from the chicory rhizosphere were initially selected for their ability to biodegrade 2,4-D; subsequent genome sequencing of Brazilian strains displaying this phenotype provided resources for investigating the genetic determinants potentially associated with degradation [11,12,13].
This review aims to provide a comprehensive overview of advances in genomic mining of BGCs as a strategy for the discovery of new bioactive metabolites, with an emphasis on compounds of antimicrobial and environmental interest. The review discusses the main microbial groups that produce BGCs, the bioinformatics tools used to identify, annotate, and characterize these gene clusters, as well as the computational approaches employed to prioritize candidates, including protein structure prediction and molecular docking. Furthermore, the review highlights the potential of integrating genomics, artificial intelligence, and molecular modeling to accelerate the rational screening of molecules with applications in health and bioremediation, offering a current perspective on the challenges and opportunities in this rapidly expanding field.
2. BGCs: Definition and Classification
A BGC is defined as a group of two or more genes that are physically close (collocated) in the genome, and that are co-regulated to synthesize a specialized metabolite. This physical proximity is a key characteristic, typically defined by “signature” genes. The genes within a cluster belong to specific functional categories, such as core genes, which encode the main mechanism for synthesizing the chemical structure of molecules, such as non-ribosomal peptide synthases (NRPSs) or polyketide synthases (PKSs). Tailoring enzymes, also referred to as modifying enzymes, introduce chemical modifications that contribute to the structural and functional diversification of the biosynthesized product, and include oxidoreductases, transferases, hydrolases, and ligases. Transport and self-resistance genes encode proteins responsible for exporting the metabolite out of the cell or for conferring immunity to the producing organism itself against the toxicity of the synthesized molecule. Finally, the clusters may also contain regulatory genes, including transcriptional regulators that modulate pathway expression [10,14,15,16].
Proteins encoded by genes within a BGC can contain conserved Pfam domains; for example, modular enzymes such as NRPSs and PKSs have complex domain architectures that operate sequentially to assemble the molecule. NRPSs typically include condensation (C), adenylation (A), and peptidyl carrier protein (PCP) domains, whereas type I PKSs typically contain ketosynthase (KS), acyltransferase (AT), and acyl carrier protein (ACP) domains, sometimes accompanied by reductive domains. Some BGCs are considered hybrids when they combine genes from different biosynthetic classes, such as PKS–NRPS systems. Although physical co-localization is the general rule, the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) 4.0 standard recognizes cases where satellite genes or distant subclusters participate in biosynthesis, although the core genes are still required to be clustered together for the BGC to be defined. Tools such as antiSMASH extend the search for accessory genes in 5-, 10-, or 20-kb flanks beyond the detected core genes [10,14,15,16]. The general organization of these functional categories, including flanking genomic regions and the possible contribution of satellite loci, is summarized schematically in Figure 1.
Historically, BGCs have attracted interest primarily for their ability to encode antibiotics and other bioactive compounds. However, it is now recognized that these gene clusters are involved in the biosynthesis of a much broader range of specialized metabolites with different ecological functions, including siderophores, lipopeptides, pigments, signaling molecules, and other compounds that may contribute to microbial adaptation and interactions with the environment. At the same time, environmental bacteria may also harbor genes and catabolic pathways with significant biotechnological potential, which must be distinguished from BGCs. A well-known example is Pseudideonella sakaiensis (formerly Ideonella sakaiensis), which produces enzymes involved in the depolymerization and subsequent assimilation of polyethylene terephthalate (PET), generating terephthalic acid and ethylene glycol. Although this system does not, in and of itself, constitute a BGC, it illustrates the broad genomic potential of environmental bacteria for biotechnological and bioremediation applications. Thus, genome-guided bioprospecting can integrate the investigation of BGCs related to the production of specialized metabolites with other genetic determinants of environmental function, while maintaining a clear distinction between biosynthetic and catabolic pathways [14,17,18,19].
The natural products synthesized by BGCs play crucial roles in communication and competition among different species of soil microorganisms, contributing to ecological balance and the maintenance of soil quality, as well as serving as a basis for the development of new agrochemicals. As a result, interest in BGC mining has expanded beyond the discovery of new antimicrobials to include the identification of molecules with potential for industrial, agricultural, and environmental applications, including bioremediation processes. BGCs also encode systems for capturing scarce resources in the environment, such as siderophores, which chelate ferric iron in the environment and transport it into microbial cells [14,17,18,19].
BGCs can be classified according to the chemical nature of the natural products they encode, and the biosynthetic machinery involved in their production (Table 1). Among the main categories of biosynthetic products represented in databases such as MIBiG are PKSs, NRPSs, ribosomally synthesized and post-translationally modified peptides (RiPPs), terpenes, saccharides, and alkaloids. Polyketides are synthesized by PKSs, which assemble polyketide chains from precursors derived from acyl-CoA units; these pathways are associated with the biosynthesis of various antibiotics and other bioactive metabolites. NRPS are produced by large modular enzyme complexes that incorporate amino acids independently of the ribosome, enabling the synthesis of structurally diverse molecules, including antibiotics, immunosuppressants, and siderophores. RiPPs comprise peptides initially synthesized by the ribosome and subsequently modified by specific enzymes, resulting in a wide diversity of natural products. Terpenes, in turn, are derived from isoprenoid precursors and exhibit broad structural and functional diversity. Saccharides and alkaloids also constitute important categories of natural products, associated, respectively, with the biosynthesis of specialized sugars and structurally diverse nitrogen-containing compounds. These major BGC classes rely on distinct biosynthetic strategies, ranging from modular assembly lines in NRPSs and type I PKSs to ribosomal precursor modification in RiPPs and isoprenoid-based terpene biosynthesis [10,14,15,16]. Representative examples of these biosynthetic strategies are summarized in Figure 2.
In addition to these main categories, different types of BGCs can be defined based on specific characteristics of the biosynthetic machinery or the structure of the products formed. Among them are BGCs involved in the production of bacteriocins, lanthipeptides, and thiopeptides; NRPS-independent siderophore BGCs, in which biosynthesis occurs via discrete enzymes that produce molecules capable of chelating iron; and hybrid PKS–NRPS BGCs, which combine the polyketide and non-ribosomal peptide synthesis machinery. Other categories include aryl polyenes, associated with the production of bacterial pigments; enediynes, which encode highly reactive compounds with potent cytotoxic activity and antitumor potential; and BGCs involved in the biosynthesis of organ arsenic metabolites, some of which have been identified through evolutionary mining (EvoMining) approaches [10,14,15,16].
Secondary metabolites have a wide range of industrial applications, particularly in the medical, agricultural, and biotechnology sectors. Medicine and pharmacology are the most significant areas of application, with approximately 70% of anti-infective drugs derived from natural products. Specific functions include antibiotics and antifungals, such as penicillin, streptomycin, and daptomycin; antitumor and anticancer agents, such as taxol, marizomib, and carfilzomib, used in the treatment of various types of cancer, including multiple myeloma and lung and breast carcinomas; immunosuppressants, used to prevent organ rejection in transplants, such as cyclosporine; and, finally, metabolic control agents, which include cholesterol-lowering drugs such as lovastatin. Secondary metabolites and their derivatives account for about 36% of all new active ingredients in pesticides registered between 1997 and 2010, such as insecticides like avermectin and spinosyn, used to paralyze insects; herbicides such as glufosinate and aspterric acid, examples of compounds that eliminate unwanted plants by inhibiting vital enzymes; fungicides, such as fenpicoxamid, which inhibits cellular respiration in pathogens; and chemical hybridization agents, such as the aforementioned aspterric acid, which can be used to specifically inhibit pollen formation, facilitating the production of hybrid seeds. In the biotechnology industry, microorganisms such as the yeast Saccharomyces cerevisiae are used as “factories” to produce complex compounds on a large scale, such as opioid precursors and the antimalarial drug artemisinin. Bacteriocins, on the other hand, are used in the food industry as probiotics and preservatives to inhibit gastrointestinal pathogens. Some compounds have also been used for cosmetic and recreational purposes, in addition to being explored for the development of new types of biofuels and industrial chemicals [17,22,23].
3. Environmental Bacteria as Reservoirs of Metabolic Diversity
Environmental bacteria are considered invaluable sources of BGCs, and the search for new bioactive products has once again turned to the exploration of the genetic diversity present in these microorganisms in their natural state. Although soil is the most common source, cultured microorganisms represent less than 0.1% of all soil microorganisms, suggesting that the natural environment still holds a vast “gold mine” of new bioactive compounds to be explored through genomic mining [19,24,25]. Although they are very common in soil bacteria, the BGCs have also been identified in the human microbiome, where they play a role in protecting against invading pathogens [14,15].
3.1. Streptomyces Genus
The genus Streptomyces, belonging to the class Actinomycetes, is widely regarded as a “gold mine” for the discovery of antibiotics and other bioactive compounds due to its prolific production of secondary metabolites and its vast, still-unexplored biosynthetic potential. Since the discovery of streptomycin in the 1940s, the genus has been the focus of intensive microbial sampling efforts. It is estimated that most of all antibiotics used in the 21st century was developed from secondary metabolites produced by Streptomyces [24,25]. In addition to streptomycin, a wide range of other antibiotics, antitumor, antivirals and antifungals compounds are produced by the genus Streptomyces. Some of them are listed in Table 2.
Streptomyces genomes contain between 25 and 70 BGCs, with recorded averages of approximately 40 BGCs per genome, a number significantly higher than that of any other genus of actinobacteria. These clusters are responsible for the production of various classes of compounds, such as NRPS, PKS, terpenes, and lanthipeptides. Even very closely related strains or strains of the same species can exhibit an extremely variable BGC composition, suggesting that sequencing at the strain level may reveal an even greater diversity of useful compounds and derivatives. Genomic mining has revealed that the biosynthetic potential of these bacteria has been underestimated for decades. This is because only a small fraction of BGCs is expressed under laboratory conditions; many remain “silent” or inactive, awaiting specific environmental stimuli to produce new bioactive compounds [24,25].
3.2. Salinispora Genus
The genus Salinispora is recognized as the first formally described genus of strictly marine actinomycetes. Belonging to the family Micromonosporeaceae, these microorganisms are Gram-positive bacteria that grow slowly and form branched mycelia that produce black spores. Originally discovered in tropical and subtropical marine sediments, members of the genus Salinispora are now widely used as models in the discovery of natural products because approximately 10% of their genomes are dedicated to the biosynthesis of secondary metabolites. Currently, the genus comprises nine described and valid species. Among them, we can highlight the species Salinispora tropica and Salinispora arenicola. S. tropica is considered the source of salinosporamide A (marizomib), a potent proteasome inhibitor that has advanced clinical trials for cancer treatment. Other associated metabolites include sporolides, salinyllactam, and the sioxanthin pigment. The species S. arenicola, on the other hand, consistently produces rifamycins and staurosporins. It is also associated with the production of arenimycin, arenicolides (macrolides), arenamides, saliniketals, and the recently characterized retimycin A. The genus is considered a source of various compounds with chemical structures that were unknown until their discovery. In addition, metabolites such as lymphostins (immunosuppressants) and siderophores such as desferrioxamines (iron chelators) are produced by multiple species within the genus [18,26,27,28]. The other natural products associated with the genus Salinispora are listed in Table 2.
3.3. Bacillus and Related Genus
Bacillus consists of Gram-positive and rod-shaped bacteria that are widely distributed across various natural habitats such as soil, water, and air [29]. The genus is a member of the phylum Firmicutes with 125 valid named species [30]. Despite the well-established capacity of Bacillus species to produce a broad range of biologically active natural products, the biosynthetic potential of this genus remains underexplored. The substantial interspecies variability in BGCs diversity, distribution, and novelty represents a major challenge for systematic genome mining and the selection of Bacillus strains for the discovery of previously uncharacterized natural products [31].
According to the data from the Natural Products Atlas 2.0 [31], more than 450 secondary metabolites have been reported from Bacillus species, encompassing diverse chemical classes such as lipopeptides and surfactins (which have been widely used as antibiotics or natural bio-preservatives), polyketides, bacteriocins, lantibiotics, siderophores, and macrolactones [32,33,34,35,36]. This remarkable chemical diversity is accompanied by a broad spectrum of biological activities, including antimicrobial, antiviral, immunosuppressive, and antitumor effects, highlighting the genus as a valuable source of bioactive natural products with significant bioprospecting potential [35,37,38]. The main bioactive and antimicrobial compounds produced by Bacillus species are listed in Table 2.
Xia et al. (2022) demonstrated that the distribution of BGCs across Bacillus genomes exhibits a considerable degree of species specificity. Through a large-scale genome mining analysis of 4,268 Bacillus genomes, the authors identified 49,671 BGCs, corresponding to an average of 11.6 BGCs per genome and highlighting the extensive, yet still largely unexplored, secondary metabolic potential of the genus. Notably, substantial differences in BGC abundance were observed among strains belonging to the same species, including the BGC-rich species Bacillus subtilis. These findings indicate that both inter- and intraspecific variation in biosynthetic capacity should be considered when selecting strains for genome mining, and that comparative analyses across Bacillus species and strains may facilitate the prioritization of promising candidates for the discovery of novel natural products [39].
The biosynthetic potential of spore-forming bacteria extends well beyond Bacillus, with related genera such as Brevibacillus, Paenibacillus and Priestia emerging as important reservoirs of unexplored BGCs and secondary [40,41,42,43,44,45]. Kim et al. (2024) identified 848 BGCs in 554 Paenibacillus genomes, of which 84.4% were classified as unknown, with the majority associated with NRPS, PKS, and bacteriocin biosynthesis; further investigation of Paenibacillus brasilensis led to the discovery of five novel BGCs and the characterization of bracidin, a new antibiotic active against Gram-positive bacteria and fungi [42]. Similarly, genome mining of Brevibacillus revealed a remarkable diversity of biosynthetic pathways, including numerous uncharacterized BGCs, while the integration of genome mining with MALDI-TOF MS enabled the discovery of the novel brevipentin family [40,41]. Other studies exploring the genomes of Priestia megaterium strains have identified BGCs for terpenes, siderophores, phosphonates, and ranthipeptides [43,45]. Collectively, these findings highlight the substantial and largely untapped biosynthetic diversity of Bacillus-related spore-forming bacteria and support their exploration as promising sources of novel secondary metabolites with biotechnological and pharmaceutical potential. The main bioactive and antimicrobial compounds produced by Bacillus species and related genera are listed in Table 2.
3.4. Non-Fermenting Gram-Negative Bacilli Genus
Non-fermenting Gram-negative bacilli comprises a heterogeneous group of microorganisms characterized by their inability to ferment carbohydrates and by their remarkable metabolic and ecological versatility [46]. Although several species have been recognized as opportunistic pathogens, members of these genera are also widely distributed in soil, water, plants, industrial environments and other ecological niches. This broad environmental distribution, together with their metabolic plasticity and genetic diversity, makes the group an interesting and comparatively underexplored resource for bioprospecting [46,47].
From a biotechnological perspective, non-fermenting Gram-negative bacilli are particularly relevant because their genomes frequently harbour diverse BGCs, including those encoding NRPSs, PKSs, siderophores, bacteriocins, terpenes and other specialized metabolites. Several compounds with antibacterial, antifungal, antiparasitic and other biological activities have already been identified from Gram-negative bacteria, demonstrating that these microorganisms possess a substantial biosynthetic repertoire that remains incompletely characterized [46].
3.4.1. Pseudomonas Genus
The genus Pseudomonas is recognized as one of the most ubiquitous and metabolically diverse groups of bacteria in nature [48,49]. With more than 300 described species [18], these microorganisms can colonize nearly all known ecological niches, including soil, water, plants, and animals. This ability to survive in such varied environmental conditions is driven by their remarkable metabolic versatility and genetic plasticity. The genus is considered a rich source of diverse natural products, many of which have been discovered or characterized using advanced genomic mining techniques. The most common types of biosynthetic systems identified include NRPS-type systems—the most prevalent and responsible for the production of siderophores, such as pyoverdine (ubiquitous in the genus), toxins, and surfactants; bacteriocins, identified as the second most common type of BGC, revealing a vast arsenal of antimicrobial peptides that remain poorly characterized; PKS-type systems, which appear only in a few specific species [48,49].
Several publicly available genomes of Pseudomonas species and subspecies, originating from a wide variety of niches, including soil, plants, aquatic environments, animals, and food, have been explored for their biosynthetic potential. The analyzed genomes ranged in size from 3 to 7 Mbp, while the number of BGCs varied from 1 to 23 per genome, highlighting the considerable diversity of the genus’s biosynthetic repertoire. Evolutionary analysis of genes associated with secondary metabolism suggests that their diversification parallels, to some extent, the evolution of Pseudomonas species, indicating the presence of specific patterns of conservation and diversification among lineages. Furthermore, the genus exhibits a broad capacity to produce short-chain cationic peptides and other metabolites with antimicrobial activity, including molecules that show no significant similarity to compounds previously characterized in databases. This broad genetic repertoire suggests that Pseudomonas constitutes an important reservoir of metabolic diversity that remains largely unexplored, including BGCs potentially associated with cryptic metabolites (those not expressed in laboratory conditions [48,49]. The main bioactive and antimicrobial compounds produced by Pseudomonas species are listed in Table 2.
3.4.2. Burkholderia Genus
The genus Burkholderia comprises metabolically versatile with remarkable ecological adaptability, encompassing species associated with soil, water, plants, animals, fungi and clinical environments [50,51]. Although the taxonomic boundaries of the genus have undergone substantial revision, resulting in the transfer of numerous former Burkholderia species to other genera within the Burkholderia sensu lato complex, Burkholderia sensu stricto remains a highly diverse and biologically important group [50]. Current taxonomic frameworks recognize several major species complexes within the genus, including the Burkholderia cepacia complex (Bcc), Burkholderia pseudomallei complex and Burkholderia glumae complex, encompassing organisms with markedly different ecological lifestyles and biological properties. Their ecological versatility is supported by large, multireplicon genomes, frequently comprising three replicons and reaching approximately 10 Mbp, which provide an extensive genetic repertoire for adaptation to different environments and interactions with eukaryotic hosts [52,53].
Beyond their ecological and clinical relevance, Burkholderia species are recognized as prolific producers of specialized metabolites with considerable potential for pharmaceutical, agricultural and biotechnological applications. Genome-mining studies have demonstrated that their genomes harbour an unusually rich repertoire of BGCs, particularly those NRPSs and PKSs, as well as hybrid NRPS–PKS systems. These biosynthetic systems are associated with the production of diverse classes of natural products, including antibiotics, antifungal compounds, siderophores, lipopeptides, bacteriocins, phenazines, polyynes, terpenes, ribosomally synthesized and RiPPs, and other structurally diverse metabolites [42,52,53,54].
Liu and Cheng (2014) identified the structural and functional diversity of natural compounds in Burkholderia genome, highlighting the genus as a source of bioactive natural products. Among those Burkholderia strains explored by the authors, B. thailandensis E264 showed 21 putative BGCs predicted with the antiSMASH program and many interesting natural products have been discovered, including betulinan/terferol analogues, capistruin, malleilactone/burkholderic acid, thailandamides and thailandamide lactone, and thailandepsins/burkholdacs [52].
Among the most prominent biosynthetic systems in Burkholderia are NRPS pathways responsible for the synthesis of siderophores and lipopeptides. Genome mining of 48 complete Burkholderia genomes identified 161 NRPS-containing clusters, with the predicted products including numerous siderophores and lipopeptides with potential applications in biocontrol and agriculture. This analysis led to the identification of previously uncharacterized biosynthetic pathways, including the cluster associated with cepaciachelin production in Burkholderia ambifaria and a novel malleobactin-like siderophore, phymabactin, in Burkholderia phymatum [54]. More recent comparative analyses have further demonstrated that siderophore-associated BGCs, including those responsible for ornibactin production, exhibit species- and lineage-associated distribution patterns, suggesting that secondary metabolism may contribute to ecological adaptation and niche differentiation within the genus [42].
The biosynthetic potential predicted from Burkholderia genomes frequently exceeds the repertoire of metabolites detected under conventional laboratory cultivation. Genome-mining approaches have revealed numerous BGCs with low or no similarity to characterized clusters, suggesting the presence of potentially novel chemical scaffolds. For example, genome-guided investigation of Burkholderia gladioli associated with the beetle Lagria villosa led to the discovery of the cryptic gladiofungin BGC and a previously uncharacterized antifungal PKS [55]. Similarly, integrated genome mining, metabolomics and isotope-labelling approaches have recently enabled the identification of lagriamide-like metabolites from Burkholderiales, demonstrating how genomic information can reveal specialized metabolites that may remain undetected by conventional activity-guided screening [56]. Webster et al. (2025) studied the genome of the biopesticidal strain B. ambifaria BCC0191 and identified 21 predicted BGCs representing 14 metabolite classes. These include pathways associated with the antimicrobial compounds of cepacin, pyrrolnitrin, phenazine and burkholdines, as well as the siderophore ornibactin, while several additional clusters remain uncharacterized. Taken together, these findings indicate that Burkholderia is not only an ecologically and metabolically versatile bacterial genus but also an important reservoir of biosynthetic diversity with significant potential for natural-product discovery [57].
3.4.3. Ralstonia Genus
The genus Ralstonia comprises metabolically versatile bacteria within the family Burkholderiaceae [18], with representatives recovered from a wide range of environments, including soil, freshwater, plants, industrial environments and clinical settings [58]. Members of the genus exhibit remarkable ecological and metabolic diversity, ranging from important plant pathogens, such as Ralstonia solanacearum, R. pseudosolanacearum and R. syzygii [58,59,60], to environmental and opportunistic human-associated species, including R. pickettii, R. mannitolilytica and R. insidiosa [61,62,63]. Comparative genomic analyses indicate that this ecological versatility is supported by substantial variation in genome content and metabolic pathways, including genes involved in carbon metabolism, transport, adaptation to environmental stress and biosynthesis of secondary metabolites [59,64,65].
Ralstonia species possess a notable capacity for the biosynthesis of specialized metabolites. Genome-mining studies have identified multiple BGCs in representatives of the genus, including clusters associated with siderophores, aryl polyenes, ribosomally synthesized and RiPPs, terpenes, β-lactones, redox cofactors, NRPSs, PKSs and hybrid NRPS–PKS systems. A comparative genomic analysis of Ralstonia strains identified several classes of secondary-metabolite BGCs that were conserved across different species, while other clusters displayed a more restricted distribution, suggesting lineage-specific diversification of the biosynthetic repertoire [59,66].
The biosynthetic potential of Ralstonia is particularly evident in the R. solanacearum species complex, in which several well-characterized specialized metabolites have been identified. Genome mining of R. solanacearum GMI1000 revealed a hybrid PKS–NRPS gene cluster associated with the biosynthesis of micacocidin, a complex yersiniabactin-like siderophore with antimycoplasmal activity. The corresponding biosynthetic locus was shown to be activated under iron-limiting conditions, demonstrating that genomic predictions can be directly linked to the production of biologically active natural products. This finding highlights the potential of Ralstonia genome mining not only to describe the presence of BGCs but also to identify previously unrecognized metabolites with antimicrobial properties [59]. Montecillo et al. (2018) studied the genome of R. solanacearum T523 strain, which can biosynthesis ralsolamycin and related ralstonins. Ralsolamycin is an antibiotic lipopeptide produced by a hybrid NRPS–PKS pathway and has been associated with antifungal activity, whereas ralstonins are unusual lipodepsipeptides exhibiting phytotoxic and chlamydospore-inducing activities. The biosynthesis of these compounds is regulated by quorum sensing, illustrating the close relationship between secondary metabolism, microbial communication and ecological interactions. Similarly, the ralfuranone/ralstonin biosynthetic system produces specialized metabolites whose expression can be stimulated by plant-derived sugars. Ralfuranones and ralstonins have been detected infected plant tissues, and disruption of their biosynthetic genes has been associated with reduced virulence, indicating that these metabolites can contribute directly to host–pathogen interactions [67].
Comparative genome analyses further suggest that the distribution of secondary-metabolite BGCs is not uniform throughout the genus. In a genomic comparison of Ralstonia species, plant-pathogenic strains contained, on average, more known secondary-metabolite gene clusters than non-pathogenic strains, with approximately 10.87 clusters per strain in the pathogenic group compared with 6.28 in the non-pathogenic group. Particularly interesting were homoserine-lactone- and furan-associated clusters detected in pathogenic strains, which exhibited less than 20% similarity to known BGCs. Such low similarity suggests that these loci may encode chemically novel metabolites and highlights the possibility that the biosynthetic diversity observed through genome mining substantially exceeds the currently characterized chemical repertoire of the genus [68]. So, the number of predicted BGCs in Ralstonia genomes does not necessarily correspond to the number of metabolites that have been experimentally detected. Many biosynthetic pathways may remain silent or poorly expressed under conventional laboratory conditions, while others may encode compounds that have not yet been associated with known chemical structures. The identification of BGCs with low similarity to characterized clusters, together with the occurrence of strain-specific and lineage-associated biosynthetic pathways, therefore suggests the existence of a substantial reservoir of cryptic metabolites within the genus. Genome mining can consequently provide an important starting point for prioritizing Ralstonia strains for antimicrobial screening and for guiding the discovery of previously uncharacterized natural products [68].
3.4.4. Stenotrophomonas Genus
The genus Stenotrophomonas comprises metabolically versatile bacteria widely distributed in environmental habitats, including soil, rhizosphere, freshwater, plants and other ecological niches [69]. Members of the genus exhibit remarkable physiological adaptability and can interact closely with plants and other microorganisms. Although Stenotrophomonas maltophilia is also recognized as an opportunistic human pathogen [70], many strains have been investigated for their ecological and biotechnological properties, particularly their ability to colonize plant tissues, promote plant growth and antagonize phytopathogenic microorganisms [71]. This combination of ecological versatility, metabolic plasticity and competitive capacity makes Stenotrophomonas an interesting target for the bioprospecting of bioactive microbial products [72]. Members of the genus can produce proteases, chitinases, glucanases, lipases and other extracellular enzymes that contribute to antagonism against competing microorganisms and may participate in plant protection. In addition, Stenotrophomonas strains have been reported to produce volatile organic compounds and specialized metabolites with antibacterial and, particularly, antifungal activities. These characteristics have stimulated considerable interest as a source of biological control agents and natural products for agricultural applications [72,73].
One of the best characterized examples of secondary metabolism is maltophilin, a macrocyclic lactam antibiotic produced by S. maltophilia R3089, originally isolated from the rhizosphere of oilseed rape (Brassica napus). Maltophilin exhibits antifungal activity against a range of saprophytic, human-pathogenic and phytopathogenic fungi, although it showed no activity against the bacterial species evaluated in the original study [74]. Another important group of metabolites comprises xanthobaccins A–C, produced by Stenotrophomonas sp. SB-K88, a rhizobacterium associated with sugar beet. These macrocyclic lactam compounds exhibit antifungal activity against several phytopathogenic fungi and were subsequently associated with the suppression of damping-off disease [75].
Genomic analyses of Stenotrophomonas have identified BGCs associated with several classes of specialized metabolites, including NRPSs, metallophore-associated pathways, aryl polyenes, terpenes, RiPPs, lassopeptides and other hybrid biosynthetic systems. Analysis of the genome of S. maltophilia strain 3A identified five predicted secondary-metabolite BGCs, including arylpolyene, terpene-precursor, RiPP-like, lassopeptide and NRP-metallophore/NRPS clusters. Several of these clusters displayed relatively low similarity to previously characterized biosynthetic pathways, suggesting the potential to produce previously uncharacterized metabolites [71].
3.4.5. Acinetobacter Genus
Among non-fermenting Gram-negative bacilli, the genus Acinetobacter stands out for its predominantly coccobacillary cell morphology, as well as its widespread environmental distribution and recognized clinical relevance. Species of the genus Acinetobacter are metabolically diverse and produce a wide range of natural products with potential for industrial, agricultural, pharmaceutical, and ecological applications. Many of these compounds are associated with environmental adaptation, plant growth promotion, interspecific competition, and virulence. Siderophores of type NRPSs or NRPSs-independent are the most conserved and common specialized products in Acinetobacter species. They play a crucial role in iron uptake under iron-limited conditions and are essential for bacterial virulence and survival. Among these, acinetobactin and baumannoferrin A/B are particularly noteworthy. Endophytic species of Acinetobacter, such as Acinetobacter endophylla and Acinetobacter pittii, possess genetic clusters with the potential to produce lipopeptides with antifungal activity against pathogens. These molecules exhibit structural similarity to known biocontrol compounds, such as fengycin, plipastatin, iturin, and bacillomycin D. Acinetobacter junii strains synthesize biosurfactants that exhibit antimicrobial, antibiofilm, and antiproliferative activities against other microbial strains, opening the door to the development of medical and industrial formulations. The genome of Acinetobacter species contains several classes of RiPPs, such as thiopeptides, compounds similar to berninamycin A, and lasso peptides such as acinetodin, produced by the species Acinetobacter gyllenbergii. Several strains of Acinetobacter spp. produce generic bioemulsifiers and biosurfactants, which facilitate the solubilization of hydrocarbons and hydrophobic oils. Their primary biotechnological application lies in the bioremediation of oil spills and in facilitating enhanced recovery of heavy oils [76,77,78].
4. Integrated Framework for BGC-Driven Bioprospecting
The diversity of environmental bacteria described in the previous section defines a broad biosynthetic search space, but translating this potential into biotechnologically relevant candidates requires a sequence of analytical decisions. BGC-driven bioprospecting can therefore be organized as a progressive workflow in which genomic information is first used to identify candidate biosynthetic loci, followed by dereplication and prioritization, selection of informative genes or proteins, structural or functional characterization when appropriate, and experimental testing.
This framework is intended to distinguish the role of each analytical stage: Genome mining identifies and annotates candidate BGCs; comparative analyses assess relatedness to characterized clusters; prioritization selects candidates for deeper investigation; protein-level analyses generate mechanistic hypotheses; and experimental approaches determine whether the predicted biosynthetic potential is translated into detectable products or biological activity.
4.1. Genome Mining of BGCs
The discovery of natural products has undergone a transition from the “Golden Age” (experimental screening) to the “Platinum Age,” driven by genomics and computational tools. The main pillars underpinning the “Platinum Age” include genomic mining, which uses tools to identify BGCs directly from sequenced genomes, revealing that many microorganisms possess biosynthetic potential far greater than what traditional cultivation techniques were able to detect; access to “dark matter” through metagenomics, enabling the mining of BGCs from uncultivable microorganisms in the environment and tapping into a massive reservoir of previously unseen chemical diversity; expression and detection technologies, such as heterologous expression in domesticated hosts and cell -free systems, are used to activate “cryptic” or silent clusters that do not express themselves under normal laboratory conditions; and finally, AI using deep learning algorithms, which enable the discovery of entirely new classes of BGCs that defy known biosynthetic rules. AI models cross-referenced structural information with biological functions and discovered the antibiotic abaucin, which is specific to Acinetobacter baumannii. The goal of the “Platinum Age” is to accelerate the development of bioproducts, such as new drugs and agrochemicals, serving as a key driver for the growth of the global bioeconomy [15,79,80].
In recent decades, various genomic mining tools for BGCs have been developed and continuously refined, employing different computational strategies to identify, classify, and characterize genomic regions associated with the biosynthesis of specialized metabolites. These approaches explore distinct characteristics of BGCs, including “signature” biosynthetic genes, conserved protein domains, gene organization, similarity to previously characterized clusters, and, more recently, evolutionary patterns and machine learning models. The evolution of these tools has kept pace with the increasing availability of genomic data and the development of new analytical strategies, shifting from methods based predominantly on the identification of known biosynthetic genes and domains to approaches capable of recognizing more divergent BGCs and estimating their novelty. Currently, genomic mining allows for a more comprehensive exploration of the biosynthetic potential of microorganisms, including BGCs that remain silent or that show little or no similarity to previously described biosynthetic pathways [14,81,82].
Among the most widely used tools, antiSMASH has established itself as one of the leading platforms for the identification and annotation of BGCs, while subsequent approaches, such as DeepBGC, have incorporated various strategies for machine learning, pattern recognition, and the prediction of natural products. At the same time, tools such as BiG-SCAPE (Biosynthetic Gene Similarity Clustering and Prospecting Engine) allow for the comparison and grouping of BGCs based on their genetic and architectural similarity, contributing to the assessment of the diversity and novelty of the identified clusters [14,15,83]. In this study, we have categorized and grouped genomic mining tools based on the logic and architecture of their algorithms. A summary of some prediction tools of BGCs can be viewed in Table 3.
4.1.1. Rule-Based Tools
Rule-based tools utilize prior biosynthetic knowledge and traditional statistical models to identify known classes of clusters, primarily using profile hidden Markov models (pHMMs) and sequence similarity searches to find signature genes and conserved domains; for example, the BLAST algorithm [15,79].
The antiSMASH tool has established itself as one of the leading platforms for the identification and annotation of BGCs. The tool was launched in 2011 with the goal of accelerating the identification of promising targets in genomes, enabling laboratory research to keep pace with the rapid pace of genomic discoveries [14,16]. The antiSMASH algorithm is trained on multiple alignments of protein sequences or protein domains that are unique to certain types of biosynthetic clusters, in addition to using manually defined rules to determine which biosynthetic functions must coexist in a genomic region for it to be classified as a BGC [14,79,83,84]. The tool also uses the “Greedy” approach, once the signature genes have been identified, the algorithm groups nearby genes (generally within 10 kb) and extends the cluster boundaries (between 5 and 20 kb, depending on the type) to capture accessory, regulatory, and transport genes. For modular clusters, such as NRPSs and PKSs, antiSMASH performs detailed predictions of substrate specificity and stereochemistry and even generates a theoretical core chemical structure in SMILES format. In recent years, antiSMASH has undergone continuous updates. The latest version (v. 8.0) has expanded detection capabilities to 101 types of BGCs, reintroduced in-depth terpene analyses based on curated pHMMs, and created a dedicated tab for tailoring enzymes, organized by Enzyme Commission categories. There has been a change in the processing of fungi; users are advised to provide external annotations from tools such as AUGUSTUS to ensure the quality of the analyses [14,16,79,83,84,85].
antiSMASH results are primarily reported to the user via an interactive web page, which allows for detailed navigation through the identified gene clusters. In addition, the tool generates several downloadable files that facilitate further analysis and manual editing. In version 8.0, similarity to known clusters is reported at three levels: “high” (≥75%), “medium” (50–75%), and “low” (15–50%). Similarities between clusters below 15% are no longer considered sufficiently similar [14]. One of antiSMASH’s major limitations is that it relies on an algorithm based on predefined rules. This means it can only identify BGCs belonging to already known and characterized biosynthetic classes. According to Dinglansan et al. (2025), the algorithm cannot predict truly novel or “atypical” cluster types that fall outside the manual models hard coded into the system. AI tools were developed specifically to address this gap [79]. Furthermore, due to the “greedy” approach used to determine cluster size, independent biosynthetic loci located close to one another in the genome may be erroneously merged into a single large “region” or “supercluster.” Neighboring genes not involved in the natural product’s biosynthesis may be included when the algorithm extends the boundaries by a fixed distance (e.g., 10 kb to 20 kb). Another critical issue is that antiSMASH identifies genetic potential but cannot determine whether the cluster is active or “cryptic” (silent) under laboratory conditions [79,83].
The BAGEL (BActeriocin GEnome mining tooL) tool deserves special mention in this category. It is widely used for the identification and visualization of BGCs of bacteriocin and RiPPs in prokaryotic genomes. The BAGEL suite emerged in 2006 (BAGEL1) as a web server focused on bacteriocin mining based on sequence similarity. Its evolution has kept pace with advancements in genomics. Currently, the tool is in version BAGEL5, which focuses on high throughput and is capable of processing up to 100,000 input files in its stand-alone version. This version supports metagenomics, utilizes the DIAMOND aligner, and introduced “Strict” and “Discovery” search modes [86,87]. The input file is in DNA FASTA format. Subsequently, the DNA is translated into six reading frames. The algorithm then searches for protein motifs using pHMMs and similarity against a core peptide database. Upon identifying these elements, an area of interest identification (AOI) is defined, typically 30 kb in size. Within this area, the tool performs open reading frame (ORF) calling, including small ORFs in intergenic regions. Genes are functionally annotated, and promoters and terminators are automatically identified to aid in understanding gene regulation. Results are presented in an interactive graphical report comprising a table of all detected AOIs, their locations, and bacteriocin classes, alongside a genomic map and peptide alignments where database homology exists. Despite its efficiency, the tool presents challenges, such as a reliance on homology and a higher false-positive rate when using the version 5 “Discovery” mode; while this mode allows for the exploration of “biosynthetic dark matter,” it can result in clusters lacking clear precursor peptides. There is also a risk of omission, where extremely small peptides may be overlooked during initial stages if alignment criteria are overly stringent [86,87].
4.1.2. Machine Learning and AI-Based Tools (Pattern-Based)
This category aims to overcome the limitations of rule-based tools by learning complex patterns in BGC sequences to detect new or atypical cluster types. The algorithms use recurrent neural networks, conditional random fields, and deep neural networks [79,85]. DeepBGC is an example of a genomic mining tool based on deep learning and natural language processing (NLP), developed for the identification and classification of BGCs in bacterial genomes. Released in 2019 by Geoffrey D. Hannigan and colleagues, the tool emerged as an innovative alternative to traditional methods based on rules and linear probabilistic models. Initially, gene prediction is performed using Prodigal, and Pfam domains are identified. Next, each Pfam domain identifier is converted into a 100-dimensional dense numerical vector using a neural network called pfam2vec. These vectors capture the contextual and functional properties of proteins based on their ordered arrangement within bacterial genomes. The dense vectors feed into a Bidirectional Long Short-Term Memory (BiLSTM) recurrent neural network, which natively captures short- and long-range dependencies between distant or adjacent biosynthetic genes. The output layer calculates a probabilistic score (between 0 and 1) for each Pfam domain to be part of a BGC. These scores are aggregated for the corresponding genes to determine the exact physical boundaries of the cluster. After cluster detection, DeepBGC employs Random Forest classifiers to determine the chemical class of the biosynthesized compound and the potential biological activity of the secondary metabolite [15,82].
Despite its innovative nature, the DeepBGC tool has some limitations, including training bias. Because it was trained on public databases of validated BGCs, such as MIBiG, it exhibits a strong bias toward thoroughly studied microorganisms, such as those of the genus Streptomyces, which can reduce accuracy when analyzing underrepresented environmental or metagenomic taxa. Furthermore, the model was designed primarily for bacterial taxa. In fungal genomes, where BGC architectures are much more complex and varied, it exhibits low accuracy and performance compared to rule-based tools. In 2022, researchers developed e-DeepBGC (an extension of DeepBGC), improving upon the original model by incorporating additional biological information from the Pfam domains and introducing data augmentation techniques, thereby reducing overfitting and significantly improving sensitivity and class prediction for rare BGCs (Hannigan et al., 2019; M. Liu et al., 2022).
Another tool worth highlighting in this category is GECCO (Gene Cluster prediction with COnditional random fields), a machine-learning-based software designed to identify novel BGCs in bacterial genomes. GECCO differs from traditional approaches based on rigid rules and from conventional deep learning neural networks. The algorithm uses a conditional random fields-based probabilistic model to predict BGCs, evaluating the context and order of specific biological features enriched in bacterial BGCs, primarily using Gene Ontology (GO) terms. Unlike tools such as DeepBGC, which must be trained on massive datasets to “learn” complex patterns from scratch, GECCO requires considerably less training data, as it utilizes and inherits predefined, structured biological features for the model. The generated prediction results in standardized formats can be directly integrated and visualized within the antiSMASH interface. GECCO was specifically designed and optimized to analyze bacterial genomic data; it addresses common practical challenges of machine learning models, such as limitations in the analysis of eukaryotes (fungi), dependence on Pfam annotations, and the lack of detailed chemical structure predictions [14,79].
4.1.3. Evolutionary Mining Tools
Unlike the previous tools, these analyze the evolutionary principles that led to the formation of clusters, such as the duplication and recruitment of genes involved in primary metabolism. The algorithms are based on comparative phylogenetics and genomic synteny anomalies [79,88]. In this review, we will discuss the EvoMining and ARTS tools.
EvoMining is an evolution-based genomic mining tool designed to discover atypical or novel BGCs by analyzing the evolutionary trajectories of enzyme families. The tool was first developed around 2016 (version 1.0) and launched as a searchable website focused on Actinobacteria. In 2019, version 2.0 was released as a standalone tool and packaged in a Docker container to facilitate distribution and reproducibility. Unlike tools that search for BGCs by direct homology to known biosynthetic genes, the EvoMining algorithm tracks how genes from primary/central metabolism have undergone duplication or horizontal gene transfer and have been “recruited” for specialized metabolism. The user provides a seed enzyme database for central metabolism and a genome database. The tool performs a similarity search (BLASTP). An enzyme family is classified as significantly expanded if at least one genome exhibits a copy number above the lineage average plus two standard deviations. The sequences obtained from the expansions are subjected to another BLASTP search against a database of experimentally characterized BGCs, such as MIBiG. The most conserved sequences are classified as copies of the central metabolism through Bidirectional Best Hits (BBH) analysis. The remaining sequences, known recruitments, and orthologs of the central metabolism are aligned, curated, and constructed into maximum likelihood phylogenetic trees using FastTree [79,88].
The predictions of EvoMining are presented through three main visual outputs: the heat plot, the interactive-colored trees, and the synteny alignments. Reliance on curated databases is one of the algorithm’s main limitations, as it must map homology to known biosynthetic genes to label “blue branches” on the trees. In biologically under-characterized groups, such as archaea, the lack of references in MIBiG hinders the ability to define the metabolic fate of the predictions. Redundancy and false positives in primary metabolism have been reported, since not every expansion of a central gene is recruited for the synthesis of natural products. Some additional copies continue to function in primary metabolism due to metabolic redundancy, morphogenesis, or cofactor adaptation, leading to false identifications of BGCs. Furthermore, if the enzymatic sequence of a recruited copy diverges drastically from that of the central metabolism copy, the initial homology search may fail, preventing that enzyme from being grouped within the biosynthetic expansion phylogeny. The tool is also unable to delineate physical boundaries or predict compounds autonomously, relying instead on external integration with tools such as antiSMASH for these tasks [79,88].
The ARTS (Antibiotic Resistant Target Seeker) tool was developed to prioritize BGCs associated with the production of antibiotics with novel or promising modes of action, driven by the need to reinvigorate the search for new antibiotics in the face of the global crisis of multidrug-resistant pathogens. The ARTS algorithm is based on the evolutionary premise that antibiotic-producing microorganisms require self-protection mechanisms to prevent their own cell death. Often, this resistance occurs through the duplication of an essential gene involved in primary metabolism. One copy remains vulnerable to maintain cellular function, while the other copy, which undergoes resistance mutations or is acquired through horizontal gene transfer (HGT), is integrated into the antibiotic-producing gene cluster to confer self-protection. The ARTS screening pipeline first maps the genome to identify secondary metabolism using antiSMASH. ARTS then checks whether essential genes or known resistance patterns are located within or very close to the physical boundaries of these BGCs. The algorithm analyzes the frequency of essential genes and sets a duplication threshold based on the mean and standard deviation observed in a set of reference genomes. Genes that exceed this threshold are flagged as duplicates. ARTS performs a phylogenetic screening by comparing individual trees of essential genes with a tree of representative species constructed using coalescence analysis. Gene trees that are inconsistent with the species tree indicate probable HGT events. If an essential gene meets multiple criteria, such as duplication, HGT phylogeny, and being physically located within a BGC, the associated cluster is given high priority for laboratory validation [89,90].
Version 1.0 of the ARTS tool was released in 2017 with a limited scope, relying solely on complete actinobacteria genomes to map duplication events and resistance phylogeny. In 2020, version 2.0 significantly expanded its taxonomic coverage to include incomplete genomes and, crucially, metagenomes, enabling the search for resistance markers in complex environmental samples and uncultivable microorganisms. Version 2.0 now integrates the BiG-SCAPE algorithm directly into multigenomic workflows, enabling comparative analyses. The results are consolidated in an interactive web-based portal structured with tabs and searchable dynamic tables, containing interactive phylogenetic trees and comparative analyses with heatmaps showing the presence/absence of BGCs, as well as interactive networks based on BiG-SCAPE for comparing resistance clusters. The inability to distinguish biological functions constitutes one of its main limitations. ARTS cannot automatically differentiate whether a homologous gene is acting as a resistance mechanism or whether it is an essential component of the cluster’s own biosynthetic machinery. For genomes that do not belong to the reference phyla or when analyzing mixed metagenomes, the phylogeny and duplication criteria cannot be applied robustly due to the absence of coherent phylogenetic species sets [89,90].
4.1.4. Clustering and Networking Tools
Clustering and network tools play a fundamental role in the modern genomic era. As thousands of genomes and metagenomes are sequenced, the volume of predicted BGCs grows exponentially. To avoid the rediscovery of known compounds and to map chemical diversity on a global scale, it is necessary to group these BGCs into Gene Cluster Families (GCFs), which bring together clusters with similar architecture and functions that are likely to produce identical or related molecules [79,83].
BiG-SCAPE is the gold-standard tool for interactive and detailed analysis of similarity networks of BGC sequences. The BiG-SCAPE algorithm converts BGCs into sets of Pfam domains. To calculate the distance/similarity between each pair of clusters, it combines three main metrics: domain content similarity, synteny, and sequence identity. Based on these pairwise distances, it generates a similarity network and applies to the Affinity Propagation (AP) clustering algorithm to define the GCFs. The current version (2.0) of BiG-SCAPE allows the analysis to focus on individual protoclusters rather than entire genomic regions. This prevents irrelevant neighboring genes from influencing the clustering. The algorithm prioritizes the alignment of common domain sequences around the cluster’s actual biosynthetic core, ignoring transport cassettes or resistance genes that often introduce noise at the boundaries of the BGC. In addition to global and local alignments, a purely local alignment has been added, which is ideal for identifying small, highly conserved biosynthetic sequences within highly divergent genomic sequences. To prevent the AP algorithm from excessively splitting very similar clusters, BiG-SCAPE 2.0 evaluates the network’s connection density. If a cluster is too compact, it adjusts its internal parameters to avoid artificial splits [83]6).
Although BiG-SCAPE is extremely accurate, the “all-vs-all” distance calculation becomes computationally infeasible when processing millions of BGCs. BiG-SLiCE (Biosynthetic Genes Super-Linear Clustering Engine) was developed to address this scalability bottleneck. Instead of comparing BGCs pairwise, BiG-SLiCE converts each BGC into a dense numerical vector in Euclidean space, based on the count of biosynthetic Pfam domains. It then clusters these vectors extremely quickly in a superlinear manner using the BIRCH algorithm. It is capable of processing millions of clusters in just a few hours. The first version of BiG-SLiCE suffered from sensitivity imbalance. BGCs classes that were shorter or had fewer domains, such as RiPPs, generated very sparse vectors and were poorly clustered compared to NRPS or PKS. BiG-SLiCE 2.0 resolved this by replacing Euclidean distance with a cosine-like distance metric, which balanced the biological accuracy across all metabolite classes. In addition, code optimizations and the replacement of the classic HMMER with PyHMMER resulted in a further 25% to 50% reduction in total clustering time [79,83].
4.1.5. BGCs Repositories
MIBiG and BiG-FAM (Biosynthetic Gene Cluster Families Database) are two of the most important pillars in the database ecosystem for genomic mining of natural products. They operate on complementary fronts. While the former focuses on the rigorous experimental characterization of individual clusters, the latter organizes the global diversity of these genes into structured families [10,79].
MIBiG is the global gold-standard repository dedicated exclusively to storing and cataloging BGCs that have been experimentally tested and validated in the laboratory. It defines the minimum guidelines for chemical and biosynthetic annotations in a standardized, computer-readable format. MIBiG serves as the foundation for training artificial intelligence and prediction tools and acts as a reference anchor in comparative analyses. Its latest version (4.0) was the result of a global collaborative effort that made over 8,300 edits to the database. MIBiG now houses 3,059 curated and validated BGCs. Updates have been occurring quarterly and on an ongoing basis (minor releases), allowing researchers to collaborate in real time. Submissions and reviews now go through a rigorous pipeline that includes the use of interactive Kanban boards (Trello), a dedicated submission portal, and screening by specialized volunteer reviewers, who categorize submissions into quality levels (high, medium, and questionable) [10,14,79].
BiG-FAM is a public database designed to organize the immense volume of BGCs stored in public genome banks and classify them into GCFs. According to its statistics available online (https://bigfam.bioinformatics.nl/stats), BiG-FAM has cataloged a total of 1,225,071 BGCs, organizing them into 29,955 GCFs. These data were extracted from 209,206 genomes from the “observable microbial universe” (including more than 188,000 reference genomes from RefSeq/GenBank and more than 20,000 assembled metagenomes from the ocean, the rumen, and the human intestine). To calculate similarities in a database containing over 1 million sequences, BiG-FAM uses the partitions calculated by BiG-SLiCE. The Job ID of the sequences processed on the antiSMASH web server can be entered directly into the search tab of BiG-FAM. The portal maps the user’s gene cluster against the mathematical models in the global database, instantly reporting whether the queried cluster is an identical homologue of a known compound or represents a completely novel metabolic pathway [79,91].
Both repositories operate in an integrated manner. BiG-FAM integrates validated entries from MIBiG (version 2.0) directly as reference metadata to group and “anchor” the GCFs. Thus, when a user performs a quick query in BiG-FAM and finds a match with a specific family, they receive dynamic, cross-referenced links that point directly to the MIBiG and antiSMASH records. This allows researchers to correlate orphan sequences from environmental genomes or metagenomes with validated laboratory structural chemical data to immediately infer the molecular skeleton generated by the cluster [91].
4.1.6. Dereplication, Candidate Prioritization, and Selection of Target Genes and Proteins
Once BGCs have been detected, the central problem becomes candidate selection. A biosynthetically rich genome can contain many candidate clusters, and experimental characterization of all of them is rarely feasible. This review therefore treats dereplication and prioritization as a distinct computational stage between BGC detection and structural or experimental characterization; their roles and limitations are discussed below. Dereplication assesses whether a predicted BGC is closely related to characterized clusters or established GCFs, thereby reducing effort devoted to likely rediscovery; prioritization then identifies candidates for deeper investigation.
Prioritization integrates multiple criteria rather than relying on a single score. In this review, these criteria are organized into four dimensions: (i) data quality and prediction confidence, including assembly completeness and the consistency of the predicted biosynthetic locus; (ii) novelty, based on similarity to characterized BGCs and gene cluster families; (iii) biological and biotechnological relevance, including the predicted metabolite class, ecological context, and the presence of genes encoding functions related to transport, regulation, tailoring, or self-resistance; and (iv) complementary evidence, such as conservation among strains, transcriptomic signals, metabolomic features, or prior phenotypic observations.
This multi-evidence approach reflects the broader shift from simple BGC cataloging to rational exploration of biosynthetic space. Large-scale comparative genomic frameworks and paired genomic–metabolomic resources illustrate how the integration of independent evidence layers can strengthen candidate selection [92,93].
A high-priority BGC should therefore be understood as a candidate selected for deeper investigation, not as evidence that a specific metabolite is produced or active. Prioritization scores, novelty estimates, predicted functions, and AI-based rankings should be interpreted as complementary evidence for candidate selection rather than as proof of BGC functionality, metabolite production, or biological activity. For workflows that proceed to structural analysis, cluster-level prioritization should be followed by a second selection step. A BGC encodes multiple gene products with different roles, including core biosynthetic enzymes, tailoring enzymes, transporters, regulators, and resistance proteins. Structural analysis therefore shifts the unit of investigation from the BGC to selected encoded proteins. The relevant question becomes which protein can best address a defined biological question.
Target selection should be hypothesis driven. A core enzyme may be selected to investigate catalytic mechanism or substrate specificity; a tailoring enzyme may help explain chemical diversification; a transporter can be investigated in relation to export; and a self-resistance protein may provide clues to the metabolite’s mode of action or molecular target. Sequence annotation, domain architecture, conserved residues, predicted localization, and similarity to experimentally characterized proteins should therefore precede three-dimensional modeling. This selection step provides the conceptual bridge between genome-level prioritization and protein-level structural characterization.
4.2. Structural Characterization of Proteins Encoded by Prioritized BGCs
In this context, computational structural biology provides an additional layer of analysis following BGC prioritization and the selection of a BGC-encoded protein whose structural characterization can address a specific functional or mechanistic question. Structure-prediction approaches have also been successfully applied to biosynthetic enzymes, supporting their use in the investigation of proteins involved in specialized-metabolite biosynthesis (Gordon et al., 202(G. Protein structure prediction can support functional annotation and the identification of putative ligand-binding sites, provided that model confidence and local structural uncertainty are critically evaluated [95].
This caution is particularly important because even high-confidence predictions may contain inaccuracies in domain orientation or local backbone and side-chain conformations, and predicted models do not inherently account for ligands, cofactors, or other environmental factors that may influence protein structure and function [96,97]. These analyses do not establish that the BGC is expressed or that its predicted metabolite is produced, since many BGCs remain silent under standard laboratory conditions and require experimental approaches to link genomic potential to metabolite production [22]. Accordingly, this review examines target selection, structure prediction and validation, molecular docking, and virtual screening, as well as the limitations of structure prediction and structure-based functional inference. Where conformational flexibility is relevant to the biological question, dynamic or ensemble-based approaches may provide complementary information beyond static structural models.
4.2.1. Selection and Functional Annotation of Target Proteins
Before structural modeling, the target protein should be annotated at sequence and domain levels. Conserved motifs, catalytic residues, domain boundaries, homologous proteins, subcellular localization, and the biological role predicted from the BGC context should be considered together. For large modular proteins, whole-protein modeling should be distinguished from domain-focused modeling, particularly when interdomain orientation is uncertain.
4.2.2. Protein Structure Prediction
Modern structure-prediction methods have substantially expanded access to three-dimensional models for proteins without experimentally solved structures. AlphaFold2 demonstrated high-accuracy protein structure prediction from sequence, whereas ColabFold made AlphaFold-based workflows more accessible by combining AlphaFold2 with accelerated multiple-sequence-alignment generation using MMseqs2 [98,99].
More recently, AlphaFold 3 extended structure prediction to biomolecular complexes involving proteins, nucleic acids, small molecules, ions, and modified residues, further expanding the range of structural questions that can be addressed computationally [100]. For BGC-encoded proteins from poorly characterized environmental organisms, low sequence identity to experimentally characterized proteins can reduce opportunities for direct homology-based comparison without, by itself, determining the reliability of an AI-based prediction. Sequence divergence, domain architecture, available evolutionary information, and local model confidence should therefore be considered together. This distinction is particularly important for large multidomain proteins, in which well-supported local folds may coexist with greater uncertainty in interdomain orientation or functionally relevant regions [95,97,98].
In the context of BGCs, structure prediction can support functional annotation, comparison with characterized enzyme families, identification of plausible catalytic or ligand-binding regions, and generation of hypotheses for downstream biochemical studies. The choice of modeling strategy should be guided by the biological question and by the available structural information rather than treated as a strict distinction between template-based and AI-based approaches. Comparative or template-based modeling remains particularly informative when experimentally determined homologs are available, because the target–template alignment, sequence coverage, missing regions, conformational state, and the presence of experimentally observed ligands, cofactors, or oligomeric assemblies can be examined explicitly [101].
Conversely, AI-based methods can provide structurally informative models even when close structural templates are unavailable. These approaches are not mutually exclusive, since AlphaFold2 can also incorporate template information, and experimentally characterized homologs remain valuable references for interpreting AI-predicted structures [95,98]. AlphaFold-based predictions, the predicted TM-score (pTM) provides information on confidence in the overall topology and relative domain arrangement, whereas, for multichain predictions, the interface predicted TM-score (ipTM) estimates confidence in the relative placement of chains and their interfaces [98,102].
Confidence estimates should therefore be interpreted at different structural levels. The predicted Local Distance Difference Test (pLDDT) provides a residue-level estimate of local model confidence, whereas Predicted Aligned Error (PAE) estimates uncertainty in the relative positioning of residue pairs and is particularly informative for assessing domain packing and interdomain orientation. Consequently, a multidomain protein may contain individually high-confidence domains while retaining substantial uncertainty in their relative arrangement. This distinction is particularly relevant to BGC-encoded proteins such as NRPSs and PKSs, which frequently comprise repeated catalytic domains connected through flexible regions. For such proteins, domain-focused models may therefore be more informative than an uncritical interpretation of the full-length prediction when the relative orientation of domains is poorly constrained [95,98].
For structure-based functional analysis, confidence should also be examined specifically in regions relevant to the biological hypothesis, including catalytic residues, putative ligand-binding pockets, domain interfaces, and substrate-access channels. A high-confidence global fold does not necessarily imply equivalent accuracy at these local functional sites. In addition, AlphaFold2 models generally lack the ligands, metal ions, and cofactors that may be required to represent a biologically relevant state of an enzyme. Homology-based approaches such as AlphaFill can complement predicted structures by transferring experimentally observed ligands and cofactors from structurally related proteins, although such information should be regarded as supporting structural context rather than direct evidence of the native bound state [96].
Thus, a structure-prediction workflow for a prioritized BGC-encoded protein should retain not only the predicted coordinates but also information on the modeling method and version, sequence and domain boundaries, template coverage when applicable, and local and interdomain confidence metrics. These elements provide the basis for the structural validation and functional interpretation steps discussed in the following section.
4.2.3. Structural Validation and Functional Interpretation
Structural validation should not be reduced to a single confidence score. Prediction confidence, stereochemical quality, structural accuracy, and biological correctness represent different levels of evidence and should not be treated as interchangeable. Confidence metrics estimate the expected reliability of a prediction; stereochemical validation evaluates the plausibility of local atomic geometry; structural accuracy refers to agreement with the true structure; and biological correctness concerns whether the predicted conformation, domain organization, oligomeric state, and molecular interactions are relevant in their biological context. None of these dimensions alone establishes that a model has been experimentally validated [95].
Stereochemical quality assessment remains complementary to prediction-specific confidence metrics. MolProbity-based analyses can identify steric clashes, unfavorable backbone conformations in the Ramachandran plot, side-chain rotamer outliers, and other local geometric inconsistencies [103]. Favorable stereochemistry indicates that a model is geometrically plausible, but it does not establish that the predicted fold, domain arrangement, oligomeric state, or molecular interface is structurally accurate or biologically correct.
Interface quality requires a separate assessment, particularly for oligomers and protein-protein complexes. Interchain PAE and ipTM should be interpreted together with predicted interface contacts, buried surface area, shape and electrostatic complementarity, conservation of interface residues, and convergence across independently generated models [102]. High confidence in the individual protein structures does not necessarily imply high confidence in the predicted orientation or biological relevance of a protein complex.
Structural similarity to experimentally characterized proteins provide a complementary line of evidence but should not be treated as proof of function. The strongest functional hypothesis is obtained when prediction confidence, stereochemical plausibility, structural similarity, domain annotation, conserved residues and catalytic motifs, BGC context, and, where available, biochemical or functional data converge. This evidence-integration approach is particularly important for uncharacterized proteins encoded in environmentally derived or poorly studied BGCs.
4.2.4. Conformational Dynamics and Ensemble-Based Analysis
Protein structure prediction provides a structural starting point for mechanistic investigation, but a single predicted model does not represent the full conformational repertoire accessible to a protein. Many biological processes depend on transitions between conformational states, collective domain motions, local rearrangements, or transient exposure of interaction sites. This distinction is particularly relevant to proteins encoded by BGCs. Modular non-ribosomal peptide synthetases (NRPSs), for example, rely on coordinated movements of carrier and catalytic domains during substrate activation, condensation, and product elongation, while modular polyketide synthases (PKSs) undergo substantial rearrangements as intermediates are transferred between catalytic centers. Structural studies of these systems therefore support a view of biosynthetic enzymes as dynamic molecular assemblies rather than rigid architectures [104,105]. More broadly, the growing availability of highly accurate predicted structures has shifted part of the structural-biology problem from identifying a plausible fold to understanding the conformational ensemble associated with function [106].
For large multidomain proteins or complexes, normal mode analysis (NMA) and elastic network models (ENMs) provide a computationally inexpensive first approximation of collective structural motions. In these approaches, the protein is represented as an elastic network around a reference conformation, allowing low-frequency modes associated with large-scale movements to be identified. The Gaussian Network Model (GNM) describes patterns and correlations of residue fluctuations, whereas the Anisotropic Network Model (ANM) additionally provides information on the directions of collective motions [107–109]. Such analyses can highlight hinge regions, coupled domain motions, opening and closing movements, or changes that may influence access to catalytic and interaction sites. Contemporary implementations such as ProDy extend these approaches to protein families, structural ensembles, and supramolecular systems [110]. Their low computational cost is attractive for large BGC-encoded enzymes for which extensive atomistic simulations may not be practical. Nevertheless, ENM-derived modes describe mechanically accessible motions around the starting structure under a simplified harmonic approximation; they should not be interpreted as complete molecular trajectories or as direct evidence that the starting model is experimentally correct.
The interpretation requires additional caution for proteins from newly described or poorly characterized organisms. In these cases, low sequence identity to experimentally characterized proteins may limit direct homology-based assessment and can make functional inference more uncertain, even when a predicted fold appears plausible. At the same time, sequence divergence may be precisely what makes such proteins attractive for bioprospecting. These two aspects should not be conflated: limited sequence identity increases the interest of unexplored sequence space, but it does not by itself establish structural novelty, nor does it justify treating a predicted structure as experimentally validated. Structural homologs may retain related folds and dynamic features despite limited sequence identity, and structure-based comparisons can therefore remain informative even when sequence-based annotation becomes weak [110]. Dynamic analyses are most defensible after the initial model has passed confidence, stereochemical, and structural-consistency assessments. In this setting, their role is to ask whether a structurally plausible model supports coherent motions relevant to the proposed mechanism, rather than to compensate for an unreliable starting structure.
When greater configurational sampling is required, coarse-grained and atomistic molecular dynamics (MD) offer complementary levels of resolution. Coarse-grained models reduce the number of explicit degrees of freedom and can therefore extend the accessible length and time scales, which is useful for examining domain rearrangements, large protein assemblies, membrane-associated systems, and other structurally demanding targets. Martini 3, for example, expanded the range of biomolecular systems accessible to coarse-grained simulation and includes applications involving protein–protein and protein–membrane interactions [111]. This reduction in resolution, however, necessarily sacrifices part of the chemical detail required to describe local hydrogen-bond networks, side-chain rearrangements, solvent organization, metal coordination, or other interactions that may be critical in catalytic sites. Atomistic MD is better suited to these questions and can be used to examine local relaxation, residue fluctuations, interactions within a pocket or interface, and conformational changes over the sampled trajectory. Its interpretation remains dependent on the force field, starting coordinates, simulation length, and extent of conformational sampling; a stable trajectory is therefore evidence of compatibility with the chosen simulation conditions, not proof that a predicted conformation is the unique or biologically predominant state [111,112].
Recent work has also begun to connect structure-prediction confidence with reduced descriptions of molecular dynamics. Jussupow and Kaila (2023) compared AlphaFold-derived confidence information with fluctuations observed in explicit MD simulations and proposed an AlphaFold-informed elastic network model in which predicted uncertainty contributes to the parameterization of the elastic network. Their results illustrate one way in which prediction-derived information can be used to guide effective dynamics and coarse-grained exploration. Importantly, the authors also note that ENMs constructed from a single reference structure can underestimate global dynamics when proteins occupy multiple prominent conformational states. This is particularly relevant to multidomain biosynthetic proteins, in which relative domain positions may differ substantially during the catalytic cycle. Approaches that combine prediction, reduced dynamic models, and more detailed simulations may therefore be useful as a hierarchical strategy, provided that the limitations inherited from each level of representation remain explicit [113].
Structure prediction can also be used to explore alternative conformations, although such applications require a different interpretation from conventional single-structure prediction. Modified AlphaFold2 workflows have recovered experimentally observed alternative states by changing the depth or composition of the multiple sequence alignment (MSA), including different functional states of transporters and receptors and alternative conformations in metamorphic proteins [114,115]. Subsampling strategies have also been investigated as a means of approximating conformational distributions, with agreement reported for specific systems evaluated against nuclear magnetic resonance data [116]. These approaches demonstrate that the AlphaFold framework contains information that can be exploited for conformational sampling, but the resulting collections of models should not automatically be interpreted as thermodynamic ensembles or as reliable estimates of state populations. Indeed, subsequent analyses have shown that the outcome of MSA-based sampling can be sensitive to the sequence-selection strategy, emphasizing the need for independent structural or dynamical evidence when alternative states are inferred computationally [117].
For BGC-oriented structural studies, these methods are therefore best used selectively and according to the biological question. NMA or ENM can provide an initial, low-cost assessment of collective flexibility; coarse-grained simulations can extend exploration to larger systems or broader rearrangements; and atomistic MD can be reserved for selected proteins, interfaces, or ligand-binding regions in which local interactions require greater chemical detail. If functionally relevant flexibility affects a binding pocket or interaction surface, representative conformations identified from these analyses can also be carried forward as a conformational ensemble for docking rather than restricting the analysis to a single rigid receptor. Conversely, post-docking simulations can address whether prioritized poses remain structurally coherent or reorganized under the selected simulation conditions. In neither case should conformational persistence be treated as experimental validation of binding or biological activity. Dynamic and ensemble-based analyses instead provide an additional mechanistic layer between structural prediction and experimental testing, helping to define which conformations and interactions merit further investigation.
4.2.5. Molecular Docking and Virtual Screening
Molecular docking can be used to explore plausible interactions between a structurally modeled or experimentally characterized BGC-encoded protein and candidate substrates, products, cofactors, inhibitors, or other ligands. AutoDock Vina is a widely used docking engine for molecular docking and virtual screening. Since its original implementation, the platform has been expanded to incorporate additional docking methods, scoring functions, and software capabilities [118,119].
Docking results must nevertheless be interpreted cautiously. Scoring functions approximate complex physical interactions and can differ in their ability to reproduce poses, rank ligands, or estimate affinity. Benchmark studies of AutoDock and AutoDock Vina illustrate why docking performance depends on the evaluation task and why a favorable score should not be equated directly with experimental binding or biological activity [120].
For BGC research, docking is most informative when it is linked to a specific biosynthetic hypothesis: the protein (receptor) preparation, ligand set, binding-site definition, cofactors, protonation states, and structural uncertainty should all be justified in relation to the biological question. Virtual screening can prioritize ligands or interaction hypotheses for experimental follow-up, but it cannot demonstrate catalytic competence, pathway activity, or cellular efficacy.
4.2.6. Limitations of Structural Inference in BGC Studies
Structural modeling provides a mechanistic layer of evidence but does not demonstrate that a BGC is expressed, that the modeled protein is produced under the conditions of interest, or that the corresponding pathway synthesizes the predicted metabolite. Likewise, a plausible docking pose does not establish binding affinity, catalytic turnover, antimicrobial activity, or environmental function. These limitations are particularly relevant for large multidomain NRPS/PKS proteins, proteins requiring oligomerization or partner interactions, membrane-associated systems, dynamic active sites, and pathways that depend on sequential enzymatic reactions. Accordingly, this review positions structural modeling and docking as hypothesis-generating and prioritization tools that should inform the design of biochemical, metabolomic, microbiological, or other experiments.
4.3. In Vitro Evaluation of Biological Activity
Genome mining reveals biosynthetic potential, whereas experimental progression requires distinct evidence for pathway expression, metabolite production, and biological activity. A predicted BGC with no expression data, an expressed pathway or BGC–metabolite association, and a chemically characterized product with demonstrated activity represent different evidence levels. Metabolomics and expression studies can connect genomic predictions to detectable products, but functional assays are required to establish the claimed phenotype.
4.3.1. Evaluation of Antimicrobial Activity
The in vitro evaluation of antimicrobial activity can be performed using a range of complementary methodologies, depending on the nature of the test substance, the microorganism under investigation, and the specific parameter to be determined. The most widely used approaches include disk-diffusion agar assays, broth dilution methods, gradient diffusion, time-kill assays, and methods for assessing antimicrobial interactions or activity against biofilms [121–125]. Each methodology has specific advantages and limitations, and the selection of an appropriate method is essential to ensure that the results are reproducible, quantitatively meaningful, and comparable between studies [125,126].
Among quantitative susceptibility testing methods, the broth dilution assay (micro or macrodilution) is one of the most widely used approaches for determining the minimum inhibitory concentration (MIC) of an antimicrobial agent. The MIC is defined as the lowest concentration of an antimicrobial that prevents visible microbial growth under the specified test conditions. In broth microdilution, serial concentrations of the test substance are prepared in a liquid growth medium, usually in the wells of a sterile 96-well microtiter plate, followed by inoculation with a standardized microbial suspension. After incubation, microbial growth can be assessed visually or by spectrophotometric or other appropriate methods. This approach allows the simultaneous testing of multiple concentrations and isolates while requiring relatively small volumes of antimicrobial compounds and culture medium. Broth microdilution is considered a reference methodology for antimicrobial susceptibility testing and is extensively standardized by the Clinical and Laboratory Standards Institute (CLSI) [127].
Agar dilution represents another quantitative approach in which defined concentrations of the antimicrobial agent are incorporated into solid agar before inoculation. The lowest concentration preventing visible growth is then determined as the MIC. Agar dilution can be particularly useful when testing large numbers of bacterial isolates under highly standardized conditions, although its preparation is more laborious than broth microdilution [127].
Agar diffusion methods are among the simplest and most commonly employed approaches for the initial screening of antimicrobial activity. In the disk diffusion assay, a standardized microbial suspension is evenly inoculated onto the surface of an appropriate agar medium, after which paper disks containing specific concentration of the test substance are placed on the agar surface. The paper disks impregnated with a defined volume or concentration of the test substance can be used. The test substance diffuses radially through the agar during incubation, producing a zone of growth inhibition around the disk. The diameter of this zone can subsequently be measured and used as an indicator of antimicrobial activity [124,127]. This approach is particularly useful for preliminary screening of antimicrobial-producing microorganisms, crude extracts and purified compounds. Nevertheless, the diameter of the inhibition zone is influenced not only by the intrinsic antimicrobial potency of the tested substance but also by its physicochemical properties, including molecular size, solubility, stability and diffusion coefficient. Consequently, inhibition-zone diameters obtained by disk diffusion should not be interpreted directly as equivalent to MIC values, particularly when comparing chemically unrelated compounds [125,126].
The agar well diffusion assay is another frequently used diffusion-based methodology. In this approach, wells are made in an agar plate previously inoculated with the test microorganism, and a defined volume of the test substance is introduced into each well. During incubation, the substance diffuses into the surrounding agar, resulting in a zone of inhibited growth. The method is particularly convenient for testing liquid samples. As with disk diffusion, however, the size of the inhibition zone depends on both antimicrobial activity and diffusion characteristics. Therefore, well diffusion is generally more appropriate as a screening method than as a stand-alone quantitative measure of antimicrobial potency [125,126].
The determination of the minimum bactericidal concentration (MBC) provides complementary information to the MIC by assessing whether an antimicrobial agent can kill the microorganism rather than merely inhibiting its growth [121]. Typically, samples from dilution wells showing no visible growth are subcultured onto antimicrobial-free agar. The MBC is defined as the lowest concentration associated with the absence of recoverable viable microorganisms according to the criteria adopted for the assay [128]. MBC determination can distinguish predominantly bacteriostatic activity from bactericidal activity. Although, the interpretation of bactericidal effects requires careful consideration of the experimental conditions and the definition of the endpoint. The minimum fungicidal concentration (MFC) follows the same idea, but culture media specific for fungi are used [129,130].
The time-kill assay provides a dynamic assessment of antimicrobial activity by examining changes in viable microbial populations over time following exposure to defined concentrations of test substance. Aliquots are collected at predetermined time points and viable microorganisms are quantified, commonly by determining colony-forming units (CFU/mL). The resulting time-kill curves can demonstrate whether a test substance produces rapid or delayed killing and whether its activity is concentration- or time-dependent. This approach can provide information that cannot be obtained from a single MIC measurement and is particularly valuable when characterizing the pharmacodynamic behavior of new antimicrobial compounds [125,126,130].
Another important consideration is the evaluation of antimicrobial activity against biofilms. Conventional MIC assays primarily assess planktonic cells and may not accurately reflect the susceptibility of microorganisms growing within a biofilm. Biofilm-associated cells can exhibit substantially increased tolerance to antimicrobial agents because of the extracellular matrix, altered metabolic states, heterogeneous microenvironments and physiological adaptations associated with surface-associated growth [70]. Biofilm susceptibility can be investigated using microtiter-plate models, Calgary Biofilm Device systems, flow-cell models and other experimental platforms. Parameters such as the minimum biofilm inhibitory concentration (MBIC) and minimum biofilm eradication concentration (MBEC) may be determined, although biofilm susceptibility testing remains less standardized than conventional planktonic susceptibility testing [131,132].
In addition to these conventional methods, several alternative approaches have been described for the rapid or mechanistic assessment of antimicrobial activity. These include flow cytometry, which can provide information on membrane integrity, cellular viability and physiological changes; bioluminescence-based assays, which can provide rapid measurements of microbial viability or metabolic activity; and thin-layer chromatography–bioautography, which can be particularly useful for identifying antimicrobial compounds within complex mixtures. Such techniques can provide information beyond simple growth inhibition, although they may require specialized instrumentation and, in some cases, further standardization before results can be directly compared with reference susceptibility methods [125,126].
For research focused on the discovery or preliminary characterization of novel antimicrobial compounds, a combined strategy is therefore recommended (Figure 3C). An initial agar diffusion assay, using either impregnated paper disks or agar wells, can provide a rapid qualitative screen of antimicrobial activity. Positive samples can subsequently be investigated using broth macro/microdilution to determine MIC values, followed by MBC or MFC and/or time-kill assays to establish whether the observed activity is inhibitory or microbicidal and to characterize its kinetics (Figure 3A). Where appropriate, additional experiments involving antimicrobial combinations, biofilms or mechanistic assays can then provide a more comprehensive characterization of biological activity.
4.3.2. Evaluation of Bioremediation Activity
In vitro validation is the stage at which the potential indicated by genome mining is tested against a measurable phenotype. In bioremediation, the presence of a gene, metabolic pathway, or BGC does not, by itself, demonstrate that the system is expressed or that it has a relevant effect on a contaminant; environmental function therefore requires experimental validation [6,133].
It is equally important to distinguish catabolic from biosynthetic potential. Monooxygenases, dioxygenases, hydrolases, reductases, efflux pumps, and transporters can participate directly in pollutant transformation without forming BGCs. By contrast, BGCs encode pathways for specialized metabolites such as siderophores, lipopeptides, NRPSs, and PKSs, which can increase substrate availability, chelate metals, or promote adaptation to chemical stress [7,9,10].
Genome-guided bioremediation should therefore be viewed as a chain of evidence: genetic potential must be linked to expression, protein or metabolite production, and ultimately to a measurable chemical transformation. The more consistently this relationship is maintained as the system progresses from controlled conditions to environmental matrices, the stronger the functional interpretation [6,133].
Linking Genetic Potential to a Reproducible In Vitro Phenotype
The first challenge is to obtain a reproducible phenotype from genomic potential that may remain silent under standard laboratory conditions. The One Strain-Many Compounds (OSMAC) strategy varies medium composition, nutrient availability, pH, salinity, temperature, aeration, or co-culture conditions to uncover alternative metabolic profiles. In Streptomyces ciscaucasicus GS2 strain, for example, such changes altered the siderophore profile detected by high-performance liquid chromatography–high-resolution mass spectrometry (HPLC-HRMS); however, relevance to bioremediation is established only when the induced product is linked to a measurable effect on the contaminant [134,135].
Standardization of the inoculum and culture medium is equally important. Cells harvested at a defined physiological stage, washed to minimize carryover of organic carbon from the preculture, and adjusted to a known initial biomass improve comparability among assays. When direct use of a xenobiotic is being tested, a defined mineral medium without an alternative organic carbon source better links growth, substrate consumption, and metabolite formation; Optical density (OD) 600 nm can be used for standardization, but it should be complemented by an additional measure when cell size, morphology, or pigmentation may compromise comparisons [136–138].
Physical culture conditions also affect the outcome. In shake flasks, liquid volume, headspace, shaking frequency, vessel geometry, and closure type influence oxygen transfer and, for volatile compounds, may alter abiotic losses. These variables need not be exhaustively investigated in every study, but they should be standardized and reported, particularly when degradation depends on aerobic enzymes or poorly stable compounds [138,139]. Using the contaminant as the sole carbon source is a strong experimental design, but it should not be treated as a universal requirement. Many xenobiotics are transformed cometabolically, in which case the cell depends on another substrate for growth. Comparing the contaminant alone with the contaminant plus a defined cosubstrate can distinguish direct metabolism from cometabolic transformation without misclassifying dependence on an additional carbon source as failure of the strain [140,141].
Validation should also follow the process over time. Measurements of biomass, parent-compound concentration, and transformation products by gas chromatography–mass spectrometry (GC-MS), High-Performance Liquid Chromatography (HPLC), or liquid chromatography-mass spectrometry (LC-MS) can reveal intermediates and distinguish primary degradation from mineralization. When complete biodegradation is claimed, carbon dioxide evolution, oxygen consumption, or carbon-balance measurements strengthen the conclusion, together with controls for adsorption, volatilization, photolysis, and other abiotic losses [136,138,142].
Functional Validation of Organic Xenobiotic Degradation
For organic xenobiotics, a convincing bioremediation claim must go beyond tolerance or disappearance of the parent compound. Ideally, microbial activity should be linked to substrate consumption, formation of transformation products, and, when relevant, reduced toxicity, because contaminant resistance is not equivalent to degradation [9,140]. 2,4-D provides a useful example of this integration. Rhizosphere isolates from chicory were phenotypically screened, and a Delftia strain was identified among bacteria capable of 2,4-D biodegradation; later, the genomes of Brucella intermedia DF13 (formerly Ochrobactrum intermedium) and Enterobacter hormaechei MG02, both Brazilian 2,4-D-degrading strains, provided genomic resources for investigating genetic determinants associated with this phenotype [11,12,13].
For pesticides, the experimental design should establish whether the compound supports growth or whether removal depends on a cosubstrate. Escherichia fergusonii and Clostridium bifermentans degraded chlorpyrifos in mineral medium with the insecticide as a carbon source, and glucose supplementation increased removal, illustrating how direct metabolism and cosubstrate-enhanced efficiency can coexist within the same system [141].
For hydrocarbons, low aqueous solubility and sorption make bioavailability part of the mechanism. Bacillus subtilis EB1 degraded phenanthrene under different pressure regimes, with transformation products detected by GC-MS and genomic determinants related to aromatic degradation and stress adaptation. In Kocuria flava IOS11, integration of genomic, transcriptomic, and metabolomic data further connected expressed genes with metabolites detected during phenanthrene degradation [143,144].
For plastics, evidence of colonization, biofilm formation, or mass loss alone is insufficient. A more robust assessment combines polymer-level changes, such as Fourier-Transform Infrared Spectroscopy (FT-IR) or gel permeation chromatography (GPC), with detection of monomers or oligomers and, ideally, evidence of mineralization or carbon assimilation. The same caution applies to dyes, pharmaceuticals, and other emerging contaminants, for which decolorization or a decrease in concentration does not necessarily demonstrate complete detoxification [9,140,142,145].
Specialized Metabolites and Indirect Mechanisms of Bioremediation
Not all bioremediation relies on direct transformation of the contaminant. Specialized metabolites can alter contaminant availability, mobility, or toxicity, making BGCs particularly relevant to indirect mechanisms. Biosurfactants, for example, can enhance hydrocarbon dispersion, but drop-collapse, oil-spreading, or emulsification assays are screening tests for surface activity and do not by themselves demonstrate improved remediation [7,8,146].
The evidence becomes stronger when surface activity is connected to chemical characterization and a measurable effect on the contaminant. Bacillus cereus NWUAB01 harbored BGCs and metal-resistance determinants and produced a biosurfactant associated with lead, cadmium, and chromium removal; Serratia sp. and Acinetobacter sp. isolates from oily sludge were evaluated using phenotypic screening, FT-IR, and genomic analysis; and Brucella pituitosa BU72 combined hydrocarbon growth, metal tolerance, and exopolysaccharide-based surfactant production [8,146,147].
Siderophores, metallophores, and exopolysaccharides extend this role to inorganic contaminants. Because metals are not biodegraded, outcomes should be described in terms of biosorption, bioaccumulation, mobilization, immobilization, or redox transformation. Mucilaginibacter pedocola TBZ30T, for example, combines metal resistance and exopolysaccharide production with zinc and cadmium removal, whereas bacteria from hydrothermal vents showed metal biosorption that depended on the element, pH, concentration, and contact time [7,148,149].
Tolerance and remediation should therefore be assessed separately. The minimum inhibitory concentration indicates resistance but does not quantify removal; when metals are involved, Inductively Coupled Plasma Optical Emission Spectrometry (ICP-OES) or Inductively Coupled Plasma-Mass Spectrometry (ICP-MS) can be used to monitor concentration, partitioning, and, when required, speciation across biomass, supernatant, and extracellular fractions. This distinction prevents survival phenotypes from being misinterpreted as remediation capacity [7,8,149].
Multi-Omics Integration and the Genotype-to-Phenotype Gap
Multi-omics integration helps narrow the gap between genetic potential and function. Genomics defines the available repertoire; transcriptomics identifies genes that respond to the contaminant; proteomics brings the analysis closer to the proteins produced; and metabolomics captures chemical intermediates and end products. This convergence is important because gene presence, expression, and metabolic flux are not equivalent [6,133].
In Kocuria flava IOS11, the combination of genomics, transcriptomics, and metabolomics linked genome annotation more closely to an experimentally supported phenanthrene-degradation pathway. Metabolic models can complement this process by suggesting bottlenecks, cofactor requirements, or nutrient dependencies, but their predictions gain value only when tested through substrate consumption, enzyme assays, expression analyses, or metabolomics [133,143].
When a specific gene or BGC is proposed as a determinant of a phenotype, correlative omics evidence should, where feasible, be strengthened by functional genetics. Gene disruption, complementation, or heterologous expression can test whether a locus is required or sufficient for degradation or metabolite production. Deletion analysis in Burkholderia xenovorans LB400 demonstrated distinct contributions of redundant benzoate-catabolic pathways, while BGC-focused workflows increasingly combine genome mining, metabolomics, and molecular biology to establish gene-to-metabolite links [150,151]. The same caution applies to cryptic or silent BGCs. Altering culture conditions may reveal products that are undetectable under standard conditions, but BGC activation alone does not establish environmental function: the metabolite must be characterized and linked to a measurable effect on the contaminant [134,135].
From Pure Cultures to Environmental Matrices
The final step is to determine whether the phenomenon observed in pure culture persists in an environmental matrix. Soil, sediment, natural water, and effluents introduce sorption, organic matter, variable pH, nutrients, oxygen availability, and microbial competition. Microcosms and mesocosms are useful because they introduce this complexity gradually while retaining some experimental control [6,152].
Pseudomonas qingdaonensis ZCR6 illustrates why this step is decisive. In vitro, the strain showed hydrocarbon degradation, metal resistance, siderophore production, and surface activity, together with a genomic repertoire consistent with these phenotypes. However, in co-contaminated soil planted with maize, bioaugmentation did not increase hydrocarbon removal relative to the control despite colonization of the plant and rhizosphere [152,153].
Consortia adds another layer of complexity but may better reflect environmental function because different species can partition steps of a degradation pathway, consume toxic intermediates, supply cofactors, or increase substrate bioavailability. In these systems, interpretation should preserve mechanistic traceability through population quantification, community-level analyses, and chemical measurements of the contaminant [6,133].
Taken together, these studies show that genome mining is better viewed as a prioritization tool than as proof of function. The strongest evidence emerges when a genomic signal is linked to a reproducible phenotype, a measurable chemical change, and, whenever possible, performance in a matrix that more closely approximates the real environment. The minimum experimental requirements for different types of genome-guided bioremediation claims are summarized in Table 4 [6,133,136,152].
5. Conclusions
The rapid expansion of genomic and computational technologies has transformed the discovery of microbial natural products. Environmental bacteria represent a particularly valuable reservoir of biosynthetic diversity, encompassing both well-characterized BGCs and others that remain largely unexplored. Genome mining has revealed that the biosynthetic potential of many microorganisms substantially exceeds the repertoire of metabolites detected through conventional cultivation and activity-based screening.
Integrating genome mining with comparative genomics, BGC similarity analysis, AI, protein structure prediction, molecular docking, metabolomics, and experimental screening provides an increasingly robust strategy for prioritizing biosynthetic systems with biotechnological potential. In this context, computational approaches can narrow the search space and generate hypotheses regarding the chemical and biological properties of candidate pathways, while structural analyses can yield additional insights into enzymatic function, substrate recognition, and potential molecular interactions. However, these approaches should be viewed as complementary and hypothesis-generating, rather than as definitive evidence of metabolite production or biological activity. In particular, predicted protein structures and docking interactions cannot independently demonstrate pathway expression, catalytic activity, binding affinity, antimicrobial efficacy, or environmental function.
The presence of a BGC does not necessarily indicate that the corresponding metabolite is expressed under laboratory conditions, and many clusters remain cryptic or silent. Strategies such as OSMAC, co-cultivation, heterologous expression, cell-free systems, metabolomics, and targeted molecular approaches can help establish the link between genetic potential and metabolite production. For antimicrobial discovery, complementary assays are required to confirm biological activity.
Overall, the future of microbial bioprospecting lies not in replacing experimental biology with computational prediction, but rather in integrating both approaches into a seamless, evidence-based workflow. Such integration can contribute not only to the discovery of new antimicrobial agents in response to the global antimicrobial resistance crisis but also to the development of agricultural, industrial, and environmental biotechnologies grounded in the metabolic diversity of environmental bacteria.
6. Future Directions
BGC-based bioprospecting will increasingly rely on the integration of complementary computational and experimental approaches. Although current genome-mining platforms have substantially expanded the number of identifiable candidate BGCs, the growing volume of genomic data also creates a bottleneck in prioritization. Future developments should focus not only on detecting more clusters but also on improving the ability to distinguish biosynthetic systems that are novel, expressed, chemically accessible, and biologically relevant from those representing redundant or inactive predictions. Combining BGC similarity networks, curated reference databases, comparative genomics, phylogenetic information, and AI-based predictions can provide increasingly powerful strategies to this end.
A key priority will be the continued development of AI and machine learning approaches capable of identifying biosynthetic systems beyond the limitations of manually defined rules. Conventional platforms, such as antiSMASH, remain highly valuable for detecting and annotating established classes of BGCs, but their rule-based architecture can limit the recognition of atypical or previously unknown biosynthetic systems. Machine learning approaches can complement these methods by learning sequence and domain patterns that are not necessarily represented by predefined biosynthetic rules. Future models will need to increasingly integrate information on sequences, genomic context, and chemical structure, as well as metabolomic, ecological, and functional data, rather than relying exclusively on sequence similarity.
Another important direction is the development of approaches capable of activating cryptic BGCs. The significant discrepancy between predicted biosynthetic potential and experimentally detected metabolites indicates that many pathways remain inaccessible under conventional laboratory conditions. Therefore, strategies such as OSMAC, co-cultivation, regulatory network manipulation, heterologous expression, and cell-free biosynthetic systems should be increasingly integrated into genome mining. Multi-omics approaches can help determine which predicted pathways are expressed, which metabolites are produced, and under what environmental or physiological conditions they are activated. A clear distinction between predicted BGCs, expressed pathways, chemically characterized metabolites, and experimentally demonstrated biological functions will improve reproducibility and facilitate comparisons across studies. The continuous development and curation of resources such as MIBiG and BGC family databases, along with standardized metadata describing genomic, chemical, ecological, and experimental information, will be essential for training increasingly robust predictive models and reducing redundancy in natural product discovery.
Taken together, these advances point toward a more integrated bioprospecting model, in which computational prediction, experimental biology, chemical analysis, and environmental validation cease to be independent stages and instead become interconnected components of a single discovery workflow. The ultimate goal is not merely to increase the number of identified BGCs, but to improve the efficiency with which novel and biologically relevant metabolites are selected, produced, characterized, and translated into practical applications.
Author Contributions
Conceptualization, J.N.R., A.S.S., A.F.S., M.L.L.B. and L.V.C.; methodology, J.N.R., A.S.S., A.F.S., M.L.L.B. and L.V.C.; validation, J.N.R., A.S.S., A.F.S., M.L.L.B. and L.V.C.; formal analysis, J.N.R., A.S.S., A.F.S., M.L.L.B. and L.V.C.; investigation, J.N.R., A.S.S., A.F.S., M.L.L.B. and L.V.C.; data curation, J.N.R., A.S.S., A.F.S., M.L.L.B. and L.V.C.; writing—original draft preparation, J.N.R., A.S.S., A.F.S., M.L.L.B. and L.V.C.
Funding
This study was funded by Fundação de Amparo à Pesquisa do Estado do Rio de Janeiro – FAPERJ (E-26/200.546/2025, E-26/210.563/2025 and E-26/204.245/2025-BOLSA) and CNPq in the form of a Productivity in Research Fellowship (PQ-C) (Process: 304877/2024-7).
Institutional Review Board Statement
Not applicable.
Data Availability Statement
No new data were created or analyzed in this study. Data sharing is not applicable to this article.
Acknowledgments
During the preparation of this manuscript, the author(s) used GPT-5.6 Luna to generate and improv the quality of the Figures. GPT-5.6 Luna was also used for English revision.
Conflicts of Interest
The authors declare no conflicts of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| 2,4-D | 2,4-Dichlorophenoxyacetic acid |
| AP | Affinity Propagation |
| AI | Artificial intelligence |
| AMR | Antimicrobial resistance |
| AOI | Area of interest identification |
| ARTS | Antibiotic Resistant Target Seeker |
| BAGEL | BActeriocin GEnome mining tooL |
| BBH | Bidirectional best hits |
| BGC | Biosynthetic gene cluster |
| BiG-SCAPE | Biosynthetic Gene Similarity Clustering and Prospecting Engine |
| BiG-SLiCE | Biosynthetic Genes Super-Linear Clustering Engine |
| BiG-FAM | Biosynthetic Gene Cluster Families Database |
| BiLSTM | Bidirectional Long Short-Term Memory |
| BLASTP | Basic local alignment search tool (protein) |
| CFU | Colony-forming units |
| CLSI | Clinical and Laboratory Standards Institute |
| DAPG | 2,4- Diacetylphloroglucinol |
| EUCAST | European Committee on Antimicrobial Susceptibility Testing |
| FT-IR | Fourier-transform infrared spectroscopy |
| GC-MS | Gas chromatography – mass spectrometry |
| GCFs | Gene Cluster Families |
| GECCO GNM |
Gene Cluster prediction with COnditional random fields Gaussian network model |
| GO | Gene Ontology |
| GPC | Gel permeation chromatography |
| HGT | Horizontal gene transfer |
| HPLC | High-performance liquid chromatography |
| HPLC-HRMS | High-performance liquid chromatography–high-resolution mass spectrometry |
| ICP-MS | Inductively Coupled Plasma-Mass Spectrometry |
| ICP-OES | Inductively Coupled Plasma Optical Emission Spectrometry |
| kb | kilobase |
| LC-MS | Liquid chromatography-mass spectrometry |
| LPSN | List of Prokaryotic names with Standing in Nomenclature |
| MALDI-TOF MS | Matrix-assisted laser desorption/ionization time-of-flight mass spectrometry |
| Mbp | megabase |
| MBC | Minimum bactericidal concentration |
| MBEC | Minimum biofilm eradication concentration |
| MBIC | Minimum biofilm inhibitory concentration |
| MFC | Minimum fungicidal concentration |
| MIBiG | Minimum Information about a Biosynthetic Gene Cluster |
| MIC MSA |
Minimum inhibitory concentration Multiple sequence alignment |
| NLP NMA |
Natural language processing Normal mode analysis |
| NRPSs | Non-ribosomal peptide synthetases |
| OD OECD |
Optical density Organization for Economic Co-operation and Development |
| ORF OSMAC |
Open reading frame One Strain-Many Compounds |
| PET | Polyethylene terephthalate |
| Pfam | Protein families |
| pHMMs | Profile hidden Markov models |
| pLDDT | predicted Local Distance Difference Test |
| PPI | Protein-protein interaction |
| RiPPs | Ribosomally synthesized and post-translationally modified peptides |
| RNA-Seq | RNA sequencing |
| SMILES | Simplified Molecular Input Line Entry System |
| T2PKSs | Type II polyketide synthases |
| T3PKSs | Type III polyketide synthases |
References
- GBD. Global burden of bacterial antimicrobial resistance 1990-2021: a systematic analysis with forecasts to 2050. Lancet 2021, 404, 1199–226. [Google Scholar]
- Ho, C.S.; Wong, C.T.H.; Aung, T.T.; Lakshminarayanan, R.; Mehta, J.S.; Rauz, S.; et al. Antimicrobial resistance: a concise update. In Lancet Microbe [Internet]; Elsevier Ltd., 2025; p. 6. [Google Scholar] [CrossRef] [PubMed]
- Brüssow, H. The antibiotic resistance crisis and the development of new antibiotics. In Microb Biotechnol [Internet]. Microb Biotechnol; 2024; p. 17. [Google Scholar] [CrossRef] [PubMed]
- da Cunha, B.R.; Fonseca, L.P.; Calado, C.R.C. Antibiotic Discovery: Where Have We Come from, Where Do We Go? Antibiotics (Basel) [Internet]. In Antibiotics (Basel); 2019. [Google Scholar] [CrossRef] [PubMed]
- Van Goethem, M.W.; Marasco, R.; Hong, P.Y.; Daffonchio, D. The antibiotic crisis: On the search for novel antibiotics and resistance mechanisms. In Microb Biotechnol [Internet]. Microb Biotechnol; 2024; p. 17. [Google Scholar] [CrossRef] [PubMed]
- Sharma, P.; Singh, S.P.; Iqbal, H.M.N.; Tong, Y.W. Omics approaches in bioremediation of environmental contaminants: An integrated approach for environmental safety and sustainability. In Environ Res [Internet]; Academic Press, 2022; Volume 211, p. 113102. [Google Scholar] [CrossRef] [PubMed]
- Ahmed, E.; Holmström, S.J.M. Siderophores in environmental research: Roles and applications. In Microb Biotechnol [Internet]; REQUESTEDJOURNAL:JOURNAL:17517915; John Wiley and Sons Ltd.: WGROUP; STRING:PUBLICATION, 2014; Volume 7, pp. 196–208. [Google Scholar] [CrossRef] [PubMed]
- Ayangbenro, A.S.; Babalola, O.O. Genomic analysis of Bacillus cereus NWUAB01 and its heavy metal removal from polluted soil. In Scientific Reports; Nature Publishing Group, 2020; Volume 10:1 10, p. 19660. [Google Scholar] [CrossRef] [PubMed]
- Kumari, S.; Das, S. Bacterial enzymatic degradation of recalcitrant organic pollutants: catabolic pathways and genetic regulations. In Environmental Science and Pollution Research; Springer, 2023; Volume 2023 30:33 30, pp. 79676–705. [Google Scholar] [CrossRef] [PubMed]
- Zdouc, M.M.; Blin, K.; Louwen, N.L.L.; Navarro, J.; Loureiro, C.; Bader, C.D.; et al. MIBiG 4.0: advancing biosynthetic gene cluster curation through global collaboration. Nucleic Acids Res [Internet]. Nucleic Acids Res. 2025, 53, D678–90. [Google Scholar] [CrossRef] [PubMed]
- Peckle, B.A.; da Silva, S.; Ribeiro, JR de A; de Oliveira, S.S.; Vianez-Júnior, J.L.S.G.; Direito, I.C.N.; et al. The Genome of Enterobacter hormaechei Strain MG02, a 2,4-Dichlorophenoxyacetic Acid-Degrading Bacterium Isolated from Brazilian Soil. In Microbiol Resour Announc [Internet]; WEBSITE:WEBSITE:ASMJ;JOURNAL:JOURNAL:GENOMEA;ISSUE:ISSUE:DOI; American Society for Microbiology, 2022; p. 11. [Google Scholar] [CrossRef] [PubMed]
- Succar, J.B.; da Silva, A.S.; Berbert, L.C.; Flores, V.R.; Ferreira, J.V.R.; Cardoso, A.M.; et al. Identificação de bactérias com a capacidade de biodegradação do herbicida ácido 2,4-diclorofenoxiacético. Análise Crítica das Ciências Biológicas e da Natureza 3 [Internet]; Atena Editora, 2019 [cited 2026 Sep 3. Available online: https://atenaeditora.com.br/catalogo/post/identificacao-de-bacterias-com-a-capacidade-de-biodegradacao-do-herbicida-acido-24-diclorofenoxiacetico (accessed on 3 Sep 2026).
- da Silva, S.; Peckle, B.A.; Ribeiro, JR de A; de Oliveira, S.S.; Bianco, K.; Clementino, M.M.; et al. The Genome Sequence of Brucella intermedia DF13, a 2,4-Dichlorophenoxyacetic Acid-Degrading Soil Bacterium Isolated in Brazil. In Microbiol Resour Announc [Internet]; American Society for Microbiology: WGROUP; STRING:PUBLICATION, 2022; p. 11. [Google Scholar] [CrossRef] [PubMed]
- Blin, K.; Shaw, S.; Vader, L.; Szenei, J.; Reitz, Z.L.; Augustijn, H.E.; et al. antiSMASH 8.0: extended gene cluster detection capabilities and analyses of chemistry, enzymology, and regulation. In Nucleic Acids Res [Internet]; Oxford Academic, 2025; Volume 53, pp. W32–8. [Google Scholar] [CrossRef] [PubMed]
- Hannigan, G.D.; Prihoda, D.; Palicka, A.; Soukup, J.; Klempir, O.; Rampula, L.; et al. A deep learning genome-mining strategy for biosynthetic gene cluster prediction. In Nucleic Acids Res [Internet]; Oxford Academic, 2019; Volume 47, pp. e110–e110. [Google Scholar] [CrossRef] [PubMed]
- Medema, M.H.; Blin, K.; Cimermancic, P.; De Jager, V.; Zakrzewski, P.; Fischbach, M.A.; et al. antiSMASH: rapid identification, annotation and analysis of secondary metabolite biosynthesis gene clusters in bacterial and fungal genome sequences. Nucleic Acids Res. [Internet] Nucleic Acids Res. 2011, 39. [Google Scholar] [CrossRef] [PubMed]
- Chen, R.; Wong, H.L.; Burns, B.P. New Approaches to Detect Biosynthetic Gene Clusters in the Environment. In Medicines [Internet]; MDPI AG, 2019; Volume 6. [Google Scholar] [CrossRef] [PubMed]
- Freese, H.M.; Meier-Kolthoff, J.P.; Sardà Carbasse, J.; Afolayan, A.O.; Göker, M. TYGS and LPSN in 2025: a Global Core Biodata Resource for genome-based classification and nomenclature of prokaryotes within DSMZ Digital Diversity. Nucleic Acids Res. [Internet] Nucleic Acids Res. 2026, 54, D884–91. [Google Scholar] [CrossRef] [PubMed]
- Hug, J.J.; Krug, D.; Müller, R. Bacteria as genetically programmable producers of bioactive natural products. Nat. Rev. Chem. [Internet] Nat. Rev. Chem. 2020, 4, 172–93. [Google Scholar] [CrossRef] [PubMed]
- Keatinge-Clay, A.T. The structures of type I polyketide synthases. Nat. Prod. Rep. [Internet] Nat. Prod. Rep. 2012, 29, 1050–73. [Google Scholar] [CrossRef] [PubMed]
- Süssmuth, R.D.; Mainz, A. Nonribosomal Peptide Synthesis-Principles and Prospects. Angew. Chem. Int. Ed. Engl. [Internet] Angew. Chem. Int. Ed. Engl. 2017, 56, 3770–821. [Google Scholar] [CrossRef] [PubMed]
- Rutledge, P.J.; Challis, G.L. Discovery of microbial natural products by activation of silent biosynthetic gene clusters. Nat. Rev. Microbiol. [Internet] Nat. Rev. Microbiol. 2015, 13, 509–23. [Google Scholar] [CrossRef] [PubMed]
- Yan, Y.; Liu, Q.; Jacobsen, S.E.; Tang, Y. The impact and prospect of natural product discovery in agriculture: New technologies to explore the diversity of secondary metabolites in plants and microorganisms for applications in agriculture. EMBO Rep. [Internet] EMBO Rep. 2018, 19. [Google Scholar] [CrossRef] [PubMed]
- Lee, N.; Hwang, S.; Kim, J.; Cho, S.; Palsson, B.; Cho, B.K. Mini review: Genome mining approaches for the identification of secondary metabolite biosynthetic gene clusters in Streptomyces. Comput Struct. Biotechnol. J. [Internet] Comput Struct. Biotechnol. J. 2020, 18, 1548–56. [Google Scholar] [CrossRef] [PubMed]
- Belknap, K.C.; Park, C.J.; Barth, B.M.; Andam, C.P. Genome mining of biosynthetic and chemotherapeutic gene clusters in Streptomyces bacteria. In Sci Rep [Internet]. Sci Rep; 2020; p. 10. [Google Scholar] [CrossRef] [PubMed]
- Creamer, K.E.; Castro-Falcón, G.; Ince, E.; Vasilat, V.; Gorbitz, D.V.; Demko, A.M.; et al. Taxonomic and biosynthetic diversity of the marine actinomycete Salinispora across spatial scales. In Appl Environ Microbiol [Internet]. Appl Environ Microbiol; 2026; p. 92. [Google Scholar] [CrossRef] [PubMed]
- Duncan, K.R.; Crüsemann, M.; Lechner, A.; Sarkar, A.; Li, J.; Ziemert, N.; et al. Molecular networking and pattern-based genome mining improves discovery of biosynthetic gene clusters and their products from salinispora species. In Chem Biol [Internet]; Elsevier Ltd., 2015; Volume 22, pp. 460–71. [Google Scholar] [CrossRef] [PubMed]
- Jensen, P.R.; Moore, B.S.; Fenical, W. The marine actinomycete genus Salinispora: a model organism for secondary metabolite discovery. Nat. Prod. Rep. [Internet] Nat. Prod. Rep. 2015, 32, 738–51. [Google Scholar] [CrossRef] [PubMed]
- Joshua, K.; Rajasulocha, P. A Review on Isolation, Identification of Bacillus and Antimicrobial Activity Detection. Ann. Romanian Soc. Cell Biol. 2021, 25, 4709–17. [Google Scholar]
- Göker, M.; Christensen, H.; Fingerle, V.; Kostovski, M.; Margos, G.; Moore, E.R.B.; et al. List of Recommended Names for bacteria of medical importance: report of the Ad Hoc Committee on Mitigating Changes in Prokaryotic Nomenclature. Int. J. Syst. Evol. Microbiol. [Internet] Microbiol. Soc. 2025, 75, 006943. [Google Scholar] [CrossRef]
- Yin, Q.J.; Ying, T.T.; Zhou, Z.Y.; Hu, G.A.; Yang, C.L.; Hua, Y.; et al. Species-specificity of the secondary biosynthetic potential in Bacillus. In Front Microbiol [Internet]; Frontiers Media SA, 2023; Volume 14. [Google Scholar] [CrossRef]
- Ongena, M.; Jacques, P. Bacillus lipopeptides: versatile weapons for plant disease biocontrol. In Trends Microbiol [Internet]; Elsevier Ltd., 2008; Volume 16, pp. 115–25. [Google Scholar] [CrossRef] [PubMed]
- Khurana, H.; Sharma, M.; Verma, H.; Lopes, B.S.; Lal, R.; Negi, R.K. Genomic insights into the phylogeny of Bacillus strains and elucidation of their secondary metabolic potential. In Genomics [Internet]; Academic Press, 2020; Volume 112, pp. 3191–200. [Google Scholar] [CrossRef] [PubMed]
- Van Santen, J.A.; Poynton, E.F.; Iskakova, D.; Mcmann, E.; Alsup, T.A.; Clark, T.N.; et al. The Natural Products Atlas 2.0: a database of microbially-derived natural products. In Nucleic Acids Res [Internet]; Oxford Academic, 2022; Volume 50, pp. D1317–23. [Google Scholar] [CrossRef] [PubMed]
- Xiao, S.; Chen, N.; Chai, Z.; Zhou, M.; Xiao, C.; Zhao, S.; et al. Secondary Metabolites from Marine-Derived Bacillus: A Comprehensive Review of Origins, Structures, and Bioactivities. Marine Drugs 2022, Vol 20, Page 567 [Internet]; Multidisciplinary Digital Publishing Institute, 2022; Volume 20. [Google Scholar] [CrossRef] [PubMed]
- Zhang, B.; Xu, L.; Ding, J.; Wang, M.; Ge, R.; Zhao, H.; et al. Natural antimicrobial lipopeptides secreted by Bacillus spp. and their application in food preservation, a critical review. In Trends Food Sci Technol [Internet]; Elsevier, 2022; Volume 127, pp. 26–37. [Google Scholar] [CrossRef]
- Morandini, L.; Caulier, S.; Bragard, C.; Mahillon, J. Bacillus cereus sensu lato antimicrobial arsenal: An overview. In Microbiol Res [Internet]; Elsevier GmbH, 2024; p. 283. [Google Scholar] [CrossRef] [PubMed]
- Sansinenea, E.; Ortiz, A. Secondary metabolites of soil Bacillus spp. Biotechnol Lett [Internet]. Biotechnol. Lett. 2011, 33, 1523–38. [Google Scholar] [CrossRef] [PubMed]
- Xia, L.; Miao, Y.; Cao, A.; Liu, Y.; Liu, Z.; Sun, X.; et al. Biosynthetic gene cluster profiling predicts the positive association between antagonism and phylogeny in Bacillus. In Nature Communications; Nature Publishing Group, 2022; Volume 2022 13:1 13, p. 1023. [Google Scholar] [CrossRef] [PubMed]
- Jähne, J.; Herfort, S.; Doellinger, J.; Lasch, P.; Tam, L.T.T.; Borriss, R.; et al. Investigation of the potential of Brevibacillus spp. for the biosynthesis of nonribosomally produced bioactive compounds by combination of genome mining with MALDI-TOF mass spectrometry. In Front Microbiol [Internet]; Frontiers Media SA, 2023; Volume 14. [Google Scholar] [CrossRef]
- Jähne, J.; Le Thi, T.T.; Blumenscheit, C.; Schneider, A.; Pham, T.L.; Le Thi, P.T.; et al. Novel Plant-Associated Brevibacillus and Lysinibacillus Genomospecies Harbor a Rich Biosynthetic Potential of Antimicrobial Compounds. Microorganisms [Internet] MDPI 2023, 11, 168. [Google Scholar] [CrossRef]
- Kim, B.; Han, S.R.; Lee, H.; Oh, T.J. Insights into group-specific pattern of secondary metabolite gene cluster in Burkholderia genus. Front Microbiol. [Internet] Front Microbiol. 2024, 14. [Google Scholar] [CrossRef] [PubMed]
- Mahmoud, F.M.; Pritsch, K.; Siani, R.; Benning, S.; Radl, V.; Kublik, S.; et al. Comparative genomic analysis of strain Priestia megaterium B1 reveals conserved potential for adaptation to endophytism and plant growth promotion. In Microbiol Spectr [Internet]; American Society for Microbiology; JOURNAL:JOURNAL:SPECTRUM; ISSUE:ISSUE:DOI, 2024; p. 12. [Google Scholar] [CrossRef] [PubMed]
- Yang, F.; Jiang, H.; Ma, K.; Hegazy, A.; Wang, X.; Liang, S.; et al. Genomic and phenotypic analyses reveal Paenibacillus polymyxa PJH16 is a potential biocontrol agent against cucumber fusarium wilt. In Front Microbiol [Internet]; Frontiers Media SA, 2024; Volume 15. [Google Scholar] [CrossRef]
- Adeniji, A.A.; Chukwuneme, C.F.; Conceição, E.C.; Ayangbenro, A.S.; Wilkinson, E.; Maasdorp, E.; et al. Unveiling novel features and phylogenomic assessment of indigenous Priestia megaterium AB-S79 using comparative genomics. In Microbiol Spectr [Internet]; JOURNAL:JOURNAL:SPECTRUM; American Society for Microbiology; WGROUP:STRING:PUBLICATION, 2025; p. 13. [Google Scholar] [CrossRef] [PubMed]
- Masschelein, J.; Jenner, M.; Challis, G.L. Antibiotics from Gram-negative bacteria: a comprehensive overview and selected biosynthetic highlights. Nat. Prod. Rep. [Internet] Nat. Prod. Rep. 2017, 34, 712–83. [Google Scholar] [CrossRef] [PubMed]
- Birkelbach, J.; Seyfert, C.E.; Walesch, S.; Müller, R. Harnessing Gram-negative bacteria for novel anti-Gram-negative antibiotics. In Microb Biotechnol [Internet]. Microb Biotechnol; 2024; p. 17. [Google Scholar] [CrossRef] [PubMed]
- Alam, K.; Islam, M.M.; Li, C.; Sultana, S.; Zhong, L.; Shen, Q.; et al. Genome Mining of Pseudomonas Species: Diversity and Evolution of Metabolic and Biosynthetic Potential. Molecules 2021, Vol 26, Page 7524 [Internet]; Multidisciplinary Digital Publishing Institute, 2021; Volume 26. [Google Scholar] [CrossRef] [PubMed]
- Saati-Santamaría, Z.; Selem-Mojica, N.; Peral-Aranega, E.; Rivas, R.; García-Fraile, P. Unveiling the genomic potential of Pseudomonas type strains for discovering new natural products. Microb. Genom. [Internet] Microb. Genom. 2022, 8. [Google Scholar] [CrossRef] [PubMed]
- Mullins, A.J.; Mahenthiralingam, E. The Hidden Genomic Diversity, Specialized Metabolite Capacity, and Revised Taxonomy of Burkholderia Sensu Lato. In Front Microbiol; Frontiers Media S.A., 2021; Volume 12. [Google Scholar] [CrossRef]
- French, C.T.; Bulterys, P.L.; Woodward, C.L.; Tatters, A.O.; Ng, K.R.; Miller, J.F. Virulence from the rhizosphere: ecology and evolution of Burkholderia pseudomallei-complex species. In Curr Opin Microbiol [Internet]; Elsevier Ltd., 2020; Volume 54, pp. 18–32. [Google Scholar] [CrossRef] [PubMed]
- Liu, X.; Cheng, Y.Q. Genome-guided discovery of diverse natural products from Burkholderia sp. J Ind Microbiol Biotechnol [Internet]. J. Ind. Microbiol. Biotechnol. 2014, 41, 275–84. [Google Scholar] [CrossRef] [PubMed]
- Petrova, Y.D.; Mahenthiralingam, E. Discovery, mode of action and secretion of Burkholderia sensu lato key antimicrobial specialised metabolites. In The Cell Surface [Internet]; Elsevier B.V., 2022; p. 8. [Google Scholar] [CrossRef] [PubMed]
- Esmaeel, Q.; Pupin, M.; Kieu, N.P.; Chataigné, G.; Béchet, M.; Deravel, J.; et al. Burkholderia genome mining for nonribosomal peptide synthetases reveals a great potential for novel siderophores and lipopeptides synthesis. Microbiologyopen [Internet] Microbiol. 2016, 5, 512–26. [Google Scholar] [CrossRef] [PubMed]
- Niehs, S.P.; Kumpfmüller, J.; Dose, B.; Little, R.F.; Ishida, K.; Flórez, L. V.; et al. Insect-Associated Bacteria Assemble the Antifungal Butenolide Gladiofungin by Non-Canonical Polyketide Chain Termination. Angew. Chem. Int. Ed. Engl. [Internet] Angew. Chem. Int. Ed. Engl. 2020, 59, 23122–6. [Google Scholar] [CrossRef] [PubMed]
- Fergusson, C.H.; Saulog, J.; Paulo, B.S.; Wilson, D.M.; Liu, D.Y.; Morehouse, N.J.; et al. Discovery of a lagriamide polyketide by integrated genome mining, isotopic labeling, and untargeted metabolomics. Chem. Sci. [Internet] Chem. Sci. 2024, 15, 8089–96. [Google Scholar] [CrossRef] [PubMed]
- Webster, G.; Mullins, A.J.; Mahenthiralingam, E. Complete genome sequence of the biopesticidal Burkholderia ambifaria strain BCC0191. Microbiol Resour Announc [Internet]. Microbiol. Resour. Announc 2025, 14. [Google Scholar] [CrossRef] [PubMed]
- Sharma, P.; Johnson, M.A.; Mazloom, R.; Allen, C.; Heath, L.S.; Lowe-Power, T.M.; et al. Meta-analysis of the Ralstonia solanacearum species complex (RSSC) based on comparative evolutionary genomics and reverse ecology. Microb Genom [Internet]. Microb. Genom. 2022, 8. [Google Scholar] [CrossRef] [PubMed]
- Kreutzer, M.F.; Kage, H.; Gebhardt, P.; Wackler, B.; Saluz, H.P.; Hoffmeister, D.; et al. Biosynthesis of a complex yersiniabactin-like natural product via the mic locus in phytopathogen Ralstonia solanacearum. Appl Environ Microbiol [Internet]. Appl. Environ. Microbiol. 2011, 77, 6117–24. [Google Scholar] [CrossRef] [PubMed]
- Safni, I.; Subandiyah, S.; Fegan, M. Ecology, Epidemiology and Disease Management of Ralstonia syzygii in Indonesia. Front Microbiol. [Internet] Front Microbiol. 2018, 9. [Google Scholar] [CrossRef] [PubMed]
- Liao, L.; Lin, D.; Liu, Z.; Gao, Y.; Hu, K. A case of meningitis caused by Ralstonia insidiosa, a rare opportunistic pathogen. BMC Infect. Dis. [Internet] BMC Infect. Dis. 2023, 23. [Google Scholar] [CrossRef] [PubMed]
- Neelambaran, K.; Velmurugan, H.; Venkatesan, S.; Thangaraju, P. Ralstonia mannitolilytica Infections: A Systematic Review of Case Reports Unveiling Clinical Patterns and Therapeutic Insights. Curr. Drug Res. Rev. [Internet] Curr. Drug Res. Rev. 2025, 17, 293–300. [Google Scholar] [CrossRef] [PubMed]
- Rodrigues da Silva, I.; Helena Simões Villas Bôas, M.; Oswaldo Cruz, F.; Veloso da Costa Fundação Oswaldo Cruz, L.; Luiz Lima Brandão Fundação Oswaldo Cruz, M. Desafios de identificação e controle de Ralstonia pickettii na indústria farmacêutica: uma revisão integrativa da literatura. Rev. Científica Do UBM [Internet] 2026, 55, 340–58. [Google Scholar]
- Ohtsubo, Y.; Fujita, N.; Nagata, Y.; Tsuda, M.; Iwasaki, T.; Hatta, T. Complete Genome Sequence of Ralstonia pickettii DTP0602, a 2,4,6-Trichlorophenol Degrader. Genome Announc [Internet] Genome Announc 2013, 1. [Google Scholar] [CrossRef] [PubMed]
- Kai, K.; Ohnishi, H.; Kiba, A.; Ohnishi, K.; Hikichi, Y. Studies on the biosynthesis of ralfuranones in Ralstonia solanacearum. Biosci. Biotechnol. Biochem [Internet] Biosci. Biotechnol. Biochem 2016, 80, 440–4. [Google Scholar] [CrossRef] [PubMed]
- Lu, C.H.; Zhang, Y.Y.; Jiang, N.; Chen, W.; Shao, X.; Zhao, Z.M.; et al. Ralstonia chuxiongensis sp. nov., Ralstonia mojiangensis sp. nov., and Ralstonia soli sp. nov., isolated from tobacco fields, are three novel species in the family Burkholderiaceae. Front Microbiol. [Internet] Front Microbiol. 2023, 14. [Google Scholar] [CrossRef] [PubMed]
- Montecillo, A.D.; Raymundo, A.K.; Papa, I.A.; Aquino, G.M.B.; Jacildo, A.J.; Stothard, P.; et al. Near-Complete Genome Sequence of Ralstonia solanacearum T523, a Phylotype I Tomato Phytopathogen Isolated from the Philippines. In Microbiol Resour Announc [Internet]; American Society for Microbiology, 2018; Volume 7, p. e01048-18. [Google Scholar] [CrossRef] [PubMed]
- Liu, J.Y.; Zhang, J.F.; Wu, H.L.; Chen, Z.; Li, S.Y.; Li, H.M.; et al. Proposal to classify Ralstonia solanacearum phylotype I strains as Ralstonia nicotianae sp. nov., and a genomic comparison between members of the genus Ralstonia. Front Microbiol. [Internet] Front Microbiol. 2023, 14. [Google Scholar] [CrossRef] [PubMed]
- Zhao, Y.; Ding, W.J.; Xu, L.; Sun, J.Q. A comprehensive comparative genomic analysis revealed that plant growth promoting traits are ubiquitous in strains of Stenotrophomonas. In Front Microbiol [Internet]; Frontiers Media SA, 2024; Volume 15. [Google Scholar] [CrossRef] [PubMed]
- Souza, P.A.; dos Santos, M.C.S.; da Silva Lage de Miranda, R.V.; da Costa, L.V.; da Silva, R.P.P.; Silva, C.; et al. Evaluation of biofilm formation, antimicrobial pattern, and typing of Stenotrophomonas maltophilia isolated from clinical sources in Brazil. Braz. J. Microbiol. [Internet] Braz. J. Microbiol. 2025, 56, 1861–71. [Google Scholar] [CrossRef] [PubMed]
- Sahu, P.K.; Nanda, K.D.; Kale, N.; Gupta, A.; Rai, N.; Singla, D.; et al. Untargeted metabolomic profiling and genome mining of endophytic Stenotrophomonas maltophilia strain 3A reveal a rich source of bioactive secondary metabolites. Front Microbiol. Front. 2026, 17, 1792452. [Google Scholar] [CrossRef] [PubMed]
- Kumar, A.; Rithesh, L.; Kumar, V.; Raghuvanshi, N.; Chaudhary, K.; Abhineet; et al. Stenotrophomonas in diversified cropping systems: friend or foe? Front Microbiol [Internet]. Front Microbiol. 2023, 14. [Google Scholar] [CrossRef] [PubMed]
- Hayward, A.C.; Fegan, N.; Fegan, M.; Stirling, G.R. Stenotrophomonas and Lysobacter: ubiquitous plant-associated gamma-proteobacteria of developing significance in applied microbiology. J. Appl. Microbiol. [Internet] J. Appl. Microbiol. 2010, 108, 756–70. [Google Scholar] [CrossRef] [PubMed]
- Jakobi, M.; Winkelmann, G.; Kaiser, D.; Kempter, C.; Jung, G.; Berg, G.; et al. Maltophilin: a new antifungal compound produced by Stenotrophomonas maltophilia R3089. J. Antibiot. 1996, 49, 1101–4. [Google Scholar] [CrossRef] [PubMed]
- Nakayama, T.; Homma, Y.; Hashidoko, Y.; Mizutani, J.; Tahara, S. Possible role of xanthobaccins produced by Stenotrophomonas sp. strain SB-K88 in suppression of sugar beet damping-off disease. Appl. Environ. Microbiol. [Internet] Appl. Environ. Microbiol. 1999, 65, 4334–9. [Google Scholar] [CrossRef] [PubMed]
- Abdelsalam, N.A.; Elhadidy, M.; Saif, N.A.; Elsayed, S.W.; Mouftah, S.F.; Sayed, A.A.; et al. Biosynthetic gene cluster signature profiles of pathogenic Gram-negative bacteria isolated from Egyptian clinical settings. In Microbiol Spectr [Internet]. Microbiol Spectr; 2023; p. 11. [Google Scholar] [CrossRef] [PubMed]
- Alejo, M.A.; Sarahi Lozano Gamboa, M.; Muñoz Gomez, B.; Hernández Magro Gil, K.G.; García-Contreras, R.; Whitaker, R.J.; et al. Resistome, virulome, mobilome, and biosynthetic gene clusters adaptations of Acinetobacter baumannii Mexican strains before and during the COVID-19 pandemic: insights from whole-genome sequencing. In Front Public Health [Internet]; Frontiers, 2026; Volume 14. [Google Scholar] [CrossRef] [PubMed]
- Mouhib, S.; Ait Si Mhand, K.; Radouane, N.; Errafii, K.; Kadmiri, I.M.; Andrade-Molina, D.; et al. Unveiling Acinetobacter endophylla sp. nov.: A Specialist Endophyte from Peganum harmala with Distinct Genomic and Metabolic Traits. Microorganisms [Internet] Microorg. 2025, 13. [Google Scholar] [CrossRef] [PubMed]
- Dinglasan, J.L.N.; Otani, H.; Doering, D.T.; Udwary, D.; Mouncey, N.J. Microbial secondary metabolites: advancements to accelerate discovery towards application. Nat. Rev. Microbiol. [Internet] Nat. Rev. Microbiol. 2025, 23, 338–54. [Google Scholar] [CrossRef] [PubMed]
- Liu, G.; Catacutan, D.B.; Rathod, K.; Swanson, K.; Jin, W.; Mohammed, J.C.; et al. Deep learning-guided discovery of an antibiotic targeting Acinetobacter baumannii. Nat. Chem. Biol. [Internet] Nat. Chem. Biol. 2023, 19, 1342–50. [Google Scholar] [CrossRef] [PubMed]
- Carroll, L.M.; Larralde, M.; Fleck, J.S.; Ponnudurai, R.; Milanese, A.; Cappio, E.; et al. Accurate de novo identification of biosynthetic gene clusters with GECCO; Cold Spring Harbor Laboratory, 2021; p. 2021.05.03.442509. [Google Scholar] [CrossRef]
- Liu, M.; Li, Y.; Li, H. Deep Learning to Predict the Biosynthetic Gene Clusters in Bacterial Genomes. J. Mol. Biol. [Internet] J. Mol. Biol. 2022, 434. [Google Scholar] [CrossRef] [PubMed]
- Draisma, A.; Loureiro, C.; Louwen, N.L.L.; Kautsar, S.A.; Navarro-Muñoz, J.C.; Doering, D.T.; et al. BiG-SCAPE 2.0 and BiG-SLiCE 2.0: scalable, accurate and interactive sequence clustering of metabolic gene clusters. In Nature Communications; Nature Publishing Group, 2026; Volume 2026 17:1, p. 17:2000. [Google Scholar] [CrossRef] [PubMed]
- Udwary, D.W.; Doering, D.T.; Foster, B.; Smirnova, T.; Kautsar, S.A.; Mouncey, N.J. The secondary metabolism collaboratory: a database and web discussion portal for secondary metabolite biosynthetic gene clusters. Nucleic Acids Res [Internet]. Nucleic Acids Res. 2025, 53, D717–23. [Google Scholar] [CrossRef] [PubMed]
- Zhu, S.; Xu, H.; Liu, Y.; Hong, Y.; Yang, H.; Zhou, C.; et al. Computational advances in biosynthetic gene cluster discovery and prediction. In Biotechnol Adv [Internet]; Elsevier Inc., 2025; p. 79. [Google Scholar] [CrossRef] [PubMed]
- Van Heel, A.J.; De Jong, A.; Song, C.; Viel, J.H.; Kok, J.; Kuipers, O.P. BAGEL4: a user-friendly web server to thoroughly mine RiPPs and bacteriocins. Nucleic Acids Res [Internet]. Nucleic Acids Res. 2018, 46, W278–81. [Google Scholar] [CrossRef] [PubMed]
- Fernandez-Cantos, M.; Kuipers, O.; de Jong, A. CHAPTER 4. BAGEL5: Improved and extended mining of bacteriocins by rapid analysis of extensive (meta-) genomic datasets. In Gut Bacteroidales: antimicrobial potencies and host-bacteria interactions; University of Groningen, 2024. [Google Scholar] [CrossRef]
- Séelem-Mojica, N.; Aguilar, C.; Gutiéerrez-García, K.; Martínez-Guerrero, C.E.; Barona-Gómez, F. EvoMining reveals the origin and fate of natural product biosynthetic enzymes. Microb. Genom. [Internet] Microb. Genom. 2019, 5. [Google Scholar] [CrossRef] [PubMed]
- Alanjary, M.; Kronmiller, B.; Adamek, M.; Blin, K.; Weber, T.; Huson, D.; et al. The Antibiotic Resistant Target Seeker (ARTS), an exploration engine for antibiotic cluster prioritization and novel drug target discovery. Nucleic Acids Res. [Internet] Nucleic Acids Res. 2017, 45, W42–8. [Google Scholar] [CrossRef] [PubMed]
- Mungan, M.D.; Alanjary, M.; Blin, K.; Weber, T.; Medema, M.H.; Ziemert, N. ARTS 2.0: feature updates and expansion of the Antibiotic Resistant Target Seeker for comparative genome mining. Nucleic Acids Res. [Internet] Nucleic Acids Res. 2020, 48, W546–52. [Google Scholar] [CrossRef] [PubMed]
- Kautsar, S.A.; Blin, K.; Shaw, S.; Weber, T.; Medema, M.H. BiG-FAM: the biosynthetic gene cluster families database. Nucleic Acids Res [Internet]. Nucleic Acids Res. 2021, 49, D490–7. [Google Scholar] [CrossRef] [PubMed]
- Doroghazi, J.R.; Albright, J.C.; Goering, A.W.; Ju, K.S.; Haines, R.R.; Tchalukov, K.A.; et al. A roadmap for natural product discovery based on large-scale genomics and metabolomics. Nat. Chem. Biol. [Internet] Nat. Chem. Biol. 2014, 10, 963–8. [Google Scholar] [CrossRef] [PubMed]
- Schorn, M.A.; Verhoeven, S.; Ridder, L.; Huber, F.; Acharya, D.D.; Aksenov, A.A.; et al. A community resource for paired genomic and metabolomic data mining. Nat. Chem. Biol. [Internet] Nat. Chem. Biol. 2021, 17, 363–8. [Google Scholar] [CrossRef] [PubMed]
- Gordon, C.H.; Hendrix, E.; He, Y.; Walker, M.C. AlphaFold Accurately Predicts the Structure of Ribosomally Synthesized and Post-Translationally Modified Peptide Biosynthetic Enzymes. In Biomolecules; Multidisciplinary Digital Publishing Institute (MDPI), 2023. [Google Scholar] [CrossRef] [PubMed]
- Akdel, M.; Pires, D.E.V.; Pardo, E.P.; Jänes, J.; Zalevsky, A.O.; Mészáros, B.; et al. A structural biology community assessment of AlphaFold2 applications. In Nature Structural & Molecular Biology; Nature Publishing Group, 2022; Volume 2022 29:11 29, pp. 1056–67. [Google Scholar] [CrossRef] [PubMed]
- Hekkelman, M.L.; de Vries, I.; Joosten, R.P.; Perrakis, A. AlphaFill: enriching AlphaFold models with ligands and cofactors. In Nature Methods; Nature Publishing Group, 2022; Volume 2022 20:2 20, pp. 205–13. [Google Scholar] [CrossRef] [PubMed]
- Terwilliger, T.C.; Liebschner, D.; Croll, T.I.; Williams, C.J.; McCoy, A.J.; Poon, B.K.; et al. AlphaFold predictions are valuable hypotheses and accelerate but do not replace experimental structure determination. In Nature Methods; Nature Publishing Group, 2023; Volume 2023 21:1 21, pp. 110–6. [Google Scholar] [CrossRef] [PubMed]
- Jumper, J.; Evans, R.; Pritzel, A.; Green, T.; Figurnov, M.; Ronneberger, O.; et al. Highly accurate protein structure prediction with AlphaFold. Nature [Internet] Nat. 2021, 596, 583–9. [Google Scholar] [CrossRef] [PubMed]
- Mirdita, M.; Schütze, K.; Moriwaki, Y.; Heo, L.; Ovchinnikov, S.; Steinegger, M. ColabFold: making protein folding accessible to all. Nat. Methods [Internet] Nat. Methods 2022, 19, 679–82. [Google Scholar] [CrossRef] [PubMed]
- Abramson, J.; Adler, J.; Dunger, J.; Evans, R.; Green, T.; Pritzel, A.; et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. In Nature; Nature Publishing Group, 2024; Volume 2024 630:8016 630, pp. 493–500. [Google Scholar] [CrossRef] [PubMed]
- Webb, B.; Sali, A. Comparative protein structure modeling using MODELLER. In Curr Protoc Bioinformatics [Internet]; JOURNAL:JOURNAL:1934340X; John Wiley and Sons Inc.: WGROUP; STRING:PUBLICATION, 2016; Volume 2016, pp. 5.6.1–5.6.37. [Google Scholar] [CrossRef] [PubMed]
- Evans, R.; O’Neill, M.; Pritzel, A.; Antropova, N.; Senior, A.; Green, T.; et al. Protein complex prediction with AlphaFold-Multimer. In bioRxiv [Internet]; Cold Spring Harbor Laboratory, 2022; p. 10.04.463034. [Google Scholar] [CrossRef]
- Williams, C.J.; Headd, J.J.; Moriarty, N.W.; Prisant, M.G.; Videau, L.L.; Deis, L.N.; et al. MolProbity: More and better reference data for improved all-atom structure validation. In Protein Science [Internet]; Blackwell Publishing Ltd.: WGROUP; STRING:PUBLICATION, 2018; Volume 27, pp. 293–315. [Google Scholar] [CrossRef] [PubMed]
- Patel, K.D.; MacDonald, M.R.; Ahmed, S.F.; Singh, J.; Gulick, A.M. Structural advances toward understanding the catalytic activity and conformational dynamics of modular nonribosomal peptide synthetases. In Nat Prod Rep [Internet]; The Royal Society of Chemistry, 2023; Volume 40, pp. 1550–82. [Google Scholar] [CrossRef] [PubMed]
- Bagde, S.R.; Kim, C.Y. Architecture of full-length type I modular polyketide synthases revealed by X-ray crystallography, cryo-electron microscopy, and AlphaFold2. In Nat Prod Rep [Internet]; The Royal Society of Chemistry, 2024; Volume 41, pp. 1219–34. [Google Scholar] [CrossRef] [PubMed]
- Bowman, G.R. AlphaFold and Protein Folding: Not Dead Yet! The Frontier Is Conformational Ensembles. In Annu Rev Biomed Data Sci [Internet]; Annual Reviews Inc., 2024; Volume 7, pp. 51–7. [Google Scholar] [CrossRef]
- Tirion, M.M. Large Amplitude Elastic Motions in Proteins from a Single-Parameter, Atomic Analysis. In Phys Rev Lett [Internet]; American Physical Society, 1996; Volume 77, p. 1905. [Google Scholar] [CrossRef] [PubMed]
- Haliloglu, T.; Bahar, I.; Erman, B. Gaussian Dynamics of Folded Proteins. In Phys Rev Lett [Internet]; American Physical Society, 1997; Volume 79, p. 3090. [Google Scholar] [CrossRef]
- Atilgan, A.R.; Durell, S.R.; Jernigan, R.L.; Demirel, M.C.; Keskin, O.; Bahar, I. Anisotropy of fluctuation dynamics of proteins with an elastic network model. In Biophys J [Internet]; Biophysical Society, 2001; Volume 80, pp. 505–15. [Google Scholar] [CrossRef] [PubMed]
- Zhang, S.; Krieger, J.M.; Zhang, Y.; Kaya, C.; Kaynak, B.; Mikulska-Ruminska, K.; et al. ProDy 2.0: increased scale and scope after 10 years of protein dynamics modelling with Python. In Bioinformatics [Internet]; Oxford Academic, 2021; Volume 37, pp. 3657–9. [Google Scholar] [CrossRef] [PubMed]
- Souza, P.C.T.; Alessandri, R.; Barnoud, J.; Thallmair, S.; Faustino, I.; Grünewald, F.; et al. Martini 3: a general purpose force field for coarse-grained molecular dynamics. In Nature Methods; Nature Publishing Group, 2021; Volume 2021 18:4 18, pp. 382–8. [Google Scholar] [CrossRef] [PubMed]
- Karplus, M.; McCammon, J.A. Molecular dynamics simulations of biomolecules. In Nature Structural Biology; Nature Publishing Group, 2002; Volume 2002 9:9 9, pp. 646–52. [Google Scholar] [CrossRef] [PubMed]
- Jussupow, A.; Kaila, V.R.I. Effective Molecular Dynamics from Neural Network-Based Structure Prediction Models. In J Chem Theory Comput [Internet]; American Chemical Society, 2023; Volume 19, pp. 1965–75. [Google Scholar] [CrossRef] [PubMed]
- Del Alamo, D.; Sala, D.; McHaourab, H.S.; Meiler, J. TITLE: Sampling alternative conformational states of transporters and receptors with AlphaFold2. In Elife; eLife Sciences Publications Ltd., 2022; p. 11. [Google Scholar] [CrossRef] [PubMed]
- Wayment-Steele, H.K.; Ojoawo, A.; Otten, R.; Apitz, J.M.; Pitsawong, W.; Hömberger, M.; et al. Predicting multiple conformations via sequence clustering and AlphaFold2. In Nature; Nature Publishing Group, 2023; Volume 2023 625:7996 625, pp. 832–9. [Google Scholar] [CrossRef] [PubMed]
- Monteiro da Silva, G.; Cui, J.Y.; Dalgarno, D.C.; Lisi, G.P.; Rubenstein, B.M. High-throughput prediction of protein conformational distributions with subsampled AlphaFold2. In Nature Communications; Nature Publishing Group, 2024; Volume 2024 15:1 15, p. 2464. [Google Scholar] [CrossRef] [PubMed]
- Schafer, J.W.; Lee, M.; Chakravarty, D.; Thole, J.F.; Chen, E.A.; Porter, L.L. Sequence clustering confounds AlphaFold2. In Nature; Nature Publishing Group, 2025; Volume 2025 638:8051 638, p. E8–12. [Google Scholar] [CrossRef] [PubMed]
- Trott, O.; Olson, A.J. AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization and multithreading. In J Comput Chem [Internet]; Wiley, 2010; Volume 31, p. 455. [Google Scholar] [CrossRef] [PubMed]
- Eberhardt, J.; Santos-Martins, D.; Tillack, A.F.; Forli, S. AutoDock Vina 1.2.0: New Docking Methods, Expanded Force Field, and Python Bindings. J. Chem. Inf. Model [Internet] J. Chem. Inf. Model 2021, 61, 3891–8. [Google Scholar] [CrossRef] [PubMed]
- Gaillard, T. Evaluation of AutoDock and AutoDock Vina on the CASF-2013 Benchmark. J. Chem. Inf. Model [Internet] J. Chem. Inf. Model 2018, 58, 1697–706. [Google Scholar] [CrossRef] [PubMed]
- Schug, A.R.; Bartel, A.; Scholtzek, A.D.; Meurer, M.; Brombach, J.; Hensel, V.; et al. Biocide susceptibility testing of bacteria: Development of a broth microdilution method. In Vet Microbiol [Internet]. Vet Microbiol; 2020; p. 248. [Google Scholar] [CrossRef] [PubMed]
- Strieth, D.; Lenz, S.; Ulber, R. In vivo and in silico screening for antimicrobial compounds from cyanobacteria. In Microbiologyopen [Internet]. Microbiologyopen; 2022; p. 11. [Google Scholar] [CrossRef] [PubMed]
- Hamidi, M.; Toosi, A.M.; Javadi, B.; Asili, J.; Soheili, V.; Shakeri, A. In vitro antimicrobial and antibiofilm screening of eighteen Iranian medicinal plants. BMC Complement Med. Ther. [Internet] BMC Complement Med. Ther. 2024, 24. [Google Scholar] [CrossRef] [PubMed]
- EUCAST. European Committee on Antimicrobial Susceptibility Testing: Breakpoint tables for interpretation of MICs and zone diameters. 2026, 16.1. [Google Scholar]
- Balouiri, M.; Sadiki, M.; Ibnsouda, S.K. Methods for in vitro evaluating antimicrobial activity: A review. J. Pharm. Anal. [Internet] Xi’an Jiaotong Univ. 2016, 6, 71–9. [Google Scholar] [CrossRef] [PubMed]
- Hossain, T.J. Methods for screening and evaluation of antimicrobial activity: A review of protocols, advantages, and limitations. Eur. J. Microbiol. Immunol. (Bp) [Internet]. Eur J Microbiol Immunol (Bp) 2024, 14, 97–115. [Google Scholar] [CrossRef] [PubMed]
- CLSI. Performance standards for antimicrobial susceptibility testing, 34th ed.; Clinical and Laboratory Standards Institute, 2024. [Google Scholar]
- Andrews, J.M. Determination of minimum inhibitory concentrations. In Journal of Antimicrobial Chemotherapy [Internet]; Oxford Academic, 2001; Volume 48, pp. 5–16. [Google Scholar] [CrossRef] [PubMed]
- Nojo, H.; Watanabe, A.; Makimura, K.; Kano, R. Comparing the Minimum Inhibitory Concentrations and Minimum Fungicidal Concentrations of Antifungal Drugs in Microsporum canis. Med. Mycol. J. [Internet] Med. Mycol. J. 2026, 67, 79–81. [Google Scholar] [CrossRef] [PubMed]
- Sóczó, G.; Kardos, G.; McNicholas, P.M.; Balogh, E.; Gergely, L.; Varga, I.; et al. Correlation of posaconazole minimum fungicidal concentration and time kill test against nine Candida species. J. Antimicrob. Chemother. [Internet] J. Antimicrob. Chemother. 2007, 60, 1004–9. [Google Scholar] [CrossRef] [PubMed]
- Ravi, N.S.; Aslam, R.F.; Veeraraghavan, B. A New Method for Determination of Minimum Biofilm Eradication Concentration for Accurate Antimicrobial Therapy. Methods Mol. Biol. [Internet] Methods Mol. Biol. 2019, 1946, 61–7. [Google Scholar] [CrossRef] [PubMed]
- Okae, Y.; Nishitani, K.; Sakamoto, A.; Kawai, T.; Tomizawa, T.; Saito, M.; et al. Estimation of Minimum Biofilm Eradication Concentration (MBEC) on In Vivo Biofilm on Orthopedic Implants in a Rodent Femoral Infection Model. In Front Cell Infect Microbiol [Internet]; Frontiers Media S.A., 2022; Volume 12. [Google Scholar] [CrossRef]
- Malik, G.; Arora, R.; Chaturvedi, R.; Paul, M.S. Implementation of Genetic Engineering and Novel Omics Approaches to Enhance Bioremediation: A Focused Review. In Bulletin of Environmental Contamination and Toxicology; Springer, 2021; Volume 2021 108:3 108, pp. 443–50. [Google Scholar] [CrossRef] [PubMed]
- Armin, R.; Zühlke, S.; Grunewaldt-Stöcker, G.; Mahnkopp-Dirks, F.; Kusari, S. Production of Siderophores by an Apple Root-Associated Streptomyces ciscaucasicus Strain GS2 Using Chemical and Biological OSMAC Approaches. Molecules 2021, Vol 26, Page 3517 [Internet]; Multidisciplinary Digital Publishing Institute, 2021; Volume 26. [Google Scholar] [CrossRef] [PubMed]
- Scherlach, K.; Hertweck, C. Mining and unearthing hidden biosynthetic potential. In Nature Communications; Nature Publishing Group, 2021; Volume 12:1 12, p. 3864. [Google Scholar] [CrossRef] [PubMed]
- OECD. Test No. 310: Ready Biodegradability - CO2 in sealed vessels (Headspace Test). In OECD Guidelines for the Testing of Chemicals, Section 3 [Internet]; OECD Publishing, 2014. [Google Scholar] [CrossRef]
- Goodhead, A.K.; Head, I.M.; Snape, J.R.; Davenport, R.J. Standard inocula preparations reduce the bacterial diversity and reliability of regulatory biodegradation tests. In Environmental Science and Pollution Research; Springer, 2013; Volume 21:16 21, pp. 9511–21. [Google Scholar] [CrossRef] [PubMed]
- Pan, X.; Lin, D.; Zheng, Y.; Zhang, Q.; Yin, Y.; Cai, L.; et al. Biodegradation of DDT by Stenotrophomonas sp. DDT-1: Characterization and genome functional analysis. In Scientific Reports; Nature Publishing Group, 2016; Volume 2016 6:1 6, p. 21332. [Google Scholar] [CrossRef] [PubMed]
- Brauneck, G.; Engel, D.; Grebe, L.A.; Hoffmann, M.; Lichtenberg, P.G.; Neuß, A.; et al. Pitfalls in Early Bioprocess Development Using Shake Flask Cultivations. In Eng Life Sci [Internet]; WEBSITE:WEBSITE:ANALYTICALSCIENCEJOURNALS; John Wiley and Sons Inc: DOI; ISSUE:ISSUE, 2025; Volume 25, p. e70001. [Google Scholar] [CrossRef] [PubMed]
- Nzila, A. Update on the cometabolism of organic pollutants by bacteria. In Environmental Pollution [Internet]; Elsevier, 2013; Volume 178, pp. 474–82. [Google Scholar] [CrossRef] [PubMed]
- Lagiso, T.L.; Woldesemayat, A.A.; Gemta, Z.B. Biodegradation of chlorpyrifos by the newly isolated Escherichia fergusonii and Clostridium bifermentans: Identification and growth optimization for bioremediation. In Scientific Reports; Nature Publishing Group, 2026; Volume 2026 16:1 16, p. 16768. [Google Scholar] [CrossRef] [PubMed]
- Obrador-Viel, T.; Zadjelovic, V.; Nogales, B.; Bosch, R.; Christie-Oleza, J.A. Assessing microbial plastic degradation requires robust methods. In Microb Biotechnol [Internet]; John Wiley and Sons Ltd.: PAGEGROUP; STRING:PUBLICATION, 2024; Volume 17, p. e14457. [Google Scholar] [CrossRef] [PubMed]
- Ganesh Kumar, A.; Sujitha, K.; Sushmita Dubey, D.; Magesh Peter, D.; Dharani, G.; Ramakrishnan, B. Genomic and transcriptomic characterization of genes expressed at 20 MPa by the marine actinobacterium Kocuria flava. Mar. Pollut. Bull. [Internet] Pergamon 2026, 232, 120022. [Google Scholar] [CrossRef] [PubMed]
- Ganesh Kumar, A.; Manisha, D.; Nivedha Rajan, N.; Sujitha, K.; Magesh Peter, D.; Kirubagaran, R.; et al. Biodegradation of phenanthrene by piezotolerant Bacillus subtilis EB1 and genomic insights for bioremediation. Mar. Pollut. Bull. [Internet] Pergamon 2023, 194, 115151. [Google Scholar] [CrossRef] [PubMed]
- Gates, E.G.; Crook, N. The biochemical mechanisms of plastic biodegradation. In FEMS Microbiol Rev [Internet]; Oxford Academic, 2024; p. 48. [Google Scholar] [CrossRef] [PubMed]
- Chafale, A.; Das, S.; Kapley, A. Valorization of oily sludge waste using biosurfactant-producing bacteria. In World Journal of Microbiology and Biotechnology; Springer, 2023; Volume 2023 39:11, p. 39:316. [Google Scholar] [CrossRef] [PubMed]
- Mahjoubi, M.; Cherif, H.; Aliyu, H.; Chouchane, H.; Cappello, S.; Neifar, M.; et al. Brucella pituitosa strain BU72, a new hydrocarbonoclastic bacterium through exopolysaccharide-based surfactant production. In International Microbiology; Springer, 2024; Volume 2024 28:2 28, pp. 299–313. [Google Scholar] [CrossRef] [PubMed]
- Fan, X.; Tang, J.; Nie, L.; Huang, J.; Wang, G. High-quality-draft genome sequence of the heavy metal resistant and exopolysaccharides producing bacterium Mucilaginibacter pedocola TBZ30T. In Standards in Genomic Sciences; BioMed Central, 2018; Volume 13:1, p. 13:34. [Google Scholar] [CrossRef] [PubMed]
- Munir Ahamed, J.; Dahms, H.U.; Huang, Y.L. Heavy metal tolerance, and metal biosorption by exopolysaccharides produced by bacterial strains isolated from marine hydrothermal vents. Chemosphere [Internet] Pergamon 2024, 351, 141170. [Google Scholar] [CrossRef] [PubMed]
- Denef, V.J.; Klappenbach, J.A.; Patrauchan, M.A.; Florizone, C.; Rodrigues, J.L.M.; Tsoi, T. V.; et al. Genetic and genomic insights into the role of benzoate-catabolic pathway redundancy in Burkholderia xenovorans LB400. In Appl Environ Microbiol [Internet]; WEBSITE:WEBSITE:ASMJ;JOURNAL:JOURNAL:AM;ISSUE:ISSUE:DOI; American Society for Microbiology, 2006; Volume 72, pp. 585–95. [Google Scholar] [CrossRef] [PubMed]
- Panter, F.; Bader, C.D.; Müller, R. Synergizing the potential of bacterial genomics and metabolomics to find novel antibiotics. In Chem Sci [Internet]; The Royal Society of Chemistry, 2021; Volume 12, pp. 5994–6010. [Google Scholar] [CrossRef] [PubMed]
- Pacwa-Płociniczak, M.; Daszkowska-Golec, A.; Gobetti, S.; Sinkkonen, A.; Płociniczak, T. Plant and soil transcriptomics reveal the consequences of bioaugmentation of co-contaminated soil with Pseudomonas qingdaonensis ZCR6 during bacteria-assisted phytoremediation. In Applied Microbiology and Biotechnology; Springer, 2026; Volume 2026 110:1, p. 110:100. [Google Scholar] [CrossRef] [PubMed]
- Chlebek, D.; Płociniczak, T.; Gobetti, S.; Kumor, A.; Hupert-Kocurek, K.; Pacwa-Płociniczak, M. Analysis of the genome of the heavy metal resistant and hydrocarbon-degrading rhizospheric pseudomonas qingdaonensis zcr6 strain and assessment of its plant-growth-promoting traits. Int. J. Mol. Sci. [Internet] MDPI 2022, 23, 214. [Google Scholar] [CrossRef]
Figure 1.
Schematic organization of a bacterial BGC. The main BGC may comprise core biosynthetic genes, tailoring or other additional biosynthetic genes, transport genes, self-resistance genes, regulatory genes, and accessory genes, flanked by neighboring genomic regions. In some biosynthetic pathways, satellite genes or distant subclusters located outside the main cluster may also contribute to metabolite biosynthesis. The schematic is illustrative; gene composition, order, orientation, and cluster boundaries may vary among BGCs. NRPS, non-ribosomal peptide synthetase; PKS, polyketide synthase [10,14].
Figure 1.
Schematic organization of a bacterial BGC. The main BGC may comprise core biosynthetic genes, tailoring or other additional biosynthetic genes, transport genes, self-resistance genes, regulatory genes, and accessory genes, flanked by neighboring genomic regions. In some biosynthetic pathways, satellite genes or distant subclusters located outside the main cluster may also contribute to metabolite biosynthesis. The schematic is illustrative; gene composition, order, orientation, and cluster boundaries may vary among BGCs. NRPS, non-ribosomal peptide synthetase; PKS, polyketide synthase [10,14].

Figure 2.
Representative biosynthetic strategies encoded by bacterial BGCs. (A) Non-ribosomal peptide synthetases (NRPSs) use modular assembly lines in which condensation (C), adenylation (A), and peptidyl carrier protein (PCP) domains coordinate amino-acid activation and peptide-chain extension; terminal thioesterase (TE) domains represent a common, but not universal, mechanism of product release. (B) Type I polyketide synthases (PKSs) use ketosynthase (KS), acyltransferase (AT), and acyl carrier protein (ACP) domains, together with optional reductive domains such as ketoreductase (KR), dehydratase (DH), and enoylreductase (ER), to assemble and process polyketide chains from acyl-CoA-derived building blocks. (C) RiPP biosynthesis begins with a ribosomally synthesized precursor peptide followed by enzyme-mediated post-translational modification. (D) Terpene biosynthesis converts isoprenoid precursors into terpene scaffolds that may subsequently undergo tailoring reactions. (E) Hybrid PKS–NRPS systems combine polyketide and non-ribosomal peptide biosynthetic logic within a single pathway. The schemes represent generalized biosynthetic architectures and do not encompass all possible domain organizations or pathway variants. RiPP, ribosomally synthesized and post-translationally modified peptide [14,20,21].
Figure 2.
Representative biosynthetic strategies encoded by bacterial BGCs. (A) Non-ribosomal peptide synthetases (NRPSs) use modular assembly lines in which condensation (C), adenylation (A), and peptidyl carrier protein (PCP) domains coordinate amino-acid activation and peptide-chain extension; terminal thioesterase (TE) domains represent a common, but not universal, mechanism of product release. (B) Type I polyketide synthases (PKSs) use ketosynthase (KS), acyltransferase (AT), and acyl carrier protein (ACP) domains, together with optional reductive domains such as ketoreductase (KR), dehydratase (DH), and enoylreductase (ER), to assemble and process polyketide chains from acyl-CoA-derived building blocks. (C) RiPP biosynthesis begins with a ribosomally synthesized precursor peptide followed by enzyme-mediated post-translational modification. (D) Terpene biosynthesis converts isoprenoid precursors into terpene scaffolds that may subsequently undergo tailoring reactions. (E) Hybrid PKS–NRPS systems combine polyketide and non-ribosomal peptide biosynthetic logic within a single pathway. The schemes represent generalized biosynthetic architectures and do not encompass all possible domain organizations or pathway variants. RiPP, ribosomally synthesized and post-translationally modified peptide [14,20,21].

Figure 3.
Methods for evaluation of antimicrobial activity of novel substances. (A) Description of primary methods for antimicrobial evaluation. (B) Description of additional or specialized approaches for antimicrobial evaluation. (C) Description of a integrated workflow for discovery and characterization of novel antimicrobial substances.
Figure 3.
Methods for evaluation of antimicrobial activity of novel substances. (A) Description of primary methods for antimicrobial evaluation. (B) Description of additional or specialized approaches for antimicrobial evaluation. (C) Description of a integrated workflow for discovery and characterization of novel antimicrobial substances.

Table 1.
Major classes of BGCs and characteristics of their associated natural metabolites.
| Classes | Characteristics of biosynthetic machinery | Main types/ subclasses | Examples of products or applications |
|---|---|---|---|
| PKSs | Polyketide synthase enzymes responsible for the assembly of polyketide chains from precursors derived from acyl-CoA units | Modular type I PKSs; iterative type I PKSs; type II PKSs | Antibiotics, antifungals, immunosuppressants, and other bioactive metabolites |
| NRPSs | Large modular enzyme complexes that incorporate amino acids independently of the ribosome, allowing for broad structural diversity | Modular NRPSs; hybrid NRPSs | Antibiotics, immunosuppressants, siderophores, and other bioactive metabolites |
| RiPPs | Peptides initially synthesized by the ribosome and subsequently modified by specific enzymes | Bacteriocins; lanthipeptides; thiopeptides; lasso peptides, among others | Antimicrobials, cytotoxic compounds, and other bioactive peptides |
| Terpenes | Biosynthesis based on isoprenoid precursors, followed by the formation and modification of terpene skeletons by specific enzymes | Mono-, sesqui-, di-, and triterpene terpenes, among others | Antimicrobial and antioxidant compounds, as well as metabolites with various ecological and biotechnological functions |
| Saccharides | Gene clusters involved in the biosynthesis, modification, and transport of specialized sugars | Polysaccharides and sugars associated with complex metabolites | Glycoconjugates, antibiotics, and other natural products |
| PKSs–NRPSs Hybrids | Integration of PKS and NRPS systems into a single biosynthetic pathway, enabling the formation of complex structures | PKS–NRPS and Related Hybrid Systems | Antibiotics, antifungals, and other bioactive metabolites |
PKSs, polyketides synthases; NRPSs, non-ribosomal peptide synthases, RiPPs, ribosomally synthesized and post-translationally modified peptides.
Table 2.
Examples of some natural products of biotechnological interest produced by the most widely studied bacterial groups that produce BGCs.
Table 2.
Examples of some natural products of biotechnological interest produced by the most widely studied bacterial groups that produce BGCs.
| Bacterial genus | Representative producer species | Predominant niches | BGC classes | Representative natural products | Main applications |
|---|---|---|---|---|---|
| Acinetobacter | A. baumannii | ubiquitous | NRPS | acinetobactin | siderophore |
|
A. endophyla A. pittii |
other | fengycin | antifungal | ||
| A. gyllenbergii | RiPPs | acinetodin | antibiotic | ||
| Bacillus | B. cereus | soil, food, industrial environments | NRPS, PKS, NRPS–PKS, RiPP | bacillibactin, cereulide, lipopeptide, petrobactin, thumolycin, zwittermicin | antibacterial, antifungal, siderophore |
| B. subtilis | soil, rhizosphere, plants, environments associated with organic matter | NRPS, PKS, NRPS–PKS, RiPP |
bacilysin, bacillibactin, bacillaene, difficidin, fengycin, subtilosin, surfactin | biocontrol, antimicrobials, agriculture, biotechnology | |
| B. thuringiensis | soil, plants, and environments associated with insects | NRPS, NRPS–PKS |
bacillibactin, thumolycin, zwittermicin | antibacterial, antifungal | |
| Brevibacillus | Brevibacillus spp. | soil, water, plants, insects, and animals | NRPS, PKS, NRPS–PKS, |
edeine, gramicidin, petrobactin, tyrocidine | antibacterials, biocontrol, bioprospecting, siderophore |
| Burkholderia | B. ambifaria | soil, water, plants, animals, fungi and clinical environments | other | cepaciachelin | siderophore |
| B. gladioli | PKSs | gladiofungin | antifungal | ||
| B. thailandensis | RiPPs | capistruin | antibiotic | ||
| Paenibacillus | Paenibacillus spp. | soil, rhizosphere, plants, environments associated with insects | NRPS, PKS, NRPS–PKS, lanthipeptides, terpenes | fusaricidins, polymyxins, paenibacillin, paenilan, paenicidin A, tridecaptin | antibiotics, antifungals, biocontrol, agriculture |
| Priestia | P. megaterium | soil, rhizosphere, plants | NRPS/NRPS-like, phosphonates, terpenes, ranthipeptides | carotenoids, synechobactins, schizokinens | biotechnology, plant growth promotion, siderophores, bioprospecting |
| Pseudomonas | P. fluorescens | ubiquitous | NRPSs | gacamide A | antibiotic |
|
P. kilonensis P. brassicacearum P. thivervalensis |
T3PKSs | DAPG | biocontrol | ||
| P. protegens | other | pyrroniltrin | antifungic | ||
| Pseudomonas spp. | NRPSs | pyoverdine | siderophore | ||
| Pseudomonas spp. | other | lankacidin A | antitumor | ||
| Ralstonia | R. solanacearum | ubiquitous | other | ralstonin | phytotoxic and chlamydospore-inducing |
| NRPSs-PKSs | ralsolamycin | antibiotic | |||
| Salinispora | S. arenicola | seawater | NRPSs | retimycin A | antibiotic/antitumor |
| S. pacifica | salinichelin | siderophore | |||
| Stenotrophomonas | S. maltophilia | ubiquitous | other | maltophilin | antifungal |
| Stenotrophomononas sp. | other | xanthobaccins A–C | antifungal | ||
| Streptomyces | S. clavuligerus | soil | other | clavulanic acid | antibiotic |
|
S. peucetius S. galilaeus S. nogalater |
PKSs | anthracyclines | antitumor | ||
| S. coelicolor | NRPSs | coelichelin | siderophore |
DAPG, 2,4-diacetilfloroglucinol; NRPSs, non-ribosomal peptide synthases, PKSs, polyketides synthases; RiPPs, post-translationally modified peptides; T2PKSs, type II polyketide synthases; T3PKSs, type III polyketide synthases.
Table 3.
Comparative summary of genomic mining tools for BGCs.
| Categories | Tools | Main algorithm | Advantages | Limitations |
|---|---|---|---|---|
| Rule-based | antiSMASH 8.0 (2025) | Trained pHMMs and manual rules for detecting biosynthetic signatures. A greedy approach to physical boundary determination. | Detects 101 types of BGCs; theoretical prediction of core chemical structure (SMILES) for NRPS/PKS; enhanced analysis of terpenes and modifying enzymes (tailoring) with dedicated tabs. | Unable to detect entirely novel or atypical BGC classes (not described in the rules); greedy boundary delineation may merge adjacent BGCs or include irrelevant genes, and removal of the internal fungal call. |
| BAGELS (2024/2026) | Search for Pfam motifs using pHMM and similarity against a core peptide database (via DIAMOND) across 6 reading frames, regardless of the initial ORFs calling. | Specializes in RiPPs and bacteriocins; high-speed processing; integrates promoter/terminator predictions and RNA-Seq expression data. Supports metagenomic data and offers high throughput (up to 100,000 files). | Strict focus on bacteriocins and RiPPs. Does not analyze PKS/NRPS. Highly dependent on homology in the database to identify new core peptides. Chemical bond representations in the alignments are merely indicative (not chemically verified). Higher false positive rate in “discovery” mode. | |
| Machine learning and AI | DeepBGC (e-DeepBGC) | Deep learning using BiLSTM trained on the Pfam2vec vector representation of Pfam domains. | Rule-free detection of BGCs from known BGCs and previously unseen classes (leave-class-out). Classification of chemical class and biological activity using Random Forest. e-DeepBGC incorporates data augmentation and clans. | Strong training bias toward thoroughly studied taxa, such as Streptomyces. Significantly lower performance on fungal genomes compared to rule-based tools. High false-positive rate if not post-processed. Does not predict fine stereochemistry or generate SMILES. |
| GECCO v.0.9.10 (2024) | A probabilistic statistical CRF model that analyzes the context and order of Pfam domains and GO terms. | Faster than DeepBGC. Requires less training data. Higher accuracy in predicting physical boundaries (boundary detection). Highly interpretable and auditable through the CRF domain weights. Integrated with antiSMASH. | Focused on and trained for prokaryotes (bacteria), with low effectiveness in complex fungi (introns and non-operon BGCs). Lacks documented practical functional validation of new BGCs identified by the tool. Does not provide predictions of stereochemistry or monomers. | |
| Evolutionary Mining | EvoMining 2.0 (2019) | Phylogenomic analysis of EF expansions and the recruitment of copies from central/primary metabolism to specialized secondary pathways. | Identifies atypical or non-canonical BGCs that defy homology rules. Enables tracking of biosynthetic evolution and customization of genome/enzyme seed libraries. | Entirely dependent on the MIBiG database to detect recruitment. Prone to false positives; some primary metabolism expansions are for redundancy rather than secondary pathways. Requires complex manual curation of seed families. Does not delineate boundaries or generate chemical predictions? |
| ARTS 2.0 (2020) | Search for self-protection factors by combining physical co-localization (via antiSMASH), duplication of essential (housekeeping) genes, and HGT phylogeny. | Highly effective in prioritizing antibiotic-producing BGCs with novel modes of action/therapeutic targets. Version 2.0 adds support for metagenomes and global taxonomy. Integrated with BiG-SCAPE. | Completely dependent on antiSMASH, where initial detection errors propagate. Unable to distinguish whether a duplicated homolog is involved in resistance or in the molecule’s own biosynthesis. Limited phylogenetic signal outside of actinobacteria. | |
| Clustering and networking | BiG-SCAPE 2.0 (2026) | Pairwise alignment of Pfam domains and similarity calculation based on three metrics (Pfam content, synteny, and identity) using AP. | Gold standard in accuracy. Version 2.0 focuses on “protoclusters,” eliminating noise from greedy edges. Core-guided alignment. AP adjusted for subgraph density. | It does not perform primary BGC detection; it relies on external annotated GenBank files. The calculation of the paired similarity matrix (all-vs-all) is computationally expensive for databases of global scale. |
| BiG-SLICE 2.0 (2026) | Vectorization of BGCs based on Pfam biosynthetic domain counts and superlinear clustering using the BIRCH algorithm. | Hyper-scalability capable of clustering millions of BGCs in a few hours. Version 2.0 introduces a cosine-based distance with L2 normalization that corrects the bias toward short BGCs, such as RiPPs. It feeds into the global BiG-FAM repository. | It does not perform primary detection. It disregards synteny and fine-tuning, resulting in lower biological accuracy for grouping divergent taxa when compared to BiG-SCAPE. |
AP, Affinity Propagation; BGCs, biosynthetic gene clusters; BiLSTM, bidirectional recurrent neural networks; CRF, Conditional Random Fields; EF, enzyme Family; GO, Gene Ontology; HGT; horizontal gene transfer; NRPS, non-ribosomal peptide synthases; ORFs, open read frames; Pfam, protein families; pHMM, profile hidden Markov models; PKS, polyketide synthases; RiPPs, post-translationally modified peptides; RNA, ribonucleic acid.
Table 4.
Minimum experimental evidence supporting genome-guided bioremediation claims.
| Claim/mechanism | Minimum experimental design and controls | Evidence that strengthens the claim |
|---|---|---|
| Direct metabolism of an organic xenobiotic | Use a washed inoculum at a defined physiological state and standardized biomass in mineral medium containing the contaminant as the sole organic carbon and energy source; include uninoculated and killed-biomass controls and sample over time. | Demonstrate growth together with depletion of the parent contaminant and formation of transformation products by GC-MS, HPLC, or LC-MS; add evidence of mineralization when complete biodegradation is claimed (OECD, 2014; Pan et al., 2016). |
| Cometabolic transformation | Compare the contaminant alone with contaminant plus a defined cosubstrate while keeping inoculum, nutrients, pH, and aeration equivalent; include controls for the cosubstrate itself. | Chemical transformation and metabolite formation should exceed the corresponding controls. Additional growth caused only by the cosubstrate is not evidence of enhanced contaminant transformation (Nzila, 2013; Lagiso et al., 2026). |
| Biosurfactant-mediated remediation | Screen for surface activity, chemically characterize the product, and then test whether it changes solubilization, bioavailability, or removal of the target contaminant. | Drop-collapse, oil-spreading, and emulsification assays are screening tools. The claim is strengthened only when the characterized product measurably changes contaminant fate (Ayangbenro and Babalola, 2020; Chafale et al., 2023; Mahjoubi et al., 2025). |
| Metal removal or immobilization | Expose standardized biomass or purified extracellular products to a defined metal concentration and evaluate pH and contact time to the proposed mechanism; distinguish tolerance from removal. | Quantify the metal before/after treatment and, when possible, determine partitioning or speciation by ICP-OES, ICP-MS, or equivalent. MIC alone demonstrates tolerance, not remediation (Ahmed and Holmström, 2014; Fan et al., 2018; Ahamed et al., 2024). |
| Plastic biodegradation | Use a defined medium, standardized polymer mass or surface area, and prolonged monitoring, with controls for abiotic changes, additives, and pretreatment. | Combine polymer-level evidence, such as FTIR or GPC, with detection of monomers or oligomers and, preferably, carbon dioxide evolution or isotope-based carbon tracing. Biofilm or mass loss alone is insufficient (Obrador-Viel et al., 2024; Gates and Crook, 2024). |
| Causal validation of a gene or BGC | Where feasible, disrupt the candidate gene or BGC and perform complementation, or express the locus in a suitable heterologous host; include matched wild-type and vector controls. | Loss of the phenotype after disruption and its recovery after complementation, or production of the expected metabolite or function in a heterologous host, provides stronger causal evidence (Denef et al., 2006; Panter et al., 2021). |
| Environmental-matrix performance | Progress from pure culture to a microcosm or mesocosm using the soil, sediment, water, or effluent of interest while preserving chemical endpoints and appropriate matrix controls. | Assess contaminant fate together with persistence or activity of the inoculum and, when relevant, the response of the resident community. Strong in vitro performance may not predict matrix-level efficacy (Chlebek et al., 2022; Pacwa-Płociniczak et al., 2026). |
GC-MS, Gas chromatography - mass spectrometry; HPLC, High-performance liquid chromatography; LC-MS, Liquid chromatography-mass spectrometry; ICP-OES, Inductively Coupled Plasma Optical Emission Spectrometry; ICP-MS, Inductively Coupled Plasma-Mass Spectrometry; MIC, Minimum inhibitory concentration; FT-IR, Fourier-transform infrared spectroscopy; GPC, Gel permeation chromatography; BGC, Biosynthetic gene clusters.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.