Preprint
Article

This version is not peer-reviewed.

The Panama Endemic Cluster A3 of Mycobacterium tuberculosis L2.2.M3 Sublineage Forms a Separate Monophyletic Cluster Within the Global Diversity

Submitted:

01 September 2026

Posted:

02 September 2026

You are already at the latest version

Abstract
Drug-sensitive Mycobacterium tuberculosis L2 strains remain highly transmissible and adapted to specific populations according to their genetic background. It is currently unknown how the L2 strains from the Province of Colón, Panama, relate to global diversity, and which genetic variations favor this strain's higher transmissibility. This study aims to compare the Mycobacterium tuberculosis L2 strains collected in Colón, Panama, with a global dataset and to describe the impact of the interested SNPs on protein function. We analyzed whole-genome sequencing (WGS) data from L2.2.M3 strains collected between 2015 and 2023 in Colón, Panama, and compared them with the entire global strains database (identified by the specific mutation at genome position 1219683 G>A) available at the NCBI Sequence Read Archive. After comparing with 1,578 L2.2.M3 global strains, the 91 isolates from Panama comprise a separate monophyletic cluster within the sublineage, which we have termed the Panama Endemic Cluster A3 (PECA3) branch. This PECA3 branch contained only three additional strains identified from non-Asian countries. Only 4.3% (4/94) of these strains had SNPs conferring drug resistance. The branches containing the closest relatives were defined as CR1, CR2, and CR3. These branches had higher proportions of drug-resistant strains – 59.3% (16/27), 42.4% (28/66), and 76.9% (10/13), respectively –, with most strains from, but not limited to, East and Southeast Asian countries. Complete genome analysis of 200 isolates from PECA3, CR1, CR2, and CR3 revealed that these four branches belong to a larger cluster characterized by 38 single-nucleotide polymorphisms (SNPs), of which only 4 SNPs are shared among the four branches. PECA3 remains mostly pan-susceptible and genetically homogeneous, defined by 18 specific SNPs, of which 9 are nonsynonymous with moderate impact, and one causes a stop codon with high impact. Our results confirm that PECA3 presents a unique branch within the global diversity of L2.2.M3 sublineage. PECA3 is mainly pan-susceptible, with limited transmission to or from other countries, and appears to be endemic to Panama. Genomic surveillance and additional analyses are warranted to monitor de novo resistance, transmission, dispersal, the evolution of virulence, and host adaptation genes in PECA3.
Keywords: 
;  ;  ;  ;  

Impact Statement

Mycobacterium tuberculosis Lineage 2 (L2), also known as the East Asian lineage and/or the Beijing genotype, has spread worldwide. These genotypes are often associated with increased drug resistance and rapid transmission. However, drug-susceptible strains can also evolve to become highly transmissible. In this work, we compare whole-genome sequencing (WGS) data from the L2.2.M3 sublineage strains collected in the province of Colón, Panama, together with publicly available L2.2.M3 strains worldwide. In silico analyses confirm that these strains, collected from TB patients between 2015 and 2023, are genetically homogeneous and mostly pan-susceptible. Thus, we have termed it Panama Endemic Cluster A3 (PECA3). PECA3 strains harbor 9 nonsynonymous single-nucleotide polymorphisms (SNPs) with moderate impact at genes hspR, Rv0430, Rv0822c, Rv0945, Rv2042c, Rv2509, Rv3729, scoA, and speE; and 1 SNP causing a stop codon with high impact at the gene Rv1453. Hence, this research provides insights into genomic changes in PECA3 that may explain the strains’ survival without acquiring drug resistance.

Introduction

Whole-genome sequencing (WGS) has become a key tool for studying the population structure and evolution of Mycobacterium tuberculosis (M. tuberculosis). Progress in genomics of this pathogen has been driven not only by the increasing number of sequenced isolates but also by the development of analytical platforms capable of handling large-scale genomic datasets. Recent approaches using data mining in publicly available M. tuberculosis genome sequences have shed light on new genetic variations, such as the identification of the new lineage 10 [1], exploring poorly defined sublineages within lineages 4, 6, and 7 [2], and improving the description of lineage 2 [3,4].
Lineage 2 (L2), also known as the East Asian lineage, has attracted attention due to the global distribution of the Beijing genotype, which is frequently associated with drug resistance and outbreaks. L2 genome-wide single-nucleotide polymorphism (SNP) genotyping has been subdivided into over 40 sublineages [4,5]. With this classification, known multidrug-resistant (MDR) strains such as Central Asia Outbreak (CAO, or L2.2M4.9.1) and B0/W148 (L2.2.M4.5) have sufficient markers to classify them into sublineages and relate them to their geographical distribution [3,5]. This allows differentiation from other circulating strains that have limited geographical spread yet cause sustained local endemic transmission.
In 2024, Panama registered a total of 2,210 tuberculosis cases, with an incidence rate of 48.96 per 100,000 inhabitants, which is a higher incidence than the 42.2 reported in 2015 [6,7]. Although the province of Colón accounts for approximately 10% of the total TB cases nationwide, this region has consistently reported more than 30% of its cases belonging to a distinct drug-susceptible cluster of M. tuberculosis L2.2.M3 (formerly classified as Asian African 3/Bmyc13/L2.2.5) since 2015. The cluster was provisionally named “Strain A” and characterized by a SNP at position 518748 G>A [8,9,10], now formally designated as Panama Endemic Cluster A3 (PECA3). The proportion of PECA3 strains reported in Colón is unusually high, compared to less than 3.5% of the reported TB cases in Panama’s western region [11]. In the Americas, only Cuba (17.2%), Peru (16%), and Colombia (5%) have reported relatively high proportions of L2 strains [12]. Previous L2 studies [5,10] have been regionally focused. Hence, this work aims to describe PECA3 within the L2.2.M3 sublineage in a global context using a thorough database, state-of-the-art tools, and an expert-based classification approach.

Results

Phylogenetic Analysis, Drug Resistance Profile, and Geographical Distribution

From the 120,000 publicly available M. tuberculosis WGS datasets (at the time of study, May 2024) in NCBI, TB-Annotator identified 1,578 M. tuberculosis L2.2.M3 genomes, which we depicted in an unrooted phylogenetic tree (Figure 1). A total of 94 PECA3 genomes are clustered in a single branch, while 106 genomes representing the closest relatives were grouped into three distinct branches. Although the analysis with the Tuberculosis Lineage Genotyping (TbLG) tool returned only 1,476 M. tuberculosis L2.2.M3 genomes due to pipeline technical differences, the number of isolates for PECA3 and close relatives branches was validated. The maximum-likelihood phylogenetic tree (Supplementary Figure 1A) revealed that the L2.2.M3 genomes form a star-like phylogeny, with the PECA3 and its close relative branches standing out at the end of the tree. The genome drug resistance profile and branch and country of isolation have been summarized in Table 1. General information for each genome, along with the detailed drug resistance profile, is also provided in Supplementary Table 1.
Figure 1A shows a tree with PECA3 strains clustered in a single branch, comprising 94 genomes with a mean internal distance of 35 ± 17 SNPs, of which 96.8% (91/94) were from Panama. The three other genomes were collected in the United States of America (1.1%, 1/91), Portugal (1.1%, 1/91), and Aruba (1.1%, 1/91). Hence, 98.9% (93/94) of the genomes were in the Americas, while 1.1% (1/94) were from Europe. TB-profiler analysis confirmed that 95.7% (90/94) of the samples belonging to the PECA3 branch are sensitive, whereas only 4.3% (4/94) of genomes showed drug resistance, and were from Panama. The drug resistance profiles of three isolates matched those previously reported (rifampicin, rifampicin/pyrazinamide, and rifampicin/isoniazid/ethionamide) [9], confirming that these genomes do not contain SNPs added in the second edition of the Catalog of mutations published by the World Health Organization (WHO) [13]. In contrast, sample ERR4403842 is now classified as RR-TB due to a p.Leu449Gln mutation in the rpoB gene; the mutation was not previously detected, as it was analyzed using a previous database (TB-Profiler v. 4.4.2 with the tbdb_c2fb9a2 database), based on the catalog’s first edition [14], which used a non-standard notation [15,16]. It’s also worth noting that the catalog warns of caution regarding mutations at this position identified by Illumina sequencing [13]. Given that this sequence has an average coverage of 53X and a variant allele frequency of 13%, our manual review of the reads revealed that it lacked the classical signals of sequencing artifacts. As a result, this sample is considered RR-DR in this study.
Finally, although the isolate ERR2517612 is reported as sourced from the Netherlands tuberculosis reference laboratory (sample name NL567 in the supplementary material) in a study [17], a representative of the Dutch National Institute for Public Health and the Environment (RIVM) confirmed that this sample was isolated in Aruba in 2016, remains unique in their local WGS database and does not cluster with any Dutch (EU) isolate (Richard M. Anthony, PhD; email, February 9th, 2026), which corroborates the findings of a maximum-likelihood tree of the PECA3 branch which shows that the Panama genomes were collected first (Supplementary Figure 1B). Aruba is a constituent country of the Kingdom of the Netherlands; hence, we classified this isolate as part of the Americas.
The closest branch, defined as Closest Relative 1 (CR1), contains 27 isolates with a mean internal distance of 77 ± 30 SNPs (Figure 1B). These genomes were collected in seven countries: Australia, Canada, China, Madagascar, Myanmar, Thailand, and Vietnam. Hence, 3.7% (1/27) of the genomes originated in Oceania, 11.1% (3/27) in the Americas, 3.7% (1/27) in Africa, and 70.4% (19/27) in East and Southeast Asia. Only 11.1% (3/27) of the cases have an unknown origin. TB-profiler analysis revealed that 40.7% (11/27) of the genomes are susceptible, while 59.3% (16/27) are drug-resistant.
The second-closest branch, defined as Closest Relative 2 (CR2), contains 66 isolates (Figure 1C). The genomes were collected in 10 countries: Australia, Cambodia, Canada, China, Guatemala, India, Myanmar, Thailand, the United States of America, and Vietnam. Hence, 12.1% (8/66) of the genomes originated in Oceania, 21.2% (14/66) in the Americas, 1.5% (1/66) in South Asia, and 36.4% (24/66) in East and Southeast Asia. In this branch, the percentage of cases of an unknown origin reached 28.8% (19/66). TB-profiler analysis revealed that 57.6% (38/66) of the genomes are sensitive, while 42.4% (28/66) are drug resistant.
Finally, the Closest Relative 3 (CR3) branch contains 13 isolates (Figure 1D), collected in four countries: Canada, China, the United States, and Vietnam. Hence, 15.4% (2/13) of the genomes were in the Americas, while 84.6% (11/13) were from East and Southeast Asia. Of these, 23.1% (3/13) of the genomes are susceptible (collected in Canada, the United States, and Vietnam), while 76.9% (10/13) are drug-resistant, all collected in China.

SNP Analysis and Impact on Amino Acids

The above comparative analysis of the 200 isolates from PECA3, CR1, CR2, and CR3 identified 37 SNPs. In addition, the PECA3-defining SNP (518748 G>A) remains unique to this group. Hence, the PECA3, CR1, CR2, and CR3 branches belong to a larger cluster characterized by 38 SNPs in total (Table 2). When analyzing SNP variation between these branches, we found that PECA3 is genetically homogeneous and defined by 18 specific SNPs. In addition, 11 SNPs are shared by CR1 and PECA3; 5 SNPs are common to CR1, CR2, and PECA3; and 4 SNPs are shared across all four groups (CR1-3 and PECA3). This demonstrates that CR3 is quite distant from the strains of interest, and the most distant ancestor can be traced at the time of this study.
A closer look at the PECA3-specific SNPs also revealed that SnpEff [18] predicted 9 SNPs causing nonsynonymous mutations with moderate impact at genes hspR, Rv0430, Rv0822c, Rv0945, Rv2042c, Rv2509, Rv3729, scoA, and speE; and 1 SNP causing a stop codon with high impact at the gene Rv1453. Results from the in silico analysis using SIFT and PAM1-based methods, along with information on mutant genes retrieved from the Mycobrowser portal and articles, are available in Supplementary Table 3. A protein-protein network analysis of these 9 genes revealed that they are not connected (Supplementary Figure 2). In contrast, a network analysis of gene Rv1453 (Figure 2) revealed that it is interconnected via co-expression links encoding transcriptional regulatory proteins.

Discussion

M. tuberculosis Lineage 2 (L2), especially the Beijing genotype, is globally distributed and often associated with increased drug resistance and rapid transmission, partly due to epistatic interactions in drug-resistant isolates and the acquisition of compensatory mutations that restore fitness [19,20]. Thawornwattana et al. (2021) analyzed 4,425 L2 isolates, primarily from East Asia, Southeast Asia, and South Asia, where L2 is endemic. For the L2.2.M3 sublineage, they found that the genomes grouped into three major clades in East and Southeast Asia, each associated with specific countries: China, Vietnam, and Thailand. The Thailand’s samples formed a large cluster, likely due to a recent MDR outbreak [5]. In our previous work, we found that L2.2.M3 is present in at least 5 countries in the Americas (Canada, Colombia, Guatemala, Peru, and Panama), and although these samples were from the same sublineage, many were distantly related [10]. In the present study, our findings (Table 1) confirm that the L2.2.M3 sublineage is not geographically restricted. Among the four branches studied here, 56.0% (112/200) of the genomes were from the Americas, 27.0% (54/200) were from East and Southeast Asian countries (China, Thailand, and Vietnam), 11% (22/200) from unknown origin, 4.5% (9/200) from Oceania (Australia), 0.5% (1/100) from Europe (Portugal), 0.5% from Africa (Madagascar) and 0.5% from South Asia (India). At a branch level, all PECA3 genomes originated in the Americas and Europe. CR1-3 branches, although they have more genomes from East and Southeast Asian countries (China, Thailand, and Vietnam), they also have genomes from different regions: the Americas (Guatemala, the United States of America, and Canada), Africa (Madagascar), Oceania (Australia), and South Asia (India). Hence, a broader study of this sublineage could reveal other drug-resistant or drug-susceptible clusters and their geographic relationships.
In this work, we describe PECA3 as a unique pan-susceptible branch within the L2.2.M3 sublineage with limited transboundary spread, with only 3 cases detected outside Panama: the Aruba (1), Portugal (1), and the USA (1). There are also known examples of drug-susceptible clusters with increasing circulation. For example, in Cape Town, South Africa, a study revealed that over a 12-year period, the incidence of a drug-susceptible Beijing L2 clade increased exponentially [21]. In this regard, studies exploring the relationship between Beijing genotypes and drug resistance have produced mixed results. One study found that the Beijing genotype could be grouped in 4 patterns: 1) endemic, not associated with drug resistance; 2) epidemic, associated with drug resistance; 3) epidemic but drug sensitive; and 4) very low level or absent [22]. Interestingly, the first pattern (endemic, susceptible) has consistently been reported in many parts of China over the years [22,23,24,25,26,27] while the second pattern (epidemic, drug-resistant) was observed more often out of East Asia (high level in Cuba, the former Soviet Union, Vietnam, and South Africa, lower level in parts of Western Europe). Hence, although the Beijing genotype has a selective advantage (higher virulence or transmissibility) over other genotypes, they are not more likely to acquire drug resistance than non-Beijing genotypes. Thus, the spread of drug-resistant strains is influenced by external factors specific to geographic settings, such as deteriorating tuberculosis control systems, host populations, socioeconomic, and nutritional status [24,25,27,28].
The in silico analysis of the 9 missense mutations and 1 stop codon mutation provides some insights into how PECA3 has restored its fitness without acquiring drug resistance. On the one hand, regarding nonsynonymous SNPs in genes hspR, Rv0430, Rv0822c, Rv0945, Rv2042c, Rv2509, Rv3729, scoA, and speE and the resulting amino acid changes, most appeared to have a significant impact on the proteins assessed. A detailed discussion of each PECA3-specific mutation and its impact on the gene is beyond the scope of this article. However, we note that two of these genes are defined as “essential genes” in Mycobrowser: Rv0430 and Rv2509. In a recent study by Datta et al. (2019), the protein Rv0430, encoded by the essential gene Rv0430, was reported to exhibit nucleoid-associated protein (NAP) features [29]. Thus, the authors have named Rv0430 as NapA (Nucleoid-associated protein with autoregulatory activity). NapA protects DNA from damaging agents and is the first gene product of a multi-gene operon that also contains virR and sodC genes, which are two regulators of M. tuberculosis virulence [29]. Hence, Rv0430 is thought to be a transcriptional and virulence regulator [30]. Given that the nonsynonymous mutation at gene Rv0430 is significant (P=0.01), it may play a role in PECA’s virulence. The second essential gene, Rv2509, encodes mycolic acid reductase and is required for growth in slow-growing mycobacteria [31]. Given this strain’s current success, these mutations are unlikely to be detrimental. In addition, at least four other mutations were described to impact protein structure and function: hspR (transcriptional regulation, virulence), Rv0822c (putative membrane proteins and membrane-bound regulatory proteins), Rv2042c, and scoA (fatty acid degradation/synthesis, lipid metabolism). Perhaps mutations in Rv3729 (cellular metabolism) and speE are also structurally important (https://mycobrowser.epfl.ch/ and https://sift.bii.a-star.edu.sg/www/SIFT_seq_submit2.html, accessed December 2025). Taking this information into account, it is noteworthy that these 9 proteins are not connected in the network, suggesting that they are involved in distinct pathways that do not act synergistically.
On the other hand, the SNP at position 1638962 C>A introduces a premature stop-codon in gene Rv1453, producing a 2/3 abridged protein Rv1453. Firstly, it has been reported that mutations at gene Rv1453 are associated with resistance to clofazimine (CFZ) [32] and mutations have arisen even when clofazimine was not part of the treatment [33]. In the study by Li et al. (2021), the Rv1453 knockout strain had an increased transcriptional level of the genes Rv1455, which encodes a protein of unknown function, and qor, which encodes quinone reductase [32]. However, in our study, the stop-codon mutation makes this gene inactive and 95.74% (90/94) of PECA3 genomes are pansusceptible. Although it is unclear how this mutation increases PECA3 fitness, we must remain vigilant in case a drug-resistant PECA3 case is reported. Secondly, Rv1453 is a regulatory protein and is part of the broad network of transcriptional regulatory proteins (Figure 2). Bacterial regulatory genes connect in networks to enable rapid, coordinated responses to complex environmental changes. By integrating multiple signals (nutrients, stress, toxins), they control entire sets of functionally related genes (operons) efficiently. Like that, they ensure survival by turning on needed pathways and shutting down others as a single unit, not one gene at a time. These networks allow bacteria to act as multiple signal processors, fine-tuning gene expression for adaptation [34]. Overall, further experimental studies are required to determine how these specific SNPs contribute to the high transmissibility and/or pathogenicity of the PECA3 strains in the province of Colón, Panama.
Besides bacterial genetics, other host factors – such as population genetics, host-pathogen interactions, or community settings that may act as sources of infection and transmission – may contribute to the persistence of this PECA3 branch. According to the Global Tuberculosis Report, undernutrition, alcohol use disorders, smoking, HIV infection, and diabetes are the main TB risk factors [35]. However, in recent years, prisons have been recognized as high-risk settings due to the convergence of factors: overcrowding, individual-level risk factors, prolonged exposure, and inadequate access to health-care services [36]. In the Latin America region, people deprived of liberty (PDL) experience an average tuberculosis incidence rate that is 26 times higher than that of the general population [37]. Colón City prison currently reports an overcrowding of 146% and frequent reports of TB cases [10]. A plausible hypothesis is that the Colón City prison could represent ideal conditions for TB transmission, like recent findings in Paraguay, where the transmission of tuberculosis was traced within and toward the surrounding community [38]. Therefore, further studies are warranted to determine whether the Colón City prison serves as a setting that contributes to the transmission of this unique PECA3 epidemic branch.

Material and Methods

Whole-Genome Sequencing Data

To investigate the position of PECA3 within the global diversity of the L2.2.M3 sublineage, we first identified the Sequence Read Archive (SRA) accession numbers of the genomes carrying the PECA3-defining SNP (518748 G>A) that were deposited in the BioProjects PRJNA1079299, PRJEB39699, PRJEB23681, and PRJEB29408. These genomes are the result of previous studies conducted between 2015 and 2023 in the Province of Colón, Panama [8,9,10,39].

Genome Analysis

We used the TB-Annotator pipeline for global genome analysis because it integrated more than 120,000 publicly available M. tuberculosis WGS datasets (at the time of study, May 2024) from the National Center for Biotechnology Information (NCBI) and their known genomic characteristics. Briefly, we provided TB-Annotator with the SRA accession numbers of interest and M. tuberculosis L2.2.M3 genomes (1219683 G>A) available at the NCBI Sequence Read Archive (Supplementary Table 2). FASTQ files were downloaded and processed using default parameters. Processed reads were mapped, and variant calling was performed using the reference genome M. tuberculosis H37Rv (NC_000962.3). Detailed information regarding its architecture and processes has been described by Senelle, et al. (2023) [40]. The resulting SNP list was used to construct an alignment, which was subsequently used to create a phylogenetic tree with RAxML-NJ [41].
To analyze the nearest populations from which the PECA3 strains originated, the same SRA accession numbers (Supplementary Table 2) were retrieved and analyzed with the TBvar v1.1.5 workflow (https://github.com/dbespiatykh/TBvar). In brief, reads were processed, mapped, and variant calling was performed using the reference genome M. tuberculosis H37Rv (NC_000962.3). Variant effects were annotated with SnpEff v5.1d [18], and genomes were genotyped with Tuberculosis Lineage Genotyping (TbLG) tool v0.1.5 (https://github.com/dbespiatykh/tblg). Both tools were used with default parameters, and information about them has been previously published by Shitikov and Bespiatykh (2023) [4]. The resulting SNP list was manually curated as follows: polymorphisms present in all PECA3 genomes relative to M. tuberculosis H37Rv (NC_000962.3) were removed, rows were sorted according to the nucleotide position where the substitution occurs, masking of PE-PPE genes and intergenic regions upstream of them, and masking of false SNPs (either a mixture or low coverage). The resulting SNP list was used to construct an alignment, which was subsequently used to create a maximum-likelihood phylogenetic tree using IQ-TREE 2 [42,43].

Drug Resistance Analysis

After assessing the results, the genomes belonging to the branches of interest were analyzed with TB-Profiler v. 6.6.5 [44,45] using database 0a909094 to determine their drug resistance profiles [2]. For each SRA, the country of isolation was determined from the metadata provided for each genome in NCBI and the supplementary material of published research articles that declared the use of the BioProjects.

Amino Acid Change Analysis

To determine the significance of SNPs causing nonsynonymous mutations – i.e., their potential role in altering protein function –, the PECA3-specific nonsynonymous SNPs were assessed in silico using Point Accepted Mutation 1 (PAM1, or 1 accepted point mutation per 100 residues) and SIFT (https://sift.bii.a-star.edu.sg/www/SIFT_seq_submit2.html, accessed December 2025) [46] with UniProtKB/Swiss-Prot + TrEMBL databases [47]. Information on M. tuberculosis genes was obtained from the Mycobrowser portal (https://mycobrowser.epfl.ch/, accessed December 2025) [48] and supplemented with a PubMed search (https://pubmed.ncbi.nlm.nih.gov/, accessed December 2025) [49] using “tuberculosis” and gene names as keywords. Finally, a gene-gene and protein interaction network was constructed using the Search Tool for the Retrieval of Interacting Genes (STRING v.12.0, https://string-db.org/, accessed December 2025) [50].

Conclusion

Our study confirms that the PECA3 from the province of Colón, Panama, presents a unique endemic branch within the global diversity of L2.2.M3 sublineage. PECA3 is mostly pan-susceptible, with limited transmission to or from other countries. Current efforts to retrospectively and prospectively sequence M. tuberculosis L2 isolates will aid in several areas, but not limited to 1) achieving a robust global phylogenetic and evolutionary analysis; 2) tracking the transmission of PECA3 to understand the nature of its high transmissibility; and 3) developing effective strategies to identify community settings that may act as sources of infection and break transmission chains.

Limitations and Future Directions

All the genomes included in this study were collected in Colón, one of the provinces in Panama, where M. tuberculosis strains have been systematically collected and cultivated for whole-genome sequencing studies since 2015. In addition, we only evaluated SNPs and their protein impact, as the genomes were sequenced using short-read technology. An evaluation of other genomic features, such as indels and changes in hard-to-sequence regions, such as the PE-PPE genes [51] could shed more light on how PECA3 has evolved to adapt to Colón’s Afrodescendant setting.

Supplementary Materials

The following supporting information can be downloaded at the website of this paper posted on Preprints.org, Supplementary: Figure 1. Global Mycobacterium maximum likelihood tree with 1,476 L2.2.M3 strains (PDF file). Supplementary: Figure 2. Network analysis of SNPs causing nonsynonymous mutations in PECA3 (PDF file). Supplementary: Table 1. General information and resistance profile of the 200 genomes belonging to the branches of interest (Excel file). Supplementary: Table 2. List of SRA accession numbers used for genome analysis (CSV file). Supplementary Table 3. PECA3-specific nonsynonymous mutations: possible impact on altering protein structure and function (PDF file).

Author Contributions

Conceptualization: J.E.K, E.S., I.M., C.S., A.G. Resources: J.E.K., F.A., E.S., P.P, D.S., D.C., M.D., J.J., L.S., O.L., C.G., I.M., C.S., A.G. Investigation: J.E.K., F.A., E.S., P.P., D.S., D.C., M.D., J.J., L.S., O.L., C.L., C.G., I.M., C.S., A.G. Software: E.S., C.G., I.M., C.S. Methodology: J.E.K, F.A., E.S., C.G., I.M., C.S., A.G. Formal analysis: J.E.K., E.S., C.G., I.M., C.S., A.G. Data curation: J.E.K., E.S., C.L., C.G., I.M., C.S. Validation: J.E.K., E.S., C.L., C.G., I.M., C.S., A.G. Visualization: J.E.K., E.S., C.G., I.M., C.S. Funding acquisition: F.A., A.G. Writing – original draft: J.E.K. Writing – review and editing: J.E.K., F.A., E.S., P.P., D.S., D.C., M.D., J.J., L.S., O.L., C.L., C.G., I.M., C.S., A.G. Supervision: F.A., A.G. Project administration: J.E.K, A.G. All authors read and approved the final manuscript.

Funding

F.A. and A.G. were funded by the Sistema Nacional de Investigación (SNI) of Panama, contract numbers 072-2022 and 022-2020, respectively. F.A. was funded by the reinsertion scholarship program Grant DDCCT no. 068-2021 and Grant FIED23-07 / 137-2023 of the Secretaría Nacional de Ciencia, Tecnología e Innovación (SENACYT). Computational analyses were performed on the Mésocentre de Franche-Comté supercomputer facilities in France; the high-performance parallel computing unit Baru High Performance Computing at INDICAST-AIP, Panama; the core facilities of the LOPUKHIN FRCC PCM “Genomics, proteomics, metabolomics” (http://rcpcm.org/?p=2806), and the St. Petersburg Pasteur Institute, Russia.

Institutional Review Board Statement

This research used publicly available data and thus required no ethics approval.

Data Availability Statement

No new data were created in this study. The data presented in this study are available in NCBI at https://www.ncbi.nlm.nih.gov/. The BioProject and SRA numbers are provided in the article and the supplementary data file.

Acknowledgments

The authors would like to thank the support of collaborators at the Ministerio de Salud and Caja de Seguro Social of Panama, including Ana de Chavez, Maritza Mayrena, Maybis Garay, William Teerán (q.d.e.p), Victoria Williams, Lizbeth Garibaldi, Isolina Martinez, Yaracelis Cuadra, Rolando González, Luz Fruto, and other healthcare workers who identified and treated tuberculosis patients in Colón City, Panama, and whose isolates were also used for this study. We also thank Joanna Rivas for insightful discussions regarding this work.

Conflicts of Interest

The authors declare that they have no relevant financial or non-financial interests to disclose.

References

  1. Guyeux, C.; Senelle, G.; Meur, A.L.; Supply, P.; Gaudin, C.; Phelan, J.E.; Clark, T.G.; Rigouts, L.; Jong, B.d.; Sola, C.; et al. Newly Identified Mycobacterium africanum Lineage 10, Central Africa. 2024. [CrossRef]
  2. Senelle, G.; Guyeux, C.; Refrégier, G.; Sola, C. Towards a better understanding of the long-lasting evolutionary history of Mycobacterium tuberculosis. Tuberculosis 2023, 143, 102374. [CrossRef]
  3. Atavliyeva, S.; Auganova, D.; Tarlykov, P. Genetic diversity, evolution and drug resistance of Mycobacterium tuberculosis lineage 2. Frontiers in Microbiology 2024, Volume 15 - 2024. [CrossRef]
  4. Shitikov, E.; Bespiatykh, D. A revised SNP-based barcoding scheme for typing Mycobacterium tuberculosis complex isolates. mSphere 2023, 8, e00169-00123. [CrossRef]
  5. Thawornwattana, Y.; Mahasirimongkol, S.; Yanai, H.; Maung, H.M.W.; Cui, Z.; Chongsuvivatwong, V.; Palittapongarnpim, P. Revised nomenclature and SNP barcode for Mycobacterium tuberculosis lineage 2. Microb Genom 2021, 7, 000697. [CrossRef]
  6. Salud, M.d. Tuberculosis en Panamá Año 2024. 2024.
  7. Salud, M.d. Tuberculosis en Panamá Años 2022, 2023, 2024 (p). 2024.
  8. Acosta, F.; Norman, A.; Sambrano, D.; Batista, V.; Mokrousov, I.; Shitikov, E.; Jurado, J.; Mayrena, M.; Luque, O.; Garay, M.; et al. Probable long-term prevalence for a predominant Mycobacterium tuberculosis clone of a Beijing genotype in Colon, Panama. Transbound Emerg Dis 2021, 68, 2229-2238. [CrossRef]
  9. Acosta, F.; Norman, A.; Sambrano, D.; Batista, V.; Mokrousov, I.; Shitikov, E.; Jurado, J.; Mayrena, M.; Luque, O.; Garay, M.; et al. Corrigendum: Probable long-term prevalence for a predominant Mycobacterium tuberculosis clone of a Beijing genotype in Colon, Panama. Transbound Emerg Dis 2022, 69, 3142-3142. [CrossRef]
  10. Acosta, F.; Candanedo, D.; Patel, P.; Llanes, A.; Ku, J.E.; Salazar, K.; Morán, M.; Sambrano, D.; Jurado, J.; Martínez, I.; et al. Endemic transmission of a Mycobacterium tuberculosis L2.2.M3 sublineage of the L2 lineage within Colon, Panama: A prospective study. Infection, Genetics and Evolution 2025, 131, 105749. [CrossRef]
  11. Acosta, F.; Saldaña, R.; Miranda, S.; Candanedo, D.; Sambrano, D.; Morán, M.; Bejarano, S.; Arriba, Y.D.; Reigosa, A.; Dixon, E.D.; et al. Heterogeneity of Mycobacterium tuberculosis Strains Circulating in Panama’s Western Region. The American Journal of Tropical Medicine and Hygiene 2023, 109, 740-747. [CrossRef]
  12. Cerezo-Cortés, M.; Rodríguez-Castillo, J.; Hernández-Pando, R.; Murcia, M. Circulation of M. tuberculosis Beijing genotype in Latin America and the Caribbean. Pathogens and Global Health 2019, 113, 336-351. [CrossRef]
  13. Catalogue of mutations in Mycobacterium tuberculosis complex and their association with drug resistance, 2nd ed. 2023.
  14. Catalogue of mutations in Mycobacterium tuberculosis complex and their association with drug resistance, 1st ed. 2021.
  15. Verboven, L.; Phelan, J.; Heupink, T.H.; Van Rie, A. TBProfiler for automated calling of the association with drug resistance of variants in Mycobacterium tuberculosis. PLOS ONE 2022, 17, e0279644. [CrossRef]
  16. Verboven, L.; Phelan, J.; Heupink, T.H.; Van Rie, A. Correction: TBProfiler for automated calling of the association with drug resistance of variants in Mycobacterium tuberculosis. PLOS ONE 2023, 18, e0293254. [CrossRef]
  17. The CRyPTIC Consortium and the 100, G.P. Prediction of Susceptibility to First-Line Tuberculosis Drugs by DNA Sequencing. New England Journal of Medicine 2018, 379, 1403-1415. [CrossRef]
  18. Cingolani, P.; Platts, A.; Wang, L.L.; Coon, M.; Nguyen, T.; Wang, L.; Land, S.J.; Lu, X.; Ruden, D.M. A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff: SNPs in the genome of Drosophila melanogaster strain w1118; iso-2; iso-3. Fly 2012, 6, 80-92. [CrossRef]
  19. Gagneux, S.; Long, C.D.; Small, P.M.; Van, T.; Schoolnik, G.K.; Bohannan, B.J.M. The Competitive Cost of Antibiotic Resistance in Mycobacterium tuberculosis. Science 2006, 312, 1944-1946. [CrossRef]
  20. Nguyen, Q.H.; Nguyen, T.V.A.; Bañuls, A.-L. Multi-drug resistance and compensatory mutations in Mycobacterium tuberculosis in Vietnam. Tropical Medicine & International Health 2025, 30, 426-436. [CrossRef]
  21. van der Spuy, G.D.; Kremer, K.; Ndabambi, S.L.; Beyers, N.; Dunbar, R.; Marais, B.J.; van Helden, P.D.; Warren, R.M. Changing Mycobacterium tuberculosis population highlights clade-specific pathogenic characteristics. Tuberculosis (Edinb) 2009, 89, 120-125. [CrossRef]
  22. Tuberculosis, E.C.A.o.N.G.G.M.a.T.f.t.E.a.C.o. Beijing/W Genotype Mycobacterium tuberculosis and Drug Resistance. Emerg Infect Dis 2006, 12, 736-743. [CrossRef]
  23. Wang, J.; Liu, Y.; Zhang, C.-L.; Ji, B.-Y.; Zhang, L.-Z.; Shao, Y.-Z.; Jiang, S.-L.; Suzuki, Y.; Nakajima, C.; Fan, C.-L.; et al. Genotypes and Characteristics of Clustering and Drug Susceptibility of Mycobacterium tuberculosis Isolates Collected in Heilongjiang Province, China▿. J Clin Microbiol 2011, 49, 1354-1362. [CrossRef]
  24. Yang, C.; Luo, T.; Sun, G.; Qiao, K.; Sun, G.; DeRiemer, K.; Mei, J.; Gao, Q. Mycobacterium tuberculosis Beijing Strains Favor Transmission but Not Drug Resistance in China. Clin Infect Dis 2012, 55, 1179-1187. [CrossRef]
  25. Li, Y.; Cao, X.; Li, S.; Wang, H.; Wei, J.; Liu, P.; Wang, J.; Zhang, Z.; Gao, H.; Li, M.; et al. Characterization of Mycobacterium tuberculosis isolates from Hebei, China: genotypes and drug susceptibility phenotypes. BMC Infectious Diseases 2016, 16, 107. [CrossRef]
  26. Liu, Y.; Zhang, X.; Zhang, Y.; Sun, Y.; Yao, C.; Wang, W.; Li, C. Characterization of Mycobacterium tuberculosis strains in Beijing, China: drug susceptibility phenotypes and Beijing genotype family transmission. BMC Infectious Diseases 2018, 18, 658. [CrossRef]
  27. Zhao, L.-L.; Li, M.-C.; Liu, H.-C.; Xiao, T.-Y.; Li, G.-L.; Zhao, X.-Q.; Liu, Z.-G.; Wan, K.-L. Beijing genotype of Mycobacterium tuberculosis is less associated with drug resistance in south China. Int J Antimicrob Agents 2019, 54, 766-770. [CrossRef]
  28. Ockenga, J.; Fuhse, K.; Chatterjee, S.; Malykh, R.; Rippin, H.; Pirlich, M.; Yedilbayev, A.; Wickramasinghe, K.; Barazzoni, R. Tuberculosis and malnutrition: The European perspective. Clinical Nutrition 2023, 42, 486-492. [CrossRef]
  29. Datta, C.; Jha, R.K.; Ganguly, S.; Nagaraja, V. NapA (Rv0430), a Novel Nucleoid-Associated Protein that Regulates a Virulence Operon in Mycobacterium tuberculosis in a Supercoiling-Dependent Manner. J Mol Biol 2019, 431, 1576-1591. [CrossRef]
  30. Santoshi, M.; Tare, P.; Nagaraja, V. Nucleoid-associated proteins of mycobacteria come with a distinctive flavor. Molecular Microbiology 2025, 123, 177-194. [CrossRef]
  31. Javid, A.; Cooper, C.; Singh, A.; Schindler, S.; Hänisch, M.; Marshall, R.L.; Kalscheuer, R.; Bavro, V.N.; Bhatt, A. The mycolic acid reductase Rv2509 has distinct structural motifs and is essential for growth in slow-growing mycobacteria. Molecular Microbiology 2020, 113, 521-533. [CrossRef]
  32. Li, Y.; Fu, L.; Zhang, W.; Chen, X.; Lu, Y. The Transcription Factor Rv1453 Regulates the Expression of qor and Confers Resistant to Clofazimine in Mycobacterium tuberculosis. Infect Drug Resist 2021, 14, 3937-3948. [CrossRef]
  33. Zhang, L.; Zhang, Y.; Li, Y.; Huo, F.; Chen, X.; Zhu, H.; Guo, S.; Fu, L.; Wang, B.; Lu, Y. Rv1453 is associated with clofazimine resistance in Mycobacterium tuberculosis. Microbiol Spectr 2023, 11, e0000223. [CrossRef]
  34. Shimada, T. Transcriptional Regulation in Bacteria. Microorganisms 2024, 12, 2514. [CrossRef]
  35. (GTB), G.P.o.T.a.L.H. Global Tuberculosis Report 2024. 2024.
  36. Cords, O.; Martinez, L.; Warren, J.L.; O’Marr, J.M.; Walter, K.S.; Cohen, T.; Zheng, J.; Ko, A.I.; Croda, J.; Andrews, J.R. Incidence and prevalence of tuberculosis in incarcerated populations: a systematic review and meta-analysis. The Lancet Public Health 2021, 6, e300-e308. [CrossRef]
  37. Sequera, G.; Aguirre, S.; Estigarribia, G.; Walter, K.S.; Horna-Campos, O.; Liu, Y.E.; Andrews, J.R.; Croda, J.; Garcia-Basteiro, A.L. Incarceration and TB: the epidemic beyond prison walls. BMJ Glob Health 2024, 9, e014722. [CrossRef]
  38. Sanabria, G.E.; Sequera, G.; Aguirre, S.; Méndez, J.; dos Santos, P.C.P.; Gustafson, N.W.; Godoy, M.; Ortiz, A.; Cespedes, C.; Martínez, G.; et al. Phylogeography and transmission of Mycobacterium tuberculosis spanning prisons and surrounding communities in Paraguay. Nat Commun 2023, 14, 303. [CrossRef]
  39. Domínguez, J.; Acosta, F.; Pérez-Lago, L.; Sambrano, D.; Batista, V.; De La Guardia, C.; Abascal, E.; Chiner-Oms, Á.; Comas, I.; González, P.; et al. Simplified Model to Survey Tuberculosis Transmission in Countries Without Systematic Molecular Epidemiology Programs. Emerg Infect Dis 2019, 25, 507-514. [CrossRef]
  40. Senelle, G.; Guyeux, C.; Refrégier, G.; Sola, C. TB-annotator: a scalable web application that allows in-depth analysis of very large sets of publicly available Mycobacterium tuberculosis complex genomes. 2023. [CrossRef]
  41. Kozlov, A.M.; Darriba, D.; Flouri, T.; Morel, B.; Stamatakis, A. RAxML-NG: a fast, scalable and user-friendly tool for maximum likelihood phylogenetic inference. Bioinformatics 2019, 35, 4453-4455. [CrossRef]
  42. Minh, B.Q.; Schmidt, H.A.; Chernomor, O.; Schrempf, D.; Woodhams, M.D.; von Haeseler, A.; Lanfear, R. IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era. Molecular Biology and Evolution 2020, 37, 1530-1534. [CrossRef]
  43. Minh, B.Q.; Schmidt, H.A.; Chernomor, O.; Schrempf, D.; Woodhams, M.D.; von Haeseler, A.; Lanfear, R. Corrigendum to: IQ-TREE 2: New Models and Efficient Methods for Phylogenetic Inference in the Genomic Era. Molecular Biology and Evolution 2020, 37, 2461-2461. [CrossRef]
  44. Phelan, J.E.; O’Sullivan, D.M.; Machado, D.; Ramos, J.; Oppong, Y.E.A.; Campino, S.; O’Grady, J.; McNerney, R.; Hibberd, M.L.; Viveiros, M.; et al. Integrating informatics tools and portable sequencing technology for rapid detection of resistance to anti-tuberculous drugs. Genome Med 2019, 11, 41. [CrossRef]
  45. Coll, F.; McNerney, R.; Preston, M.D.; Guerra-Assunção, J.A.; Warry, A.; Hill-Cawthorne, G.; Mallard, K.; Nair, M.; Miranda, A.; Alves, A.; et al. Rapid determination of anti-tuberculosis drug resistance from whole-genome sequences. Genome Med 2015, 7, 51. [CrossRef]
  46. Ng, P.C.; Henikoff, S. Predicting deleterious amino acid substitutions. Genome Res 2001, 11, 863-874. [CrossRef]
  47. Mokrousov, I.; Slavchev, I.; Solovieva, N.; Dogonadze, M.; Vyazovaya, A.; Valcheva, V.; Masharsky, A.; Belopolskaya, O.; Dimitrov, S.; Zhuravlev, V.; et al. Molecular Insight into Mycobacterium tuberculosis Resistance to Nitrofuranyl Amides Gained through Metagenomics-like Analysis of Spontaneous Mutants. Pharmaceuticals 2022, 15, 1136. [CrossRef]
  48. Kapopoulou, A.; Lew, J.M.; Cole, S.T. The MycoBrowser portal: A comprehensive and manually annotated resource for mycobacterial genomes. Tuberculosis 2011, 91, 8-13. [CrossRef]
  49. White, J. PubMed 2.0. Medical Reference Services Quarterly 2020, 39, 382-387. [CrossRef]
  50. Szklarczyk, D.; Kirsch, R.; Koutrouli, M.; Nastou, K.; Mehryary, F.; Hachilif, R.; Gable, A.L.; Fang, T.; Doncheva, Nadezhda T.; Pyysalo, S.; et al. The STRING database in 2023: protein–protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Res 2023, 51, D638-D646. [CrossRef]
  51. D’Souza, C.; Kishore, U.; Tsolaki, A.G. The PE-PPE Family of Mycobacterium tuberculosis: Proteins in Disguise. Immunobiology 2023, 228, 152321. [CrossRef]
Figure 1. Global Mycobacterium phylogeny of sublineage L2.2.M3 isolates built using RAxML-NJ. L2.2.M3 isolates being featured in each group are colored pink. A) Panama Endemic Cluster A (PECA3) branch contains 94 isolates with origin in 4 countries. B) Closest Relatives 1 (CR1) branch contains 27 isolates in total, 24 isolates with origin in 7 countries, and three isolates with unknown origin. C) Closest Relatives 2 (CR2) branch contains 66 isolates in total, 47 isolates with origin in 10 countries, and 19 isolates with unknown origin. D) Closest Relatives 3 (CR3) branches contain 13 isolates from 4 countries.
Figure 1. Global Mycobacterium phylogeny of sublineage L2.2.M3 isolates built using RAxML-NJ. L2.2.M3 isolates being featured in each group are colored pink. A) Panama Endemic Cluster A (PECA3) branch contains 94 isolates with origin in 4 countries. B) Closest Relatives 1 (CR1) branch contains 27 isolates in total, 24 isolates with origin in 7 countries, and three isolates with unknown origin. C) Closest Relatives 2 (CR2) branch contains 66 isolates in total, 47 isolates with origin in 10 countries, and 19 isolates with unknown origin. D) Closest Relatives 3 (CR3) branches contain 13 isolates from 4 countries.
Preprints 231162 g001
Figure 2. Network analysis of gene Rv1453. This gene, inactivated in the PECA3 strain, revealed that all genes located above Rv1453 are interconnected via co-expression links and encode the following transcriptional regulatory proteins: cmr, embR, moxR2, moaR1, Rv3736, Rv3167c, Rv0894, and Rv3840. PPI enrichment p-value: 1.88 e-10; this network has significantly more interactions than expected. Network created with STRING tool v.12.0.
Figure 2. Network analysis of gene Rv1453. This gene, inactivated in the PECA3 strain, revealed that all genes located above Rv1453 are interconnected via co-expression links and encode the following transcriptional regulatory proteins: cmr, embR, moxR2, moaR1, Rv3736, Rv3167c, Rv0894, and Rv3840. PPI enrichment p-value: 1.88 e-10; this network has significantly more interactions than expected. Network created with STRING tool v.12.0.
Preprints 231162 g002
Table 1. Drug resistance profile by branches and country of isolation.
Table 1. Drug resistance profile by branches and country of isolation.
Countries by branch Sensitive
(%)
HR-TB
(%)
RR-TB
(%)
MDR-TB
(%)
Pre-XDR-TB
(%)
Other
(%)
PECA3 Branch (N = 94) 90 (95.7) 3 (3.2) 1 (1.1)
  Aruba (1) 1 (1.1)
  Panama (91) 87 (92.6) 3 (3.2) 1 (1.1)
  Portugal (1) 1 (1.1)
  USA (1) 1 (1.1)
  
CR1 Branch (N = 27) 11 (40.7) 4 (14.8) 1 (3.7) 6 (22.2) 3 (11.1) 2 (7.4)
  Australia (1) 1 (3.7)
  Canada (3) 1 (3.7) 2 (7.4)
  China (2) 1 (3.7) 1 (3.7)
  Madagascar (1) 1 (3.7)
  Myanmar (4) 1 (3.7) 3 (11.1)
  Origin unknown (3) 1 (3.7) 1 (3.7) 1 (3.7)
  Thailand (5) 2 (7.4) 1 (3.7) 2 (7.4)
  Vietnam (8) 5 (18.5) 1 (3.7) 2 (7.4)
  
CR2 Branch (N = 66) 38 (57.6) 2 (3) 4 (6.1) 1 (1.5) 21 (31.8)
  Australia (8) 5 (7.6) 3 (4.5)
  Cambodia (1) 1 (1.5)
  Canada (5) 3 (4.5) 2 (3)
  China (3) 2 (3) 1 (1.5)
  Guatemala (3) 1 (1.5) 2 (3)
  India (1) 1 (1.5)
  Myanmar (1) 1 (1.5)
  Origin unknown (19) 10 (15.2) 3 (4.5) 6 (9.1)
  Thailand (1) 1 (1.5)
  USA (6) 5 (7.6) 1 (1.5)
  Vietnam (18) 9 (13.6) 1 (1.5) 8 (12.1)
  
CR3 Branch (N = 13) 3 (23.1) 6 (46.2) 4 (30.8)
  Canada (1) 1 (7.7)
  China (10) 6 (46.2) 4 (30.8)
  USA (1) 1 (7.7)
  Vietnam (1) 1 (7.7)
PECA3: A3. Panama Endemic Cluster A3; CR1: Closest Relatives 1; CR2: Closest Relatives 2; CR3: Closest Relatives 3; USA: United States of America; HR-TB: resistant to isoniazid only; RR-TB: resistant to rifampicin only; MDR-TB: resistant to rifampicin and isoniazid; Pre-XDR-TB: MDR and resistant to a fluoroquinolone (levofloxacin, moxifloxacin, ciprofloxacin, ofloxacin); Other: resistant to any drug but none of the previous categories. The drug resistance profile by genome is available in Supplementary Material 1.
Table 2. List of SNPs found in PECA3 and close relatives.
Table 2. List of SNPs found in PECA3 and close relatives.
# POS REF Allele Annotation Putative Impact Gene Name Gene ID Transcript
Biotype
Sample
Count
PECA3 CR1 CR2 CR3
1 3827921 G A synonymous_variant LOW choD Rv3409c protein_coding 200 Yes Yes Yes Yes
2 797692 G A missense_variant MODERATE Rv0697 Rv0697 protein_coding 200 Yes Yes Yes Yes
3 1077920 G T intergenic_region MODIFIER Rv0966c-csoR Rv0966c-Rv0967 200 Yes Yes Yes Yes
4 1086544 G A missense_variant MODERATE accD2 Rv0974c protein_coding 199 Yes Yes Yes Yes*
5 2999257 G A synonymous_variant LOW dxs1 Rv2682c protein_coding 187 Yes Yes Yes
6 963851 C T synonymous_variant LOW mog Rv0865 protein_coding 187 Yes Yes Yes
7 2254060 A G missense_variant MODERATE otsB1 Rv2006 protein_coding 187 Yes Yes Yes
8 2501931 G A missense_variant MODERATE Rv2228c Rv2228c protein_coding 187 Yes Yes Yes
9 3720616 G A missense_variant MODERATE Rv3333c Rv3333c protein_coding 187 Yes Yes Yes
10 4380053 G T missense_variant MODERATE eccC2 Rv3894c protein_coding 121 Yes Yes
11 4390916 A G missense_variant MODERATE esxF Rv3905c protein_coding 121 Yes Yes
12 3297984 G A synonymous_variant LOW fadD22 Rv2948c protein_coding 121 Yes Yes
13 887719 G A synonymous_variant LOW lpdB Rv0794c protein_coding 121 Yes Yes
14 1544135 C T synonymous_variant LOW Rv1371 Rv1371 protein_coding 121 Yes Yes
15 1895809 G C missense_variant MODERATE Rv1669 Rv1669 protein_coding 121 Yes Yes
16 2539961 G T intergenic_region MODIFIER Rv2265-cyp124 Rv2265-Rv2266 121 Yes Yes
17 3034206 G C missense_variant MODERATE Rv2721c Rv2721c protein_coding 121 Yes Yes
18 3063391 C T synonymous_variant LOW Rv2750 Rv2750 protein_coding 121 Yes Yes
19 3243400 C G intergenic_region MODIFIER Rv2929-fadD26 Rv2929-Rv2930 121 Yes Yes
20 3136737 G A synonymous_variant LOW vapC22 Rv2829c protein_coding 121 Yes Yes
21 2185586 C T synonymous_variant LOW fadE17 Rv1934c protein_coding 94 Yes
22 423976 T A missense_variant MODERATE hspR Rv0353 protein_coding 94 Yes
23 3479530 C A synonymous_variant LOW moaC1 Rv3111 protein_coding 94 Yes
24 438757 G A intergenic_region MODIFIER Rv0360c-Rv0361 Rv0360c-Rv0361 94 Yes
25 518748 G A missense_variant MODERATE Rv0430 Rv0430 protein_coding 94 Yes**
26 717501 C T synonymous_variant LOW Rv0625c Rv0625c protein_coding 94 Yes
27 915653 C T missense_variant MODERATE Rv0822c Rv0822c protein_coding 94 Yes
28 1054953 C A missense_variant MODERATE Rv0945 Rv0945 protein_coding 94 Yes
29 1638962 C A stop_gained HIGH Rv1453 Rv1453 protein_coding 94 Yes
30 2266598 G C intergenic_region MODIFIER Rv2019-Rv2021c Rv2019-Rv2021c 94 Yes
31 2266624 G T intergenic_region MODIFIER Rv2019-Rv2021c Rv2019-Rv2021c 94 Yes
32 2288504 A C missense_variant MODERATE Rv2042c Rv2042c protein_coding 94 Yes
33 2729274 C T synonymous_variant LOW Rv2434c Rv2434c protein_coding 94 Yes
34 2824855 A G missense_variant MODERATE Rv2509 Rv2509 protein_coding 94 Yes
35 4178801 A G missense_variant MODERATE Rv3729 Rv3729 protein_coding 94 Yes
36 2819689 A C missense_variant MODERATE scoA Rv2504c protein_coding 94 Yes
37 2929910 C G missense_variant MODERATE speE Rv2601 protein_coding 94 Yes
38 1062869 G C synonymous_variant LOW sucC Rv0951 protein_coding 94 Yes
Total SNPs per branch 38 20 9 4
Branch-specific SNPs shared with PECA3 N/A 11 5 4
SNPs specific to PECA3 18
*Only 4831617. lacks the SNP at 1086544 G>A. **PECA3 defining SNP. SNP, Single-nucleotide polymorphism; PECA3, Panama Endemic Cluster A3; CR1, Closest Relatives 1; CR2, Closest Relatives 2; CR3, Closest Relatives 3; N/A, not applicable.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.