Preprint
Article

This version is not peer-reviewed.

Beyond Clinical Prediction: Unexplained Tuberculosis Treatment Outcomes Reveal a Need for Molecular Investigation in Rural South Africa

Submitted:

18 August 2026

Posted:

20 August 2026

You are already at the latest version

Abstract
Tuberculosis (TB) treatment outcomes may differ among individuals with similar routinely recorded demographic and clinical characteristics, suggesting limitations in the explanatory scope of programme surveillance data. This study evaluated the capacity of routinely collected demographic, clinical, socioeconomic, and behavioural variables to explain TB treatment outcomes and used identified evidence gaps to define priorities for complementary molecular research. We conducted a retrospective secondary analysis of anonymised routine TB programme data from 422 adults treated in public healthcare facilities in rural OR Tambo District Municipality, Eastern Cape, South Africa, between January 2018 and December 2022. Treatment success was defined as cure or treatment completion, whereas unsuccessful outcomes included death, treatment failure, or loss to follow-up. The explanatory and predictive performance of logistic regression, Poisson regression, and Random Forest models was evaluated, alongside a structured evidence-gap analysis comparing variables available in routine surveillance with molecular determinants of treatment response and drug resistance identified in the literature. Previous TB treatment was associated with lower odds of treatment success in univariable analysis (OR 0.48, 95% CI 0.29–0.78; p=0.003), but the association attenuated after adjustment (OR 0.64, 95% CI 0.38–1.09; p=0.10). HIV status was not independently associated with treatment success in logistic (OR 0.80, 95% CI 0.49–1.29; p=0.36) or Poisson regression (aRR 0.95, 95% CI 0.66–1.37; p=0.79). Random Forest achieved a higher average precision than logistic regression (0.907 versus 0.810), demonstrating improved prediction from routinely available variables; however, molecular mechanisms could not be evaluated because genomic, transcriptomic, immunological, and pharmacological measures were absent from the dataset. These findings define an important boundary between programme-level prediction and mechanistic explanation. We propose a program-to-molecular translational framework in which routine surveillance identifies clinically relevant patterns and knowledge gaps that can subsequently be investigated using genomic, transcriptomic, and functional molecular approaches. This integration may strengthen mechanistic understanding of heterogeneous TB treatment responses and inform future precision-oriented TB research.
Keywords: 
;  ;  ;  ;  ;  ;  

1. Introduction

Tuberculosis (TB) remains a major cause of infectious disease morbidity and mortality worldwide despite substantial advances in diagnosis, treatment, and prevention [1,2]. According to the World Health Organization (WHO), an estimated 10.8 million people developed TB in 2023 [2]. Drug-resistant tuberculosis (DR-TB) continues to pose a major challenge to global TB control because of its more complex treatment requirements, poorer outcomes, and potential for transmission of resistant strains [3]. The burden is particularly pronounced in low- and middle-income countries, where health systems face persistent challenges related to delayed diagnosis, constrained laboratory capacity, treatment adherence, and socioeconomic inequalities [4]. South Africa remains among the countries with a high burden of TB and HIV, with rural and resource-constrained settings such as the Eastern Cape experiencing substantial challenges related to TB and drug resistance [5].
Routine TB surveillance systems are essential for monitoring disease burden, evaluating treatment outcomes, identifying populations at increased risk, and informing public health interventions [6]. Programme datasets provide valuable real-world epidemiological information on demographic characteristics, HIV co-infection, previous TB treatment, treatment outcomes, socioeconomic circumstances, and patterns of healthcare utilisation [7]. These data have contributed substantially to TB programme monitoring, identification of vulnerable populations, resource allocation, and evidence-informed policy development. However, routine surveillance systems are primarily designed for programme monitoring and service delivery rather than mechanistic investigation. Consequently, they contain limited information on the biological processes that may influence disease progression, treatment response, bacterial persistence, and the emergence of drug resistance [8].
Advances in molecular microbiology have demonstrated that treatment response in Mycobacterium tuberculosis infection involves biological processes extending beyond routinely measured demographic and clinical characteristics [9]. Bacterial adaptive responses, including changes in gene expression, activation of stress-response pathways, regulation of efflux systems, metabolic reprogramming, and persistence mechanisms, may contribute to antimicrobial tolerance and treatment response [10]. Resistance-associated mutations in genes such as rpoB, katG, and inhA are important determinants of resistance to key anti-TB drugs [11]; however, phenotypic resistance and bacterial survival may also be influenced by transcriptional regulation, physiological state, metabolic adaptation, and environmental pressures during antimicrobial exposure [12]. A comprehensive understanding of heterogeneous TB treatment responses may therefore require complementary integration of epidemiological information with genomic, transcriptomic, and functional molecular data.
Numerous epidemiological studies have investigated demographic, socioeconomic, behavioural, and clinical predictors of TB treatment outcomes. Factors including age, sex, HIV status, previous TB treatment, socioeconomic circumstances, and behavioural characteristics have been associated with treatment outcomes across different populations [13,14]. However, the magnitude and consistency of these associations vary considerably across settings, and routinely captured variables explain only part of the heterogeneity observed in treatment outcomes. Machine-learning approaches offer an opportunity to model nonlinear relationships and complex interactions among routinely collected variables [15]. Nevertheless, the biological interpretation of such models remains inherently constrained by the information contained in the underlying dataset; analytical sophistication cannot directly characterize biological mechanisms that were not measured [16]. Thus, improved prediction using routine data should not necessarily be interpreted as improved mechanistic understanding of treatment response.
Evidence from rural South Africa further illustrates this distinction between prediction and biological explanation. A retrospective analysis of TB patients in the Eastern Cape reported a high prevalence of HIV co-infection, but HIV status was not independently associated with treatment success following adjustment for other variables [17]. Related analyses have demonstrated that demographic, clinical, socioeconomic, and contextual characteristics can contribute useful information for predicting TB outcomes, while still leaving important dimensions of treatment response inadequately characterised [18]. These observations highlight an important limitation of programme-based epidemiological research: routinely collected data can identify statistical associations and predictive patterns but cannot directly characterise unmeasured biological processes, including bacterial transcriptional responses, host–pathogen interactions, metabolic adaptation, and other molecular determinants potentially relevant to treatment response [19].
This limitation represents an important translational gap between population-level epidemiology and mechanistic laboratory research. Epidemiological surveillance can identify disease burden, vulnerable populations, treatment patterns, and programme-level determinants of outcomes, whereas molecular investigations can interrogate biological processes underlying antimicrobial resistance, bacterial persistence, and therapeutic response. Integrating these complementary approaches may therefore provide a more comprehensive understanding of heterogeneous TB treatment outcomes. Technologies such as whole-genome sequencing, transcriptomic profiling, and functional molecular assays provide opportunities to investigate biological mechanisms that cannot be directly assessed using routine programme data alone.
Importantly, unexplained variation in epidemiological models should not itself be interpreted as evidence that molecular mechanisms are responsible for treatment outcomes. Such variation may also reflect unmeasured clinical, behavioural, socioeconomic, healthcare-system, adherence, environmental, or biological factors. Rather, the absence of molecular information defines an important boundary of what can legitimately be inferred from routine surveillance data. Identifying this boundary can be scientifically valuable because it enables population-level observations to generate testable hypotheses for subsequent mechanistic investigation.
Accordingly, the novelty of the present study lies not primarily in identifying additional clinical predictors of TB treatment outcomes, but in examining the explanatory boundaries of routine TB programme data and translating the identified evidence gaps into priorities for molecular investigation. This program-to-molecular perspective positions routine surveillance as a starting point for translational research: epidemiological analyses identify patterns and unresolved questions, while subsequent laboratory investigations can determine whether genomic, transcriptomic, immunological, or other biological mechanisms contribute to those observations. Such integration may ultimately support the development of biologically informed risk stratification, biomarkers, and precision-oriented approaches to TB research.
Therefore, this study aimed to evaluate the explanatory capacity of routinely collected demographic, clinical, socioeconomic, and behavioural variables for TB treatment outcomes using secondary programme data from a rural South African setting and to systematically identify biological information absent from routine surveillance. By integrating conventional epidemiological modelling, machine-learning prediction, and a structured molecular evidence-gap assessment, the study further aimed to develop a translational framework linking programme-level observations with priorities for future genomic, transcriptomic, and functional molecular investigation. Within this framework, emerging experimental approaches, including photobiomodulation, are considered hypothesis-generating avenues for future laboratory investigation rather than interventions whose efficacy can be inferred from the present epidemiological data.

2. Materials and Methods

2.1. Study Design

This study employed a retrospective secondary analysis of routinely collected tuberculosis (TB) programme data. A translational epidemiological approach was adopted to evaluate the explanatory capacity of routinely available demographic, clinical, socioeconomic, and behavioural variables for TB treatment outcomes and to identify dimensions of treatment response that could not be investigated using routine programme data alone.
Rather than focusing primarily on the identification of additional clinical predictors, the analysis examined the extent to which routinely captured variables could account for observed differences in treatment outcomes using conventional statistical modelling and machine-learning approaches. The study subsequently incorporated a structured evidence-gap assessment to identify biological domains absent from the programme dataset, including genomic, transcriptomic, immunological, pharmacological, and functional molecular information.
The study was based on the premise that routine epidemiological surveillance and mechanistic laboratory research provide complementary levels of evidence. Routine programme data can characterise population-level patterns, associations, and predictive relationships but cannot directly determine molecular mechanisms that were not measured. Accordingly, unexplained variation in treatment outcomes was not assumed to have a molecular origin; it may reflect unmeasured clinical, behavioural, socioeconomic, healthcare-system, environmental, or biological factors.
The identified limitations of routine surveillance were therefore used to define testable questions and priorities for subsequent mechanistic research rather than to infer specific biological pathways. This programme-to-molecular translational approach provides a framework for linking epidemiological observations with future laboratory investigations of bacterial gene expression, genomic and transcriptomic variation, host–pathogen interactions, and other biological determinants of TB treatment response. Emerging experimental approaches, including photobiomodulation, were considered only as potential areas for future hypothesis-driven laboratory investigation and were neither evaluated nor assumed to have therapeutic efficacy in the present study.

2.2. Study Setting

The study used routinely collected tuberculosis (TB) programme data from public healthcare facilities within the OR Tambo District Municipality in the Eastern Cape Province of South Africa. OR Tambo District is predominantly rural and serves communities characterised by substantial socioeconomic disadvantage and a high burden of TB and HIV. Healthcare delivery across the district is provided through a network of primary healthcare clinics, community health centres, and hospitals operating within the South African public health system and implementing the National Tuberculosis Programme. TB services within these facilities include screening, diagnosis, treatment initiation, treatment monitoring, HIV testing and linkage to HIV care, and documentation of treatment outcomes as part of routine programme implementation. Patient-level information generated through these services forms part of routine TB programme surveillance and was the source of data used in the present secondary analysis. The rural context of OR Tambo District is associated with health-system and geographical challenges that may influence access to TB services, including long travel distances, transportation barriers, constrained access to specialised diagnostic services, and limitations in healthcare resources. This setting therefore provides an important real-world context in which to examine the explanatory capacity and information boundaries of routinely collected TB programme data.

2.3. Data Source

Secondary data were obtained from an existing retrospective cohort comprising routinely collected tuberculosis (TB) programme records for adult patients who received TB treatment between January 2018 and December 2022 in public healthcare facilities within OR Tambo District Municipality, Eastern Cape Province, South Africa. The dataset contained anonymised patient-level information derived from routine TB and HIV programme records and had previously been used to investigate epidemiological characteristics and treatment outcomes among patients with TB in the rural Eastern Cape. For the present study, the existing dataset was re-examined from a translational perspective to evaluate the explanatory and predictive capacity of routinely collected programme variables and to identify categories of biological information not captured through routine surveillance. The available data included demographic, socioeconomic, behavioural, clinical, and treatment-related characteristics routinely recorded during programme implementation. No new participant recruitment, clinical assessment, specimen collection, or laboratory investigation was undertaken for the present secondary analysis.

2.4. Study Population

The source population comprised adults aged ≥18 years with documented TB who initiated anti-TB treatment during the study period. The analytical cohort included patients with documented HIV status and a recorded final TB treatment outcome. Patients recorded as transferred out, those who remained on treatment at the time of data extraction, and those without sufficient information to determine a definitive treatment outcome were excluded from the outcome analysis. The final analytical dataset comprised 422 patients who met the eligibility criteria. Because the analysis was based on an existing programme dataset, no additional sampling or participant recruitment was undertaken. Restricting the analytical cohort to records with ascertainable final treatment outcomes was necessary for classification of the primary outcome; however, the potential for selection bias from excluding records with incomplete outcome information was considered when interpreting the findings.

2.5. Variables

The primary outcome was final TB treatment outcome. In accordance with the definitions used in the TB program, outcomes were dichotomized as treatment success (cure or treatment completion) or unsuccessful treatment (death, treatment failure, or loss to follow-up). Explanatory variables available in the routine programme dataset included demographic, socioeconomic, behavioural, and clinical characteristics recorded during patient management. These comprised age, sex, HIV status, previous TB treatment history, TB type (pulmonary or extrapulmonary), educational level, employment status, income source, smoking status, alcohol use, and available healthcare-related characteristics. The dataset did not contain molecular, genomic, transcriptomic, immunological, pharmacological, or detailed functional microbiological measurements. Consequently, biological processes such as bacterial gene-expression patterns, resistance-associated genomic variants, transcriptional regulation, host immune responses, oxidative stress responses, metabolic adaptation, efflux activity, bacterial persistence mechanisms, and host–pathogen molecular interactions could not be directly evaluated.
The absence of these measurements was treated as an information boundary of the routine dataset rather than evidence that any unmeasured biological mechanism caused the observed treatment outcomes. These information gaps were subsequently examined through the structured evidence-gap assessment to identify biological domains that may warrant investigation in future mechanistic studies.

2.6. Conceptual Framework

A programme-to-molecular translational framework was used to structure the study and distinguish between information that can be obtained from routine epidemiological surveillance and biological questions requiring complementary mechanistic investigation. The framework was developed to organise the analytical process rather than to infer molecular causation from epidemiological associations.
The framework comprised six sequential components: (i) routine TB programme surveillance; (ii) characterisation of routinely available demographic, socioeconomic, behavioural, and clinical variables; (iii) assessment of their explanatory and predictive capacity for treatment outcomes; (iv) identification of biological information not captured by routine surveillance; (v) formulation of priorities and testable hypotheses for future molecular research; and (vi) potential future translation of molecular evidence into improved TB diagnosis, risk stratification, treatment, and programme implementation (Figure 1).
Stage 1: Routine TB programme surveillance
The first component represents routine programme surveillance through which information on patient characteristics, TB diagnosis, HIV status, treatment initiation, treatment monitoring, and final treatment outcomes is generated. These data provide the epidemiological foundation for describing patient populations, monitoring programme performance, and evaluating treatment outcomes under routine healthcare conditions.
Stage 2: Routinely available predictors
The second component represents demographic, socioeconomic, behavioural, and clinical characteristics available for epidemiological analysis. These included age, sex, HIV status, previous TB treatment, TB type, education, employment, income source, smoking, and alcohol use. These variables constituted the routinely measured information available for evaluating variation in treatment outcomes.
Stage 3: Assessment of explanatory and predictive capacity
The third component involved evaluating the extent to which routinely available variables were associated with and could predict TB treatment outcomes. Conventional regression approaches and machine-learning analyses were used for this purpose. Importantly, predictive performance was distinguished from mechanistic explanation: an algorithm may discriminate between outcome groups based on measured characteristics without identifying the biological mechanisms underlying those outcomes.
Stage 4: Identification of biological evidence gaps
The fourth component involved systematically identifying biological domains that were not represented in the routine programme dataset. These included resistance-associated genomic variation, bacterial gene expression and transcriptional regulation, oxidative stress responses, efflux activity, dormancy and persistence pathways, metabolic adaptation, host immune responses, host–pathogen interactions, and bacterial population heterogeneity.
The absence of these variables was not interpreted as evidence that they explained the residual variation in the outcome. Rather, it identified biological questions that could not be addressed using the available epidemiological data and therefore represented potential areas for complementary mechanistic investigation.
Stage 5: Prioritization of future molecular research
The fifth component translated identified information gaps into testable priorities for future research. Potential approaches include whole-genome sequencing, transcriptomic profiling, RNA sequencing, host-response studies, and functional molecular experiments designed to investigate mechanisms of bacterial adaptation, antimicrobial resistance, persistence, and treatment response.
Emerging experimental interventions, including photobiomodulation (PBM), were situated within this stage solely as potential subjects for hypothesis-driven laboratory investigation. PBM exposure was not measured in the present dataset, and the framework does not assume or establish its molecular or therapeutic efficacy.
Stage 6: Potential translational pathway
The final component represents the prospective translation of evidence generated by future molecular investigations. Molecular findings, if independently validated, could potentially contribute to biomarker discovery, improved risk stratification, precision diagnostics, therapeutic development, and refinement of TB programme strategies.
Accordingly, the framework positions routine surveillance as a source of real-world epidemiological observations and clinically relevant research questions, while molecular and functional investigations provide complementary approaches for testing the biological mechanisms underlying those observations. Figure 1 summarises this progression from programme surveillance through evidence-gap identification and hypothesis generation to future mechanistic investigation and potential clinical translation.

2.7. Statistical Analysis

Descriptive statistics were used to characterise the demographic, clinical, socioeconomic, and behavioural characteristics of the analytical cohort. Continuous variables were summarised using medians and interquartile ranges (IQR), while categorical variables were presented as frequencies and percentages. The present study undertook a secondary evaluation of findings from previously developed regression and machine-learning models based on the same routine programme dataset. Multivariable logistic regression estimates were examined to assess the independent associations between routinely collected explanatory variables and the binary treatment outcome. Adjusted odds ratios (ORs) with 95% confidence intervals (CIs) and corresponding p-values were used to describe the magnitude and precision of these associations. Findings from Poisson regression models were additionally examined using adjusted risk ratios (aRRs) and 95% CIs to assess the robustness and interpretability of associations with treatment outcome.
Predictive performance was further evaluated using previously developed exploratory Random Forest analyses and compared with conventional logistic regression. Because treatment-success categories were imbalanced, particular emphasis was placed on precision–recall performance and average precision (AP), rather than classification accuracy alone. Feature-importance estimates from the Random Forest model were examined to identify the variables that contributed most strongly to prediction. However, feature importance was interpreted as a measure of predictive contribution and not as evidence of an independent association or causal effect.
The statistical evaluation distinguished association, prediction, and mechanistic explanation as conceptually different forms of inference. Regression models were used to assess statistical associations between measured variables and treatment outcomes, whereas Random Forest analysis evaluated predictive performance using patterns in the available data. Neither approach was considered capable of identifying biological mechanisms that were not directly measured. Accordingly, interpretation was based not only on statistical significance but also on effect estimates, confidence intervals, consistency across modelling approaches, predictive performance, and the scope of variables represented in the dataset. A two-sided p-value <0.05 was considered statistically significant for inferential analyses. Residual or unexplained variation in treatment outcomes was not attributable to any specific biological mechanism because the dataset lacked the measurements required to test them.

2.8. Assessment of Molecular Evidence Gaps

A structured evidence-gap assessment was undertaken to determine which categories of information relevant to TB treatment response were represented in routine programme surveillance and which biological domains could not be evaluated using the available dataset. Variables contained within the programme database were organised into demographic, socioeconomic, behavioural, clinical, and programme-related domains.
These routinely available domains were then conceptually compared with biological processes relevant to TB treatment response and antimicrobial resistance described in the contemporary molecular literature. The molecular domains considered included resistance-associated genomic variation, bacterial gene-expression profiles, transcriptional regulation, oxidative stress responses, efflux activity, dormancy and persistence pathways, metabolic adaptation, host immune responses, host–pathogen interactions, transcriptomic biomarkers, and genomic signatures associated with antimicrobial resistance.
The purpose of this assessment was not to determine whether any unmeasured molecular factor caused the observed treatment outcomes, but to identify biological questions that could not be addressed using routine programme data. A molecular domain was therefore considered an evidence gap when it was relevant to the biological understanding of TB treatment response but was not represented by corresponding measurements in the programme dataset.
The resulting evidence gaps were used to define priorities and testable hypotheses for future laboratory-based research incorporating genomic, transcriptomic, immunological, pharmacological, and functional molecular approaches. This analytical distinction prevented the absence of measured variables from being interpreted as evidence of their causal contribution to treatment outcomes.

2.9. Translational Interpretation

Findings were interpreted within a programme-to-molecular translational framework linking population-level epidemiological surveillance with hypothesis-driven mechanistic research. Epidemiological associations and predictive patterns were first used to establish what could be inferred from routinely collected programme variables. Areas that remained biologically uncharacterised were subsequently mapped to questions requiring direct laboratory investigation. Importantly, no causal molecular mechanisms were inferred from the epidemiological data, and unexplained variation in outcomes was not assumed to reflect biological factors alone. Unmeasured clinical characteristics, treatment adherence, healthcare access, health-system factors, environmental exposures, socioeconomic conditions, measurement error, and other unrecorded determinants may also contribute to differences in treatment outcomes. The translational interpretation, therefore, focused on identifying research questions that could be investigated using complementary approaches, including RNA sequencing, whole-genome sequencing, host-response profiling, and functional molecular experiments. Photobiomodulation was considered only as an example of an emerging intervention that could be evaluated in future hypothesis-driven laboratory studies; it was not measured in the present dataset, and no conclusions regarding its molecular effects, safety, or therapeutic efficacy were drawn from this analysis. The study consequently represents a hypothesis-generating bridge between routine programme epidemiology and future mechanistic research rather than direct evidence for precision therapeutic interventions. Molecular and functional validation would be required before identified biological hypotheses could inform diagnostic, prognostic, or therapeutic applications in drug-resistant TB.

2.9. Ethical Clearance

This study involved a secondary analysis of an existing, fully anonymized tuberculosis (TB) program dataset. Ethical approval for the original study and the use of the data was obtained from the Walter Sisulu University Research Ethics and Biosafety Committee (Reference No. 140/2025; 02 July 2025) and the Eastern Cape Department of Health (Reference No. EC_202507_022; 11 July 2025). No additional participant recruitment, direct participant contact, or prospective data collection was undertaken for the present analysis. All records were de-identified prior to analysis, and no information that could identify individual participants was accessed or reported. Participant confidentiality and data security were maintained throughout the study in accordance with applicable institutional and ethical requirements governing the secondary use of routinely collected healthcare data. The study was conducted in accordance with the principles of the Declaration of Helsinki and relevant institutional and national ethical guidelines.

3. Results

3.1. Characteristics of the Routine Programme Dataset

The secondary analysis included 422 adult patients with TB managed through routine TB services in the rural Eastern Cape, South Africa. The programme dataset comprised demographic, socioeconomic, behavioural, and basic clinical characteristics routinely captured during programme implementation, including age, sex, HIV status, TB type, previous TB treatment, education, employment status, income source, smoking, alcohol use, and final treatment outcomes. The dataset did not contain molecular, genomic, transcriptomic, immunological, pharmacological, or detailed functional microbiological measurements. Specifically, information on bacterial gene-expression profiles, resistance-associated genomic variation, transcriptional activity, bacterial stress responses, metabolic signatures, host inflammatory or immune biomarkers, and functional molecular measurements was unavailable. Photobiomodulation (PBM) exposure was also not recorded. Consequently, these biological domains could not be directly evaluated in relation to treatment outcomes using the available programme data.

3.2. Associations Between Routinely Collected Variables and TB Treatment Outcomes

In univariable logistic regression, previous TB treatment was associated with lower odds of treatment success among patients undergoing retreatment (PT1) compared with those with new TB (OR = 0.48, 95% CI 0.29–0.78; p = 0.003). A similar but non-significant association was observed among patients with more than one previous treatment episode (PT2) (OR = 0.36, 95% CI 0.08–1.57; p = 0.174). Reporting no income showed a non-significant association with higher odds of treatment success (OR = 1.84, 95% CI 0.92–3.67; p = 0.085). Following multivariable adjustment, previous TB treatment was no longer significantly associated with treatment success (OR = 0.64, 95% CI 0.38–1.09; p = 0.10). This finding was consistent with the Poisson regression analy+sis (aRR = 0.87, 95% CI 0.69–1.10; p = 0.24). HIV status was also not independently associated with treatment success in logistic regression (OR = 0.80, 95% CI 0.49–1.29; p = 0.36) or Poisson regression (aRR = 0.95, 95% CI 0.66–1.37; p = 0.79). Age, sex, pulmonary TB, education, income, and employment status similarly showed no statistically significant independent associations with treatment outcome, as shown in Table 1.

3.3. Predictive Performance of Machine-Learning and Regression Models

The previously developed Random Forest (RF) model achieved higher average precision than logistic regression (LR), with AP values of 0.907 and 0.810, respectively. Precision–recall analysis similarly indicated better classification performance for RF across most recall levels. Feature-importance analysis identified age as the strongest contributor to RF prediction (importance = 0.555), followed by sex (0.066), HIV status (0.053), no income (0.033), secondary education (0.032), primary education (0.032), tertiary education (0.029), unemployment (0.026), and government employment (0.026). Importantly, predictive importance did not necessarily correspond to an independent statistical association. For example, age contributed substantially to RF prediction but was not independently associated with treatment success in multivariable logistic regression. Similarly, HIV status contributed to RF prediction but was not independently associated with treatment success in either logistic or Poisson regression, as seen in Table 2.

3.4. Biological Information Absent from Routine Surveillance

The structured evidence-gap assessment identified several biological domains relevant to TB treatment response and antimicrobial resistance that were not represented in the routine programme dataset. These included bacterial transcriptional responses and gene-expression profiles, resistance-associated genomic variation, oxidative stress responses, efflux activity, dormancy and persistence pathways, metabolic adaptation, host inflammatory and immune responses, host–pathogen molecular interactions, and genomic and transcriptomic biomarkers. The absence of these measurements limited the ability of the programme dataset to investigate biological mechanisms potentially relevant to heterogeneous treatment responses and antimicrobial resistance. Importantly, the absence of these variables was not interpreted as evidence that the corresponding biological mechanisms caused the observed treatment outcomes. Rather, the evidence-gap assessment established that these mechanisms could not be directly evaluated using the available routine surveillance data. The identified gaps therefore define specific biological questions requiring complementary genomic, transcriptomic, immunological, and functional molecular investigation as noted in Table 3.

3.5. Translation of Epidemiological Evidence into Molecular Research Priorities

Mapping the identified evidence gaps onto the programme-to-molecular translational framework highlighted several priorities for subsequent mechanistic investigation. These include genomic characterisation of resistance-associated variation in Mycobacterium tuberculosis; transcriptomic profiling to characterise bacterial gene-expression responses during treatment; investigation of transcriptional regulation under antimicrobial exposure; assessment of bacterial persistence, oxidative stress responses, efflux activity, and metabolic adaptation; and investigation of host–pathogen interactions associated with heterogeneous treatment responses. Future studies should also evaluate whether independently validated genomic, transcriptomic, and other molecular biomarkers provide complementary information when integrated with routinely collected epidemiological and clinical variables. These priorities are intended to generate testable molecular hypotheses from limitations identified at the programme level, rather than to imply that the unmeasured biological processes account for the treatment outcomes observed in the present cohort. Molecular and functional studies will therefore be required to determine whether these mechanisms contribute to treatment response, bacterial persistence, or antimicrobial resistance. The framework consequently positions routine epidemiological surveillance and molecular investigation as complementary approaches, with programme data identifying clinically relevant patterns and evidence gaps and laboratory studies providing the means to interrogate their potential biological basis as indicated in Table 4.

3.6. Evidence Gap Relevant to Future Photobiomodulation Research

Photobiomodulation (PBM) exposure and its biological effects were not assessed in the routine programme dataset. Consequently, the present analysis provides no direct evidence regarding the effects of PBM on Mycobacterium tuberculosis, antimicrobial resistance, treatment response, or clinical outcomes. However, the structured evidence-gap assessment identified bacterial gene expression, oxidative stress responses, metabolic adaptation, and transcriptional regulation as biological domains that were not captured by routine surveillance data. These domains represent potential molecular endpoints for future hypothesis-driven studies investigating whether PBM induces measurable biological responses in M. tuberculosis. Such investigations would require controlled laboratory experiments incorporating appropriate molecular and functional assays. Accordingly, the present findings identify PBM as a hypothesis-generating area for future mechanistic investigation rather than providing evidence of its biological or therapeutic efficacy.

4. Discussion

The present secondary analysis demonstrates the value of routinely collected tuberculosis (TB) programme data for characterizing treatment outcomes and evaluating population-level determinants of programme performance. At the same time, the findings delineate an important boundary between epidemiological prediction and mechanistic explanation. Routinely available demographic, socioeconomic, behavioural, and clinical characteristics provided useful information on treatment outcomes, but few variables demonstrated independent associations after multivariable adjustment. Although machine-learning analysis improved predictive performance relative to conventional logistic regression, the underlying dataset lacked the molecular, genomic, transcriptomic, immunological, or pharmacological measurements needed to investigate the biological mechanisms underlying treatment response. Rather than attributing residual outcome heterogeneity to any specific unmeasured mechanism, the present study uses these information gaps to define priorities for complementary molecular investigation. This programme-to-molecular perspective provides a translational framework through which real-world epidemiological observations can generate testable questions for subsequent mechanistic research.

4.1. Routine Surveillance Remains Indispensable but Biologically Incomplete

Routine TB surveillance systems are fundamental to national and global TB control programmes. They provide standardised information on disease burden, treatment outcomes, HIV co-infection, previous treatment, and programme performance, thereby supporting evidence-informed policy development, resource allocation, and monitoring of public health interventions [20]. In high-burden settings, including rural areas of the Eastern Cape, such surveillance is particularly important because it captures programme performance under real-world healthcare conditions. Evidence from high-TB-burden countries demonstrates the importance of strengthening routine surveillance systems through timely case detection and notification, complete and accurate patient-level information, interoperability with laboratory and pharmacy systems, accessibility across healthcare sectors, and the generation of actionable information for programme management [21]. Evidence from the Eastern Cape similarly emphasises the importance of active TB surveillance and linkage to care for earlier case identification and reduction of ongoing transmission, particularly in high-burden settings [22].
Despite these strengths, routine surveillance systems are designed primarily for programme monitoring and public health decision-making rather than mechanistic biological investigation [23,24]. They can characterise patient populations, treatment pathways, clinical outcomes, and healthcare utilisation but generally contain limited information on bacterial physiology, genomic variation, transcriptional responses, host immunity, pharmacological exposures, or molecular adaptation during treatment. Consequently, routine surveillance alone cannot determine the biological mechanisms underlying differences in treatment response among patients with otherwise similar routinely recorded characteristics [25]. This distinction is important. The inability of routine programme data to characterise molecular mechanisms does not diminish their epidemiological value; rather, it defines the level of inference that can reasonably be made from such data. Programme surveillance and molecular investigation should therefore be considered complementary components of TB research rather than interchangeable approaches.

4.2. Routine Clinical Predictors Explain Only Part of Treatment Variability

One of the principal findings of the present analysis was that few routinely collected demographic, socioeconomic, and clinical characteristics demonstrated independent associations with TB treatment outcomes after multivariable adjustment. The contribution of individual demographic and clinical predictors to TB treatment outcomes has varied considerably across populations and healthcare settings [16,26,27].
Previous TB treatment was associated with reduced odds of treatment success in the univariable analysis, but the association was attenuated after adjustment for other measured characteristics. Similarly, HIV status, age, sex, TB type, education level, employment status, and income source were not independently associated with treatment outcomes in the multivariable analysis.
These findings partly correspond with evidence from two rural clinics in the Eastern Cape, where HIV status, sex, previous TB treatment, employment, and social history were not independently associated with treatment outcomes [26]. However, that study reported that increasing age was associated with lower odds of treatment success and that pulmonary TB independently predicted successful treatment [26]. A systematic review and meta-analysis of TB treatment outcomes in Africa also reported associations between HIV co-infection, previous treatment, and unsuccessful outcomes, whereas age, sex, and TB classification were not consistently associated across the included studies [28].
Such differences highlight the context-dependent nature of epidemiological predictors. Associations may vary according to patient characteristics, disease severity, healthcare access, treatment regimens, adherence, programme implementation, outcome definitions, sample size, and the covariates incorporated into statistical models [26,27,28,29]. Therefore, the absence of statistically significant independent associations in the present cohort should not be interpreted as evidence that these factors are clinically unimportant.
More importantly, routinely recorded clinical and socioeconomic variables represent only one level of the determinants influencing TB treatment outcomes. HIV infection, previous treatment, socioeconomic circumstances, healthcare access, and other programme-level characteristics can influence treatment pathways and outcomes [28,30], but routine datasets may not capture the full range of patient-, pathogen-, treatment-, healthcare-system-, and environmental-level determinants involved. Residual variation in statistical models should therefore not automatically be attributed to molecular processes. It may reflect unmeasured clinical characteristics, adherence patterns, treatment exposure, healthcare-system factors, socioeconomic circumstances, environmental influences, measurement error, biological heterogeneity, or combinations of these factors [31,32,33]. The present findings consequently identify an information boundary of the available programme data rather than establishing a specific biological explanation for treatment heterogeneity.

4.3. Machine Learning Improves Prediction, but Cannot Replace Biological Information

Machine-learning approaches have increasingly been investigated for predicting TB outcomes because of their ability to model complex, nonlinear relationships and interactions among multiple variables [34,35]. In the present analysis, the Random Forest model achieved a higher average precision than conventional logistic regression, increasing from 0.810 to 0.907. This finding is consistent with evidence showing that machine-learning approaches can outperform conventional statistical models in some TB prediction settings. Lee et al. reported that XGBoost achieved an AUROC of 0.818 and an AUPRC of 0.432 for predicting unfavourable TB treatment outcomes, compared with an AUROC of 0.758 and an AUPRC of 0.251 for conventional logistic regression [36]. Adeboye et al. similarly reported better predictive performance of a Random Survival Forest model than a conventional Cox regression model for TB mortality, including a lower Integrated Brier Score and higher Integrated AUC [34]. However, improved performance of machine learning over conventional regression is not universal. A study from China reported that logistic regression outperformed Random Forest in distinguishing latent from active TB, achieving an AUROC of 0.957 and AUPRC of 0.977 compared with 0.903 and 0.922, respectively, for Random Forest [37]. These differences reinforce the importance of dataset characteristics, outcome definition, sample size, predictor availability, model development, and validation strategy when evaluating machine-learning performance. An important observation in the present analysis was that variables with high predictive importance did not necessarily demonstrate independent statistical associations. Age, for example, contributed substantially to Random Forest prediction but was not independently associated with treatment success in multivariable logistic regression. This distinction reflects the different purposes of explanatory and predictive modelling. Regression estimates are useful for quantifying conditional statistical associations, whereas machine-learning algorithms can exploit nonlinearities and interactions to optimise prediction without necessarily identifying causal relationships. Crucially, neither analytical approach can identify biological mechanisms that are absent from the underlying data [38,39]. Machine-learning models can detect complex patterns among measured variables, but increasing algorithmic sophistication does not compensate for the absence of relevant biological measurements [40]. The higher average precision achieved by Random Forest should therefore be interpreted as improved prediction from the available programme variables rather than evidence that the biological basis of treatment response has been explained. These observations have implications for artificial intelligence applications in infectious disease research. Machine learning and molecular biology should not be regarded as competing approaches. Future research should instead investigate whether combining epidemiological and clinical information with independently validated genomic, transcriptomic, proteomic, metabolomic, pharmacological, and immunological biomarkers improves prediction, biological interpretation, and clinical utility.

4.4. Bridging the Gap Between Epidemiology and Molecular Microbiology

A central contribution of this study is the identification and systematic characterisation of an information gap between routine epidemiological surveillance and molecular investigation. Comparison of routinely collected programme variables with contemporary understanding of multidrug-resistant tuberculosis (MDR-TB) biology demonstrated that the programme dataset did not capture several biological domains relevant to treatment response and antimicrobial resistance. These included resistance-associated genomic variation, bacterial gene-expression profiles, transcriptional regulation, oxidative stress responses, efflux activity, dormancy and persistence pathways, metabolic adaptation, host immune responses, host–pathogen molecular interactions, and genomic and transcriptomic biomarkers [41,42].
Molecular methods provide information about biological processes that cannot be directly inferred from conventional demographic and clinical surveillance variables [43]. Under antimicrobial exposure, Mycobacterium tuberculosis can undergo transcriptional and physiological adaptations, including changes in gene expression, stress responses, and metabolic activity, that may contribute to drug tolerance and bacterial persistence [43,44,45]. Genomic analyses can additionally identify mutations and other genetic variation associated with antimicrobial resistance and transmission [41,42,46]. Efflux mechanisms and other adaptive responses may further influence antimicrobial susceptibility and bacterial survival [47]. Experimental molecular studies have demonstrated that M. tuberculosis can dynamically alter transcriptional and physiological programmes under environmental and antimicrobial pressures, supporting the importance of investigating bacterial adaptation alongside conventional resistance-associated genomic variation [43,44,45,46,47,48]. However, the present study did not measure these mechanisms. Their biological relevance in the wider TB literature should therefore be distinguished from evidence generated directly by this cohort. The current analysis establishes that these domains were absent from the routine dataset, not that they caused unsuccessful outcomes among the participants studied.
This distinction provides the basis for a more defensible programme-to-molecular research strategy. Epidemiological surveillance can identify disease burden, population-level patterns, treatment outcomes, and clinically relevant questions, while appropriately designed molecular studies can subsequently test biological mechanisms underlying those observations. Integration across these levels of investigation may strengthen both epidemiological interpretation and translational TB research.

4.5. Translating Epidemiological Evidence into Molecular Research

The principal conceptual innovation of this study is the development of a programme-to-molecular translational framework linking routine TB surveillance with hypothesis-driven laboratory investigation. Rather than treating surveillance as the endpoint of research, the framework positions routine programme data as a source of real-world observations from which unresolved clinical and biological questions can be identified. Within this framework, programme data first characterise disease burden, patient characteristics, treatment outcomes, and epidemiological patterns. Statistical and machine-learning approaches can then determine the extent to which available variables are associated with or predict treatment outcomes. Biological domains that are not represented in the dataset are subsequently identified as evidence gaps rather than presumed causes of residual outcome variation. These gaps can be translated into specific, testable priorities for future research. Whole-genome sequencing could characterise resistance-associated genomic variation [41,42,46], while transcriptomic approaches could investigate bacterial gene-expression and regulatory responses during antimicrobial exposure [43,44,45,48]. Functional molecular studies could further examine persistence pathways, efflux activity, oxidative stress responses, metabolic adaptation, and host–pathogen interactions [42,43,44,45,46,47,48]. Ultimately, future studies could investigate whether independently validated molecular biomarkers add explanatory or predictive information when combined with routinely collected epidemiological variables. Such work would provide a stronger empirical basis for developing integrated risk-stratification approaches, biomarkers, diagnostic strategies, and precision-oriented therapeutic research.

4.6. Implications for Future Photobiomodulation Research

Photobiomodulation (PBM) represents an emerging experimental area that may warrant mechanistic investigation within the broader programme-to-molecular framework. However, PBM exposure and PBM-associated biological responses were not measured in the present dataset. The current study therefore provides no direct evidence that PBM affects M. tuberculosis, antimicrobial resistance, treatment response, or clinical outcomes.
The relevance of PBM to the present framework arises from the biological evidence gaps identified through the analysis. Bacterial gene expression, oxidative stress responses, transcriptional regulation, metabolic adaptation, and other functional molecular processes were not represented in routine programme surveillance. These domains provide potential molecular endpoints through which emerging interventions could be experimentally evaluated.
Accordingly, any investigation of PBM within this framework should begin with controlled laboratory studies to determine whether PBM produces reproducible molecular or physiological effects in M. tuberculosis. Appropriate approaches could include transcriptomic profiling, functional molecular assays, metabolic measurements, and other validated experimental endpoints. Only if reproducible biological effects are demonstrated would subsequent investigation of therapeutic relevance be justified.
Thus, the present study positions PBM as a hypothesis-generating direction for future mechanistic research rather than as an intervention supported by the epidemiological findings. This distinction is particularly important because the references used to establish the molecular evidence gaps [41,42,43,44,45,46,47,48] principally address TB molecular biology, antimicrobial resistance, transcriptional adaptation, and related mechanisms; they should not be interpreted as direct evidence for PBM unless supported by PBM-specific experimental literature.

4.7. Recommendations

Based on the findings and evidence gaps identified in this secondary analysis, several priorities are proposed for future research.
  • Strengthen the linkage between routine surveillance and molecular data. Future TB research platforms should explore mechanisms to link routinely collected epidemiological and clinical information to validated molecular data, including resistance-associated genomic variants and relevant molecular biomarkers. Such integration should be evaluated prospectively before incorporation into routine surveillance systems.
  • Conduct prospective multi-omics investigations of treatment response. Future studies should integrate epidemiological and clinical information with whole-genome sequencing, transcriptomics, proteomics, metabolomics, pharmacologic measurements, and host immunologic profiling to investigate determinants of bacterial persistence, antimicrobial resistance, and heterogeneous treatment responses.
  • Develop and externally validate integrated predictive models. Future machine-learning studies should evaluate whether validated molecular biomarkers provide incremental predictive value when added to routine demographic and clinical variables. Model development should incorporate appropriate internal and external validation to determine generalisability and clinical utility.
  • Evaluate PBM through a mechanistic laboratory investigation. Before any consideration of clinical translation, controlled experimental studies should determine whether PBM produces measurable and reproducible effects on M. tuberculosis gene expression, oxidative stress responses, transcriptional regulation, metabolic activity, or other relevant molecular pathways.
  • Strengthen interdisciplinary translational collaboration. Greater collaboration among TB programmes, clinicians, epidemiologists, molecular microbiologists, bioinformaticians, laboratory scientists, and biomedical researchers may facilitate the translation of program-level observations into mechanistically testable research questions and, where appropriately validated, the translation of molecular discoveries back into clinical and public health applications.
Collectively, these recommendations provide a pathway for moving from routine epidemiological observation to hypothesis-driven molecular investigation while maintaining a clear distinction between associations identified in programme data and biological mechanisms requiring direct experimental validation.

4.8. Strengths and Limitations

A major strength of this study is its translational perspective. Rather than undertaking a conventional secondary analysis focused solely on identifying additional epidemiological predictors, the study evaluated the explanatory and informational boundaries of routinely collected TB programme data and used these boundaries to develop a framework linking population-level surveillance with future mechanistic research. The integration of conventional regression, machine-learning prediction, and a structured biological evidence-gap assessment provides a multidimensional approach for distinguishing epidemiological association and prediction from mechanistic explanation.
The use of real-world programme data from a rural, high-burden setting is another strength because it reflects routine TB service delivery and treatment outcomes under operational healthcare conditions. The comparison between conventional regression and Random Forest analysis also illustrates that improved predictive performance does not necessarily translate into improved mechanistic understanding.
Several limitations should nevertheless be acknowledged. First, the analysis relied on an existing retrospective programme dataset and was restricted to variables collected during routine service delivery. Important clinical, behavioural, adherence, healthcare-system, environmental, and treatment-exposure variables may therefore have been unavailable or incompletely measured.
Second, the analytical cohort was restricted to patients with documented HIV status and ascertainable final treatment outcomes. Exclusion of transferred patients, individuals remaining on treatment, and records with incomplete outcome information may have introduced selection bias if excluded individuals differed systematically from those retained in the analysis.
Third, molecular, genomic, transcriptomic, immunological, pharmacological, and detailed functional microbiological measurements were unavailable. Consequently, the study could identify these domains as evidence gaps but could not determine whether they explained treatment outcomes in the present cohort.
Fourth, the regression and machine-learning analyses were based on routinely collected variables from a relatively modest sample of 422 patients. The predictive performance of the Random Forest model should therefore be interpreted cautiously, particularly in the absence of independent external validation. Feature importance should likewise not be interpreted as evidence of causal effects.
Fifth, the structured molecular evidence gap assessment was informed by existing literature rather than by direct molecular measurements from cohort participants. The proposed program-to-molecular framework is consequently evidence-informed and hypothesis-generating, and its utility requires evaluation in prospective studies incorporating epidemiological and molecular data.
Sixth, PBM was not evaluated experimentally or clinically in this study. No conclusions can therefore be drawn regarding its molecular effects, safety, efficacy, or therapeutic value in TB or DR-TB. Its inclusion in the framework is limited to identifying a potential direction for controlled mechanistic investigation.
Finally, because the study was conducted using programme data from a rural South African setting, the findings and proposed framework may not be directly generalizable to settings with different epidemiological profiles, healthcare systems, treatment practices, or data infrastructure. External evaluation in other populations and prospective datasets will be required.
Despite these limitations, the study makes a conceptual and methodological contribution by demonstrating how routine surveillance data can be used not only to characterize TB program outcomes but also to identify questions that surveillance alone cannot answer, thereby prioritizing subsequent mechanistic investigation.

4.9. Conclusions

This secondary analysis demonstrates that routine TB programme surveillance provides an essential foundation for epidemiological monitoring, characterization of treatment outcomes, and identification of population-level patterns, but it has limited capacity to interrogate biological mechanisms that are not directly measured. Routinely collected demographic, clinical, socioeconomic, and behavioural variables provided useful epidemiological information, while machine learning improved prediction performance compared with conventional logistic regression. Nevertheless, improved predictive performance should not be equated with a mechanistic explanation.
Importantly, residual variation in treatment outcomes cannot be assumed to reflect molecular processes alone. Unmeasured clinical, behavioural, treatment-related, healthcare-system, environmental, socioeconomic, and biological factors may all contribute. The principal contribution of this study is therefore the identification of the information boundary between routine programme surveillance and mechanistic investigation and the development of a programme-to-molecular translational framework for addressing that boundary.
Integrating routine epidemiological surveillance with prospectively collected genomic, transcriptomic, immunological, pharmacological, and functional molecular data may provide a more comprehensive understanding of the heterogeneity of TB treatment responses and antimicrobial resistance. Such integration could subsequently support the development and validation of molecular biomarkers, biologically informed predictive models, improved risk stratification, and precision-oriented approaches to TB research. Within this framework, gene expression profiling and emerging experimental approaches, such as photobiomodulation, represent hypothesis-generating directions that require direct laboratory validation, rather than interventions supported by the present epidemiological analysis. Future prospective and mechanistic studies are therefore required to determine whether the biological pathways identified as evidence gaps contribute to treatment response and whether targeting these pathways has diagnostic, prognostic, or therapeutic relevance for drug-resistant TB.

Abbreviation

Abbreviation Full term
AI Artificial Intelligence
AP Average Precision
aRR Adjusted Risk Ratio
CI Confidence Interval
DR-TB Drug-Resistant Tuberculosis
HIV Human Immunodeficiency Virus
IQR Interquartile Range
LR Logistic Regression
MDR-TB Multidrug-Resistant Tuberculosis
ML Machine Learning
OR Odds Ratio
PBM Photobiomodulation
RF Random Forest
RNA Ribonucleic Acid
RNA-seq RNA Sequencing
TB Tuberculosis
WHO World Health Organization
WGS Whole-Genome Sequencing
WSU Walter Sisulu University

Author Contributions

Luzuko Mkono: Conceptualization, Investigation, Writing – Original Draft Preparation. Ntandazo Dlatu: Conceptualization, Methodology, Writing – Review & Editing. Ncomeka Sineke: Conceptualization, Methodology, Writing – Review & Editing. Mandlenkosi Manika: Methodology, Writing – Review & Editing. Siphosihle Conham: Methodology, Writing – Review & Editing. Bulela Sonka: Methodology, Writing – Review & Editing. Nande Ndamase: Methodology, Writing – Review & Editing. Hloniphani Guma: Methodology, Writing – Review & Editing. Lubabalo Macingwana: Writing – Review & Editing. Lindiwe Modest Faye: Conceptualization, Methodology, Supervision, Writing – Review & Editing. All authors reviewed and approved the final version of the manuscript and agree to be accountable for their respective contributions to the work.

Funding

The author(s) declared that no financial support was received for this work and/or its publication. Walter Sisulu University School of Pathology funds for study conduct and School of Public Health funds for APC.

Data Availability Statement

In accordance with institutional and national data protection regulations, the individual-level patient data utilized in this study are restricted and not openly accessible. These restrictions are imposed by the Walter Sisulu University Research Ethics and Biosafety Committee and the Eastern Cape Department of Health to safeguard patient privacy, given the sensitive nature of TB treatment records. Researchers seeking access to the de-identified dataset may submit a formal request to the corresponding author. All requests will be evaluated on a case-by-case basis, and access will be granted only upon receipt of written approval from the relevant ethics review boards and execution of a standard data-sharing agreement that ensures confidentiality and appropriate use of the data. The authors are committed to transparency and will fully facilitate data sharing permitted by ethical and legal frameworks.

Acknowledgments

The authors gratefully acknowledge the Honours and Master’s students from the School of Pathology and Laboratory Medicine, the TB Research Group, and Walter Sisulu University for their valuable support and assistance during the data collection process. The authors also thank all colleagues and collaborators whose contributions, guidance, and support facilitated the successful completion of this study and preparation of the manuscript.

Competing Interests

The authors declare no competing interests, financial or non-financial, that could have influenced the work reported in this manuscript.

Generative AI Statement

The authors declare that generative artificial intelligence (AI) was not used in the conceptualization, analysis, interpretation of data, or preparation of the scientific content of this manuscript.

References

  1. Farnia, P.; Velayati, A.A.; Ghanavi, J.; Farnia, P. An ongoing global threat. In Proteins in Mycobacterium Tuberculosis: Functions and Therapeutic Advances; Springer Nature Switzerland: Cham, Switzerland, 2025; pp. 1–31. [Google Scholar]
  2. Lee, H.; Kim, J.; Kim, J.; Park, Y.J. Review of the global burden of tuberculosis in 2023: Insights from the WHO Global Tuberculosis Report 2024. Public Health Wkly. Rep. 2025, 18, 55. [Google Scholar]
  3. Yapa, H.M.; et al. Drug-resistant tuberculosis: A priority pathogen for enhanced public health research and practice. Clin. Microbiol. Rev. 2025, 38, e00064-25. [Google Scholar] [CrossRef] [PubMed]
  4. Mahlangu, J.; Diop, S.; Lavin, M. Diagnosis and treatment challenges in lower resource countries: State-of-the-art. Haemophilia 2024, 30, 78–85. [Google Scholar] [CrossRef] [PubMed]
  5. Faye, L.M.; Hosu, M.C.; Apalata, T. Drug-resistant tuberculosis in rural Eastern Cape, South Africa: A study of patients’ characteristics in selected healthcare facilities. Int. J. Environ. Res. Public Health 2024, 21, 1594. [Google Scholar] [CrossRef] [PubMed]
  6. Litvinjenko, S.; Magwood, O.; Wu, S.; Wei, X. Burden of tuberculosis among vulnerable populations worldwide: An overview of systematic reviews. Lancet Infect. Dis. 2023, 23, 1395–1407. [Google Scholar] [CrossRef] [PubMed]
  7. Ponson, T.A. A Comparative Study of Predictive Modeling in Human Immunodeficiency Virus/Tuberculosis Co-Infection Antiretroviral Treatment Adherence in Mozambique . Doctoral Dissertation, Mercer University, Macon, GA, USA, 2025. [Google Scholar]
  8. Gopalaswamy, R.; Subbian, S. The power of resistance: Mechanisms of antimicrobial resistance in Mycobacterium tuberculosis and its impact on tuberculosis management. Clin. Microbiol. Rev. 2026, 39, e00194-25. [Google Scholar] [CrossRef] [PubMed]
  9. Dookie, N.; et al. The changing paradigm of drug-resistant tuberculosis treatment: Successes, pitfalls, and future perspectives. Clin. Microbiol. Rev. 2022, 35, e00180-19. [Google Scholar] [CrossRef] [PubMed]
  10. Ahmad, M.; Aduru, S.V.; Smith, R.P.; Zhao, Z.; Lopatkin, A.J. The role of bacterial metabolism in antimicrobial resistance. Nat. Rev. Microbiol. 2025, 23, 439–454. [Google Scholar] [CrossRef] [PubMed]
  11. Napier, G.; Campino, S.; Phelan, J.E.; Clark, T.G. Large-scale genomic analysis of Mycobacterium tuberculosis reveals extent of target and compensatory mutations linked to multidrug-resistant tuberculosis. Sci. Rep. 2023, 13, 623. [Google Scholar] [CrossRef] [PubMed]
  12. Acierno, C.; et al. Metabolic rewiring of bacterial pathogens in response to antibiotic pressure—A molecular perspective. Int. J. Mol. Sci. 2025, 26, 5574. [Google Scholar] [CrossRef] [PubMed]
  13. Faye, L.M.; Hosu, M.C.; Apalata, T. Drug-resistant tuberculosis hotspots in Oliver Reginald Tambo District Municipality, Eastern Cape, South Africa. Infect. Dis. Rep. 2024, 16, 1197–1213. [Google Scholar] [CrossRef] [PubMed]
  14. Ndamase, N.; Faye, L.M.; Dlatu, N.; Apalata, T.; Hosu, M.C. Hierarchical risk profiles in tuberculosis treatment outcomes: The role of drug resistance, age, and socio-economic factors. Microbiol. Res. 2026, 17, 42. [Google Scholar] [CrossRef]
  15. Zhu, J.J.; Yang, M.; Ren, Z.J. Machine learning in environmental research: Common pitfalls and best practices. Environ. Sci. Technol. 2023, 57, 17671–17689. [Google Scholar] [CrossRef] [PubMed]
  16. McBride, J.M.; Eckmann, J.P.; Tlusty, T. General theory of specific binding: Insights from a genetic-mechano-chemical protein model. Mol. Biol. Evol. 2022, 39, msac217. [Google Scholar] [CrossRef] [PubMed]
  17. Mkono, L.; et al. Epidemiology and treatment outcomes of TB–HIV co-infection in a rural South African setting: An exploratory analysis of retreatment and contextual factors. Front. Public Health 2026, 14, 1792518. [Google Scholar] [CrossRef] [PubMed]
  18. Nasir, M.; Summerfield, N.S.; Carreiro, S.; Berlowitz, D.; Oztekin, A. A machine learning approach for diagnostic and prognostic predictions, key risk factors and interactions. Health Serv. Outcomes Res. Methodol. 2025, 25, 1–28. [Google Scholar] [CrossRef] [PubMed]
  19. Messina, F.; et al. Molecular exploration of host-pathogen interactions in severe Pseudomonas aeruginosa infection through a multi-level data integration approach. Front. Med. 2025, 12, 1600509. [Google Scholar] [CrossRef] [PubMed]
  20. World Health Organization. Global Tuberculosis Report 2025; World Health Organization: Geneva, Switzerland, 2025. [Google Scholar]
  21. Oga-Omenka, C.; et al. Implementation of digital tuberculosis information systems: Perspectives from 10 high TB burden countries. BMC Infect. Dis. 2026, 26, 1030. [Google Scholar] [CrossRef] [PubMed]
  22. Ajudua, F.I.; Mash, R.J. Implementing active surveillance for TB: A descriptive survey of healthcare workers in the Eastern Cape, South Africa. Afr. J. Prim. Health Care Fam. Med. 2024, 16, 1–12. [Google Scholar] [CrossRef] [PubMed]
  23. Fiorillo, R.M.; et al. Optimizing use of data and evidence-related tools to complement routine surveillance systems for tuberculosis program planning. Front. Tuberc. 2025, 3, 1622167. [Google Scholar] [CrossRef]
  24. Ajudua, F.I.; Mash, R.J. Implementing active surveillance for tuberculosis: The experiences of healthcare workers at four sites in two provinces in South Africa. S. Afr. Fam. Pract. 2022, 64, 5514. [Google Scholar] [CrossRef] [PubMed]
  25. Limenh, L.W.; et al. Tuberculosis treatment outcomes and associated factors among tuberculosis patients treated at healthcare facilities of Motta Town, Northwest Ethiopia: A five-year retrospective study. Sci. Rep. 2024, 14, 7695. [Google Scholar] [CrossRef] [PubMed]
  26. Nxumalo, E.L.; Sineke, N.; Dlatu, N.; Apalata, T.; Faye, L.M. Treatment outcomes of tuberculosis in the Eastern Cape: Clinical and socio-demographic predictors from two rural clinics. Int. J. Environ. Res. Public Health;DOI 2025, 22, 1804. [Google Scholar] [CrossRef] [PubMed]
  27. Faye, L.M.; Hosu, M.C.; Dlatu, N.; Iruedo, J.; Apalata, T. Predicting treatment adherence in patients with drug-resistant tuberculosis: Insights from socioeconomic, demographic, and clinical factors of patients in the rural Eastern Cape. Front. Tuberc.;DOI 2025, 3. [Google Scholar] [CrossRef]
  28. Teferi, M.Y.; et al. Tuberculosis treatment outcome and predictors in Africa: A systematic review and meta-analysis. Int. J. Environ. Res. Public Health 2021, 18, 10678. [Google Scholar] [CrossRef] [PubMed]
  29. Zenatti, G.; et al. High variability in tuberculosis treatment outcomes across 15 health facilities in a semi-urban area in central Ethiopia. J. Clin. Tuberc. Other Mycobact. Dis. 2023, 30, 100344. [Google Scholar] [CrossRef] [PubMed]
  30. Goossens, S.N.; Sampson, S.L.; Van Rie, A. Mechanisms of drug-induced tolerance in Mycobacterium tuberculosis. Clin. Microbiol. Rev. 2020, 34, e00141-20. [Google Scholar] [CrossRef] [PubMed]
  31. Kengo, A. Nonlinear Mixed-Effects Modelling of Drug-Drug Interactions between Antiretroviral Therapy and Tuberculosis Treatment. University of Cape Town, Cape Town, South Africa, 2025; Available online; University of Cape Town Institutional Repository; (accessed on 6 August 2026).
  32. Mboowa, G. Reimagining tuberculosis control in the era of genomics: The case for global investment in Mycobacterium tuberculosis genomic surveillance. Pathogens 2025, 14, 975. [Google Scholar] [CrossRef] [PubMed]
  33. Warner, D.F.; Barczak, A.K.; Gutierrez, M.G.; Mizrahi, V. Mycobacterium tuberculosis biology, pathogenicity and interaction with the host. Nat. Rev. Microbiol. 2025. [Google Scholar] [CrossRef] [PubMed]
  34. Adeboye, A.; Georgeleen, O.; Olusesan, A.; Colin, N. Machine learning prediction of tuberculosis mortality: A comparative analysis of random survival forest and Cox regression models. BMC Infect. Dis. 2026, 26, 829. [Google Scholar] [CrossRef] [PubMed]
  35. Chinagudaba, S.N.; Gera, D.; Dasu, K.K.V.; Singarajpure, A.; Chadda, V.K. Predictive analysis of tuberculosis treatment outcomes using machine learning: A Karnataka TB data study at a scale. arXiv 2024, arXiv:2403.08834. [Google Scholar]
  36. Lee, T.; et al. Predicting unfavorable tuberculosis outcomes using machine learning: A prospective cohort. Trop. Med. Health 2026, 54, 114. [Google Scholar] [CrossRef] [PubMed]
  37. Chen, L.; et al. The performance of VCS (volume, conductivity, light scatter) parameters in distinguishing latent tuberculosis and active tuberculosis by using machine learning algorithm. BMC Infect. Dis. 2023, 23, 881. [Google Scholar] [CrossRef] [PubMed]
  38. Lecca, P. Machine learning for causal inference in biological networks: Perspectives of this challenge. Front. Bioinform. 2021, 1, 746712. [Google Scholar] [CrossRef] [PubMed]
  39. Lawrence, E.; et al. Understanding biology with machine learning: Compression, intelligibility, and dependency. Artif. Intell. Life Sci. 2026, 100161. [Google Scholar] [CrossRef]
  40. Greener, J.G.; Kandathil, S.M.; Moffat, L.; Jones, D.T. A guide to machine learning for biologists. Nat. Rev. Mol. Cell Biol. 2022, 23, 40–55. [Google Scholar] [CrossRef] [PubMed]
  41. Gul, M.; et al. Genetic characterization of first-line drug-resistance mutations in multidrug-resistant Mycobacterium tuberculosis. Pathogens 2026, 15, 455. [Google Scholar] [CrossRef] [PubMed]
  42. Tiwari, S.; Singh, A.K.; Acharya, A.; Ghosh, J.; Singh, R.P. Multidrug-resistant tuberculosis: A comprehensive review of pathogenesis, drug resistance, current treatment and future prospects. Arch. Microbiol. 2026, 208, 395. [Google Scholar] [CrossRef] [PubMed]
  43. Poonawala, H.; et al. Transcriptomic responses to antibiotic exposure in Mycobacterium tuberculosis. Antimicrob. Agents Chemother. 2024, 68, e01185-23. [Google Scholar] [CrossRef] [PubMed]
  44. Hazra, D.; et al. Impact of whole-genome sequencing of Mycobacterium tuberculosis on treatment outcomes for MDR-TB/XDR-TB: A systematic review. Pharmaceutics 2023, 15, 2782. [Google Scholar] [CrossRef] [PubMed]
  45. Wynn, E.A.; et al. Transcriptional adaptation of drug-tolerant Mycobacterium tuberculosis in mice. bioRxiv 2023. [Google Scholar] [CrossRef] [PubMed]
  46. Culviner, P.; et al. Evolution of. Mycobacterium tuberculosis transcription regulation is associated with increased transmission and drug resistance 2025. [Google Scholar] [CrossRef]
  47. Long, Y.; et al. Overexpression of efflux pump genes is one of the mechanisms causing drug resistance in Mycobacterium tuberculosis. Microbiol. Spectr. 2024, 12, e02510-23. [Google Scholar] [CrossRef] [PubMed]
  48. Li, S.; et al. CRISPRi chemical genetics and comparative genomics identify genes mediating drug potency in Mycobacterium tuberculosis. Nat. Microbiol. 2022, 7, 766–779. [Google Scholar] [CrossRef] [PubMed]
Figure 1. Translational framework linking routine tuberculosis programme surveillance with future molecular research.
Figure 1. Translational framework linking routine tuberculosis programme surveillance with future molecular research.
Preprints 228929 g001
Table 1. Associations between routinely collected programme variables and TB treatment outcomes. 
Table 1. Associations between routinely collected programme variables and TB treatment outcomes. 
Variable Univariable OR (95% CI) Multivariable OR (95% CI) p-value Independent association Interpretation
Age (per year) 1.01 (0.99–1.03) 0.21 No No significant independent association
Sex (male vs female) 1.08 (0.66–1.77) 1.08 (0.66–1.77) 0.77 No No significant independent association
HIV status 0.80 (0.49–1.29) 0.36 No Not independently associated with treatment success
Pulmonary TB 1.12 (0.71–1.78) 0.62 No No significant independent association
Previous TB treatment 0.48 (0.29–0.78) 0.64 (0.38–1.09) 0.10 No Univariable association attenuated after adjustment
Education level 1.07 (0.67–1.71) 0.78 No No significant independent association
Income source 1.84 (0.92–3.67) 1.36 (0.74–2.49) 0.32 No Univariable association attenuated after adjustment
Employment status 1.22 (0.72–2.05) 0.46 No No significant independent association
OR, odds ratio; CI, confidence interval; TB, tuberculosis. The reported p-values correspond to the multivariable logistic regression estimates.
Table 2. Comparative performance and interpretation of the predictive models. 
Table 2. Comparative performance and interpretation of the predictive models. 
Model/analysis Principal result Interpretation Main limitation
Logistic regression AP = 0.810 Provided conventional prediction and estimates of independent associations Restricted to variables available in routine programme data
Random Forest AP = 0.907 Higher predictive performance than logistic regression Does not establish causal or biological mechanisms
RF feature importance Age = 0.555; sex = 0.066; HIV = 0.053 Identified variables contributing to prediction Feature importance should not be interpreted as an independent causal effect
Multivariable logistic regression No listed variable remained a statistically significant independent predictor at p < 0.05 Limited evidence of independent associations among the measured variables Unmeasured determinants could not be evaluated
Poisson regression HIV: aRR = 0.95 (95% CI 0.66–1.37); previous TB treatment: aRR = 0.87 (95% CI 0.69–1.10) Findings were consistent with the multivariable logistic analysis Restricted to routinely recorded variables
AP, average precision; RF, Random Forest; LR, logistic regression; aRR, adjusted risk ratio; CI, confidence interval.
Table 3. Biological evidence gaps identified in routine TB programme surveillance. 
Table 3. Biological evidence gaps identified in routine TB programme surveillance. 
Biological domain Captured in routine dataset? What could not be determined Potential future approach
Resistance-associated genomic variation No Presence and distribution of resistance-associated variants Whole-genome sequencing
Bacterial gene expression No Differential expression associated with treatment exposure or response RNA sequencing/transcriptomics
Transcriptional regulation No Regulatory responses during antimicrobial stress Transcriptomic and functional assays
Oxidative stress response No Bacterial responses to oxidative stress Gene-expression and functional studies
Efflux activity No Contribution of efflux mechanisms to drug tolerance/resistance Molecular and functional assays
Dormancy and persistence pathways No Biological mechanisms associated with bacterial persistence Transcriptomic and functional studies
Metabolic adaptation No Metabolic states associated with antimicrobial exposure or persistence Metabolic and molecular profiling
Host immune responses No Relationship between host immunity and treatment response Immunological/host-response profiling
Host–pathogen interactions No Molecular interactions associated with heterogeneous outcomes Integrated host–pathogen studies
Transcriptomic biomarkers No Molecular signatures associated with treatment response RNA sequencing and biomarker validation
PBM-associated molecular responses No Whether PBM alters bacterial molecular or physiological responses Controlled laboratory experiments
Table 4. Programme-to-molecular translational research priorities. 
Table 4. Programme-to-molecular translational research priorities. 
Priority Epidemiological evidence gap Proposed molecular investigation Potential future contribution
1. Genomic characterisation Resistance-associated genomic information unavailable Whole-genome sequencing of M. tuberculosis isolates Characterisation of resistance-associated variants
2. Transcriptomic profiling Gene-expression responses unavailable RNA sequencing under defined treatment conditions Identification of treatment-associated transcriptional responses
3. Persistence and stress responses Physiological adaptation not measured Functional molecular and transcriptomic assays Characterisation of persistence and stress-response pathways
4. Host–pathogen biology Host molecular responses unavailable Host-response and integrated host–pathogen profiling Improved understanding of heterogeneous treatment responses
5. Integrated prediction Current models contain routine programme variables only Combine validated molecular biomarkers with epidemiological variables in future cohorts Test whether multimodal data improve prediction and risk stratification
6. Experimental PBM research No PBM exposure or molecular response data available Controlled in vitro PBM experiments with molecular endpoints Determine whether PBM produces measurable biological effects before considering therapeutic applications
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.