Preprint
Review

This version is not peer-reviewed.

Generative Artificial Intelligence in Health Economics and Outcomes Research: A Critical Review and a Capability–Governance Framework for Health Technology Assessment in Low- and Middle-Income Countries

Submitted:

26 August 2026

Posted:

27 August 2026

You are already at the latest version

Abstract
Background. Health technology assessment (HTA) is how health systems turn a fixed budget into coverage decisions they can defend. The evidence it runs on is slow and expensive to produce, and generative artificial intelligence (AI) has been proposed as a way out of that constraint. Between 2022 and 2026 the field acquired a dense layer of methodological and reporting guidance. Almost all of it was written by, and for, well-resourced high-income agencies. Objectives. This review has three aims: to appraise what is actually known about generative AI performance across the health economics and outcomes research (HEOR) value chain; to explain why the available governance instruments do not transfer cleanly to resource-constrained settings; and to propose a framework that such agencies can use. Methods. Critical narrative review, with a structured search of PubMed and MEDLINE to August 2026, supplemented by targeted retrieval of agency and multilateral guidance and by backward citation tracking. Sources were appraised for transferability rather than pooled. Findings. Reported performance is strong, but it clusters in bounded tasks where a human can check the answer. Screening and extraction accuracy in HEOR-specific systems is high, and prompt optimisation improves recall across model families for a few dollars. The tasks that actually move an incremental cost-effectiveness ratio, including comparative effect estimation, extrapolation, parameter selection and model structuring, remain largely unvalidated. Reference fabrication, non-determinism, proxy-induced bias and privacy exposure are unresolved. Existing instruments (PALISADE, CHEERS-AI, ELEVATE-GenAI, the NICE position statement, and the RAISE-aligned position of the evidence-synthesis collaborations) govern reporting well, but they assume institutional preconditions, including dense registry data, local value sets, deep analytic benches and an explicit cost-effectiveness threshold, that most low- and middle-income country (LMIC) agencies do not have. Conclusions. I propose REACH-AI: five readiness domains (Readiness, Evidence-fitness, Accountability, Capability, Harm mitigation) coupled to a three-tier task–risk stratification in which oversight is set by how reversible the error is, not by how sophisticated the tool is. The claim underneath the framework is a simple one. Automation speeds up whatever decision rule is already in place, so an agency without a locally estimated threshold and a national reference case will use these tools to reach undefined conclusions faster.
Keywords: 
;  ;  ;  ;  ;  ;  ;  

Introduction

Every health system spends within a limit. That means every decision to fund a new technology is also, silently, a decision to stop funding something else. Health technology assessment exists to drag that second decision into the open, so that displacement happens deliberately and can be argued about, rather than happening by default and going unnoticed. Over four decades the discipline has built settled reporting conventions (Husereau et al., 2022) and, more recently, a real institutional presence outside the high-income world (MacQuilkan et al., 2018).
The awkward part is that this expansion is arriving just as the discipline's own production capacity runs short. Published clinical evidence keeps growing. The range of things requiring appraisal has widened well beyond pharmaceuticals, to devices, digital tools and whole health programmes (Prinja et al., 2021). And the window in which a decision is still useful has narrowed. A single systematic review supporting an HTA submission takes months; a de novo economic model takes longer.
When assessment capacity cannot keep up with the flow of candidate technologies, the consequence is not that decisions get postponed. Decisions get made anyway, without assessment. That is the failure state worth keeping in view, because it sets the comparator for everything that follows.
Generative artificial intelligence has been offered as a structural answer rather than a marginal efficiency. Large language models can draft search strategies, screen abstracts, pull structured data out of unstructured text, write analytic code and summarise deliberative material. Between them, those tasks account for a large share of the labour buried in an HTA. The professional response has been unusually quick: in roughly three years the field produced a good-practices checklist for machine learning (Padula et al., 2022), a reporting extension for economic evaluations of AI-enabled technologies (Elvidge, Hawksworth, et al., 2024), a policy assessment for HTA agencies (Fleurence et al., 2025c), a working taxonomy of applications (Fleurence et al., 2025a), reporting guidelines for large language model–assisted studies (Fleurence et al., 2025b), a formal position from a national HTA body (National Institute for Health and Care Excellence [NICE], 2024), and a joint statement from the major evidence-synthesis collaborations (Flemyng et al., 2025). Multilateral ethical guidance arrived alongside it (World Health Organization [WHO], 2024).
By the standards of most scientific fields that is an impressive governance response. It is also, once you look at who wrote it, a geographically narrow one. The agencies behind these documents work with linked administrative data, national utility value sets, established reference cases, published willingness-to-pay thresholds and analytic teams numbering in the dozens. Their instruments are correspondingly reporting-centric. They specify what must be disclosed when AI is used, on the quiet assumption that an institution capable of disclosing is also capable of checking. That assumption carries a great deal of weight, and it does not hold everywhere.
My concern is not that LMIC agencies will be slow to adopt these tools. The economics run the other way. Where an agency has three analysts rather than thirty, a system that screens ten thousand abstracts overnight is far more attractive than it would be in a well-staffed unit in London or Ottawa. The concern is that adoption will outpace the verification infrastructure needed to notice when the tool is wrong, and that the resulting errors will be quiet ones, because the same capacity shortfall that makes automation attractive is what prevents independent checking. An assessment that is fast, internally consistent, professionally presented and materially wrong is worse than no assessment, since it carries the authority of the process without the warrant.
So this review does three things. It appraises what is actually known, as distinct from what is claimed, about generative AI across the HEOR value chain, separating tasks with validation evidence from those without. It sets out the specific structural reasons why high-income governance instruments do not transfer intact, using India as the working case, since India is one of the few LMICs with an institutionalised HTA function and a published evidence base against which claims can be checked rather than merely asserted. And it proposes REACH-AI, a capability–governance framework, together with a research agenda for testing it.
What I have deliberately not written is another reporting checklist. Adequate checklists already exist, and adding one would not help. The gap is at an earlier point, in the question a director of a small agency has to answer before any checklist becomes relevant: given what this institution actually has, in data, in skills, in spare capacity for oversight, which of these tasks can it responsibly hand over, and on what conditions?

Methods

Review Approach

This is a critical review in the established sense: it aims at interpretive synthesis and conceptual contribution, and it appraises sources for what they contribute to an argument rather than pooling them into a summary estimate. It is not a systematic review and I make no claim to exhaustiveness. The search was nonetheless structured, and I report it in full so that readers can judge whether it was adequate and so that others can extend it.

Search Strategy and Source Selection

I searched PubMed and MEDLINE from inception to 23 August 2026, combining terms across three concept axes: the technology (artificial intelligence, machine learning, large language models, generative artificial intelligence, foundation models), the discipline (health economics, outcomes research, health technology assessment, economic evaluation, cost-effectiveness, evidence synthesis, systematic review) and the setting (low- and middle-income countries, India, universal health coverage, priority setting). Each axis was searched separately as well as in combination, because the three-way intersection is thin and searching only there would have excluded the foundational work in each strand.
Three supplementary strategies filled predictable gaps. Agency and multilateral guidance is often not indexed bibliographically, so I retrieved it directly from institutional repositories. I tracked citations backward from the ISPOR working group reports and the NICE position statement, which between them function as hubs for this literature. And I tracked forward through related-article algorithms to pick up empirical validation studies that the formal guidance had not yet caught up with.
I retained sources that reported empirical performance of AI methods in an HEOR or evidence-synthesis task, that articulated a methodological, ethical or governance position from a recognised agency or collaboration, that established methodological standards for economic evaluation, or that characterised the structural conditions of HTA in resource-constrained settings. Commentary was kept only where it came from a body with standard-setting authority. Table 1 sets out the full search strategy, the supplementary strategies and the inclusion and exclusion criteria. Every citation in this article was checked against indexed bibliographic metadata for author list, title, journal, volume, issue, pagination and digital object identifier.

Constraints on Inference

Two features of this literature limit what can honestly be concluded from it, and both are worth stating before the findings rather than after.
The evidence base is young. Most validation studies appeared between 2024 and 2026, several describe a single team's implementation of a system that team also built, and independent replication is close to absent. On top of that, publication incentives here favour demonstrations that something works over reports that it did not, so the aggregate picture is almost certainly flattering. Both points argue for the risk-stratified posture I develop in Section 7, rather than for either enthusiasm or blanket refusal.

Generative Artificial Intelligence Across the HEOR Value Chain

Generative AI does not act on HTA as a single undifferentiated thing. It acts on particular tasks in a sequence, and what happens when it fails depends heavily on where in that sequence the failure occurs. Figure 1 sets out the value chain, the points at which generative AI has been proposed or demonstrated, and the gradient that organises the rest of this review: the further downstream an output sits, the harder an undetected error is to reverse. A badly built search string produces a review that is visibly thin and gets sent back. A misextracted hazard ratio produces a model that is confidently wrong and gets approved.

Evidence Synthesis and Systematic Literature Review

Evidence synthesis is where generative AI has been tested most heavily, and for good reason. Its tasks are bounded, high in volume, and easy to score against a human reference standard.
Table 2 summarises the studies discussed in this section, their designs, their reported performance and the principal caveat attaching to each. The numbers are genuinely impressive. Li et al. (2025) built a five-module human-in-the-loop system for HTA submissions and reported mean screening sensitivity of 90%, accuracy of 89% and Cohen's kappa of 0.71 against human reviewers, with an F1 score of 93 for data extraction. Notably, the design needed no manually annotated training data at all, relying instead on in-context learning. Lee et al. (2025) evaluated an eight-module agentic framework across two clinical areas and reported screening F1 scores between 0.917 and 0.977, kappa values of 0.844 to 0.906 for risk-of-bias assessment, and F-scores from 0.96 to 0.998 for extraction of patient characteristics, outcomes, economic model parameters and cost-effectiveness data. Reason et al. (2024) demonstrated scale rather than accuracy alone, pulling population, intervention, comparator and outcome elements from 682,667 abstracts in under three hours for roughly US$3,390 in total, with human verification of 350 randomly selected abstracts confirming accurate and complete extraction in 98% of cases.
Two findings from Józwiak et al. (2026) deserve separate attention, because they overturn assumptions that would otherwise drive procurement in exactly the wrong direction. Applying automated reflective prompt optimisation under an asymmetric loss function that penalises false negatives (which is the correct specification for screening, where missing an eligible study costs far more than admitting an irrelevant one), they improved recall across all nine models tested, by between 3.7 and 37.1 percentage points, for a total optimisation cost of US$6.36. Their best-performing model reached 91% accuracy at roughly one twenty-fifth the per-abstract cost of a substantially worse alternative. Price, in other words, did not predict performance. And a prompt optimised on a single open-source model generalised to all nine without retraining. For an agency weighing an expensive proprietary service against an open-weight model it could host itself, that second result matters a great deal.
The optimism needs bounding in three respects, though.
First, these accuracies come from purpose-built architectures with retrieval, verification loops and human checkpoints in them. They are not what a general-purpose chatbot produces when handed a research question. Second, performance is reported against internal reference standards built by the same teams that built the systems, and external validation on unfamiliar corpora is scarce. Third, and most relevant here, the benchmark corpora are overwhelmingly English-language literature from high-income settings. That is not the corpus an Indian or sub-Saharan African agency has to search when the relevant evidence sits in regional journals, national programme reports and grey literature.

Real-World Data and Parameterisation

The second domain is real-world data: turning routinely collected clinical and administrative records into parameters a model can use. Generative AI has been proposed for phenotyping from unstructured notes, cohort construction, harmonisation across heterogeneous sources, and retrieval of cost and utility parameters (Fleurence et al., 2025a).
The demand is real and rising. Real-world evidence now turns up routinely in regulatory and HTA submissions, and its acceptance is contested in ways that bear directly on whether patients get access. Zong et al. (2025) reviewed oncology approvals across the European Medicines Agency and three national HTA bodies and found real-world evidence used mainly as an external control arm for indirect comparison or contextualisation, and mostly rejected on methodological grounds, with acceptance diverging not only between the regulator and the HTA bodies but among the HTA bodies themselves. Geissler et al. (2023) argue that in precision oncology this produces a structural impasse: where molecular stratification leaves too few eligible patients for a conventional randomised trial, insisting on randomised evidence forecloses access to therapies for which observational data are the best obtainable evidence. Sarri and Hernandez (2024) mapped 46 guidance documents across regulatory and HTA agencies and found broad agreement on design and quality principles sitting alongside persistent inconsistency in terminology and appraisal tools, a fragmentation that imposes real costs on anyone submitting in more than one jurisdiction.
Automation meets this contested terrain in a way that should give pause. The methodological expectations for observational evidence to support a causal claim are demanding, and rightly so: a protocol and statistical analysis plan registered before any inferential analysis, formal identification of the causal estimand, systematic enumeration of confounders through explicit causal graphs, target trial emulation to control selection bias, and negative-control and falsification analyses to detect residual confounding (Cucherat et al., 2025). Those are acts of causal reasoning about a particular clinical situation. A language model can accelerate the mechanical work around them, assembling covariate lists, drafting analysis code, retrieving candidate parameters. It cannot supply the domain judgement that decides whether a given variable is a confounder, a mediator or a collider. Automating the surface of a causal analysis while leaving its logic unexamined gets you evidence that is faster to produce and no more believable.

Health Economic Modelling

The third domain is economic modelling itself, where generative AI has been proposed across the whole sequence from conceptualisation through parameterisation, coding, validation and reporting (Fleurence et al., 2025c).
There are antecedents in the machine learning literature. The ISPOR machine learning task force identified five areas where such methods could improve HEOR, namely cohort selection, feature and covariate identification, predictive analytics, causal inference through targeted maximum likelihood or double-debiased estimation, and the reduction of structural, parameter and sampling uncertainty in cost-effectiveness analysis, while warning that opacity in feature selection and prediction, particularly under unsupervised conditions, shifts risk onto the decision maker (Padula et al., 2022). Glynn et al. (2024) showed the integration concretely, combining machine learning estimates of treatment-effect heterogeneity with a decision-analytic model and using policy-tree algorithms to define subgroups from model output; both the incremental net health benefit of a one-size-fits-all policy and the gains available from stratification changed under this approach. Hamaya et al. (2025) laid out the corresponding estimation frameworks for individualised treatment effects and note that doubly robust and R-learner approaches are the better choices in high-dimensional settings.
For generative AI in modelling, the situation differs in kind rather than degree, and the distinction is worth making carefully.
Code generation is verifiable. A model either runs, reproduces a known benchmark and survives internal validation, or it does not, and the analyst finds out quickly. Structural conceptualisation is not verifiable in anything like the same way. The choice of health states, which transitions are permitted, cycle length, the comparator set, the time horizon: these encode a clinical theory of the disease and a normative theory of what counts as a consequence worth counting. An error here does not announce itself. It produces a well-behaved model of the wrong problem, and the model will pass every internal validation check you throw at it, because internal validation tests whether the model does what you told it to do.
I could not find a single published study validating generative AI–proposed model structures against expert-derived structures with respect to their effect on the incremental cost-effectiveness ratio. If that gap has been filled since I searched, I would be glad to be corrected. As things stand it is the most consequential absence in this literature, and it sits precisely where the technology's advocates locate the greatest promise.

Appraisal, Dossier Production and Deliberation

The last domain covers assessment reports, reimbursement dossiers and support for deliberative processes. Drafting, cross-jurisdictional adaptation, plain-language summarising for lay committee members and consistency checking against reporting standards are all plausible uses, and the pressures behind them are well documented: only around half of newly approved medicines receive unrestricted reimbursement recommendations, and mismatch between regulatory and HTA evidence expectations is a leading cause of delay (Cowie et al., 2022).
Two cautions apply, and the first is subtle enough to be worth spelling out. Cross-jurisdictional adaptation is a domain where fluency actively misleads. Asked to adapt a dossier from one jurisdiction to another, a model will produce text reading as though the local reference case, threshold, perspective and discount rate have all been correctly applied, because the surface features of these documents are highly patterned and easy to imitate. Whether the substantive parameters were in fact changed is a separate question, and the fluency of the output is precisely what makes that question easy to skip.
The second caution concerns deliberation. An appraisal committee is not an information-processing device. Committees weigh severity, unmet need, equity and uncertainty in ways that are contested by design, and part of what makes their conclusions legitimate is the visible fact that identifiable people exercised judgement and can be asked to account for it. Summarisation that smooths real disagreement into consensus prose does not help deliberation. It hides the tensions the committee exists to resolve.

Cross-Cutting Methodological and Ethical Risks

Five failure modes run across every application domain. I treat them together because mitigating them is an institutional responsibility rather than a task-specific one, and because an agency that handles them well in one workflow and badly in another has not really handled them at all.

Fabrication and Referential Integrity

The best-documented failure is the confident production of references that do not exist. McGowan et al. (2023) ran two widely used systems through a psychiatry literature search and found that of 35 citations produced, two were real. Twelve resembled genuine manuscripts but carried the wrong authors, journal or year. The remaining twenty-one were composites, assembled out of several existing papers. Their argument that the usual term ‘hallucination’ is a misnomer seems right to me: these are not perceptual errors but fabrications, produced by a system optimised for plausibility rather than truth, and the word ‘hallucination’ makes the failure sound involuntary in a way that obscures what is actually happening.
The problem has not gone away in systems marketed for clinical and academic use. Bridges (2024) compared a large language model against a purpose-built diagnostic decision support system across 201 cases and found that although the model produced 145 correct references out of 165 (87.9%), only 52 of the accompanying digital object identifiers resolved correctly (31.5%). The pattern is worth dwelling on. The reference resembles a real paper closely enough to survive a glance, while the one element that would allow verification is the element that fails.
Architectural mitigation works. Gorenshtein et al. (2025) built a multi-agent system that searches MEDLINE directly through the PubMed application programming interface and uses bidirectional inter-agent communication to confirm citations, reaching 99.82% citation accuracy and 96.81% referencing accuracy (the latter measuring whether in-text claims match source metadata), while drawing exclusively on peer-reviewed sources, against a comparator that returned 35.60% non-academic content. The general lesson goes beyond citations: reliability comes from grounding generation in a verifiable retrieval layer, not from a better generator.
For an agency the operational implication is blunt. Any AI-assisted reference must resolve to an indexed record before it enters a submission, and that check has to be automated rather than left to reviewer diligence, because this particular failure mode is calibrated to survive exactly the kind of visual inspection a busy reviewer performs at the end of a long day.

Reproducibility and Non-Determinism

Economic evaluation rests on an implicit contract: a competent third party, given the same inputs, can reproduce the result. Generative AI strains that contract in ways that differ from ordinary analytic variation. Outputs vary between runs at non-zero temperature. They vary with rephrasings of the prompt that carry no substantive difference. And they vary because providers update models without notice and retire earlier versions, so a specification that ran in one quarter may be unreproducible in the next through no fault of the analyst.
Yarar et al. (2026), reviewing 42 systematic reviews, found technical capability concerns (reliability and reproducibility prominent among them) to be the most frequently raised category, ahead of ethical, legal and societal concerns and ahead of cost. They also judged the underlying reviews to be of very poor methodological quality, which is itself a finding about how mature this evidence base is not.
The governance response so far has been reporting-based. ELEVATE-GenAI specifies ten domains including model characteristics, accuracy, reproducibility, and fairness and bias (Fleurence et al., 2025b). CHEERS-AI adds ten AI-specific items to the 28 CHEERS 2022 items and elaborates eight existing ones (Elvidge, Hawksworth, et al., 2024; Elvidge & Dawoud, 2024). Both are necessary and well built. Neither is sufficient, for a reason that has nothing to do with the quality of the drafting: disclosing a model version does not restore reproducibility once that version has been withdrawn. Real reproducibility for determinative analyses needs either version-pinned self-hosted models or the archiving of complete input–output pairs alongside the analysis.

Bias, Fairness and the Proxy Problem

Bias in health algorithms is mostly not a property of model architecture. It is a property of the target variable and of who is represented in the training data. Obermeyer et al. (2019) gave the canonical demonstration: an algorithm applied to millions of patients showed severe racial bias, not because race was an input, but because healthcare cost was used as a proxy for health need. Since less had historically been spent on Black patients at equivalent morbidity, the proxy encoded that disparity as lower predicted need. Fixing the target variable would have raised the share of Black patients flagged for additional support from 17.7% to 46.5%.
That mechanism transfers directly to HEOR, and it should worry anyone applying these methods in an LMIC. Utilisation and expenditure data from settings with substantial unmet need systematically understate the need of people who did not present for care, or could not. In a system where the poorest households face the steepest access barriers, an algorithm trained on utilisation will conclude that they need less. Where that algorithm feeds cohort construction, cost estimation or subgroup identification inside an economic evaluation, the analysis ends up encoding the existing distribution of access as though it were a distribution of need, and doing so underneath a layer of quantitative authority that makes the judgement hard to argue with.
A separate exposure comes from foundation models trained overwhelmingly on high-income, English-language corpora. Applied to disease epidemiology, treatment pathways, resource use or cost structures in an LMIC, such models default to the distributions they have seen. The digital divide in AI capability is well documented (Gulumbe et al., 2023), and equitable adoption is constrained by infrastructure and governance gaps that fall unevenly across countries (Ahmed et al., 2025). So the risk is not simply that LMIC agencies get access later. It is that when access arrives, the tools carry assumptions drawn from elsewhere and state them with undifferentiated confidence.

Privacy, Synthetic Data and Sovereignty

Cloud-hosted commercial models require data to leave the institution and usually the jurisdiction, which is hard to square with data-protection obligations over patient-level records. Synthetic data has been proposed as a way around this, but the proposal contains a tension nobody has resolved: synthetic records must stay faithful enough to preserve analytic utility while differing enough to prevent re-identification, and there is no consensus standard for generating them or for evaluating that trade-off (Arora et al., 2025). Using synthetic data to parameterise a model that will inform a national coverage decision, without such standards, moves an unquantified risk into the decision itself.
WHO guidance on large multi-modal models addresses these questions head-on, with over forty recommendations covering data governance, inclusive design, transparency and the regulatory capacity governments need in order to assess health applications at all (WHO, 2024). The guidance is explicit that a governance gap runs between countries able to regulate these technologies and countries that are not, which is the same asymmetry this review approaches from the evidence-production side.

Deskilling and the Erosion of Verification Capacity

The fifth risk gets less attention because it arrives slowly and nobody notices the day it starts.
If junior analysts learn to prompt rather than to model, the institution loses its ability to detect when the prompt has returned something wrong. This is not hypothetical in agencies where the analytic bench is already thin. Model conceptualisation, structural assumption testing and the interpretation of probabilistic sensitivity analysis are skills built by repeated practice on problems whose answers nobody knows in advance, which is exactly the kind of practice that automation removes first because it is the most tedious. An institution that automates the practice while keeping responsibility for the judgement is accumulating a liability that stays invisible until it fails. The ordering matters more than the tooling: modelling competence should come first.

The Governance Landscape and What It Omits

Taken together, the available instruments amount to a coherent architecture, and I want to give it its due before criticising it. PALISADE offers considerations for judging whether machine learning adds anything over conventional approaches, and deals with transparency and explainability directly (Padula et al., 2022). CHEERS-AI extends the dominant reporting standard for economic evaluation to AI-enabled interventions (Elvidge, Hawksworth, et al., 2024). ELEVATE-GenAI supplies a ten-domain framework and checklist for reporting large language model–assisted HEOR research (Fleurence et al., 2025b). The NICE position statement requires that AI use be justified, transparent, aligned to established checklists, and augmentative of human involvement rather than substitutive, and flags cybersecurity exposures including data poisoning and prompt injection (NICE, 2024). The joint position of Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence, aligned to the RAISE recommendations, holds that synthesists remain ultimately accountable, that AI may be used only where it demonstrably does not compromise methodological rigour, that human oversight is mandatory, and that any AI use making or suggesting judgements must be reported transparently (Flemyng et al., 2025). Governance commentary has argued for balancing innovation against ethical constraint in HTA specifically (Carapinha et al., 2024).
Three features of this architecture seem to me to limit its usefulness outside the settings that produced it.
It is overwhelmingly reporting-centric. These instruments specify what must be disclosed, not what must be verified, by whom, or to what standard. Disclosure is necessary for external scrutiny but it is not scrutiny, and it presupposes a reader with both the capacity and the standing to act on what has been disclosed.
It was written inside a small set of well-resourced institutions. The ISPOR working groups, NICE and the evidence-synthesis collaborations work in environments with deep analytic benches and mature data infrastructure. Their recommendations suit those environments. They are not framed as conditional on them, which is where the trouble starts when someone elsewhere tries to apply them.
Most consequentially, the architecture is undifferentiated with respect to consequence. The same broad expectations of transparency and human oversight apply whether the AI drafted a search string or estimated a comparative treatment effect. As a floor that is defensible. As a ceiling it is wasteful, because it imposes identical procedural cost on tasks whose errors differ by orders of magnitude in how reversible they are. In an agency where oversight capacity is the binding constraint, uniform requirements mean scarce scrutiny gets spread evenly instead of being concentrated where it would do the most good. That is the problem the framework in Section 7 is built to address.

Why High-Income Guidance Does Not Transfer

Four structural asymmetries explain why an LMIC agency cannot simply adopt the architecture described above and expect it to work.

The Threshold Problem

Cost-effectiveness analysis produces a ratio, and the ratio means nothing without a threshold to judge it against. Many LMIC evaluations still invoke one to three times gross domestic product per capita, a convention that never had a strong theoretical basis and now has a substantial empirical case against it.
Pichon-Riviere et al. (2023) derived thresholds for 174 countries from health expenditure per capita and life expectancy, finding values from US$87 to US$95,958 per quality-adjusted life year and, more importantly, that the threshold fell below 0.5 GDP per capita in 96% of low-income and 76% of lower-middle-income countries, and below 1 GDP per capita in 168 of the 174 countries examined. Vu et al. (2024) approached the question from an entirely different direction, synthesising 64 studies that elicited willingness to pay directly, and found a median of 0.53 GDP per capita, with most countries' values sitting below 1 GDP per capita. Two methodologically independent approaches converging well below the conventional benchmark is not easy to wave away.
The practical consequence is severe. An agency using a 1× GDP threshold where the opportunity-cost-based threshold is closer to 0.5× GDP will systematically approve technologies that displace more health than they generate, and will do so while believing it is being rigorous. The Indian evidence shows how sensitive conclusions are: Gupta et al. (2021) found temozolomide for glioblastoma multiforme cost approximately ₹212,020 per QALY gained, with only a 4.7% probability of being cost-effective at a 1× per capita GDP threshold, rising to 80% only if the price fell by 90%. Conclusions like that invert entirely under modest threshold revision.
The same logic scales up from single drugs to programme-level decisions, such as the introduction of novel tuberculosis vaccines across low- and middle-income countries, where modelled cost and cost-effectiveness estimates have to be read against whichever threshold the analyst has adopted (Portnoy et al., 2023).
This asymmetry is logically prior to anything about AI. A tool that accelerates the production of ICERs in a system without a defensible threshold accelerates the production of uninterpretable numbers. That is why the foundation layer in Figure 2 is a threshold and a reference case rather than a technology.

Data Infrastructure

High-income guidance assumes linked administrative datasets, national utility value sets, established costing databases and interoperable clinical records. Where these exist in LMICs they are usually partial, and this is not a gap that AI closes. A model trained to extract structured data from clinical notes needs clinical notes in retrievable electronic form. Where records are on paper, or fragmented across unlinked private providers, no amount of model capability substitutes for the missing substrate.
India is instructive here precisely because it has invested seriously. National EQ-5D-5L population norms now exist, drawn from 3,548 respondents across five states, reporting a mean utility of 0.848 and a mean visual analogue scale score of 75.18, with values falling across age and differing by sex (Jyani et al., 2023). Condition-specific utility data are accumulating (Wadasadawala et al., 2023). A national health system cost database supports costing (Gupta et al., 2021). That is a real achievement, and it is still incomplete relative to what the guidance takes for granted. Most LMICs are considerably further back. The comparative analysis of HTA development across China, India and South Africa identified weak health data infrastructure as a shared barrier alongside minimal HTA expertise, rising costs, fragmented delivery and a growing non-communicable disease burden (MacQuilkan et al., 2018).

The Capacity Paradox

Capacity constraints work in two directions at once, and this is the crux of the whole argument. Thin analytic capacity makes automation more attractive while making verification less possible. Those are not independent problems; they are the same problem seen from two sides.
The internationally supported capacity-building programme for evidence-informed priority setting in India was designed around exactly this, targeting capabilities at each level of the health system on the premise that effective HTA depends on a functioning ecosystem rather than on technical methods alone (Downey et al., 2020).
Adaptive HTA has emerged as one pragmatic response, offering a structured way to identify and run an optimised analysis where a full assessment is not feasible inside the decision window, and functioning both as a topic-prioritisation tool and as a way of holding onto methodological rigour under time pressure (Kumar S. et al., 2026). The parallel with my argument is close. Both adaptive HTA and risk-stratified AI governance accept that assessment intensity should be proportionate to what the decision requires and to what the institution can actually deliver, and neither treats the full high-income protocol as the only legitimate form of assessment.

India as an Instructive Case

India deserves specific treatment because it offers something rare: an evidence base against which claims about HTA in LMICs can be tested rather than merely asserted.
HTA has been partially institutionalised under the Department of Health Research within the Ministry of Health and Family Welfare (Chatterjee et al., 2025). An Indian Reference Case was developed through a four-step process comprising review of international guidelines, systematic review of adherence in three countries, empirical analysis of how alternative assumptions change cost-effectiveness results, and stakeholder consultation, issuing twelve recommendations across eleven principles of economic evaluation with deliberate room for context-specific flexibility (Sharma et al., 2023). Methodological adaptation for devices and health programmes, covering multiple usage, learning curves, long causal pathways and the valuation of false positives, has been worked out from the Indian experience (Prinja et al., 2021).
Most importantly for my argument, the return on this investment has been quantified. Bahuguna et al. (2025) evaluated three assessments commissioned between 2018 and 2020, using a framework that nets out the benefits obtainable without HTA and accounts for the opportunity cost of the assessment process itself. They found a return of 9:1, rising to a potential 71:1 with full implementation of recommendations, with variation from 5:1 to 40:1 across interventions.
That variability is at least as informative as the headline figure, and it points somewhere uncomfortable for the AI case. If returns vary that much between topics, the value of HTA sits largely in which questions get asked and whether the answers get implemented, not in how many assessments an agency can produce. A technology that increases throughput without improving prioritisation or uptake is therefore addressing the less important constraint. That is not an argument against using it; it is an argument against expecting throughput alone to deliver value.
The wider system context raises the stakes. Publicly funded insurance under Ayushman Bharat Pradhan Mantri Jan Aarogya Yojana has been associated with substantially lower out-of-pocket expenditure and a fall in catastrophic health expenditure from 45.9% to 13.3% among beneficiaries, with residual burden concentrated in non-medical costs and drug availability (Kumar et al., 2025). Coverage decisions inside a scheme of that size are consequential at population scale, which is the strongest argument available for governing the evidence that informs them.

The REACH-AI Framework

REACH-AI is my proposal for a capability–governance framework suited to HTA agencies in resource-constrained settings. It is not a reporting checklist and is not meant to replace ELEVATE-GenAI, CHEERS-AI or PALISADE, all of which should be applied wherever they are relevant. It addresses the question those instruments assume has already been answered: given what this institution actually possesses, which tasks can it responsibly delegate, and on what conditions?
The framework has three components. Five readiness domains (Figure 2), a three-tier task–risk stratification that sets oversight intensity (Figure 3), and a foundation requirement without which the rest does not function. Table 3 translates the five domains into the readiness questions an agency has to answer and the evidence that would show each answer is adequate.

Foundation Requirement

The foundation is a locally estimated cost-effectiveness threshold and a national reference case. I state this as a precondition rather than as a domain, because automation accelerates whatever decision rule is already in place. Where the rule is undefined or miscalibrated, faster assessment simply produces faster misallocation. An agency without an explicit threshold and reference case should treat establishing them as strictly prior to any investment in AI-assisted evidence generation, and the evidence in Section 6.1 indicates that falling back on a 1× GDP convention does not meet this requirement.

The Five Readiness Domains

Readiness: Data and Digital Infrastructure

The agency needs to be able to say what proportion of relevant clinical activity is captured in retrievable electronic form, whether local cost and utility sources exist or whether values are being borrowed from other jurisdictions, whether coding is standardised enough for automated extraction to work, and whether compute, connectivity and licensing are sustainable rather than propped up by time-limited donor funding. Where a data domain is simply absent, applications depending on it should be classified as not yet feasible. Calling them aspirational is a way of avoiding the decision.

Evidence-Fitness: Task and Risk Stratification

Every proposed application should be assigned to a tier before deployment, with oversight matched to that tier. Performance thresholds need to be pre-specified rather than reported afterwards, and benchmarking has to happen on locally relevant material, because performance on high-income English-language literature establishes nothing about performance on regional journals, national programme reports or multilingual grey literature.

Accountability: Human-In-The-Loop Assurance

Every AI-assisted output needs a named human guarantor who is professionally answerable for it. Prompts, model identifiers, versions, parameters and random seeds should be logged well enough that the analysis can be defended when challenged, which is a higher bar than logging them well enough to satisfy a reviewer. Determinative outputs require dual review. Disclosure in submissions has to be specific enough for an external assessor to work out what was automated and what was actually checked.

Capability: Skills, Retention and Sovereignty

Modelling competence must come before tooling. An agency that cannot build a Markov model unaided cannot evaluate one built with assistance, and no amount of procedural governance compensates for that. Analyst retention deserves treatment as an explicit governance objective rather than as a human-resources matter, since institutional memory is the substrate verification runs on. Where patient-level data are involved, domestic hosting of open-weight models is preferable, and the finding that optimised prompts transfer across model families makes this considerably more feasible than it was (Józwiak et al., 2026).

Harm Mitigation: Equity, Bias and Privacy

Inputs need auditing for representativeness, with particular attention to whether utilisation-based data are encoding access barriers as though they were preferences or need. Distributional impact assessment belongs wherever AI-derived parameters feed subgroup analysis. Consent and de-identification arrangements should be documented, and synthetic data should not be used to parameterise decision-relevant models until generation and evaluation standards exist (Arora et al., 2025). Environmental and opportunity costs of compute are worth stating openly, particularly where they compete with service delivery for the same budget line.

Task–Risk Stratification

The stratification in Figure 3 rests on one organising idea: oversight intensity should follow the reversibility of the error, not the sophistication of the tool. That inverts an intuition embedded in a good deal of current practice, which tends to scrutinise unfamiliar methods more heavily than familiar ones regardless of what actually turns on the result.
Tier 1, assistive, covers outputs that feed into human work and get verified as a by-product of ordinary practice: search-string drafting, reference formatting, translation, plain-language summarising, de-duplication. A single reviewer plus disclosure in the methods section is enough. Demanding more here diverts scrutiny away from where it is needed.
Tier 2, substitutive, covers outputs that replace a task a trained analyst would otherwise do in full: abstract and full-text screening, structured data extraction, risk-of-bias pre-assessment, cohort phenotyping from free text. Here I would require verification of a pre-specified random sample, with sensitivity reported under an asymmetric loss function favouring recall, following the specification Józwiak et al. (2026) validated, together with full prompt, model and version logging.
Tier 3, determinative, covers outputs that materially shape the ICER or the recommendation itself: comparative effect estimation, survival and cost extrapolation, utility and cost parameter selection, model structure, causal inference. The minimum here is independent dual human derivation, a pre-registered analysis plan, sensitivity analysis against a non-AI reference method, and explicit committee sight of what the AI contributed. Given the absence of validation evidence noted in Section 3.3, an agency choosing to use generative AI at Tier 3 should treat the exercise as methodological research that ought to be published, not as routine practice.
The efficiency argument for structuring things this way deserves stating plainly, because it is what makes the framework adoptable rather than merely admirable. Uniform oversight spreads finite scrutiny evenly across tasks whose consequences differ enormously. Stratification concentrates it where errors cannot be undone. For an agency with three analysts, that is the only arrangement under which these tools can be used safely at all.

Positioning an Agency: Readiness Against Adoption Intensity

Figure 4 locates an agency along two axes: institutional readiness across the REACH-AI domains, and intensity of generative AI adoption. The four quadrants describe different risk profiles rather than different levels of virtue.
Constrained baseline is the manual, slow, under-resourced agency whose evidence backlog grows faster than it can absorb; the operative risk there is that decisions get made with no HTA at all. Deliberate capacity describes an agency with governance, data and skills in place that has paced adoption cautiously. It is safe, and it is leaving efficiency on the table. Governed acceleration is the target, where automation is matched by verification, local benchmarking and retained human modelling competence. Ungoverned acceleration is rapid uptake without verification, benchmarking or local validation, and it is the most dangerous position precisely because its failures are quiet: the outputs are fast, internally consistent and confidently wrong.
None of this is an argument for slower adoption. It is an argument for staying on the diagonal, with readiness and intensity advancing together and readiness slightly ahead. The jump from constrained baseline straight to ungoverned acceleration is the path of least resistance under capacity pressure, and interrupting that jump is what the framework is for.

A Research Agenda

Six questions strike me as priorities, and Table 4 pairs each with a proposed design and the cost of leaving it unanswered. I have ordered them by what it costs to keep not knowing the answer, rather than by how easy they would be to answer.
The first and most urgent: do generative AI–proposed economic model structures differ materially from expert-derived structures in their effect on the ICER? This is the largest gap I identified and it is directly testable. Have independent teams build models for the same decision problem with and without AI assistance, then compare structures, ICERs and where the decision uncertainty ends up sitting.
Second, does screening and extraction performance hold on LMIC-relevant corpora? Published benchmarks rest on high-income, English-language literature. Until someone replicates them on regional journals, national programme reports and multilingual grey literature, agencies elsewhere are relying on performance claims that were never tested on their material.
Third, what is the cost-effectiveness of AI adoption inside an HTA agency? Since the return on HTA investment has been quantified for India at 9:1 (Bahuguna et al., 2025), the marginal return on AI-assisted assessment can be estimated with the same framework. It should be, before agencies commit budget on the strength of vendor claims.
Fourth, does AI-assisted evidence generation change equity outcomes? Where utilisation-based data underpin automated cohort construction or parameter estimation, the mechanism Obermeyer et al. (2019) identified implies a testable prediction about distributional bias. As far as I can tell, nobody has examined it in an economic evaluation context.
Fifth, does the ordering of capability and tooling affect output quality? The deskilling hypothesis in Section 4.5 is tractable through longitudinal comparison of analyst competence in agencies that adopted AI before versus after establishing modelling capability.
Sixth, is REACH-AI usable? It needs prospective evaluation in at least two agencies with different readiness profiles, looking at whether the tier classifications get applied consistently, whether the oversight requirements turn out to be proportionate in practice, and whether the readiness domains actually discriminate between institutions in ways that predict output quality.

Limitations

Several things qualify what I have argued here.

This is a critical rather than a systematic review. No protocol was registered, screening was not duplicated, and what got included reflects a judgement about relevance to an argument. Relevant work will have been missed, particularly work published in languages other than English and in venues the databases I searched do not index, which is an irony I am aware of given that the review itself complains about corpus coverage.
The empirical base is young and thin. Most validation studies date from 2024 to 2026, several report a single team's implementation, and independent replication is largely absent. Reported performance is probably flattering, since feasibility demonstrations publish more readily than failures. Everything I have said about performance should be read as provisional.
The framework is a conceptual proposal that has not been evaluated prospectively. Its domains and tiers came from the reviewed literature and from reasoning about institutional constraint, not from a Delphi process or any other consensus method, and its usability, reliability and predictive validity are unestablished. I offer it as a hypothesis about how to govern these tools, for testing rather than for adoption on my say-so.
It also leans heavily on the Indian experience, which is unrepresentative in one important respect: India has an institutionalised HTA function, a published reference case, national utility norms and quantified return-on-investment evidence. Agencies without those will find the foundation requirement considerably more demanding than the Indian case makes it look, and generalisation should be handled accordingly.
Finally, the technology will move faster than this literature. The specific performance figures I have cited will date quickly. The structural argument, that oversight should track reversibility and that readiness should lead adoption, is meant to survive that movement, but readers should check current capability against current evidence rather than against numbers reported in 2025 and 2026.

Conclusions

Generative AI is going to become a routine part of health economics and outcomes research, and arguing about whether that should happen seems to me a poor use of anyone's time. The evidence reviewed here shows real, quantified gains in bounded and verifiable tasks. Screening and extraction accuracy in HEOR-specific systems is high. Prompt optimisation improves recall across model families for almost nothing. And the findings that price does not predict performance, and that optimised prompts transfer between models, make responsible adoption substantially more achievable for agencies without large procurement budgets than it looked two years ago.
The same evidence shows that those gains sit where errors are cheap and visible, and that the tasks carrying the most weight in a coverage decision remain largely unvalidated. Reference fabrication, non-determinism, proxy-induced bias and privacy exposure are unresolved rather than solved. The governance instruments now available are well built, but they govern disclosure rather than verification, and they assume institutional conditions most LMIC agencies do not meet.
REACH-AI is my answer to that mismatch. Its central claim is simple and its implications are awkward: automation accelerates whatever decision rule is already in place. An agency without a locally estimated threshold and a national reference case that invests in AI-assisted evidence generation will produce uninterpretable numbers more quickly than before. An agency with a thin analytic bench that automates determinative tasks without verification will produce errors it has no way of catching. In both cases the failure is silent, because the output reads well.
Which makes the most consequential recommendation in this review also the least technological. Establish the threshold. Establish the reference case. Build modelling competence before buying tools. Then stratify by how reversible the error is, and put the scrutiny where it cannot be undone. Agencies that follow that order will capture most of the available gain at a fraction of the risk. Agencies that reverse it will find out about the risk only after a decision has already been made.

Funding

This research received no specific grant from any funding agency in the public, commercial or not-for-profit sectors.

Institutional Review Board Statement

Not required. This review used only published literature and publicly available agency documents; no human participants or identifiable patient data were involved.

Data Availability Statement

No primary data were generated. All sources cited are publicly available and identified by digital object identifier in the reference list.

Acknowledgments

A large language model was used to support literature identification, verification of bibliographic metadata against indexed records, and drafting of the manuscript. All retrieved sources were independently verified by the author against indexed bibliographic records for author list, title, journal, volume, issue, pagination and digital object identifier. All figures were originally constructed for this article; no artwork was reproduced or adapted from third-party sources. The author reviewed, revised and verified all content and takes full responsibility for the accuracy, integrity and interpretation of the work, consistent with the position that accountability for AI-assisted output rests with the named human author (Flemyng et al., 2025; National Institute for Health and Care Excellence, 2024).

Conflicts of Interest

The author declares no competing interests.

References

  1. Ahmed, M. M.; Okesanya, O. J.; Olaleke, N. O.; Adigun, O. A.; Adebayo, U. O.; Oso, T. A.; Eshun, G.; Lucero-Prisno, D. E. Integrating digital health innovations to achieve universal health coverage: Promoting health outcomes and quality through global public health equity. Healthcare 2025, 13(9), 1060. [Google Scholar] [CrossRef]
  2. Arora, A.; Wagner, S. K.; Carpenter, R.; Jena, R.; Keane, P. A. The urgent need to accelerate synthetic data privacy frameworks for medical research. Lancet Digit. Health 2025, 7(2), e157–e160. [Google Scholar] [CrossRef]
  3. Bahuguna, P.; Baker, P. A.; Briggs, A.; Gulliver, S.; Hesselgreaves, H.; Mehndiratta, A.; Ruiz, F.; Tyagi, K.; Wu, O.; Guzman, J.; Grieve, E. Is health technology assessment value for money? Estimating the return on investment of health technology assessment in India (HTAIn). BMJ Evid.-Based Med. 30 2025, Suppl 2, s29–s37. [Google Scholar] [CrossRef]
  4. Bridges, J. M. Computerized diagnostic decision support systems — A comparative performance study of Isabel Pro vs. ChatGPT4. Diagnosis 2024, 11(3), 250–258. [Google Scholar] [CrossRef]
  5. Carapinha, J. L.; Botes, D.; Carapinha, R. Balancing innovation and ethics in AI governance for health technology assessment. J. Med. Econ. 2024, 27(1), 754–757. [Google Scholar] [CrossRef]
  6. Chatterjee, K.; Dangi, A.; Chauhan, V. S. Health technology assessment for mental health in India. Med. J. Armed Forces India 2025, 81(6), 626–629. [Google Scholar] [CrossRef]
  7. Cowie, M. R.; Bozkurt, B.; Butler, J.; Briggs, A.; Kubin, M.; Jonas, A.; Adler, A. I.; Patrick-Lake, B.; Zannad, F. How can we optimise health technology assessment and reimbursement decisions to accelerate access to new cardiovascular medicines? Int. J. Cardiol. 365 2022, 61–68. [Google Scholar] [CrossRef]
  8. Cucherat, M.; Demarcq, O.; Chassany, O.; Le Jeunne, C.; Borget, I.; Collignon, C.; Diebolt, V.; Feuilly, M.; Fiquet, B.; Leyrat, C.; Naudet, F.; Porcher, R.; Schmidely, N.; Simon, T.; Roustit, M. Methodological expectations for demonstration of health product effectiveness by observational studies. Therapie 2025, 80(1), 47–59. [Google Scholar] [CrossRef]
  9. Downey, L. E.; Dabak, S.; Eames, J.; Teerawattananon, Y.; De Francesco, M.; Prinja, S.; Guinness, L.; Bhargava, B.; Rajsekar, K.; Asaria, M.; Rao, N. V.; Selvaraju, V.; Mehndiratta, A.; Culyer, A.; Chalkidou, K.; Cluzeau, F. A. Building capacity for evidence-informed priority setting in the Indian health system: An international collaborative experience. Health Policy OPEN 1 2020, 100004. [Google Scholar] [CrossRef]
  10. Elvidge, J.; Dawoud, D. Reporting standards to support cost-effectiveness evaluations of AI-driven health care. Lancet Digit. Health 2024, 6(9), e602–e603. [Google Scholar] [CrossRef]
  11. Elvidge, J.; Hawksworth, C.; Avşar, T. S.; Zemplenyi, A.; Chalkidou, A.; Petrou, S.; Petykó, Z.; Srivastava, D.; Chandra, G.; Delaye, J.; Denniston, A.; Gomes, M.; Knies, S.; Nousios, P.; Siirtola, P.; Wang, J.; Dawoud, D. Consolidated Health Economic Evaluation Reporting Standards for interventions that use artificial intelligence (CHEERS-AI). Value Health 2024, 27(9), 1196–1205. [Google Scholar] [CrossRef]
  12. Flemyng, E.; Noel-Storr, A.; Macura, B.; Gartlehner, G.; Thomas, J.; Meerpohl, J. J.; Jordan, Z.; Minx, J.; Eisele-Metzger, A.; Hamel, C.; Jemioło, P.; Porritt, K.; Grainger, M. Position statement on artificial intelligence (AI) use in evidence synthesis across Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence 2025. Campbell Syst. Rev. 2025, 21(4), e70074. [Google Scholar] [CrossRef]
  13. Fleurence, R. L.; Wang, X.; Bian, J.; Higashi, M. K.; Ayer, T.; Xu, H.; Dawoud, D.; Chhatwal, J. A taxonomy of generative artificial intelligence in health economics and outcomes research: An ISPOR Working Group report. Value Health 2025a, 28(11), 1601–1610. [Google Scholar] [CrossRef]
  14. Fleurence, R. L.; Dawoud, D.; Bian, J.; Higashi, M. K.; Wang, X.; Xu, H.; Chhatwal, J.; Ayer, T. ELEVATE-GenAI: Reporting guidelines for the use of large language models in health economics and outcomes research — An ISPOR Working Group report. Value Health 2025b, 28(11), 1611–1625. [Google Scholar] [CrossRef]
  15. Fleurence, R. L.; Bian, J.; Wang, X.; Xu, H.; Dawoud, D.; Higashi, M.; Chhatwal, J. Generative artificial intelligence for health technology assessment: Opportunities, challenges, and policy considerations — An ISPOR Working Group report. Value Health 2025c, 28(2), 175–183. [Google Scholar] [CrossRef]
  16. Geissler, J.; Makaroff, L. E.; Söhlke, B.; Bokemeyer, C. Precision oncology medicines and the need for real world evidence acceptance in health technology assessment: Importance of patient involvement in sustainable healthcare. Eur. J. Cancer 193 2023, 113323. [Google Scholar] [CrossRef]
  17. Glynn, D.; Giardina, J.; Hatamyar, J.; Pandya, A.; Soares, M.; Kreif, N. Integrating decision modeling and machine learning to inform treatment stratification. Health Econ. 2024, 33(8), 1772–1792. [Google Scholar] [CrossRef]
  18. Gorenshtein, A.; Shihada, K.; Sorka, M.; Aran, D.; Shelly, S. LITERAS: Biomedical literature review and citation retrieval agents. Comput. Biol. Med. 2025, 192 Pt B, 110363. [Google Scholar] [CrossRef]
  19. Gulumbe, B. H.; Yusuf, Z. M.; Hashim, A. M. Harnessing artificial intelligence in the post-COVID-19 era: A global health imperative. Trop. Dr. 2023, 53(4), 414–415. [Google Scholar] [CrossRef]
  20. Gupta, N.; Prinja, S.; Patil, V.; Bahuguna, P. Cost-effectiveness of temozolamide for treatment of glioblastoma multiforme in India. JCO Glob. Oncol. 7 2021, 108–117. [Google Scholar] [CrossRef]
  21. Hamaya, R.; Hara, K.; Manson, J. E.; Rimm, E. B.; Sacks, F. M.; Xue, Q.; Qi, L.; Cook, N. R. Machine-learning approaches to predict individualized treatment effect using a randomized controlled trial. Eur. J. Epidemiol. 2025, 40(2), 151–166. [Google Scholar] [CrossRef]
  22. Husereau, D.; Drummond, M.; Augustovski, F.; de Bekker-Grob, E.; Briggs, A. H.; Carswell, C.; Caulley, L.; Chaiyakunapruk, N.; Greenberg, D.; Loder, E.; Mauskopf, J.; Mullins, C. D.; Petrou, S.; Pwu, R.-F.; Staniszewska, S. Consolidated Health Economic Evaluation Reporting Standards 2022 (CHEERS 2022) statement: Updated reporting guidance for health economic evaluations. Value Health 2022, 25(1), 3–9. [Google Scholar] [CrossRef]
  23. Józwiak, Á.; Imre, A.; Hagymásy, J.; Tittmann, J.; Nagy, Á.; Kovács, S.; Kardas, P.; van Boven, J. F. M.; Mommers, I.; Ágh, T. REFLECTIVE-TIAB: Cost-effective prompt optimization for large language model-based title and abstract screening in literature reviews. Expert Rev. Pharmacoeconomics Outcomes Res. 2026, 26(6), 747–755. [Google Scholar] [CrossRef]
  24. Jyani, G.; Prinja, S.; Garg, B.; Kaur, M.; Grover, S.; Sharma, A.; Goyal, A. Health-related quality of life among Indian population: The EQ-5D population norms for India. J. Glob. Health 13 2023, 04018. [Google Scholar] [CrossRef]
  25. Kumar, A. P.; Yerram, A.; Chugh, Y.; Rana, S.; Mudgal, D.; Prinja, S.; Muraleedharan, V. R. Impact of India's publicly funded health insurance scheme on financial risk protection: A case-control study from Haryana state in India. BMJ Open 2025, 15(9), e093304. [Google Scholar] [CrossRef]
  26. Kumar S., S.; Loganathan, V.; Drake, T.; Mehndiratta, A.; Demeshko, A.; Pramesh, C. S.; Sengar, M.; Kar, S. S. Adaptive health technology assessment — Value and limitations to inform resource allocation in India. In Health Economics, Policy and Law; Advance online publication; 2026. [Google Scholar] [CrossRef]
  27. Lee, K.; Paek, H.; Ofoegbu, N.; Rube, S.; Higashi, M. K.; Dawoud, D.; Xu, H.; Shi, L.; Wang, X. A4SLR: An agentic artificial intelligence-assisted systematic literature review framework to augment evidence synthesis for health economics and outcomes research and health technology assessment. Value Health 2025, 28(11), 1655–1664. [Google Scholar] [CrossRef]
  28. Li, Y.; Datta, S.; Rastegar-Mojarad, M.; Lee, K.; Paek, H.; Glasgow, J.; Liston, C.; He, L.; Wang, X.; Xu, Y. Enhancing systematic literature reviews with generative artificial intelligence: Development, applications, and performance evaluation. J. Am. Med. Inform. Assoc. 2025, 32(4), 616–625. [Google Scholar] [CrossRef]
  29. MacQuilkan, K.; Baker, P.; Downey, L.; Ruiz, F.; Chalkidou, K.; Prinja, S.; Zhao, K.; Wilkinson, T.; Glassman, A.; Hofman, K. Strengthening health technology assessment systems in the global south: A comparative analysis of the HTA journeys of China, India and South Africa. Glob. Health Action 2018, 11(1), 1527556. [Google Scholar] [CrossRef]
  30. McGowan, A.; Gui, Y.; Dobbs, M.; Shuster, S.; Cotter, M.; Selloni, A.; Goodman, M.; Srivastava, A.; Cecchi, G. A.; Corcoran, C. M. ChatGPT and Bard exhibit spontaneous citation fabrication during psychiatry literature search. Psychiatry Res. 326 2023, 115334. [Google Scholar] [CrossRef]
  31. National Institute for Health and Care Excellence. Use of AI in evidence generation: NICE position statement. 2024. Available online: https://www.nice.org.uk/corporate/ecd11.
  32. Obermeyer, Z.; Powers, B.; Vogeli, C.; Mullainathan, S. Dissecting racial bias in an algorithm used to manage the health of populations. Science 2019, 366(6464), 447–453. [Google Scholar] [CrossRef]
  33. Padula, W. V.; Kreif, N.; Vanness, D. J.; Adamson, B.; Rueda, J.-D.; Felizzi, F.; Jonsson, P.; IJzerman, M. J.; Butte, A.; Crown, W. Machine learning methods in health economics and outcomes research — The PALISADE checklist: A good practices report of an ISPOR task force. Value Health 2022, 25(7), 1063–1080. [Google Scholar] [CrossRef]
  34. Pichon-Riviere, A.; Drummond, M.; Palacios, A.; Garcia-Marti, S.; Augustovski, F. Determining the efficiency path to universal health coverage: Cost-effectiveness thresholds for 174 countries based on growth in life expectancy and health expenditures. Lancet Glob. Health 2023, 11(6), e833–e842. [Google Scholar] [CrossRef]
  35. Portnoy, A.; Clark, R. A.; Quaife, M.; Weerasuriya, C. K.; Mukandavire, C.; Bakker, R.; Deol, A. K.; Malhotra, S.; Gebreselassie, N.; Zignol, M.; Sim, S. Y.; Hutubessy, R. C. W.; Baena, I. G.; Nishikiori, N.; Jit, M.; White, R. G.; Menzies, N. A. The cost and cost-effectiveness of novel tuberculosis vaccines in low- and middle-income countries: A modeling study. PLoS Med. 2023, 20(1), e1004155. [Google Scholar] [CrossRef]
  36. Prinja, S.; Jyani, G.; Gupta, N.; Rajsekar, K. Adapting health technology assessment for drugs, medical devices, and health programs: Methodological considerations from the Indian experience. Expert Rev. Pharmacoeconomics Outcomes Res. 2021, 21(5), 859–868. [Google Scholar] [CrossRef]
  37. Reason, T.; Langham, J.; Gimblett, A. Automated mass extraction of over 680,000 PICOs from clinical study abstracts using generative AI: A proof-of-concept study. Pharm. Med. 2024, 38(5), 365–372. [Google Scholar] [CrossRef]
  38. Sarri, G.; Hernandez, L. G. The maze of real-world evidence frameworks: From a desert to a jungle! An environmental scan and comparison across regulatory and health technology assessment agencies. J. Comp. Eff. Res. 2024, 13(9), e240061. [Google Scholar] [CrossRef]
  39. Sharma, D.; Prinja, S.; Aggarwal, A. K.; Rajsekar, K.; Bahuguna, P. Development of the Indian Reference Case for undertaking economic evaluation for health technology assessment. Lancet Reg. Health – Southeast Asia 16 2023, 100241. [Google Scholar] [CrossRef]
  40. Vu, A. N.; Hoang, M. V.; Lindholm, L.; Sahlen, K. G.; Nguyen, C. T. T.; Sun, S. A systematic review on the direct approach to elicit the demand-side cost-effectiveness threshold: Implications for low- and middle-income countries. PLoS ONE 2024, 19(2), e0297450. [Google Scholar] [CrossRef]
  41. Wadasadawala, T.; Mohanty, S. K.; Sen, S.; Khan, P. K.; Pimple, S.; Mane, J. V.; Sarin, R.; Gupta, S.; Parmar, V. Health-related quality of life (HRQoL) using EQ-5D-5L: Value set derived for Indian breast cancer cohort. Asian Pac. J. Cancer Prev. 2023, 24(4), 1199–1207. [Google Scholar] [CrossRef]
  42. World Health Organization. Ethics and governance of artificial intelligence for health: Guidance on large multi-modal models. 2024. Available online: https://www.who.int/publications/i/item/9789240084759.
  43. Yarar, F.; Addis, P.; Fairweather, M.; Craig, D.; O'Keefe, H. Concerns of using large language models in health care research and practice: Umbrella review. J. Med. Internet Res. 28 2026, e87804. [Google Scholar] [CrossRef]
  44. Zong, J.; Rojubally, A.; Pan, X.; Wolf, B.; Greenfeder, S.; Upton, A.; Gdovin Bergeson, J. A review and comparative case study analysis of real-world evidence in European regulatory and health technology assessment decision making for oncology medicines. Value Health 2025, 28(1), 31–41. [Google Scholar] [CrossRef]
Figure 1. The HEOR and HTA Evidence Value Chain and Entry Points for Generative Artificial Intelligence. HEOR = health economics and outcomes research; HTA = health technology assessment; PICO = population, intervention, comparator, outcome. The consequence gradient represents the declining reversibility of an undetected error as outputs move downstream toward the recommendation. Figure constructed by the author.
Figure 1. The HEOR and HTA Evidence Value Chain and Entry Points for Generative Artificial Intelligence. HEOR = health economics and outcomes research; HTA = health technology assessment; PICO = population, intervention, comparator, outcome. The consequence gradient represents the declining reversibility of an undetected error as outputs move downstream toward the recommendation. Figure constructed by the author.
Preprints 230306 g001
Figure 2. The REACH-AI Capability–Governance Framework for Health Technology Assessment in Low- and Middle-Income Countries. REACH-AI = Readiness, Evidence-fitness, Accountability, Capability, Harm mitigation. The foundation layer is stated as a precondition rather than a domain, because the framework's central claim is that automation accelerates whatever decision rule is already in place. Figure constructed by the author.
Figure 2. The REACH-AI Capability–Governance Framework for Health Technology Assessment in Low- and Middle-Income Countries. REACH-AI = Readiness, Evidence-fitness, Accountability, Capability, Harm mitigation. The foundation layer is stated as a precondition rather than a domain, because the framework's central claim is that automation accelerates whatever decision rule is already in place. Figure constructed by the author.
Preprints 230306 g002
Figure 3. Tiered Task–Risk Stratification and Corresponding Minimum Oversight Requirements. Tier assignment is determined by the reversibility of an undetected error rather than by the technical sophistication of the method employed. Requirements are stated as minima and are additive across tiers. Figure constructed by the author.
Figure 3. Tiered Task–Risk Stratification and Corresponding Minimum Oversight Requirements. Tier assignment is determined by the reversibility of an undetected error rather than by the technical sophistication of the method employed. Requirements are stated as minima and are additive across tiers. Figure constructed by the author.
Preprints 230306 g003
Figure 4. Institutional Readiness and Adoption Intensity: Four Positions for a Health Technology Assessment Agency. The diagonal represents alignment between institutional readiness and adoption intensity. Movement from constrained baseline directly to ungoverned acceleration is the path of least resistance under capacity pressure and is the trajectory the framework is designed to interrupt. Figure constructed by the author.
Figure 4. Institutional Readiness and Adoption Intensity: Four Positions for a Health Technology Assessment Agency. The diagonal represents alignment between institutional readiness and adoption intensity. Movement from constrained baseline directly to ungoverned acceleration is the path of least resistance under capacity pressure and is the trajectory the framework is designed to interrupt. Figure constructed by the author.
Preprints 230306 g004
Table 1. Structured Search Strategy and Sources of Evidence.
Table 1. Structured Search Strategy and Sources of Evidence.
Element Specification
Databases PubMed and MEDLINE, searched from inception to 23 August 2026
Concept axis 1
(technology)
artificial intelligence; machine learning; large language models; generative artificial intelligence; foundation models
Concept axis 2
(discipline)
health economics; outcomes research; health technology assessment; economic evaluation; cost-effectiveness; evidence synthesis; systematic review
Concept axis 3
(setting)
low- and middle-income countries; India; universal health coverage; priority setting
Supplementary
strategies
Direct retrieval of agency and multilateral guidance not indexed bibliographically; backward citation tracking from ISPOR working group reports and the NICE position statement; forward tracking via related-article algorithms
Inclusion
criteria
Empirical performance of AI methods in an HEOR or evidence-synthesis task; methodological, ethical or governance position from a recognised agency or collaboration; methodological standards for economic evaluation; characterisation of HTA conditions in resource-constrained settings
Exclusion
criteria
Commentary without standard-setting authority; studies of AI as a clinical intervention without an economic or assessment component
Verification Every citation verified against indexed bibliographic metadata for author list, title, journal, volume, issue, pagination and digital object identifier
Note. HEOR = health economics and outcomes research; HTA = health technology assessment; ISPOR = Professional Society for Health Economics and Outcomes Research; NICE = National Institute for Health and Care Excellence. Concept axes were searched both in combination and separately, since the intersection of all three axes is sparse.
Table 2. Reported Performance of Generative Artificial Intelligence in Evidence-Synthesis Tasks.
Table 2. Reported Performance of Generative Artificial Intelligence in Evidence-Synthesis Tasks.
Study Task and design Reported performance Principal caveat
Li et al. (2025) Five-module human-in-the-loop system for HTA submissions; screening and extraction Screening sensitivity 90%; accuracy 89%; Cohen's κ 0.71; extraction F1 93 Internal reference standard; no manually annotated training data required
Lee et al. (2025) Eight-module agentic framework; two clinical areas Screening F1 0.917–0.977; risk-of-bias κ 0.844–0.906; extraction F 0.96–0.998 Single-team implementation; no independent external replication
Reason et al. (2024) Mass PICO extraction from 682,667 abstracts 98% accurate and complete extraction on a random sample of 350; <3 h; ~US$3,390 total Proof of concept; verification sample small relative to corpus
Józwiak et al. (2026) Reflective prompt optimisation under asymmetric loss; nine models; 8,520 records Recall improved 3.7–37.1 percentage points across all models; best model 91% accuracy, F1 81.6; optimisation cost US$6.36 Single clinical domain; model price did not predict performance
Gorenshtein et al. (2025) Multi-agent retrieval grounded in the PubMed application programming interface; five medical fields Citation accuracy 99.82%; referencing accuracy 96.81%; 0% non-academic sources Lower median journal impact factor than a proprietary comparator
Note. F1 and F denote harmonic means of precision and recall. HTA = health technology assessment; PICO = population, intervention, comparator, outcome. All performance figures are as reported by the original investigators; none has been independently replicated at the time of writing.
Table 3. Operationalising the REACH-AI Domains: Indicative Readiness Questions and Verification Evidence.
Table 3. Operationalising the REACH-AI Domains: Indicative Readiness Questions and Verification Evidence.
Domain Readiness question the agency must answer Evidence that the answer is adequate
Readiness What proportion of relevant clinical activity is captured in retrievable electronic form, and do local cost and utility sources exist? Documented data-coverage audit; national or condition-specific value set; costing database with stated provenance
Evidence-fitness Has every proposed application been assigned to a tier, with performance thresholds pre-specified? Written tier register; pre-specified sensitivity and precision thresholds; benchmarking conducted on locally relevant corpora
Accountability Is a named individual professionally answerable for each AI-assisted output, and can the analysis be reconstructed under challenge? Guarantor register; complete prompt, model, version and seed logs; dual-review records for determinative outputs
Capability Can the analytic team build and critique the relevant model without AI assistance? Demonstrated de novo modelling capability; analyst retention data; documented hosting and licensing arrangements
Harm mitigation Have inputs been audited for representativeness, and are consent and de-identification arrangements documented? Representativeness audit; distributional impact assessment; data-protection documentation; stated compute and opportunity cost
Note. REACH-AI = Readiness, Evidence-fitness, Accountability, Capability, Harm mitigation. Questions are stated so that a negative answer identifies a specific remediable gap rather than a general deficiency. Where a readiness question cannot be answered affirmatively, applications depending on that domain should be classified as not yet feasible.
Table 4. Priority Research Questions, Proposed Designs and Consequence of Continued Uncertainty.
Table 4. Priority Research Questions, Proposed Designs and Consequence of Continued Uncertainty.
Priority Question Proposed design Consequence if unresolved
1 Do generative AI–proposed model structures alter the incremental cost-effectiveness ratio relative to expert-derived structures? Parallel independent model construction for a common decision problem, with and without AI assistance Determinative use proceeds without any evidence of structural validity
2 Does screening and extraction performance hold on LMIC-relevant corpora? Replication of published benchmarks on regional journals, programme reports and multilingual grey literature Performance claims derived elsewhere are assumed to transfer
3 What is the return on investment of AI adoption within an HTA agency? Application of the established HTA return-on-investment framework to AI-assisted assessment Scarce agency budget is committed on the basis of vendor claims
4 Does AI-assisted evidence generation alter equity outcomes? Distributional analysis of AI-derived versus conventionally derived parameters in matched evaluations Access barriers are encoded as need and reproduced at scale
5 Does the ordering of capability and tooling affect output quality? Longitudinal comparison of analyst competence across agencies adopting AI before versus after establishing modelling capability Verification capacity erodes invisibly until failure
6 Is REACH-AI itself usable and discriminating? Prospective evaluation in at least two agencies with differing readiness profiles The framework remains an untested proposal
Note. HTA = health technology assessment; LMIC = low- and middle-income country; REACH-AI = Readiness, Evidence-fitness, Accountability, Capability, Harm mitigation. Priorities are ordered by the consequence of continued uncertainty rather than by feasibility.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.