Preprint
Review

This version is not peer-reviewed.

Artificial Intelligence Readiness in Clinical Trial Operations: A Narrative Review and Site-Level Governance Framework

Submitted:

17 July 2026

Posted:

20 July 2026

You are already at the latest version

Abstract
Background/Objectives: Artificial intelligence (AI) is increasingly introduced into clinical trial operations, but operational usefulness does not imply regulatory or site-level readiness. This review synthesizes evidence on AI applications across clinical trial operations and proposes a site-level governance model for responsible implementation in healthcare organizations. Methods: We conducted a structured narrative review with evidence mapping of peer-reviewed literature, regulatory documents and contextual sources (2020–2026). AI use cases were mapped by trial lifecycle stage, evidence maturity, autonomy level, trial impact and governance implications using a transparent 0–8 evidence-maturity score. Results: The strongest evidence concerns patient–trial matching and eligibility assessment, which nonetheless reached only moderate maturity because the systems were evaluated retrospectively or in simulated screening rather than inside a live trial; no use case reached the highest band. Other applications remain less mature or context-dependent. Recurrent risks include hallucination, automation bias, weak local validation, limited auditability, model drift and unclear accountability. We propose a site-level readiness and deployment-decision model linking evidence maturity, AI autonomy, trial impact and site implementation capacity. Conclusions: AI readiness in clinical trial operations should be assessed at the level of the AI-enabled workflow rather than the model alone. Safe adoption requires context-specific validation, human accountability, auditability, lifecycle monitoring and alignment with Good Clinical Practice.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

Clinical trials remain the cornerstone of evidence generation in medicine, yet they are widely regarded as slow, expensive and operationally fragile. Rising protocol complexity, difficulty in identifying and enrolling eligible participants, growing volumes of heterogeneous data, and the progressive decentralization of trial activities have all increased the operational burden placed on sponsors, contract research organizations and, in particular, investigational sites and the healthcare organizations that host them [1,2]. Site burden is now a measurable phenomenon: survey-based work quantifying protocol-related site burden across hundreds of sites has linked excessive procedural and administrative load to slower performance [66]. Against this background, artificial intelligence (AI) — and especially machine learning (ML) and, more recently, large language models (LLMs) and generative AI — has been proposed as a way to relieve specific bottlenecks in trial design, recruitment, monitoring, safety surveillance and documentation [2,3,4]. These developments build on broader early appraisals of GPT-4 in healthcare, which described its potential role in medical education, triage, documentation and clinical communication while emphasizing persistent limitations and the need for professional oversight [72].
Much of the existing literature frames AI in clinical trials as a collection of promising but isolated applications. However, AI is now moving from the automation of discrete tasks toward workflow-embedded systems that support, augment or partially automate activities across the entire trial lifecycle [1,3]. This transition changes the relevant questions. When AI performs a one-off calculation, the central issue is accuracy. When AI becomes embedded in operational workflows that influence who is screened, which sites are selected, how deviations are detected and what participants are told, the central issue becomes governance: validation, human oversight, auditability, accountability and alignment with Good Clinical Practice (GCP) [4,5].
In this review, we use the term AI as an operational layer to describe workflow-embedded AI systems that support or partially automate activities across protocol design, feasibility assessment, site selection, recruitment, eligibility assessment, trial conduct, monitoring, safety surveillance, documentation and regulatory oversight. This framing does not imply autonomous replacement of investigators, sponsors or trial staff; rather, it emphasizes that the operational value of AI is inseparable from the conditions under which its outputs are validated, supervised, documented and governed.
The operational rationale for AI adoption is real but contextual rather than evidentiary. Industry analyses indicate persistent pressure on clinical development productivity [41], high costs associated with trial delays — with one estimate placing the mean cost of a delay day in drug development at approximately USD 40,000 [42] — and increasing strain on investigational sites [43,66]. Commercial forecasts predict rapid growth in AI-enabled trial solutions but vary substantially across sources [46,47], and registry analyses indicate that AI-related trials are now globally distributed yet unevenly reported [54,55]. These signals help explain why adoption is accelerating, but they do not establish that AI is ready for routine operational use.
The aim of this narrative review is to synthesize current evidence on AI in clinical trial operations, to identify areas of relative operational maturity, to discuss major risks and failure modes, and to propose a site-level governance and readiness framework aligned with GCP, participant safety, data integrity, human oversight and emerging regulatory expectations. We address five questions: (i) which AI applications show the greatest current operational maturity; (ii) what are the main operational, ethical, data-related and regulatory risks; (iii) how can the level of AI autonomy be linked to governance requirements; (iv) which site-level readiness elements should be assessed before implementation; and (v) how can AI readiness be aligned with GCP, human accountability and evolving regulatory expectations. The scope is deliberately restricted to AI in clinical trial operations rather than AI in medicine at large; we exclude, except where directly relevant, AI in drug discovery, routine clinical diagnosis, imaging, robotic surgery, hospital operations unrelated to trials, and AI as a regulated medical device [6].

1.1. Positioning Relative to Existing Literature

Several recent frameworks and reviews have addressed AI adoption in clinical trials or in healthcare more broadly. Mateen and colleagues proposed a lifecycle-level framework for the effective adoption of AI in clinical trials [57], and You and colleagues proposed a clinical-trials-informed pathway for the real-world implementation and deployment of AI in healthcare organizations [62]. Broader guidance for trustworthy and deployable healthcare AI, such as FUTURE-AI, spans the entire lifecycle from design and validation to regulation, deployment and monitoring [63], and practical frameworks for the appropriate implementation and review of healthcare AI have been proposed [71]. Established reporting guidelines improve transparency when AI is the object of study: SPIRIT-AI and CONSORT-AI extend trial protocol and reporting standards to AI interventions [58,59], DECIDE-AI addresses the early-stage clinical evaluation of AI-driven decision-support systems [60], and TRIPOD+AI updates reporting for AI prediction models [61].
These contributions leave a specific and consequential gap. Existing AI-in-clinical-trials reviews primarily provide use-case taxonomies of where AI may be used, and existing AI reporting guidelines primarily address situations in which AI is the intervention, decision-support system or prediction model under evaluation. Many AI systems likely to enter clinical trial operations, however, will function as workflow infrastructure rather than study interventions: they may generate prescreening lists, summarize source data, prioritize queries, draft participant-facing explanations or flag safety signals, thereby influencing regulated trial conduct without appearing in the protocol as the object of evaluation. This produces an operational AI governance gap at the level of the investigational site — the need to determine when such tools are mature enough to use, what degree of human oversight is required, what documentation must be retained, and how an inspector could reconstruct an AI-influenced decision. In short, the literature no longer lacks awareness that AI can be used in clinical trials; it lacks a site-level implementation logic for determining when, how and under what governance conditions AI-enabled workflows should be used. The distinctive contribution of this review is to address that gap with a site-level decision model linking four dimensions rarely considered together — evidence maturity, AI autonomy, trial impact and site implementation capacity — to inspection-ready, GCP-aligned deployment decisions.
What this review adds
• It distinguishes AI as a clinical-trial intervention from AI as operational trial infrastructure.
• It maps AI use cases across the trial lifecycle by evidence maturity, autonomy level and trial impact.
• It argues that AI readiness should be assessed at the level of the AI-enabled workflow, not the model alone.
• It introduces inspection-ready AI as a practical standard for regulated trial operations.
• It provides a site-level deployment-decision model and implementation dossier for healthcare organizations hosting clinical trial sites.

2. Materials and Methods

We conducted a structured narrative review with evidence mapping. It is intended as a critical, transparent synthesis rather than a quantitative meta-analysis, and no pooled effect estimates or formal risk-of-bias scoring were performed. PubMed/MEDLINE served as the primary indexed database. It was supplemented by Scopus, Web of Science Core Collection and Google Scholar, which were used for citation chaining and for grey literature, and by targeted retrieval from the websites of the U.S. Food and Drug Administration (FDA), the European Medicines Agency (EMA) and the International Council for Harmonisation (ICH). The search covered literature published between 2020 and 2026, and older sources were retained only where conceptually necessary. The structured database search was executed on 13 July 2026. Full search strings, database yields and the eligibility criteria are reported in Supplementary Table S4.
The search was run at two levels of specificity, and the difference between them is itself informative. A broad query combining AI concepts (artificial intelligence, machine learning, large language model, generative AI) with clinical-trial concepts returned 6,877 records in PubMed/MEDLINE. That set proved dominated by studies in which AI is a clinical intervention, a diagnostic model or an outcome-prediction model, rather than infrastructure used to run a trial. A second query restricted the trial concepts to operational trial-conduct terms (eligibility criteria, trial recruitment, trial matching, risk-based monitoring, protocol deviations, clinical data management) and returned 815 records. This operationally specific set formed the pool from which peer-reviewed evidence was drawn.
We included peer-reviewed empirical studies, reviews, regulatory and reporting-guidance documents, legal instruments and implementation-science frameworks that addressed AI applied to the conduct of clinical trials. Records were excluded for the following pre-specified reasons: AI evaluated as a clinical intervention, a diagnostic or a prognostic model rather than as trial infrastructure; AI for drug discovery, molecular design or preclinical research; papers on trial recruitment or eligibility criteria with no AI component; AI applications outside clinical trial operations; trial reports in which machine learning served only as a statistical method for outcome analysis; and editorials, letters or conference abstracts without substantive content. Screening, full-text assessment and final selection were performed by the first author and verified by a second author, and disagreements were resolved by discussion. Two explicitly identified preprints were retained only as emerging evidence and were never used as sole support for a maturity classification. Industry, market and registry sources were used strictly as contextual, non-evidentiary material and are reported separately (Supplementary Table S1).
Sources were classified by evidentiary role: empirical operational evidence; review or synthesis evidence; regulatory, legal or reporting guidance; implementation-science framework; and contextual market or registry evidence. Only the first four categories were used to support evidence-maturity or governance conclusions; market and registry sources were used solely to characterize adoption pressure. A narrative design was chosen because the evidence base spans heterogeneous empirical studies, technical validations, regulatory documents, reporting guidelines, implementation frameworks and contextual sources that are not amenable to quantitative synthesis but do require conceptual integration.
Selection within the eligible pool was purposive rather than exhaustive, and this is the principal methodological difference between this review and a systematic review. We did not apply a PRISMA-style, count-based screening flow, and we do not claim to have appraised every eligible record. We report the search strings, the search date and the yield at each level of specificity so that the coverage of the search, and the residual risk of selective sourcing, can be judged by the reader. In total, 93 sources informed the manuscript: peer-reviewed empirical studies and reviews, regulatory, legal and reporting-guidance documents, contextual industry and registry sources, two explicitly identified preprints, and methodological or implementation-science references. The identification, selection and classification process is summarized in Supplementary Figure S2.
For each source we extracted the publication year, article type, AI technology, trial lifecycle stage, operational use case, reported benefits, risks or limitations, validation approach, implementation context and relevance to site-level governance. To make the maturity assessment transparent and comparable, we applied a 0–8 evidence-maturity score summing four criteria, each rated 0–2 (Table 1): evidence type, data realism, workflow integration and governance reporting. Scores mapped to four classes: 0–2, emerging; 3–4, limited; 5–6, moderate; 7–8, high. Component subscores are reported in Supplementary Table S2.
The four criteria were not selected arbitrarily. They correspond to the four questions that determine whether an operational AI result can be acted upon at a trial site: what kind of study produced the result (evidence type); whether it was obtained on data resembling the site’s own (data realism); whether the tool was tested inside a workflow rather than offline (workflow integration); and whether oversight, auditability and validation were reported at all (governance reporting). These dimensions map onto established implementation-outcome constructs, in particular feasibility, fidelity and appropriateness [76], onto the workflow and people dimensions of sociotechnical models of health information technology [77], and onto the applicability and reporting expectations of appraisal tools for AI prediction models [79]. We did not include model accuracy as a criterion, because accuracy is the quantity whose operational sufficiency is in question.
Equal weighting was applied deliberately, and it is a simplification that should be stated plainly. We had no empirical basis for differential weights, and any weighting we imposed would have encoded our own assumptions about which dimension matters most. Equal weights keep the mapping legible and auditable. To test whether the conclusions depend on that choice, we recomputed the ranking under two alternative weightings: one that doubled the weight of evidence type, and one that doubled the weight of workflow integration (Supplementary Table S6). The two most mature use cases retained their rank under both. Under the workflow-weighted variant, data cleaning and patient-facing consent AI rose into the same band as safety surveillance, but no use case reached the high band under any weighting. The substantive conclusions are therefore not an artefact of equal weighting.
Scores were assigned at the use-case level, and they reflect the dominant body of evidence for a use case rather than the single best study within it. Scoring was performed by the first author and verified by a second author, with disagreements resolved by discussion. The score is an evidence-mapping tool that makes the authors’ interpretation explicit and comparable. It is not a validated instrument and not a formal risk-of-bias assessment, and it should not be read as one. Because most of the underlying evidence comes from retrospective analyses and technical validations rather than from prospective use inside live trials, no use case in this corpus reached the high band; the maximum observed score was 6. Although this was not a systematic review, the manuscript was structured to address key quality expectations for narrative reviews, including justification of importance, clarity of aims, transparent literature search, appropriate referencing, scientific reasoning and relevant presentation of data [75]. Generative AI tools were used to support language editing, conceptual organization and the formatting of tables and figures; they were not used to generate original data or to perform independent evidence selection, and all AI-assisted content was reviewed and verified by the authors. The principal limitations of the method are those inherent to narrative reviews: the absence of a formal, exhaustive search protocol and quantitative bias assessment, potential selection bias and rapid obsolescence.

3. Results

3.1. Overview of the Evidence Corpus

The final synthesis included 93 sources spanning empirical AI clinical-trial studies, reviews and scoping reviews, regulatory and reporting-guidance documents, implementation-science and sociotechnical frameworks, and contextual registry or market sources. Empirical evidence was concentrated in recruitment and eligibility assessment, whereas patient-facing AI, digital twins, synthetic controls and several monitoring applications were supported by more heterogeneous or early-stage evidence. Consistent with the source classification described in the Methods, only empirical, review, regulatory and implementation-science sources were used to support the evidence-maturity map and governance conclusions that follow, while market and registry sources were used solely to contextualize adoption pressure. The characteristics of the principal evidentiary sources are summarized in Supplementary Table S5, and the identification and classification process is shown in Supplementary Figure S2.

3.2. AI Across the Clinical Trial Lifecycle

AI applications were identified at every stage of the clinical trial lifecycle, from protocol design and feasibility to documentation and regulatory oversight (Figure 1). Rather than a single dominant application, the evidence describes a distributed set of use cases, each with a distinct value proposition, dominant risk and governance requirement (Table 2). At the design stage, LLMs and information-extraction methods simplify or assess eligibility criteria and other protocol elements [3,14]. In feasibility and site selection, ML models trained on real-world and historical operational data estimate patient pools and rank sites and investigators [17,18,19]. In recruitment and eligibility, LLM-based systems support patient-trial matching and criterion-level screening [9,10,11]. During conduct and monitoring, AI supports risk-based quality management, protocol-deviation detection and site risk scoring [7,18,22]; in data management it supports query generation and data cleaning [20,21]. In safety, natural-language methods support adverse-event and adverse-drug-event detection [33,34,35], and in patient-facing processes, generative systems support informed-consent documents and participant education [23,24].
Two observations follow. First, the breadth of applications supports the characterization of AI as an emerging operational layer rather than a peripheral analytical tool. Second, risk is highly stage-dependent: an error in a draft administrative summary is not equivalent to an error in an eligibility decision, a safety signal or a statement to a participant. A lifecycle map must therefore precede any general claim about readiness.

3.3. Evidence Maturity of AI Use Cases

Applying the scoring rubric (Table 1; component subscores in Supplementary Table S2), operational maturity is uneven (Table 3; Figure 2). The comparatively strongest evidence base concerns recruitment and eligibility assessment, which scored 6 of 8 (moderate), supported by multiple empirical systems, EHR-based validation and a randomized evaluation of human–AI teaming [9,10,11,12,14]. It is important to be precise about what that rating does and does not mean. It did not reach the high band, because the reviewed systems were evaluated in user studies or simulated screening rather than inside a live trial workflow, and governance reporting was partial. No use case in this corpus reached the high band, and the maximum observed score was 6. Site selection and feasibility also scored moderate (5), on the strength of externally validated site-risk models, although the validation is retrospective [17,18,19]. Safety surveillance scored limited (4): the underlying evidence is real-world but comes predominantly from offline pharmacovigilance benchmarks and reviews rather than from trial conduct [33,34,35,36]. Risk-based monitoring, data management and patient-facing consent tools each scored limited (3) [7,20,21,22,23,24,25], and digital twins and synthetic control arms remain emerging (2) [30,31,32]. Therefore, maturity should determine the intended role of AI: cautious pilots with mandatory human review in the more mature areas, and strictly bounded, supervised experimentation or deferral in the less mature ones. A higher maturity rating indicates stronger supporting evidence, not a recommendation to adopt.

3.4. Recruitment and Eligibility Assessment as the Leading Use Case

Recruitment and eligibility assessment represent the most operationally mature use case. TrialGPT, an LLM framework combining retrieval, criterion-level matching and ranking, achieved 87.3% accuracy in criterion-level eligibility judgments and reduced screening time by 42.6% in a user study [9]. Whether such gains transfer to routine site workflows has not yet been shown. End-to-end systems such as TrialMatchAI process structured data and clinical narratives and produce explainable, retrieval-augmented outputs suitable for local deployment [10]. PRISM was validated against real-world electronic health record (EHR) data in an oncology center. It showed that a smaller, fine-tuned model can be competitive with larger general-purpose models, which is directly relevant to privacy and local deployment [11]. These systems were, however, developed and evaluated predominantly on English-language clinical records [9,10,11].
Upstream, AutoCriteria extracts granular inclusion and exclusion criteria from protocols [14], while work converting eligibility criteria into standardized database queries shows both the promise of automation and its characteristic failure modes, including hallucination and mapping errors [15]. Broader reviews indicate that recruitment and retention are among the most frequently addressed operational targets for AI [8]. Crucially, the safest configuration is human–AI teaming: a randomized evaluation using retrospective EHR data found that combining AI prescreening with human review improved both accuracy and efficiency relative to either alone [12]. Sociotechnical and economic analyses emphasize that value depends on workflow integration, cost and accountability, not on model accuracy alone [13], and emerging preprint benchmarks should be treated as indicative rather than confirmatory [16]. In practice, recruitment and eligibility tools are best suited to controlled site-level pilots with mandatory human verification rather than to autonomous inclusion or exclusion. At the site level this often translates into a small extra step: a coordinator confirms a model-prescreened patient against the source record before that patient is entered into the screening log. The two possible errors are not symmetrical. A false negative, meaning an eligible patient the model does not surface, is a recruitment loss and is largely invisible. A false positive that leads to the enrolment of an ineligible participant is an eligibility violation and, in inspection terms, a major protocol deviation. Under GCP the determination of eligibility remains an investigator responsibility and is not delegable to a system [39]. This asymmetry is the practical reason why a documented human decision should sit between the model output and the screening log (Figure 3).

3.5. AI in Trial Conduct, Monitoring and Data Integrity

A scoping review of AI in clinical trial risk assessment identified a substantial literature (142 studies) spanning safety, efficacy and operational risk, but documented fragmentation and limited prospective validation [7]. ML models for site risk prediction have been developed and externally validated for site qualification [18], and methods combining statistical outlier estimation with generative-AI content analysis have been proposed to characterize protocol deviations [22]. In data management, the literature describes a transition toward AI-supported clinical data science, with data cleaning and query generation that may reduce site burden and accelerate database lock [20,21]. These applications illustrate a recurring tension: each efficiency gain raises, rather than lowers, the governance burden — alerts require calibrated thresholds and escalation; AI-assisted data cleaning requires documented human review, traceability and validation; and any model used in oversight must itself be auditable. In practice, AI in monitoring and data management is best positioned as risk-alerting and quality tooling embedded within the existing quality management system rather than as autonomous oversight.

3.6. Patient-Facing AI: Consent, Education and Engagement

Patient-facing AI is among the most sensitive applications because it enters directly into communication with participants. LLMs have been evaluated as tools to improve the readability and actionability of informed-consent documents while preserving accuracy [23] and can generate educational materials that support understanding [26]. Conversational systems such as RESPECT show that retrieval-augmented, source-grounded consent assistants can be evaluated for safety and utility, including the capacity to decline out-of-scope questions [24]. Yet evaluations of LLM performance in the consent process reveal a gap: outputs may be highly accurate factually yet fall short on empathy, tone and contextual sensitivity, so an accurate answer is not necessarily acceptable [25]. Ethical analyses frame AI-supported consent as a balance between improved comprehension and the risks of undue influence and diffusion of responsibility [27]. International guidance on large multi-modal models emphasizes the governance obligations that arise in health communication [70]. Patient-facing AI should also be assessed against its effect on participant trust, comprehension and willingness to engage [44]. Patient-facing AI should therefore be treated as high-impact communication support requiring bounded, approved and source-grounded content, disclosure, escalation to site staff, and institutional review board and legal review. In a consent conversation this usually means that an AI system may draft or explain material, while a study nurse or investigator remains the person who answers the participant’s questions and obtains the signature.

3.7. Safety Surveillance and Decentralized Data

AI extends into safety surveillance and connects in-trial monitoring with post-marketing signal detection. Reviews describe AI, including LLMs, for adverse-event and adverse-drug-event detection, signal detection and reporting automation, while underscoring limits related to interpretability, data quality and model variability [33,34,35,36]. Because much of this evidence comes from pharmacovigilance and adjacent safety-surveillance contexts, its direct transfer to interventional trial operations should be made cautiously. In parallel, the growing use of digital health technologies and decentralized designs — evidenced by a systematic review of 262 rare-disease trials — generates new, continuous data streams that both enable and complicate AI-supported oversight [28], and regulator alignment on the validation of digital measures offers directly transferable lessons [29]. For this reason, safety-related AI should remain a triage or prioritization layer under investigator and safety-physician review, and the provenance and quality of decentralized data must be assured before such data drive AI-based safety decisions.

3.8. Risks and Failure Modes

The principal danger of AI in clinical trial operations is not simply that a model may err, but that its error may go unnoticed, undocumented or wrongly accepted by a human operator. The reviewed literature converges on a recognizable set of failure modes. Hallucination and mapping errors can produce confident but incorrect outputs — for example, incorrect protocol interpretation or inaccurate participant-facing information [15,25]. Automation bias may lead coordinators or investigators to over-trust recommendations, undermining the human oversight that is meant to make AI safe [12]; in AI-driven clinical decision support it has been identified as a plausible mechanism of patient harm when users over-rely on imperfect AI outputs [78]. Data bias can systematically disadvantage patients with sparser or lower-quality documentation, distorting eligibility and site-selection decisions [11,17]. Weak local validation means that a model performing well in publication may perform poorly at a given site, while a weak audit trail makes a decision impossible to reconstruct or defend during inspection [15,20]. Finally, model drift degrades performance after deployment, vendor lock-in and cybersecurity exposure introduce external dependencies and privacy risk, and accountability gaps leave it unclear who is responsible for an AI-influenced decision [4,5,32].
One further failure mode falls outside every governed pathway described above, and it may currently be the most common. Unsanctioned use, for example a coordinator pasting an excerpt of a source document into a publicly available chatbot in order to summarize or translate it, leaves no audit trail and is invisible to the sponsor and to monitoring. It also cannot be validated, because the site does not know which system or model version produced the output. Where the text contains participant data, such use may additionally constitute processing and transfer of personal data without a lawful basis under Regulation (EU) 2016/679 [85], and it departs from GCP expectations for confidentiality and data integrity [39]. Validating sanctioned tools does not address this failure mode at all. It calls instead for an explicit acceptable-use policy, stating which AI systems may be used, with which categories of data, by whom and for which tasks. Such a policy is likely to work better when it is accompanied by a sanctioned alternative for the tasks that staff would otherwise delegate to an unapproved tool. A governance framework that covers only approved AI may therefore cover the smaller part of the problem.

3.9. AI Autonomy, Trial Impact and Governance Intensity

Risk depends jointly on how much independent action an AI system takes and on how consequential that action is; we therefore treat autonomy and trial impact as two distinct axes and combine them (Table 4). Governance should be proportionate to autonomy, from administrative support through decision support and semi-automated workflows to fully autonomous action. Fully autonomous action is generally not recommended for critical tasks. Consistent with reporting standards for AI decision-support systems, each level should also specify the user role, the human–AI interaction, and explicit override and escalation criteria [60]. Autonomy alone, however, is insufficient: two tools at the same autonomy level may differ greatly in impact, as an internal administrative summary and a participant-facing consent explanation illustrate. Combining the axes yields a simple organizing principle — governance intensity should be determined by the interaction between AI autonomy and trial impact — so that high-impact uses, and any use approaching autonomous action on eligibility, safety or regulatory documentation, demand the most stringent validation, auditability and human accountability.

3.10. A Site-Level AI Readiness Framework, Deployment Model and Lifecycle

Sites — and the healthcare organizations that operate them — are not passive users of AI; they are the point at which AI outputs must be verified, documented and translated into accountable trial decisions. The following framework is proposed as a synthesis-derived implementation aid rather than a validated instrument. We propose a site-level readiness framework organized around eleven domains: intended use; data governance; AI-ready data quality; local validation; human oversight; workflow integration; documentation and audit trail; training; vendor qualification; safety escalation; and lifecycle monitoring. Because AI readiness begins with data rather than with the model, we separate general data governance from a distinct AI-ready data-quality domain, reflecting evidence that dataset completeness, standardization and fitness for purpose are prerequisites for reliable AI outputs [65]. For each domain the framework specifies a site-level question, the responsible role, the required documentation and the risk if the domain is neglected; the minimum documentation a site should assemble is summarized in Box 1, which operationalizes the domains and supports inspection-readiness. In practice, the dossier would be owned jointly by the site or healthcare organization’s research-governance function and the sponsor or contract research organization, with final accountability for trial conduct remaining aligned with GCP-defined sponsor and investigator responsibilities.
Box 1. Minimum site-level AI implementation dossier (readiness domains operationalized)
• Intended use: context-of-use and use-case description; risk classification (autonomy level and trial impact)
• Data governance: data-flow map and legal/ethical basis; acceptable-use policy defining permitted AI systems, permitted data categories and permitted users, including prohibition of unsanctioned tools
• AI-ready data quality: data-readiness assessment and quality metrics
• Local validation: validation plan and report, test cases and error analysis; change control and documented release for use
• Human oversight: who signs off; decision log; override and escalation SOP
• Workflow integration: process map and updated standard operating procedure
• Documentation / audit trail: model and version information; source references; versioning; AI use agreed with the sponsor and filed in the investigator site file and trial master file
• Training: training log and competency check; AI-related tasks reflected in the delegation log
• Vendor qualification: vendor and security assessment
• Safety escalation: defined escalation pathway
• Lifecycle monitoring: drift/performance review, incident log, update and retirement criteria
These domains become decision-relevant when combined with the preceding axes. Table 5 translates the framework into a concrete bridge between use case and action, mapping representative use cases — through their evidence maturity, autonomy and trial impact — to a recommended deployment status ranging from research-only to controlled or restricted operational deployment. Figure 4 shows the corresponding stage-gated pipeline (detailed in Table 6). Consistent with lifecycle-oriented guidance for trustworthy AI [63], the framework spans three phases: a pre-deployment phase (intended use, maturity, risk classification, AI-ready data, local validation and vendor qualification); a deployment phase (human oversight, SOPs, workflow integration and documentation); and a post-deployment phase (drift and performance monitoring, incident logging, periodic review and defined update or retirement criteria).
A framework is only useful at a site if it states, without ambiguity, what has to be true before an AI tool may be used on a real trial. We therefore specify minimum conditions for each transition, from research use to a supervised pilot and from a pilot to controlled operational use (Table 6). Three points in that table are deliberately strict. First, a pilot runs in parallel with the existing process and does not replace it, so that the AI can be wrong without harming the trial. Second, acceptance thresholds are agreed before testing begins, because a threshold chosen after seeing the results is not a threshold. Third, the conditions for withdrawing a tool are defined at the same time as the conditions for deploying it; a system that cannot be switched off in a defined way should not be switched on.
A tool that does not meet the conditions for a stage remains at the preceding stage. Autonomous action on eligibility, safety reporting or endpoint adjudication is outside this table and is not recommended.
The site cannot satisfy these conditions alone, and a framework that pretends otherwise will not be implemented. Trial sites operate inside a sponsor–CRO–site relationship in which accountability for trial conduct is defined by GCP and cannot be reallocated by contract [39]. Table 7 therefore sets out who decides, who executes and who verifies for each readiness domain. Two consequences deserve emphasis. Under GCP, tasks are delegated to people and recorded on a delegation log; an AI system is not a delegatee but a tool used by a delegated person, so accountability does not move to the model or to its vendor. And because a single site typically runs studies for several sponsors with different positions on AI, a site is usually better served by one internal AI policy calibrated to its most restrictive sponsor than by a separate arrangement for each study.
Accountability for trial conduct remains with the sponsor and the investigator under ICH E6(R3) [39]. Nothing in this allocation transfers that accountability to a contract research organization or to a technology vendor.

3.11. Regulatory, GCP and Reporting Alignment

The readiness of AI for clinical trial operations must be judged against the regulatory and GCP framework. The FDA draft guidance on AI to support regulatory decision-making introduces a defined context of use and a risk-based credibility assessment [37]. The joint FDA–EMA guiding principles articulate human-centric design, risk-based approaches, data governance, performance assessment and lifecycle management [38]. ICH E6(R3) provides the overarching, technology-agnostic GCP framework for quality by design, data integrity and sponsor/investigator responsibilities [39]. ICH E8(R1), with its emphasis on quality by design and critical-to-quality factors, maps particularly well onto AI readiness understood as quality planning [69]. The EMA reflection paper offers a European lifecycle perspective [40]. The EU Artificial Intelligence Act (in force from 1 August 2024) introduces risk-tiered obligations. These become relevant wherever operational AI approaches medical-device or high-risk status, and a need for sector-specific guidance in healthcare has been recognized [67]. The treatment of AI as a medical device clarifies the boundary of the present scope [6].
One practical consequence of this alignment is easily overlooked. Operational AI used in a trial is a computerised system, and computerised systems in regulated research already sit inside a mature compliance framework. ICH E6(R3) sets expectations for computerised systems used in trials, including risk-proportionate validation, system release, security, access control, audit trails and data-life-cycle integrity [39]. In the United States, electronic records and signatures relied upon in clinical investigations fall under 21 CFR Part 11 [81], and current FDA guidance addresses electronic systems, records and signatures in clinical investigations specifically [82]. In the European Union, Annex 11 sets corresponding expectations for computerised systems and is itself under revision [83], while risk-based validation of computerised systems in regulated environments is codified in GAMP 5 [84]. Treating AI as an operational layer therefore need not create a new compliance regime; it brings AI inside an existing one. In concrete terms, an AI tool used in trial conduct should have a documented intended use and risk classification, a validation plan and report, defined access control, an audit trail, change control and a documented release for use. These are the artefacts expected of any other regulated computerised system, and they are the artefacts that Box 1 assembles. What is new is not the requirement but its difficulty. A deterministic system can be validated against fixed expected outputs, whereas a generative model may not return the same output twice, and its behaviour may shift as the underlying data or the model itself changes. Validation therefore has to be specified at the level of the AI-enabled workflow, with pre-defined acceptance criteria, a human checkpoint and periodic re-qualification, rather than at the level of a single model output.
The applicability of the EU Artificial Intelligence Act to operational trial AI is narrower, and more specific, than is often assumed [86]. The Regulation does not apply to AI systems developed and put into service for the sole purpose of scientific research and development (Article 2(6)), nor to research, testing or development activity prior to placing on the market (Article 2(8)). Those exclusions cover AI as the object of research. They do not cover AI used as a tool to conduct research: Recital 25 provides that any other AI system that may be used for the conduct of a research activity remains subject to the Regulation. A commercial language model used to prescreen participants is therefore within scope. Whether it is high-risk is a separate question, and on the present text the answer is usually no. The Annex I route requires the system to be a product, or the safety component of a product, covered by sectoral legislation such as the Medical Device Regulation, and operational trial AI is normally not a device on the intended-purpose test [93]. The Annex III list does not clearly capture prescreening of trial participants: its eligibility provision is confined to essential public assistance benefits and services and to systems used by, or on behalf of, public authorities. AI used to monitor or evaluate the performance of site staff, by contrast, would be captured.
The practical consequence is uncomfortable. Most operational trial AI sits inside the AI Act but outside its high-risk regime, so the obligations that actually bite are the AI-literacy duty on deployers, which includes sponsors, contract research organizations and sites, and the transparency duties requiring that a person be told when they are interacting with an AI system, which is directly relevant to participant-facing consent tools. That is a thin set of duties for a technology that can influence who is offered access to a trial. It is one concrete sense in which an operational AI governance gap exists: the instrument most often invoked in discussion of AI in trials imposes, in this setting, less than is commonly assumed, while the instruments that do impose substantive obligations are GCP, data-protection law and the rules for computerised systems.
A further governance gap concerns reporting standards. Established guidelines — SPIRIT-AI and CONSORT-AI for trials of AI interventions [58,59], DECIDE-AI for early clinical evaluation of AI decision-support [60], and TRIPOD+AI for AI prediction models [61] — were designed for studies in which the AI system is the object of evaluation. For AI-based feasibility, site-risk or monitoring models in particular, such reporting standards should be complemented by appraisal tools such as PROBAST+AI, because transparent reporting alone does not establish low risk of bias or local applicability [79]. Most AI tools discussed here are not the intervention; they operate inside the trial as recruitment, monitoring, documentation or safety-triage aids, and no comparable, widely adopted standard yet governs the reporting and qualification of such operational AI workflows. The consistent message is that AI becomes acceptable only when it has a defined context of use, is validated for that context, is auditable, is subject to lifecycle management and remains under meaningful human oversight — precisely the properties the site-level framework is designed to demonstrate at the clinical trial site level.

3.12. Data Protection and Cybersecurity

Data protection is not a side condition of operational AI in trials. For many sites it is the binding constraint, and three questions determine whether an AI-enabled workflow is lawful at all. The first is the allocation of roles. The sponsor is normally a controller; the site or investigator is a controller, or a joint controller, for the source record; a contract research organization ordinarily acts as a processor; and an external AI vendor is a processor only for as long as it processes strictly on documented instructions [88]. The moment a vendor uses trial data to train or improve its own models, it pursues its own purpose and becomes a controller, which requires its own lawful basis for processing health data and disclosure to participants [85]. This cannot be argued away in a data-processing agreement. It has to be contractually excluded and then verified. The European Data Protection Board has addressed data-protection aspects of AI models, including when a model may be considered anonymous and when legitimate interest may be relied upon [87], but we are not aware of guidance addressing third-party AI vendors in clinical trial operations specifically.
The second question is where the data go and where they rest. Sending source data to an external AI service is a transfer, and where the service processes outside the European Economic Area it engages the transfer rules of the General Data Protection Regulation, which require an adequacy decision or appropriate safeguards together with an assessment of the destination. Data residency, retention and deletion must be specified. A site that cannot state where its prompts are stored, for how long they are kept, and whether they can be deleted, cannot demonstrate compliance. Given the scale of processing and the special category of the data, a data protection impact assessment is in practice mandatory for AI prescreening of trial participants. The third question is automated decision-making. Where an investigator genuinely reviews each candidate, the prohibition on solely automated decisions is normally not engaged. Where a model silently excludes patients who are never surfaced to any human, there is no human decision at all for those individuals; and the Court of Justice of the European Union has held that an automated score may itself constitute the decision where the person acting on it draws strongly upon it [92]. Whether non-enrolment in a trial amounts to a similarly significant effect has not been tested, and we do not assert that it does. But a site that cannot demonstrate meaningful human review is exposed to the question.
Cybersecurity belongs in the same discussion, because an external AI vendor is a supply-chain dependency. Under the NIS2 Directive, healthcare providers and entities carrying out research and development of medicinal products fall within the sectors of high criticality, and the risk-management obligations extend explicitly to supply-chain security and carry incident-reporting timelines [89]. For a trial site this converts vendor qualification from good practice into a legal expectation. Certification against information-security and AI-management standards, in particular ISO/IEC 27001 [94] and ISO/IEC 42001 [90], is a reasonable procurement signal, although it should not be mistaken for compliance with the AI Act, which has its own conformity route. Looking further ahead, the European Health Data Space will govern the secondary use of health data, including reuse for training and validating models, through data permits and secure processing environments [91]. It does not govern the primary use of data to run a trial, and the two should not be conflated.

4. Discussion

This review supports three main conclusions. First, AI in clinical trial operations is genuinely transitioning from isolated task automation toward a workflow-embedded operational layer [1,3,7]. Second, operational maturity is markedly uneven: recruitment and eligibility assessment are comparatively mature [9,10,11,12], whereas patient-facing consent tools, digital twins and synthetic controls remain emerging or high-risk [24,25,30,31]. Third, the barriers to responsible adoption are predominantly organizational and governance-related rather than purely technical [4,5,15].
The review’s contribution is to connect four dimensions rarely considered together — evidence maturity, AI autonomy, trial impact and site readiness — into a single deployment-decision model, situated within the broader movement toward trustworthy, lifecycle-governed healthcare AI [57,62,63,68].
A recurring theme is that implementation readiness cannot be inferred from model performance alone; clinical adoption requires a multi-faceted implementation evaluation rather than accuracy metrics in isolation [64]. The relevant unit of validation is the AI-enabled workflow — the model together with its source data, user role, standard operating procedure, escalation rules, documentation pathway and final human decision — rather than the model in isolation. Validating a model without validating the workflow around it provides false reassurance, because most failure modes in a regulated setting arise at the interfaces between the model and the humans, data and procedures that surround it. Earlier benchmark-style evaluations of large language models on medical examination tasks, including a national medical specialization examination, illustrate this distinction: the ability to answer domain-specific questions is not equivalent to readiness for regulated clinical-workflow deployment [73]. A similar translational-readiness problem has been described in AI-enabled obesity care, where AI was framed as part of a broader operational architecture requiring pathway integration, human oversight and governance rather than predictive performance alone [74].
This reframing implies a more demanding standard than explainability. For clinical trial operations, the operative requirement is not only explainable AI but inspection-ready AI. At any point, a monitor, auditor or inspector should be able to reconstruct the following: which model version produced a given output; what data were used; what was generated; who reviewed it; what the final decision was; what entered the trial documentation; and whether the output was modified or rejected. Inspection-readiness maps onto GCP expectations for data integrity and accountability and can be operationalized through the readiness domains, deployment mapping, pathway and dossier proposed here (Table 4 and Table 5, Figure 4, Box 1), and it requires a clear accountability chain across sponsor, site, investigator, vendor, monitor and data-protection roles [68]. The elements that make a workflow inspection-ready are shown in Figure 5.
Framed organizationally, healthcare organizations that host clinical trial sites are not merely physical locations for sponsor-led studies; they are complex adaptive sociotechnical systems in which AI-enabled workflows involve people, processes, technology, governance and organizational culture, not technology alone [77]. They provide the data infrastructure, workforce, privacy controls, quality systems, patient-facing communication channels and governance culture within which AI-enabled trial workflows are implemented. Site-level AI readiness is therefore also an organizational capability, and much of the readiness framework proposed here overlaps with the broader digital-health governance responsibilities of the hosting organization. Consequently, AI may reduce task-level burden while simultaneously increasing system-level governance burden: each efficiency gain introduces new obligations — validation, training, standard operating procedures, audit trails, vendor qualification, data security and drift monitoring — that fall disproportionately on sites and their workforce. Operational readiness is therefore not equivalent to regulatory readiness, and AI readiness in clinical trial operations should be assessed at the level of the workflow, not the model. Consistent with the human–AI teaming evidence, the framework treats AI as decision support that culminates in a documented human decision, not as an autonomous substitute for trial staff [12]. The risk, moreover, is not only that AI tools fail validation, but that they are adopted in pilots and later abandoned because they do not fit site workflows, organizational capacity or sustainability requirements — a pattern well described for health and care technologies more broadly [80].
This review does not claim that the proposed framework is sufficient for regulatory acceptance, nor that AI-enabled workflows should be adopted wherever evidence maturity is higher. Rather, it provides a structured way to identify the minimum site-level conditions that should be satisfied before operational use is considered.

4.1. Market Enthusiasm Versus Operational Readiness

Commercial and operational data help explain the pace of AI adoption without establishing readiness [42,45,46], and they raise the risk not only of slow adoption but of premature adoption. A structured comparison of market momentum against evidence maturity by domain is provided in Supplementary Table S3. Three layers follow from it. Market evidence, together with registry-level counts of AI-related trials (Supplementary Figure S1), explains why adoption is accelerating. Peer-reviewed evidence shows where AI is mature or immature. Regulatory, GCP and reporting principles define what must be in place before adoption becomes responsible. The relevant question is therefore whether AI-enabled workflows can be validated, audited and operated safely at the site level.
The market for AI-enabled clinical-trial solutions is expanding rapidly, but market growth should not be confused with operational maturity or regulatory readiness.

4.2. Implementation Outcomes for AI-Enabled Trial Workflows

Because clinical adoption depends on more than model performance [64], evaluation should extend to implementation outcomes. Such outcomes — distinct from model-performance outcomes — have been defined to include acceptability, adoption, appropriateness, feasibility, fidelity, implementation cost, penetration and sustainability [76]. Future evaluations should report implementation outcomes alongside technical metrics: time saved; the additional verification workload created for site staff; the number and proportion of AI outputs overridden or rejected; false-positive and false-negative eligibility flags; time to query resolution; the number of safety alerts escalated; staff acceptability and trust; training burden; equity and unintended consequences; and any inspection findings related to AI use. Reporting these outcomes would move the field from demonstrating that a model can perform a task to demonstrating that an AI-enabled workflow can be operated safely, sustainably and accountably at a site. A consolidated evidence-gap and research agenda is set out in Table 8.

4.3. Practical Implications for Healthcare Organizations and Trial Sites

For healthcare organizations and the trial sites they host, several practical implications follow. First, no AI-enabled workflow should be introduced without an explicitly defined context of use and risk classification. Second, AI should not be used for eligibility or safety decisions without a documented human sign-off. Third, models should be validated locally against the site’s own population and workflow before operational use. Fourth, every AI-influenced decision should leave an audit trail sufficient to reconstruct it during monitoring or inspection. Fifth, performance drift and the additional verification workload placed on staff should be monitored after deployment. These implications map directly onto the readiness domains (Box 1) and the deployment pathway (Figure 4) and are intended to be actionable at the level of a research-governance function. For a coordinator or investigator they are often felt as small but concrete changes to daily work — for example, a required sign-off field in the screening system before a prescreened patient is entered into the screening log, or a note in the trial master file recording which model version produced a flagged eligibility match.
Two further constraints determine whether any of this is actually done. The first is that a site is rarely a free agent. A single site may run many trials for several sponsors, each with its own position on AI, while the sponsor retains oversight of trial conduct and of delegated activities [39]. The use of an AI tool in a given trial should therefore be agreed with the sponsor before it is used. It should also be recorded in the investigator site file and trial master file, and reflected in the delegation and training records of the staff who operate it. The second constraint is cost. Local validation, vendor qualification, training, additional verification steps and drift monitoring consume staff time that is not currently a line item in most clinical trial agreements or site budgets. Where these activities are neither funded nor assigned to a named role, they are unlikely to be performed. Efficiency gained at the level of the model may then be offset, or even reversed, at the level of the site. This concern is consistent with reports of rising site burden [66], although we are not aware of empirical data quantifying the net effect of AI adoption on site workload.

4.4. Limitations

As a structured narrative review, this work does not provide an exhaustive, protocol-driven search or a quantitative synthesis and is subject to selection bias and to the rapid evolution of the field. The evidence-maturity score is a pragmatic mapping tool, not a validated instrument. The readiness framework, deployment mapping and pathway should likewise be read as a conceptual implementation aid rather than a validated tool. Future empirical work should test whether the framework improves implementation quality, audit-readiness, staff confidence, participant safety or inspection outcomes at real sites. In particular, the framework was not developed through a Delphi process, stakeholder consultation or empirical testing with site personnel, sponsors, monitors or regulators. The evidence supporting the strongest maturity rating also has a linguistic and geographical boundary. The recruitment and eligibility systems reviewed here were developed and evaluated predominantly on English-language clinical records and in a small number of centres. Source documents at many trial sites are written in the local language and use local abbreviations, templates and coding practices. Whether the reported performance transfers to non-English electronic health records has not been demonstrated, so these findings should not be assumed to generalize across languages and documentation cultures without local validation. Database yields are reported for PubMed/MEDLINE only; the other sources were used for citation chaining and grey-literature retrieval rather than as separately counted search streams, and this limits the reproducibility of the search beyond the primary database. Some supporting evidence derives from preprints or adjacent literatures such as pharmacovigilance and is used with caution, and market and registry figures are used only as contextual signals. The regulatory and data-protection analysis reflects the position at the time of writing; the implementation timetable and guidance for the EU Artificial Intelligence Act continue to evolve, and readers should verify the current position before relying on it. Further research should also develop standardized reporting for operational AI workflows and invest in AI literacy for the research workforce.

5. Conclusions

Artificial intelligence is moving from isolated task automation toward an operational layer that can influence how clinical trials are designed, recruited, monitored, documented and overseen. Its greatest current value lies in recruitment and eligibility assessment, where empirical evidence is strongest, while other applications remain emerging or context-dependent. Extending prior lifecycle-level frameworks, the distinctive contribution of this review is a site-level decision model linking evidence maturity, AI autonomy, trial impact and site implementation capacity to inspection-ready, GCP-aligned deployment decisions. Across all use cases, the decisive factor is not technological sophistication but governance. AI in clinical trial operations should therefore be treated as a managed, inspection-ready and GCP-aligned operational layer, not as an autonomous substitute for investigators and trial staff. In practice this means validating the whole workflow for a defined context of use, keeping it auditable, monitoring it across its lifecycle and embedding it in human accountability. Much of this does not require a new compliance regime so much as the disciplined application of an existing one. In clinical trial operations, AI readiness should be assessed at the level of the workflow, not the model, and closing that gap at the level of the trial site and its host healthcare organization is the central practical challenge for the responsible adoption of AI in clinical research.

Author Contributions

Conceptualization, S.W. and J.D.-K.; methodology, S.W.; investigation, S.W. and A.R.; writing—original draft preparation, S.W.; writing—review and editing, A.R. and J.D.-K.; supervision, J.D.-K. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable (review article not involving human participants or animals).

Data Availability Statement

No new data were created or analyzed in this study.

Use of Generative AI

Claude (Anthropic; Claude Opus 4.8), accessed July 2026, was used to support language editing, conceptual organization and the formatting of tables and figures. It was not used to generate original data or to perform independent evidence selection. All AI-assisted content was reviewed, verified and approved by the authors, who take full responsibility for the accuracy and integrity of the manuscript.

Conflicts of Interest

The authors declare no conflict of interest.

Abbreviations

AI, artificial intelligence; ML, machine learning; LLM, large language model; RAG, retrieval-augmented generation; EHR, electronic health record; RBQM, risk-based quality management; ADE, adverse drug event; DHT, digital health technology; SOP, standard operating procedure; GCP, Good Clinical Practice; FDA, U.S. Food and Drug Administration; EMA, European Medicines Agency; ICH, International Council for Harmonisation; OMOP CDM, Observational Medical Outcomes Partnership Common Data Model; DCT, decentralized clinical trial.

References

  1. Knapen, D.G.; van Kruchten, M.; de Groot, D.J.A.; Broekman, K.E.; Fehrmann, R.S.N. Artificial intelligence for clinical trial design, conduct, and analysis: a narrative review. ESMO Real. World Data Digit. Oncol. 2026, 11, 100682. [Google Scholar] [CrossRef] [PubMed]
  2. Olawade, D.B.; Fidelis, S.C.; Marinze, S.; Egbon, E.; Osunmakinde, A.; Osborne, A. Artificial intelligence in clinical trials: a comprehensive review of opportunities, challenges, and future directions. Int. J. Med. Inform. 2026, 206, 106141. [Google Scholar] [CrossRef] [PubMed]
  3. Lin, A.; Wang, Z.; Jiang, A.; et al. Large language models in clinical trials: applications, technical advances, and future directions. BMC Med. 2025, 23, 563. [Google Scholar] [CrossRef] [PubMed]
  4. Badani, A.; de Moraes, F.Y.; Vollmuth, P.; Chung, C.; Mansouri, A. AI and innovation in clinical trials. npj Digit. Med. 2025, 8, 683. [Google Scholar] [CrossRef] [PubMed]
  5. Sergi, C.M.; Sesso, H.D. Artificial intelligence and the future of clinical trials. Contemp. Clin. Trials Commun. 2025, 47, 101545. [Google Scholar] [CrossRef] [PubMed]
  6. Shanmugam, U.; Rajendran, M.K.; Natarajan, J.; Karri, V.V.S.R. Clinical trial design and regulatory requirements for artificial intelligence as a medical device: a PRISMA-ScR-guided scoping review of global guidance and evidence (2017--2025). J. Clin. Med. 2026, 15, 1937. [Google Scholar] [CrossRef] [PubMed]
  7. Teodoro, D.; Naderi, N.; Yazdani, A.; Zhang, B.; Bornet, A. A scoping review of artificial intelligence applications in clinical trial risk assessment. npj Digit. Med. 2025, 8, 486. [Google Scholar] [CrossRef] [PubMed]
  8. Lu, X.; Yang, C.; Liang, L.; Hu, G.; Zhong, Z.; Jiang, Z. Artificial intelligence for optimizing recruitment and retention in clinical trials: a scoping review. J. Am. Med. Inform. Assoc. 2024, 31, 2749–2759. [Google Scholar] [CrossRef] [PubMed]
  9. Jin, Q.; Wang, Z.; Floudas, C.S.; et al. Matching patients to clinical trials with large language models. Nat. Commun. 2024, 15, 9074. [Google Scholar] [CrossRef] [PubMed]
  10. Abdallah, M.; Nakken, S.; Georges, M.; et al. TrialMatchAI: an end-to-end AI-powered clinical trial recommendation system to streamline patient-to-trial matching. Nat. Commun. 2026, 17, 4472. [Google Scholar] [CrossRef] [PubMed]
  11. Gupta, S.; Basu, A.; Nievas, M.; et al. PRISM: Patient Records Interpretation for Semantic clinical trial Matching system using large language models. npj Digit. Med. 2024, 7, 305. [Google Scholar] [CrossRef] [PubMed]
  12. Parikh, R.B.; Kolla, L.; Beothy, E.A.; et al. Human-AI teaming to improve accuracy and efficiency of eligibility criteria prescreening for oncology trials. Nat. Commun. 2026, 17, 2306. [Google Scholar] [CrossRef] [PubMed]
  13. Qian, Q. Large language models in clinical trial recruitment: sociotechnical and economic framework development study. JMIR AI 2026, 5, e95899. [Google Scholar] [CrossRef] [PubMed]
  14. Datta, S.; Lee, K.; Paek, H.; et al. AutoCriteria: a generalizable clinical trial eligibility criteria extraction system powered by large language models. J. Am. Med. Inform. Assoc. 2024, 31, 375–385. [Google Scholar] [PubMed]
  15. Lee, K.H.; Jang, S.; Kim, G.J.; et al. Large language models for automating clinical trial criteria conversion to OMOP CDM queries: validation and evaluation study. JMIR Med. Inform. 2025, 13, e71252. [Google Scholar] [CrossRef] [PubMed]
  16. Authors. Retrieval-augmented large language models for evidence localization in clinical trial recruitment from longitudinal EHR narratives. arXiv 2026, arXiv:2604.05190. [Google Scholar]
  17. Hulstaert, L.; Twick, I.; Sarsour, K.; Verstraete, H. Enhancing site selection strategies in clinical trial recruitment using real-world data modeling. PLoS ONE 2024, 19, e0300109. [Google Scholar] [CrossRef] [PubMed]
  18. Yang, Z. Machine learning for site risk prediction in clinical trials: development, external validation, and operational application in site qualification. Int. J. Med. Inform. 2026, 211, 106314. [Google Scholar] [CrossRef] [PubMed]
  19. Gao, J.; Xiao, C.; Glass, L.M.; Harrison, E.M.; Sun, J. Matching clinicians with clinical trials using AI. Nat. Health 2026, 1, 290–299. [Google Scholar] [CrossRef]
  20. Musik, S.; et al. Bridging the past and future of clinical data management: the transformative impact of artificial intelligence. Open Access J. Clin. Trials 2025, 17. [Google Scholar] [CrossRef]
  21. Purri, M.; Patel, A.; Deurell, E. Leveraging AI to accelerate medical data cleaning: a comparative study of AI-assisted vs. traditional methods. arXiv 2025, arXiv:2508.05519. [Google Scholar]
  22. Spyroglou, I.I.; Řeháková, M.; Klapka, R.; et al. Protocol deviation outlier estimation combined with generative AI. BMC Res. Notes 2026, 19, 226. [Google Scholar] [CrossRef] [PubMed]
  23. Shi, Q.; Luzuriaga, K.; Allison, J.J.; et al. Transforming informed consent generation using large language models: mixed methods study. JMIR Med. Inform. 2025, 13, e68139. [Google Scholar] [CrossRef] [PubMed]
  24. Giorgi, S.; Ryan, K.; Kim, J.P. RESPECT: a conversational AI system for informed consent with accuracy, safety, and stakeholder-centered evaluation. npj Digit. Med. 2026, 9, 2691. [Google Scholar] [CrossRef] [PubMed]
  25. Moscatel, R.; Aryal, K.; Chen, D.; Detsky, A.S.; Quinn, K.L. Performance of a large language model in the informed consent process for participation in a clinical trial. npj Digit. Med. 2026. [Google Scholar] [CrossRef] [PubMed]
  26. Waters, M. AI meets informed consent: a new era for clinical trial communication. JNCI Cancer Spectr. 2025, 9, pkaf028. [Google Scholar] [CrossRef] [PubMed]
  27. Allen, J.W.; Schaefer, O.; Mann, S.P.; Earp, B.D.; Wilkinson, D. Augmenting research consent: should large language models be used for informed consent to clinical research? Res. Ethics 2025, 21, 644–670. [Google Scholar] [PubMed]
  28. Mao, X.; Zeng, D.; Wang, X.; et al. Digital health technology use in clinical trials of rare diseases: a systematic review. Commun. Med. 2025, 5, 449. [Google Scholar] [CrossRef] [PubMed]
  29. Hill, D.L.; Carroll, C.; Belfiore-Oshan, R.; Stephenson, D. Aligning with regulatory agencies for the use of digital health technologies in drug development: a case study from Parkinson’s disease. Front. Digit. Health 2025, 7, 1415202. [Google Scholar] [CrossRef] [PubMed]
  30. Akbarialiabad, H.; Pasdar, A.; Murrell, D.F.; et al. Enhancing randomized clinical trials with digital twins. npj Syst. Biol. Appl. 2025, 11, 110. [Google Scholar] [CrossRef] [PubMed]
  31. Delleani, M.; et al. Synthetic data for clinical research and innovation: opportunities, challenges and future directions. ESMO Real. World Data Digit. Oncol. 2025. [Google Scholar] [CrossRef]
  32. Pasculli, G.; Virgolin, M.; Myles, P.; et al. Synthetic data in healthcare and drug development: definitions, regulatory frameworks, issues. CPT Pharmacomet. Syst. Pharmacol. 2025, 14, 840–852. [Google Scholar] [CrossRef]
  33. Zitu, M.M.; Owen, D.; Manne, A.; Wei, P.; Li, L. Large language models for adverse drug events: a clinical perspective. J. Clin. Med. 2025, 14, 5490. [Google Scholar] [CrossRef] [PubMed]
  34. Algarvio, R.C.; Conceição, J.; Rodrigues, P.P.; Ribeiro, I.; Ferreira-da-Silva, R. Artificial intelligence in pharmacovigilance: a narrative review and practical experience with an expert-defined Bayesian network tool. Int. J. Clin. Pharm. 2025, 47, 932–944. [Google Scholar] [CrossRef] [PubMed]
  35. Schreier, O.; Yazdani, A.; Galdadas, I.; et al. Application of language models for the analysis of adverse drug events in pharmaceutical research and development: scoping review. JMIR AI 2026, 5, e77732. [Google Scholar] [CrossRef] [PubMed]
  36. Elbiach, O.; Grissette, H.; Nfaoui, E.H. Benchmarking large language models for adverse drug reaction extraction in social media and clinical texts. Results Eng. 2025, 28, 107362. [Google Scholar] [CrossRef]
  37. U.S. Food and Drug Administration. Considerations for the use of artificial intelligence to support regulatory decision-making for drug and biological products (draft guidance). In FDA; 2025. [Google Scholar]
  38. European Medicines Agency and U.S. Food and Drug Administration. Guiding principles of good AI practice in drug development. In EMA/FDA; 2026. [Google Scholar]
  39. International Council for Harmonisation. ICH harmonised guideline: Good Clinical Practice (GCP) E6(R3), Step 4; ICH, 2025. [Google Scholar]
  40. European Medicines Agency. Reflection paper on the use of artificial intelligence in the medicinal product lifecycle. In EMA; 2024. [Google Scholar]
  41. IQVIA Institute for Human Data Science. Global R&D Trends 2026. 2026. Available online: https://www.iqvia.com/insights/the-iqvia-institute/reports-and-publications/reports/global-r-and-d-trends-2026 (accessed on 9 July 2026).
  42. Smith, Z.P.; DiMasi, J.A.; Getz, K.A. New estimates on the cost of a delay day in drug development. 2024. Available online: https://csdd.tufts.edu (accessed on 9 July 2026).
  43. Society for Clinical Research Sites. Global Site Landscape Survey 2025. 2025. Available online: https://myscrs.org (accessed on 9 July 2026).
  44. Center for Information and Study on Clinical Research Participation. Perceptions & Insights Study 2025. 2025. Available online: https://www.ciscrp.org/2025-perceptions-and-insights-study (accessed on 9 July 2026).
  45. Grand View Research. Clinical Trials Market Size, Share & Trends Analysis Report. 2025. Available online: https://www.grandviewresearch.com/industry-analysis/global-clinical-trials-market (accessed on 9 July 2026).
  46. MarketsandMarkets. AI in Clinical Trials Market -- Global Forecast. 2024. Available online: https://www.marketsandmarkets.com/Market-Reports/ai-in-clinical-trials-market-42687548.html (accessed on 9 July 2026).
  47. Fortune Business Insights. Artificial Intelligence in Clinical Trials Market. 2025. Available online: https://www.fortunebusinessinsights.com/ai-in-clinical-trials-market-114081 (accessed on 9 July 2026).
  48. Fortune Business Insights. Decentralized Clinical Trials Market. 2025. Available online: https://www.fortunebusinessinsights.com/decentralized-clinical-trials-market-117303 (accessed on 9 July 2026).
  49. McKinsey (Ed.) Faster, smarter trials: modernizing biopharma’s R&D IT applications. 2025. Available online: https://www.mckinsey.com/industries/life-sciences/our-insights/faster-smarter-trials-modernizing-biopharmas-r-and-d-it-applications (accessed on 9 July 2026).
  50. Boston Consulting Group. Agentic AI in biopharma: game-changing efficiency. 2025. Available online: https://www.bcg.com/publications/2025/agentic-ai-in-biopharma-game-changing-efficiency (accessed on 9 July 2026).
  51. Deloitte. Tech Trends 2025: a life sciences & health care perspective. 2025. Available online: https://www.deloitte.com (accessed on 9 July 2026).
  52. Accenture. Reinventing R&D in the age of AI. 2025. Available online: https://www.accenture.com/us-en/insights/life-sciences/the-rd-opportunity (accessed on 9 July 2026).
  53. Choudhury, K.; Mahatole, S. AI-led selloff in contract research firms may be misjudging disruption risk. Reuters Wire 2026. [Google Scholar] [CrossRef]
  54. Kwon, C.-Y. Beyond the growth: a registry-based analysis of global imbalances in artificial intelligence clinical trials. Healthcare 2025, 13, 2018. [Google Scholar] [CrossRef] [PubMed]
  55. Maru, S.; Kuwatsuru, R.; Matthias, M.D.; Simpson, R.J. Public disclosure of results from artificial intelligence/machine learning research in health care: comprehensive analysis of ClinicalTrials.gov, PubMed, and Scopus data (2010--2023). J. Med. Internet Res. 2025, 27, e60148. [Google Scholar] [CrossRef] [PubMed]
  56. ClinicalTrials.gov; U.S. National Library of Medicine. Descriptive keyword search (“artificial intelligence”; “large language model”). 2026. Available online: https://clinicaltrials.gov/api/v2/studies (accessed on 9 July 2026).
  57. Mateen, B.A.; Moorthy, V.; Labrique, A.; Farrar, J. Artificial intelligence and clinical trials: a framework for effective adoption. Lancet Digit. Health 2025, 7, 100898. [Google Scholar] [CrossRef] [PubMed]
  58. Cruz Rivera, S.; Liu, X.; Chan, A.W.; Denniston, A.K.; Calvert, M.J. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI extension. Nat. Med. 2020, 26, 1351–1363. [Google Scholar] [CrossRef] [PubMed]
  59. Liu, X.; Cruz Rivera, S.; Moher, D.; Calvert, M.J.; Denniston, A.K. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat. Med. 2020, 26, 1364–1374. [Google Scholar] [CrossRef] [PubMed]
  60. Vasey, B.; Nagendran, M.; Campbell, B.; et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat. Med. 2022, 28, 924–933. [Google Scholar] [CrossRef] [PubMed]
  61. Collins, G.S.; Moons, K.G.M.; Dhiman, P.; et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024, 385, e078378. [Google Scholar] [CrossRef] [PubMed]
  62. You, J.G.; Hernandez-Boussard, T.; Pfeffer, M.A.; Landman, A.; Mishuris, R.G. Clinical trials informed framework for real world clinical implementation and deployment of artificial intelligence applications. npj Digit. Med. 2025, 8, 107. [Google Scholar] [CrossRef] [PubMed]
  63. Lekadir, K.; Frangi, A.F.; Porras, A.R.; et al. FUTURE-AI: international consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ 2025, 388, e081554. [Google Scholar] [CrossRef] [PubMed]
  64. van de Sande, D.; Chung, E.F.F.; Oosterhoff, J.; van Bommel, J.; Gommers, D.; van Genderen, M.E. To warrant clinical adoption AI models require a multi-faceted implementation evaluation. npj Digit. Med. 2024, 7, 58. [Google Scholar] [CrossRef] [PubMed]
  65. Domagalski, M.J.; Lu, Y.; Pilozzi, A.; et al. Preparing clinical research data for artificial intelligence readiness: insights from the NIDDK data-centric challenge. J. Am. Med. Inform. Assoc. 2025, 32, 1609–1616. [Google Scholar] [CrossRef] [PubMed]
  66. Florez, M.; Smith, Z.; Olah, Z.; Martin, M.; Getz, K. Quantifying site burden to optimize protocol performance. Ther. Innov. Regul. Sci. 2024, 58, 347–356. [Google Scholar] [CrossRef] [PubMed]
  67. van Kolfschooten, H.; van Oirschot, J. The EU Artificial Intelligence Act (2024): implications for healthcare. Health Policy 2024, 149, 105152. [Google Scholar] [CrossRef] [PubMed]
  68. Angus, D.C.; Khera, R.; Lieu, T.; et al. AI, health, and health care today and tomorrow: the JAMA Summit report on artificial intelligence. JAMA 2025, 334, 1650–1664. [Google Scholar] [CrossRef] [PubMed]
  69. International Council for Harmonisation. ICH harmonised guideline: general considerations for clinical studies E8(R1), Step 4. ICH, 2021. [Google Scholar]
  70. World Health Organization. Ethics and governance of artificial intelligence for health: guidance on large multi-modal models; WHO: Geneva, 2024. [Google Scholar]
  71. Wells, B.J.; Nguyen, H.M.; McWilliams, A.; et al. A practical framework for appropriate implementation and review of artificial intelligence (FAIR-AI) in healthcare. npj Digit. Med. 2025, 8, 514. [Google Scholar] [CrossRef] [PubMed]
  72. Wójcik, S.; Rulkiewicz, A.; Pruszczyk, P.; Lisik, W.; Poboży, M.; Domienik-Karłowicz, J. Beyond ChatGPT: what does GPT-4 add to healthcare? The dawn of a new era. Cardiol. J. 2023, 30, 1018–1025. [Google Scholar] [CrossRef] [PubMed]
  73. Wójcik, S.; Rulkiewicz, A.; Pruszczyk, P.; Lisik, W.; Poboży, M.; Domienik-Karłowicz, J. Reshaping medical education: performance of ChatGPT on a PES medical examination. Cardiol. J. 2024, 31, 442–450. [Google Scholar] [CrossRef] [PubMed]
  74. Wójcik, S.; Tomaszewska, M.; Rulkiewicz, A. Artificial intelligence-based risk stratification in obesity care: from diagnosis to personalised treatment pathways. Diagnostics 2026, 16, 1461. [Google Scholar] [CrossRef] [PubMed]
  75. Baethge, C.; Goldbeck-Wood, S.; Mertens, S. SANRA—a scale for the quality assessment of narrative review articles. Res. Integr. Peer Rev. 2019, 4, 5. [Google Scholar] [CrossRef] [PubMed]
  76. Proctor, E.; Silmere, H.; Raghavan, R.; et al. Outcomes for implementation research: conceptual distinctions, measurement challenges, and research agenda. Adm. Policy Ment. Health 2011, 38, 65–76. [Google Scholar] [CrossRef] [PubMed]
  77. Sittig, D.F.; Singh, H. A new sociotechnical model for studying health information technology in complex adaptive healthcare systems. Qual. Saf. Health Care 2010, 19, i68–i74. [Google Scholar] [CrossRef] [PubMed]
  78. Khera, R.; Simon, M.A.; Ross, J.S. Automation bias and assistive AI: risk of harm from AI-driven clinical decision support. JAMA 2023, 330, 2255–2257. [Google Scholar] [CrossRef] [PubMed]
  79. Moons, K.G.M.; Damen, J.A.A.; Kaul, T.; et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ 2025, 388, e082505. [Google Scholar] [CrossRef] [PubMed]
  80. Greenhalgh, T.; Wherton, J.; Papoutsi, C.; et al. Beyond adoption: a new framework for theorizing and evaluating nonadoption, abandonment, and challenges to the scale-up, spread, and sustainability of health and care technologies. J. Med. Internet Res. 2017, 19, e367. [Google Scholar] [CrossRef] [PubMed]
  81. U.S. Food and Drug Administration. 21 CFR Part 11: Electronic records; electronic signatures. Code of Federal Regulations, Title 21, Part 11, 1997. Available online: https://www.ecfr.gov/current/title-21/chapter-I/subchapter-A/part-11 (accessed on 9 July 2026).
  82. U.S. Food and Drug Administration. Electronic systems, electronic records, and electronic signatures in clinical investigations: questions and answers. Guidance for industry. FDA. 2024. Available online: https://www.federalregister.gov/documents/2024/10/02/2024-22562 (accessed on 9 July 2026).
  83. European Commission. EudraLex — The rules governing medicinal products in the European Union. Good manufacturing practice, Annex 11: Computerised systems. European Commission. 2011, Volume 4. Available online: https://health.ec.europa.eu/system/files/2016-11/annex11_01-2011_en_0.pdf (accessed on 9 July 2026).
  84. International Society for Pharmaceutical Engineering. GAMP 5: A risk-based approach to compliant GxP computerized systems, 2nd ed.; ISPE: Tampa, FL, USA, 2022. [Google Scholar]
  85. European Parliament and Council of the European Union. Regulation (EU) 2016/679 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data (General Data Protection Regulation). Off. J. Eur. Union L 119 2016, 4.5.2016, 1–88. Available online: https://eur-lex.europa.eu/eli/reg/2016/679/oj (accessed on 9 July 2026).
  86. European Parliament and Council of the European Union. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Off. J. Eur. Union L, 2024/1689, 12.7. 2024. Available online: https://eur-lex.europa.eu/eli/reg/2024/1689/oj (accessed on 9 July 2026).
  87. European Data Protection Board. Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models; EDPB: Brussels, Belgium, 2024; Available online: https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf (accessed on 9 July 2026).
  88. European Data Protection Board. Opinion 3/2019 concerning the questions and answers on the interplay between the Clinical Trials Regulation and the General Data Protection Regulation; EDPB: Brussels, Belgium, 2019; Available online: https://www.edpb.europa.eu/our-work-tools/our-documents/opinion-board-art-70/opinion-32019-concerning-questions-and-answers_en (accessed on 9 July 2026).
  89. European Parliament and Council of the European Union. Directive (EU) 2022/2555 on measures for a high common level of cybersecurity across the Union (NIS2 Directive). Off. J. Eur. Union L 333. 2022, p. 80--152. Available online: https://eur-lex.europa.eu/eli/dir/2022/2555/oj (accessed on 9 July 2026).
  90. International Organization for Standardization and International Electrotechnical Commission. ISO/IEC 42001:2023 --- Information technology --- Artificial intelligence --- Management system; ISO/IEC. Geneva, Switzerland, 2023.
  91. European Parliament and Council of the European Union. Regulation (EU) 2025/327 on the European Health Data Space. Off. J. Eur. Union L, 2025/327, 5.3. 2025. Available online: https://eur-lex.europa.eu/eli/reg/2025/327/oj (accessed on 9 July 2026).
  92. Court of Justice of the European Union. ECLI:EU:C:2023:957; Case C-634/21, SCHUFA Holding (Scoring). CJEU, judgment of 7 December 2023.
  93. Medical Device Coordination Group. MDCG 2019-11 Rev.1: Guidance on qualification and classification of software in Regulation (EU) 2017/745 (MDR) and Regulation (EU) 2017/746 (IVDR). European Commission. 2023. Available online: https://health.ec.europa.eu/document/download/1ffc5b6a-2c81-4b6f-9c1a-9b0a1a5b9e51_en (accessed on 9 July 2026).
  94. International Organization for Standardization and International Electrotechnical Commission. ISO/IEC 27001:2022 --- Information security, cybersecurity and privacy protection --- Information security management systems --- Requirements; ISO/IEC. Geneva, Switzerland, 2022.
Figure 1. Artificial intelligence as an operational layer in clinical trial operations. The clinical trial lifecycle (top layer) rests on an AI operational layer of workflow-embedded systems that support, augment or partially automate trial activities, which in turn rests on a foundation of regulatory and Good Clinical Practice requirements (ICH E6(R3), data integrity, human oversight, auditability, data protection and validation). AI is the connective tissue between lifecycle activities, not the foundation.
Figure 1. Artificial intelligence as an operational layer in clinical trial operations. The clinical trial lifecycle (top layer) rests on an AI operational layer of workflow-embedded systems that support, augment or partially automate trial activities, which in turn rests on a foundation of regulatory and Good Clinical Practice requirements (ICH E6(R3), data integrity, human oversight, auditability, data protection and validation). AI is the connective tissue between lifecycle activities, not the foundation.
Preprints 223723 g001
Figure 2. Evidence-maturity of AI use cases in clinical trial operations, positioned along the four maturity bands (emerging, limited, moderate, high). Market and research momentum is concentrated in use cases that have not yet reached the highest band: reaching a score of 7–8 requires prospective evaluation inside a live trial workflow with explicit governance reporting, and no use case in the current evidence corpus meets that standard (high band shown empty). Scores correspond to Table 3.
Figure 2. Evidence-maturity of AI use cases in clinical trial operations, positioned along the four maturity bands (emerging, limited, moderate, high). Market and research momentum is concentrated in use cases that have not yet reached the highest band: reaching a score of 7–8 requires prospective evaluation inside a live trial workflow with explicit governance reporting, and no use case in the current evidence corpus meets that standard (high band shown empty). Scores correspond to Table 3.
Preprints 223723 g002
Figure 3. The asymmetry of AI errors in patient recruitment. A false negative (an eligible patient the model does not surface) is a recruitment loss that is largely invisible; a false positive that leads to enrolment of an ineligible participant is an eligibility violation and a major protocol deviation subject to regulatory inspection. Because the determination of eligibility remains a non-delegable investigator responsibility under Good Clinical Practice, a documented human decision should sit between the AI prescreening output and the official screening log.
Figure 3. The asymmetry of AI errors in patient recruitment. A false negative (an eligible patient the model does not surface) is a recruitment loss that is largely invisible; a false positive that leads to enrolment of an ineligible participant is an eligibility violation and a major protocol deviation subject to regulatory inspection. Because the determination of eligibility remains a non-delegable investigator responsibility under Good Clinical Practice, a documented human decision should sit between the AI prescreening output and the official screening log.
Preprints 223723 g003
Figure 4. Stage-gated operationalization of an AI-enabled workflow. A use case moves from research use only, through a supervised parallel-run pilot, to controlled operational use, passing a tollgate at each transition (local validation; met acceptance thresholds). Defined off-ramps return the tool to a previous stage or retire it on threshold breach, performance drift, an unrevalidated model or vendor change, a security incident, or withdrawal of sponsor agreement. Detailed conditions are given in Table 6.
Figure 4. Stage-gated operationalization of an AI-enabled workflow. A use case moves from research use only, through a supervised parallel-run pilot, to controlled operational use, passing a tollgate at each transition (local validation; met acceptance thresholds). Defined off-ramps return the tool to a previous stage or retire it on threshold breach, performance drift, an unrevalidated model or vendor change, a security incident, or withdrawal of sponsor agreement. Detailed conditions are given in Table 6.
Preprints 223723 g004
Figure 5. The workflow-validation paradigm. Because a generative model is non-deterministic and may not return the same output twice, traditional software validation of the model alone is insufficient. The unit of validation is the AI-enabled workflow: the model together with defined source data, a human checkpoint, an escalation standard operating procedure and an audit trail. Validated as a whole, the workflow remains inspection-ready even when the model errs or drifts, because the human checkpoints and audit trail prevent an unreviewed output from contaminating the trial.
Figure 5. The workflow-validation paradigm. Because a generative model is non-deterministic and may not return the same output twice, traditional software validation of the model alone is insufficient. The unit of validation is the AI-enabled workflow: the model together with defined source data, a human checkpoint, an escalation standard operating procedure and an audit trail. Validated as a whole, the workflow remains inspection-ready even when the model errs or drifts, because the human checkpoints and audit trail prevent an unreviewed output from contaminating the trial.
Preprints 223723 g005
Table 1. Evidence-maturity scoring rubric (0–8). Class bands: 0–2 emerging; 3–4 limited; 5–6 moderate; 7–8 high. A score of 7–8 requires prospective or real-world evaluation, multi-site or external data, testing inside an operational workflow, and explicit governance reporting; no use case in this corpus met that standard.
Table 1. Evidence-maturity scoring rubric (0–8). Class bands: 0–2 emerging; 3–4 limited; 5–6 moderate; 7–8 high. A score of 7–8 requires prospective or real-world evaluation, multi-site or external data, testing inside an operational workflow, and explicit governance reporting; no use case in this corpus met that standard.
Criterion 0 points 1 point 2 points
Evidence type Conceptual / preprint Retrospective / technical validation Prospective / user / real-world evaluation
Data realism Synthetic / demonstration Single-centre real-world Multi-site / external data
Workflow integration Offline model only Simulated workflow Operational workflow testing
Governance reporting Absent Partial Explicit oversight / auditability / validation
Table 2. AI applications across the clinical trial lifecycle.
Table 2. AI applications across the clinical trial lifecycle.
Stage AI use case Potential value Main risk Required governance
Protocol design Eligibility simplification, protocol assessment Better feasibility, reduced complexity Misinterpretation, hallucination Expert review, version control
Feasibility / site selection Patient-pool, site and investigator ranking Better site choice, realistic recruitment Historical/data bias Transparent criteria, local validation
Recruitment / eligibility Patient-trial matching, criterion-level screening Faster screening, better matching False positives / false negatives Human verification, audit trail
Conduct / monitoring RBQM, deviation detection, site risk scoring Proactive oversight Alert fatigue, missed issues SOP thresholds, escalation pathways
Data management Data cleaning, query generation Reduced site burden, faster database lock Erroneous queries Human review, traceability
Safety Adverse-event / ADE signal support Earlier signal detection Over- or under-reporting Safety physician review
Consent / education Consent-form generation, chatbot support Readability, accessibility Hallucination, undue influence Approved content boundaries, escalation
Table 3. Evidence-maturity map of AI use cases (scored per Table 1; subscores in Supplementary Table S2).
Table 3. Evidence-maturity map of AI use cases (scored per Table 1; subscores in Supplementary Table S2).
Use case Score (0–8) Class Evidence basis Key sources
Recruitment / eligibility 6 Moderate Multiple empirical systems; randomized human–AI teaming evaluation on retrospective data; not yet tested inside a live trial workflow [9,10,11,12,14]
Site selection / feasibility 5 Moderate ML ranking and externally validated site-risk models; no prospective operational deployment [17,18,19]
Safety surveillance / ADE 4 Limited Largely offline benchmarks and reviews from pharmacovigilance rather than trial conduct [33,34,35,36]
RBQM / monitoring / deviations 3 Limited Scoping review plus targeted methods papers; no prospective evaluation [7,22]
Data management / cleaning 3 Limited Efficiency evidence including a preprint; governance reporting absent [20,21]
Patient-facing consent AI 3 Limited Early LLM/RAG evaluations of high-stakes communication [23,24,25]
Digital twins / synthetic controls 2 Emerging Conceptual and review evidence; limited routine use [30,31,32]
Table 4. AI autonomy, trial impact and governance intensity.
Table 4. AI autonomy, trial impact and governance intensity.
Axis and level Description / example Governance requirement
Autonomy — L1 (administrative support) Protocol summary, draft communication Human review, version control
Autonomy — L2 (decision support) List of potentially eligible patients Documented final human decision, validation, override rule
Autonomy — L3 (semi-automated) AI prescreening with escalation rules SOP, thresholds, audit trail, escalation pathway
Autonomy — L4 (autonomous action) AI includes/excludes a participant without review Generally not recommended for critical tasks without exceptional validation and regulatory justification
Trial impact — Low Internal administrative summaries, scheduling, non-regulatory drafts Light-touch human review
Trial impact — Moderate Prescreening lists, query prioritization, site risk alerts Documented human decision, validation
Trial impact — High Eligibility support, consent communication, safety triage, regulatory documentation Validation, audit trail, explicit oversight
Trial impact — Critical Autonomous inclusion/exclusion, SAE reporting decisions, endpoint adjudication Generally not delegated to AI without exceptional justification
Table 5. From AI use case to site-level deployment recommendation (a synthesis-derived aid, not a validated instrument).
Table 5. From AI use case to site-level deployment recommendation (a synthesis-derived aid, not a validated instrument).
Use case Evidence maturity AI autonomy Trial impact Deployment recommendation
Patient–trial matching / eligibility Moderate Decision support Moderate–high Controlled pilot with mandatory human verification; not autonomous inclusion/exclusion
Site selection / feasibility Moderate Decision support Moderate Decision support only; transparent criteria and local validation
RBQM / risk alerts Limited Semi-automated alerting Moderate–high Risk alerting only; no autonomous oversight
AI-assisted data cleaning Limited Semi-automated Moderate Human-reviewed and audit-trailed; validated pipeline
AI-generated consent / education Limited Administrative support High Bounded content, IRB/legal review, disclosure and escalation
Safety / ADE triage Limited Decision support High Triage/prioritization layer under safety-physician review
Digital twins / synthetic controls Emerging Decision support High Research use only; not routine site workflow
Table 6. Minimum conditions for moving an AI-enabled workflow from research use to pilot and to controlled operational use.
Table 6. Minimum conditions for moving an AI-enabled workflow from research use to pilot and to controlled operational use.
Stage Evidence and validation Error tolerance Human oversight Exit / rollback criteria
Research use only Published or internal evidence of any maturity; no local validation required Not applicable: outputs must not influence any trial decision or enter trial documentation Outputs are not used in trial conduct Not applicable
Supervised pilot (parallel run) Local validation on the site’s own retrospective data against a documented reference standard; acceptance thresholds agreed with the sponsor before testing begins Pre-specified and impact-dependent. For eligibility, the tolerated number of ineligible participants entering the screening log is zero; model sensitivity and specificity are reported and compared with the site’s current process. For moderate-impact tasks, thresholds are set locally and justified 100% human review of every AI output; the AI runs alongside the existing process and does not replace it Pilot stops if acceptance thresholds are not met, if any unreviewed output reaches trial documentation, or if an eligibility- or safety-relevant error occurs
Controlled operational use Acceptance thresholds met in the pilot and confirmed on a prospective sample at the site; released for use under change control Continuous monitoring against the same pre-specified thresholds; a breach reverts the tool to pilot conditions Risk-proportionate: full review for high-impact outputs; documented sampling for moderate-impact outputs, with full review of all flagged and all discordant cases Defined triggers for suspension and retirement: threshold breach, performance drift, model or vendor change without revalidation, security incident, or withdrawal of sponsor agreement
Table 7. Allocation of responsibility for AI-enabled trial workflows across site, sponsor, CRO and technology vendor.
Table 7. Allocation of responsibility for AI-enabled trial workflows across site, sponsor, CRO and technology vendor.
Readiness domain Site / investigator Sponsor CRO Technology vendor
Intended use and risk classification Proposes the use case; the investigator confirms clinical relevance and acceptability Agrees to the use for the specific trial; confirms fit with the protocol Advises; reflects the agreed use in the monitoring plan States the intended purpose, the operating limits and the known failure modes
Data protection and legal basis Controller or joint controller for the source record; performs the impact assessment Controller; determines the purposes of processing Processor for the sponsor, unless acting for its own purposes Processor only; no training or model improvement on trial data without a separate lawful basis
Local validation Executes validation on its own data and population Reviews and accepts the validation report May perform validation on the sponsor’s behalf Supplies performance data, model and version information, and known limitations
Human oversight The investigator remains accountable; oversight is not delegable to a system Defines the minimum oversight required in the protocol and monitoring plan Verifies at monitoring visits that the required review actually occurred Provides override, escalation and audit functions
Documentation and audit trail Files the dossier in the investigator site file; retains the decision log Files in the trial master file; retains accountability for records Ensures trial master file completeness Guarantees an audit trail and the export of records
Training and delegation Trains staff; AI-related tasks appear in the delegation and training records Confirms at the site initiation visit Checks training records during monitoring Provides training materials and version-specific documentation
Incidents and CAPA Reports AI errors as quality issues within the existing quality management system Performs root-cause analysis; owns corrective and preventive actions Escalates and tracks to closure Reports defects, model changes and security incidents
Lifecycle and retirement Monitors drift; suspends the tool when a trigger is met Approves continued use; approves retirement Reports performance and incidents through monitoring Notifies model updates, deprecation and end of support
Table 8. Evidence gaps and research agenda for AI-enabled clinical trial operations.
Table 8. Evidence gaps and research agenda for AI-enabled clinical trial operations.
Domain Current evidence gap Recommended study design Minimum outcomes to report
Eligibility matching Limited prospective site-level validation Prospective multi-site implementation study Accuracy, screen-failure rate, staff workload, override rate
Consent AI Limited participant-level evidence Controlled comprehension and usability study Comprehension, trust, escalation, undue influence
RBQM / risk alerts Limited impact on monitoring quality Prospective operational evaluation False-alert rate, missed issues, deviation detection
Data cleaning Preprint- and efficiency-heavy evidence Controlled workflow evaluation Query accuracy, time saved, human-correction rate
Safety triage Uncertain transfer from pharmacovigilance Trial-specific safety-workflow study Escalation appropriateness, false-negative rate
Digital twins / synthetic data Limited routine operational use Regulatory-grade validation study Validity, bias, acceptability, regulatory acceptance
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings