Preprint
Review

This version is not peer-reviewed.

Governance of Artificial Intelligence in Clinical and Therapeutic Research: A Systematic Evaluation of Principles, Accountability, and the Implementation Gap

Submitted:

29 April 2026

Posted:

30 April 2026

You are already at the latest version

Abstract
Artificial intelligence is fast expanding in clinical research and medicinal development. In response, a considerable governance literature has arisen, characterised by ambitious theoretical frameworks but persisting gaps in practical implementation. This critical analysis assesses the underlying assumptions, organisational constraints, and institutional flaws that undermine responsible AI governance in healthcare and clinical research. The analysis combines findings from AI ethics, organisational governance, computational toxicology, clinical trial methodology, and patient safety science. The core thesis is that, despite significant agreement among governments, corporations, and academia on stated objectives, the responsible AI field has persistently failed to bridge the gap between normative goals and organisational realities. In clinical settings, this failure has direct consequences for patient safety. The analysis is structured around five interconnected critiques: the conceptual inadequacy of the performance-centric evaluation paradigm, which conflates statistical reliability with clinical safety; the inadequacy of explainability methods as substitutes for genuine accountability; the practical unimplementability of principled administrative frameworks in most healthcare research institutions; and the characterisation of regulatory fragmentation as a political economy. Drawing on a large body of research, the review suggests that solving the governance gap in clinical AI requires facing more fundamental assumptions than current studies acknowledges.
Keywords: 
;  ;  ;  ;  ;  

1. Introduction

The rapid integration of artificial intelligence into clinical care and therapeutic research has resulted in a fundamental dilemma. Although AI systems are becoming increasingly adept at forecasting adverse drug reactions, monitoring safety signals in clinical trials, assisting diagnostic decision-making, and producing clinical documentation at scale, current regulatory and governance frameworks are insufficient to ensure their safe, equitable, and accountable implementation. While this issue is well acknowledged in the responsible AI field [1,2,3,4,5], many evaluations fail to carefully evaluate the sufficiency of the frameworks.
The dominant narrative portrays the governance gap as a temporary institutional delay in which technical innovation outpaces regulatory adjustment while anticipating harmonisation through improved principles, clearer directions, and strengthened implementation mechanisms. This review argues that such an interpretation is overly simplistic. In clinical settings, the consequences of this misinterpretation are far more serious than is typically recognised in the literature.
Rather than being primarily caused by institutional delays, the governance gap in clinical AI reflects deeper conceptual limits in how responsible AI is defined, operationalised, and evaluated across the lifecycle of high-risk clinical and biomedical systems. Papagiannidis et al. [1] define responsible AI governance as a set of structural, procedural, and relational procedures that include decision-making authority allocation, data governance and compliance systems, and stakeholder engagement processes. While this tripartite approach, based on known IT governance theory [6,7], provides a valuable organisational framework, its deployment in clinical and therapeutic research settings is hindered by a number of structural restrictions. These include the irreversibility of patient harm, the difficulty of determining causality in AI-assisted decision-making, the unequal distribution of model faults across patient groups, and the restricted interpretability of large-scale models in clinical workflows.
Clinical AI governance is based on two major academic lineages. The first, from organisational information systems research, focuses on governance structures, antecedents, and outcomes [1], which provides useful insight into institutional implementation processes but is less appropriate for the prescriptive and high-stakes character of healthcare regulation. The second, derived from AI ethics, focuses on normative values like as transparency, justice, responsibility, and human oversight. However, empirical research [8,9] demonstrate that, while these principles are broadly agreed upon, there is significant debate about how they should be implemented, enforced, and resolved when values conflict. As a result, the primary problem is not to choose between these traditions, but to solve the institutional, epistemological, and distributional conditions that require governance to promote patient safety rather than just institutional conformity.
Despite these limitations, the responsible AI governance literature has achieved significant advances. The convergence of key ethical principles, as reported by Jobin et al. [9], Fjeld et al. [10], and Floridi and Cowls [11], is a significant normative achievement. The organisational framework presented by Papagiannidis et al. [1] takes implementation-oriented thinking beyond previous principle-based methods. Arrieta et al. [12] and Adadi and Berrada [13] examined developments in explainable AI, which have enhanced system openness, while regulatory proposals such as the EU AI Act [14,15] reflect major institutional advances. However, this analysis contends that these developments do not adequately address the deeper structural and epistemological issues that continue to limit the effectiveness of clinical AI governance.
The investigation is organised into five interconnected elements of governance failure in clinical AI. It initially criticises the performance-centric evaluation paradigm for its insufficiency as a framework for patient safety. It then investigates the limitations of explainability as an alternative to genuine accountability. It takes into account the organisational restrictions that hinder the practical application of responsible AI governance frameworks in the majority of clinical research contexts. It then analyses regulatory fragmentation as a matter of political economy rather than technological coordination. Finally, it discusses the distributional repercussions of governance failures, namely the disproportionate burden felt by various patient populations when AI systems fail. The assessment concludes by highlighting the requirements for a more appropriate clinical AI governance structure that goes beyond present assumptions in the area.

2. The Performance Evaluation Paradigm and Its Patient Safety Implications

2.1. An Epistemological Foundation with Structural Flaws

At the heart of clinical AI evaluation is an epistemic commitment to performance measurement using held-out test datasets. Deployment decisions are based on aggregate measures such as area under the ROC curve, sensitivity, specificity, positive predictive value, and comparable statistical summaries of classification success. This idea is derived directly from machine learning benchmarks, in which test set performance is considered the gold standard for model quality.
When transplanted into clinical governance, however, that norm is not just imperfect, but structurally incompatible with what patient safety requires. The test set is not the actual deployment environment. It is a statistical sample from a distribution that, while steady at the time of data collection, will predictably and continually diverge from the distribution found in actual clinical practice [16,17,18].
This mismatch is not a new discovery. Chekroud et al. [16] have established a larger trend of "illusory generalizability", the systematic overestimation of real-life clinical prediction performance because test data is too similar to training data across many clinical domains and prediction tasks. Consider medication toxicity predictions. Models trained on benchmark datasets like SIDER [19] and validated by systems like toxCSM [20] or CSM Toxin [21] produce performance metrics that instill actual clinical trust. However, their real-world performance in patient populations under-represented in those standards may be significantly lower. Hepatotoxicity models [22,23,24,25,26], nephrotoxicity predictors [27], cardiotoxicity classifiers [28], and cutaneous reaction models [29] all exhibit the same characteristics: sophisticated methodology optimised for benchmark performance, limited validation outside training distributions, and administrative frameworks that equate benchmark metrics with deployment safety.
What distinguishes this from a technical constraint is how performance loss emerges in clinical implementation. When an AI model trained on a relatively homogeneous patient population is used in a more diversified clinical scenario, it does not indicate that its performance is declining. It continues to generate outcomes with the same apparent confidence as in the verified environment [17]. A clinician who receives a confident AI prediction of minimal hepatotoxicity risk for a patient outside the model's training distribution has no way of knowing based only on model output that this prediction was derived under epistemic degradation conditions. The governance structure that permitted deployment based on benchmark performance did not provide a mechanism for detecting this scenario. The clinical institution receives no routine notification that its installed AI system is operating outside of validated parameters. This is not a failure of any one component of the governance system. It is a structural flaw in a governance design that regards test set performance as an appropriate proxy for deployment safety [30,31,32]. Figure 1 below depicts the governance gap in clinical AI, from principles to practice.

2.2. Calibration: The Neglected Patient Safety Dimension

Calibration, defined as the connection between asserted confidence and empirical chance of being true, has received surprisingly little governance attention, despite its obvious implications for healthcare decision making. Calibration is conceptually independent of discrimination. However, it has a substantial impact on how clinicians interact with AI outcomes [17,33]. A model that consistently assigns 85% confidence to empirically correct predictions just 60% of the time misinforms physicians about individual forecasts, as well as the epistemic standing of all AI recommendations. It constantly influences the balance of human-AI collaborative decision making in ways that aggregate accuracy assessments cannot detect.
The technical robustness dimension of responsible AI governance [1,34,35] includes exactness and trustworthiness as subdimensions of technical safety. However, both formulations fail to distinguish between discriminative accuracy (the capacity to correctly rank cases) and calibration accuracy (the relationship between reported confidence and empirical probability). The clinical AI literature has acquired this conflation from machine learning practice, where calibration has traditionally been viewed as a secondary concern. Systematic miscalibration has serious implications for clinical governance. A badly calibrated adverse event prediction algorithm produces mistakes with misleading confidence features, distorting clinical judgement throughout the whole patient group it examines [36,37].
Uncertainty measurement methods, such as Bayesian neural networks, conformal prediction, and ensemble variance estimation, provide technical approaches to calibration, and the AI development literature has made significant advances in this area [38,39,40]. The governance gap is not technical but rather institutional. There are no defined calibration evaluation requirements for the chemical spaces, demographic distributions, and clinical situations essential to deployment. In clinical AI publications, calibration is not required to be reported with discrimination measurements. There is no post-deployment monitoring infrastructure that can detect systematic miscalibration in ordinary clinical practice. The epistemic infrastructure needed to control calibration as a patient safety feature simply does not exist, and responsible AI administrative frameworks that do not expressly mandate its development remain aspirational rather than operational [18,30].

2.3. How Aggregate Metrics Conceal Equity-Relevant Failures

A third issue with the performance evaluation paradigm, which clinical AI administrative frameworks have not adequately addressed, is the systematic concealing of clinically and ethically significant subgroup performance variation inside aggregate measures. The responsible AI principles of diversity, non-discrimination, and equity outlined across the global landscape of AI ethics guidelines [9,10] and operationalised through sub dimensions such as accessibility and lack of unfair bias [1,34] state that AI systems should not reproduce the discriminatory patterns encoded in historical clinical data. The operationalisation challenge is that aggregate performance measures can remain stable, or even improve, while performance for specific demographic, pharmacogenomic, or comorbidity defined subgroups deteriorates, exactly for the populations that responsible AI principles are most concerned with protecting.
The technique of concealing is systemic, not incidental. Clinical AI models are typically trained on datasets that under-represent women, the elderly, racial and ethnic minorities, and patients from low-income nations. This reflects decades of systematic under-representation in clinical research [41,42,43]. When such models are assessed on test data from the same under-represented source, aggregate metrics will fail to detect the performance degradation that under-represented populations incur during deployment. Adverse event reporting adds another layer of systematic bias: reporting rates differ depending on drug type, patient demographics, geographic region, and clinical environment [42,43]. Models trained on pharmacovigilance data may thus learn from systematically skewed outcome labels, in which the absence of a documented adverse event signals reporting failure rather than clinical safety. Regulatory frameworks that rely on aggregate performance measures based on such data are more than just technically inadequate. They are systematically producing false safety proof for the same patient populations whose protection is the most pressing moral argument for clinical AI oversight.
Feuerriegel, Dolata, and Schwabe's research of fair AI [44] uncovers additional complexity that aggregate metric frameworks habitually obfuscate. In many realistic deployment scenarios, the key mathematical concepts of algorithmic equity, demographic parity, equalised chances, individual fairness, collective calibration, and counterfactual fairness are clearly incompatible. Choosing a fairness criterion is thus not a technical optimisation challenge, but rather a values decision that necessitates political debate on whose conception of fairness should guide AI deployment in certain therapeutic contexts. This is a value decision that technical performance appraisal frameworks are unprepared to make, and administrative frameworks have yet to build suitable procedures to navigate [45,46].

3. The Accountability Deficit: Structural Issues and the Limitations of Technical Remedies

3.1. Distributed Causation and the Inadequacy of Existing Liability Frameworks

Few concerns in clinical AI governance are as practical and institutionally unresolved as this one: who is responsible when an AI system contributes to a clinical adverse event? The causal chain linking an AI system's output to a patient result is usually spread over a complicated sociotechnical system. There are algorithm developers who made important design decisions during training. There are data curators who helped shape the training dataset. There are validation teams that have developed deployment approval requirements.
Some institutional administrators have decided to integrate AI into medical workflows. Clinical informaticists have configured the deployment environment. And there are clinicians who acted on AI results, typically under time constraints that prevented them from critically evaluating the recommendations they got [47,48]. The current frameworks for medical professional liability were established to address the carelessness of individual practitioners who made identifiable clinical decisions with recorded justifications for specific patients. They are structurally inadequate for situations where harm is caused by the interaction of multiple agents in a distributed system where no single agent made the decision that caused the harm, and where the AI system's contribution cannot be reconstructed without audit infrastructure, which most clinical institutions lack.
This lack of responsibility is more than a legal technicality that needs to be addressed. It has direct patient safety implications that the organisational governance literature has not adequately examined. When institutions are unable to properly link AI-related bad occurrences to specific actors, systems, or design decisions, the institutional learning mechanisms that ordinarily drive patient safety improvement are disrupted. Adverse events involving AI contributions may be attributed to the clinical decision maker who acted on the AI recommendation who may not have realised that their judgement was determined by an AI system operating outside of its validated conditions rather than to systematic AI performance issues that would emerge from aggregate analysis if AI-related incidents were systematically identified, coded, and compared [47,49]. This systematic event detection architecture underpins the accountability measures identified by Papagiannidis et al. [1] as important to responsible AI governance, including audit trails, incident reporting, role and responsibility assignment, and AI oversight committee assessment. In its absence, they serve as administrative formalities rather than patient safety procedures.

3.2. Explainability as a Substitute for Accountability

The primary technical answer to the clinical AI accountability issue has been the invention and institutional promotion of explainability approaches such as SHAP, LIME, attention visualisation, counterfactual explanations, and their variants. These are intended to make AI decision making more open to clinical inspection while also providing the transparency that accountability requires. The substantial technical literature on explainable AI [12,13] is a genuine methodological achievement; these tools provide previously unavailable information about how models generate specific outputs from specific inputs, and they contribute meaningfully to understanding AI system behaviour. The transparency principle, as formulated by Papagiannidis et al. [1] across many responsible AI frameworks [34,35], correctly recognises explainability as a major governance component via the subdimensions of explicability, traceability, and communication.
The main issue is that explainability and accountability are frequently confused in both technical and governance literatures. Transparency in the sense of releasing information about model operations does not imply accountability in the sense of establishing institutional processes for using that knowledge to challenge, rectify, or seek recourse for AI-generated suggestions that cause harm. These are separate properties. Conflating them results in administrative frameworks that misinterpret the provision of technical knowledge for the institutional architecture that accountability demands [50,51,52].
Consider a SHAP explanation that attributes a hepatotoxicity risk prediction to specific molecular characteristics of a medication candidate. It describes the model's internal logic for that particular prediction. What it does not provide is information about whether that logic has been clinically validated as a biomarker of the predicted toxicity, whether the model is operating within the demographic distribution for which its logic has been validated, whether the confidence attached to the prediction reflects actual empirical probability, or whether the patient's clinical context introduces considerations that the model's training data did not represent [4,53]. A therapist equipped with this SHAP explanation cannot do the critical evaluation required for accountability.
This is not due to personal deficiency. It is because the information presented is insufficient for such review, and clinical practice's cognitive conditions do not allow for persistent critical engagement with algorithm-based reasoning, as explainability frameworks presuppose [54,55].
The empirical literature on human-AI interaction in therapeutic settings provides scant support for the critically engaged clinician's vision of using explainability data to make informed override decisions. Automation bias-systematic over reliance on AI outputs even when those outputs are plainly erroneous, particularly under time constraints and cognitive demand is well reported in a variety of clinical situations and AI system types [33,54,55].
A parallel challenge to algorithm aversion is the systematic underuse of AI guidance by clinicians who fear AI in ways that restrict the realisation of actual AI advantages [56]. The governance challenge of calibrating clinical trust in AI is not addressed by providing better explanations for AI outputs. It involves intervention at the levels of human AI interface design, clinical training, institutional culture, and incentive structures, all of which are recognised as crucial in the governance literature but have yet to be translated into implementable standards [57,58,59].

3.3. Large Language Models and the Emergence of Irreducible Accounting Complexity

The rapid adoption of large language models in clinical procedures for clinical record generation, electronic health record summarisation, patient communication support, and clinical reasoning assistance presents accountability challenges that far outweigh those posed by supervised classification models. Existing responsible AI administrative frameworks have yet to develop appropriate tools to solve them. The hallucination phenomenon-the generation of factually incorrect information in linguistically fluent, contextually plausible, and syntactically coherent form represents a type of clinical AI failure with governance properties not seen in any previous clinical AI system [48,60].
Unlike errors in supervised classification models, which produce discrete outputs with associated certainty scores that can be statistically tracked, LLM hallucinations occur at the level of individual linguistic tokens embedded within complex clinical narratives. They generate errors that are indistinguishable from proper outputs under the cursory examination that medical processes frequently allow. They are incorporated into otherwise correct clinical text. Furthermore, they propagate through subsequent judgements without producing the statistical signals that classification error monitoring may detect [60]. A faked laboratory value in an AI-generated clinical summary, or an improper drug dose in an AI-drafted prescription suggestion, could influence therapeutic decisions before any doctor notices the mistake. In healthcare institutions without methods to link AI-generated content to subsequent clinical decisions, the audit trail for this inaccuracy may be irrecoverable.
The accountability techniques identified as critical to responsible AI governance by Papagiannidis et al. [1] audit trails, incident reporting, and oversight committee review were conceptualised in respect to AI systems that generate discrete, attributable outputs with clear decision logic. Extending them to language models that generate continuous clinical narratives at a scale and rate that precludes systematic human review necessitates analytical work that neither the organisational governance literature nor the clinical AI safety literature has yet performed with the rigour that the clinical stakes require. The responsible AI governance framework must face the reality that its current conceptual architecture was not created for this type of AI system. Extending it needs more than simply applying current principles to a new technology type; it also necessitates the development of new governance concepts that are appropriate to the special failure modes, accountability issues, and trust dynamics of LLMs in healthcare settings [61,62,63].

4. Implementation Deficit: Organisational Conditions for Governance Failure

4.1. Principled Frameworks Generate Compliance Without Protection

The scoping study by Papagiannidis et al. [1] offers a significant and undervalued contribution by recognising the translation gap the failure to translate high-level responsible AI ideas into concrete organisational practices as the key difficulty in responsible AI governance research. They give far more organisational specificity than principles-only frameworks by introducing structural, procedural, and relational practice categories. Their synthesis reveals what previous responsible AI literature, including Jobin et al. [9] and Fjeld et al. [10], had not addressed: knowing what responsible AI requires does not necessarily translate into knowing how to implement it in organisations operating under real constraints.
The important point their methodology does not fully address, however, is why the implementation gap continues among businesses that officially support responsible AI principles, have established governance mechanisms, and employ specialists whose job descriptions include AI ethics. Answering this question reveals structural aspects of responsible AI governance that more advanced models cannot address.
The implementation gap is not primarily caused by organisational ignorance, resource restrictions, or insufficient implementation guidance, though all of these factors contribute. It reflects a more fundamental issue: responsible AI principles are formulated at a level of abstraction that consistently underdetermines the specific design choices, procurement criteria, validation requirements, monitoring systems, and institutional authority structures required for governance to function substantively rather than symbolically.
The principle of accountability, even when elaborated through the sub dimensions of auditability and responsibility as Papagiannidis et al. [1] propose following European Commission guidelines [34], does not specify what audit infrastructure clinical research organisations must build, who has institutional authority to access it and act on its findings, what information it must capture at what granularity, or what consequences must follow from audit findings that reveal Each of these specifications necessitates interpretive and political decisions that organisations must make on their own, in the absence of standardised methodologies, established clinical AI governance expertise, and in competitive environments where responsible AI governance can boost reputation but not profitability.
The end effect, as demonstrated experimentally in practitioner studies of responsible AI implementation [2], evaluations of explainable AI deployment [4], and syntheses of responsible AI for digital health [65], is governance that is more symbolic than meaningful. Organisations demonstrate their commitment to responsible AI principles through formal governance structures such as ethics committees, responsible AI policies, and explainability tool deployment, but none of these activities provide meaningful assurance of patient safety in the sense of detecting, preventing, or learning from the types of clinical AI failure that governance frameworks are intended to address. The implementation gap is not a problem of organisational capacity in general, but of governance framework design: the frameworks do not specify the institutional investments required to put their principles into action, and in their absence, organisations implement the cheapest activities that meet surface-level compliance requirements [5,8,66].

4.2. Structural Practices and Organisational Architecture Problem

The structural dimension of responsible AI governance examined by Papagiannidis et al. [1] focuses on the assignment of roles, duties, and decision-making authority for AI governance within and across organisational borders. The proposal that organisations form AI oversight committees with cross-functional representation, implement decision-making protocols that ensure consistency and accountability, and create documentation frameworks for rights and responsibilities reflects a truly important institutional logic diverse representation and formal processes create conditions for detecting governance problems that any single professional perspective might miss [67,68,69]. The essential issue in clinical research contexts is not the appropriateness of these structural practices, but the organisational conditions under which they might work as true governance mechanisms rather than just legitimation devices.
An AI oversight committee that lacks institutional authority to halt AI deployments pending safety review, access to the technical information necessary to evaluate AI system behaviour in real deployment, protected time for members to perform meaningful deliberation rather than rubber stamping decisions made elsewhere in the organization, and institutional protection for members who raise safety concerns that conflict with organisational commercial interests is not a govern It is a compliance tool that gives the illusion of monitoring while providing none of the patient protection that good governance necessitates [70,71].The study goal of Papagiannidis et al. [1] properly identifies the multilevel structure of responsible AI governance and the necessity for both vertical and horizontal coordination within companies as crucial understudied areas. The clinical AI governance challenge is more specific: it necessitates institutional authority structures in which AI oversight has sufficient organisational independence and power to constrain commercial and operational pressures that would otherwise deprioritize patient safety governance a structural independence that voluntary corporate AI governance frameworks, by definition, cannot provide.
The inter-organisational dimension of structural practices identified by Papagiannidis et al. [1] is especially important in clinical research contexts, where AI development is frequently distributed across pharmaceutical companies, contract research organisations, academic medical centres, regulatory agencies, and technology vendors in contractual relationships that create complex accountability architectures. Existing responsible AI governance frameworks do not provide a clear answer to the question of who governs AI systems that cross organisational boundaries—whose oversight committee has authority over a monitoring AI developed by a technology vendor, configured by a contract research organisation, and deployed in a sponsor-funded trial conducted at academic medical centres across multiple jurisdictions [64,72,73].

4.3. Relational Practices and Competency Gap

The Papagiannidis et al. [1] paradigm emphasises the relational practices part of responsible AI governance, which includes workforce education in responsible AI literacy, stakeholder engagement in AI development processes, and cross-functional collaboration for AI governance. This reflects an accurate understanding that governance performance is dependent on people competencies distributed throughout organisational levels, not just formal structures and documented processes [2,74,75]. The identification of responsible AI literacy as a critical organisational capability, as well as the recognition that developing it necessitates education in technical AI limitations, algorithmic bias, uncertainty quantification, and human AI interaction dynamics, is both theoretically sound and practical.
The major challenge for clinical research institutions is that the competency profile required for truly meaningful AI governance far outstrips what most healthcare organisations can realistically achieve through training programs and responsible AI literacy initiatives. The management of calibration across demographically different deployment groups necessitates statistical competence, which most clinical governance personnel lack. The detection of distributional shifts in real-time deployment monitoring necessitates both technical data science expertise and extensive clinical domain knowledge, which seldom coincide in the same person. The essential examination of explainability outputs for clinical validity necessitates concurrent expertise in machine learning, clinical pharmacology, and biostatistics, which most clinical AI users lack and no existing training method consistently generates. The relational practices framework correctly identifies responsible AI literacy as central to governance; however, it underestimates the investment required to develop competency profiles adequate to the specific governance tasks that clinical AI safety requires, as well as the structural barriers clinical workforce constraints, training opportunity costs, and career incentive misalignment that prevent most institutions from making that investment [18,33,55].
The stakeholder participation feature of relational activities complicates clinical environments. The vision of inclusive, participatory AI development, which includes patient perspectives, clinical user input, and community representation throughout the AI design, development, and deployment process, reflects genuine ethical commitments to who should have a say in decisions that affect them. The practical challenge, as documented in analyses of algorithmic accountability [75,76], is that meaningful stakeholder participation necessitates both the organisational infrastructure to elicit and process input from diverse stakeholders and the institutional commitment to allowing that input to constrain AI development decisions, including those with commercial or operational consequences. Organisations that develop stakeholder engagement processes without institutional mechanisms to make stakeholder input meaningful are engaging in consultation theatre, which may satisfy procedural governance requirements while providing no substantive protection for the patient populations that those procedures nominally serve [2,5].

5. Regulatory Fragmentation: Political Economy Instead of Coordination Failure

5.1. The Architecture of Judicial Incoherence

The global regulatory landscape for AI in clinical research is marked by jurisdictional fragmentation. The literature on responsible AI governance constantly acknowledges this fragmentation, but it frequently mischaracterises the root causes. The standard account, which is reflected in the governance literature [1,77,78,79], views regulatory fragmentation as a coordination issue: different jurisdictions have developed different frameworks that could be aligned through international harmonisation efforts if technical and institutional barriers to coordination were overcome through better process design. This account misidentifies the root cause of fragmentation, with practical implications for the governance initiatives it encourages.
The European Union's AI Act, the FDA's Software as a Medical Device regulatory pathway, and the NIST AI Risk Management Framework are all technical standards for the same governance purpose. They embody various political philosophies about the appropriate role of the state in governing powerful technologies, theories about the relationship between technological innovation and precautionary regulation, and institutional assumptions about whose interests AI governance is primarily intended to protect [14,15,80]. The EU AI Act's rights-based, mandatory, precautionary approach reflects a European political tradition and specific institutional judgements about the risks of market-driven AI development that are not only technically different from the American approach, but also politically contested at a level that international technical standards bodies ISO/IEC technical committees, WHO harmonisation initiatives, and ICH working groups lack the authority to resolve [5,77,78].
The practical implications of this political economics reality for international clinical AI governance are serious and underappreciated. Clinical trial networks that operate across multiple regulatory jurisdictions at the same time must navigate AI safety monitoring systems that generate regulatory obligations reporting timelines, documentation requirements, safety signal thresholds, and corrective action obligations that may be structurally incompatible across participating jurisdictions [31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82]. An AI system that detects a potential drug-induced hepatotoxicity signal in a global phase III trial creates concurrent obligations under EU, FDA, and national regulatory frameworks, which may specify different thresholds for signal significance, reporting timelines, stakeholder notification requirements, and corrective action expectations. The AI system generates a single safety signal; the regulatory reaction that signal elicits is jurisdictionally diverse in ways that the existing governance infrastructure clinical, regulatory, and organizational is ill-equipped to manage coherently [30,36,37].

5.2. Geographic Concentration and the Global Governance Equity Problem

The systematic geographic concentration of governance framework development and its equity implications for global clinical AI governance are two aspects of regulatory fragmentation that the Papagiannidis et al. [1] framework acknowledges through its identification of cultural and ethical variations as antecedents of responsible AI governance but does not pursue with sufficient analytical force. The frameworks that have had the greatest influence on shaping responsible AI governance globally were developed primarily by institutions in Europe and North America, reflecting institutional capacities, legal traditions, regulatory philosophies, and social values that do not represent the majority of the world's research institutions, patient populations, or clinical environments [9,10,80].
When these frameworks are applied in clinical research contexts across Africa, South and Southeast Asia, and Latin America where AI adoption in research institutions is accelerating, where regulatory infrastructure is differently configured, where AI training datasets are substantially under-represented, and where the consequences of AI governance failure fall on patient populations already subject to significantly worse health outcomes the mismatch between framework assumptions
The technical robustness criteria of key governance frameworks imply institutional capabilities for prospective validation, real-time distribution monitoring, and ongoing model recalibration, which most clinical research institutions in resource-constrained settings lack. The accountability mechanisms they propose are based on legal frameworks, independent regulatory agencies, and professional liability systems whose structures differ greatly between jurisdictions in ways that framework harmonisation cannot easily accommodate [52,62,80]. The result is a global governance landscape in which the patient populations most likely to be harmed by clinical AI failures those from under-represented demographic groups in training data, those receiving care in institutions with inadequate oversight infrastructure, and those from jurisdictions with limited regulatory capacity are also the least well served by the governance frameworks ostensibly intended to protect them. This is not an incidental observation concerning governance geography. It is a structural element of how responsible AI governance has evolved, reflecting power dynamics that influence which institutional conditions are used as the design baseline for global governance frameworks [9,72,80].

6. The Distributional Dimension of Clinical AI Governance Failures

6.1. The Structured Distribution of Clinical AI Governance Failures

The presence of clinical AI governance challenges across patient populations is not coincidental. The combination of training data demographics, deployment institutional characteristics, and governance infrastructure capabilities results in predictable structural patterns. Existing responsible AI frameworks, such as the Papagiannidis et al. [1] synthesis, address distributive issues by emphasising fairness, nondiscrimination, and societal and environmental well-being as governance components. These articulations are important, but they treat distributional consequences as one of several governance concerns, rather than as the moral crux of the clinical AI governance quandary.
Patients from demographic groups under-represented in training data are the most vulnerable to clinical AI governance failures models perform consistently worse for populations about whom they know less [41,42]. They are people receiving care in underfunded institutions with inadequate supervisory infrastructure because governance procedures designed for well-funded university medical centers provide governance theatre rather than real protection in resource-constrained contexts [80]. They originate from jurisdictions with weaker regulatory frameworks, as regulation only provides distributional safeguards to those with regulatory authority [9,77]. And it is those with limited institutional voice who must demand accountability because the political economy of governance responds to organised interests with institutional power, and patient populations harmed by AI systems are among the least organised and institutionally powerful actors in the clinical AI ecosystem [50,52].
This distributional pattern is not a required component of present inadequate AI systems, which will be abolished as technology advances. It is a structural consequence of how AI systems are built using data generated by existing healthcare systems with systematic inequities and deployed in clinical environments with varying levels of human oversight and institutional accountability required to detect and correct AI failures before they cause population-level harm. A governance framework that fails to explicitly center this distributional component and design its criteria accordingly is more than just technically deficient. It implicitly acknowledges a pattern of government failure that disproportionately affects groups already underserved by existing healthcare disparities [44,45,83].

6.2. The Political Economy of Corporate AI Governance

The responsible AI governance landscape includes voluntary corporate governance ethics frameworks, responsible AI policies, algorithmic impact assessments, and AI oversight committees established by companies that develop and employ AI technologies. According to earlier research [66,84], Papagiannidis et al. [1] suggest that competent AI governance can deliver organisational value by improving firm reputation, building consumer trust, and increasing employee morale. The critical question raised by this positive framing is whether governance mechanisms whose primary organisational function is reputation management and regulatory compliance rather than patient protection can be expected to produce substantive safety outcomes when reputational and safety interests diverge.
The systemic conflict of interest in corporate AI governance is not the result of organisational bad faith. It represents the institutional logic of businesses that must provide rewards for stakeholders while simultaneously mitigating AI dangers. The governance mechanisms that organisations design for themselves, using criteria they determine, overseen by committees whose authority they control, and reported in terms they choose, operate within this institutional logic in ways that systematically favour interpretations and implementations of responsible AI governance that minimise constraints on commercial AI deployment [5,8,66]. Patients bear the costs when this institutional logic produces governance failure when AI systems are deployed before adequate safety validation because commercial timelines take precedence, when adverse outcomes in under-represented patient populations are not systematically monitored because monitoring infrastructure requires investment that does not generate direct return, when governance processes are designed to manage reputational risk rather than identify and prevent systematic AI failures have no institutional standing within the corporate governance processes that made these decisions [2,3,83].

7. Critical Synthesis of Unresolved Issues in Responsible AI Governance in Clinical Research

7.1. The Epistemological Infrastructure Problem

The study in the preceding sections reveals a structural issue that extends through all of the specific governance failures addressed. Responsible AI governance frameworks systematically specify what properties AI systems should have and what processes should govern their deployment, while assuming without sufficient examination the epistemological infrastructure data collection, linkage, analytical, and institutional capacity necessary to verify that those specifications have been met and to monitor them continuously in deployment conditions.
The Papagiannidis et al. [1] framework's procedural practices for compliance monitoring and the technical robustness requirements in the European Commission [34] and Singapore Government [35] frameworks assume organisational access to disaggregated outcome data by patient subgroup, real-time model performance monitoring against deployment distribution, audit trails linking AI outputs to clinical decisions and patient outcomes, and the statistical expertise to interpret these. Most clinical research institutions do not usually have this infrastructure, and governance frameworks do not force them to construct it as a prerequisite for AI implementation [18,32,33].
This epistemic infrastructure deficit cannot be remedied by more complex governance frameworks since it is a gap in institutional reality rather than framework content. Governance frameworks can prescribe calibration assessment requirements without establishing the data infrastructure required to assess calibration in real-world clinical settings. They may require fairness monitoring without developing the outcome data gathering tools necessary to discover differences in outcomes based on patient demographics. They can need audit trails without describing the technical architecture, data linkage, or institutional access rights needed to recreate those audit trails in the case of an adverse event. The responsible AI governance field needs to be more explicit that certain governance claims about fairness in deployment, about calibration across patient subgroups, about accountability for adverse outcomes cannot be substantiated without specific institutional infrastructure, and that the absence of that infrastructure means that the governance properties have not been achieved, but that the governance framework cannot determine whether they have been achieved [3,16,17].

7.2. The Self-Referential Limitation of Benchmark-Based Governance

A subtle but significant issue that penetrates the clinical AI governance landscape is the self-referential nature of the evaluation systems on which governance is based. The benchmark datasets used to evaluate clinical AI systems SIDER for drug side effects [19], toxCSM [20], CSM Toxin [21], and clinical trial outcome repositories are all products of the same clinical data infrastructure that contains the biases, under-representation patterns, and documentation inconsistencies that governance frameworks identify as sources of AI safety risk. A benchmark based solely on adverse event data obtained in Western clinical settings cannot determine if toxicity prediction methods apply to non-Western patient populations. A benchmark based on prospectively obtained trial data cannot assess model performance in retrospective EHR applications [41,85].
When governance frameworks use benchmark performance as a criterion for deployment approval which they effectively do by recognising that benchmark validation is adequate evidence of deployment safety, they inherit the benchmarks' limitations. The governance system's ability to detect the most serious clinical AI safety issues is hampered by the same data and methodological constraints that limit the benchmarks, resulting in a governance architecture that is systematically blind to the types of AI failure that have the greatest impact on the most vulnerable patient populations. This self-referential constraint is structural, and it cannot be addressed by enhancing current standards inside their existing conceptual and data infrastructure frameworks [16,18,70].

8. Conclusions

This review contends that the regulation of artificial intelligence in clinical and therapeutic research faces more complex issues than are commonly acknowledged in the current literature. The governance gap is more than just a matter of regulatory delay or institutional adaptation; it represents fundamental limits in how safety, accountability, and justice are currently defined and applied in clinical AI systems. Existing frameworks are mainly based on performance indicators and general ethical ideals, but these methods frequently fail to meet the realities of clinical practice, where mistakes can cause lasting patient suffering. Many healthcare organisations also lack the practical infrastructure required for good governance, such as clear accountability processes, open oversight, and dependable event reporting systems. At the same time, regulatory complexity and unequal institutional capacities make consistent governance challenging, and the repercussions of governance failure frequently fall disproportionately on disadvantaged patient populations. Although responsible AI frameworks have made significant progress in developing ethical and organisational guidance, they are still unsuitable for the unique needs of clinical contexts. Moving forward, clinical AI governance must adopt more robust, context-specific approaches that prioritise practical accountability, patient safety, and equitable protection, ensuring that healthcare AI innovation is matched with governance structures capable of appropriately managing its dangers.

Author Contributions

C.T.Z: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Supervision, Resources, Writing – original draft, Writing – review & editing.

Funding

Not applicable.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

Not applicable.

Conflicts of interest

The authors declare no conflicts of interest.

Declaration of Generative AI and AI-Assisted Technologies in the Writing Process

During the preparation of this manuscript, the authors used ChatGPT solely to assist with language refinement, including improvements in grammar, clarity, and readability. All outputs generated with AI assistance were carefully reviewed, validated, and edited by the authors. The authors take full responsibility for the accuracy, integrity, and originality of the content presented in this work.

References

  1. Papagiannidis, E.; Mikalef, P.; Conboy, K. Responsible artificial intelligence governance: A review and research framework. J. Strateg. Inf. Syst. 2025, 34, 101885. [Google Scholar] [CrossRef]
  2. Rakova, B.; et al. Where responsible AI meets reality: Practitioner perspectives on enablers for shifting organizational practices. Proc. ACM Hum.-Comput. Interact. 2021, 5, 1–23. [Google Scholar] [CrossRef]
  3. Mikalef, P.; Conboy, K.; Lundström, J.E.; Popovič, A. Thinking responsibly about responsible AI and the dark side of AI. Eur. J. Inf. Syst. 2022, 31, 257–268. [Google Scholar] [CrossRef]
  4. Meske, C.; et al. Explainable artificial intelligence: Objectives, stakeholders, and future research opportunities. Inf. Syst. Manag. 2022, 39, 53–63. [Google Scholar] [CrossRef]
  5. Schiff, D.; et al. Explaining the principles to practices gap in AI. IEEE Technol. Soc. Mag. 2021, 40, 81–94. [Google Scholar] [CrossRef]
  6. Van Grembergen, W.; De Haes, S.; Guldentops, E. Structures, processes and relational mechanisms for IT governance. In Strategies for Information Technology Governance; IGI Global, 2004; pp. 1–36. [Google Scholar]
  7. Tallon, P.P.; Ramirez, R.V.; Short, J.E. The information artifact in IT governance. J. Manag. Inf. Syst. 2013, 30, 141–178. [Google Scholar] [CrossRef]
  8. Hagendorff, T. The ethics of AI ethics: An evaluation of guidelines. Minds Mach. 2020, 30, 99–120. [Google Scholar] [CrossRef]
  9. Jobin, A.; Ienca, M.; Vayena, E. The global landscape of AI ethics guidelines. Nat. Mach. Intell. 2019, 1, 389–399. [Google Scholar] [CrossRef]
  10. Fjeld, J.; et al. Principled Artificial Intelligence: Mapping consensus in ethical and rights-based approaches to principles for AI; Berkman Klein Center, 2020. [Google Scholar]
  11. Floridi, L.; Cowls, J. A unified framework of five principles for AI in society. In Ethics, Governance, and Policies in Artificial Intelligence; Springer, 2021; pp. 5–17. [Google Scholar]
  12. Arrieta, A.B.; et al. Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Inf. Fusion. 2020, 58, 82–115. [Google Scholar] [CrossRef]
  13. Adadi, A.; Berrada, M. Peeking inside the black-box: A survey on explainable artificial intelligence (XAI). IEEE Access. 2018, 6, 52138–52160. [Google Scholar] [CrossRef]
  14. Smuha, N.A. From a race to AI to a race to AI regulation. Law. Innov. Technol. 2021, 13, 1–29. [Google Scholar] [CrossRef]
  15. Taeihagh, A. Governance of artificial intelligence. Policy Soc. 2021, 40, 137–157. [Google Scholar] [CrossRef]
  16. Chekroud, A.M.; et al. Illusory generalizability of clinical prediction models. Science 2024, 383, 164–167. [Google Scholar] [CrossRef]
  17. Desai, R.J.; et al. Individualized prediction of the effect of machine learning treatment. NEJM Evid. 2024, 3, EVIDoa2300041. [Google Scholar] [CrossRef] [PubMed]
  18. Vollmer, S.; et al. Machine learning and artificial intelligence research for patient benefit: 20 critical questions. BMJ 2020, 368, l6927. [Google Scholar] [CrossRef] [PubMed]
  19. Kuhn, M.; Letunic, I.; Jensen, L.J.; Bork, P. The SIDER database of drugs and side effects. Nucleic Acids Res. 2016, 44, D1075–D1079. [Google Scholar] [CrossRef]
  20. de Sá, A.G.C.; Long, Y.; Portelli, S.; Pires, D.E.V.; Ascher, D.B. toxCSM: complete prediction of small molecule toxicity profiles. Brief. Bioinform. 2022, 23, bbac337. [Google Scholar] [CrossRef]
  21. Morozov, V.; Rodrigues, C.H.M.; Ascher, D.B. CSM-Toxin: a web server for protein toxicity prediction. Pharmaceutics 2023, 15, 431. [Google Scholar] [CrossRef]
  22. Kim, E.; Nam, H. Prediction models for drug-induced hepatotoxicity using weighted molecular fingerprints. BMC Bioinform. 2017, 18, 227. [Google Scholar] [CrossRef] [PubMed]
  23. Lysenko, A.; Sharma, A.; Boroevich, K.A.; Tsunoda, T. An integrative machine learning approach for toxicity-related drug safety prediction. Life Sci. Alliance 2018, 1, e201800098. [Google Scholar] [CrossRef]
  24. He, S.; et al. An in silico model to predict drug-induced hepatotoxicity. Int. J. Mol. Sci. 2019, 20, 1897. [Google Scholar] [CrossRef]
  25. Nguyen-Vo, T.H.; et al. Prediction of drug-induced liver injury using a convolutional neural network. ACS Omega 2020, 5, 25432–25439. [Google Scholar] [CrossRef]
  26. Lim, S.; et al. Supervised exploration of chemical graphs improves the prediction of drug-induced liver injury. iScience 2023, 26, 105677. [Google Scholar] [CrossRef]
  27. Gong, Y.; et al. In silico prediction of potential drug-induced nephrotoxicity using machine learning methods. J. Appl. Toxicol. 2022, 42, 1639–1650. [Google Scholar] [CrossRef]
  28. Zhao, H.; et al. A novel graphical attention model for predicting drug side effect frequencies. Brief. Bioinform. 2021, 22, bbab239. [Google Scholar] [CrossRef]
  29. Chong, L.H.; et al. Integration of a microfluidic multicellular co-culture network with machine learning analysis to predict adverse skin reactions to drugs. Lab. A Chip 2022, 22, 1890–1904. [Google Scholar] [CrossRef] [PubMed]
  30. Agrafiotis, D.K.; et al. Risk-based clinical trial monitoring: an integrative approach. Clin. Ther. 2018, 40, 1204–1212. [Google Scholar] [CrossRef] [PubMed]
  31. Barnes, B.; et al. Risk-based monitoring in clinical trials: past, present and future. Ther. Innov. Regul. Sci. 2021, 55, 899. [Google Scholar] [CrossRef] [PubMed]
  32. Koneswarakantha, B.; Ménard, T.; Rolo, D.; Barmaz, Y.; Bowling, R. Harnessing the power of quality assurance data. Ther. Innov. Regul. Sci. 2020, 54, 1227–1235. [Google Scholar] [CrossRef]
  33. Weissler, E.H.; et al. The role of machine learning in clinical research: transforming the future of evidence production. Trials 2021, 22, 537. [Google Scholar] [CrossRef]
  34. European Commission. Ethics guidelines for trustworthy AI. High.-Lev. Expert Group Artif. Intell. 2019. [Google Scholar]
  35. Singapore Government. Model AI Governance Framework, 2nd ed.; Personal Data Protection Commission, 2020. [Google Scholar]
  36. Fneish, F.; Schaarschmidt, F.; Fortwengel, G. Improving risk assessment in clinical trials. Curr. Ther. Res. 2021, 95, 100643. [Google Scholar] [CrossRef] [PubMed]
  37. Yao, B.; Zhu, L.; Jiang, Q.; Xia, H.A. Safety monitoring in clinical trials. Pharmaceutics 2013, 5, 94–106. [Google Scholar] [CrossRef]
  38. Wu, W.; Huang, T.; Gong, K. Ethical principles and governance technology development of AI in China. Engineering 2020, 6, 302–309. [Google Scholar] [CrossRef]
  39. Peng, Y.; Zhang, Z.; Jiang, Q.; Guan, J.; Zhou, S. TOP: A deep mixture representation learning method to improve molecular toxicity prediction. Methods 2020, 179, 55–64. [Google Scholar] [CrossRef]
  40. de Lomana, M.; et al. ChemBioSim: Improving the conformal prediction of in vivo toxicity. J. Chem. Inf. Model. 2021, 61, 3255–3272. [Google Scholar] [CrossRef]
  41. Badwan, B.A.; et al. Machine learning approaches to predict drug efficacy and toxicity in oncology. Cell. Rep. Methods 2023, 3, 100413. [Google Scholar] [CrossRef] [PubMed]
  42. Ménard, T.; Barmaz, Y.; Koneswarakantha, B.; Bowling, R.; Popko, L. Enabling data-driven clinical quality assurance. Drug. Saf. 2019, 42, 1045–1053. [Google Scholar] [CrossRef]
  43. Yazdani, A.; et al. An evaluation criterion for predicting adverse drug events from clinical trial results. Sci. Data 2025, 12, 424. [Google Scholar] [CrossRef]
  44. Feuerriegel, S.; Dolata, M.; Schwabe, G. Fair AI: Challenges and opportunities. Bus. Inf. Syst. Eng. 2020, 62, 379–384. [Google Scholar] [CrossRef]
  45. Varona, D.; Suárez, J.L. Discrimination, bias, fairness, and trustworthy AI. Appl. Sci. 2022, 12, 5826. [Google Scholar] [CrossRef]
  46. Jakesch, M.; et al. How different groups prioritize ethical values for responsible AI. Proceedings of FAccT, 2022; pp. 310–323. [Google Scholar]
  47. Askin, S.; Burkhalter, D.; Calado, G.; El Dakrouni, S. Artificial intelligence applied to clinical trials: opportunities and challenges. Health Technol. 2023, 13, 203–213. [Google Scholar] [CrossRef]
  48. Harrer, S.; Shah, P.; Antony, B.; Hu, J. Artificial intelligence for clinical trial design. Trends Pharmacol. Sci. 2019, 40, 577–591. [Google Scholar] [CrossRef]
  49. Nebeker, J.R.; Barach, P.; Samore, M.H. Clarification of adverse drug events: a clinician's guide. Ann. Intern. Med. 2004, 140, 795–801. [Google Scholar] [CrossRef]
  50. Yeung, K.; Howes, A.; Pogrebna, G. AI governance by human rights-centered design, deliberation, and oversight. In The Oxford Handbook of Ethics of AI; Oxford University Press, 2020. [Google Scholar]
  51. Winfield, A.F.T.; Jirotka, M. Ethical governance is essential to building trust in robotics and AI systems. Philos. Trans. R. Soc. A 2018, 376, 20180085. [Google Scholar] [CrossRef] [PubMed]
  52. Theodorou, A.; Dignum, V. Towards ethical and socio-legal governance in AI. Nat. Mach. Intell. 2020, 2, 10–12. [Google Scholar] [CrossRef]
  53. Mangione, W.; Falls, Z.; Samudrala, R. Efficient holistic characterization of small molecule effects using heterogeneous biological networks. Front. Pharmacol. 2023, 14, 1113007. [Google Scholar] [CrossRef]
  54. Moingeon, P.; Kuenemann, M.; Guedj, M. Design and development of AI-enhanced drugs. Drug. Discov. Today 2022, 27, 215–222. [Google Scholar] [CrossRef] [PubMed]
  55. Alowais, S.A.; et al. Revolutionizing healthcare: the role of artificial intelligence in clinical practice. BMC Med. Educ. 2023, 23, 689. [Google Scholar] [CrossRef]
  56. Tolmeijer, S.; et al. Capable but amoral? Comparing AI and human expert collaboration in ethical decision making. Proceedings of CHI, 2022; pp. 1–17. [Google Scholar]
  57. Brendel, A.B.; et al. Ethical management of artificial intelligence. Sustainability 2021, 13, 1974. [Google Scholar] [CrossRef]
  58. Shneiderman, B. Human-centered artificial intelligence: Three fresh ideas. AIS Trans. Hum.-Comput. Interact. 2020, 12, 109–124. [Google Scholar] [CrossRef]
  59. Yerlikaya, S.; Erzurumlu, Y.O. Artificial intelligence in public sector: A framework to address opportunities and challenges. In The Fourth Industrial Revolution; Springer, 2021; pp. 201–216. [Google Scholar]
  60. Ghim, J.L.; Ahn, S. Transforming clinical trials: the emerging roles of major language models. Transl. Clin. Pharmacol. 2023, 31, 131–138. [Google Scholar] [CrossRef]
  61. Dignum, V. Responsible Artificial Intelligence: How to Develop and Use AI in a Responsible Way; Springer, 2019. [Google Scholar]
  62. Ghallab, M. Responsible AI: Requirements and challenges. AI Perspect. 2019, 1, 1–7. [Google Scholar] [CrossRef]
  63. Helbing, D. Towards Digital Enlightenment; Springer, 2019. [Google Scholar]
  64. Mäntymäki, M.; et al. Defining organizational AI governance. AI Ethics 2022, 2, 603–609. [Google Scholar] [CrossRef]
  65. Trocin, C.; Mikalef, P.; Papamitsiou, Z.; Conboy, K. Responsible AI for digital health: A synthesis and a research agenda; Information Systems Frontiers, 2021. [Google Scholar]
  66. de Laat, P.B. Companies committed to responsible AI: From principles towards implementation and regulation? Philos. Technol. 2021, 34, 1135–1193. [Google Scholar] [CrossRef] [PubMed]
  67. Janssen, M.; et al. Data governance: Organizing data for trustworthy Artificial Intelligence. Gov. Inf. Q. 2020, 37, 101493. [Google Scholar] [CrossRef]
  68. Radu, R. Steering the governance of artificial intelligence: National strategies in perspective. Policy Soc. 2021, 40, 178–193. [Google Scholar] [CrossRef]
  69. Too, E.G.; Weaver, P. The management of project management: A conceptual framework for project governance. Int. J. Proj. Manag. 2014, 32, 1382–1394. [Google Scholar] [CrossRef]
  70. Matthews, J. Patterns and antipatterns, principles and pitfalls: Accountability and transparency in artificial intelligence. AI Mag. 2020, 41, 82–89. [Google Scholar] [CrossRef]
  71. Raji, I.D.; et al. Closing the AI accountability gap: defining an end-to-end framework for internal algorithmic auditing. Proceedings of FAccT, 2020; pp. 33–44. [Google Scholar]
  72. Wirtz, B.W.; Weyerer, J.C.; Sturm, B.J. The dark sides of artificial intelligence: An integrated AI governance framework for public administration. Int. J. Public Adm. 2020, 43, 818–829. [Google Scholar] [CrossRef]
  73. Butcher, J.; Beridze, I. What is the state of artificial intelligence governance globally? RUSI J. 2019, 164, 88–96. [Google Scholar] [CrossRef]
  74. Aldoseri, A.; Al-Khalifa, K.; Hamouda, A. A road map for integrating automation with process optimization for AI-powered digital transformation. Preprints 2023. [Google Scholar]
  75. Edwards, L.; Veale, M. Enslaving the algorithm: From a right to an explanation to a right to better decisions? IEEE Secur. Priv. 2018, 16, 46–54. [Google Scholar] [CrossRef]
  76. Felzmann, H.; et al. Towards transparency by design for artificial intelligence. Sci. Eng. Ethics 2020, 26, 3333–3361. [Google Scholar] [CrossRef] [PubMed]
  77. de Almeida, P.G.R.; et al. Artificial intelligence regulation: a framework for governance. Ethics Inf. Technol. 2021, 23, 505–525. [Google Scholar] [CrossRef]
  78. Butcher, J.; Beridze, I. What is the state of artificial intelligence governance globally? RUSI J. 2019, 164, 88–96. [Google Scholar] [CrossRef]
  79. Papagiannidis, et al. (as above).
  80. Nzobonimpa, S.; Savard, J.F. Ready but irresponsible? Analysis of the Government AI Readiness Index. Policy Internet 2023, 15, 397–414. [Google Scholar] [CrossRef]
  81. Mentz, R.J.; et al. Good clinical practices and pragmatic clinical trials. Circulation 2016, 133, 872–880. [Google Scholar] [CrossRef] [PubMed]
  82. Simović, M.; Nikolić, N. Challenges of risk-based clinical trial monitoring. Clin. Res. Regul. Aff. 2015, 32, 83–87. [Google Scholar] [CrossRef]
  83. Akter, S.; et al. Algorithmic bias in data-driven innovation in the age of AI. Int. J. Inf. Manag. 2021, 60, 102387. [Google Scholar] [CrossRef]
  84. Martin, K.; Waldman, A. Are algorithmic decisions legitimate? J. Bus. Ethics 2023, 183, 653–670. [Google Scholar] [CrossRef]
  85. Reymond, J.L. The chemical space project. Acc. Chem. Res. 2015, 48, 722–730. [Google Scholar] [CrossRef] [PubMed]
Figure 1. The governance gap in clinical AI: from principles to practice.
Figure 1. The governance gap in clinical AI: from principles to practice.
Preprints 211033 g001
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings