Preprint
Review

This version is not peer-reviewed.

Artificial Intelligence in Radiology: Why Retrospective Performance Does Not Ensure Clinical Utility

Submitted:

23 July 2026

Posted:

24 July 2026

You are already at the latest version

Abstract
Purpose: Artificial intelligence is increasingly developed and deployed to support radiological image interpretation, triage, segmentation and workflow prioritisation. However, high performance reported in retrospective studies does not necessarily translate into safe and useful deployment in routine practice. This narrative review examines why this translational gap persists and what radiology departments, researchers and developers should address before clinical implementation. Methods: We synthesised methodological, empirical and regulatory literature on the development, validation, deployment and governance of radiology AI. Radiology was used as the primary setting, with adjacent imaging fields discussed only where they illustrate broader mechanisms relevant to medical image analysis. Results: Recurrent barriers include limited data quality and representativeness, scanner- and protocol-related domain shift, shortcut learning, hidden stratification, insufficient external and prospective validation, poor calibration, over-reliance on global metrics and weak workflow integration. Model performance can be shaped by scanner protocols, reconstruction methods, local reporting practices, patient selection and institutional workflows beyond the underlying pathology. In Europe, GDPR-based data governance, the Medical Device Regulation, the AI Act and the European Health Data Space further shape clinical deployment. Conclusions: Emerging approaches, including foundation models, generative AI, multimodal systems, federated learning and local adaptation, may address selected technical constraints, but they introduce new uncertainties and do not replace independent validation in the intended clinical setting. Radiology AI should be evaluated not only as an algorithmic model, but as a clinical decision-support component embedded in radiology practice, requiring external and preferably prospective validation, subgroup analysis, calibration assessment, human oversight, workflow integration and post-deployment monitoring.
Keywords: 
;  ;  ;  ;  ;  ;  

Introduction

In recent years, artificial intelligence methods, particularly deep learning techniques, have become one of the most intensively researched areas in radiology and medical imaging [1,2]. Common applications include image classification, lesion detection, segmentation, triage, workflow prioritisation and quantitative image analysis [1,2]. Selected examples from adjacent imaging fields are used only as supporting analogies [3]. The importance of radiology as one of the main areas for AI implementation is confirmed by an analysis of AI/ML devices authorised by the FDA. Sivakumar et al. demonstrated that, of the 950 AI/ML devices authorised by the FDA up to June 2024, 723 – or 76 per cent – were related to radiology [4]. At the same time, the scale of authorisation for such tools should not be confused with full confirmation of their usefulness in actual clinical practice. Of the 717 radiological devices for which documentation was available, only 33, or 5%, had been prospectively tested, 56, or 8%, incorporated a human-in-the-loop element, and 208, or 29%, had been clinically tested [4]. The FDA maintains a public list of AI-enabled medical devices authorised for marketing in the US, but notes that this is not a comprehensive register of all such solutions [5]. Although some models achieved performance comparable to that of healthcare professionals in selected retrospective diagnostic tasks, these findings should be treated as hypothesis-generating evidence of translational potential rather than proof of clinical readiness [6]. Such results do not provide a sufficient basis for routine clinical implementation, particularly when studies are retrospective, inadequately reported or lack external validation [6,7,8]. Furthermore, the current publication ecosystem often disproportionately rewards positive outcomes, which can lead to inflated expectations regarding real-world AI utility. At the same time, research on the generalisability of models, including the study by Zech et al. and the review by Yu et al., indicates that high metric values obtained under internal validation conditions may lead to an overestimation of the model’s effectiveness when applied in a different centre, patient population or with a different data acquisition protocol [8,9].
One of the reasons for these differences is the way in which imaging data is generated. The outcome of an imaging procedure depends not only on the patient’s condition, but also on how the procedure is carried out, the equipment and software used, and local procedures for reporting and classifying cases [10,11,12]. A model trained on such data can utilise not only the visual characteristics of the disease, but also technical, organisational or demographic factors that co-occur with the diagnosis [8,9,10,11,12]. Such dependencies may limit the model’s ability to generalise beyond the environment from which the training data was obtained [8,9,11]. This review discusses why AI systems for radiological image analysis may fail to retain clinical utility when transferred from retrospective development datasets to real radiology practice. The focus is on mechanisms directly relevant to radiology departments: data quality, scanner- and protocol-related domain shift, shortcut learning, insufficient external validation, calibration, workflow integration, regulation and post-deployment monitoring [8,9,11]. Selected examples from digital pathology are included where they illustrate broader imaging-AI mechanisms such as sample preparation, domain shift, labelling variability and generalisability [3,13]. These challenges are particularly relevant for European radiology departments, where AI deployment must be considered alongside GDPR-based data governance, the MDR framework for medical device software, the AI Act and emerging data-sharing infrastructure such as the European Health Data Space [14,15,16,17]. A schematic overview of the successive stages at which radiology AI systems may lose clinical utility is presented in Figure 1.

Methodology and Selection of Literature

This article is a narrative review. Literature was selected through targeted searches of PubMed and Google Scholar, supplemented by reference chaining from systematic reviews, meta-analyses, methodological papers, reporting guidelines, empirical validation studies and regulatory documents. Searches combined terms related to artificial intelligence, radiology, medical imaging, external validation, domain shift, shortcut learning, hidden stratification, dataset shift, calibration, workflow integration, medical devices, foundation models, generative AI, multimodal AI and federated learning. Preference was given to publications from 2018 onward, with older sources included when they provided important conceptual or methodological context.
The aim was not to perform a systematic review or meta-analysis, but to synthesise mechanisms that can explain why retrospective model performance may fail to translate into routine radiology practice. Because of the narrative design, no formal risk-of-bias assessment was conducted. The final selection, interpretation and verification of the cited sources were performed by the authors.

Limitations Regarding Data Quality and Representativeness

The quality of data in AI models used in medical imaging goes beyond the technical parameters of the image. According to the framework proposed by Schwabe et al. as part of METRIC, the assessment of medical data quality should relate to the specific application of the model and encompass multiple dimensions of the training dataset’s suitability, rather than merely its size or basic technical parameters [10]. In practice, this means that it is necessary to assess the consistency of the data acquisition and processing workflow, the reliability of labels, the completeness of metadata, and the ability to relate the image to its clinical context. The reliability of labels is particularly important, as a reference result in medicine may come from various sources, such as a radiological report, a diagnosis code, a clinical decision, an expert assessment or a histopathological test result. Ambiguous or inconsistent labels increase the risk that the model will reproduce local documentation and diagnostic practices, rather than stable features associated with the actual clinical condition. Technical and organisational inconsistencies can have a similar effect. In radiology, these include differences in scanner vendor, acquisition protocol, reconstruction kernel, slice thickness, contrast phase, DICOM metadata completeness and local reporting practices [10,12,18]. In digital pathology, analogous issues may arise from specimen preparation, staining and digitisation, but these are used here mainly as supporting examples of broader imaging-AI data-quality problems [3,13]. For this reason, dataset quality should be assessed in relation to the entire process of data generation, including patient eligibility, examination performance, image acquisition, reporting, labelling and subsequent use for model training and validation [10]. The evaluation of training data for medical AI should therefore take into account not only its volume, but also its completeness, consistency, reliability, representativeness and suitability for a specific clinical application [10].
Data quality is directly related to its representativeness, understood as the relevance of the training dataset to the population and the conditions in which the model is to be applied. It is not only the number of cases that is important, but also the demographic, technical and clinical diversity of the dataset. Training data should reflect the real-world conditions of the model’s intended application as accurately as possible. A model trained on data from a single centre or a single healthcare system may accurately reproduce the local data structure, whilst performing worse in an environment with different patient characteristics, medical equipment or diagnostic procedures [8,9,19]. Representativeness is also linked to data balance, i.e. the appropriate representation of both common cases and rarer but clinically significant diagnoses. A highly unbalanced dataset can lead to a bias towards the dominant class and an overestimation of overall metrics, whilst resulting in poor detection of rare cases [20].
The completeness of the data is also important. The absence of a reference result, metadata or clinical information may limit the ability to train, evaluate and interpret the model correctly [10]. These gaps are often not random, but result from the organisation of diagnostic procedures, the severity of the patient’s condition, the availability of additional tests, or clinical decisions, which may reduce the representativeness of the training set in relation to the target population [10]. Limited data quality, representativeness and completeness may therefore lead to systematic errors, i.e. persistent biases resulting in the model performing differently across patient groups, diagnoses and centres [19]. Consequently, the model’s average performance may not reflect its performance in populations that are under-represented in the training data. Therefore, the development of medical AI systems requires not only large datasets, but above all multi-centre data that is well-documented and assessed for potential sources of systematic error. This issue also relates to the geographical origin of the data: the over-representation of datasets from the United States and China may limit the generalisability of models to other populations and healthcare systems, including those in Europe [19].
However, it is difficult to obtain data that meet these criteria, as in medicine the issue of data quality and representativeness is directly linked to limitations in access to, sharing of, and standardisation of clinical resources.

Regulatory and Organisational Barriers to Radiology AI Deployment

Regulatory and organisational barriers affect radiology AI at two connected levels: access to representative multicentre data and transition from a research algorithm to deployable medical software. Existing medical data are often underused because they remain in institutional silos and privacy concerns restrict access to them [21]. Under the GDPR, health data belong to special categories of personal data and require an appropriate legal basis and safeguards such as pseudonymisation, encryption and other appropriate technical and organisational measures [16]. These requirements are essential for patient privacy, but they may increase the time and complexity of multicentre data curation, external validation and clinical deployment [16,21]. The European Health Data Space aims to establish a common framework for the use, exchange and secondary reuse of electronic health data across the European Union [17].
Data access alone is not sufficient. Reliable radiology AI also requires standardised, traceable, correctly annotated and de-identified imaging data, together with curation, harmonisation and quality control [18]. These processes are costly and often depend on local reporting conventions, scanner protocols and institutional workflows. As a result, models may still be developed on datasets that are smaller, more local and less representative than clinical practice requires [18,21].
A separate barrier is the transition from a research algorithm to software intended for diagnostic use. Depending on intended purpose and risk classification, radiology AI may qualify as medical device software and may also be subject to high-risk AI requirements [14,15,22]. These frameworks shift evaluation from model performance alone to risk management, technical documentation, clinical evaluation, human oversight and post-market monitoring [14,15,22]. Therefore, high retrospective performance cannot be assumed to indicate readiness for clinical deployment without separate evidence of safety, performance, analytical validity, clinical validity and clinical usefulness in the intended setting [14,15,22].
Technical and organisational approaches can reduce some barriers but do not remove the need for data governance and local validation. Controlled access, common data standards and federated learning can support multicentre collaboration without centralising all raw patient data [17,21,23]. However, federated learning does not by itself solve heterogeneity in acquisition protocols, labels, metadata or local clinical practice [18,23]. Even after data access and software requirements have been addressed, radiology AI still has to be evaluated in the technical, population and workflow conditions of the site where it will be used.

Domain Shift

Domain shift refers to a situation in which the data used to train a model differs statistically from the data on which the model is subsequently applied. Finlayson et al. describe this problem as one of the causes of AI systems malfunctioning following clinical deployment, when the relationship between the development data and the target data changes [11]. In medical imaging, domain shift arises from the fact that an image is dependent on the local environment in which data is acquired, processed and interpreted. In their review on domain adaptation in medical image analysis, Guan and Liu emphasised that models often encounter the problem of distribution differences between source and target data, arising, amongst other things, from different scanners, acquisition parameters and patient cohorts [12]. A change to any of these factors may result in a systematic discrepancy between the training data and the target data, even if both sets relate to the same medical condition [12]. In radiology, domain shift is caused primarily by variations in equipment, scanner vendor, acquisition protocols, reconstruction techniques and patient cohorts [12]. Similar mechanisms have been described in adjacent imaging fields such as digital pathology, where specimen preparation, staining and digitisation can induce measurable domain shift and reduce deep learning performance [13]. These differences may be of little significance from an expert’s point of view, yet they can significantly alter the distribution of the input data analysed by the model [12,13].
A consequence of domain shift is the model’s limited ability to generalise predictions beyond the training environment. This problem is reflected in external validation studies. Zech et al. demonstrated variable generalisation of a model for detecting pneumonia on chest X-rays across different centres, whilst Yu et al., in a systematic review, described a frequent decline in the performance of deep learning algorithms in radiological diagnosis when applied to external data [8,9]. This is because the model may learn characteristics specific to the training environment, rather than solely the stable features of the disease. In the study by Zech et al., it was noted that the models may have utilised signals associated with the hospital centre or system, which limited their transferability between clinical settings [9]. Domain shift is not limited to the appearance of the image. It may also encompass population structure, patient eligibility criteria, label definitions and the relationship between the image and the clinical outcome. Subbaswamy and Saria emphasise that the issue of the transferability of health models should be analysed not only as a change in the distribution of input data, but also as a change in the mechanisms linking data, clinical decisions and health outcomes [24]. Therefore, a model that performs well on one dataset may need to be re-evaluated, calibrated or adapted before being used in a different clinical setting. In this context, Guan and Liu describe domain adaptation as a family of methods designed to mitigate distribution differences between related but distinct domains of medical data [12]. A drop in performance following a change in the clinical environment may be due not only to the difference between the datasets themselves, but also to the fact that the model has learnt to recognise spurious signals present in the training dataset [9].
This problem may also be time-related. Clinical data change as a result of equipment replacement, software updates, modifications to diagnostic protocols, changes in the patient population, new treatment regimens or the reorganisation of the department’s work. Finlayson et al. point out that clinicians using AI systems should consider whether local clinical practice differs from the conditions under which the system was developed, and whether it has changed over time [11]. Such data drift may gradually reduce the model’s effectiveness after deployment, even if the model was successfully validated at the time of deployment [11,24]. For this reason, the evaluation of an AI model should not be limited to a single internal test. It is essential to test the model on external data – ideally from multiple centres – and, following deployment, to monitor the model’s performance over time and periodically assess whether the input data still corresponds to the conditions under which the model was trained and validated [8,11,24].

Shortcut Learning and Hidden Stratification

Shortcut learning occurs when a model uses correlations that are predictive in the development dataset but do not reflect stable or clinically meaningful disease features. Geirhos et al. define shortcuts as decision rules that enable good performance under standard test conditions but fail to transfer to more challenging or altered conditions of application [25]. In radiology, this problem is clinically relevant because technical, organisational or population-related signals may co-vary with the target label, allowing high internal performance without genuine disease recognition [9,25,26].
In radiology, shortcuts may arise from site-specific acquisition patterns, image framing, dataset composition or differences in disease prevalence between datasets [9,26]. In a study by DeGrave et al., models classifying COVID-19 on chest X-rays achieved seemingly high accuracy, but relied on confounding factors related to data origin rather than only pathology-related signal [26]. Similar non-disease signals may also occur in adjacent imaging fields such as digital pathology, where specimen preparation, staining and digitisation vary between laboratories. In this review, these examples are used only to illustrate broader imaging-AI mechanisms [13].
Shortcut learning is closely related to domain shift. A shortcut may appear valid during internal validation if the same spurious association is present in both the training and test data, but may fail after transfer to another centre, scanner protocol or patient population [9,25,26]. In such cases, high performance does not prove that the model is addressing the intended clinical problem. It may only show that the model has reproduced the local structure of the development dataset [9,26].
Hidden stratification is a related problem in which a model achieves good overall performance but performs substantially worse in clinically important subgroups or disease subtypes. Oakden-Rayner et al. describe hidden stratification as a clinically meaningful failure mode in medical imaging, especially when relevant subgroups are not explicitly represented in the labels used for model training and evaluation [27]. This may involve rare, atypical, under-represented or label-ambiguous cases. As a result, global metrics may suggest that the model performs well even though it fails in a subgroup where accurate detection is particularly important [27].
Detection of shortcuts and hidden stratification can be supported by subgroup analysis, error analysis, external validation and cautious use of interpretability methods. DeGrave et al. used explainability methods to show that chest X-ray models could rely on confounding factors rather than pathology-related signal [26]. However, saliency maps should not be treated as sufficient evidence of model validity, as Adebayo et al. showed that some saliency methods may be insensitive to model parameters or data [28]. Therefore, interpretability methods should be treated as supportive tools, while the primary assessment of robustness should rely on external validation, subgroup analysis and testing under varied data conditions [8,27,28].

Limitations on the Validation and Interpretation of Model Results

Standard internal validation does not always show whether a model will remain reliable in the target clinical setting. In the systematic review by Yu et al., most analysed deep learning algorithms for radiological diagnosis performed worse during external validation than during internal validation [8]. Even when a test set is formally separated from the training set, it may share the same acquisition protocols, reporting practices, patient-selection mechanisms and site-specific factors. Internal validation can therefore overestimate generalisability if the deployment setting differs from the development environment [8,24,29].
Data partitioning errors may be an additional source of overestimation of effectiveness. In medical imaging, partitioning at the patient, study or centre level is of particular importance, rather than merely the random partitioning of individual images. Yagis et al. demonstrated that, in 3D MRI classification using 2D models, the division of data at the level of individual slices can lead to information leakage between the training and test sets and a significant overestimation of the reported model performance [30]. If images of the same patient, similar slices from the same scan, or data from the same series end up in both the training and test sets, the model may utilise information that is not available in real-world applications, leading to an overestimation of validation results [30].
For this reason, external validation is crucial because it uses data from outside the environment in which the model was trained. Riley et al. describe external validation as a stage that allows the model’s performance to be assessed on new data, taking into account not only predictive metrics but also data quality and the model’s potential clinical utility [29]. Datasets covering different centres, equipment, protocols, patient populations and time periods help determine whether performance remains stable under changed clinical conditions and whether local calibration, fine-tuning or revalidation is needed before implementation [8,29]. Prospective validation provides stronger evidence because it captures routine workflow, incomplete data, real-world eligibility criteria and user interaction. For AI-based intervention trials, CONSORT-AI and SPIRIT-AI further require transparent reporting of system integration, input and output data, human–AI interaction and error cases [31,32].
Another limitation of many studies is an over-reliance on global metrics. In classification, metrics such as accuracy or AUC are often reported, whilst in segmentation, the Dice coefficient or Hausdorff distance are used. These metrics are useful, but they do not describe the full clinical value of the model. Riley et al. emphasise that the evaluation of predictive models should encompass not only predictive performance, but also calibration and clinical utility [29]. A high AUC indicates the model’s ability to distinguish between cases, but does not determine whether the chosen decision threshold is clinically justified, what the false-positive rate will be in a population with a low prevalence of the disease, or whether the effectiveness remains similar across different patient subgroups [27,29,33]. Similarly, a high average Dice coefficient may hide significant segmentation errors in small, complex or clinically important structures. For this reason, the evaluation of the model should include not only the average result, but also an analysis of subgroups, estimation uncertainty, the distribution of errors and the implications of the chosen decision threshold [27,29,33].
Calibration is essential when model outputs are interpreted as probabilities or used to trigger clinical action. A model may rank cases correctly while producing poorly calibrated risk estimates, which can lead to inappropriate thresholds, false alarms or missed cases in routine practice [29,33]. Evaluation should therefore include not only discrimination, but also calibration, uncertainty, threshold selection and the clinical consequences of false-positive and false-negative results [29,33].
A reliable assessment of medical AI models requires a combination of external validation, subgroup analysis, calibration assessment and, where justified by the intended use of the system, prospective validation in real-world practice [8,27,29]. It is also important to provide a clear description of how the data were split, the number of cases, any missing data, the inclusion and exclusion criteria, the characteristics of the population, and the metrics used. The absence of such information makes it difficult to assess the risk of systematic error and to compare results across studies [33,34]. For this reason, reporting standards should be selected according to the type of study: CLAIM applies to AI studies in medical imaging, TRIPOD+AI to studies developing or evaluating predictive models, CONSORT-AI to reports on clinical trials of interventions with an AI component, and SPIRIT-AI to the protocols for such trials. Together, they emphasise the need for a detailed description of the data, validation, metrics, limitations and clinical context of AI-based systems [31,32,33,34].
The key factors limiting the clinical translation of radiology AI systems are summarised in Table 1.

Clinical Objectives and Workflow Integration in Radiology Practice

Even a rigorously validated model may have limited value if its intended purpose does not correspond to an actual radiology decision. Kelly et al. emphasise that delivering clinical impact requires moving beyond technical effectiveness to the assessment of real-world use, workflow and patient benefit [35]. Although AI studies often frame tasks as classification, detection or segmentation, radiology decisions usually involve urgency, further diagnostics, treatment planning or comparison with prior imaging. A technically accurate model may therefore still have limited value unless its output is explicitly linked to a specific clinical decision or action [35].
Different intended uses, such as triage, prioritisation, measurement, second reading or diagnostic support, require different validation criteria, thresholds, user interfaces and safety controls [31,32,35]. A triage algorithm does not necessarily need to provide a final diagnosis. Overly broad objectives may increase error risk and reduce usefulness when a simpler, controlled function would be clinically sufficient [35].
Radiological predictions rarely depend on the image alone. Interpretation depends on indication, prior imaging, medical history, treatment and disease course. An image-only model may be technically correct but clinically incomplete when similar findings require different decisions depending on context. Its value therefore depends on whether the output is embedded in a real decision-making scenario [35].
Interdisciplinary collaboration is therefore needed from the stage of defining the model’s objective. Kelly et al. point out that translation into clinical practice requires attention not only to algorithm design, but also to validation, regulation, implementation and work organisation [35]. Clinicians help identify the clinical need, while the technical team translates it into data preparation, training, validation and integration with existing infrastructure.
Another key aspect of usability is workflow integration. Tejani et al. emphasise that scalable AI implementation in radiology requires standardised interoperability and integration with the existing working environment, because non-standard integrations may increase operational burden and implementation risk [36]. Even an effective algorithm may have limited value if it requires a separate login, manual data export or switching between systems. Outputs should be available within PACS, RIS, HIS or EHR systems and should not substantially prolong reporting [36]. A system that increases alert burden, reporting time or cognitive workload may fail to provide clinical value despite high standalone accuracy. Evaluation should therefore include usability, reporting time, human–AI interaction, alert burden and effects on clinical decisions [31,32,35,36].
The issue of integration is also linked to interpretability, trust and responsibility. Tonekaboni et al. point out that clinicians’ explainability needs depend on the context of use and on how the output supports decision-making [37]. The user should understand whether the system is used for triage, prioritisation, detection, measurement, second reading or broader decision support, because acceptable error rates, supervision and responsibility depend on this role [31,32,37,38,39]. Asan et al. emphasise that clinicians’ trust is necessary for safe use, but must be calibrated to avoid both rejection of useful tools and over-reliance on automation [38]. Amann et al. point out that explainability in medical AI also has medical, legal, ethical and social dimensions, supporting the need for multidisciplinary implementation [39].
Ultimately, the clinical value of a model depends not only on accuracy, but on whether the system supports the actual diagnostic and treatment process. Kelly et al. emphasise that clinical impact requires a shift from research findings to solutions assessed in the context of real-world application, work organisation and patient benefit [35]. If a tool performs well on a technical task but fails to address a key clinical question or integrate with practice, its usefulness remains limited despite high metrics [35,36]. Radiology AI therefore requires not only improved algorithms and validation, but also sustained integration of clinical, technical and organisational environments that define intended use and clinical utility [35,39].

Emerging Approaches: Foundation Models, Generative AI and Local Adaptation

One important direction in medical AI is the shift from narrow task-specific models towards foundation models pre-trained on large and diverse datasets. Moor et al. describe this trend as the development of ‘generalist medical AI’, i.e. systems capable of performing diverse medical tasks with little or no task-specific labelled data [40]. According to Moor et al., such models may provide reusable representations of medical data, support different tasks and incorporate multiple modalities, including imaging, medical text, laboratory results and electronic health records [40]. However, this approach does not eliminate problems of data quality, domain shift, systematic error, interpretability and validation in the target clinical environment [10,11,29,40]. On the contrary, the increased complexity and opacity of foundation models can exacerbate these challenges, making auditability harder and masking dataset-specific shortcuts behind superficially impressive general capabilities.
Generative models and synthetic data represent another emerging approach that may address selected data limitations in radiology. Koetzier et al. describe synthetic medical imaging data as a potential way to augment training datasets, support contrast synthesis, reconstruct missing sequences and prepare educational material [41]. This is attractive in radiology because large, diverse and well-annotated datasets are often limited by privacy, annotation cost, rare diagnoses and barriers to data sharing [41]. At the same time, Koetzier et al. emphasise that synthetic images should be evaluated beyond visual realism, including their diversity, clinical plausibility, inherited bias and impact on model performance and generalisability [41]. The limitations of general-purpose generative tools illustrate this risk: an evaluation of anatomical images produced with Midjourney found frequent omissions of structures, disruptions of anatomical order and AI hallucinations, precluding their uncritical use as anatomical atlases [42]. Although general-purpose tools differ from dedicated medical image-synthesis models, the same caution applies: synthetic data should be used only after expert validation and evidence that it improves performance in the intended clinical task [41].
Acosta et al. point out that the growing availability of medical imaging, electronic health records, omics, wearable and other biomedical data enables multimodal models that better reflect the complexity of health and disease [43]. This is relevant to radiology because decisions rarely depend on a single image alone. Models combining images with reports, clinical records, laboratory results or prior history may better reflect the diagnostic process and partly address the lack of clinical context in image-only systems [35,43]. However, Acosta et al. also point out challenges related to data, modelling, privacy and interpretability. Multimodal models may also inherit errors and biases from each modality [19,43].
Rieke et al. describe federated learning as a way to train models on data distributed across institutions without centralised data collection, while Guan et al. discuss its applications and challenges in medical image analysis [21,23]. In this approach, data remain local, while model updates rather than raw patient data are exchanged between centres, which may support multicentre collaboration and reduce direct data transfer [21,23]. However, Guan et al. emphasise that federated learning in medical image analysis must still address data heterogeneity, acquisition differences and communication or organisational constraints [23]. If local datasets differ in acquisition protocols, labelling, population structure or documentation quality, the model may still be affected by heterogeneity and domain shift [10,12,23]. Federated training therefore requires not only technical infrastructure, but also data standards, quality control, consistent annotation and cross-centre validation [18,23,29].
At the same time, local model adaptation is becoming increasingly important and includes fine-tuning, calibration and post-deployment monitoring. In the context of medical imaging, Guan and Liu describe domain adaptation as an attempt to reduce distribution differences between related but distinct data domains, whilst Riley et al. emphasise the importance of evaluating model performance and calibration on new data [12,29]. Even a model pre-trained on a large dataset may need to be adapted to the local population, equipment, imaging protocols, reporting practices and organisational procedures [11,12,29,40]. However, local adaptation should not be understood as simply training the model on a small number of cases, but rather as part of a broader evaluation process, including validation in the target population and monitoring of the system’s performance following implementation [11,24,29]. This involves assessing effectiveness in the target population, conducting subgroup analyses, evaluating calibration, determining the decision threshold, and monitoring whether data drift occurs over time [11,27,29].
These approaches may reduce selected barriers related to data access, missing clinical context and domain shift, but they do not eliminate the need for local validation, performance monitoring and assessment of clinical impact [11,29,35,40,43]. Radiology AI should therefore be developed not as a standalone algorithm, but as part of the clinical process, requiring data quality assurance, interdisciplinary collaboration, local validation and post-implementation oversight [10,29,35,39].

Limitations of this Paper

A limitation of this review is its narrative nature and the deliberate selection of the literature. This study does not include a comprehensive, systematic identification of all publications on AI in radiology and medical imaging, nor does it include a meta-analysis or a formal assessment of the risk of error in all included studies. The conclusions should therefore be regarded as a qualitative synthesis of the most important translational issues, rather than a quantitative comparison of the effectiveness of specific models, architectures, or commercial AI systems.

Summary

Artificial intelligence has substantial potential to support radiological image interpretation, triage, segmentation, quantitative assessment and workflow organisation. However, high retrospective performance should not be equated with readiness for routine clinical use. The clinical value of radiology AI depends not only on the algorithm, but also on data quality, representativeness, validation, calibration, intended use and workflow integration.
The main translational barriers arise because medical imaging data are generated within specific technical, organisational and population contexts. Models trained on local or incomplete datasets may lose performance when transferred to other centres, scanner protocols, patient populations or workflows. Domain shift, data drift, shortcut learning, hidden stratification and poor subgroup performance may therefore remain invisible when evaluation relies mainly on internal validation and global metrics.
Reliable assessment requires external and, where appropriate, prospective validation, patient-, study- or centre-level data splitting, subgroup analysis, calibration assessment, transparent reporting, and interpretation of decision thresholds in relation to clinical consequences. Equally important is alignment with the clinical process. A model may be technically accurate but of limited value if it does not address a relevant radiology decision, lacks clinical context, fails to integrate with PACS/RIS/HIS/EHR infrastructure or increases workload.
Foundation models, generative AI, multimodal systems, federated learning and local adaptation may reduce selected barriers, but they do not replace local validation, post-deployment monitoring, data quality assurance, human oversight and assessment of clinical impact. The central question is therefore not only whether an algorithm performs well on a curated retrospective dataset, but whether it remains reliable, interpretable and clinically beneficial when deployed in the setting in which it will actually be used.

Funding

This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.

Acknowledgments

The authors have no acknowledgements to declare.

Conflicts of Interest

The authors declare no conflicts of interest.

Declaration of generative AI and AI-assisted technologies in the manuscript preparation process

During the preparation of this work, the authors used LeapSpace, DeepL, and OpenAI’s ChatGPT to support literature exploration, language editing, readability, and manuscript structuring. The authors reviewed and edited the output as needed and take full responsibility for the content of the published article.

References

  1. Aggarwal, R.; Sounderajah, V.; Martin, G.; Ting, D.S.W.; Karthikesalingam, A.; King, D.; Ashrafian, H.; Darzi, A. Diagnostic accuracy of deep learning in medical imaging: a systematic review and meta-analysis. npj Digit. Med. 2021, 4, 65. [Google Scholar] [CrossRef]
  2. Najjar, R. Redefining radiology: a review of artificial intelligence integration in medical imaging. Diagnostics 2023, 13, 2760. [Google Scholar] [CrossRef] [PubMed]
  3. McGenity, C.; Clarke, E.L.; Jennings, C.; Matthews, G.; Cartlidge, C.; Freduah-Agyemang, H.; Stocken, D.D.; Treanor, D. Artificial intelligence in digital pathology: a systematic review and meta-analysis of diagnostic test accuracy. npj Digit. Med. 2024, 7, 114. [Google Scholar] [CrossRef] [PubMed]
  4. Sivakumar, R.; Lue, B.; Kundu, S. FDA approval of artificial intelligence and machine learning devices in radiology: A systematic review. JAMA Netw. Open 2025, 8, e2542338. [Google Scholar] [CrossRef] [PubMed]
  5. U.S. Food and Drug Administration. Artificial Intelligence-Enabled Medical Devices. n.d. Available online: https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devices (accessed on 16 June 2026).
  6. Liu, X.; Faes, L.; Kale, A.U.; Wagner, S.K.; Fu, D.J.; Bruynseels, A.; Mahendiran, T.; Moraes, G.; Shamdas, M.; Kern, C. A comparison of deep learning performance against health-care professionals in detecting diseases from medical imaging: a systematic review and meta-analysis. Lancet Digit. Health 2019, 1, e271–e297. [Google Scholar] [CrossRef] [PubMed]
  7. Nagendran, M.; Chen, Y.; Lovejoy, C.A.; Gordon, A.C.; Komorowski, M.; Harvey, H.; Topol, E.J.; Ioannidis, J.P.A.; Collins, G.S.; Maruthappu, M. Artificial intelligence versus clinicians: systematic review of design, reporting standards, and claims of deep learning studies. BMJ 2020, 368, m689. [Google Scholar] [CrossRef] [PubMed]
  8. Yu, A.C.; Mohajer, B.; Eng, J. External validation of deep learning algorithms for radiologic diagnosis: a systematic review. Radiol. Artif. Intell. 2022, 4, e210064. [Google Scholar] [CrossRef] [PubMed]
  9. Zech, J.R.; Badgeley, M.A.; Liu, M.; Costa, A.B.; Titano, J.J.; Oermann, E.K. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLoS Med. 2018, 15, e1002683. [Google Scholar] [CrossRef] [PubMed]
  10. Schwabe, D.; Becker, K.; Seyferth, M.; Klaß, A.; Schaeffter, T. The METRIC-framework for assessing data quality for trustworthy AI in medicine: a systematic review. npj Digit. Med. 2024, 7, 203. [Google Scholar] [CrossRef] [PubMed]
  11. Finlayson, S.G.; Subbaswamy, A.; Singh, K.; Bowers, J.; Kupke, A.; Zittrain, J.; Kohane, I.S.; Saria, S. The clinician and dataset shift in artificial intelligence. N. Engl. J. Med. 2021, 385, 283–286. [Google Scholar] [CrossRef] [PubMed]
  12. Guan, H.; Liu, M. Domain adaptation for medical image analysis: a survey. IEEE Trans. Biomed. Eng. 2022, 69, 1173–1185. [Google Scholar] [CrossRef] [PubMed]
  13. Stacke, K.; Eilertsen, G.; Unger, J.; Lundström, C. Measuring domain shift for deep learning in histopathology. IEEE J. Biomed. Health Inform. 2021, 25, 325–336. [Google Scholar] [CrossRef] [PubMed]
  14. European Parliament and Council of the European Union. Regulation (EU) 2024/1689 — Artificial Intelligence Act. n.d. Available online: https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng (accessed on 13 June 2026).
  15. European Parliament and Council of the European Union. Regulation (EU) 2017/745 — Medical Device Regulation. n.d. Available online: https://eur-lex.europa.eu/eli/reg/2017/745/oj/eng (accessed on 13 June 2026).
  16. European Parliament and Council of the European Union. Regulation (EU) 2016/679 — General Data Protection Regulation, (2016). Available online: https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng (accessed on 13 June 2026).
  17. European Parliament and Council of the European Union. Regulation (EU) 2025/327 on the European Health Data Space, (2025). Available online: https://eur-lex.europa.eu/eli/reg/2025/327/oj/eng (accessed on 13 June 2026).
  18. Galbusera, F.; Cina, A. Image annotation and curation in radiology: an overview for machine learning practitioners. Eur. Radiol. Exp. 2024, 8 11. [Google Scholar]
  19. Celi, L.A.; Cellini, J.; Charpignon, M.-L.; Dee, E.C.; Dernoncourt, F.; Eber, R.; Mitchell, W.G.; Moukheiber, L.; Schirmer, J.; Situ, J. Sources of bias in artificial intelligence that perpetuate healthcare disparities—A global review. PLoS Digit Health 2022, 1, e0000022. [Google Scholar] [CrossRef] [PubMed]
  20. Salmi, M.; Atif, D.; Oliva, D.; Abraham, A.; Ventura, S. Handling imbalanced medical datasets: review of a decade of research. Artif. Intell. Rev. 2024, 57, 273. [Google Scholar] [CrossRef]
  21. Rieke, N.; Hancox, J.; Li, W.; Milletari, F.; Roth, H.R.; Albarqouni, S.; Bakas, S.; Galtier, M.N.; Landman, B.A.; Maier-Hein, K. The future of digital health with federated learning. npj Digit. Med. 2020, 3, 119. [Google Scholar] [CrossRef] [PubMed]
  22. International Medical Device Regulators Forum, Software as a Medical Device (SaMD): Clinical Evaluation, (2017). Available online: https://www.imdrf.org/documents/software-medical-device-samd-clinical-evaluation (accessed on 13 June 2026).
  23. Guan, H.; Yap, P.-T.; Bozoki, A.; Liu, M. Federated learning for medical image analysis: A survey. Pattern Recognit. 2024, 151, 110424. [Google Scholar] [CrossRef] [PubMed]
  24. Subbaswamy, S. Saria, From development to deployment: dataset shift, causality, and shift-stable models in health AI. Biostatistics 2020, 21, 345–352. [Google Scholar] [PubMed]
  25. Geirhos, R.; Jacobsen, J.-H.; Michaelis, C.; Zemel, R.; Brendel, W.; Bethge, M.; Wichmann, F.A. Shortcut learning in deep neural networks. Nat. Mach. Intell. 2020, 2, 665–673. [Google Scholar] [CrossRef]
  26. DeGrave, A.J.; Janizek, J.D.; Lee, S.-I. AI for radiographic COVID-19 detection selects shortcuts over signal. Nat. Mach. Intell. 2021, 3, 610–619. [Google Scholar] [CrossRef]
  27. Oakden-Rayner, L.; Dunnmon, J.; Carneiro, G.; Ré, C. Hidden stratification causes clinically meaningful failures in machine learning for medical imaging. Proc. ACM Conf. Health Inference Learn 2020, 151–159. [Google Scholar] [CrossRef] [PubMed]
  28. Adebayo, J.; Gilmer, J.; Muelly, M.; Goodfellow, I.; Hardt, M.; Kim, B. Sanity checks for saliency maps. Adv. Neural Inf. Process. Syst. 2018, 31. [Google Scholar]
  29. Riley, R.D.; Archer, L.; Snell, K.I.E.; Ensor, J.; Dhiman, P.; Martin, G.P.; Bonnett, L.J.; Collins, G.S. Evaluation of clinical prediction models (part 2): how to undertake an external validation study. BMJ 2024, 384, e074820. [Google Scholar] [CrossRef] [PubMed]
  30. Yagis, E.; Atnafu, S.W.; García Seco de Herrera, A.; Marzi, C.; Scheda, R.; Giannelli, M.; Tessa, C.; Citi, L.; Diciotti, S. Effect of data leakage in brain MRI classification using 2D convolutional neural networks. Sci. Rep. 2021, 11, 22544. [Google Scholar] [CrossRef] [PubMed]
  31. Liu, X.; Cruz Rivera, S.; Moher, D.; Calvert, M.J.; Denniston, A.K.; SPIRIT-AI and CONSORT-AI Working Group. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat. Med. 2020, 26, 1364–1374. [Google Scholar] [CrossRef] [PubMed]
  32. Rivera, S.C.; Liu, X.; Chan, A.-W.; Denniston, A.K.; Calvert, M.J.; SPIRIT-AI and CONSORT-AI Working Group. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI Extension. BMJ 2020, 370, m3210. [Google Scholar] [CrossRef] [PubMed]
  33. Collins, G.S.; Moons, K.G.M.; Dhiman, P.; Riley, R.D.; Beam, A.L.; Van Calster, B.; Ghassemi, M.; Liu, X.; Reitsma, J.B.; Van Smeden, M. TRIPOD+ AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024, 385, e078378. [Google Scholar] [CrossRef] [PubMed]
  34. Tejani, A.S.; Klontzas, M.E.; Gatti, A.A.; Mongan, J.T.; Moy, L.; Park, S.H.; C.E. Kahn, C., Jr. 2024 U. Panel, Checklist for artificial intelligence in medical imaging (CLAIM): 2024 update. Radiol. Artif. Intell. 2024, 6, e240300. [Google Scholar] [CrossRef] [PubMed]
  35. Kelly, C.J.; Karthikesalingam, A.; Suleyman, M.; Corrado, G.; King, D. Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 2019, 17, 195. [Google Scholar] [CrossRef] [PubMed]
  36. Tejani, A.S.; Cook, T.S.; Hussain, M.; Sippel Schmidt, T.; O’Donnell, K.P. Integrating and adopting AI in the radiology workflow: a primer for standards and integrating the healthcare enterprise (IHE) profiles. Radiology 2024, 311, e232653. [Google Scholar] [CrossRef] [PubMed]
  37. Tonekaboni, S.; Joshi, S.; McCradden, M.D.; Goldenberg, A. What clinicians want: contextualizing explainable machine learning for clinical end use. Machine Learning for Healthcare Conference, PMLR, 2019; pp. 359–380. [Google Scholar]
  38. Asan; A.E. Bayrak, A. Choudhury, Artificial intelligence and human trust in healthcare: focus on clinicians. J. Med. Internet Res. 2020, 22, e15154. [Google Scholar] [CrossRef] [PubMed]
  39. Amann, J.; Blasimme, A.; Vayena, E.; Frey, D.; V.I. Madai, P. Consortium, Explainability for artificial intelligence in healthcare: a multidisciplinary perspective. BMC Med. Inform. Decis. Mak. 2020, 20, 310. [Google Scholar] [CrossRef] [PubMed]
  40. Moor, M.; Banerjee, O.; Abad, Z.S.H.; Krumholz, H.M.; Leskovec, J.; Topol, E.J.; Rajpurkar, P. Foundation models for generalist medical artificial intelligence. Nature 2023, 616, 259–265. [Google Scholar] [CrossRef] [PubMed]
  41. Koetzier, L.R.; Wu, J.; Mastrodicasa, D.; Lutz, A.; Chung, M.; Koszek, W.A.; Pratap, J.; Chaudhari, A.S.; Rajpurkar, P.; Lungren, M.P. Generating synthetic data for medical imaging. Radiology 2024, 312, e232471. [Google Scholar] [CrossRef] [PubMed]
  42. Kozioł, M.; Kozioł, T.S.; Batko, J.; Paniec, P.; Wajda, J.; Depukat, P.; Moryś, J.; Klejbor, I.; Walocha, J.; Koziej, M. AI-natomy: human anatomy through the eyes of artificial intelligence. Is there a distinction between reality and imagination? Folia Morphol. (Warsz) 2025. [Google Scholar] [CrossRef] [PubMed]
  43. Acosta, J.N.; Falcone, G.J.; Rajpurkar, P.; Topol, E.J. Multimodal biomedical AI. Nat. Med. 2022, 28, 1773–1784. [Google Scholar] [CrossRef] [PubMed]
Figure 1. Translational gap in radiology AI. The diagram shows stages at which model performance may lose clinical relevance: retrospective data collection, model training, internal validation, external validation, local calibration, prospective validation, clinical deployment and post-deployment monitoring.
Figure 1. Translational gap in radiology AI. The diagram shows stages at which model performance may lose clinical relevance: retrospective data collection, model training, internal validation, external validation, local calibration, prospective validation, clinical deployment and post-deployment monitoring.
Preprints 224642 g001
Table 1. Mechanisms limiting the clinical translation of radiology AI systems.
Table 1. Mechanisms limiting the clinical translation of radiology AI systems.
Mechanism Source of the problem Possible clinical consequence Methods of risk mitigation
Limited data quality and completeness inconsistent labels, missing metadata, incomplete clinical context, differences in data preparation and processing the model may replicate local practices relating to documentation, diagnosis or data preparation rather than the stable characteristics of the disease assessment of data quality in relation to clinical application, label checking, complete metadata, data standardisation and curation
Lack of representativeness and data silos data from a single centre, a limited population or a single healthcare system, difficulties in sharing data between institutions good average performance may not reflect the model’s performance in other populations, centres or patient subgroups multicentre data, description of the target population, subgroup analysis, controlled access to data, common data standards
Domain shift and data drift differences in scanner vendor, acquisition protocols, reconstruction methods, patient populations, or changes in these factors over time a decline in effectiveness when the model is applied in a different environment, or a gradual deterioration in performance following implementation external validation, local calibration, domain adaptation, performance monitoring and periodic revalidation
Shortcut learning and hidden stratification technical, organisational or population-related factors associated with the label; unspecified subtypes of the disease or patient subgroups A high overall performance metric may mask the use of spurious signals or poor efficacy in a clinically important subgroup external validation, error analysis, subgroup analysis, robustness testing, cautious use of explainability methods
Limitations on the validation and interpretation of metrics internal validation using similar data, data leakage between the training and test sets, reliance on global metrics without assessing calibration and decision thresholds overestimation of the model’s performance and the risk that a high AUC, accuracy or Dice score will not translate into clinical utility data splitting at the level of the patient, examination, study, series or centre; external validation; assessment of calibration, uncertainty, decision thresholds and results in subgroups
Mismatch between the model’s objective and the clinical decision reducing a clinical problem to a technical task without specifying how the system’s output is intended to support medical decision-making a model may be technically sound, but it may not answer a key clinical question or change diagnostic or therapeutic management a clear definition of the intended use, the target population and the clinical decision; an interdisciplinary definition of the system’s purpose
Regulatory barriers and integration into practice the need to meet the requirements for a medical device or SaMD; lack of integration with PACS/RIS/HIS/EHR; an excessive number of alerts; unclear user responsibility good results in a publication do not necessarily mean the system is ready for clinical use; the system may increase the workload on staff or may not be adopted in practice definition of the intended use, risk management, clinical evaluation, human oversight, integration into the workflow and assessment of the impact on staff work
Limitations of emerging AI approaches foundation, generative, multimodal, federated and locally adapted models continue to depend on data quality, representativeness, validation and clinical integration a larger scale or more advanced architecture may reduce some of the barriers, but does not guarantee safety, effectiveness or usability in the local clinical environment local validation, subgroup analysis, calibration assessment, data quality control, post-implementation performance monitoring and assessment of the impact on clinical practice
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings