Submitted:
23 July 2026
Posted:
24 July 2026
You are already at the latest version
Abstract

Keywords:
Introduction
Methodology and Selection of Literature
Limitations Regarding Data Quality and Representativeness
Regulatory and Organisational Barriers to Radiology AI Deployment
Domain Shift
Shortcut Learning and Hidden Stratification
Limitations on the Validation and Interpretation of Model Results
Clinical Objectives and Workflow Integration in Radiology Practice
Emerging Approaches: Foundation Models, Generative AI and Local Adaptation
Limitations of this Paper
Summary
Funding
Acknowledgments
Conflicts of Interest
Declaration of generative AI and AI-assisted technologies in the manuscript preparation process
References
- Aggarwal, R.; Sounderajah, V.; Martin, G.; Ting, D.S.W.; Karthikesalingam, A.; King, D.; Ashrafian, H.; Darzi, A. Diagnostic accuracy of deep learning in medical imaging: a systematic review and meta-analysis. npj Digit. Med. 2021, 4, 65. [Google Scholar] [CrossRef]
- Najjar, R. Redefining radiology: a review of artificial intelligence integration in medical imaging. Diagnostics 2023, 13, 2760. [Google Scholar] [CrossRef] [PubMed]
- McGenity, C.; Clarke, E.L.; Jennings, C.; Matthews, G.; Cartlidge, C.; Freduah-Agyemang, H.; Stocken, D.D.; Treanor, D. Artificial intelligence in digital pathology: a systematic review and meta-analysis of diagnostic test accuracy. npj Digit. Med. 2024, 7, 114. [Google Scholar] [CrossRef] [PubMed]
- Sivakumar, R.; Lue, B.; Kundu, S. FDA approval of artificial intelligence and machine learning devices in radiology: A systematic review. JAMA Netw. Open 2025, 8, e2542338. [Google Scholar] [CrossRef] [PubMed]
- U.S. Food and Drug Administration. Artificial Intelligence-Enabled Medical Devices. n.d. Available online: https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devices (accessed on 16 June 2026).
- Liu, X.; Faes, L.; Kale, A.U.; Wagner, S.K.; Fu, D.J.; Bruynseels, A.; Mahendiran, T.; Moraes, G.; Shamdas, M.; Kern, C. A comparison of deep learning performance against health-care professionals in detecting diseases from medical imaging: a systematic review and meta-analysis. Lancet Digit. Health 2019, 1, e271–e297. [Google Scholar] [CrossRef] [PubMed]
- Nagendran, M.; Chen, Y.; Lovejoy, C.A.; Gordon, A.C.; Komorowski, M.; Harvey, H.; Topol, E.J.; Ioannidis, J.P.A.; Collins, G.S.; Maruthappu, M. Artificial intelligence versus clinicians: systematic review of design, reporting standards, and claims of deep learning studies. BMJ 2020, 368, m689. [Google Scholar] [CrossRef] [PubMed]
- Yu, A.C.; Mohajer, B.; Eng, J. External validation of deep learning algorithms for radiologic diagnosis: a systematic review. Radiol. Artif. Intell. 2022, 4, e210064. [Google Scholar] [CrossRef] [PubMed]
- Zech, J.R.; Badgeley, M.A.; Liu, M.; Costa, A.B.; Titano, J.J.; Oermann, E.K. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLoS Med. 2018, 15, e1002683. [Google Scholar] [CrossRef] [PubMed]
- Schwabe, D.; Becker, K.; Seyferth, M.; Klaß, A.; Schaeffter, T. The METRIC-framework for assessing data quality for trustworthy AI in medicine: a systematic review. npj Digit. Med. 2024, 7, 203. [Google Scholar] [CrossRef] [PubMed]
- Finlayson, S.G.; Subbaswamy, A.; Singh, K.; Bowers, J.; Kupke, A.; Zittrain, J.; Kohane, I.S.; Saria, S. The clinician and dataset shift in artificial intelligence. N. Engl. J. Med. 2021, 385, 283–286. [Google Scholar] [CrossRef] [PubMed]
- Guan, H.; Liu, M. Domain adaptation for medical image analysis: a survey. IEEE Trans. Biomed. Eng. 2022, 69, 1173–1185. [Google Scholar] [CrossRef] [PubMed]
- Stacke, K.; Eilertsen, G.; Unger, J.; Lundström, C. Measuring domain shift for deep learning in histopathology. IEEE J. Biomed. Health Inform. 2021, 25, 325–336. [Google Scholar] [CrossRef] [PubMed]
- European Parliament and Council of the European Union. Regulation (EU) 2024/1689 — Artificial Intelligence Act. n.d. Available online: https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng (accessed on 13 June 2026).
- European Parliament and Council of the European Union. Regulation (EU) 2017/745 — Medical Device Regulation. n.d. Available online: https://eur-lex.europa.eu/eli/reg/2017/745/oj/eng (accessed on 13 June 2026).
- European Parliament and Council of the European Union. Regulation (EU) 2016/679 — General Data Protection Regulation, (2016). Available online: https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng (accessed on 13 June 2026).
- European Parliament and Council of the European Union. Regulation (EU) 2025/327 on the European Health Data Space, (2025). Available online: https://eur-lex.europa.eu/eli/reg/2025/327/oj/eng (accessed on 13 June 2026).
- Galbusera, F.; Cina, A. Image annotation and curation in radiology: an overview for machine learning practitioners. Eur. Radiol. Exp. 2024, 8 11. [Google Scholar]
- Celi, L.A.; Cellini, J.; Charpignon, M.-L.; Dee, E.C.; Dernoncourt, F.; Eber, R.; Mitchell, W.G.; Moukheiber, L.; Schirmer, J.; Situ, J. Sources of bias in artificial intelligence that perpetuate healthcare disparities—A global review. PLoS Digit Health 2022, 1, e0000022. [Google Scholar] [CrossRef] [PubMed]
- Salmi, M.; Atif, D.; Oliva, D.; Abraham, A.; Ventura, S. Handling imbalanced medical datasets: review of a decade of research. Artif. Intell. Rev. 2024, 57, 273. [Google Scholar] [CrossRef]
- Rieke, N.; Hancox, J.; Li, W.; Milletari, F.; Roth, H.R.; Albarqouni, S.; Bakas, S.; Galtier, M.N.; Landman, B.A.; Maier-Hein, K. The future of digital health with federated learning. npj Digit. Med. 2020, 3, 119. [Google Scholar] [CrossRef] [PubMed]
- International Medical Device Regulators Forum, Software as a Medical Device (SaMD): Clinical Evaluation, (2017). Available online: https://www.imdrf.org/documents/software-medical-device-samd-clinical-evaluation (accessed on 13 June 2026).
- Guan, H.; Yap, P.-T.; Bozoki, A.; Liu, M. Federated learning for medical image analysis: A survey. Pattern Recognit. 2024, 151, 110424. [Google Scholar] [CrossRef] [PubMed]
- Subbaswamy, S. Saria, From development to deployment: dataset shift, causality, and shift-stable models in health AI. Biostatistics 2020, 21, 345–352. [Google Scholar] [PubMed]
- Geirhos, R.; Jacobsen, J.-H.; Michaelis, C.; Zemel, R.; Brendel, W.; Bethge, M.; Wichmann, F.A. Shortcut learning in deep neural networks. Nat. Mach. Intell. 2020, 2, 665–673. [Google Scholar] [CrossRef]
- DeGrave, A.J.; Janizek, J.D.; Lee, S.-I. AI for radiographic COVID-19 detection selects shortcuts over signal. Nat. Mach. Intell. 2021, 3, 610–619. [Google Scholar] [CrossRef]
- Oakden-Rayner, L.; Dunnmon, J.; Carneiro, G.; Ré, C. Hidden stratification causes clinically meaningful failures in machine learning for medical imaging. Proc. ACM Conf. Health Inference Learn 2020, 151–159. [Google Scholar] [CrossRef] [PubMed]
- Adebayo, J.; Gilmer, J.; Muelly, M.; Goodfellow, I.; Hardt, M.; Kim, B. Sanity checks for saliency maps. Adv. Neural Inf. Process. Syst. 2018, 31. [Google Scholar]
- Riley, R.D.; Archer, L.; Snell, K.I.E.; Ensor, J.; Dhiman, P.; Martin, G.P.; Bonnett, L.J.; Collins, G.S. Evaluation of clinical prediction models (part 2): how to undertake an external validation study. BMJ 2024, 384, e074820. [Google Scholar] [CrossRef] [PubMed]
- Yagis, E.; Atnafu, S.W.; García Seco de Herrera, A.; Marzi, C.; Scheda, R.; Giannelli, M.; Tessa, C.; Citi, L.; Diciotti, S. Effect of data leakage in brain MRI classification using 2D convolutional neural networks. Sci. Rep. 2021, 11, 22544. [Google Scholar] [CrossRef] [PubMed]
- Liu, X.; Cruz Rivera, S.; Moher, D.; Calvert, M.J.; Denniston, A.K.; SPIRIT-AI and CONSORT-AI Working Group. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nat. Med. 2020, 26, 1364–1374. [Google Scholar] [CrossRef] [PubMed]
- Rivera, S.C.; Liu, X.; Chan, A.-W.; Denniston, A.K.; Calvert, M.J.; SPIRIT-AI and CONSORT-AI Working Group. Guidelines for clinical trial protocols for interventions involving artificial intelligence: the SPIRIT-AI Extension. BMJ 2020, 370, m3210. [Google Scholar] [CrossRef] [PubMed]
- Collins, G.S.; Moons, K.G.M.; Dhiman, P.; Riley, R.D.; Beam, A.L.; Van Calster, B.; Ghassemi, M.; Liu, X.; Reitsma, J.B.; Van Smeden, M. TRIPOD+ AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024, 385, e078378. [Google Scholar] [CrossRef] [PubMed]
- Tejani, A.S.; Klontzas, M.E.; Gatti, A.A.; Mongan, J.T.; Moy, L.; Park, S.H.; C.E. Kahn, C., Jr. 2024 U. Panel, Checklist for artificial intelligence in medical imaging (CLAIM): 2024 update. Radiol. Artif. Intell. 2024, 6, e240300. [Google Scholar] [CrossRef] [PubMed]
- Kelly, C.J.; Karthikesalingam, A.; Suleyman, M.; Corrado, G.; King, D. Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 2019, 17, 195. [Google Scholar] [CrossRef] [PubMed]
- Tejani, A.S.; Cook, T.S.; Hussain, M.; Sippel Schmidt, T.; O’Donnell, K.P. Integrating and adopting AI in the radiology workflow: a primer for standards and integrating the healthcare enterprise (IHE) profiles. Radiology 2024, 311, e232653. [Google Scholar] [CrossRef] [PubMed]
- Tonekaboni, S.; Joshi, S.; McCradden, M.D.; Goldenberg, A. What clinicians want: contextualizing explainable machine learning for clinical end use. Machine Learning for Healthcare Conference, PMLR, 2019; pp. 359–380. [Google Scholar]
- Asan; A.E. Bayrak, A. Choudhury, Artificial intelligence and human trust in healthcare: focus on clinicians. J. Med. Internet Res. 2020, 22, e15154. [Google Scholar] [CrossRef] [PubMed]
- Amann, J.; Blasimme, A.; Vayena, E.; Frey, D.; V.I. Madai, P. Consortium, Explainability for artificial intelligence in healthcare: a multidisciplinary perspective. BMC Med. Inform. Decis. Mak. 2020, 20, 310. [Google Scholar] [CrossRef] [PubMed]
- Moor, M.; Banerjee, O.; Abad, Z.S.H.; Krumholz, H.M.; Leskovec, J.; Topol, E.J.; Rajpurkar, P. Foundation models for generalist medical artificial intelligence. Nature 2023, 616, 259–265. [Google Scholar] [CrossRef] [PubMed]
- Koetzier, L.R.; Wu, J.; Mastrodicasa, D.; Lutz, A.; Chung, M.; Koszek, W.A.; Pratap, J.; Chaudhari, A.S.; Rajpurkar, P.; Lungren, M.P. Generating synthetic data for medical imaging. Radiology 2024, 312, e232471. [Google Scholar] [CrossRef] [PubMed]
- Kozioł, M.; Kozioł, T.S.; Batko, J.; Paniec, P.; Wajda, J.; Depukat, P.; Moryś, J.; Klejbor, I.; Walocha, J.; Koziej, M. AI-natomy: human anatomy through the eyes of artificial intelligence. Is there a distinction between reality and imagination? Folia Morphol. (Warsz) 2025. [Google Scholar] [CrossRef] [PubMed]
- Acosta, J.N.; Falcone, G.J.; Rajpurkar, P.; Topol, E.J. Multimodal biomedical AI. Nat. Med. 2022, 28, 1773–1784. [Google Scholar] [CrossRef] [PubMed]

| Mechanism | Source of the problem | Possible clinical consequence | Methods of risk mitigation |
|---|---|---|---|
| Limited data quality and completeness | inconsistent labels, missing metadata, incomplete clinical context, differences in data preparation and processing | the model may replicate local practices relating to documentation, diagnosis or data preparation rather than the stable characteristics of the disease | assessment of data quality in relation to clinical application, label checking, complete metadata, data standardisation and curation |
| Lack of representativeness and data silos | data from a single centre, a limited population or a single healthcare system, difficulties in sharing data between institutions | good average performance may not reflect the model’s performance in other populations, centres or patient subgroups | multicentre data, description of the target population, subgroup analysis, controlled access to data, common data standards |
| Domain shift and data drift | differences in scanner vendor, acquisition protocols, reconstruction methods, patient populations, or changes in these factors over time | a decline in effectiveness when the model is applied in a different environment, or a gradual deterioration in performance following implementation | external validation, local calibration, domain adaptation, performance monitoring and periodic revalidation |
| Shortcut learning and hidden stratification | technical, organisational or population-related factors associated with the label; unspecified subtypes of the disease or patient subgroups | A high overall performance metric may mask the use of spurious signals or poor efficacy in a clinically important subgroup | external validation, error analysis, subgroup analysis, robustness testing, cautious use of explainability methods |
| Limitations on the validation and interpretation of metrics | internal validation using similar data, data leakage between the training and test sets, reliance on global metrics without assessing calibration and decision thresholds | overestimation of the model’s performance and the risk that a high AUC, accuracy or Dice score will not translate into clinical utility | data splitting at the level of the patient, examination, study, series or centre; external validation; assessment of calibration, uncertainty, decision thresholds and results in subgroups |
| Mismatch between the model’s objective and the clinical decision | reducing a clinical problem to a technical task without specifying how the system’s output is intended to support medical decision-making | a model may be technically sound, but it may not answer a key clinical question or change diagnostic or therapeutic management | a clear definition of the intended use, the target population and the clinical decision; an interdisciplinary definition of the system’s purpose |
| Regulatory barriers and integration into practice | the need to meet the requirements for a medical device or SaMD; lack of integration with PACS/RIS/HIS/EHR; an excessive number of alerts; unclear user responsibility | good results in a publication do not necessarily mean the system is ready for clinical use; the system may increase the workload on staff or may not be adopted in practice | definition of the intended use, risk management, clinical evaluation, human oversight, integration into the workflow and assessment of the impact on staff work |
| Limitations of emerging AI approaches | foundation, generative, multimodal, federated and locally adapted models continue to depend on data quality, representativeness, validation and clinical integration | a larger scale or more advanced architecture may reduce some of the barriers, but does not guarantee safety, effectiveness or usability in the local clinical environment | local validation, subgroup analysis, calibration assessment, data quality control, post-implementation performance monitoring and assessment of the impact on clinical practice |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).