Submitted:
03 July 2026
Posted:
06 July 2026
You are already at the latest version
Abstract
Keywords:
Introduction
Methods
Study Design and Registration
Eligibility Criteria
Literature Search
Data Extraction
| Tier | Feature Name | Description and Extraction Method |
|---|---|---|
| 1 | internal_auc | Area under the ROC curve on internal/held-out validation set; extracted from primary results table |
| 1 | log_train_size | Natural logarithm of the number of training images reported; log-transformation applied to normalise right-skewed distribution |
| 1 | hdi_gap | Absolute difference in Human Development Index between training-site country and external validation site country; sourced from UNDP 2024 database |
| 2 | architecture_type | Categorical: ResNet family, VGG family, EfficientNet, custom CNN, Transformer, other; extracted from methods section |
| 2 | imaging_modality | Categorical: ophthalmology fundus, chest X-ray, computed tomography, other; extracted from methods section |
| 2 | class_balance_ratio | Ratio of majority to minority class in training set; extracted from dataset description; imputed as 1.0 if not reported |
| 2 | multi_center_train | Binary: training data drawn from ≥2 institutions (1) or single centre (0) |
| 2 | publication_year | Calendar year of publication; included as proxy for temporal methodological maturity |
| 3 | augmentation_applied | Binary: any data augmentation strategy explicitly described (1) or not mentioned (0) |
| 3 | external_dataset_size | Number of images in external validation dataset; log-transformed |
| 3 | domain_shift_type | Categorical: geographic shift, temporal shift, scanner shift, demographic shift, mixed; coded from methods/validation description |
| 4 | preprocessing_standardised | Binary: explicit image normalisation or standardisation pipeline described (1) or absent (0) |
| 4 | demographic_reported | Binary: age, sex, or ethnicity distribution of training set explicitly reported (1) or not (0) |
Outcome Variable
Statistical Analysis
Progressive Metadata Tier Regression
Non-Parametric Bootstrap Replicability Audit
Modality Sensitivity Analysis
Results
Study Identification and Inclusion
Study Characteristics
| Characteristic | Fundus (n=32) | Chest X-Ray (n=35) | CT (n=33) |
|---|---|---|---|
| Publication year ≥ 2020, n (%) | 22 (69) | 24 (69) | 22 (67) |
| Multicenter training, n (%) | 17 (53) | 22 (63) | 19 (58) |
| Augmentation reported, n (%) | 26 (81) | 28 (80) | 25 (76) |
| Demographic data reported, n (%) | 11 (34) | 14 (40) | 10 (30) |
| Preprocessing standardised, n (%) | 14 (44) | 17 (49) | 15 (45) |
| Median train size (IQR) | 9,800 (3,100–32,000) | 12,400 (4,800–41,000) | 11,200 (4,000–36,000) |
| Median internal AUC (IQR) | 0.91 (0.87–0.95) | 0.89 (0.85–0.93) | 0.92 (0.88–0.96) |
| Median ΔAUC (IQR) | 0.09 (0.06–0.12) | 0.09 (0.05–0.12) | 0.09 (0.05–0.12) |
Information Sufficiency Audit: Progressive Metadata Tier Analysis
Non-Parametric Bootstrap Replicability Audit
Validation of Cross-Validated R2 Estimates
| Tier | Description | Features Added | CV R2 | 95% CI |
|---|---|---|---|---|
| 1 | Baseline Metrics | internal_auc, log_train_size, hdi_gap | −0.3438 | −0.44 to −0.25 |
| 2 | Standard Published Metadata | +architecture_type, imaging_modality, class_balance_ratio, multi_center_train, publication_year | −0.4002 | −0.51 to −0.30 |
| 3 | Study Design Variables | +augmentation_applied, external_dataset_size, domain_shift_type | −0.5025 | −0.61 to −0.39 |
| 4 | Full 13-Feature Expansion | +preprocessing_standardised, demographic_reported | −0.6053 | −0.72 to −0.49 |
Modality Sensitivity Profile
Discussion
Principal Findings
Interpretation of Negative Cross-Validated R2
Minimum Viable Reporting Set
Modality-General Nature of the Explanatory Deficit
Comparison with Prior Literature
Limitations
Future Directions
Conclusion
| Feature | Tier | Selection Freq (%) | Exceeds 5% Threshold | Classification |
| augmentation_applied | 3 | ≈65% | Yes | Replicable signal |
| class_balance_ratio | 2 | ≈23% | Yes | Replicable signal |
| architecture_type | 2 | ≈15% | Yes | Replicable signal |
| multi_center_train | 2 | ≈13% | Yes | Replicable signal |
| internal_auc | 1 | ≈6% | Marginal | Near-threshold |
| log_train_size | 1 | ≈4% | No | Noise level |
| hdi_gap | 1 | ≈3% | No | Noise level |
| publication_year | 2 | — | No | Not extracted |
| imaging_modality | 2 | — | No | Not extracted |
| preprocessing_standardised | 4 | — | No | Not extracted |
| demographic_reported | 4 | — | No | Not extracted |
| external_dataset_size | 3 | — | No | Not extracted |
| domain_shift_type | 3 | — | No | Not extracted |
Funding
Institutional Review Board Statement
Data Availability Statement
Conflicts of Interest
References
- Nagendran, M.; Chen, Y.; Lovejoy, C.A.; et al. Artificial intelligence versus clinicians: systematic review of design, reporting standards, and claims of deep learning studies in medical imaging. BMJ. 2020, 368, m689. [Google Scholar] [CrossRef] [PubMed]
- Liu, X.; Faes, L.; Kale, A.U.; et al. A comparison of deep learning performance against health-care professionals in detecting diseases from medical imaging: a systematic review and meta-analysis. Lancet Digit. Health 2019, 1, e271–e297. [Google Scholar] [CrossRef] [PubMed]
- Collins, G.S.; Dhiman, P.; Andaur Navarro, C.L.; et al. Protocol for development of a reporting guideline (TRIPOD-AI) and risk of bias tool (PROBAST-AI) for diagnostic and prognostic prediction model studies based on artificial intelligence. BMJ Open. 2021, 11, e048008. [Google Scholar] [CrossRef] [PubMed]
- Muhammad, W.; Hart, G.R.; Nartowt, B.; et al. Inclusive machine learning for medical imaging applications: avoiding barriers in data curation and model evaluation. Lancet Digit. Health 2022, 4, e243–e244. [Google Scholar]
- Hand, D.J. Measuring classifier performance: a coherent alternative to the area under the ROC curve. Mach. Learn. 2009, 77, 103–123. [Google Scholar] [CrossRef]
- Breiman, L. Random forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef]
- Robnik-Sikonja, M.; Kononenko, I. Explaining classifications for individual instances. IEEE Trans. Knowl. Data Eng. 2008, 20, 589–600. [Google Scholar] [CrossRef]
- Esteva, A.; Kuprel, B.; Novoa, R.A.; et al. Dermatologist-level classification of skin cancer with deep neural networks. Nature 2017, 542, 115–118. [Google Scholar] [CrossRef] [PubMed]
- Rajpurkar, P.; Irvin, J.; Ball, R.L.; et al. Deep learning for chest radiograph diagnosis: a retrospective comparison of the CheXNeXt algorithm to practicing radiologists. PLoS Med. 2018, 15, e1002686. [Google Scholar] [CrossRef] [PubMed]
- Gulshan, V.; Peng, L.; Coram, M.; et al. Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA 2016, 316, 2402–2410. [Google Scholar] [CrossRef] [PubMed]
- Kumari, S.; Bhattacharjee, A.; Bhattacharya, A. Domain adaptation in medical imaging: a systematic review. IEEE Trans. BioMed Eng. 2024, 71, 1–16. [Google Scholar] [CrossRef]
- Gargeya, R.; Leng, T. Automated identification of diabetic retinopathy using deep learning. Ophthalmology 2017, 124, 962–969. [Google Scholar] [CrossRef] [PubMed]
- Pooch, E.H.P.; Ballester, P.; Barros, R.C. Can we trust deep learning based diagnosis? The impact of domain shift in chest radiograph classification. Lect. Notes Comput Sci. 2020, 12102, 74–83. [Google Scholar] [CrossRef]
- Yang, J.; Shi, R.; Wei, D.; et al. MedMNIST v2: a large-scale lightweight benchmark for 2D and 3D biomedical image classification. Sci. Data 2023, 10, 41. [Google Scholar] [CrossRef] [PubMed]
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [PubMed]
- van der Velden, B.H.M.; Kuijf, H.J.; Gilhuijs, K.G.A.; Viergever, M.A. Explainable artificial intelligence (XAI) in deep learning-based medical image analysis. Med. Image Anal. 2022, 79, 102470. [Google Scholar] [CrossRef] [PubMed]
- Woerl, A.-C.; Eckstein, M.; Geiger, J.; et al. Deep learning predicts molecular subtype of muscle-invasive bladder cancer from conventional histopathological slides. Eur. Urol. 2023, 83, 558–567. [Google Scholar]
- Seyyed-Kalantari, L.; Zhang, H.; McDermott, M.B.A.; Chen, I.Y.; Ghassemi, M. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat. Med. 2021, 27, 2176–2182. [Google Scholar] [CrossRef] [PubMed]
- Woerl, A.-C.; Brieu, N.; Ziegler, K.; et al. Deep learning-based diagnosis of lung cancer subtypes on histopathological images. Diagnostics 2022, 12, 1465. [Google Scholar]
- Irvin, J.; Rajpurkar, P.; Ko, M.; et al. CheXpert: a large chest radiograph dataset with uncertainty labels and expert comparison. AAAI 2019, 33, 590–597. [Google Scholar] [CrossRef]





Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).