Submitted:
03 August 2026
Posted:
04 August 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Materials and Methods
2.1. Study Design
2.2. Case Selection and Preparation
- •
- Demographics and Clinical Context: Age, sex, body mass index (BMI), clinical indication for MPS (e.g., “exertional chest pain,” “preoperative risk stratification”), and relevant cardiac history (prior MI, revascularization, known CAD).
- •
- Stress Protocol: Type of stress (exercise treadmill Bruce protocol vs. pharmacologic with adenosine), peak heart rate, blood pressure response, and symptoms during stress.
- •
- Perfusion Findings (Rest): Segmental perfusion status for each vascular territory (left anterior descending [LAD], left circumflex [LCx], right coronary artery [RCA]) using a 17-segment model, categorized as normal, mildly reduced, moderately reduced, or severely reduced.
- •
- Perfusion Findings (Stress): Segmental perfusion status under stress, with explicit notation of reversibility (reversible, fixed, partially reversible) for each abnormal segment.
- •
- Left Ventricular Function: Ejection fraction (%), end-diastolic volume (mL), end-systolic volume (mL), regional wall motion (normal, hypokinetic, akinetic, dyskinetic), and wall thickening assessment.
- •
- Technical Quality Indicators: Presence of motion artifacts, attenuation artifacts (breast, diaphragmatic, or body habitus-related), submaximal stress achievement, or tracer uptake issues.
2.3. ChatGPT-4 Reporting Procedure
2.3.1. Technical Specifications and Reproducibility
2.3.2. Input Feature Construction
2.3.3. Processing Mechanism of Structured Clinical Data by the LLM
2.4. Selection of Evaluation Criteria
2.5. Physician Comparison and Evaluation Process
2.6. Data Analysis
2.7. Ethical Considerations
3. Results
3.1. Overall Performance Evaluation
3.2. Performance by Clinical Scenario
3.3. Qualitative Observations
3.4. Statistical Results
4. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Notghi, A.; Low, C.S. Myocardial perfusion scintigraphy: past, present and future. Br. J. Radiol. 2011, 84(Spec Iss 3), S229–S236. [Google Scholar] [CrossRef] [PubMed]
- Liga, R.; Vontobel, J.; Rovai, D.; Marinelli, M.; Caselli, C.; Pietila, M.; et al. Multicentre multi-device hybrid imaging study of coronary artery disease: results from the EValuation of INtegrated Cardiac Imaging for the Detection and Characterization of Ischaemic Heart Disease (EVINCI) hybrid imaging population. Eur. Heart J. Cardiovasc Imaging 2016, 17(9), 951–960. [Google Scholar] [CrossRef] [PubMed]
- Lopes, L.; Lopez-Montes, A.; Chen, Y.; Koller, P.; Rathod, N.; Blomgren, A.; et al. The evolution of artificial intelligence in nuclear medicine. Semin Nucl. Med. 2025, 55(3), 313–327. [Google Scholar] [CrossRef] [PubMed]
- Schopow, N.; Osterhoff, G.; Baur, D. Applications of the natural language processing tool ChatGPT in clinical practice: comparative study and augmented systematic review. JMIR Med. Inf. 2023, 11, e48933. [Google Scholar] [CrossRef] [PubMed]
- Rao, S.J.; Isath, A.; Krishnan, P.; Tangsrivimol, J.A.; Virk, H.U.H.; Wang, Z.; et al. ChatGPT: a conceptual review of applications and utility in the field of medicine. J. Med. Syst. 2024, 48(1), 59. [Google Scholar] [CrossRef] [PubMed]
- Khullar, D.; Wang, X.; Wang, F. Large language models in health care: charting a path toward accurate, explainable, and secure AI. J. Gen. Intern Med. 2024, 39(7), 1239–1241. [Google Scholar] [CrossRef] [PubMed]
- Busch, F.; Hoffmann, L.; Dos Santos, D.P.; Makowski, M.R.; Saba, L.; Prucker, P.; et al. Large language models for structured reporting in radiology: past, present, and future. Eur. Radiol. 2025, 35(5), 2589–2602. [Google Scholar] [CrossRef] [PubMed]
- Adams, L.C.; Truhn, D.; Busch, F.; Kader, A.; Niehues, S.M.; Makowski, M.R.; Bressem, K.K. Leveraging GPT-4 for post hoc transformation of free-text radiology reports into structured reporting: a multilingual feasibility study. Radiology 2023, 307(4), e230725. [Google Scholar] [CrossRef] [PubMed]
- Biswas, S. ChatGPT and the future of medical writing. Radiology 2023, 307(2), e223312. [Google Scholar] [CrossRef] [PubMed]
- Klang, E.; Collins, J.D.; Glicksberg, B.S.; Korfiatis, P.; Nadkarni, G.; Sorin, V. Large language models in radiology reporting: a systematic review of performance, limitations, and clinical implications. medRxiv 2025. [Google Scholar] [CrossRef]
- Choi, H.; Lee, D.; Kang, Y.K.; Suh, M. Empowering PET imaging reporting with retrieval-augmented large language models and reading reports database: a pilot single center study. Eur. J. Nucl. Med. Mol. Imaging 2025, 52(7), 2452–2462. [Google Scholar] [CrossRef] [PubMed]
- Hirata, K.; et al. Generative AI and large language models in nuclear medicine: current status and future prospects. Ann. Nucl. Med. 2024, 38(11), 853–864. [Google Scholar] [CrossRef] [PubMed]
- Licu, R.A.; et al. Reliability of Gemini 2.5 Pro, ChatGPT 4.1, DeepSeek V3, and Claude Opus 4 in generating standardized CMR protocols. Eur. Radiol. Exp. 2026, 10(1), 7. [Google Scholar] [CrossRef] [PubMed]
- Jeblick, K.; Schachtner, B.; Dexl, J.; Mittermeier, A.; Stüber, A.T.; Topalis, J.; et al. ChatGPT makes medicine easy to swallow: an exploratory case study on simplified radiology reports. Eur. Radiol. 2024, 34(5), 2817–2825. [Google Scholar] [CrossRef] [PubMed]
- Skryd, A.; Lawrence, K. ChatGPT as a tool for medical education and clinical decision-making on the wards: case study. JMIR Form. Res. 2024, 8, e51346. [Google Scholar] [CrossRef] [PubMed]
- Hirosawa, T.; Kawamura, R.; Harada, Y.; Mizuta, K.; Tokumasu, K.; Kaji, Y.; et al. ChatGPT-generated differential diagnosis lists for complex case-derived clinical vignettes: diagnostic accuracy evaluation. JMIR Med. Inf. 2023, 11, e48808. [Google Scholar] [CrossRef] [PubMed]
- Magalhães Araujo, S.; Cruz-Correia, R. Incorporating ChatGPT in medical informatics education: mixed methods study on student perceptions and experiential integration proposals. JMIR Med. Educ. 2024, 10, e51151. [Google Scholar] [CrossRef] [PubMed]
- Wang, C.; Liu, S.; Yang, H.; Guo, J.; Wu, Y.; Liu, J. Ethical considerations of using ChatGPT in health care. J. Med. Internet Res. 2023, 25, e48009. [Google Scholar] [CrossRef] [PubMed]
- Zhang, J.; Zhang, Z.M. Ethics and governance of trustworthy medical artificial intelligence. BMC Med. Inf. Decis. Mak. 2023, 23(1), 7. [Google Scholar] [CrossRef] [PubMed]
| Case Category | n | Description | Key Features |
| Normal perfusion | 6 | No perfusion defects | Symmetric uptake, normal LVEF |
| Single-vessel ischemia (LAD) | 6 | Reversible defect in LAD territory | Anterior/anteroseptal reversibility |
| Single-vessel ischemia (RCA/LCx) | 4 | Reversible defect in RCA or LCx territory | Inferior/inferolateral or lateral reversibility |
| Multivessel ischemia | 4 | Reversible defects in ≥2 territories | Widespread reversible perfusion abnormalities |
| Old myocardial infarction | 4 | Fixed perfusion defects | Fixed defects with wall motion abnormalities |
| Attenuation/technical artifact | 4 | Breast, diaphragmatic, or motion artifact | Basal inferior or anterior hypoperfusion with artifact suspicion |
| Mixed/complex pathology | 2 | Combination of ischemia + infarct or artifact + ischemia | Multiple competing findings |
| Evaluation Criterion | Description |
| Clinical Accuracy | Accuracy in interpreting findings and disease characterization |
| Terminological Appropriateness | Use of relevant and meaningful medical terminology |
| Report Structure | Presentation of the report in a professional and organized format |
| Overall Comprehensibility | Clarity, simplicity, and logical flow of language |
| Evaluation Criteria | ChatGPT-4 Median (IQR) | Physician Median (IQR) | p-value | Effect Size (r) |
| Clinical Accuracy | 4 (3-4) | 5 (4-5) | 0.002 | 0.52 |
| Terminological Appropriateness | 4 (4-5) | 5 (4-5) | 0.18 | 0.22 |
| Report Structure | 5 (4-5) | 5 (5-5) | 0.11 | 0.28 |
| Overall Comprehensibility | 5 (4-5) | 5 (4-5) | 0.42 | 0.14 |
| Case Category | n | Clinical Accuracy Median (IQR) | Physician Median (IQR) |
| Normal perfusion | 6 | 5 (4-5) | 5 (5-5) |
| Single-vessel ischemia (LAD) | 6 | 4 (4-5) | 5 (5-5) |
| Single-vessel ischemia (RCA/LCx) | 4 | 4 (4-5) | 5 (5-5) |
| Multivessel ischemia | 4 | 3 (3-4) | 5 (4-5) |
| Old myocardial infarction | 4 | 4 (3-4) | 5 (5-5) |
| Attenuation/technical artifact | 4 | 3 (3-4) | 5 (4-5) |
| Mixed/complex pathology | 2 | 3 (3-3) | 4 (4-5) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).