Submitted:
16 December 2024
Posted:
18 December 2024
You are already at the latest version
Abstract
Background/Objectives: The role of artificial intelligence (AI) in radiological image analysis is rapidly evolving. This study evaluates the diagnostic performance of Chat Generative Pre-trained Transformer (ChatGPT) in detecting intracranial hemorrhages (ICH) on non-contrast computed tomography (NCCT) images, along with its ability to classify hemorrhage type, stage, anatomical location, and associated findings. Methods: A retrospective study was conducted using 240 cases, comprising 120 ICH cases and 120 controls with normal findings. Five consecutive NCCT slices per case were selected by radiologists and analyzed by ChatGPT-4o using a standardized prompt with nine questions. Diagnostic accuracy, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) were calculated by comparing the model’s results with radiologists’ assessments (gold standard). After a two-week interval, the same dataset was re-evaluated to assess intra-observer reliability and consistency. Results: ChatGPT-4o achieved 100% accuracy in identifying imaging modality type. For ICH detection, the model demonstrated 68.3% diagnostic accuracy, sensitivity of 79.2%, specificity of 57.5%, PPV of 65.1%, and NPV of 73.4%. It correctly classified 34.0% of hemorrhage types and 7.3% of localizations. All ICH-positive cases were identified as acute phase (100%). In the second evaluation, diagnostic accuracy improved to 73.3%, with sensitivity of 86.7% and specificity of 60%. The Cohen’s Kappa coefficient for intraobserver agreement in ICH detection indicated moderate agreement (κ=0.469). Conclusions: ChatGPT-4o shows promise in identifying imaging modalities and ICH presence but demonstrates limitations in localization and hemorrhage type classification. These findings highlight its potential for improvement through targeted training for medical applications.
Keywords:
1. Introduction
2. Materials and Methods
2.1. Study Design and Case Selection
2.2. Inclusion and Exclusion Criteria
2.3. Brain NCCT Images
2.4. Image selection and evaluation
2.5. Prompt selection and testing
2.6. ChatGPT interaction and prompting
2.7. Executors and readers
2.8. Statistical Analysis
3. Results
4. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Van Asch, C.J.; Luitse, M.J.; Rinkel, G.J.; van der Tweel, I.; Algra, A.; Klijn, C.J. Incidence, case fatality, and functional outcome of intracerebral haemorrhage over time, according to age, sex, and ethnic origin: a systematic review and meta-analysis. Lancet Neurol. 2010, 9, 167–176. [Google Scholar] [CrossRef]
- Heit, J.J.; Iv, M.; Wintermark, M. Imaging of intracranial hemorrhage. J. Stroke 2016, 19, 11. [Google Scholar] [CrossRef]
- Ginat, D.T. Analysis of head CT scans flagged by deep learning software for acute intracranial hemorrhage. Neuroradiology 2020, 62, 335–340. [Google Scholar] [CrossRef] [PubMed]
- Kidwell, C.S. , Chalela J.A., Saver J.L., Starkman S, Hill M.D., Demchuk A.M., et al. Comparison of MRI and CT for detection of acute intracerebral hemorrhage. JAMA 2004, 292, 1823–1830. [Google Scholar] [CrossRef] [PubMed]
- Zahuranec, D.B. , Lisabeth L.D., Sánchez B.N., Smith M.A., Brown D.L., Garcia N.M., et al. Intracerebral hemorrhage mortality is not changing despite declining incidence. Neurology 2014, 82, 2180–2186. [Google Scholar] [CrossRef] [PubMed]
- Elliott, J.; Smith, M. The acute management of intracerebral hemorrhage: a clinical review. Anesth. Analg. 2010, 110, 1419–1427. [Google Scholar] [CrossRef]
- Fugate, J.E.; Rabinstein, A.A. Absolute and relative contraindications to IV rt-PA for acute ischemic stroke. Neurohospitalist 2015, 5, 110–121. [Google Scholar] [CrossRef]
- Qureshi, A.I.; Mendelow, A.D.; Hanley, D.F. Intracerebral haemorrhage. Lancet 2009, 373, 1632–1644. [Google Scholar] [CrossRef] [PubMed]
- Morotti, A.; Goldstein, J.N. Diagnosis and management of acute intracerebral hemorrhage. Emerg. Med. Clin. North. Am. 2016, 34, 883–899. [Google Scholar] [CrossRef]
- McDonald, R.J. , Schwartz K.M., Eckel L.J., Diehn F.E., Hunt C.H., Bartholmai B.J., et al. The effects of changes in utilization and technological advancements of cross-sectional imaging on radiologist workload. Acad. Radiol. 2015, 22, 1191–1198. [Google Scholar] [CrossRef]
- Sellers, A.; Hillman, B.J.; Wintermark, M. Survey of after-hours coverage of emergency department imaging studies by US academic radiology departments. J. Am. Coll. Radiol. 2014, 11, 725–730. [Google Scholar] [CrossRef]
- Spitler K, Vijayasarathi A, Salehi B, Dua S, Azizyan A, Cekic M, et al. Neuroradiologist coverage improves resident perception of educational experience, referring physician satisfaction, and turnaround time. Curr. Probl. Diagn. Radiol. 2020, 49, 168–172. [Google Scholar] [CrossRef]
- Hosny, A.; Parmar, C.; Quackenbush, J.; Schwartz, L.H.; Aerts, H.J.W.L. Artificial intelligence in radiology. Nat. Rev. Cancer 2018, 18, 500–510. [Google Scholar] [CrossRef]
- Almeida, L.C.; Farina, E.M.J.M.; Kuriki, P.E.A.; Abdala, N.; Kitamura, F.C. Performance of ChatGPT on the Brazilian radiology and diagnostic imaging and mammography board examinations. Radiol. Artif. Intell. 2024, 6, e230103. [Google Scholar] [CrossRef] [PubMed]
- Lee, J.Y.; Kim, J.S.; Kim, T.Y.; Kim, Y.S. Detection and classification of intracranial haemorrhage on CT images using a novel deep-learning algorithm. Sci. Rep. 2020, 10, 20546. [Google Scholar] [CrossRef]
- Yun, T.J. , Choi J.W., Han M, Jung W.S., Choi S.H., Yoo R.E., et al. Deep learning based automatic detection algorithm for acute intracranial haemorrhage: a pivotal randomized clinical trial. NPJ Digit. Med. 2023, 6, 61. [Google Scholar] [CrossRef]
- Dawud, A.M.; Yurtkan, K.; Oztoprak, H. Application of deep learning in neuroradiology: brain haemorrhage classification using transfer learning. Comput. Intell. Neurosci. 2019, 2019, 4629859. [Google Scholar] [CrossRef] [PubMed]
- Kaluarachchi, T.; Reis, A.; Nanayakkara, S. A review of recent deep learning approaches in human-centered machine learning. Sensors 2021, 21, 2514. [Google Scholar] [CrossRef]
- Lewick, T.; Kumar, M.; Hong, R.; Wu, W. Intracranial hemorrhage detection in ct scans using deep learning. Proc - 2020 IEEE 6th Int Conf Big Data Comput Serv Appl BigDataService 2020. 2020 Aug 1;169–72.
- Heit, J.J. , Coelho H, Lima F.O., Granja M, Aghaebrahim A, Hanel R, et al. Automated cerebral hemorrhage detection using RAPID. Am. J. Neuroradiol. 2021, 42, 273–278. [Google Scholar] [CrossRef]
- Kuo, W.; Hӓne, C.; Mukherjee, P.; Malik, J.; Yuh, E.L. Expert-level detection of acute intracranial hemorrhage on head computed tomography using deep learning. Proc. Natl. Acad. Sci. USA 2019, 116, 22737–22745. [Google Scholar] [CrossRef] [PubMed]
- Voter, A.F.; Meram, E.; Garrett, J.W.; Yu, J.P.J. Diagnostic accuracy and failure mode analysis of a deep learning algorithm for the detection of intracranial hemorrhage. J. Am. Coll. Radiol. 2021, 18, 1143–1152. [Google Scholar] [CrossRef]
- Rahsepar, A.A.; Tavakoli, N.; Kim, G.H.J.; Hassani, C.; Abtin, F.; Bedayat, A. How AI responds to common lung cancer questions: ChatGPT versus Google Bard. Radiology 2023, 307, e230922. [Google Scholar] [CrossRef] [PubMed]
- Temperley, H.C. , O’Sullivan N.J., Mac Curtain B.M., Corr A, Meaney J.F., Kelly M.E., et al. Current applications and future potential of ChatGPT in radiology: a systematic review. J. Med. Imaging Radiat. Oncol. 2024, 68, 257–264. [Google Scholar] [CrossRef] [PubMed]
- Haver, H.L. , Bahl M, Doo F.X., Kamel P.I., Parekh V.S., Jeudy J, et al. Evaluation of multimodal ChatGPT (GPT-4V) in describing mammography image features. Can. Assoc. Radiol. J. 2024, 75, 947–949. [Google Scholar] [CrossRef] [PubMed]
- Mert S, Stoerzer P, Brauer J, Fuchs B, Haas-Lützenberger E. M., Demmer W, et al. Diagnostic power of ChatGPT 4 in distal radius fracture detection through wrist radiographs. Arch. Orthop. Trauma. Surg. 2024, 144, 2461–2467. [Google Scholar] [CrossRef]
- Dehdab R, Brendlin A, Werner S, Almansour H, Gassenmaier S, Brendel J. M., et al. Evaluating ChatGPT-4V in chest CT diagnostics: a critical image interpretation assessment. Jpn. J. Radiol. 2024, 42, 1168–1177. [Google Scholar] [CrossRef] [PubMed]
- Mongan, J.; Moy, L.; Kahn, C.E. Checklist for artificial intelligence in medical imaging (CLAIM): a guide for authors and Reviewers. Radiol. Artif. Intell. 2020, 2, e200029. [Google Scholar] [CrossRef] [PubMed]
- Image inputs for ChatGPT - FAQ | OpenAI Help Center [Internet]. Available from: https://help.openai.com/en/articles/8400551-image-inputs-for-chatgpt-faq.
- Arbabshirani, M.R. , Fornwalt B.K., Mongelluzzo G.J., Suever J.D., Geise B.D., Patel A.A., et al. Advanced machine learning in action: identification of intracranial hemorrhage on computed tomography scans of the head with clinical workflow integration. NPJ Digit. Med, 2018, 1, 9. [Google Scholar] [CrossRef] [PubMed]
- Ojeda, P.; Zawaideh, M.; Mossa-Basha, M.; Haynor, D. The utility of deep learning: evaluation of a convolutional neural network for detection of intracranial bleeds on non-contrast head computed tomography studies. Proc. SPIE, 2019; 10949, 899–906. [Google Scholar] [CrossRef]
- Dave, T.; Athaluri, S.A.; Singh, S. ChatGPT in medicine: an overview of its applications, advantages, limitations, future prospects, and ethical considerations. Front. Artif. Intell, 2023, 6, 1169595. [Google Scholar] [CrossRef] [PubMed]
- Kundisch A, Hönning A, Mutze S, Kreissl L, Spohn F, Lemcke J, et al. Deep learning algorithm in detecting intracranial hemorrhages on emergency computed tomographies. PLoS One, 2021, 16, e0260560. [Google Scholar]
- Morales, H. Pitfalls in the Imaging Interpretation of Intracranial Hemorrhage. Semin. Ultrasound CT MR 2018, 39, 457–468. [Google Scholar] [CrossRef]



| ICH Group (n=120) |
Healthy Control Group (n=120) | P value | |
|---|---|---|---|
| Gender *, n (%) | |||
| Female | 39 (32.5) | 40 (33.3) | 0.891 |
| Male | 81 (67.5) | 80 (66.7) | |
| Age**, years, Mean ± SD | 63.95 ± 19.58 | 63.73 ± 18.81 | 0.877 |
| n (%) | |
|---|---|
| Hemorrhage Type (Hemorrhagic areas, n=150) | |
| Subdural | 78 (52) |
| Epidural | 12 (8) |
| Intraventricular | 7 (4.7) |
| Intraparenchymal | 35 (23.3) |
| Subarachnoid | 18 (12) |
| Hemorrhage Location (Cases, n=120) | |
| Unifocal | 92 (76.7) |
| Multifocal | 28 (23.3) |
| Associated Pathologies (n=85) | |
| Cerebral edema | 44 (51.8) |
| Midline shift | 23 (27) |
| Right lateral ventricle compression | 8 (9.4) |
| Left lateral ventricle compression | 10 (11.8) |
| Hemorrhage Stage |
Subdural n (%) |
Epidural n (%) |
Intraventricular n (%) | Intraparenchymal n (%) |
Subaraknoid n (%) |
Total n (%) |
|---|---|---|---|---|---|---|
| Acute | 38 (25.3) | 10 (6.7) | 7 (4.7) |
32 (21.3) | 18 (12) | 105 (70) |
| Subacute | 11 (7.3) |
1 (0.7) | 0 (0) |
1 (0.7) |
0 (0) |
13 (8.7) |
| Chronic |
24 (16) |
1 (0.7) |
0 (0) |
2 (1.3) |
0 (0) |
27 (18) |
| Acute-Subacute |
2 (1.3) | 0 (0) |
0 (0) |
0 (0) |
0 (0) |
2 (1.3) |
| Acute-Chronic | 3 (2) | 0 (0) | 0 (0) | 0 (0) | 0 (0) | 3 (2) |
| Total | 78 (52) | 12 (8) | 7 (4.7) | 35 (23.3) | 18 (12) | 150 (100) |
| True Positive, n (%) |
False Positive, n (%) |
True Negative, n (%) |
False Negative, n (%) |
Sensitivity, % | Specificity, % | PPD, % | NPD, % | Diagnostic Accuracy, % |
|
|---|---|---|---|---|---|---|---|---|---|
| ChatGPT-4o,Round 1 | 95 (79.2) |
51 (42.5) | 69 (57.5) |
25 (20.8) |
79.2 | 57.5 | 65.1 | 73.4 | 68.3 |
| ChatGPT-4o,Round 2 | 104 (86.7) | 48 (40) | 72 (60) |
16 (13.3) |
86.7 | 60.0 | 68.4 | 81.8 | 73.3 |
| Hemorrhage Presence Assessment | ||||||||
|---|---|---|---|---|---|---|---|---|
| ICH group (n=120) | Healthy Control Group (n=120) | |||||||
| ChatGPT-4o, Round 1 | ChatGPT-4o, Round 1 | |||||||
| Negative | Positive | Negative | Positive | |||||
| ChatGPT- 4o, Round 2 |
Negative | 9 | 7 | ChatGPT-4o, Round 2 | Negative | 52 | 20 | |
| Positive | 16 | 88 | Positive | 17 | 31 | |||
| Hemorrhage Type (n=150) |
Correct, n (%) |
Partially Correct, n (%) | Incorrect, n (%) |
False Negative Cases, n(%) | Total, n (%) |
|---|---|---|---|---|---|
| Subdural | 21 (14) | 7 (4.7) | 26 (17.3) | 24 (16) | 78 (52) |
| Epidural | 1 (0.7) | 0 (0) | 10 (6.6) | 1 (0.7) | 12 (8) |
| Intraventricular | 3 (2) | 0 (0) | 4 (2.7) | 0 (0) | 7 (4.7) |
| Intraparenchymal | 25 (16.6) | 0 (0) | 9 (6) | 1 (0.7) | 35 (23.3) |
| Subarachnoid | 1 (0.7) | 0 (0) | 15 (10) | 2 (1.3) | 18 (12) |
| Total | 51 (34) | 7 (4.7) | 64 (42.6) | 28 (18.7) | 150 (100) |
|
Associated Pathologies |
True Positive, n (%) |
False Negative, n (%) |
|---|---|---|
| Midline shift, n=23 | 11 (47.8) | 12 (52.2) |
| Cerebral edema, n=44 | 29 (65.9) | 15 (34.1) |
| Right lateral ventricle compression, n=8 | 0 (0) | 8 (100) |
| Left lateral ventricle compression, n=10 | 0 (0) | 10 (100) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).