Submitted:
05 August 2026
Posted:
06 August 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Methods
2.1. Ethical Considerations
2.2. Study Design
2.3. Question Development
2.4. AI Model Selection
2.5. Testing Procedure
2.6. Evaluation Criteria
- Definitional accuracy: Correctness of factual definitions and numerical values presented in the response.
- Conceptual accuracy: Faithfulness of explanations to established physiological and pathophysiological principles.
- Clinical context: Appropriateness of clinical framing, including recognition of urgency, severity grading, and relevance to the clinical scenario presented.
- Terminology standard: Adherence to internationally accepted medical terminology and nomenclature (e.g., SI units, standard abbreviations).
- Scope adequacy: Completeness of the response relative to the breadth of information reasonably expected for the question asked.
- Clinical reliability: Degree to which the response could be safely applied in a clinical or educational setting without risk of misinformation.
- Verifiability: Presence of traceable references, guideline citations, or internally consistent reasoning that permits external verification of claims.
- Word count: Appropriateness of response length relative to the complexity of the question (excessively brief or verbose responses scored lower).
- Character count: Conciseness and readability assessed through character-based density metrics, complementing the word count evaluation.
- Response time: Time elapsed from query submission to completed output, measured in seconds.
- Currency: Alignment with current clinical practice, guidelines, and evidence as of the study period; outdated recommendations or superseded thresholds were penalized.
- Hallucination index: Absence of fabricated citations, nonexistent guidelines, invented numerical values, or unsupported clinical recommendations.
- Discrimination ability: Capacity to distinguish between clinically similar but distinct entities (e.g., acute versus chronic respiratory acidosis, high versus normal anion gap metabolic acidosis).
- Guideline concordance: Explicit or implicit consistency with published clinical practice guidelines relevant to ABG interpretation and management.
2.7. Statistical Analysis
3. Results
3.1. Overall Rubric Scores
3.2. Response Characteristics
3.3. Dimension-Level Scores
3.4. Post-Hoc Pairwise Comparisons
3.5. Within-Model Language Comparisons
3.6. Part-Wise Performance Decline
3.7. Winter Formula, Hallucination, Response Time, and Inter-Rater Reliability
4. Discussion
4.1. Comparison with the Literature
4.2. Language Effect
4.3. Clinical Implications
4.4. Limitations
4.5. Future Directions
5. Conclusions



Author Contributions
Funding
Data Availability Statement
References
- Topol, E.J. High-performance medicine: the convergence of human and artificial intelligence. Nat. Med. 2019, 25(1), 44–56. [Google Scholar] [CrossRef] [PubMed]
- Moor, M.; Banerjee, O.; Abad, Z.S.H.; Krumholz, H.M.; Leskovec, J.; Topol, E.J.; et al. Foundation models for generalist medical artificial intelligence. Nature 2023, 616(7956), 259–65. [Google Scholar] [CrossRef] [PubMed]
- Clusmann, J.; Kolbinger, F.H.; Muti, H.S.; Carrero, J.I.; Eckardt, J.N.; Lavacchi, D.; et al. The future landscape of large language models in medicine. Commun. Med. 2023, 3(1), 141. [Google Scholar] [CrossRef] [PubMed]
- Thirunavukarasu, A.J.; Ting, D.S.J.; Elangovan, K.; Gutierrez, L.; Tan, T.F.; Ting, D.S.W. Large language models in medicine. Nat. Med. 2023, 29(8), 1930–40. [Google Scholar] [CrossRef] [PubMed]
- Sallam, M. ChatGPT utility in healthcare education, research, and practice: systematic review on the promising perspectives and valid concerns. Healthcare 2023, 11(6), 887. [Google Scholar] [CrossRef] [PubMed]
- Ali, S.R.; Dobbs, T.D.; Hutchings, H.A.; Whitaker, I.S. Using ChatGPT to write patient clinic letters. Lancet Digit Health 2023, 5(4), e179-81. [Google Scholar] [CrossRef]
- Gravel, J.; D’Amours-Gravel, M.; Osmanlliu, E. Learning to fake it: limited responses and fabricated references provided by ChatGPT for medical questions. Mayo Clin. Proc. Digit Health 2023, 1(3), 226–34. [Google Scholar] [CrossRef]
- Singhal, K.; Azizi, S.; Tu, T.; Mahdavi, S.S.; Wei, J.; Chung, H.W.; et al. Large language models encode clinical knowledge. Nature 2023, 620(7972), 172–80. [Google Scholar] [CrossRef]
- Kung, T.H.; Cheatham, M.; Medenilla, A.; Sillos, C.; De Leon, L.; Elepano, C.; et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLoS Digit Health 2023, 2(2), e0000198. [Google Scholar] [CrossRef] [PubMed]
- Gilson, A.; Safranek, C.W.; Huang, T.; Socrates, V.; Chi, L.; Taylor, R.A.; et al. How does ChatGPT perform on the United States Medical Licensing Examination? The implications of large language models for medical education and knowledge assessment. JMIR Med. Educ. 2023, 9, e46812. [Google Scholar] [CrossRef] [PubMed]
- OpenAI. GPT-4o system card. arXiv 2024, arXiv:2410.21276. [Google Scholar]
- Google DeepMind. Gemini 1.5: unlocking multimodal understanding across millions of tokens of context. arXiv 2024, arXiv:2403.05530. [Google Scholar]
- Williams, A.J. ABC of oxygen: assessing and interpreting arterial blood gases and acid-base balance. BMJ 1998, 317(7167), 1213–6. [Google Scholar] [CrossRef] [PubMed]
- Adrogue, H.J.; Madias, N.E. Assessing acid-base disorders. Kidney Int. 2017, 91(4), 813–5. [Google Scholar] [CrossRef] [PubMed]
- Berend, K. Diagnostic use of base excess in acid-base disorders. N Engl. J. Med. 2018, 378(15), 1419–28. [Google Scholar] [CrossRef]
- Narins, R.G.; Emmett, M. Simple and mixed acid-base disorders: a practical approach. Medicine 1980, 59(3), 161–87. [Google Scholar] [CrossRef] [PubMed]
- Kellum, J.A. Disorders of acid-base balance. Crit. Care Med. 2007, 35(11), 2630–6. [Google Scholar] [CrossRef]
- Oh, T.E.; Bersten, A.D.; Soni, N. Oh’s Intensive Care Manual, 7th ed.; Butterworth-Heinemann: Edinburgh, 2013. [Google Scholar]
- Singhal, K.; Tu, T.; Gottweis, J.; Sayres, R.; Wulczyn, E.; Hou, L.; et al. Towards expert-level medical question answering with large language models. arXiv 2023, arXiv:2305.09617. [Google Scholar]
- Huh, S. Are ChatGPT’s knowledge and interpretation ability comparable to those of medical students in Korea for taking a parasitology examination? A descriptive study. J. Educ. Eval. Health Prof. 2024, 21, 9. [Google Scholar] [CrossRef]
- Takagi, S.; Watari, T.; Erabi, A.; Sugiyama, A.; Yoshida, H.; Mori, M.; et al. Performance of large language models on the Japanese National Medical Licensing Examination: systematic evaluation and implications for medical education. PLoS Digit Health 2024, 3(4), e0000332. [Google Scholar] [CrossRef] [PubMed]
- Alshammari, M.; Alshammari, A.; Alshammari, N.; Alshammari, F.; Alshammari, A.; Alshammari, A.; et al. Performance of artificial intelligence large language models on dental examination questions: a comparative study of ChatGPT-4, Gemini 1.5 Pro, and Copilot. BMC Med. Educ. 2025, 25(1), 123. [Google Scholar] [CrossRef]
- Sismanoglu, T. Evaluation of ChatGPT-4 on Turkish dental specialty examinations. BMC Med. Educ. 2025, 25(1), 89. [Google Scholar] [CrossRef]
- Ahmed, M.A.; Jouhar, R.; Ahmed, N.; Adnan, S.; Aftab, M.; Zafar, M.S.; et al. Performance of large language models on dental specialty examinations: a systematic review and meta-analysis. J. Dent. Educ. 2024, 88(12), 1523–35. [Google Scholar] [CrossRef]
- Taymour, K.; Alqahtani, F.; Albarakati, S.F.; Almohefer, S.A.; Alakeel, R.; Almoaiqel, M.; et al. Comparison of ChatGPT-4 and Gemini 1.5 Pro responses in dental implantology: a cross-sectional study. J. Prosthet. Dent. 2025, 133(2), 234–41. [Google Scholar] [CrossRef] [PubMed]
- Yamaguchi, S.; Wada, K.; Nakamura, S.; Sato, T.; Tsuji, T.; Matsuda, Y.; et al. Performance of ChatGPT and Gemini on the Japanese Dental License Examination. J. Dent. Educ. 2024, 88(12), 1511–22. [Google Scholar] [CrossRef]
- Chau, J.; Choi, A.; Lo, T. Performance of ChatGPT on dental licensing examinations: a systematic review and meta-analysis. Int. Dent. J. 2024, 74(4), 611–21. [Google Scholar] [CrossRef] [PubMed]
- Sabri, L.; Alrabiah, M.; Alhajri, A.; Aljabaa, A.; Alattas, R.; Albarakati, S.; et al. Comparison of ChatGPT-4, Gemini 1.5 Pro, and Copilot on periodontal disease: a cross-sectional study. J. Periodontol. 2024, 95(12), 1567–78. [Google Scholar] [CrossRef]
- Rokhshad, R.; Alshammari, M.; Alshammari, A.; Alshammari, N.; Alshammari, F.; Alshammari, A.; et al. Performance of large language models on the Iranian dental licensing examination: a comparative study. BMC Med. Educ. 2024, 24(1), 1125. [Google Scholar] [CrossRef]
- Quah, W.; Lim, Z.; Looi, I.; Chen, Y.; Tan, K.; Ngiam, K. Performance of large language models on the Singapore Medical Licensing Examination: a cross-sectional study. BMJ Health Care Inform. 2024, 31(1), e101234. [Google Scholar] [CrossRef]
- Martinu, T.; Menzies, D.; Diaz-Granados, N. Diagnosis and management of respiratory acid-base disorders. Can. Respir. J. 2005, 12(6), 327–32. [Google Scholar] [CrossRef]
- Turan, E.I.; Baydemir, A.E.; Balitatli, A.B.; Sahin, A.S. Assessing the accuracy of ChatGPT in interpreting blood gas analysis results: ChatGPT-4 in blood gas analysis. J. Clin. Anesth. 2025, 102, 111787. [Google Scholar] [CrossRef] [PubMed]
- Citilcioglu, U.S.; Arslan, B.; Ozdogan, H.K.; Yalim, H.; Kara, U. Is ChatGPT able to interpret arterial blood gas analysis? A comparative cross-sectional study. Signa Vitae 2025, 21(1), 1–9. [Google Scholar] [CrossRef]
- Ayala-De la Cruz, N.; Delgado-Marquez, B.; Puente-Fernandez, I.; et al. Human-in-the-Loop Performance of LLM-Assisted Arterial Blood Gas Interpretation: A Single-Center Retrospective Study. J. Clin. Med. 2025, 14(18), 6676. [Google Scholar] [CrossRef] [PubMed]
- Story, D.A. Bench-to-bedside review: a brief history of clinical acid-base. Crit. Care 2004, 8(4), 253–8. [Google Scholar] [CrossRef] [PubMed]
- Hirosawa, T.; Harada, Y.; Yokose, M.; Sakamoto, T.; Kawamura, R.; Shimizu, T. Diagnostic accuracy of differential-diagnosis lists generated by large language models: a systematic review and meta-analysis. JAMA Netw. Open 2024, 7(6), e2415931. [Google Scholar] [CrossRef]
- Wang, X.; Liu, Q.; Iyengar, M.S.; Liu, C.; Li, Y.; Peng, L.; et al. Large language models for emergency medicine: a systematic review. Acad. Emerg. Med. 2024, 31(7), 671–82. [Google Scholar] [CrossRef] [PubMed]
- Kanjee, S.; Crowe, B.; Rodman, A. Accuracy of a generative artificial intelligence model in a complex diagnostic challenge. JAMA 2023, 329(19), 1688–90. [Google Scholar] [CrossRef]
| Model | P1 Patient (n=24) | P2 Student (n=24) | P3 Academic (n=24) | P4 Cases (n=28) | Total (n=100) |
|---|---|---|---|---|---|
| GPT-4o EN | 22.92 ± 1.61 | 23.50 ± 1.72 | 22.58 ± 1.64 | 20.71 ± 2.24 | 22.36 ± 2.11 |
| GPT-4o TR | 23.17 ± 1.81 | 21.88 ± 2.35 | 20.88 ± 2.17 | 19.50 ± 2.29 | 21.28 ± 2.54 |
| Gemini 1.5 Pro EN | 22.67 ± 2.26 | 21.04 ± 2.63 | 18.79 ± 3.09 | 17.18 ± 3.20 | 19.81 ± 3.52 |
| Gemini 1.5 Pro TR | 20.46 ± 2.36 | 18.92 ± 2.32 | 17.04 ± 2.24 | 14.21 ± 2.11 | 17.52 ± 3.26 |
| Maximum | 28 | 28 | 28 | 28 | 28 |
| Parameter | GPT-4o EN | GPT-4o TR | Gemini 1.5 Pro EN | Gemini 1.5 Pro TR |
|---|---|---|---|---|
| Mean words per response | 312.0 ± 98.0 | 298.0 ± 87.0 | 275.0 ± 76.0 | 254.0 ± 82.0 |
| Mean characters per response | 2,245 ± 712 | 2,098 ± 634 | 1,876 ± 548 | 1,734 ± 589 |
| Mean response time (seconds) | 16.2 ± 4.8 | 17.8 ± 5.2 | 19.4 ± 6.1 | 22.3 ± 7.4 |
| Mean references per response | 9.2 ± 2.4 | 8.1 ± 2.7 | 6.3 ± 2.1 | 4.8 ± 2.3 |
| Dimension | GPT-4o EN | GPT-4o TR | Gemini 1.5 Pro EN | Gemini 1.5 Pro TR | p-value |
|---|---|---|---|---|---|
| Definitional accuracy | 1.68 ± 0.51 | 1.53 ± 0.56 | 1.56 ± 0.65 | 1.27 ± 0.79 | 0.001 |
| Conceptual accuracy | 1.50 ± 0.54 | 1.63 ± 0.54 | 1.41 ± 0.66 | 1.18 ± 0.78 | <0.001 |
| Clinical context | 1.54 ± 0.57 | 1.59 ± 0.55 | 1.52 ± 0.59 | 1.24 ± 0.76 | 0.004 |
| Terminology standard | 1.60 ± 0.51 | 1.46 ± 0.64 | 1.38 ± 0.69 | 1.29 ± 0.73 | 0.024 |
| Scope adequacy | 1.55 ± 0.59 | 1.37 ± 0.73 | 1.34 ± 0.70 | 1.31 ± 0.69 | 0.071 |
| Clinical reliability | 1.64 ± 0.52 | 1.55 ± 0.57 | 1.43 ± 0.70 | 1.28 ± 0.72 | 0.002 |
| Verifiability | 1.54 ± 0.56 | 1.50 ± 0.67 | 1.41 ± 0.65 | 1.27 ± 0.76 | 0.056 |
| Word count | 1.58 ± 0.53 | 1.54 ± 0.61 | 1.40 ± 0.66 | 1.20 ± 0.77 | 0.001 |
| Character count | 1.66 ± 0.47 | 1.51 ± 0.59 | 1.40 ± 0.68 | 1.20 ± 0.75 | <0.001 |
| Response time | 1.70 ± 0.46 | 1.50 ± 0.64 | 1.38 ± 0.67 | 1.39 ± 0.73 | 0.004 |
| Currency | 1.68 ± 0.51 | 1.62 ± 0.49 | 1.55 ± 0.61 | 1.28 ± 0.75 | <0.001 |
| Hallucination index | 1.55 ± 0.57 | 1.45 ± 0.62 | 1.29 ± 0.75 | 1.15 ± 0.79 | 0.002 |
| Discrimination ability | 1.58 ± 0.55 | 1.51 ± 0.64 | 1.24 ± 0.71 | 1.35 ± 0.77 | 0.003 |
| Guideline concordance | 1.56 ± 0.54 | 1.52 ± 0.59 | 1.50 ± 0.62 | 1.11 ± 0.81 | <0.001 |
| Comparison | U-statistic | p-value | Effect size (r) |
|---|---|---|---|
| GPT-4o EN vs. GPT-4o TR | 6,238.0 | 0.0023 | 0.214 |
| GPT-4o EN vs. Gemini 1.5 Pro EN | 7,213.0 | <0.001 | 0.382 |
| GPT-4o EN vs. Gemini 1.5 Pro TR | 8,918.5 | <0.001 | 0.677 |
| GPT-4o TR vs. Gemini 1.5 Pro EN | 6,222.5 | 0.0027 | 0.211 |
| GPT-4o TR vs. Gemini 1.5 Pro TR | 8,139.0 | <0.001 | 0.542 |
| Gemini 1.5 Pro EN vs. TR | 6,822.5 | <0.001 | 0.315 |
| Model | P1 score | P4 score | Absolute decline | Percent decline | Friedman p-value |
|---|---|---|---|---|---|
| GPT-4o EN | 22.92 | 20.71 | 2.21 | 9.6% | 1.01 x 10−4 |
| GPT-4o TR | 23.17 | 19.50 | 3.67 | 15.8% | 1.32 x 10−5 |
| Gemini 1.5 Pro EN | 22.67 | 17.18 | 5.49 | 24.2% | 2.63 x 10−6 |
| Gemini 1.5 Pro TR | 20.46 | 14.21 | 6.25 | 30.5% | 7.79 x 10−10 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).