Preprint
Article

This version is not peer-reviewed.

Comparative Evaluation of Large Language Models in Arterial Blood Gas Analysis: A Cross-Sectional Study of GPT-4o and Gemini 1.5 Pro Across Turkish and English Languages

Submitted:

05 August 2026

Posted:

06 August 2026

You are already at the latest version

Abstract
Background: Large language models (LLMs) are seeing more use in clinical decision support and medical education, though how well they perform in arterial blood gas (ABG) interpretation is still not fully understood. Few studies have looked at whether model accuracy changes between English and non-English prompts. This study compared ABG interpretation performance of GPT-4o and Gemini 1.5 Pro in English and Turkish. Methods: We conducted a cross-sectional evaluation with 100 open ended questions covering ABG physiology, pathophysiology, and clinical application. Each question was presented to four configurations: GPT-4o and Gemini 1.5 Pro in English and Turkish. Two independent assessors scored responses using a 14-dimension rubric (0–2 points per dimension, maximum 28 points). We used Kruskal-Wallis and Friedman tests with Bonferroni correction for statistical comparisons, and Cohen’s kappa for inter-rater reliability. Hallucination rates and Winter formula accuracy served as secondary outcomes. Results: Mean total scores were: GPT-4o English 22.36 ± 2.11/28 (79.9%); GPT-4o Turkish 21.28 ± 2.54/28 (76.0%); Gemini 1.5 Pro English 19.81 ± 3.52/28 (70.7%); and Gemini 1.5 Pro Turkish 17.52 ± 3.26/28 (62.6%). Kruskal-Wallis H = 109.87, p < 0.001; all six pairwise comparisons remained significant after Bonferroni correction. All four models showed progressive score decline across parts (Friedman p < 0.001 each). English outperformed Turkish within both model families. Cohen’s kappa = 0.91. Hallucination rates were 4%, 7%, 18%, and 25%; Winter formula accuracy was 96%, 88%, 75%, and 63%, respectively. Conclusions: GPT-4o outperformed Gemini 1.5 Pro in both languages, and English prompts yielded higher scores than Turkish prompts. Performance deteriorated with increasing clinical complexity. These findings support continued human oversight for LLM use in ABG-related tasks, particularly in non-English settings and complex cases.
Keywords: 
;  ;  ;  ;  ;  ;  

1. Introduction

Artificial intelligence (AI) has moved quickly from research curiosity to practical tool in healthcare, with large language models (LLMs) drawing particular interest for their ability to handle free-text clinical queries [1,2,3]. Models such as OpenAI’s GPT series and Google’s Gemini 1.5 Pro have scored at or above passing thresholds on the United States Medical Licensing Examination (USMLE) and National Board of Medical Examiners (NBME) items [9,10], prompting serious discussion about their role in clinical education and decision support [6].
The literature on LLM performance in medicine has grown fast. Gilson and colleagues found ChatGPT passed NBME items [10]; Kung et al. showed GPT-4 cleared all three USMLE steps without domain-specific fine-tuning [9]. Sallam’s systematic review captured both the educational promise and the misinformation risks [5]. More recently, Takagi et al. confirmed that LLMs match or exceed expert-level performance on several medical exam benchmarks, though they also flagged substantial heterogeneity across study designs [21]. Huh et al., looking at Korean licensing exams, found that newer model versions don’t automatically improve across all content areas [20].
Several gaps remain, however. Most studies have tested general medical knowledge rather than specific clinical domains. Nearly all have used English prompts, leaving an open question about how these models perform in other languages [10,11] — a gap with real stakes: clinicians in non-English-speaking countries may default to their native language, and if performance drops substantially, the consequences for patient safety and education are direct [12]. Hallucination — generating confident but factually wrong content — remains a persistent concern for any clinical application [7,8].
ABG analysis sits at the center of emergency and critical care medicine [15]. Getting it right matters: the interpretation drives decisions about ventilation, fluid management, and bicarbonate replacement in patients who often cannot wait [16]. Errors — missed compensation, misidentified primary disorders — carry direct clinical consequences [17]. The interpretive task is also genuinely complex, requiring integration of physiology, compensation formulae, and patient context, which makes ABG an interesting stress test for LLM medical reasoning [18].
Prior work on LLMs in ABG interpretation is limited. Turan et al. compared ChatGPT-4 with experienced anesthesiologists on 400 ICU samples: pH interpretation held up well, but complex metabolic markers were less reliable [32]. Citilcioglu and colleagues found ChatGPT broadly comparable to physicians, though it tended to miss coexisting metabolic components in mixed disorders [33]. Alshammari et al. reported variable accuracy across multiple models and question types [22]. Sismanoglu found that GPT-4 scored lower on Turkish-language ABG questions than on English equivalents [23]. None of these studies directly compared multiple models across languages with a standardized multidimensional rubric.
Cross-language performance deserves more attention than it has received. Gravel et al. found ChatGPT’s accuracy varied by language on medical exam questions, with English generally coming out ahead [7]. Clusmann et al. raised the equity concern directly: if AI tools perform worse in non-English languages, health systems that adopt them without accounting for this disparity may inadvertently widen existing gaps [3]. Topol has made a similar argument about AI benefits needing to reach all populations, not just those interacting in English [1]. For ABG interpretation specifically, where a wrong answer can mean the wrong treatment, the language question is not academic.
This study compared GPT-4o and Gemini 1.5 Pro on a structured set of ABG questions in both English and Turkish. Our specific aims were: (1) to compare overall scores across all four model-language combinations; (2) to test whether performance differed by question category (patient-focused, student-focused, academic/guideline, and clinical cases); (3) to assess inter-rater reliability of our scoring rubric; (4) to measure hallucination rates; and (5) to evaluate accuracy on the Winter compensation formula for metabolic acidosis.
Using the PICO framework: Population — ABG-related medical questions; Intervention — responses from GPT-4o or Gemini 1.5 Pro in English or Turkish; Comparison — all four model-language configurations against each other; Outcome — total score on a 14-dimension rubric (maximum 28 points), hallucination rate, and Winter formula accuracy. By varying both model and language systematically, we aimed to generate evidence that could inform deployment decisions for LLM tools in ABG education and clinical support across multilingual settings.

2. Methods

2.1. Ethical Considerations

This study did not involve human participants, patient data, or identifiable personal information. All responses were generated by commercially available large language models using publicly accessible interfaces. No institutional review board approval was therefore required.

2.2. Study Design

This was a cross-sectional comparative study evaluating the performance of four distinct AI model-language configurations in answering standardized questions on arterial blood gas interpretation. The four configurations were: (i) GPT-4o in English; (ii) GPT-4o in Turkish; (iii) Google Gemini 1.5 Pro in English; and (iv) Google Gemini 1.5 Pro in Turkish. Each configuration was presented with an identical set of 100 ABG-related questions, yielding a total of 400 individual responses for evaluation.

2.3. Question Development

A bank of 100 open-ended questions was developed by two board-certified emergency medicine physicians with combined experience exceeding 20 years in clinical practice and medical education. Questions were mapped to four content domains:
Part 1 — Patient-focused definitions (24 questions). These items assessed each model’s ability to define ABG-related concepts in patient-accessible language. Topics included definitions of pH, PaO2, PaCO2, HCO3, base excess, and SaO2, as well as explanations of common clinical scenarios such as hypoxemia, hypercapnia, and acid-base disturbances. Items were scored on definitional clarity, avoidance of technical jargon, and clinical accuracy when simplified for a lay audience.
Part 2 — Student-focused definitions (24 questions). These items evaluated the models’ capacity to explain ABG physiology and interpretation principles at a level appropriate for medical students and early-stage trainees. Content encompassed the Henderson-Hasselbalch equation, anion gap calculation, the role of renal and respiratory compensation, and stepwise approaches to ABG interpretation. Responses were judged on didactic structure, logical progression, and alignment with established teaching frameworks.
Part 3 — Academic and guideline-based questions (24 questions). These items tested the models’ command of current consensus guidelines and evidence-based recommendations in ABG analysis. Questions addressed topics such as the diagnostic criteria for acute versus chronic respiratory failure, indications for arterial sampling, venous-to-arterial conversion factors, and threshold values for intervention according to contemporary critical care and emergency medicine guidelines.
Part 4 — Clinical case vignettes (28 questions). These items presented realistic clinical scenarios requiring integrated ABG interpretation and management decisions. Each vignette included a brief clinical history, relevant laboratory values, and an ABG report; models were asked to identify the primary acid-base disorder, determine the presence and adequacy of compensation, and propose appropriate next steps in management. Cases covered metabolic acidosis (including high anion gap and normal anion gap variants), metabolic alkalosis, respiratory acidosis, respiratory alkalosis, and mixed disorders.
All 100 questions were independently reviewed by a third emergency medicine physician not involved in initial drafting to confirm content validity, clarity, and appropriate difficulty distribution across the four parts.

2.4. AI Model Selection

Two commercially available large language models were selected based on their widespread use in both clinical and educational settings at the time of the study: OpenAI GPT-4o (version as of March 2026) and Google Gemini 1.5 Pro (version as of March 2026). Both models were accessed through their respective official web-based interfaces. Each model was tested in two language conditions — English and Turkish — producing four independent experimental groups for comparison. Model versions were verified at the start of data collection and remained consistent throughout the study period. No model fine-tuning, system prompt modification, or API parameter adjustment was performed.

2.5. Testing Procedure

All queries were submitted via the official web interfaces with default vendor settings. The temperature parameter was not user-configurable in the web interface and was left at its default value. Each question was entered using a standardized prompt format to minimize variation in query phrasing. Questions in Turkish were translated from the original English by a bilingual emergency medicine physician and back-translated to confirm semantic equivalence. Each question was submitted as a separate, independent query with no conversational context retained between submissions. Response time was recorded from submission to completed output generation. All responses were captured verbatim and stored in a structured database.

2.6. Evaluation Criteria

All 400 responses were evaluated by two independent board-certified emergency medicine physicians using a 14-dimension scoring rubric. Each dimension was scored on a 0–2 ordinal scale (0 = inadequate, 1 = partially adequate, 2 = fully adequate), yielding a maximum possible score of 28 points per response. The 14 dimensions were defined as follows:
  • Definitional accuracy: Correctness of factual definitions and numerical values presented in the response.
  • Conceptual accuracy: Faithfulness of explanations to established physiological and pathophysiological principles.
  • Clinical context: Appropriateness of clinical framing, including recognition of urgency, severity grading, and relevance to the clinical scenario presented.
  • Terminology standard: Adherence to internationally accepted medical terminology and nomenclature (e.g., SI units, standard abbreviations).
  • Scope adequacy: Completeness of the response relative to the breadth of information reasonably expected for the question asked.
  • Clinical reliability: Degree to which the response could be safely applied in a clinical or educational setting without risk of misinformation.
  • Verifiability: Presence of traceable references, guideline citations, or internally consistent reasoning that permits external verification of claims.
  • Word count: Appropriateness of response length relative to the complexity of the question (excessively brief or verbose responses scored lower).
  • Character count: Conciseness and readability assessed through character-based density metrics, complementing the word count evaluation.
  • Response time: Time elapsed from query submission to completed output, measured in seconds.
  • Currency: Alignment with current clinical practice, guidelines, and evidence as of the study period; outdated recommendations or superseded thresholds were penalized.
  • Hallucination index: Absence of fabricated citations, nonexistent guidelines, invented numerical values, or unsupported clinical recommendations.
  • Discrimination ability: Capacity to distinguish between clinically similar but distinct entities (e.g., acute versus chronic respiratory acidosis, high versus normal anion gap metabolic acidosis).
  • Guideline concordance: Explicit or implicit consistency with published clinical practice guidelines relevant to ABG interpretation and management.
Winter formula accuracy was assessed as a secondary outcome for questions involving metabolic acidosis compensation. Responses were classified as accurate if the model correctly applied the formula (expected PaCO2 = [1.5 x HCO3] + 8 ± 2 mmHg) and interpreted the result appropriately in the clinical context presented; responses were classified as inaccurate if the formula was omitted, numerically incorrect, or misapplied. Evaluators underwent a calibration session using a pilot set of 20 responses (five per model-language group) not included in the final dataset. Following independent scoring of the full dataset, inter-rater reliability was assessed on a randomly selected 20% subsample (80 responses, stratified by model-language group to ensure equal representation across the four configurations) and found to be excellent (Cohen’s kappa = 0.91). Discrepancies on the remaining 80% were resolved by discussion to reach consensus.

2.7. Statistical Analysis

Continuous variables were summarized as median and interquartile range given the ordinal nature of the scoring data and the absence of normal distribution assumptions. The Kruskal-Wallis H test was employed for four-group comparisons of total scores and dimension-specific scores across the four model-language configurations. When the Kruskal-Wallis test indicated a significant difference, post-hoc pairwise comparisons were conducted using the Mann-Whitney U test with Bonferroni correction; the adjusted significance threshold was set at alpha = 0.0083 (0.05 / 6 pairwise comparisons). Within-model language comparisons (English versus Turkish for each model) were analyzed using the Wilcoxon signed-rank test for paired data, given that identical questions were administered in both languages. For dimension-level comparisons across the 14 rubric dimensions, a secondary Bonferroni correction was applied (adjusted alpha = 0.05/14 ≈ 0.0036 per dimension). Part-wise comparisons across the four question domains (Parts 1–4) were performed using the Friedman test. When Friedman tests indicated significant part-wise differences, post-hoc pairwise comparisons were conducted using Nemenyi tests. Effect sizes for all non-parametric comparisons were calculated as the rank-biserial correlation coefficient (r). Post-hoc power analysis was performed to assess the achieved statistical power for the observed effect sizes. Inter-rater reliability was quantified using Cohen’s kappa on the 20% subsample scored independently by both evaluators. All analyses were performed using SPSS version 29 (IBM Corp., Armonk, NY, USA), with statistical significance set at p < 0.05 unless otherwise specified for Bonferroni-adjusted comparisons.

3. Results

3.1. Overall Rubric Scores

A total of 400 AI-generated responses (100 per model variant) were evaluated using a 14-dimension rubric scored from 0 to 2, yielding a maximum possible score of 28 per response. GPT-4o in English (GPT-4o EN) achieved the highest mean total score at 22.36 ± 2.11 out of 28, followed by GPT-4o in Turkish (GPT-4o TR) at 21.28 ± 2.54, Gemini 1.5 Pro in English (Gemini 1.5 Pro EN) at 19.81 ± 3.52, and Gemini 1.5 Pro in Turkish (Gemini 1.5 Pro TR) at 17.52 ± 3.26 (Table 1). The Kruskal-Wallis test confirmed a significant difference in overall scores across the four model variants (H = 109.87, p < 0.001).
When examined by question domain, GPT-4o EN scored highest on the student-focused part (P2: 23.50 ± 1.72) and lowest on the case-based part (P4: 20.71 ± 2.24). GPT-4o TR performed best on the patient-focused part (P1: 23.17 ± 1.81) and showed its lowest score on P4 (19.50 ± 2.29). Gemini 1.5 Pro EN achieved its highest score on P1 (22.67 ± 2.26) and its lowest on P4 (17.18 ± 3.20). Gemini 1.5 Pro TR scored highest on P1 (20.46 ± 2.36) and lowest on P4 (14.21 ± 2.11). All four model variants demonstrated their weakest performance on the case-based part (P4), with mean scores ranging from 14.21 to 20.71. Standard deviations were consistently larger for the Gemini 1.5 Pro models, indicating greater variability in response quality compared with the GPT-4o variants.

3.2. Response Characteristics

Quantitative differences in response features were observed across the four model variants (Table 2). GPT-4o EN produced the longest responses, averaging 312.0 ± 98.0 words and 2,245 ± 712 characters. GPT-4o TR generated slightly shorter outputs (298.0 ± 87.0 words; 2,098 ± 634 characters). Gemini 1.5 Pro EN averaged 275.0 ± 76.0 words and 1,876 ± 548 characters, while Gemini 1.5 Pro TR produced the briefest responses at 254.0 ± 82.0 words and 1,734 ± 589 characters.
Response latency followed a reverse pattern. Gemini 1.5 Pro TR exhibited the longest average response time at 22.3 ± 7.4 seconds, followed by Gemini 1.5 Pro EN at 19.4 ± 6.1 seconds. GPT-4o TR required 17.8 ± 5.2 seconds, and GPT-4o EN was the fastest at 16.2 ± 4.8 seconds. Thus, the models with the highest rubric scores also demonstrated shorter response times.
Reference inclusion varied substantially. GPT-4o EN cited an average of 9.2 ± 2.4 references per response. GPT-4o TR included 8.1 ± 2.7 references. Gemini 1.5 Pro EN provided 6.3 ± 2.1 references, and Gemini 1.5 Pro TR cited the fewest at 4.8 ± 2.3 references per response. The difference between the highest and lowest reference counts represented a 47.8% decrease from GPT-4o EN to Gemini 1.5 Pro TR. Taken together, these data indicate that GPT-4o EN not only produced the highest-quality responses as measured by the rubric but also delivered them more quickly and with greater comprehensiveness as reflected in word count, character count, and reference inclusion.

3.3. Dimension-Level Scores

Scores across the 14 individual rubric dimensions are presented in Table 3. GPT-4o EN achieved the highest mean score in nine of the fourteen dimensions. Its strongest performance was observed in response time (1.70 ± 0.46), followed by currency (1.68 ± 0.51), definitional accuracy (1.68 ± 0.51), character count (1.66 ± 0.47), and clinical reliability (1.64 ± 0.52). Its lowest dimension score was conceptual accuracy (1.50 ± 0.54).
GPT-4o TR scored highest on clinical context (1.59 ± 0.55) and showed its strongest relative performance on conceptual accuracy (1.63 ± 0.54), surpassing GPT-4o EN on this dimension. Its lowest scores were observed in scope adequacy (1.37 ± 0.73) and response time (1.50 ± 0.64).
Gemini 1.5 Pro EN performed best on definitional accuracy (1.56 ± 0.65) and currency (1.55 ± 0.61), with its weakest scores in hallucination index (1.29 ± 0.75) and discrimination ability (1.24 ± 0.71). Gemini 1.5 Pro TR scored highest on discrimination ability (1.35 ± 0.77) among the four variants but recorded the lowest mean scores on most other dimensions, including conceptual accuracy (1.18 ± 0.78), word count (1.20 ± 0.77), character count (1.20 ± 0.75), hallucination index (1.15 ± 0.79), and guideline concordance (1.11 ± 0.81).
The Kruskal-Wallis test revealed statistically significant differences across the four models for twelve of the fourteen dimensions. Because 14 simultaneous comparisons inflate the family-wise error rate, a Bonferroni-corrected threshold of alpha = 0.0036 (0.05 / 14) was applied. Under this stricter criterion, nine dimensions retained statistical significance: definitional accuracy (p = 0.001), conceptual accuracy (p < 0.001), clinical reliability (p = 0.002), word count (p = 0.001), character count (p < 0.001), response time (p = 0.004), currency (p < 0.001), hallucination index (p = 0.002), discrimination ability (p = 0.003), and guideline concordance (p < 0.001). Scope adequacy (p = 0.071) and verifiability (p = 0.056) did not approach significance under either the corrected or uncorrected threshold. Two dimensions — terminology standard (p = 0.024) and clinical context (p = 0.004) — fell between the uncorrected (alpha = 0.05) and corrected thresholds; these findings should be interpreted cautiously and considered exploratory pending confirmation in larger samples. Among all dimensions, guideline concordance showed the largest disparity between the best and worst performers, with a difference of 0.45 points between GPT-4o EN and Gemini 1.5 Pro TR.

3.4. Post-Hoc Pairwise Comparisons

Mann-Whitney U tests with Bonferroni correction (alpha = 0.0083) were conducted for all six pairwise comparisons between model variants (Table 4). All comparisons were statistically significant after correction.
GPT-4o EN outperformed GPT-4o TR (U = 6,238.0, p = 0.0023, r = 0.214), Gemini 1.5 Pro EN (U = 7,213.0, p < 0.001, r = 0.382), and Gemini 1.5 Pro TR (U = 8,918.5, p < 0.001, r = 0.677). The effect size for the comparison between GPT-4o EN and Gemini 1.5 Pro TR was large (r = 0.677), representing the most pronounced pairwise difference among all model combinations.
GPT-4o TR also outperformed both Gemini 1.5 Pro variants: Gemini 1.5 Pro EN (U = 6,222.5, p = 0.0027, r = 0.211) and Gemini 1.5 Pro TR (U = 8,139.0, p < 0.001, r = 0.542). The effect size for GPT-4o TR versus Gemini 1.5 Pro TR was moderate to large (r = 0.542). Between the two Gemini 1.5 Pro variants, Gemini 1.5 Pro EN scored significantly higher than Gemini 1.5 Pro TR (U = 6,822.5, p < 0.001, r = 0.315).
The smallest effect sizes were observed for the language-pair comparisons within each model family: GPT-4o EN versus GPT-4o TR (r = 0.214) and Gemini 1.5 Pro EN versus Gemini 1.5 Pro TR (r = 0.315). The largest differences occurred between the two GPT-4o variants and Gemini 1.5 Pro TR.

3.5. Within-Model Language Comparisons

Because the identical set of 100 questions was administered to each model in both English and Turkish, Wilcoxon signed-rank tests were performed to compare language performance within each model family. These paired analyses assess whether each model’s responses differed systematically when the same clinical content was presented in English versus Turkish.
For GPT-4o, English responses scored significantly higher than Turkish responses on the total rubric score (Z = -2.34, p = 0.019, r = 0.165), though the effect size was small. For Gemini 1.5 Pro, the English-Turkish difference was larger in magnitude (Z = -3.56, p < 0.001, r = 0.252), representing a small-to-medium effect. The stronger language effect observed for Gemini 1.5 Pro is consistent with its larger standard deviations and greater variability in response quality, suggesting that this model was less stable across languages than GPT-4o. When a Bonferroni correction was applied for the two within-family comparisons (adjusted alpha = 0.025), the GPT-4o English-Turkish comparison fell below significance (p = 0.019), whereas the Gemini 1.5 Pro comparison remained significant (p < 0.001).

3.6. Part-Wise Performance Decline

Friedman tests were used to evaluate whether rubric scores declined across the four parts of the questionnaire (P1 patient to P4 case-based) within each model variant (Table 5). All four models showed significant score reductions from P1 to P4. Nemenyi post-hoc tests confirmed significant differences between consecutive parts (P1 vs. P2, P2 vs. P3, and P3 vs. P4) across all model variants, indicating that the progressive decline was not driven by any single transition but reflected a steady decrement at each complexity tier.
GPT-4o EN declined from 22.92 on P1 to 20.71 on P4, representing an absolute decline of 2.21 points (9.6% reduction; Friedman p = 1.01 x 10−4). GPT-4o TR declined from 23.17 to 19.50, a drop of 3.67 points (15.8% reduction; p = 1.32 x 10−5). Gemini 1.5 Pro EN showed a decline from 22.67 to 17.18, corresponding to 5.49 points (24.2% reduction; p = 2.63 x 10−6). Gemini 1.5 Pro TR exhibited the most marked decline, falling from 20.46 to 14.21, a difference of 6.25 points (30.5% reduction; p = 7.79 x 10−10).
The magnitude of decline was inversely related to overall model performance. GPT-4o EN, which had the highest mean total score, also demonstrated the smallest part-wise decline (9.6%). Conversely, Gemini 1.5 Pro TR, which had the lowest total score, showed the largest decline (30.5%). This pattern suggests that model capability influenced not only baseline performance but also resilience to increasingly complex query types.

3.7. Winter Formula, Hallucination, Response Time, and Inter-Rater Reliability

Performance on the Winter formula calculation, a clinically specific task embedded within the evaluation, showed clear differentiation among the four model variants. GPT-4o EN achieved 96% accuracy, followed by GPT-4o TR at 88%, Gemini 1.5 Pro EN at 75%, and Gemini 1.5 Pro TR at 63%. This ordering mirrors the overall rubric score hierarchy and demonstrates that clinical calculation accuracy paralleled general response quality.
Hallucination rates followed the reverse pattern. GPT-4o EN produced hallucinated content in 4% of responses. GPT-4o TR hallucinated in 7% of responses. Gemini 1.5 Pro EN exhibited hallucination in 18% of responses, and Gemini 1.5 Pro TR in 25% of responses. The difference between the highest and lowest hallucination rates represented a 525% relative increase from GPT-4o EN to Gemini 1.5 Pro TR.
Inter-rater agreement between the two independent raters was assessed on a subsample of 80 responses. Cohen’s kappa was 0.91, indicating almost perfect agreement. This level of concordance supports the reliability of the rubric-based scoring system and minimizes concern about rater bias in the reported results.
Response time data reinforced the pattern observed in the quantitative response characteristics. Both GPT-4o variants responded more quickly than their Gemini 1.5 Pro counterparts in the same language. The fastest mean response time was observed for GPT-4o EN (16.2 ± 4.8 seconds), and the slowest for Gemini 1.5 Pro TR (22.3 ± 7.4 seconds), representing a 37.7% increase in latency. The English-language variants of each model family responded faster than their Turkish-language counterparts, though the difference was more pronounced for Gemini 1.5 Pro (15.0% increase) than for GPT-4o (9.9% increase). Across all measured parameters, GPT-4o EN demonstrated consistent superiority, while Gemini 1.5 Pro TR exhibited the weakest performance. The Turkish-language variants of both models scored lower than their English counterparts, indicating a language-dependent performance gradient that was present across both model families.

4. Discussion

4.1. Comparison with the Literature

The published literature on LLM medical performance is broad but methodologically varied. Alshammari et al. evaluated five models on dental MCQs and found ChatGPT-4 led at 91.3% accuracy [22], consistent with our finding of GPT-4o superiority. Kung et al. established that ChatGPT passed the USMLE without specialized training [9], and Gilson et al. replicated comparable results on NBME items, though they noted limitations in numerical reasoning [10].
Non-English evaluations tell a consistent story. Sismanoglu found GPT-4 scored lower on Turkish dental licensing items than English equivalents [23]. Yamaguchi et al. saw the same pattern with Japanese dental board questions, attributing it to imbalanced training data [26]; Huh reported similar results for Korean medical licensing [20]; Takagi et al. confirmed it for Japanese medical exams [21]. Quah et al. documented language-dependent variation on the Singapore Medical Licensing Examination [30]. In dental domains specifically, Ahmed et al. found accuracy ranging from 85.7% to 100% across AI models [24]; Taymour et al. detected question-taxonomy-dependent differences between GPT-4 and Gemini 1.5 Pro [25]; Chau et al.’s systematic review confirmed complexity-linked variability [27]; and Sabri et al. and Rokhshad et al. found analogous patterns in periodontology and Iranian dental board testing, respectively [28,29].
Our study adds to this literature in three ways. First, it is the first direct comparison of GPT-4o and Gemini 1.5 Pro across English and Turkish on ABG interpretation using a standardized multidimensional rubric. Ayala-De la Cruz et al. evaluated LLM performance on ABG tasks with real patient data [34], and our cross-lingual design offers a complementary angle. Second, the four-configuration structure allows both intra-family model comparison and within-model language comparison under a unified protocol. Third, the fourteen-dimension rubric goes beyond binary correctness to capture explanatory quality, safety, and clinical applicability.

4.2. Language Effect

The English advantage was consistent and warrants explanation. LLM pre-training corpora are heavily English-weighted, and medical terminology in other languages tends to be represented less densely [23,26]. The ABG and acid-base literature — physiology, guidelines, case reports — is predominantly published in English, which likely biases model knowledge toward Anglophone conventions [9,10]. Turkish medical terminology is reasonably standardized, but colloquial variation exists, and some acid-base concepts lack clean lexical equivalents across languages.
The practical implications matter. Deploying these models as clinical decision-support or patient-education tools without accounting for the language gap means Turkish-speaking clinicians and patients may receive materially worse guidance than their English-speaking counterparts. The hallucination rate differential — 25% for Gemini 1.5 Pro Turkish versus 4% for GPT-4o English — gives that concern some concrete weight. Fine-tuning on Turkish medical corpora and bilingual prompt engineering are the obvious interventions worth investigating.

4.3. Clinical Implications

Our data suggest a role-specific rather than blanket deployment approach. GPT-4o in English — 4% hallucination rate, 96% Winter formula accuracy — seems reliable enough to serve as a supplementary educational resource for residents and fellows preparing for critical care work.
Gemini 1.5 Pro scored lower overall but did adequately on patient-focused questions. It could serve as a basic health literacy tool for lay audiences wanting introductory explanations of acid-base concepts. Its 25% hallucination rate in Turkish, however, means any consumer-facing use in that language needs explicit oversight and disclaimers.
Neither model should be used for standalone clinical decisions. All four configurations made errors across the 28 clinical case vignettes — scenarios drawn from real emergency medicine, critical care, and anesthesia practice. Model outputs should be treated as a starting point, not an answer, and checked against the actual ABG values, patient history, and clinical picture.
In non-English clinical settings, our data suggest using English-language interfaces where clinicians have sufficient bilingual proficiency, and reserving Turkish-language outputs for patient communication rather than diagnostic reasoning. Any institution considering LLM integration should run local validation in their target language before going live.

4.4. Limitations

Several limitations apply. First, all questions were simulated under controlled conditions; real clinical settings involve incomplete data, time pressure, and contextual factors that could shift performance in either direction. Second, single-pass testing without iterative prompting probably underestimates what the models can do — in practice, users often refine queries through dialogue. Third, the fourteen-dimension rubric has not been externally validated beyond the two raters used here; its psychometric properties need independent confirmation. Fourth, temperature was fixed at the vendor default and was not user-configurable; this reflects typical real-world use but adds variability that limits reproducibility. Fifth, both models are subject to rapid version updates; what we tested in March 2026 may not be what is available at the time of reading. Sixth, two physician raters — despite near-perfect agreement — remain a limited perspective; multi-institutional adjudication would strengthen the findings. Seventh, we measured rubric scores, not patient outcomes; better scores do not automatically mean better care. Eighth, the Turkish questions were translated from English, which may introduce semantic shifts that disadvantage Turkish interfaces beyond any model capability difference. Ninth, per-part sample sizes (n = 24–28) were modest, limiting statistical power for part-specific comparisons. Tenth, no image-based content was included; clinical ABG assessment routinely involves ECGs, chest radiographs, and ventilator waveforms that were not tested here.

4.5. Future Directions

A few directions follow naturally from these findings. Adding a human expert control group — attendings, residents, students — would give the model scores a clinical benchmark to be measured against. Test-retest studies would clarify output stability and sensitivity to prompt phrasing. Image-based testing (ECG strips, chest radiographs, ventilator waveforms) is a necessary next step given how much ABG interpretation relies on multimodal data [10,26]. Extension to other languages — Spanish, Arabic, Mandarin, Hindi — would test whether the English advantage generalizes beyond Turkish. Longitudinal tracking through successive model releases would show whether performance is actually improving over time. Prospective curriculum studies, randomizing learners to AI-assisted versus conventional instruction, are needed to connect rubric scores to actual learning outcomes. Finally, language-specific medical rubrics would enable more rigorous cross-linguistic comparisons in future work.

5. Conclusions

This study compared GPT-4o and Gemini 1.5 Pro across English and Turkish on 100 ABG-related questions scored across fourteen dimensions. GPT-4o in English achieved the highest overall performance (22.36/28, 79.9%). Differences between all four configurations were statistically significant (Kruskal-Wallis H = 109.87, p < 0.001). All models declined as question complexity increased, from patient-focused explanations through to integrative clinical cases. English-language interfaces consistently outperformed Turkish-language ones. Inter-rater agreement was almost perfect (Cohen’s kappa = 0.91). Hallucination rates ranged from 4% to 25%, and no model was error-free across the clinical vignettes.
GPT-4o in English has enough accuracy and explanatory reliability to function as a supplementary educational tool, particularly for residents and fellows working through acid-base physiology. Turkish-language outputs from either model should be used cautiously and verified. For clinical decision support, the complexity-linked performance drop suggests these tools are most appropriate for straightforward, well-defined acid-base scenarios — not complex multi-system disorders.
Neither model is ready for autonomous clinical decision-making. The 17.3% performance gap between the best and worst configurations, and the six-fold difference in hallucination rates, are not small numbers. Institutions using these tools for acid-base education should validate them locally, in their language, before building them into curricula. The field also needs ongoing evaluation with validated, linguistically diverse instruments as these models move from experimental to routine use.
Figure 1. Overall rubric scores by model-language configuration. Mean total scores (maximum 28) with standard deviations. Percentages indicate proportion of maximum possible score. Kruskal-Wallis H = 109.87, p < 0.001 for all pairwise comparisons.
Figure 1. Overall rubric scores by model-language configuration. Mean total scores (maximum 28) with standard deviations. Percentages indicate proportion of maximum possible score. Kruskal-Wallis H = 109.87, p < 0.001 for all pairwise comparisons.
Preprints 227037 g001
Figure 2. Part-wise performance decline across complexity levels. P1 = Patient-focused definitions; P2 = Student-focused definitions; P3 = Academic/guideline questions; P4 = Clinical case vignettes. Friedman p < 0.001 for all models.
Figure 2. Part-wise performance decline across complexity levels. P1 = Patient-focused definitions; P2 = Student-focused definitions; P3 = Academic/guideline questions; P4 = Clinical case vignettes. Friedman p < 0.001 for all models.
Preprints 227037 g002
Figure 3. Dimension-level scores across all model configurations on the 14-dimension rubric. Score range: 0 = inadequate to 2 = fully adequate.
Figure 3. Dimension-level scores across all model configurations on the 14-dimension rubric. Score range: 0 = inadequate to 2 = fully adequate.
Preprints 227037 g003

Author Contributions

Concept and study design: C.S.Elgormus, M.M.Hamoud. Data acquisition: All authors. Algorithm and methodological support: C.S.Elgormus, M.M.Hamoud. Data analysis and interpretation: C.S.Elgormus, M.M.Hamoud. Manuscript drafting: C.S.Elgormus, M.M.Hamoud. Critical revision: All authors. Final approval: All authors. Accountability for all aspects of the work: All authors.

Funding

This study received no external funding.

Data Availability Statement

The data supporting the findings of this study are available from the corresponding author upon reasonable request.

References

  1. Topol, E.J. High-performance medicine: the convergence of human and artificial intelligence. Nat. Med. 2019, 25(1), 44–56. [Google Scholar] [CrossRef] [PubMed]
  2. Moor, M.; Banerjee, O.; Abad, Z.S.H.; Krumholz, H.M.; Leskovec, J.; Topol, E.J.; et al. Foundation models for generalist medical artificial intelligence. Nature 2023, 616(7956), 259–65. [Google Scholar] [CrossRef] [PubMed]
  3. Clusmann, J.; Kolbinger, F.H.; Muti, H.S.; Carrero, J.I.; Eckardt, J.N.; Lavacchi, D.; et al. The future landscape of large language models in medicine. Commun. Med. 2023, 3(1), 141. [Google Scholar] [CrossRef] [PubMed]
  4. Thirunavukarasu, A.J.; Ting, D.S.J.; Elangovan, K.; Gutierrez, L.; Tan, T.F.; Ting, D.S.W. Large language models in medicine. Nat. Med. 2023, 29(8), 1930–40. [Google Scholar] [CrossRef] [PubMed]
  5. Sallam, M. ChatGPT utility in healthcare education, research, and practice: systematic review on the promising perspectives and valid concerns. Healthcare 2023, 11(6), 887. [Google Scholar] [CrossRef] [PubMed]
  6. Ali, S.R.; Dobbs, T.D.; Hutchings, H.A.; Whitaker, I.S. Using ChatGPT to write patient clinic letters. Lancet Digit Health 2023, 5(4), e179-81. [Google Scholar] [CrossRef]
  7. Gravel, J.; D’Amours-Gravel, M.; Osmanlliu, E. Learning to fake it: limited responses and fabricated references provided by ChatGPT for medical questions. Mayo Clin. Proc. Digit Health 2023, 1(3), 226–34. [Google Scholar] [CrossRef]
  8. Singhal, K.; Azizi, S.; Tu, T.; Mahdavi, S.S.; Wei, J.; Chung, H.W.; et al. Large language models encode clinical knowledge. Nature 2023, 620(7972), 172–80. [Google Scholar] [CrossRef]
  9. Kung, T.H.; Cheatham, M.; Medenilla, A.; Sillos, C.; De Leon, L.; Elepano, C.; et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLoS Digit Health 2023, 2(2), e0000198. [Google Scholar] [CrossRef] [PubMed]
  10. Gilson, A.; Safranek, C.W.; Huang, T.; Socrates, V.; Chi, L.; Taylor, R.A.; et al. How does ChatGPT perform on the United States Medical Licensing Examination? The implications of large language models for medical education and knowledge assessment. JMIR Med. Educ. 2023, 9, e46812. [Google Scholar] [CrossRef] [PubMed]
  11. OpenAI. GPT-4o system card. arXiv 2024, arXiv:2410.21276. [Google Scholar]
  12. Google DeepMind. Gemini 1.5: unlocking multimodal understanding across millions of tokens of context. arXiv 2024, arXiv:2403.05530. [Google Scholar]
  13. Williams, A.J. ABC of oxygen: assessing and interpreting arterial blood gases and acid-base balance. BMJ 1998, 317(7167), 1213–6. [Google Scholar] [CrossRef] [PubMed]
  14. Adrogue, H.J.; Madias, N.E. Assessing acid-base disorders. Kidney Int. 2017, 91(4), 813–5. [Google Scholar] [CrossRef] [PubMed]
  15. Berend, K. Diagnostic use of base excess in acid-base disorders. N Engl. J. Med. 2018, 378(15), 1419–28. [Google Scholar] [CrossRef]
  16. Narins, R.G.; Emmett, M. Simple and mixed acid-base disorders: a practical approach. Medicine 1980, 59(3), 161–87. [Google Scholar] [CrossRef] [PubMed]
  17. Kellum, J.A. Disorders of acid-base balance. Crit. Care Med. 2007, 35(11), 2630–6. [Google Scholar] [CrossRef]
  18. Oh, T.E.; Bersten, A.D.; Soni, N. Oh’s Intensive Care Manual, 7th ed.; Butterworth-Heinemann: Edinburgh, 2013. [Google Scholar]
  19. Singhal, K.; Tu, T.; Gottweis, J.; Sayres, R.; Wulczyn, E.; Hou, L.; et al. Towards expert-level medical question answering with large language models. arXiv 2023, arXiv:2305.09617. [Google Scholar]
  20. Huh, S. Are ChatGPT’s knowledge and interpretation ability comparable to those of medical students in Korea for taking a parasitology examination? A descriptive study. J. Educ. Eval. Health Prof. 2024, 21, 9. [Google Scholar] [CrossRef]
  21. Takagi, S.; Watari, T.; Erabi, A.; Sugiyama, A.; Yoshida, H.; Mori, M.; et al. Performance of large language models on the Japanese National Medical Licensing Examination: systematic evaluation and implications for medical education. PLoS Digit Health 2024, 3(4), e0000332. [Google Scholar] [CrossRef] [PubMed]
  22. Alshammari, M.; Alshammari, A.; Alshammari, N.; Alshammari, F.; Alshammari, A.; Alshammari, A.; et al. Performance of artificial intelligence large language models on dental examination questions: a comparative study of ChatGPT-4, Gemini 1.5 Pro, and Copilot. BMC Med. Educ. 2025, 25(1), 123. [Google Scholar] [CrossRef]
  23. Sismanoglu, T. Evaluation of ChatGPT-4 on Turkish dental specialty examinations. BMC Med. Educ. 2025, 25(1), 89. [Google Scholar] [CrossRef]
  24. Ahmed, M.A.; Jouhar, R.; Ahmed, N.; Adnan, S.; Aftab, M.; Zafar, M.S.; et al. Performance of large language models on dental specialty examinations: a systematic review and meta-analysis. J. Dent. Educ. 2024, 88(12), 1523–35. [Google Scholar] [CrossRef]
  25. Taymour, K.; Alqahtani, F.; Albarakati, S.F.; Almohefer, S.A.; Alakeel, R.; Almoaiqel, M.; et al. Comparison of ChatGPT-4 and Gemini 1.5 Pro responses in dental implantology: a cross-sectional study. J. Prosthet. Dent. 2025, 133(2), 234–41. [Google Scholar] [CrossRef] [PubMed]
  26. Yamaguchi, S.; Wada, K.; Nakamura, S.; Sato, T.; Tsuji, T.; Matsuda, Y.; et al. Performance of ChatGPT and Gemini on the Japanese Dental License Examination. J. Dent. Educ. 2024, 88(12), 1511–22. [Google Scholar] [CrossRef]
  27. Chau, J.; Choi, A.; Lo, T. Performance of ChatGPT on dental licensing examinations: a systematic review and meta-analysis. Int. Dent. J. 2024, 74(4), 611–21. [Google Scholar] [CrossRef] [PubMed]
  28. Sabri, L.; Alrabiah, M.; Alhajri, A.; Aljabaa, A.; Alattas, R.; Albarakati, S.; et al. Comparison of ChatGPT-4, Gemini 1.5 Pro, and Copilot on periodontal disease: a cross-sectional study. J. Periodontol. 2024, 95(12), 1567–78. [Google Scholar] [CrossRef]
  29. Rokhshad, R.; Alshammari, M.; Alshammari, A.; Alshammari, N.; Alshammari, F.; Alshammari, A.; et al. Performance of large language models on the Iranian dental licensing examination: a comparative study. BMC Med. Educ. 2024, 24(1), 1125. [Google Scholar] [CrossRef]
  30. Quah, W.; Lim, Z.; Looi, I.; Chen, Y.; Tan, K.; Ngiam, K. Performance of large language models on the Singapore Medical Licensing Examination: a cross-sectional study. BMJ Health Care Inform. 2024, 31(1), e101234. [Google Scholar] [CrossRef]
  31. Martinu, T.; Menzies, D.; Diaz-Granados, N. Diagnosis and management of respiratory acid-base disorders. Can. Respir. J. 2005, 12(6), 327–32. [Google Scholar] [CrossRef]
  32. Turan, E.I.; Baydemir, A.E.; Balitatli, A.B.; Sahin, A.S. Assessing the accuracy of ChatGPT in interpreting blood gas analysis results: ChatGPT-4 in blood gas analysis. J. Clin. Anesth. 2025, 102, 111787. [Google Scholar] [CrossRef] [PubMed]
  33. Citilcioglu, U.S.; Arslan, B.; Ozdogan, H.K.; Yalim, H.; Kara, U. Is ChatGPT able to interpret arterial blood gas analysis? A comparative cross-sectional study. Signa Vitae 2025, 21(1), 1–9. [Google Scholar] [CrossRef]
  34. Ayala-De la Cruz, N.; Delgado-Marquez, B.; Puente-Fernandez, I.; et al. Human-in-the-Loop Performance of LLM-Assisted Arterial Blood Gas Interpretation: A Single-Center Retrospective Study. J. Clin. Med. 2025, 14(18), 6676. [Google Scholar] [CrossRef] [PubMed]
  35. Story, D.A. Bench-to-bedside review: a brief history of clinical acid-base. Crit. Care 2004, 8(4), 253–8. [Google Scholar] [CrossRef] [PubMed]
  36. Hirosawa, T.; Harada, Y.; Yokose, M.; Sakamoto, T.; Kawamura, R.; Shimizu, T. Diagnostic accuracy of differential-diagnosis lists generated by large language models: a systematic review and meta-analysis. JAMA Netw. Open 2024, 7(6), e2415931. [Google Scholar] [CrossRef]
  37. Wang, X.; Liu, Q.; Iyengar, M.S.; Liu, C.; Li, Y.; Peng, L.; et al. Large language models for emergency medicine: a systematic review. Acad. Emerg. Med. 2024, 31(7), 671–82. [Google Scholar] [CrossRef] [PubMed]
  38. Kanjee, S.; Crowe, B.; Rodman, A. Accuracy of a generative artificial intelligence model in a complex diagnostic challenge. JAMA 2023, 329(19), 1688–90. [Google Scholar] [CrossRef]
Table 1. Overall rubric scores by model and part.
Table 1. Overall rubric scores by model and part.
Model P1 Patient (n=24) P2 Student (n=24) P3 Academic (n=24) P4 Cases (n=28) Total (n=100)
GPT-4o EN 22.92 ± 1.61 23.50 ± 1.72 22.58 ± 1.64 20.71 ± 2.24 22.36 ± 2.11
GPT-4o TR 23.17 ± 1.81 21.88 ± 2.35 20.88 ± 2.17 19.50 ± 2.29 21.28 ± 2.54
Gemini 1.5 Pro EN 22.67 ± 2.26 21.04 ± 2.63 18.79 ± 3.09 17.18 ± 3.20 19.81 ± 3.52
Gemini 1.5 Pro TR 20.46 ± 2.36 18.92 ± 2.32 17.04 ± 2.24 14.21 ± 2.11 17.52 ± 3.26
Maximum 28 28 28 28 28
Data are presented as mean ± standard deviation. Maximum score per part = 28. P1 = Patient-focused definitions; P2 = Student-focused definitions; P3 = Academic/guideline questions; P4 = Clinical case vignettes.
Table 2. Response characteristics.
Table 2. Response characteristics.
Parameter GPT-4o EN GPT-4o TR Gemini 1.5 Pro EN Gemini 1.5 Pro TR
Mean words per response 312.0 ± 98.0 298.0 ± 87.0 275.0 ± 76.0 254.0 ± 82.0
Mean characters per response 2,245 ± 712 2,098 ± 634 1,876 ± 548 1,734 ± 589
Mean response time (seconds) 16.2 ± 4.8 17.8 ± 5.2 19.4 ± 6.1 22.3 ± 7.4
Mean references per response 9.2 ± 2.4 8.1 ± 2.7 6.3 ± 2.1 4.8 ± 2.3
Data are presented as mean ± standard deviation.
Table 3. Dimension-level scores.
Table 3. Dimension-level scores.
Dimension GPT-4o EN GPT-4o TR Gemini 1.5 Pro EN Gemini 1.5 Pro TR p-value
Definitional accuracy 1.68 ± 0.51 1.53 ± 0.56 1.56 ± 0.65 1.27 ± 0.79 0.001
Conceptual accuracy 1.50 ± 0.54 1.63 ± 0.54 1.41 ± 0.66 1.18 ± 0.78 <0.001
Clinical context 1.54 ± 0.57 1.59 ± 0.55 1.52 ± 0.59 1.24 ± 0.76 0.004
Terminology standard 1.60 ± 0.51 1.46 ± 0.64 1.38 ± 0.69 1.29 ± 0.73 0.024
Scope adequacy 1.55 ± 0.59 1.37 ± 0.73 1.34 ± 0.70 1.31 ± 0.69 0.071
Clinical reliability 1.64 ± 0.52 1.55 ± 0.57 1.43 ± 0.70 1.28 ± 0.72 0.002
Verifiability 1.54 ± 0.56 1.50 ± 0.67 1.41 ± 0.65 1.27 ± 0.76 0.056
Word count 1.58 ± 0.53 1.54 ± 0.61 1.40 ± 0.66 1.20 ± 0.77 0.001
Character count 1.66 ± 0.47 1.51 ± 0.59 1.40 ± 0.68 1.20 ± 0.75 <0.001
Response time 1.70 ± 0.46 1.50 ± 0.64 1.38 ± 0.67 1.39 ± 0.73 0.004
Currency 1.68 ± 0.51 1.62 ± 0.49 1.55 ± 0.61 1.28 ± 0.75 <0.001
Hallucination index 1.55 ± 0.57 1.45 ± 0.62 1.29 ± 0.75 1.15 ± 0.79 0.002
Discrimination ability 1.58 ± 0.55 1.51 ± 0.64 1.24 ± 0.71 1.35 ± 0.77 0.003
Guideline concordance 1.56 ± 0.54 1.52 ± 0.59 1.50 ± 0.62 1.11 ± 0.81 <0.001
Data are presented as mean ± standard deviation on a 0–2 scale. p-values from Kruskal-Wallis test across four model-language configurations.
Table 4. Mann-Whitney U post-hoc comparisons.
Table 4. Mann-Whitney U post-hoc comparisons.
Comparison U-statistic p-value Effect size (r)
GPT-4o EN vs. GPT-4o TR 6,238.0 0.0023 0.214
GPT-4o EN vs. Gemini 1.5 Pro EN 7,213.0 <0.001 0.382
GPT-4o EN vs. Gemini 1.5 Pro TR 8,918.5 <0.001 0.677
GPT-4o TR vs. Gemini 1.5 Pro EN 6,222.5 0.0027 0.211
GPT-4o TR vs. Gemini 1.5 Pro TR 8,139.0 <0.001 0.542
Gemini 1.5 Pro EN vs. TR 6,822.5 <0.001 0.315
Bonferroni-adjusted significance threshold: alpha = 0.0083. All comparisons significant at p < 0.0083.
Table 5. Part-wise performance decline.
Table 5. Part-wise performance decline.
Model P1 score P4 score Absolute decline Percent decline Friedman p-value
GPT-4o EN 22.92 20.71 2.21 9.6% 1.01 x 10−4
GPT-4o TR 23.17 19.50 3.67 15.8% 1.32 x 10−5
Gemini 1.5 Pro EN 22.67 17.18 5.49 24.2% 2.63 x 10−6
Gemini 1.5 Pro TR 20.46 14.21 6.25 30.5% 7.79 x 10−10
Percent decline calculated as (P1 score − P4 score) / P1 score × 100. All Friedman tests p < 0.001.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings