Submitted:
19 August 2026
Posted:
21 August 2026
You are already at the latest version
Abstract
Background: Persistently, antimicrobial resistance is one of the major global public health challenges, associated with 4.95 million deaths in 2019, of which 1.27 million were directly attributable to resistant bacterial infections. Among all types, urinary tract infections (UTIs) are among the most common worldwide, with 4.49 billion cases in 2021, and they affect older adults the most. In this context, previous studies suggest an exacerbation in the prevalence of resistant pathogens (such as extended-spectrum β-lactamase (ESBL)-producing Enterobacteriaceae), a circumstance associated with delayed initiation of appropriate treatment, increased mortality, prolonged hospital stays, and selective pressure. These circumstances are more difficult to confront if the limited access to infectious disease specialists is considered, particularly in low- and middle-income countries. In this situation, artificial intelligence (AI) tools have emerged as a potential strategy to support antimicrobial decision-making.Methods: This observational cross-sectional concordance study evaluated the antimicrobial recommendations generated by three AI models: a human-in-the-loop machine learning (ML-HTL) for clinical decision support and two general-purpose language models (GPT-4.0 and Gemini 3.0). A total of 88 cases with positive urine cultures and complete clinical data were included. The AI-generated recommendations for each case were contrasted with the consensus response of an independent panel of three infectious disease specialists. Agreement was assessed using overall concordance, sensitivity, and Cohen’s kappa (κ), with 95% confidence intervals (CI).Results: The mean age of patients was 65.3 years; 81.8% were women, and 80.7% had uncomplicated UTIs. Escherichia coli was the most frequent pathogen (76.1%). Among isolates, 37.5% were ESBL-producing, 53.4% were fluoroquinolone-resistant, and 44.3% were multidrug-resistant. The ML-HTL system showed the highest agreement with the expert´s consensus (90.9%; 95% CI: 83.1–95.3), followed by GPT-4 (63.6%; 95% CI: 53.2–72.9) and Gemini (37.5%; 95% CI: 28.1–47.9). Also, ML-HTL showed high sensitivity and strong agreement (87.5%, κ = 0.89), whereas GPT-4 showed low agreement (κ = 0.32), and Gemini showed less than chance agreement (κ = −0.18).Conclusions: The ML-HTL system, specifically designed for antimicrobial optimization, may better align with experts' consensus than general-purpose language models, supporting their role in clinical decision-making.
Keywords:
artificial intelligence
; antimicrobial stewardship
; urinary tract infections
; bacterial resistance
; antibiotics
Background
Despite clinical and pharmaceutical advances, antimicrobial resistance remains one of the most alarming challenges to global public health, with an estimated 4.95 million deaths worldwide in 2019, of which 1.27 million were directly attributed to bacterial antimicrobial resistance [1]. This circumstance is particularly relevant in frequent bacterial infections, such as urinary tract infections (UTIs), which accounted for 4.49 billion cases worldwide in 2021 [2], and are more prevalent in older adults. In this context, the situation is further aggravated by the notable increase in the prevalence of UTIs caused by extended-spectrum beta-lactamase (ESBL)-producing and carbapenemase-producing enterobacteria, which complicates the empirical selection of antimicrobials and, consequently, the appropriate initiation of antibiotic therapy [3].
To ensure an appropriate therapeutic start, it is essential to adjust treatment not only according to specific patient factors but also in alignment with the rational use of antibiotics and local epidemiological considerations [4]. Moreover, inappropriate therapeutic initiation is associated with a 1.5-2-fold increase in mortality, longer hospital stays, and, as a side effect, greater selective pressure [5,6,7].
Given these circumstances, antimicrobial stewardship programs (ASPs) have become essential in healthcare institutions to improve antimicrobial prescribing and combat resistance [8,9]. Despite their importance, a major limitation of ASPs is the limited availability of infectious disease specialists [10], as well as organizational deficiencies that are more pronounced in resource-limited settings [11].
Consequently, this limitation has increased interest in data management fields, particularly artificial intelligence (AI), to expand access to expert-level antimicrobial guidance. Recent articles have shown that AI-based systems demonstrated an overall diagnostic accuracy of approximately 87.5%, close to that of infectious disease specialists (90.3%) and higher than that of students during their years of specialization (77.8%) [12]. In relation to antimicrobial therapy support, one study reports that machine learning-based software that generates therapeutic recommendations increased appropriate empirical antibiotic coverage from 68.93% in routine medical prescriptions to over 90%, with organism-targeted therapy accuracy exceeding 95% [13].
Regarding the use of AI in this type of scenario, two main approaches have emerged. The first is based on traditional machine learning models trained with clinical and microbiological data (electronic health records, laboratory results, and antibiograms) within the paradigm of big data-based machine learning in healthcare, with adaptations to the contexts of antimicrobial prescribing and optimization [14]. In this regard, Human-in-the-loop (ML-HTL) machine learning models have also been reported and validated both internally and externally, yielding significant results [15]. Within this group, it is also worth mentioning hybrid tree learning algorithms, which have used demographic data, comorbidities, and patient culture results to generate personalized recommendations that comply with institutional protocols [16,17].
A second scenario involves large language models (LLMs) based on transformer architectures, which have enabled the interpretation of unstructured clinical narratives (including progress notes, consultation reports, and guideline texts) to support clinical reasoning, antimicrobial optimization, and adherence to infectious disease management guidelines [18,19]. The LLM-based systems have been reported to provide decision support comparable to specialist recommendations in certain infectious disease scenarios, thereby enhancing access to expert guidance on antimicrobials [20].
The LLM systems (such as GPT-4, Claude, Gemini, among others) are trained on extensive text corpora, such as medical literature and clinical guidelines [21], enabling them to generate complex reasoning and contextually relevant responses [22].
However, their assessments remain mixed and may present inaccurate data, inconsistent adherence to institutional protocols, and limited incorporation of data on local resistance [23].
Previous research has shown that machine learning-based clinical decision support systems (ML-CDSS) can facilitate the interpretation of antimicrobial susceptibility and improve empirical and targeted antimicrobial therapy in multiple infectious disease scenarios. [14,24,25]. However, comparisons between ML-CDSS and LLM platforms for antimicrobial decision support remain scarce and may exhibit higher error rates, security issues, and inconsistent performance when applied to complex antimicrobial decision-making tasks [26,27].
It remains unclear which AI implementation strategies are most appropriate in this medical field [21,27]. Evaluations of medical AI systems should be based on independent expert assessment rather than author-determined or self-validated results, as emphasized in previous evaluations of AI implementation in healthcare [23,28]. Therefore, this study evaluated the concordance of an ML-HTL algorithm [15] with two of the most widely used general-purpose LLMs (GPT-4.0 and Gemini 3.0) regarding the antimicrobial treatment recommendations they provided for clinical cases of UTIs. The aim was to describe and compare the optimization of antimicrobial use through evidence-based AI and to evaluate how specialized ML algorithms and general LLMs can complement each other to improve antimicrobial prescribing.
Methods
Study Design and Setting
This study used a cross-sectional observational concordance design. It was conducted at a private reference laboratory in Lima, Peru. Both the performance of a clinical decision support system based on an ML-HTL (OneChoice by Arkstone, North Carolina, USA) and two LLMs (GPT-4.0 (OpenAI, California, USA) and Gemini 3.0 (Google, California, USA)) were evaluated against the judgment of infectious disease experts to provide antimicrobial therapy recommendations in real cases of UTI. Chronologically, the clinical cases employed occurred from October 21 to December 15 of 2024. Using this data for secondary purposes, the group collected variables for analysis (demographic, clinical, and microbiological), generated clinical vignettes, and incorporated them into the AI’s prompts for each case from November 1 to November 30, 2025. Subsequently, the relationship between AI performance and the expert judgment was evaluated from December 1 to December 31, 2025.
This study was reported following the STROBE guidelines for cross-sectional studies and the GRRAS recommendations for concordance analysis.
Case Selection
Inclusion Criteria
- Patients with positive urine culture with isolation of a single urologically pathogenic microorganism with bacterial growth of at least 100,000 colony-forming units per milliliter.
- Availability of clinical data previously reviewed by the OneChoice ML-HTL platform.
- Complete antimicrobial susceptibility test results for the isolated organism.
Exclusion Criteria
- Urine culture showed polymicrobial growth (≥2 organisms).
- Incomplete clinical or microbiological data.
- Microorganism isolated without validated cut-off points to interpret antimicrobial susceptibility.
Standard Reference
A panel of independent experts was formed, consisting of three physicians certified in infectious diseases, with training or with a certified subspecialty in antimicrobial stewardship. Two of them reside in Peru, and one in the United States. Based on their expertise, the reference standard for each case was established by simple majority (agreement when at least 2 of the 3 experts agreed on the therapeutic recommendation). This approach was implemented to address clinical scenarios with multiple plausible treatment options, consistent with methods used in other research on clinical decision support tools [29,30].
Index Test
- a.
- OneChoice ML-HTL Platform (Arkstone): The OneChoice system is a cloud-based platform that supports clinical decision-making. It uses a hybrid machine learning-human interaction (ML-HTL) algorithm that integrates patient demographic information, clinical data, and microbiological culture results (such as age, sex, site of infection, renal function, allergies, and culture-specific susceptibility profiles) to generate personalized antimicrobial treatment suggestions (15). Employing these structured clinical variables, the algorithm prioritizes treatment options based on predicted effectiveness, institutional protocols, and spectrum coverage.
- b.
- Large Language Models (LLMs) - ChatGPT: a broad language model based on transformers, trained on a large-scale text corpus using self-supervised learning, with capabilities for medical knowledge synthesis and simulated clinical reasoning tasks, although it may have limitations in terms of security and reliability [31].
- c.
- Large Language Models (LLMs) - Gemini 3.0: a multimodal large language model based on transformers, trained on large-scale text corpora using self-supervised next token prediction. It was accessed via the standard web interface during the study period.
Evaluation of AI Systems
A standardized clinical vignette was constructed for each case, containing the following elements: a) the patient demographics data (age, sex), b) relevant medical history and comorbidities, c) patient’s signs and symptoms, d) laboratory results (including a complete urinalysis), and e) the urine culture results (which included the identification of the microorganism and its antibiogram). Later, every single clinical vignette was integrated into a standardized prompt, with the final instruction to request a justified antimicrobial treatment recommendation. These last prompts were analyzed by the three models simultaneously, and all requests were conducted during the study period (to ensure comparability, without iterative refinement or optimization of the instructions). All cases and prompts are available in Supplement 1.
Afterward, each clinical case was then presented to the expert panel via a secure electronic link using a standardized format. For each case, the experts received:
- The integrated prompt (the clinical vignette plus the standardized prompt) is the same as the instructions given to the AI systems.
- The three treatment recommendations of the models, designated as Recommendation #1, Recommendation #2, and Recommendation #3.
Thanks to these tags, the experts were unaware of the origin of each recommendation. Finally, with all the information brought, the experts independently evaluated the responses and selected the most appropriate one based on their clinical judgment, compliance with current guidelines, and the specific patient context. Any communication between panelists during the evaluation process was strictly prohibited. For details on the expert’s review and the conflict-of-interest statements, see Supplement 2.
Estimated Outcomes
The primary outcome was the overall agreement rate between each system’s recommendation and the expert panel. This is defined as the proportion of cases in which the AI system’s recommendation matched the expert panel’s consensus. In addition, the agreement was stratified by pathogenic organism type, resistance phenotype, and infection severity.
Statistical Analysis
The statistical analysis was performed with Stata version 17.0 (StataCorp LLC, College Station, TX). The study’s concordance rates were reported as proportions. Also, the study employed Wilson’s scoring method to estimate 95% confidence intervals (CIs). Common measurements such as sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) were estimated, along with 95% CIs. The McNemar test for paired proportions was used to compare AI systems, with p-values < 0.05 considered statistically significant. Finally, the Inter-rater reliability among the three experts was assessed using Fleiss’s Kappa coefficient, interpreted as: <0.20 poor, 0.21-0.40 fair, 0.41-0.60 moderate, 0.61-0.80 substantial, and 0.81-1.00 almost perfect agreement [32].
Results
Eighty-eight UTI cases that met the inclusion criteria were included. Most patients were women (81.8%), and the mean age was 65.3 years (SD: 20.9). The 80.7% of UTIs were uncomplicated, and 19.3% were complicated. The most frequently isolated microorganisms were Escherichia coli (76.1%), Klebsiella pneumoniae (8%), Enterococcus sp. (6.8%), and Proteus mirabilis (4.5%). Likewise, 37.5% of isolates were ESBL-producing, 53.4% were fluoroquinolone-resistant, and 44.3% were multidrug-resistant (MDR) (Table 1).
Compared with the reference standard (expert consensus), Onechoice showed the highest overall agreement (90.9%; 95% CI: 83.1–95.3), followed by GPT-4 (63.0%; 95% CI: 53.2–72.9) and Gemini (37.5%; 95% CI: 28.1 - 47.9). Likewise, Onechoice showed high sensitivity (87.5%; 95% CI: 77.2 – 93.5), with an almost perfect agreement coefficient (κ = 0.89), while GPT-4 showed moderate sensitivity (50.0%; 95% CI: 38.1–61.9), with poor agreement (κ = 0.32), and Gemini showed low sensitivity (14.1%; 95% CI: 7.6–24.6), with poor agreement (κ = 0.18). All systems showed high positive predictive value (100%), although negative predictive value varied considerably (Table 2, Figure 1).
Depending on the complexity of the case, the AI system’s performance varied significantly. In cases of uncomplicated UTIs, Onechoice showed high concordance (95.8%), significantly outperforming GPT-4 (69.0%) and Gemini (39.4%), while in complicated UTIs, agreement decreased for all AI systems (OneChoice (70.6%), GPT-4 (41.2%), and Gemini (29.4%), p: 0.074), and no statistically significant differences were observed between them. Regarding the resistance profile, Onechoice maintained high concordance rates with susceptible microorganisms (93.3%), ESBL producers (84.8%), and MDR (89.7%), which were statistically significantly higher than those observed with GPT-4 and Gemini. Likewise, OneChoice achieved a concordance of 95.5% for Escherichia coli infections, significantly higher than the 70.1% of GPT-4 and 46.3% of Gemini (p<0.001), and for other microorganisms, the concordance of OneChoice was 76.2%, compared to 42.9% for GPT-4 and 9.5% for Gemini (p<0.001) (Table 3).
The agreement among the three members of the expert panel was low, as indicated by Cohen’s kappa coefficients ranging from 0.062 to 0.147, corresponding to a slight level of agreement. Overall reliability among the expert panel was poor, with a Fleiss Kappa of -0.004 (95% CI: -0.069 – 0.061) (Table 4). Likewise, moderate agreement was observed between GPT-4 and Gemini (78.4%), while Onechoice showed high agreement with the expert panel (90.9%) and lower agreement with Gemini (60.2%) and GPT-4 (53.4%) (Figure 2).
Figure 3 summarizes the comparative performance of AI-based antimicrobial recommendation systems using an alluvial diagram that illustrates the agreement among the therapeutic recommendations evaluated by the expert panel. OneChoice showed greater agreement with the expert panel in 80 of 88 recommendations (90.9%). GPT-4 showed intermediate performance, with 56 correct recommendations (63.6%) and 32 rejected (36.4%). Gemini performed the worst, with 33 recommendations (37.5%) approved and 55 (62.5%) considered inappropriate.
Discussion
The main objective of the current study was to assess the performance of three artificial intelligence systems (the ML-HTL “OneChoice” model, GPT-4, and Gemini) as potential antimicrobial treatment references for UTIs, and to compare their recommendations with the consensus of infectious disease experts. We found that the concordance of the ML-HTL algorithm with expert consensus was 90.9%, compared to 63.6% for GPT-4 and 37.5% for Gemini. It also observed higher agreement in cases of UTI caused by microorganisms with ESBL- and MDR-resistant phenotypes (p < 0.01), even in complicated UTIs, although the difference was not statistically significant.
These findings support published evidence demonstrating the clinical utility of specific ML-CDSS in treating infectious diseases. A review of 60 ML-CDSS systems for infectious diseases found that those using structured clinical data achieved high diagnostic accuracy in 95% of the cases studied [33]. They also reported a concordance rate of 90.9% [33], which was lower than the rate reported by the ML-HTL model in earlier validation studies [15].
This study shows that 62.5% of GPT-4 errors and 63.6% of Gemini errors were attributable to inappropriate antibiotic selection, including incorrect use of the therapeutic range or selection of antibiotics unsuitable for specific pathogens. This is likely because LLMs often produce hallucinations, lack the contextual understanding necessary to make detailed diagnostic and treatment decisions, and rely on data and training methods that are difficult to evaluate [31]. Some factors that may explain the performance gap observed between LLM and ML-HTL systems is the latter integrates demographic data such as age, gender, allergies, and ICD-10, and its training has been based on approvals from the Food and Drug Administration (FDA), the Centers for Disease Control and Prevention (CDC), the guidelines of the Infectious Diseases Society of America (IDSA), and other bibliographic sources verified by specialists.
In addition, it undergoes human expert intervention when the algorithm does not provide 100% certainty, integrating structured data, including patient demographics, clinical parameters, microbiological culture and susceptibility results, and institutional antibiogram patterns (using a specially designed hybrid tree-learning algorithm) [15].
The concordance of the ML-HTL (90.9%) and its concordance coefficient (κ = 0.89) suggest that this tool could be more suitable for optimizing antimicrobial use, in contrast with the 62.5% discordance rate observed with Gemini, indicating concordance below chance (κ = −0.18), which suggests that current general-purpose LLMs are not suitable without a substantial human oversight. Although some authors have suggested that GPT-4 could be used for “curbside consultations”[34], our data indicate that this should be approached with great caution, particularly for complex, resistant infections.
Subgroup analysis revealed clear differences in performance across the clinical scenarios evaluated. The ML-HTL model showed significant agreement across all groups analyzed. In contrast, the language model’s performance progressively declined as the clinical complexity of the cases increased. In uncomplicated urinary tract infections, the agreement of the ML-HTL model (95.8%) was significantly higher than that of GPT-4 (69.0%; p < 0.001). In complicated urinary tract infections, although the absolute difference in agreement was comparable, the small sample size (n = 17) likely limited statistical power, which could explain the result (p = 0.074).
We observed lower agreement across all systems evaluated for infections caused by BLEE-positive and MDR microorganisms. This particularity is highlighted by Tamma et al. [36], who noted that treating antimicrobial-resistant Gram-negative infections is challenging and requires nuanced decisions that incorporate resistance mechanisms, available therapeutic options, and pharmacokinetic-pharmacodynamic considerations. In a comparable manner, the GPT-4 and Gemini’s performance in evaluating ESBL-positive and MDR cases (48.5% and 51.3%, respectively) suggests that both may be structurally constrained in addressing resistance patterns, underscoring the need to upgrade specialized clinical reasoning and advanced contextual integration in these models. Meanwhile, the ML-HTL model performed better against Escherichia coli infections than non-Escherichia coli pathogens (95.5% and 76.2%, respectively), which may be related to the predominance of this pathogen in UTIs and its well-characterized resistance patterns. We note a significant decline in Gemini’s performance against non-Escherichia coli pathogens (9.5%), further highlighting the weakness of LLM models’ recommendations in less common clinical scenarios.
The results of the following study may be relevant to healthcare systems as a source for considering the implementation of AI to optimize antimicrobial use. According to the IDSA, the optimization programs encounter considerable barriers, particularly the limited availability of infectious disease specialists [8,11]. In 2017, 79.5% of counties in the United States had only one infectious disease specialist, and more than 200 million people lived in counties with few or no infectious disease specialists [31]. Similarly, in Latin America, there have been reports of shortages of specialists and substantial deficiencies in infrastructure for optimizing antimicrobial use [32].
The study has limitations. First, it was conducted at a single reference laboratory in Peru, which may limit the generalizability of our findings to other clinical settings. Second, our evaluation included only two LLMs (GPT-4 and Gemini), which are continually updated, and the results collected today could differ. Third, LLM indications were standardized without iterative refinement or optimization, which may have underestimated the achievable performance of LLMs. Fourth, our study evaluated AI recommendations against experts’ consensus and did not consider clinical outcomes, such as mortality or hospital length of stay. For superior studies, randomized controlled trials or prospective studies comparing ML-CDSS-guided prescribing with standard care or with other LLMs are suggested to establish clinical efficacy beyond diagnostic accuracy. Likewise, continuous validation and updating of AI systems are required to maintain their clinical relevance as resistance patterns evolve. Rather than expanding the number of experts, it is also important to implement a prior consensus and calibration process (e.g., Delphi rounds or training with pilot cases and explicit criteria) to standardize evaluation criteria and improve inter-rater reliability. Finally, priority should be given to implementation research that addresses barriers to the adoption of AI-CDSS in resource-limited settings, given the greater potential impact in such contexts [32,33].
Conclusions
In this comparison of AI systems for antimicrobial treatment recommendations in UTIs, the ML-HTL platform achieved statistically significant results relative to general-purpose LLMs, achieving 90.9% agreement with expert consensus, compared with 63.6% for GPT-4 and 37.5% for Gemini. This superiority was consistent across all clinical subgroups, including resistant phenotypes, where LLM performance declined significantly. These findings support the clinical utility of ML-CDSS, specifically designed to optimize antimicrobial use.
List of Abbreviations
UTI: urinary tract infections, ESBL: Extended-spectrum β-lactamases, AI: artificial intelligence, ML-HTL: Human-in-the-loop machine learning, LLM: Large language models, ASP: antimicrobial stewardship program, κ: Cohen’s kappa, CI: confidence intervals, PPV: positive predictive value, NPV: negative predictive value, MDR: multidrug-resistant, ML-CDSS: machine learning-based clinical decision support systems, IDSA: Infectious Diseases Society of America, CDC: Centers for Disease Control and Prevention.
Ethics Approval and Consent for Publication
The study was conducted in accordance with the Declaration of Helsinki and local regulations. The study’s protocol was approved by the Research Ethics Committee of the Private University of Tacna, under code “FACSA-CEI_168-10-25,” on October 21, 2025. Their verdict exempted this study from obtaining informed consent because of its observational, retrospective design and secondary data analysis.
Availability of Data and Materials
The datasets generated and analyzed during the current study are available in Mendeley Data, a cloud-based communal repository, available in this link: https://data.mendeley.com/datasets/6tjwss4g3w/1.
Competing Interests
The authors declare that they have no competing interests.
Funding
This study received partial financial support from Arkstone Medical Solutions. This funding was allocated exclusively to cover the expert panel’s costs for the study. All other study-related expenses, including design, data collection, analysis, and manuscript preparation, were fully self-funded by the research team.
Authors’ Contributions
The study’s conceptualization was carried out by JCG, AF, and MHZ. The data curation was carried out by JCG and JAC. The Formal Analysis was carried out by JCG, MHZ, and CCL. Funding acquisition was carried out by JCG, AF, and AR. The investigation was carried out by JCG, JAC, MHZ, and DMV. The methodology was carried out by JCG, AR, and MHZ. The Project Administration was carried out by JCG and MHZ. The Resources topic was carried out by JCG and CCL. The Software-related topics were carried out by MHZ. The Supervision of the study was carried out by JCG, AF, and AR. The Validation process was carried out by JCG, MHZ, and CCL. The Visualization topic was carried out by MHZ. The original draft of the manuscript was prepared by JCG, MHZ, and DMV. The review and editing of the manuscript were carried out by JCG, MHZ, Frenkel A., and JAC. All authors read and approved the final manuscript.
Acknowledgments
Not applicable.
References
- Murray, C.J.L.; Ikuta, K.S.; Sharara, F.; Swetschinski, L.; Aguilar, G.R.; Gray, A.; et al. Global burden of bacterial antimicrobial resistance in 2019: a systematic analysis. The Lancet 2022, 399, 629–55. [Google Scholar] [CrossRef] [PubMed]
- He, Y.; Zhao, J.; Wang, L.; Han, C.; Yan, R.; Zhu, P.; et al. Epidemiological trends and predictions of urinary tract infections in the global burden of disease study 2021. Sci. Rep. 2025, 15, 4702. [Google Scholar] [CrossRef] [PubMed]
- Zilberberg, M.D.; Shorr, A.F. Secular Trends in Gram-Negative Resistance among Urinary Tract Infection Hospitalizations in the United States, 2000–2009. Infect. Control Hosp. Epidemiol. 2013, 34, 940–6. [Google Scholar] [CrossRef] [PubMed]
- Gupta, K.; Hooton, T.M.; Naber, K.G.; Wullt, B.; Colgan, R.; Miller, L.G.; et al. International Clinical Practice Guidelines for the Treatment of Acute Uncomplicated Cystitis and Pyelonephritis in Women: A 2010 Update by the Infectious Diseases Society of America and the European Society for Microbiology and Infectious Diseases. Clin. Infect. Dis. 2011, 52, e103–20. [Google Scholar] [CrossRef] [PubMed]
- Allel, K.; Stone, J.; Undurraga, E.A.; Day, L.; Moore, C.E.; Lin, L.; et al. The impact of inpatient bloodstream infections caused by antibiotic-resistant bacteria in low- and middle-income countries: A systematic review and meta-analysis. PLoS Med. 2023, 20, e1004199. [Google Scholar] [CrossRef] [PubMed]
- Mo, Y.; Oonsivilai, M.; Lim, C.; Niehus, R.; Cooper, B.S. Implications of reducing antibiotic treatment duration for antimicrobial resistance in hospital settings: A modelling study and meta-analysis. PLoS Med. 2023, 20, e1004013. [Google Scholar] [CrossRef] [PubMed]
- George, N.A.; Pan, D.; Silva, L.; Baggaley, R.F.; Irizar, P.; Divall, P.; et al. The prevalence and risk of mortality associated with antimicrobial resistance within nosocomial settings—a global systematic review and meta-analysis of over 20,000 patients. eClinicalMedicine 2025, 87. [Google Scholar] [CrossRef] [PubMed]
- Barlam, T.F.; Cosgrove, S.E.; Abbo, L.M.; MacDougall, C.; Schuetz, A.N.; Septimus, E.J.; et al. Implementing an Antibiotic Stewardship Program: Guidelines by the Infectious Diseases Society of America and the Society for Healthcare Epidemiology of America. Clin. Infect. Dis. 2016, 62, e51–77. [Google Scholar] [CrossRef] [PubMed]
- Dellit, T.H.; Owens, R.C.; McGowan, J.E.; Gerding, D.N.; Weinstein, R.A.; Burke, J.P.; et al. Infectious Diseases Society of America and the Society for Healthcare Epidemiology of America Guidelines for Developing an Institutional Program to Enhance Antimicrobial Stewardship. Clin. Infect. Dis. 2007, 44, 159–77. [Google Scholar] [CrossRef] [PubMed]
- Rhee, C.; Chiotos, K.; Cosgrove, S.E.; Heil, E.L.; Kadri, S.S.; Kalil, A.C.; et al. Infectious Diseases Society of America Position Paper: Recommended Revisions to the National Severe Sepsis and Septic Shock Early Management Bundle (SEP-1) Sepsis Quality Measure. Clin. Infect. Dis. 2021, 72, 541–52. [Google Scholar] [CrossRef] [PubMed]
- Doernberg, S.B.; Abbo, L.M.; Burdette, S.D.; Fishman, N.O.; Goodman, E.L.; Kravitz, G.R.; et al. Essential Resources and Strategies for Antibiotic Stewardship Programs in the Acute Care Setting. Clin. Infect. Dis. 2018, 67, 1168–74. [Google Scholar] [CrossRef] [PubMed]
- Zhan, L.; Dang, X.; Xie, Z.; Zeng, C.; Wu, W.; Zhang, X.; et al. Evaluating GPT-4o in infectious disease diagnostics and management: A comparative study with residents and specialists on accuracy, completeness, and clinical support potential. Digit Health 2025, 11, 20552076251355797. [Google Scholar] [CrossRef] [PubMed]
- Tejeda, M.I.; Fernández, J.; Valledor, P.; Almirall, C.; Barberán, J.; Romero-Brufau, S. Retrospective validation study of a machine learning-based software for empirical and organism-targeted antibiotic therapy selection. Antimicrob. Agents Chemother. 2024, 68, e0077724. [Google Scholar] [CrossRef] [PubMed]
- Giacobbe, D.R.; Marelli, C.; Guastavino, S.; Signori, A.; Mora, S.; Rosso, N.; et al. Artificial intelligence and prescription of antibiotic therapy: present and future. Expert Rev. Anti-Infect. Ther. 2024, 22, 819–33. [Google Scholar] [CrossRef] [PubMed]
- Frenkel, A.; Rendon, A.; Chavez-Lencinas, C.; Gomez De la Torre, J.C.; MacDermott, J.; Gross, C.; et al. Internal Validation of a Machine Learning-Based CDSS for Antimicrobial Stewardship. Life 2025, 15, 1123. [Google Scholar] [CrossRef] [PubMed]
- Beaudoin, M.; Kabanza, F.; Nault, V.; Valiquette, L. Evaluation of a machine learning capability for a clinical decision support system to enhance antimicrobial stewardship programs. Artif. Intell. Med. 2016, 68, 29–36. [Google Scholar] [CrossRef] [PubMed]
- Kullar, R.; Goff, D.A.; Schulz, L.T.; Fox, B.C.; Rose, W.E. The “Epic” Challenge of Optimizing Antimicrobial Stewardship: The Role of Electronic Medical Records and Technology. Clin. Infect. Dis. 2013, 57, 1005–13. [Google Scholar] [CrossRef] [PubMed]
- Cho, H.N.; Jun, T.J.; Kim, Y.-H.; Kang, H.; Ahn, I.; Gwon, H.; et al. Task-Specific Transformer-Based Language Models in Health Care: Scoping Review. JMIR Med. Inform. 2024, 12, e49724. [Google Scholar] [CrossRef] [PubMed]
- Omar, M.; Brin, D.; Glicksberg, B.; Klang, E. Utilizing natural language processing and large language models in the diagnosis and prediction of infectious diseases: A systematic review. Am. J. Infect. Control. 2024, 52, 992–1001. [Google Scholar] [CrossRef] [PubMed]
- Lorenzoni, G.; Garbin, A.; Brigiari, G.; Papappicco, C.A.M.; Manfrin, V.; Gregori, D. Large Language Models in Action: Supporting Clinical Evaluation in an Infectious Disease Unit. Healthcare 2025, 13, 879. [Google Scholar] [CrossRef] [PubMed]
- Han, J.; Qiu, W.; Lichtfouse, E. ChatGPT in Scientific Research and Writing: A Beginner’s Guide; 2024. [Google Scholar] [CrossRef]
- Thirunavukarasu, A.J.; Ting, D.S.J.; Elangovan, K.; Gutierrez, L.; Tan, T.F.; Ting, D.S.W. Large language models in medicine. Nat. Med. 2023, 29, 1930–40. [Google Scholar] [CrossRef] [PubMed]
- Mello, M.M.; Guha, N. ChatGPT and Physicians’ Malpractice Risk. JAMA Health Forum 2023, 4, e231938. [Google Scholar] [CrossRef] [PubMed]
- Ardila, C.M.; González-Arroyave, D.; Tobón, S. Machine learning for predicting antimicrobial resistance in critical and high-priority pathogens: A systematic review considering antimicrobial susceptibility tests in real-world healthcare settings. PLoS ONE 2025, 20, e0319460. [Google Scholar] [CrossRef] [PubMed]
- AlGain, S.; Marra, A.R.; Kobayashi, T.; Marra, P.S.; Celeghini, P.D.; Hsieh, M.K.; et al. Can we rely on artificial intelligence to guide antimicrobial therapy? A systematic literature review. Antimicrob. Steward. Healthc. Epidemiol. 2025, 5, e90. [Google Scholar] [CrossRef] [PubMed]
- Bienvenu, A.-L.; Ducrocq, J.-M.; Augé-Caumon, M.-J.; Baseilhac, E. Clinical decision support system to guide antimicrobial selection: a narrative review from 2019 to 2023. J. Hosp. Infect. 2025, 162, 140–52. [Google Scholar] [CrossRef] [PubMed]
- Singhal, K.; Azizi, S.; Tu, T.; Mahdavi, S.S.; Wei, J.; Chung, H.W.; et al. Large language models encode clinical knowledge. Nature 2023, 620, 172–80. [Google Scholar] [CrossRef] [PubMed]
- Topol, E.J. High-performance medicine: the convergence of human and artificial intelligence. Nat. Med. 2019, 25, 44–56. [Google Scholar] [CrossRef] [PubMed]
- Kellerhuis, B.E.; Jenniskens, K.; Kusters, M.P.T.; Schuit, E.; Hooft, L.; Moons, K.G.M.; et al. Expert panel as reference standard procedure in diagnostic accuracy studies: a systematic scoping review and methodological guidance. Diagn. Progn. Res. 2025, 9, 12. [Google Scholar] [CrossRef] [PubMed]
- Kea, B.; Sun, B.C.-A. Consensus development for healthcare professionals. Intern Emerg. Med. 2015, 10, 373–83. [Google Scholar] [CrossRef] [PubMed]
- Schwartz, I.S.; Link, K.E.; Daneshjou, R.; Cortés-Penfield, N. Black Box Warning: Large Language Models and the Future of Infectious Diseases Consultation. Clin. Infect. Dis. 2024, 78, 860–6. [Google Scholar] [CrossRef] [PubMed]
- Fabre, V.; Cosgrove, S.E.; Secaira, C.; Torrez, J.C.T.; Lessa, F.C.; Patel, T.S.; et al. Antimicrobial stewardship in Latin America: Past, present, and future. Antimicrob. Steward. Healthc. Epidemiol. 2022, 2, e68. [Google Scholar] [CrossRef] [PubMed]
- Peiffer-Smadja, N.; Rawson, T.M.; Ahmad, R.; Buchard, A.; Georgiou, P.; Lescure, F.-X.; et al. Machine learning for clinical decision support in infectious diseases: a narrative review of current applications. Clin. Microbiol. Infect. 2020, 26, 584–95. [Google Scholar] [CrossRef] [PubMed]
- Lee, P.; Bubeck, S.; Petro, J. Benefits, Limits, and Risks of GPT-4 as an AI Chatbot for Medicine. N. Engl. J. Med. 2023, 388, 1233–9. [Google Scholar] [CrossRef] [PubMed]
Figure 1.
Comparative diagnostic performance of AI systems vs expert panel.

Figure 2.
Pairwise concordance heatmap between AI systems and Expert panel.

Figure 3.
Alluvial diagram of the flow of antibiotic recommendations of AI systems to the expert panel.
Figure 3.
Alluvial diagram of the flow of antibiotic recommendations of AI systems to the expert panel.

Table 1.
Baseline Characteristics of Urinary Tract Infection cases (n=88).
| Variable | n (%) or Mean ± SD |
| Demographics | |
| - Age (years), mean ± SD | 65.3 ± 20.9 |
| - Female sex | 72 (81.8) |
| UTI Classification | |
| - Uncomplicated | 71 (80.7) |
| - Complicated | 17 (19.3) |
| Isolated Pathogens* | |
| - Escherichia coli | 67 (76.1) |
| - Klebsiella pneumoniae | 7 (8.0) |
| - Enterococcus spp. | 6 (6.8) |
| - Proteus mirabilis | 4 (4.5) |
| - Staphylococcus spp. | 2 (2.3) |
| - Pseudomonas aeruginosa | 1 (1.1) |
| - Citrobacter spp. | 1 (1.1) |
| Antimicrobial Resistance Profile | |
| - ESBL-producing organisms** | 33 (37.5) |
| - Fluoroquinolone resistance | 47 (53.4) |
| - MDR (≥3 antimicrobial classes) | 39 (44.3) |
Abbreviations: UTI, urinary tract infection; ESBL, extended-spectrum β-lactamase; MDR, multidrug-resistant; SD, standard deviation. *Pathogens identified from urine culture with significant growth. **ESBL-producing organisms were defined as isolates demonstrating resistance to ≥2 third-generation cephalosporins (ceftriaxone, cefotaxime, ceftazidime, or cefixime).
Table 2.
Diagnostic performance of AI systems compared with Expert Panel.
| AI System | Concordance (95% CI) | Sensitivity (95% CI) | Specificity (95% CI) | PPV (95% CI) |
NPV (95% CI) |
κ (95% CI) |
| OneChoice | 90.9 (83.1–95.3) |
87.5 (77.2–93.5) |
100.0 (86.2–100) |
100.0 | 75.0 (57.9–86.7) |
0.89 (0.89–0.89) |
| GPT-4 | 63.6 (53.2–72.9) |
50.0 (38.1–61.9) |
100.0 (86.2–100.0) |
100.0 | 42.9 (30.8–55.9) |
0.32 (0.10–0.54) |
| Gemini | 37.5 (28.1–47.9) |
14.1 (7.6–24.6) |
100.0 (86.2–100.0) |
100.0 | 30.4 (21.3–41.2) |
−0.18 (−0.41–0.06) |
Abbreviations: CI, confidence interval; ML-HTL, machine learning hybrid tree-learning; PPV, positive predictive value; NPV, negative predictive value; κ, Gwet’s AC1 agreement coefficient.
Table 3.
Concordance stratified by case complexity.
| Subgroup | n | OneChoice (%) | GPT-4 (%) | Gemini (%) | p-value* |
| UTI Type | |||||
| Uncomplicated | 71 | 95.8 (88.3 – 98.6) | 69.0 (57.5 – 78.6) | 39.4 (28.9 – 51.1) | <0.001 |
| Complicated | 17 | 70.6 (46.9 – 86.7) | 41.2 (21.6 – 64.0) | 29.4 (13.3 – 53.1) | 0.074 |
| Resistance profile | |||||
| Susceptible | 45 | 93.3 (82.1 – 97.7) | 77.8 (63.7 – 87.5) | 28.9 (17.7 – 43.4) | <0.001 |
| ESBL+ | 33 | 84.8 (69.1 – 93.3) | 48.5(32.5 – 64.8) | 48.5 (32.5 – 64.8) | 0.002 |
| MDR | 39 | 89.7 (76.4 – 95.9) | 51.3 (36.2 – 66.1) | 51.3 (36.2 – 66.1) | <0.001 |
| Pathogen | |||||
| E. coli | 67 | 95.5 (87.6 – 98.5) | 70.1 (58.3 – 79.8) | 46.3 (34.9 – 58.1) | <0.001 |
| Non-E. coli | 21 | 76.2 (54.9 – 89.4) | 42.9 (24.5 – 63.5) | 9.5 (2.7 – 28.9) | <0.001 |
Table 4.
Agreement betwwen experts panel.
| Expert Pair | Cohen’s κ | 95% CI | Interpretation |
| Expert 1 vs Expert 2 | 0.085 | −0.001–0.170 | Slight |
| Expert 1 vs Expert 3 | 0.147 | 0.051–0.245 | Slight |
| Expert 2 vs Expert 3 | 0.062 | 0.005–0.122 | Slight |
| Fleiss’ κ (global) | −0.004 | −0.069–0.061 | Poor |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.