Preprint
Article

This version is not peer-reviewed.

Assessing the Performance of Artificial Intelligence on Anesthesiology In-Training Examinations and Applicability in Medical Education

Submitted:

22 September 2026

Posted:

23 September 2026

You are already at the latest version

Abstract
Introduction: Large language models (LLMs) have demonstrated substantial performance on medical licensing and board examinations, but their application to anesthesiology remains less well studied. Prior evaluations have not compared current-generation models from different developers or assessed performance on figure-dependent questions that earlier models lacked the capability to interpret. This study evaluates the performance of two current-generation LLMs on a comprehensive anesthesiology in-training examination (ITE) review question bank, including figure-dependent items. Methods: A total of 1,001 single-best-answer questions from an anesthesiology ITE review text, spanning 11 content chapters, were administered to Claude Opus 5 and GPT-5.6. No questions were excluded. Each item was presented once in a fresh, stateless context with no tool access, retrieval, answer key, or explanation provided in the prompt. Performance was analyzed by question format, and all 27 figure-dependent items were administered with their published figures supplied. Accuracy between models was compared using McNemar’s test. Results: Claude Opus 5 answered 947/1,001 items correctly (94.6%; 95% CI, 93.0–95.8), and GPT-5.6 answered 946/1,001 correctly (94.5%; 95% CI, 92.9–95.8); the difference was not significant (McNemar p = 1.00). Accuracy was similar across most question formats. On standard questions, accuracy was 94.3% and 94.6% for Claude Opus 5 and GPT-5.6, respectively; on the 18 figure-based questions with figures supplied, accuracy was 94.4% and 83.3%, respectively; and both models achieved 100% accuracy on image-option items. Withholding figures from the same 18 figure-based questions reduced pooled accuracy from 88.9% to 52.8% (exact McNemar p = 0.03 and p = 0.04 for Claude Opus 5 and GPT-5.6, respectively). The models agreed on 957/1,001 items (95.6%). Of the 33 items both models answered incorrectly, they selected the same incorrect option on 32 (97%). Discussion: Current-generation LLMs achieved approximately 95% accuracy, substantially improving on our prior results with earlier-generation models and placing both models within the highest band of the ITE normative framework, corresponding approximately to the 99th percentile across training levels. By comparison, our prior work with previous models placed ChatGPT-3.5 at the 52nd, 3rd, and 1st percentiles and ChatGPT-4.0 at the 99th, 95th, and 84th percentiles across increasing levels of training. Multimodal capability now permits successful interpretation of many figure-dependent questions, although performance declines markedly when required visual information is unavailable, and highly concordant errors remain. These findings support an increasingly promising role for LLMs as accessible, on-demand educational adjuncts for anesthesiology trainees, provided complete source material is supplied and outputs are critically reviewed.
Keywords: 
;  ;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.