Submitted:
19 November 2025
Posted:
20 November 2025
You are already at the latest version
Abstract

Keywords:
Introduction
Benchmarking Medical LLMs
Real-World Evaluation

TEHAI
Conclusions
References
- Vellum AI. LLM Leaderboard. Available from: https://www.vellum.ai/llm-leaderboard [Accessed 20 Nov 2025].
- Artificial Analysis. Models Leaderboard. Available from: https://artificialanalysis.ai/models [Accessed 20 Nov 2025].
- Balloccu S, Schmidtová P, Lango M, et al. Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs. In: Proceedings of the 19th Conference of the European Chapter of the ACL (EACL 2024). 2024. Available from: https://leak-llm.github.io/ [Accessed 20 Nov 2025].
- Yang Z, Li Y, Wang L, et al. Benchmark Data Contamination of Large Language Models: A Survey. arXiv. 2024. arXiv:2406.04244. [CrossRef]
- Bordt S, Singh A, Gokaslan A, et al. How Much Can We Forget about Data Contamination? In: International Conference on Learning Representations (ICLR 2025). 2025. Available from: https://openreview.net/forum?id=8ivK2TngIW.
- Zhang Y, Chen X, Wang M, et al. How Contaminated Is Your Benchmark? Measuring Dataset Leakage in Large Language Models with Kernel Divergence. arXiv. 2025. arXiv:2502.00678. [CrossRef]
- Ashfri NS, Wijaya R, Chen WJ. Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation. arXiv. 2025. arXiv:2505.24263. [CrossRef]
- Zhou Z, Zhang J, Wang Y, et al. Benchmarking Benchmark Leakage in Large Language Models. arXiv. 2024. arXiv:2404.18824. [CrossRef]
- Topsakal O, Edell CJ, Harper JB. Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard. arXiv. 2024. arXiv:2407.07796. [CrossRef]
- Topsakal O, Harper JB. Benchmarking Large Language Model (LLM) Performance for Game Playing via Tic-Tac-Toe. Electronics. 2024;13(8):1532. [CrossRef]
- Intuition Labs. Large Language Model Benchmarks – Life Sciences Overview. Available from: https://intuitionlabs.ai/articles/large-language-model-benchmarks-life-sciences-overview [Accessed 20 Nov 2025].
- Emergent Mind. Medical LLM Benchmarks. Available from: https://www.emergentmind.com/topics/medical-llm-benchmarks [Accessed 20 Nov 2025].
- Shool S, Adimi S, Amleshi RS, et al. A systematic review of large language model (LLM) evaluations in clinical medicine. BMC Med Inform Decis Mak. 2025;25(1):117. [CrossRef]
- Budler LC, Chen H, Chen A, Topaz M, Tam W, Bian J, Stiglic G. A brief review on benchmarking for large language models evaluation in healthcare. WIREs Data Mining Knowl Discov. 2025;15(2):e70010. [CrossRef]
- Wang H, Liu J, Zhang Y, et al. Large Language Models in Healthcare: A Comprehensive Benchmark. medRxiv. 2024. [CrossRef]
- Zhang M, Chen L, Wang Q, et al. LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation. arXiv. 2025. arXiv:2506.04078. [CrossRef]
- Omar M, Agbareia R, Glicksberg BS, et al. Benchmarking the Confidence of Large Language Models in Answering Clinical Questions: Cross-sectional Evaluation Study. JMIR Med Inform. 2025;13:e66917. [CrossRef]
- Mutisya J, Kamau P, Ochieng D, et al. Mind the Gap: Evaluating the Representativeness of Quantitative Medical Language Reasoning LLM Benchmarks for African Disease Burdens. arXiv. 2025. arXiv:2507.16322. [CrossRef]
- Moëll B, Hertzberg L, Aronsson E, et al. Swedish Medical LLM Benchmark: development and evaluation of a framework for assessing large language models in the Swedish medical domain. Front Artif Intell. 2025;10:1557920. [CrossRef]
- Li Y, Zhang H, Wang C, et al. Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions. arXiv. 2024. arXiv:2402.18060. [CrossRef]
- Qiu P, Wu C, Liu S, et al. Quantifying the reasoning abilities of LLMs on clinical cases. Nat Commun. 2025;16:9799. [CrossRef]
- Goh E, Gallo R, Hom J, et al. Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Netw Open. 2024;7(10):e2440969. [CrossRef]
- Chen S, Fiscella K, Rucci A, et al. The effect of using a large language model to respond to patient messages in primary care. Lancet Digit Health. 2024. [CrossRef]
- Laverde N, et al. Integrating large language model-based agents into a virtual patient chatbot for clinical anamnesis training. 2025. Available from: https://www.sciencedirect.com/science/article/pii/S2001037025001850. [CrossRef]
- Gao C, Li N, Li M, et al. Large language models empowered agent-based modeling and simulation. Humanit Soc Sci Commun. 2024. [CrossRef]
- Reddy S, Rogers W, Makinen V-P, et al. Evaluation framework to guide implementation of AI systems into healthcare settings. BMJ Health Care Inform. 2021;28(1):e100444. [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).