Submitted:
17 July 2026
Posted:
20 July 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Materials and Methods
2.1. System Overview
2.2. Domain Specialization and Structured Output
2.3. Deterministic Clinical Tool Suite
- patient-record retrieval and clinical data access;
- cardiovascular risk assessment, including CHA2DS2-VASc, HAS-BLED, GRACE, and TIMI-related calculations;
- renal-function assessment using CKD-EPI and Cockcroft–Gault formulations, with medication-relevant interpretation;
- medication-safety checks, including drug–drug interactions and recommendation validation;
- cardiovascular treatment support, including target evaluation and dose-adjustment logic;
- acute-care and safety functions, including critical-value detection and acute coronary syndrome assessment;
- ESC guideline consultation through either a structured internal knowledge tool or the RAG module.
2.4. Retrieval-Augmented Guideline Module
2.5. Parallel Report-Generation Pipeline
2.6. Audit Trail, Numerical Consistency Checking, and Operational XAI
2.7. Evaluation Datasets
Open clinical cases.
Cardiology multiple-choice set.
2.8. Ablation Configurations
- L0, generic LLM: zero-shot prompting without cardiology specialization, patient-record access, tools, or RAG;
- L1, specialized LLM: cardiology-specific system prompting, but without external tools or RAG;
- Complete ReAct agent: domain-specialized LLM with patient-record retrieval, deterministic tools, and ESC guideline RAG.
2.9. Outcome Measures and Statistical Analysis
3. Results
3.1. Performance on Nine Open Clinical Cases
| Metric | Result |
|---|---|
| Cases evaluated | 9 |
| Correct / partial / incorrect | 9 / 0 / 0 |
| Weighted accuracy | 100% |
| 95% Wilson CI for observed accuracy | 70.1–100% |
| Normalized keyword recall | 93.5% |
| Strict keyword recall | 88.9% |
| Expected-tool utilization | 88.9% |
| Mean number of tools per query | 2.2 |
| Mean latency | 17.7 s |
| Median latency | 8.9 s |
3.2. Ablation Study
3.3. Latency and Tool-Use Behavior
3.4. Cardiology Multiple-Choice Set
| Evaluation | Result | Interpretation |
| Cardiology MCQ set | 10/10 expected options; 95% Wilson CI 72.2–100% | Controlled, in-domain, tool-guided experiment; not the official MedQA benchmark |
| Expected-tool use in MCQ set | 10/10 | Protocol-guided rather than fully autonomous selection |
| MCQ latency | mean 2.2 s; median 2.3 s | Measures the direct tool-guided pipeline, not full ReAct orchestration |
| Safety-oriented cases | 5/5 classified correct with expected tools | Demonstrative functional check, not clinical safety validation |
| Anomalous-input robustness | 4/4 handled as intended | Empty, nonsensical, missing-patient, and insufficient-data inputs |
| Correct numerical claims | 9/9 confirmed | Limited to explicit values represented in the synthetic record |
| Intentionally altered numerical claims | 5/5 detected | BNP, INR, GFR, heart rate, and potassium inconsistencies |
3.5. Error Analysis, Robustness, and Operational Traceability
4. Discussion
4.1. Limitations
4.2. Recommended Validation Pathway
- an external set of at least 100–300 cardiology cases spanning multiple disease areas, complexity levels, and missing-data patterns;
- blinded ratings by at least two independent cardiologists, with adjudication and inter-rater agreement;
- a factorial ablation separating prompt specialization, patient context, deterministic tools, structured guideline rules, document RAG, numerical checking, and action-trace formatting;
- repeated runs across model versions and at least one alternative LLM to assess stability and architecture dependence;
- retrieval-specific metrics, including context recall, context precision, citation correctness, answer faithfulness, and contradiction sensitivity;
- safety-oriented challenge cases containing ambiguous indications, contraindications, conflicting guideline statements, prompt injection, missing critical information, and out-of-distribution values;
- workflow metrics, including unnecessary tool calls, repeated calls, total token use, cost, latency, and clinician time saved or added.
5. Conclusion
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Ethics Approval and Consent to Participate.
Use of Generative AI in Manuscript Preparation.
References
- Singhal, K.; Azizi, S.; Tu, T.; Mahdavi, S.S.; Wei, J.; Chung, H.W.; Scales, N.; Tanwani, A.; Cole-Lewis, H.; Pfohl, S.; et al. Large Language Models Encode Clinical Knowledge. Nature 2023, 620, 172–180. [Google Scholar] [CrossRef] [PubMed]
- Nori, H.; King, N.; McKinney, S.M.; Carignan, D.; Horvitz, E. Capabilities of GPT-4 on Medical Challenge Problems. arXiv 2023, arXiv:2303.13375. [Google Scholar]
- Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Kuttler, H.; Lewis, M.; Yih, W.t.; Rocktaschel, T.; et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Proceedings of the Advances in Neural Information Processing Systems, 2020, Vol. 33, pp. 9459–9474.
- Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; Cao, Y. ReAct: Synergizing Reasoning and Acting in Language Models. In Proceedings of the International Conference on Learning Representations, 2023. [Google Scholar]
- Schick, T.; Dwivedi-Yu, J.; Dessi, R.; Raileanu, R.; Lomeli, M.; Hambro, E.; Zettlemoyer, L.; Cancedda, N.; Scialom, T. Toolformer: Language Models Can Teach Themselves to Use Tools. In Proceedings of the Advances in Neural Information Processing Systems, 2023; Vol. 36. [Google Scholar]
- Gaber, F.; Shaik, M.; Allega, F.; Bilecz, A.J.; Busch, F.; Goon, K.; Franke, V.; et al. Evaluating Large Language Model Workflows in Clinical Decision Support for Triage and Referral and Diagnosis. npj Digit. Med. 2025, 8, 263. [Google Scholar] [CrossRef] [PubMed]
- Li, R.; Wang, X.; Berlowitz, D.; et al. CARE-AD: A Multi-Agent Large Language Model Framework for Alzheimer’s Disease Prediction Using Longitudinal Clinical Notes. npj Digit. Med. 2025, 8, 541. [Google Scholar] [CrossRef] [PubMed]
- Nguyen, N.T.H.; et al. Med.ai ASK: An Agentic System for Biomedical Question Answering. J. Am. Med. Inform. Assoc. 2026, 33, 1134–1145. [Google Scholar] [CrossRef] [PubMed]
- Liu, Y.; Carrero, Z.I.; Jiang, X.; et al. Benchmarking Large Language Model-Based Agent Systems for Clinical Decision Tasks. npj Digit. Med. 2026, 9, 259. [Google Scholar] [CrossRef] [PubMed]
- LangChain, A.I. LangGraph: Build Stateful, Multi-Actor Applications with LLMs. Softw. Repos. Accessed. 2024. (accessed on July 2026). [Google Scholar]
- McDonagh, T.A.; Metra, M.; Adamo, M.; Gardner, R.S.; Baumbach, A.; Bohm, M.; Burri, H.; Butler, J.; Celutkiene, J.; Chioncel, O.; et al. 2021 ESC Guidelines for the Diagnosis and Treatment of Acute and Chronic Heart Failure. Eur. Heart J. 2021, 42, 3599–3726. [Google Scholar] [CrossRef] [PubMed]
- Hindricks, G.; Potpara, T.; Dagres, N.; Arbelo, E.; Bax, J.J.; Blomstrom-Lundqvist, C.; Boriani, G.; Castella, M.; Dan, G.A.; Dilaveris, P.E.; et al. 2020 ESC Guidelines for the Diagnosis and Management of Atrial Fibrillation. Eur. Heart J. 2021, 42, 373–498. [Google Scholar] [CrossRef] [PubMed]
- Byrne, R.A.; Rossello, X.; Coughlan, J.J.; Barbato, E.; Berry, C.; Chieffo, A.; Claeys, M.J.; Dan, G.A.; Dweck, M.R.; Galbraith, M.; et al. 2023 ESC Guidelines for the Management of Acute Coronary Syndromes. Eur. Heart J. 2023, 44, 3720–3826. [Google Scholar] [CrossRef] [PubMed]
- Mach, F.; Baigent, C.; Catapano, A.L.; Koskinas, K.C.; Casula, M.; Badimon, L.; Chapman, M.J.; De Backer, G.G.; Delgado, V.; Ference, B.A.; et al. 2019 ESC/EAS Guidelines for the Management of Dyslipidaemias: Lipid Modification to Reduce Cardiovascular Risk. Eur. Heart J. 2020, 41, 111–188. [Google Scholar] [CrossRef] [PubMed]
- Turpin, M.; Michael, J.; Perez, E.; Bowman, S.R. Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. Adv. Neural Inf. Process. Syst. 2023, 36. [Google Scholar]
- Wilson, E.B. Probable Inference, the Law of Succession, and Statistical Inference. J. Am. Stat. Assoc. 1927, 22, 209–212. [Google Scholar] [CrossRef]
- Zhao, L.; Liu, S.; Xin, T.; Tan, J.; Wang, X.; Li, Y.; Bian, Z.; Chen, Y.; Kong, F.; Bian, J.; et al. AI Agent in Healthcare: Applications, Evaluations, and Future Directions. npj Artif. Intell. 2026, 2, 31. [Google Scholar] [CrossRef]
- Wiens, J.; Saria, S.; Sendak, M.; Ghassemi, M.; Liu, V.X.; Doshi-Velez, F.; Jung, K.; Heller, K.; Kale, D.; Saeed, M.; et al. Do No Harm: A Roadmap for Responsible Machine Learning for Health Care. Nat. Med. 2019, 25, 1337–1340. [Google Scholar] [CrossRef] [PubMed]


| ID | Clinical scenario | Outcome | Keywords | Expected tool used | Latency (s) |
|---|---|---|---|---|---|
| EVAL_01 | Heart failure and SGLT2 inhibitor therapy | Correct | 4/4 | Yes | 13.1 |
| EVAL_02 | Atrial fibrillation and anticoagulation | Correct | 3/3 | Yes | 82.3 |
| EVAL_03 | Warfarin–ibuprofen interaction | Correct | 2/3 | No | 6.5 |
| EVAL_04 | CHA2DS2-VASc assessment | Correct | 3/3 | Yes | 8.1 |
| EVAL_05 | Renal function and GFR | Correct | 3/3 | Yes | 10.6 |
| EVAL_06 | Ischaemic heart disease and LDL target | Correct | 2/3 | Yes | 8.9 |
| EVAL_07 | Resistant hypertension | Correct | 1/2 | Yes | 13.4 |
| EVAL_08 | HAS-BLED assessment | Correct | 3/3 | Yes | 7.6 |
| EVAL_09 | Urgent STEMI management | Correct | 3/3 | Yes | 8.7 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).