Submitted:
07 August 2026
Posted:
07 August 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Data for EHR Foundation Models
2.1. Data Landscape

| Category | Dataset | Size | EHR variables | Other modalities | Access | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| Codes | Labs | Vitals | Notes | Image | Omics | Wearable | ||||
| Integrated Health System EHR | STARR [15] | ∼3.4M | ✓ | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | Institutional |
| NYU Langone EHR [43] | not reported | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | Institutional | |
| SIPPS Dataset [44] | ∼2k | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | Institutional | |
| Mount Sinai Data Warehouse [45] | ∼12M | ✓ | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | Institutional | |
| Mayo Clinic Platform_Discover [46] | ∼10M | ✓ | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | Commercial | |
| SickKids (Hospital for Sick Children) [47] | ∼1.9M | ✓ | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | Institutional | |
| UCSF Clinical Data Warehouse [18] | ∼4.3M | ✓ | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | Institutional | |
| Mass General Brigham (RPDR) [16] | ∼7M | ✓ | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | Institutional | |
| VA EHR (VistA / CDW) [17] | ∼25M | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | Gov restricted | |
| DoD Military Health System (MDR/MHS) [48] | ∼9.6M | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | Gov restricted | |
| HCA Healthcare [49] | ∼47M encounters/year | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | Partnership | |
| Commercial RWD Platforms | Truveta Data [19] | 130M | ✓ | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | Partnership |
| Cerner Real-World Data (now Oracle Health) [20] | ∼117M | ✓ | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | Commercial | |
| IBM Explorys (now Merative) [50] | ∼64M | ✓ | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | Commercial | |
| Flatiron Health [21] | ∼5M | ✓ | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | Partnership | |
| HealthVerity [51] | ∼330M | ✓ | ✓ | ✕ | ✕ | ✕ | ✕ | ✕ | Commercial | |
| Tempus [22] | ∼8.5M records | ✓ | ✓ | ✕ | ✓ | ✓ | ✓ | ✕ | Partnership | |
| IQVIA Real-World Data [52] | ∼1.2B | ✓ | ✓ | ✕ | ✕ | ✕ | ✕ | ✕ | Commercial | |
| Veradigm (Allscripts/Practice Fusion) [53] | ∼80M | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | Commercial | |
| Komodo Health [54] | ∼330M | ✓ | ✓ | ✕ | ✕ | ✕ | ✕ | ✕ | Commercial | |
| MSK-CHORD (MSKCC) [55] | ∼1.7M | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | Institutional | |
| Federated Research Networks | PCORnet [23] | ∼100M | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | ✕ | Partnership |
| TriNetX [27] | ∼250M | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | ✕ | Commercial | |
| OHDSI Network [24] | ∼810M | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | ✕ | Partnership | |
| OpenSAFELY [26] | ∼58M | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | Gov restricted | |
| Epic Cosmos [56] | ∼300M | ✓ | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | Institutional | |
| OneFlorida+ Data Trust [57] | ∼26M | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | Academic app. | |
| N3C (National COVID Cohort Collaborative) [25] | ∼22.8M | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | Academic app. | |
| Sentinel System (FDA) [58] | 540M | ✓ | ✓ | ✕ | ✕ | ✕ | ✕ | ✕ | Gov restricted | |
| HCSRN (Health Care Systems Research Network) [59] | ∼25M | ✓ | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | Academic app. | |
| ICU Research Repositories | MIMIC-III [28] | ∼47k | ✓ | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | Open Public |
| MIMIC-IV [29] | ∼365k | ✓ | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | Open Public | |
| eICU-CRD [30] | ∼139k | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | Open Public | |
| AmsterdamUMCdb [60] | ∼20k | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | Open Public | |
| HiRID [31] | ∼34k records | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | ✕ | Open Public | |
| SICdb (Salzburg ICU) [61] | ∼27k records | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | ✕ | Open Public | |
| Primary Care Registries | CPRD (UK) [32] | ∼21M | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | Academic app. |
| THIN (The Health Improvement Network) [33] | ∼17M | ✓ | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | Commercial | |
| EHR-Linked Biobanks | UK Biobank [34] | ∼500k | ✓ | ✓ | ✓ | ✕ | ✓ | ✓ | ✓ | Academic app. |
| All of Us (NIH) [35] | ∼848k | ✓ | ✓ | ✓ | ✕ | ✕ | ✓ | ✓ | Academic app. | |
| MVP (Million Veteran Program) [62] | ∼1M | ✓ | ✓ | ✕ | ✕ | ✕ | ✓ | ✕ | Gov restricted | |
| FinnGen [36] | ∼500k | ✓ | ✓ | ✕ | ✕ | ✕ | ✓ | ✕ | Academic app. | |
| China Kadoorie Biobank (CKB) [63] | ∼512k | ✓ | ✓ | ✕ | ✕ | ✕ | ✓ | ✕ | Academic app. | |
| BioBank Japan (BBJ) [64] | ∼200k | ✓ | ✓ | ✕ | ✕ | ✕ | ✓ | ✕ | Academic app. | |
| KPRB (Kaiser Permanente Research Bank) [65] | ∼440k | ✓ | ✓ | ✓ | ✕ | ✕ | ✓ | ✕ | Institutional | |
| Danish National Biobank [66] | ∼3M | ✓ | ✓ | ✕ | ✕ | ✕ | ✓ | ✕ | Academic app. | |
| BioVU (Vanderbilt) [37] | ∼300k | ✓ | ✓ | ✓ | ✕ | ✕ | ✓ | ✕ | Institutional | |
| eMERGE (Electronic Medical Records and Genomics) [38] | ∼1.5M | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | Partnership | ||
| Administrative Claims Databases | Medicare [39] | ∼65M | ✓ | ✕ | ✕ | ✕ | ✕ | ✕ | ✕ | Gov restricted |
| Medicaid (T-MSIS/MAX) [40] | ∼80M | ✓ | ✕ | ✕ | ✕ | ✕ | ✕ | ✕ | Gov restricted | |
| MarketScan (Merative) [41] | ∼273M | ✓ | ✕ | ✕ | ✕ | ✕ | ✕ | ✕ | Commercial | |
| Optum Clinformatics / Market Clarity [42] | ∼100M | ✓ | ✓ | ✕ | ✕ | ✕ | ✕ | ✕ | Commercial | |
| HCUP-NIS (National Inpatient Sample) [67] | 7–8M records/year | ✓ | ✕ | ✕ | ✕ | ✕ | ✕ | ✕ | Academic app. | |
2.2. Data Characteristics and Challenges
3. Model Architectures and Training Strategies
3.1. Structured-Sequence Models
3.1.1. From Representation Learning to Generative Forecasting
3.1.2. From Discrete Ordering to Continuous-Time Risk Modeling
3.1.3. Scaling to Long Clinical Histories
3.2. Clinical Language Models
3.2.1. Raw Clinical Note Modeling
3.2.2. Textualization of Structured Data
3.2.3. Structuring Unstructured Data
3.3. Multimodal EHR Foundation Models
3.3.1. Multimodal Representation, Alignment and Basic Fusion
3.3.2. Knowledge-Guided and Biologically Grounded Learning
4. Downstream Tasks
4.1. Task Families and Problem Formulations
4.1.1. Clinical Outcome Prediction
4.1.2. Patient State Understanding Tasks
4.1.3. EHR-Grounded Generative and Interactive Tasks
4.1.4. Simulation and Synthetic Data Tasks
4.1.5. Causal Inference and Counterfactual Reasoning
4.2. Metrics for Downstream Tasks
5. Toward Clinical Implementation: Barriers and Requirements
5.1. Clinical Applications
5.2. Intended Use and Risk
5.3. Transparency, Reporting, and Lifecycle Documentation
5.4. Reliability Barriers in Real-World Settings
5.5. Verification, Interpretability, and Clinician Oversight
5.6. Governance, Privacy, and Equity
5.7. Deployment, Monitoring, and Revalidation
6. Conclusion
Acknowledgments
Conflicts of Interest
References
- Kristiina Häyrinen, Kaija Saranto, and Pirkko Nykänen. Definition, structure, content, use and impacts of electronic health records: a review of the research literature. International journal of medical informatics, 77(5):291–304, 2008. [CrossRef]
- Guthrie S Birkhead, Michael Klompas, and Nirav R Shah. Uses of electronic health records for public health surveillance to advance public health. Annual review of public health, 36(1):345–359, 2015. [CrossRef]
- Denis Agniel, Isaac S Kohane, and Griffin M Weber. Biases in electronic health record data due to processes within the healthcare system: retrospective observational study. Bmj, 361, 2018.
- Inci M Baytas, Cao Xiao, Xi Zhang, Fei Wang, Anil K Jain, and Jiayu Zhou. Patient subtyping via time-aware lstm networks. In Proceedings of the 23rd ACM SIGKDD international conference on knowledge discovery and data mining, pages 65–74, 2017.
- Yikuan Li, Shishir Rao, José Roberto Ayala Solares, Abdelaali Hassaine, Rema Ramakrishnan, Dexter Canoy, Yajie Zhu, Kazem Rahimi, and Gholamreza Salimi-Khorshidi. Behrt: transformer for electronic health records. Scientific reports, 10(1):7155, 2020. [CrossRef]
- Michael Moor, Oishi Banerjee, Zahra Shakeri Hossein Abad, Harlan M Krumholz, Jure Leskovec, Eric J Topol, and Pranav Rajpurkar. Foundation models for generalist medical artificial intelligence. Nature, 616(7956):259–265, 2023. [CrossRef]
- Lin Lawrence Guo, Jason Fries, Ethan Steinberg, Scott Lanyon Fleming, Keith Morse, Catherine Aftandilian, Jose Posada, Nigam Shah, and Lillian Sung. A multi-center study on the adaptability of a shared foundation model for electronic health records. NPJ digital medicine, 7(1):171, 2024. [CrossRef]
- Ziad Obermeyer, Brian Powers, Christine Vogeli, and Sendhil Mullainathan. Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464):447–453, 2019. [CrossRef]
- Ben Van Calster, Maarten van Smeden, Wouter van Amsterdam, Maarten Coemans, Laure Wynants, and Ewout W Steyerberg. The enemies of reliable and useful clinical prediction models: a review of statistical and scientific challenges. Annual Review of Statistics and Its Application, 13, 2025.
- Miguel A Hernán, Wei Wang, and David E Leaf. Target trial emulation: a framework for causal inference from observational data. Jama, 328(24):2446–2447, 2022. [CrossRef]
- Zeljko Kraljevic, Dan Bean, Anthony Shek, et al. Foresight—a generative pretrained transformer for modelling of patient timelines using electronic health records: a retrospective modelling study. The Lancet Digital Health, 6(4):e281–e290, 2024a. [CrossRef]
- Ethan Steinberg, Jason Alan Fries, Yizhe Xu, and Nigam Shah. MOTOR: A time-to-event foundation model for structured medical records. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=NialiwI2V6.
- Florian Markowetz. All models are wrong and yours are useless: making clinical prediction models impactful for patients. NPJ precision oncology, 8(1):54, 2024. [CrossRef]
- Michael Wornow, Yizhe Xu, Rahul Thapa, Birju Patel, Ethan Steinberg, Scott Fleming, Michael A Pfeffer, Jason Fries, and Nigam H Shah. The shaky foundations of large language models and foundation models for electronic health records. npj digital medicine, 6(1):135, 2023a. [CrossRef]
- Somalee Datta, Jose Posada, Garrick Engstrom, Janos Hajagos, Matthew Moran, Brett Beaulieu-Jones, Steven Hershman, Auston Bostwick, Jennifer Hughes, Adam Wilcox, et al. A new paradigm for accelerating clinical data science at Stanford Medicine. arXiv preprint arXiv:2003.10534, 2020.
- Mass General Brigham Research Computing. Research patient data registry (rpdr), 2025. URL https://rc.partners.org/about/who-we-are-risc/research-patient-data-registry. Accessed: 2026-03-10.
- Lori E Price, Kate Shea, and Sheila Gephart. Prevalence and costs of chronic conditions in the VA health care system. Medical Care Research and Review, 72(5):509–530, 2015.
- University of California, San Francisco. Ucsf clinical data for research. https://data.ucsf.edu/research/ucsf-data, 2026. Accessed: 2026-04-29.
- Truveta Research. De-identified EHR data at scale: Truveta. White Paper, 2025.
- Louis Ehwerhemuepha, Kimberly Carlson, Ryan Moog, Ben Bondurant, Cheryl Akridge, Tatiana Moreno, Gary Gasperino, and William Feaster. Cerner real-world data (crwd)-a de-identified multicenter electronic health records database. Data in brief, 42:108120, 2022. [CrossRef]
- Tamara Snow, Jeremy Snider, Leah Comment, Stella Stergiopoulos, Virginia Fisher, Margaret McCusker, and Cheryl Cho-Phan. Comparison of population characteristics in real-world clinical oncology databases in the us: Flatiron health-foundation medicine clinico-genomic databases, flatiron health research databases, and the national cancer institute seer population-based cancer registry. medRxiv, pages 2023–01, 2023.
- Tempus AI, Inc. Real-world data | tempus. https://www.tempus.com/life-sciences/real-world-data/, 2025. Accessed: 2026-04-29.
- Laura Goettinger Qualls, Thomas A Phillips, Bradley G Hammill, James Topping, Darcy M Louzao, Jeffrey S Brown, Lesley H Curtis, and Keith Marsolo. Evaluating foundational data quality in the national patient-centered clinical research network (pcornet®). Egems, 6(1):3, 2018. [CrossRef]
- George Hripcsak, Jon D Duke, Nigam H Shah, Christian G Reich, Vojtech Huser, Martijn J Schuemie, Marc A Suchard, Rae Woong Park, Ian Chi Kei Wong, Peter R Rijnbeek, et al. Observational health data sciences and informatics (ohdsi): opportunities for observational researchers. Studies in health technology and informatics, 216:574, 2015.
- Melissa A Haendel, Christopher G Chute, Tellen D Bennett, David A Eichmann, Justin Guinney, Warren A Kibbe, Philip R O Payne, Emily R Pfaff, Peter N Robinson, Joel H Saltz, Heidi Spratt, Christine Suver, John Wilbanks, Adam B Wilcox, Amy E Williams, Chunhua Wu, Sunghwan Yoo, Xiaohan Zhang, Clair Blacketer, Russ Bradford, et al. The national covid cohort collaborative (n3c): rationale, design, infrastructure, and deployment. Journal of the American Medical Informatics Association, 28(3):427–443, 2021. [CrossRef]
- Linda Nab, Andrea L Schaffer, William Hulme, Nicholas J DeVito, Iain Dillingham, Milan Wiedemann, Colm D Andrews, Helen Curtis, Louis Fisher, Amelia Green, et al. Opensafely: A platform for analysing electronic health records designed for reproducible research. Pharmacoepidemiology and drug safety, 33(6):e5815, 2024. [CrossRef]
- Umit Topaloglu and Matvey B Palchuk. Using a federated network of real-world data to optimize clinical trials operations. JCO clinical cancer informatics, 2:1–10, 2018. [CrossRef]
- Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. MIMIC-III, a freely accessible critical care database. Scientific Data, 3(1):1–9, 2016. [CrossRef]
- Alistair EW Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Shammout, Steven Horng, Tom J Pollard, Sicheng Hao, Benjamin Moody, Brian Gow, et al. MIMIC-IV, a freely accessible electronic health record dataset. Scientific Data, 10(1):1, 2023. [CrossRef]
- Tom J Pollard, Alistair EW Johnson, Jesse D Raffa, Leo A Celi, Roger G Mark, and Omar Badawi. The eICU Collaborative Research Database, a freely available multi-center database for critical care research. Scientific Data, 5(1):1–13, 2018. [CrossRef]
- Stephanie L Hyland, Martin Faltys, Matthias Hüser, Xinrui Lyu, Thomas Gumbsch, Cristóbal Esteban, Christian Bock, Max Horn, Michael Moor, Bastian Rieck, et al. Early prediction of circulatory failure in the intensive care unit using machine learning. Nature Medicine, 26(3):364–373, 2020. [CrossRef]
- Emily Herrett, Aisling M Gallagher, Krishnan Bhaskaran, Harriet Forbes, Rohini Mathur, Tjeerd van Staa, and Liam Smeeth. Data resource profile: Clinical Practice Research Datalink (CPRD). International Journal of Epidemiology, 44(3):827–836, 2015. [CrossRef]
- Betina T Blak, Mary Thompson, Hassy Dattani, and Alison Bourke. Generalisability of The Health Improvement Network (THIN) database: demographics, chronic disease prevalence and mortality rates. Informatics in Primary Care, 19(4):251–255, 2011. [CrossRef]
- Cathie Sudlow, John Gallacher, Naomi Allen, Valerie Beral, Paul Burton, John Danesh, Paul Downey, Paul Elliott, Jane Green, Martin Landray, et al. UK Biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age. PLoS Medicine, 12(3):e1001779, 2015. [CrossRef]
- Andrea H Ramirez, Lina Sulieman, David J Schlueter, Alyssa Halber, Lorenzo D Botto, Maria Botello-Harbaum, et al. The All of Us Research Program: data quality, utility, and diversity. Patterns, 3(8):100570, 2022. [CrossRef]
- Mitja I Kurki, Juha Karjalainen, Priit Palta, Timo P Sipilä, Kati Kristiansson, Kati M Donner, Mary P Reeve, Hannele Laivuori, Mervi Aaltonen, Susanna Lemmelä, et al. FinnGen provides genetic insights from a well-phenotyped isolated population. Nature, 613(7944):508–518, 2023. [CrossRef]
- Dan M Roden, Jill M Pulley, Melissa A Basford, Gordon R Bernard, Ellen W Clayton, Jeffrey R Balser, and Dan R Masys. Development of a large-scale de-identified dna biobank to enable personalized medicine. Clinical Pharmacology & Therapeutics, 84(3):362–369, 2008. [CrossRef]
- Omri Gottesman, Helena Kuivaniemi, Gerard Tromp, W Andrew Faucett, Rongling Li, Teri A Manolio, Saskia C Sanderson, Joseph Kannry, Randi Zinberg, Melissa A Basford, et al. The eMERGE Network: a consortium of biorepositories linked to electronic medical records data for conducting genomic studies. BMC Medical Genomics, 6(Suppl 2):S1, 2013.
- Centers for Medicare & Medicaid Services (CMS) Medicare Claims Data. https://www.cms.gov/data-research/statistics-trends-and-reports/basic-stand-alone-medicare-claims-public-use-files, 2024. Accessed: 2025-12.
- Nick Williams, Craig S Mayer, and Vojtech Huser. Data characterization of medicaid: legacy and new data formats in the cms virtual research data center. AMIA Summits on Translational Science Proceedings, 2021:644, 2021.
- David M Adamson, Stella Chang, and Leigh G Hansen. Health research data for the real world: the marketscan databases. New York: Thompson Healthcare, page b28, 2008.
- Paul J Wallace, Nilay D Shah, Taylor Dennen, Paul A Bleicher, and William H Crown. Optum labs: building a novel node in the learning health care system. Health affairs, 33(7):1187–1194, 2014. [CrossRef]
- Lavender Yao Jiang, Xujin Chris Liu, Nima Pour Nejatian, Mustafa Nasir-Moin, Duo Wang, Anas Abidin, Kevin Eaton, Howard Antony Riina, Ilya Laufer, Paawan Punjabi, et al. Health system-scale language models are all-purpose prediction engines. Nature, 619(7969):357–362, 2023. [CrossRef]
- Angeela Acharya, Sulabh Shrestha, Anyi Chen, Joseph Conte, Sanja Avramovic, Siddhartha Sikdar, Antonios Anastasopoulos, and Sanmay Das. Clinical risk prediction using language models: benefits and considerations. Journal of the American Medical Informatics Association, 31(9):1856–1864, 2024. [CrossRef]
- Riccardo Miotto, Li Li, Brian A Kidd, and Joel T Dudley. Deep patient: an unsupervised representation to predict the future of patients from the electronic health records. Scientific reports, 6(1):26094, 2016. [CrossRef]
- Mayo Clinic Platform. Discovery – Mayo Clinic Platform, 2024. URL https://www.mayoclinicplatform.org/discover/. Accessed: 2026-04-26.
- Lin Lawrence Guo, Maryann Calligan, Emily Vettese, Sadie Cook, George Gagnidze, Oscar Han, Jiro Inoue, Joshua Lemmon, Johnson Li, Medhat Roshdi, et al. Development and validation of the sickkids enterprise-wide data in azure repository (sedar). Heliyon, 9(11), 2023. [CrossRef]
- Tracey Perez Koehlmoos, Jessica Korona-Bailey, Jared Elzey, Brandeis Marshall, and Lea A Shanley. Ethical use of big data for healthy communities and a strong nation: unique challenges for the military health system. In BMC proceedings, volume 18, page 21. Springer, 2024.
- HCA Healthcare. Our technology. https://www.hcahealthcare.com/about/our-technology, 2026. Accessed: 2026-04-29.
- Mike D Rinderknecht and Yannick Klopfenstein. Predicting critical state after covid-19 diagnosis: model development using a large us electronic health record dataset. NPJ digital medicine, 4(1):113, 2021. [CrossRef]
- William Murk, Monica Gierada, Michael Fralick, Andrew Weckstein, Reyna Klesh, and Jeremy A Rassen. Diagnosis-wide analysis of COVID-19 complications: an exposure-crossover study. CMAJ, 193(1):E10–E18, 2021.
- IQVIA. Real world data and insights, 2024. URL https://www.iqvia.com/solutions/real-world-evidence/real-world-data-and-insights. Accessed: 2026-03-10.
- Veradigm. Veradigm network ehr data solutions. https://veradigm.com/real-world-data-solutions/, 2025. Accessed: 2026-04-29.
- Komodo Health. Komodo health data: Healthcare map, 2026. URL https://www.komodohealth.com/komodo-health-data/. Accessed: 2026-03-10.
- Zekai Chen, Arda Pekis, and Kevin Brown. Building the ehr foundation model via next event prediction. arXiv preprint arXiv:2509.25591, 2025.
- Yasir Tarabichi, Adam Frees, Steven Honeywell, Courtney Huang, Andrew M Naidech, Jason H Moore, and David C Kaelber. The cosmos collaborative: a vendor-facilitated electronic health record data aggregation platform. ACI open, 5(01):e36–e46, 2021. [CrossRef]
- Elizabeth Shenkman, Myra Hurt, William Hogan, Olveen Carrasquillo, Steven Smith, Andrew Brickman, and David Nelson. Oneflorida clinical research consortium: linking a clinical and translational science institute with a community-based distributive medical education model. Academic Medicine, 93(3):451–455, 2018. [CrossRef]
- Jeffrey S Brown, Aaron B Mendelsohn, Young Hee Nam, Judith C Maro, Noelle M Cocoros, Carla Rodriguez-Watson, Catherine M Lockhart, Richard Platt, Robert Ball, Gerald J Dal Pan, et al. The us food and drug administration sentinel system: a national resource for a learning health system. Journal of the American Medical Informatics Association, 29(12):2191–2200, 2022. [CrossRef]
- Todd R Ross, Deborah Ng, Jeffrey S Brown, Richard Pardee, Mark C Hornbrook, G Hart, and John F Steiner. The hmo research network virtual data warehouse: a public data model to support collaboration. EGEMS, 2(1):1049, 2014. [CrossRef]
- Patrick J Thoral, Jan M Peppink, Ronald H Driessen, Eric JG Sijbrands, Erwin JO Kompanje, Lewis Kaplan, Heatherlee Bailey, Jozef Kesecioglu, Maurizio Cecconi, Matthew Churpek, et al. Sharing ICU patient data responsibly under the Society of Critical Care Medicine/European Society of Intensive Care Medicine Joint Data Science Collaboration: the Amsterdam University Medical Centers Database (AmsterdamUMCdb) example. Critical Care Medicine, 49(6):e563–e577, 2021. [CrossRef]
- Nikolaus Rodemund, Bernhard Wernly, Christian Jung, Crispiana Cozowicz, and Andreas Koköfer. The Salzburg Intensive Care database (SICdb): an openly available critical care dataset. Intensive Care Medicine, 49(6):700–702, 2023. [CrossRef]
- J Michael Gaziano, John Concato, Mary Brophy, Louis Fiore, Saiju Pyarajan, James Breeling, Stacey Whitbourne, Jennifer Deen, Colleen Shannon, Donald Humphries, et al. Million Veteran Program: a mega-biobank to study genetic influences on health and disease. Journal of Clinical Epidemiology, 70:214–223, 2016. [CrossRef]
- Zhengming Chen, Junshi Chen, Rory Collins, Yu Guo, Richard Peto, Fan Wu, and Liming Li. China Kadoorie Biobank of 0.5 million people: survey methods, baseline characteristics and long-term follow-up. International Journal of Epidemiology, 40(6):1652–1666, 2011. [CrossRef]
- Akiko Nagai, Makoto Hirata, Yoichiro Kamatani, Kaori Muto, Koichi Matsuda, Yutaka Kiyohara, Toshiharu Ninomiya, Akiko Tamakoshi, Zentaro Yamagata, Taisei Mushiroda, et al. Overview of the BioBank Japan Project: study design and profile. Journal of Epidemiology, 27(3S):S2–S8, 2017. [CrossRef]
- Cathy Schaefer, RPGEH GO Project Collaboration, et al. C-a3-04: the kaiser permanente research program on genes, environment and health: a resource for genetic epidemiology in adult health and aging. Clinical Medicine & Research, 9(3-4):177–178, 2011. [CrossRef]
- Kasper Laugesen, Erzsébet Sørensen, Christian Erikstrup, Søren Brunak, Thomas Werge, Henrik Ullum, and Henrik Toft Sørensen. A review of major Danish biobanks: advantages and possibilities of health research in Denmark. Clinical Epidemiology, 15:213–239, 2023. [CrossRef]
- Jonah J Stulberg and Elliott R Haut. Practical guide to surgical data sets: healthcare cost and utilization project national inpatient sample (nis). JAMA surgery, 153(6):586–587, 2018. [CrossRef]
- Zachary C Lipton, David Kale, and Randall Wetzel. Modeling missing data in clinical time series with rnns. In Machine Learning for Healthcare Conference, pages 253–270. PMLR, 2016.
- Ofir Ben Shoham and Nadav Rappoport. Cpllm: Clinical prediction with large language models. PLOS Digital Health, 3(12):e0000680, 2024. [CrossRef]
- Laila Rasmy, Yang Xiang, Ziqian Xie, Cui Tao, and Degui Zhi. Med-bert: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction. npj Digital Medicine, 4(1):86, 2021. [CrossRef]
- Michael Wornow, Rahul Thapa, Ethan Steinberg, Jason Fries, and Nigam Shah. Ehrshot: An ehr benchmark for few-shot evaluation of foundation models. Advances in Neural Information Processing Systems, 36:67125–67137, 2023b. [CrossRef]
- Chao Pang, Xinzhuo Jiang, Nishanth Parameshwar Pavinkurve, Krishna S Kalluri, Elise L Minto, Jason Patterson, Linying Zhang, George Hripcsak, Gamze Gürsoy, Noémie Elhadad, et al. Cehr-gpt: Generating electronic health records with chronological patient timelines. arXiv preprint arXiv:2402.04400, 2024.
- Shane Waxler, Paul Blazek, Davis White, Daniel Sneider, Kevin Chung, Mani Nagarathnam, Patrick Williams, Hank Voeller, Karen Wong, Matthew Swanhorst, et al. Generative medical event models improve with scale. arXiv preprint arXiv:2508.12104, 2025.
- Artem Shmatko, Alexander Wolfgang Jung, Kumar Gaurav, Søren Brunak, Laust Hvas Mortensen, Ewan Birney, Tom Fitzgerald, and Moritz Gerstung. Learning the natural history of human disease with generative transformers. Nature, pages 1–9, 2025.
- Charles Gadd, Krishna Gokhale, Aditya Acharya, Jennifer Cooper, Francesca Crowe, Leah Fitzsimmons, Thomas Jackson, Krishnarajah Nirantharakumar, Christopher Yau, and OPTIMAL collaborative. Survivehr: a competing risks, time-to-event foundation model for multiple long-term conditions from primary care electronic health records. medRxiv, pages 2025–08, 2025.
- Junyu Luo, Muchao Ye, Cao Xiao, and Fenglong Ma. Hitanet: Hierarchical time-aware attention networks for risk prediction on electronic health records. In Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining, pages 647–656, 2020.
- Adibvafa Fallahpour, Mahshid Alinoori, Wenqian Ye, Xu Cao, Arash Afkanpour, and Amrit Krishnan. Ehrmamba: Towards generalizable and scalable foundation models for electronic health records. In Stefan Hegselmann, Helen Zhou, Elizabeth Healey, Trenton Chang, Caleb Ellington, Vishwali Mhasawade, Sana Tonekaboni, Peniel Argaw, and Haoran Zhang, editors, Proceedings of the 4th Machine Learning for Health Symposium, volume 259 of Proceedings of Machine Learning Research, pages 291–307. PMLR, 15–16 Dec 2025. URL https://proceedings.mlr.press/v259/fallahpour25a.html.
- Xinsong Du, Zhengyang Zhou, Yifei Wang, Ya-Wen Chuang, Yiming Li, Richard Yang, Wenyu Zhang, Xinyi Wang, Xinyu Chen, Hao Guan, et al. Testing and evaluation of generative large language models in electronic health record applications: a systematic review. Journal of the American Medical Informatics Association, page ocaf233, 2026.
- Kexin Huang, Jaan Altosaar, and Rajesh Ranganath. Clinicalbert: Modeling clinical notes and predicting hospital readmission. arXiv preprint arXiv:1904.05342, 2019.
- Tzu-Ying Chen, Ting-Yun Huang, and Yung-Chun Chang. Using a clinical narrative-aware pre-trained language model for predicting emergency department patient disposition and unscheduled return visits. Journal of Biomedical Informatics, 155:104657, 2024. [CrossRef]
- Yikuan Li, Ramsey M Wehbe, Faraz S Ahmad, Hanyin Wang, and Yuan Luo. Clinical-longformer and clinical-bigbird: Transformers for long clinical sequences. arXiv preprint arXiv:2201.11838, 2022.
- William R Small, Ryan J Crowley, Chloe Pariente, Jeff Zhang, Kevin P Eaton, Lavender Yao Jiang, Eric Oermann, and Yindalon Aphinyanaphongs. Enhancing the prediction of hospital discharge disposition with extraction-based language model classification. npj Health Systems, 3(1):4, 2026. [CrossRef]
- Yinghao Zhu, Junyi Gao, Zixiang Wang, Weibin Liao, Xiaochen Zheng, Lifang Liang, Miguel O Bernabeu, Yasha Wang, Lequan Yu, Chengwei Pan, et al. Clinicrealm: Re-evaluating large language models with conventional machine learning for non-generative clinical prediction tasks. arXiv preprint arXiv:2407.18525, 2024.
- Parvati Naliyatthaliyazchayil, Raajitha Muthyala, Judy Wawira Gichoya, and Saptarshi Purkayastha. Evaluating the reasoning capabilities of large language models for medical coding and hospital readmission risk stratification: Zero-shot prompting approach. Journal of medical Internet research, 27:e74142, 2025. [CrossRef]
- Hejie Cui, Zhuocheng Shen, Jieyu Zhang, Hui Shao, Lianhui Qin, Joyce C Ho, and Carl Yang. Llms-based few-shot disease predictions using ehr: A novel approach combining predictive agent reasoning and critical agent instruction. In AMIA Annual Symposium Proceedings, volume 2024, page 319, 2025.
- Nikita Makarov, Maria Bordukova, Papichaya Quengdaeng, Daniel Garger, Raul Rodriguez-Esteban, Fabian Schmich, and Michael P Menden. Large language models forecast patient health trajectories enabling digital twins. npj Digital Medicine, 8(1):588, 2025. [CrossRef]
- Zhenbang Wu, Anant Dadu, Mike Nalls, Faraz Faghri, and Jimeng Sun. Instruction tuning large language models to understand electronic health records. Advances in neural information processing systems, 37:54772–54786, 2024. [CrossRef]
- Zeljko Kraljevic, Joshua Au Yeung, Daniel Bean, James Teo, and Richard J Dobson. Large language models for medical forecasting–foresight 2. arXiv preprint arXiv:2412.10848, 2024b.
- Yusheng Liao, Chaoyi Wu, Junwei Liu, Shuyang Jiang, Pengcheng Qiu, Haowen Wang, Yun Yue, Shuai Zhen, Jian Wang, Qianrui Fan, et al. Ehr-r1: A reasoning-enhanced foundational language model for electronic health record analysis. arXiv preprint arXiv:2510.25628, 2025.
- Zachariah Zhang, Jingshu Liu, and Narges Razavian. Bert-xml: Large scale automated icd coding using bert pretraining. arXiv preprint arXiv:2006.03685, 2020.
- Weimin Lyu, Xinyu Dong, Rachel Wong, Songzhu Zheng, Kayley Abell-Hart, Fusheng Wang, and Chao Chen. A multimodal transformer: Fusing clinical notes with structured ehr data for interpretable in-hospital mortality prediction. In AMIA Annual Symposium Proceedings, volume 2022, page 719, 2023.
- Sanjib Raj Pandey, Joy Dooshima Tile, and Mahdi Maktab Dar Oghaz. Predicting 30-day hospital readmissions using clinicalt5 with structured and unstructured electronic health records. PLoS One, 20(9):e0328848, 2025. [CrossRef]
- Simon A Lee, Sujay Jain, Alex Chen, Kyoka Ono, Arabdha Biswas, Ákos Rudas, Jennifer Fang, and Jeffrey N Chiang. Clinical decision support using pseudo-notes from multiple streams of ehr data. npj Digital Medicine, 8(1):394, 2025. [CrossRef]
- Mohammad Al Olaimat and Serdar Bozdag. Caat-ehr: Cross-attentional autoregressive transformer for multimodal electronic health record embeddings. arXiv preprint arXiv:2501.18891, 2025.
- Yuewen Sun, Lingjing Kong, Guangyi Chen, Loka Li, Gongxu Luo, Zijian Li, Yixuan Zhang, Yujia Zheng, Mengyue Yang, Petar Stojanov, et al. Causal representation learning from multi-modal biomedical observations. ArXiv, pages arXiv–2411, 2025.
- Julián N Acosta, Guido J Falcone, Pranav Rajpurkar, and Eric J Topol. Multimodal biomedical ai. Nature medicine, 28(9):1773–1784, 2022. [CrossRef]
- Shih-Cheng Huang, Anuj Pareek, Saeed Seyyedi, Imon Banerjee, and Matthew P Lungren. Fusion of medical imaging and electronic health records using deep learning: a systematic review and implementation guidelines. NPJ digital medicine, 3(1):136, 2020. [CrossRef]
- Weijieying Ren, Jingxi Zhu, Zehao Liu, Tianxiang Zhao, and Vasant Honavar. A comprehensive survey of electronic health record modeling: From deep learning approaches to large language models. arXiv preprint arXiv:2507.12774, 2025a.
- Xiaorui Su, Shvat Messica, Yepeng Huang, Ruth Johnson, Lukas Fesser, Shanghua Gao, Faryad Sahneh, and Marinka Zitnik. Multimodal medical code tokenizer. arXiv preprint arXiv:2502.04397, 2025.
- Meiling Wang, Wei Shao, Shuo Huang, and Daoqiang Zhang. Hypergraph-regularized multimodal learning by graph diffusion for imaging genetics based alzheimer’s disease diagnosis. Medical image analysis, 89:102883, 2023. [CrossRef]
- Benedikt Boecking, Naoto Usuyama, Shruthi Bannur, Daniel C Castro, Anton Schwaighofer, Stephanie Hyland, Maria Wetscherek, Tristan Naumann, Aditya Nori, Javier Alvarez-Valle, et al. Making the most of text semantics to improve biomedical vision–language processing. In European conference on computer vision, pages 1–21. Springer, 2022.
- Tin Lai. Interpretable medical imagery diagnosis with self-attentive transformers: a review of explainable ai for health care. BioMedInformatics, 4(1):113–126, 2024. [CrossRef]
- Radha Ambalavanan, R Sterling Snead, Julia Marczika, Gideon Towett, Alex Malioukis, and Mercy Mbogori-Kairichi. Ontologies as the semantic bridge between artificial intelligence and healthcare. Frontiers in Digital Health, 7:1668385, 2025. [CrossRef]
- Jonathan Amar, Edward Liu, Alessandra Breschi, Liangliang Zhang, Pouya Kheradpour, Sylvia Li, Lisa Soleymani Lehmann, Alessandro Giulianelli, Matt Edwards, Yugang Jia, et al. Integrating genomics into multimodal ehr foundation models. bioRxiv, pages 2025–10, 2025.
- Zhichao Yang, Avijit Mitra, Weisong Liu, Dan Berlowitz, and Hong Yu. Transformehr: transformer-based encoder-decoder generative model to enhance prediction of disease outcomes using electronic health records. Nature communications, 14(1):7857, 2023. [CrossRef]
- Junmo Kim, Joo Seong Kim, Ji-Hyang Lee, Min-Gyu Kim, Taehyun Kim, Chaeeun Cho, Rae Woong Park, and Kwangsoo Kim. Pretrained patient trajectories for adverse drug event prediction using common data model-based electronic health records. Communications Medicine, 5(1):232, 2025. [CrossRef]
- Yukang Jiang, Bingxin Zhao, Xiaopu Wang, Borui Tang, Huiyang Peng, Zidan Luo, Yue Shen, Zheng Wang, Zhiwen Jiang, Jie Wang, et al. Ukb-mdrmf: a multi-disease risk and multimorbidity framework based on uk biobank data. Nature Communications, 16(1):3767, 2025. [CrossRef]
- Su Xian, Monika E Grabowska, Iftikhar J Kullo, Yuan Luo, Jordan W Smoller, Wei-Qi Wei, Gail Jarvik, Sean Mooney, and David Crosslin. Language-model-based patient embedding using electronic health records facilitates phenotyping, disease forecasting, and progression analysis. Research Square, pages rs–3, 2024.
- Emily Alsentzer, Matthew J Rasmussen, Romy Fontoura, Alexis L Cull, Brett Beaulieu-Jones, Kathryn J Gray, David W Bates, and Vesela P Kovacheva. Zero-shot interpretable phenotyping of postpartum hemorrhage using large language models. NPJ digital medicine, 6(1):212, 2023. [CrossRef]
- Jiajun Qiu, Yao Hu, Li Li, Abdullah Mesut Erzurumluoglu, Ingrid Braenne, Charles Whitehurst, Jochen Schmitz, Jatin Arora, Boris Alexander Bartholdy, Shrey Gandhi, et al. Deep representation learning for clustering longitudinal survival data from electronic health records. Nature Communications, 16(1):2534, 2025. [CrossRef]
- Augustin Toma, Patrick R Lawler, Jimmy Ba, Rahul G Krishnan, Barry B Rubin, and Bo Wang. Clinical camel: An open expert-level medical language model with dialogue-based knowledge encoding. arXiv preprint arXiv:2305.12031, 2023.
- Zhichao Yang, Avijit Mitra, Sunjae Kwon, and Hong Yu. Clinicalmamba: A generative clinical language model on longitudinal clinical notes. In Proceedings of the 6th Clinical Natural Language Processing Workshop, pages 54–63, 2024.
- Weijieying Ren, Tianxiang Zhao, Lei Wang, Tianchun Wang, and Vasant Honavar. Diallms: Ehr enhanced clinical conversational system for clinical test recommendation and diagnosis prediction. arXiv preprint arXiv:2506.20059, 2025b.
- Jinsung Yoon, Michel Mizrahi, Nahid Farhady Ghalaty, Thomas Jarvinen, Ashwin S Ravi, Peter Brune, Fanyu Kong, Dave Anderson, George Lee, Arie Meir, et al. Ehr-safe: generating high-fidelity and privacy-preserving synthetic electronic health records. NPJ digital medicine, 6(1):141, 2023. [CrossRef]
- Vasileios C Pezoulas, Dimitrios I Zaridis, Eugenia Mylona, Christos Androutsos, Kosmas Apostolidis, Nikolaos S Tachos, and Dimitrios I Fotiadis. Synthetic data generation methods in healthcare: A review on open-source tools and methods. Computational and structural biotechnology journal, 23:2892–2910, 2024. [CrossRef]
- Aldren Gonzales, Guruprabha Guruswamy, and Scott R Smith. Synthetic data in health care: A narrative review. PLOS Digital Health, 2(1):e0000082, 2023. [CrossRef]
- Manlio De Domenico, Luca Allegri, Guido Caldarelli, Valeria d’Andrea, Barbara Di Camillo, Luis M Rocha, Jordan Rozum, Riccardo Sbarbati, and Francesco Zambelli. Challenges and opportunities for digital twins in precision medicine from a complex systems perspective. npj Digital Medicine, 8(1):37, 2025. [CrossRef]
- Simeone Marino, Ruth Cassidy, Joseph Nanni, Yuxuan Wang, Yipeng Liu, Mingyi Tang, Yuan Yuan, Toby Chen, Anik Sinha, Balaji Pandian, et al. Medical data sharing and synthetic clinical data generation–maximizing biomedical resource utilization and minimizing participant re-identification risks. npj Digital Medicine, 8(1):526, 2025. [CrossRef]
- Mattia Prosperi, Yi Guo, Matt Sperrin, James S Koopman, Jae S Min, Xing He, Shannan Rich, Mo Wang, Iain E Buchan, and Jiang Bian. Causal inference and counterfactual prediction in machine learning for actionable healthcare. Nature Machine Intelligence, 2(7):369–375, 2020. [CrossRef]
- Edward H Kennedy. Towards optimal doubly robust estimation of heterogeneous causal effects. Electronic Journal of Statistics, 17(2):3008–3049, 2023. [CrossRef]
- Valentyn Melnychuk, Dennis Frauen, and Stefan Feuerriegel. Causal transformer for estimating counterfactual outcomes. In International conference on machine learning, pages 15293–15329. PMLR, 2022.
- Miguel A Hernán and James M Robins. Using big data to emulate a target trial when a randomized trial is not available. American journal of epidemiology, 183(8):758–764, 2016. [CrossRef]
- Haoyang Li, Chengxi Zang, Zhenxing Xu, Weishen Pan, Suraj Rajendran, Yong Chen, and Fei Wang. Federated target trial emulation using distributed observational data for treatment effect estimation. NPJ Digital Medicine, 8(1):387, 2025. [CrossRef]
- Marc Lipsitch, Eric Tchetgen Tchetgen, and Ted Cohen. Negative controls: a tool for detecting confounding and bias in observational studies. Epidemiology, 21(3):383–388, 2010. [CrossRef]
- Tyler J VanderWeele and Peng Ding. Sensitivity analysis in observational research: introducing the e-value. Annals of internal medicine, 167(4):268–274, 2017. [CrossRef]
- Andrew J Vickers and Ford Holland. Decision curve analysis to evaluate the clinical benefit of prediction models. The Spine Journal, 21(10):1643–1648, 2021. [CrossRef]
- Ryutaro Tanno, David GT Barrett, Andrew Sellergren, Sumedh Ghaisas, Sumanth Dathathri, Abigail See, Johannes Welbl, Charles Lau, Tao Tu, Shekoofeh Azizi, et al. Collaboration between clinicians and vision–language models in radiology report generation. Nature Medicine, 31(2):599–608, 2025. [CrossRef]
- Gary S Collins, Karel GM Moons, Paula Dhiman, Richard D Riley, Andrew L Beam, Ben Van Calster, Marzyeh Ghassemi, Xiaoxuan Liu, Johannes B Reitsma, Maarten Van Smeden, et al. Tripod+ ai statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. bmj, 385, 2024. [CrossRef]
- U.S. Food and Drug Administration. Considerations for the use of artificial intelligence to support regulatory decision-making for drug and biological products: Guidance for industry and other interested parties. Draft guidance, January 2025. URL https://www.fda.gov/regulatory-information/search-fda-guidance-documents/considerations-use-artificial-intelligence-support-regulatory-decision-making-drug-and-biological. Accessed: 2026-04-26.
- Sonya Makhni, Paul Cerrato, Jose Rico, Shehzad Niazi, Jack O’Horo, Steve Peters, Vijay Shah, and John Halamka. Meeting the challenges of electronic health record (ehr) optimization. npj Digital Medicine, 2025. [CrossRef]
- Stephen Gilbert, Jakob Nikolas Kather, and Aidan Hogan. Augmented non-hallucinating large language models as medical information curators. NPJ digital medicine, 7(1):100, 2024. [CrossRef]
- Monica Agrawal, Irene Y Chen, Freya Gulamali, and Shalmali Joshi. The evaluation illusion of large language models in medicine. npj Digital Medicine, 8(1):600, 2025. [CrossRef]
- Kabir Jalal, Alexandre Charest, Xiaohuan Wu, et al. The ICD-9 to ICD-10 transition has not improved identification of rapidly progressing stage 3 and stage 4 chronic kidney disease patients: a diagnostic test study. BMC Nephrology, 25:55, 2024. [CrossRef]
- John H Holmes, James Beinlich, Mary R Boland, Kathryn H Bowles, Yong Chen, Tessa S Cook, George Demiris, Michael Draugelis, Laura Fluharty, Peter E Gabriel, et al. Why is the electronic health record so challenging for research and clinical care? Methods of information in medicine, 60(01/02):032–048, 2021. [CrossRef]
- Ke Zhu, Rima Izem, Peng Yang, Ying Yuan, Herbert Pang, Mark van der Laan, Lei Nie, Birol Emir, Pallavi Mishra-Kalyani, Hana Lee, and Shu Yang. Externally controlled trials: A review of design and borrowing through a causal lens. arXiv preprint arXiv:2605.03282, 2026.
- Shuang Zhou, Mingquan Lin, Sirui Ding, Jiashuo Wang, Canyu Chen, Genevieve B Melton, James Zou, and Rui Zhang. Explainable differential diagnosis with dual-inference large language models. npj Health Systems, 2(1):12, 2025. [CrossRef]
- James C Douglas, Yidong Gan, Ben Hachey, and Jonathan K Kummerfeld. Less is more: Explainable and efficient icd code prediction with clinical entities. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 30835–30847, 2025.
- Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B. Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel. Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21), pages 2633–2650. USENIX Association, 2021.
- Laleh Seyyed-Kalantari, Haoran Zhang, Matthew BA McDermott, Irene Y Chen, and Marzyeh Ghassemi. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nature medicine, 27(12):2176–2182, 2021. [CrossRef]
- U.S. Food and Drug Administration, Health Canada, and Medicines and Healthcare products Regulatory Agency. Good machine learning practice for medical device development: Guiding principles, October 2021. URL https://www.fda.gov/media/153486/download. Accessed: 2026-04-26.
- Cheryl D. Stults, Sien Deng, Meghan C. Martinez, Joseph Wilcox, Nina Szwerinski, Kevin H. Chen, Stephanie Driscoll, Joanna Washburn, and Veena G. Jones. Evaluation of an ambient artificial intelligence documentation platform for clinicians. JAMA Network Open, 8(5):e258614, 2025. [CrossRef]
- Matthew J Duggan, Julietta Gervase, Anna Schoenbaum, William Hanson, John T Howell, Michael Sheinberg, and Kevin B Johnson. Clinician experiences with ambient scribe technology to assist with documentation burden and efficiency. JAMA Network Open, 8(2):e2460637–e2460637, 2025. [CrossRef]
- Arjun Mahajan and Dylan Powell. Transforming healthcare delivery with conversational AI platforms. npj Digital Medicine, 8:581, 2025. [CrossRef]
- Lisa M. Koch, Christian F. Baumgartner, and Philipp Berens. Distribution shift detection for the postmarket surveillance of medical AI algorithms: a retrospective simulation study. npj Digital Medicine, 7:120, 2024. [CrossRef]
- Suzanne Bakken. Ai in health: keeping the human in the loop, 2023.
- Isabelle-Emmanuella Nogues, Jun Wen, Yucong Lin, Molei Liu, Sara K Tedeschi, Alon Geva, Tianxi Cai, and Chuan Hong. Weakly semi-supervised phenotyping using electronic health records. Journal of biomedical informatics, 134:104175, 2022. [CrossRef]
- Jun Wen, Jue Hou, Clara-Lea Bonzel, Yihan Zhao, Victor M. Castro, Vivian S. Gainer, Dana Weisenfeld, Tianrun Cai, Yuk-Lam Ho, Vidul A. Panickan, Lauren Costa, Chuan Hong, J. Michael Gaziano, Katherine P. Liao, Junwei Lu, Kelly Cho, and Tianxi Cai. Latte: Label-efficient incident phenotyping from longitudinal electronic health records. Patterns, 5(1):100906, 2024. [CrossRef]
- Nicola Rieke, Jonny Hancox, Wenqi Li, Fausto Milletari, Holger R Roth, Shadi Albarqouni, Spyridon Bakas, Mathieu N Galtier, Bennett A Landman, Klaus Maier-Hein, et al. The future of digital health with federated learning. NPJ digital medicine, 3(1):119, 2020. [CrossRef]
- Jiayi Tong, Chongliang Luo, Md Nazmul Islam, Natalie E Sheils, John Buresh, Mackenzie Edmondson, Peter A Merkel, Ebbing Lautenbach, Rui Duan, and Yong Chen. Distributed learning for heterogeneous clinical data with application to integrating covid-19 data across 230 sites. NPJ digital medicine, 5(1):76, 2022. [CrossRef]



| Subcategory | Representative task | Typical metric |
|---|---|---|
| A. Clinical outcome prediction | ||
| General clinical prediction | In-hospital mortality prediction | AUROC, AUPRC, F1 score, Brier score, Expected calibration error, Precision@k, Recall@k, Mean reciprocal rank |
| 30-day readmission prediction | ||
| Adverse drug event prediction | ||
| Next event prediction | ||
| Time-to-event modeling | Time-to-death prediction | Harrell’s C-index, Time-dependent AUROC, Integrated Brier score, Dynamic prediction error, CIF calibration error |
| Time-to-deterioration prediction | ||
| Time-to-ICU transfer prediction | ||
| Competing-risk event modeling | ||
| B. Patient state understanding tasks | ||
| Predefined Clinical State Identification | Rare disease case identification | AUROC, AUPRC, Recall at fixed FPR, Positive predictive value, F1 score, Calibration slope/intercept, Decision curve net benefit |
| Treatment responder identification | ||
| Cohort eligibility matching | ||
| High-risk phenotype detection | ||
| Latent Phenotype and Progression Modeling | Patient subgroup discovery | Silhouette coefficient, Adjusted Rand Index, Normalized mutual information, Cluster stability score, Macro-F1 score, AUROC, Weighted Cohen’s kappa |
| De novo clinical subtyping | ||
| Progression trajectory derivation | ||
| Longitudinal pattern discovery | ||
| C. EHR-grounded generative and interactive tasks | ||
| Documentation assistance | Discharge summary drafting | ROUGE-L, BLEU score, BERTScore, RadGraph F1, Entity precision/recall, Clinical concept coverage, Expert rating score |
| Radiology report generation | ||
| Clinical note generation | ||
| Chart summarization | ||
| Conversational reasoning and copilots | Natural-language EHR interaction | Top-k diagnostic accuracy, Guideline adherence rate, Evidence attribution accuracy, Task success rate, Abstention rate, Clinical safety violation rate, Expert panel rating |
| Test recommendation | ||
| Guideline-based reasoning | ||
| Clinical question answering | ||
| D. Simulation and synthetic data tasks | ||
| Long-horizon trajectory simulation | Future visit sequence generation | Sequence negative log-likelihood, Dynamic time warping, Mean absolute error, Root mean squared error, Continuous ranked probability score, Trajectory coverage |
| Disease trajectory generation | ||
| Long-term comorbidity modeling | ||
| Care pathway generation | ||
| Synthetic cohort generation | Synthetic patient generation | Hellinger distance, Maximum mean discrepancy, TSTR, AUROC/AUPRC, Membership inference attack rate, Attribute disclosure risk, Privacy–utility frontier |
| Privacy-preserving data generation | ||
| Synthetic cohort sharing | ||
| Rare event data augmentation | ||
| E. Causal inference and counterfactual reasoning | ||
| Treatment effect estimation | Average treatment effect estimation | PEHE in semi-synthetic benchmarks, Treatment-effect calibration, Agreement with trial or quasi-experimental estimates, Overlap and positivity diagnostics, Sensitivity analyses including E-values |
| Conditional treatment effects | ||
| Doubly robust estimation | ||
| Subgroup heterogeneity assessment | ||
| Counterfactual trajectory modeling | Time-varying counterfactuals | Counterfactual trajectory error in simulated settings, Time-dependent calibration, Negative-control outcome consistency, Coverage of effect intervals |
| Counterfactual digital twins | ||
| Treatment trajectory comparison | ||
| Synthetic control prediction | ||
| Policy evaluation and target trial emulation | Off-policy treatment evaluation | Policy value and regret, Importance-weighted estimator variance, Target-trial emulation agreement with trial benchmarks, Sensitivity bounds under unmeasured confounding |
| Dynamic treatment regimes | ||
| Target trial operationalization | ||
| Comparative effectiveness | ||
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).