Submitted:
15 June 2026
Posted:
29 June 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- We conduct the first cross-functional survey of LLMs in public health, covering the full pipeline from surveillance to governance with a population-health boundary test.
- We introduce a task × role taxonomy that reveals systematic gaps: the field is heavily concentrated in extraction-oriented sensing (Sensor dominates T1 and T3, which together account for over half the corpus), while anticipatory roles (Predictor, Simulator) and equity-focused auditing remain underexplored.
- We quantify structural imbalances, including that 80.7% of papers use English-only data and most depend on proprietary models, and distill four gap-driven research directions grounded in the taxonomy.
2. Background & Preliminaries
3. LLM Role Dimension
4. A taxonomy of LLMs in Public Health
4.1. Population Health Surveillance (T1)
4.2. Epidemic Forecasting & Modeling (T2)
4.3. Infodemiology & Health Communication (T3)
4.4. Health Equity, SDOH & Global Health (T4)
4.5. Population Intervention & Practice (T5)
4.6. Governance, Ethics & Policy (T6)
5. Discussion and Conclusions
Appendix A
Appendix A.1. Additional Background on Large Language Models
Appendix A.2. Survey Methodology and Literature Selection
- 1.
- Institutional attribution test: Does the work’s primary application reside within the public health system (e.g., CDC, WHO, health departments, schools of public health, epidemiological registries)?
- 2.
- Core function test: Does the core task fall within one of the recognized core public health functions—surveillance, disease prevention, health communication, health equity, population intervention, or governance—as defined by established frameworks such as the CDC’s 10 Essential Public Health Services?
- 3.
- Substitution test: If removed from the public health context, would the work retain independent significance in another domain? If yes, its primary identity is likely outside public health.
Appendix A.3. The Public Health Data Ecosystem
| Data source category | % of papers |
|---|---|
| Social media (Twitter/X, Reddit, Weibo) | 25.5 |
| Scientific literature & guidelines | 16.1 |
| Simulated / agent-based data | 11.5 |
| Clinical text (EHR notes, discharge summaries) | 9.9 |
| Official reports (WHO DONs, CDC bulletins) | 9.4 |
| Epidemiological time-series | 6.8 |
| Death certificates & verbal autopsy narratives | 4.2 |
| Policy documents (e.g., OxCGRT) | 3.6 |
| Genomic / environmental (wastewater, sequences) | 3.1 |
| Consumer reviews | 2.6 |
| Surveys & questionnaires | 7.3 |
| Language coverage | |
| English-only | 80.7 |
| Monolingual non-English | 14.1 |
| Genuinely multilingual (≥3 languages) | 5.2 |
References
- Abate, Andrea, Elisa Poncato, Maria Antonietta Barbieri, Greg Powell, Andrea Rossi, Simay Peker, Anders Hviid, Andrew Bate, and Maurizio Sessa. 2025. Off-the-shelf large language models for causality assessment of individual case safety reports: A proof-of-concept with covid-19 vaccines: A. abate et al. Drug Safety 48, 7: 805–820. [Google Scholar] [PubMed]
- Abroms, Lorien C, Christina N Wysota, Artin Yousefi, Tien-Chin Wu, and David A Broniatowski. 2025. Chatgpt-based chatbot for help quitting smoking via text messaging: An interventional study. JMIR formative research 9, e79402. [Google Scholar]
- Abroms, Lorien C, Artin Yousefi, Christina N Wysota, Tien-Chin Wu, and David A Broniatowski. 2025. Assessing the adherence of chatgpt chatbots to public health guidelines for smoking cessation: content analysis. Journal of medical Internet research 27, e66896. [Google Scholar] [CrossRef]
- Achiam, Josh, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, and et al. 2023. Gpt-4 technical report. arXiv arXiv:2303.08774. [Google Scholar]
- Acosta, Jose A. 2025. Perspective: advancing public health education by embedding ai literacy. Frontiers in digital health 7, 1584883. [Google Scholar] [CrossRef]
- Adames, Ariel Guerra, Marta Avalos, Océane Doremus, Cédric Gil-Jardiné, and Emmanuel Lagarde. 2025. Uncovering judgment biases in emergency triage: A public health approach based on large language models. Proceedings of Machine Learning Research 259, 420–439. [Google Scholar]
- Afane, Mohamed, Ying Wang, and Juntao Chen. 2025. Can llms help allocate public health resources? a case study on childhood lead testing. arXiv arXiv:2511.18239. [Google Scholar]
- Ahmad, Syed Talal, Haohui Lu, Sidong Liu, Annie Lau, Amin Beheshti, Mark Dras, and Usman Naseem. 2025. Vaxguard: A multi-generator, multi-type, and multi-role dataset for detecting llm-generated vaccine misinformation. arXiv arXiv:2503.09103. [Google Scholar]
- Al Ghadban, Yasmina, Huiqi Lu, Uday Adavi, Ankita Sharma, Sridevi Gara, Neelanjana Das, Bhaskar Kumar, Renu John, Praveen Devarsetty, and Jane E Hirst. 2023. Transforming healthcare education: Harnessing large language models for frontline health worker capacity building using retrieval-augmented generation. medRxiv, 2023–12. [Google Scholar]
- Angyal, Viola, Ádám Bertalan, Péter Domján, Helga Judit Feith, and Elek Dinya. 2025. Development of a questionnaire for assessing the use of chatgpt in primary and secondary disease prevention. Frontiers in Public Health 13, 1709611. [Google Scholar] [CrossRef]
- Anik, Anirban Saha, Xiaoying Song, Elliott Wang, Bryan Wang, Bengisu Yarimbas, and Lingzi Hong. 2025. Multi-agent retrieval-augmented framework for evidence-based counterspeech against health misinformation. arXiv arXiv:2507.07307. [Google Scholar]
- Annan, Augustine, Amanda L Eiden, Dong Wang, Jingcheng Du, Majid Rastegar-Mojarad, Varun Kumar Nomula, and Xiaoyan Wang. 2025. Evaluating large language models for sentiment analysis and hesitancy analysis on vaccine posts from social media: Qualitative study. JMIR Formative Research 9, e64723. [Google Scholar] [CrossRef]
- Aoki, Goshi, and Navid Ghaffarzadegan. 2026. Ai agents as policymakers in simulated epidemics. arXiv arXiv:2601.04245. [Google Scholar]
- Bann, David, Ed Lowther, Liam Wright, and Yevgeniya Kovalchuk. 2026. Why can’t epidemiology be automated (yet)? International Journal of Epidemiology 55, 1: dyaf210. [Google Scholar] [CrossRef] [PubMed]
- Becker, Mike, Sy Hwang, Emily Schriver, Caryn Douma, Caoimhe Duffy, Joshua Atkins, Caitlyn McShane, Jason Lubken, Asaf Hanish, John D McGreevey, and et al. 2025. Automatically identifying event reports of workplace violence and communication failures using large language models. AMIA Summits on Translational Science Proceedings; p. 74. [Google Scholar]
- Belova, Anna, Raquel A Silva, Dylan M Vorndran, and Natalie R Sampson. 2025. Using large language models to learn from recent climate change discourse in public health. PLoS One 20, 4: e0321309. [Google Scholar] [CrossRef] [PubMed]
- Bhadelia, Nahid, Ioannis Ch Paschalidis, John S Brownstein, and Britta Lassman. 2026. A beacon for novel disease threats: Leveraging artificial intelligence for informal event-based outbreak surveillance. [Google Scholar] [PubMed]
- Boatman, Dannell, Abby Starkey, Lori Acciavatti, Zachary Jarrett, Amy Allen, and Stephenie Kennedy-Rea. 2024. Using social listening for digital public health surveillance of human papillomavirus vaccine misinformation online: Exploratory study. JMIR infodemiology 4, e54000. [Google Scholar] [CrossRef]
- Bommasani, Rishi, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arber, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, and et al. 2021. On the opportunities and risks of foundation models. arXiv arXiv:2108.07258. [Google Scholar]
- Bouzoubaa, Layla, Elham Aghakhani, and Rezvaneh Rezapour. 2024. Words matter: Reducing stigma in online conversations about substance use with large language models. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing; pp. 9139–9156. [Google Scholar]
- Brown, Tom, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, and et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33, 1877–1901. [Google Scholar]
- Bui, Nhat, Giang Nguyen, Nguyen Nguyen, Bao Vo, Luan Vo, Tom Huynh, Arthur Tang, Van Nhiem Tran, Tuyen Huynh, Huy Quang Nguyen, and et al. 2025. Fine-tuning large language models for improved health communication in low-resource languages. Computer Methods and Programs in Biomedicine 263, 108655. [Google Scholar]
- Burstein, Roy, Eric Mafuta, and Joshua L Proctor. 2025. Large language models for analyzing open text in global health surveys: why children are not accessing vaccine services in the democratic republic of the congo. International health 17, 5: 843–852. [Google Scholar] [CrossRef] [PubMed]
- Carpenter, Kristy A, and Russ B Altman. 2023. Using gpt-3 to build a lexicon of drugs of abuse synonyms for social media pharmacovigilance. Biomolecules 13, 2: 387. [Google Scholar] [CrossRef] [PubMed]
- Carshon-Marsh, Ronald, Richard Wen, Thomas Kai Sze Ng, Rajeev Kamadod, Isaac Bogoch, Susan J Bondy, Theodore J Witek, and Prabhat Jha. 2026. Comparison of verbal autopsy using a large language model to biologically confirmed causes of death for malaria and other communicable diseases among children in six sub-saharan african countries. Malaria Journal 25, 1: 77. [Google Scholar] [CrossRef] [PubMed]
- Chang, Crystal T, Neha Srivathsa, Charbel Bou-Khalil, Akshay Swaminathan, Mitchell R Lunn, Kavita Mishra, Sanmi Koyejo, and Roxana Daneshjou. 2025. Evaluating anti-lgbtqia+ medical bias in large language models. PLOS Digital Health 4, 9: e0001001. [Google Scholar] [CrossRef] [PubMed]
- Chen, Cai, Shu-Le Li, Anthony D So, Yao-Yang Xu, Zhao-Feng Guo, Xinbing Wang, David W Graham, and Yong-Guan Zhu. 2025. Using large language models to assist antimicrobial resistance policy development: Integrating the environment into health protection planning. Environmental Science & Technology 59, 2: 1243–1252. [Google Scholar] [CrossRef]
- Chen, Haichao, Dian Zeng, Yiming Qin, Zeyue Fan, Faye Ng Yu Ci, David C Klonoff, John S Ji, Shuyang Zhang, Kwesi Nyan Amissah-Arthur, Michelle María Jiménez de Tavárez, and et al. 2025. Large language models and global health equity: a roadmap for equitable adoption in lmics. The Lancet Regional Health–Western Pacific 63. [Google Scholar] [CrossRef]
- Chen, Yiqun T, Tyler H McCormick, Li Liu, and Abhirup Datta. 2025. Lava: Language model assisted verbal autopsy for cause-of-death determination. arXiv arXiv:2509.09602. [Google Scholar]
- Choi, Soyeon, Kangwook Lee, Oliver Sng, and Joshua M Ackerman. 2025. Infected smallville: How disease threat shapes sociality in llm agents. arXiv arXiv:2506.13783. [Google Scholar]
- Chu, Bianca, Natansh D Modi, Bradley D Menz, Stephen Bacchi, Ganessan Kichenadasse, Catherine Paterson, Joshua G Kovoor, Imogen Ramsey, Jessica M Logan, Michael D Wiese, and et al. 2025. Generative ai’s healthcare professional role creep: a cross-sectional evaluation of publicly accessible, customised health-related gpts. Frontiers in Public Health 13, 1584348. [Google Scholar]
- Consoli, Sergio, Pietro Coletti, Peter V Markov, Lia Orfei, Indaco Biazzo, Lea Schuh, Nicolas Stefanovitch, Lorenzo Bertolini, Mario Ceresa, and Nikolaos I Stilianakis. 2025. An epidemiological knowledge graph extracted from the world health organization’s disease outbreak news. Scientific Data 12, 1: 970. [Google Scholar] [CrossRef] [PubMed]
- Consoli, Sergio, Peter Markov, Nikolaos I Stilianakis, Lorenzo Bertolini, Antonio Puertas Gallardo, and Mario Ceresa. 2024. Epidemic information extraction for event-based surveillance using large language models. In International Congress on Information and Communication Technology. Springer: pp. 241–252. [Google Scholar]
- Coutinho, Isabel, Gonçalo M Correia, Bruno Martins, Afonso Moreira, and André Peralta-Santos. 2026. Icd coding of death certificates with generative language models. PLOS Digital Health 5, 2: e0001245. [Google Scholar] [CrossRef] [PubMed]
- Dai, Haixing, Yiwei Li, Zhengliang Liu, Lin Zhao, Zihao Wu, Suhang Song, Shen Ye, Dajiang Zhu, Xiang Li, Sheng Li, and et al. 2025. Ad-autogpt: An autonomous gpt for alzheimer’s disease infodemiology. PLOS Global Public Health 5, 5: e0004383. [Google Scholar] [PubMed]
- Daluwatte, Chathuri, Alena Khromava, Yuning Chen, Laurence Serradell, Anne-Laure Chabanon, Anthony Chan-Ou-Teung, Cliona Molony, and Juhaeri Juhaeri. 2024. Application of a language model tool for covid-19 vaccine adverse event monitoring using web and social media content: Algorithm development and validation study. JMIR infodemiology 4, e53424. [Google Scholar] [CrossRef]
- Datta, Rituparna, Zihan Guan, Baltazar Espinoza, Yiqi Su, Priya Pitre, Srini Venkatramanan, Naren Ramakrishnan, and Anil Vullikanti. 2026. Agentic framework for epidemiological modeling. arXiv arXiv:2602.00299. [Google Scholar]
- De Angelis, Luigi, Francesco Baglivo, Guglielmo Arzilli, Gaetano Pierpaolo Privitera, Paolo Ferragina, Alberto Eugenio Tozzi, and Caterina Rizzo. 2023. Chatgpt and the rise of large language models: the new ai-driven infodemic threat in public health. Frontiers in public health 11, 1166120. [Google Scholar]
- Deiner, Michael S, Vlad Honcharov, Jiawei Li, Tim K Mackey, Travis C Porco, and Urmimala Sarkar. 2024. Large language models can enable inductive thematic analysis of a social media corpus in a single prompt: human validation study. JMIR infodemiology 4, 1: e59641. [Google Scholar] [CrossRef] [PubMed]
- DENG, OU, and Qun Jin. 2025. Position: Public health systems should embrace a multi-layered epidemic early-warning with llm agents and local knowledge enhancement. [Google Scholar] [CrossRef] [PubMed]
- Devlin, Jacob, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies volume 1: 4171–4186. [Google Scholar] [CrossRef]
- Dhanuka, Utsav, Soham Poddar, and Saptarshi Ghosh. 2025. Utilising large language models for generating effective counter arguments to anti-vaccine tweets. arXiv arXiv:2510.16359. [Google Scholar]
- Dixit, Narendra M. 2025. Leveraging large language models for pandemic preparedness: Computational epidemiology. Nature Computational Science 5, 6: 438–439. [Google Scholar] [PubMed]
- Du, Hongru, Yang Zhao, Jianan Zhao, Shaochong Xu, Xihong Lin, Yiran Chen, Lauren M Gardner, and Hao ‘Frank’ Yang. 2025. Advancing real-time infectious disease forecasting using large language models. Nature Computational Science 5, 6: 467–480. [Google Scholar] [CrossRef] [PubMed]
- Dudley, Carson, Reiden Magdaleno, Christopher Harding, Ananya Sharma, Emily Martin, and Marisa Eisenberg. 2025. Mantis: A simulation-grounded foundation model for disease forecasting. arXiv arXiv:2508.12260. [Google Scholar]
- Ekanayake, Vinu, Md Sultan Al Nahian, and Ramakanth Kavuluru. 2025. Mining social media for barriers to opioid recovery with llms. Proceedings of the Second Workshop on Patient-Oriented Language Processing (CL4Health); pp. 83–99. [Google Scholar]
- Elmitwalli, Sarah, John Mehegan, Allen Gallagher, and Rasha Alebshehy. 2024. Enhancing sentiment and intent analysis in public health via fine-tuned large language models on tobacco and e-cigarette-related tweets. Frontiers in Big Data 7, 1501154. [Google Scholar] [CrossRef] [PubMed]
- Espinosa, Laura, Djilani Kebaili, Sergio Consoli, Kyriaki Kalimeri, Yelena Mejova, and Marcel Salathé. 2025. Open-source solution for evaluation and benchmarking of large language models for public health. medRxiv, 2025–03. [Google Scholar]
- Espinosa, Laura, and Marcel Salathé. 2024. Use of large language models as a scalable approach to understanding public health discourse. PLOS Digital Health 3, 10: e0000631. [Google Scholar] [CrossRef] [PubMed]
- Funnell, Arthur J, Panayiotis Petousis, Fabrice Harel-Canada, Ruby Romero, Alex AT Bui, Adam Koncsol, Hritika Chaturvedi, Chelsea Shover, and David Goodman-Meza. 2026. Improving drug identification in overdose death surveillance by using clinical natural language processing models. Journal of Forensic Sciences. [Google Scholar] [CrossRef] [PubMed]
- Gabriel, Rodney A, Onkar Litake, Sierra Simpson, Brittany N Burton, Ruth S Waterman, and Alvaro A Macias. 2024. On the development and validation of large language model-based classifiers for identifying social determinants of health. Proceedings of the National Academy of Sciences 121, 39: e2320716121. [Google Scholar] [CrossRef]
- Gangavarapu, Agasthya. 2024. Introducing l2m3, a multilingual medical large language model to advance health equity in low-resource regions. arXiv arXiv:2404.08705. [Google Scholar]
- Goecks, Vinicius G., and Nicholas R. Waytowich. 2023. Disasterresponsegpt: Large language models for accelerated plan of action development in disaster response scenarios. Workshop on Challenges in Deployable Generative AI at International Conference on Machine Learning (ICML), Honolulu, Hawaii, USA. [Google Scholar]
- Goncalves, Andre R, Jose Cadena Pico, Yeping Hu, David Schlessinger, John Greene, Liam O’suilleabhain, Heather Clancy, Michael Vollmer, Vincent Liu, Tom Bates, and et al. 2025. Ai-enabled diagnostic prediction within electronic health records to enhance biosurveillance and early outbreak detection. In Medrxiv. [Google Scholar]
- Gong, Chenghua, Rui Sun, Yuhao Zheng, Juyuan Zhang, Tianjun Gu, Liming Pan, and Linyuan Lv. 2025. Epillm: unlocking the potential of large language models in epidemic forecasting. arXiv arXiv:2505.12738. [Google Scholar]
- Govathson, Caroline, Candice Chetty-Makkan, Ross Greener, Sasha Frade, Dino Rech, Sarah Morris, Yohann Richard, Rouella Mendonca, Natalie Maricich, Lawrence Long, and et al. 2026. Breaking barriers: harnessing artificial intelligence for a stigma-free, efficient hiv prevention assessment among adults in south africa. Frontiers in Digital Health 7, 1731002. [Google Scholar] [CrossRef]
- Grattafiori, Aaron, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and et al. 2024. The llama 3 herd of models. arXiv arXiv:2407.21783. [Google Scholar]
- Gu, Zifan, Lesi He, Awais Naeem, Pui Man Chan, Asim Mohamed, Hafsa Khalil, Yujia Guo, Jingwei Huang, Ismael Villanueva-Miranda, Ying Ding, and et al. 2025. Sbdh-reader: a large language model-powered method for extracting social and behavioral determinants of health from clinical notes. Journal of the American Medical Informatics Association 32, 10: 1570–1580. [Google Scholar] [PubMed]
- Gullison, Lucinda, and Feng Fu. 2025. Working with large language models to enhance messaging effectiveness for vaccine confidence. arXiv arXiv:2504.09857. [Google Scholar]
- Guo, Daya, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, and et al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv arXiv:2501.12948. [Google Scholar]
- Haider, Batool, Atmika Gorti, Aman Chadha, and Manas Gaur. 2025. Mental health equity in llms: Leveraging multi-hop question answering to detect amplified and silenced perspectives. arXiv arXiv:2506.18116. [Google Scholar]
- Hakim, Joe B, Jeffery L Painter, Darmendra Ramcharran, Vijay Kara, Greg Powell, Paulina Sobczak, Chiho Sato, Andrew Bate, and Andrew Beam. 2025. The need for guardrails with large language models in pharmacovigilance and other medical safety critical settings. Scientific Reports 15, 1: 27886. [Google Scholar] [CrossRef] [PubMed]
- Han, Eileen, Miao Feng, and Pamela Ling. 2025. Building an analytical framework for tobacco-related information on social media: an exploratory analysis with generative ai assistance. BMC Public Health 25, 1: 3635. [Google Scholar] [CrossRef] [PubMed]
- Harris, Joshua, Fan Grayson, Felix Feldman, Timothy Laurence, Toby Nonnenmacher, Oliver Higgins, Leo Loman, Selina Patel, Thomas Finnie, Samuel Collins, and et al. 2025. Healthy llms? benchmarking llm knowledge of uk government public health information. arXiv arXiv:2505.06046. [Google Scholar]
- Hattab, Georges, Christopher Irrgang, Nils Körber, Denise Kühnert, and Katharina Ladewig. 2025. The way forward to embrace artificial intelligence in public health. [Google Scholar] [CrossRef] [PubMed]
- Hong, Jaeff, Duong Dung, Danielle Hutchinson, Zubair Akhtar, Rosalie Chen, Rebecca Dawson, Aditya Joshi, Samsung Lim, C Raina MacIntyre, and Deepti Gurdasani. 2023. Relation extraction from news articles (rena): A tool for epidemic surveillance. arXiv arXiv:2311.01472. [Google Scholar]
- Hou, Zhiyuan, Zhengdong Wu, Zhiqiang Qu, Liubing Gong, Hui Peng, Mark Jit, Heidi J Larson, Joseph T Wu, and Leesa Lin. 2025. A vaccine chatbot intervention for parents to improve hpv vaccination uptake among middle school girls: a cluster randomized trial. Nature Medicine 31, 6: 1855–1862. [Google Scholar] [CrossRef] [PubMed]
- Hu, Edward J, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. International Conference on Learning Representations. [Google Scholar]
- Huang, Jingyi, Yuyi Yang, Mengmeng Ji, Charles Alba, Sheng Zhang, and Ruopeng An. 2025. Use of retrieval-augmented large language model agent for long-form covid-19 fact-checking. arXiv arXiv:2512.00007. [Google Scholar]
- Huang, Weihong, Wudi Wei, Xiaotao He, Baili Zhan, Xiaoting Xie, Meng Zhang, Shiyi Lai, Zongxiang Yuan, Jingzhen Lai, Rongfeng Chen, and et al. 2025. Chatgpt-assisted deep learning models for influenza-like illness prediction in mainland china: time series analysis. Journal of Medical Internet Research 27, e74423. [Google Scholar]
- Humphries, Hilton, Lindani Msimango, Zimasa Tshawe, Natasha Gcelu, Kurt Ferreira, Jacqueline Pienaar, Elise M van der Elst, Danielle Giovenco, Don Operario, Eduard J Sanders, and et al. 2026. A qualitative study assessing the acceptability of a multi-agent ai chatbot for providing hiv and mental health support among men who have sex with men and transgender women in kwazulu-natal, south africa. Transactions of The Royal Society of Tropical Medicine and Hygiene 120, 2: 160–174. [Google Scholar] [PubMed]
- Hussain, Ayana, Patrick Zhao, and Nicholas Vincent. 2025. An audit and analysis of llm-assisted health misinformation jailbreaks against llms. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, vol. 8, pp. 1290–1301. [Google Scholar]
- Ito, Jumpei, Adam Strange, Wei Liu, Gustav Joas, Spyros Lytras, and Kei Sato. 2025. A protein language model for exploring viral fitness landscapes. Nature communications 16, 1: 4236. [Google Scholar] [CrossRef] [PubMed]
- Jannah, Saidah Zahrotul, Elyanah Aco, Shaowen Peng, Shoko Wakamiya, and Eiji Aramaki. 2025. Multilingual symptom detection on social media: enhancing health-related fact-checking with llms. Proceedings of the Eighth Fact Extraction and VERification Workshop (FEVER); pp. 54–68. [Google Scholar]
- Ji, Yuelyu, Wenhe Ma, Sonish Sivarajkumar, Hang Zhang, Eugene M Sadhu, Zhuochun Li, Xizhi Wu, Shyam Visweswaran, and Yanshan Wang. 2025. Mitigating the risk of health inequity exacerbated by large language models. npj Digital Medicine 8, 1: 246. [Google Scholar] [CrossRef] [PubMed]
- Jo, Eunkyung, Daniel A Epstein, Hyunhoon Jung, and Young-Ho Kim. 2023. Understanding the benefits and challenges of deploying conversational ai leveraging large language models for public health intervention. Proceedings of the 2023 CHI conference on human factors in computing systems; pp. 1–16. [Google Scholar]
- Joseph, Jeena, Binny Jose, and Jobin Jose. 2025. The generative illusion: how chatgpt-like ai tools could reinforce misinformation and mistrust in public health communication. Frontiers in Public Health 13, 1683498. [Google Scholar]
- Kalahasti, Suprabhath, Benjamin Faucher, Boxuan Wang, Claudio Ascione, Ricardo Carbajal, Maxime Enault, Christophe Vincent Cassis, Titouan Launay, Caroline Guerrisi, Pierre-Yves Boëlle, and et al. 2025. Foundation time series models for forecasting and policy evaluation in infectious disease epidemics. medRxiv, 2025–02. [Google Scholar]
- Kaur, Jasleen, and Zahid Ahmad Butt. 2025. Ai-driven epidemic intelligence: the future of outbreak detection and response. Frontiers in Artificial Intelligence 8, 1645467. [Google Scholar] [CrossRef]
- Keloth, Vipina K, Salih Selek, Qingyu Chen, Christopher Gilman, Sunyang Fu, Yifang Dang, Xinghan Chen, Xinyue Hu, Yujia Zhou, Huan He, and et al. 2025. Social determinants of health extraction from clinical notes across institutions using large language models. npj Digital Medicine 8, 1: 287. [Google Scholar] [CrossRef] [PubMed]
- Khademi, Sedigh, Jim Black, Christopher Palmer, Muhammad Javed, Hazel Clothier, Jim Buttery, and Gerardo Luis Dimaguila. 2025. Enhancing vaccine safety surveillance: Extracting vaccine mentions from emergency department triage notes using fine-tuned large language models. arXiv arXiv:2507.07599. [Google Scholar]
- Khademi, Sedigh, Christopher Palmer, Gerardo Luis Dimaguila, Muhammad Javed, and Jim Buttery. 2024. Exploring large language models for detecting online vaccine reactions. In Health. Innovation. Community: It Starts With Us. IOS Press: pp. 30–35. [Google Scholar]
- Kim, Kwanho, and Soojong Kim. 2025. Large language models’ accuracy in emulating human experts’ evaluation of public sentiments about heated tobacco products on social media: evaluation study. Journal of Medical Internet Research 27, e63631. [Google Scholar] [CrossRef]
- Kim, Soojong, Kwanho Kim, and Claire Wonjeong Jo. 2024. Accuracy of a large language model in distinguishing anti-and pro-vaccination messages on social media: The case of human papillomavirus vaccination. Preventive Medicine Reports 42, 102723. [Google Scholar]
- Klempir, Ondrej, Ladislav Dusek, Radim Krupicka, Gleb Donin, Jan Zigmond, Radka Storchova, and Ales Tichopad. 2025. Evaluating large language models for natural-language-to-code generation on aggregate czech public health data analysis. medRxiv, 2025–12. [Google Scholar]
- Kwok, Kin On, Tom Huynh, Wan In Wei, Samuel YS Wong, Steven Riley, and Arthur Tang. 2024. Utilizing large language models in infectious disease transmission modelling for public health preparedness. Computational and structural biotechnology journal 23, 3254–3257. [Google Scholar] [CrossRef]
- Laily, Alfu, Laura M Schwab-Reese, Megan Davish, Emily Cahue, Kathryn J LaRoche, Natalia M Rodriguez, Robert J Duncan, Randolph D Hubach, and Monica L Kasting. 2026. Examining artificial intelligence chatbots’ responses in providing human papillomavirus vaccine information for young adults: Qualitative content analysis. JMIR Public Health and Surveillance 12, e79720. [Google Scholar] [CrossRef]
- Lau, Max SY, C Jessica E Metcalf, Zewen Liu, Bryan T Grenfell, and Wei Jin. 2026. Toward ai foundation models for epidemics: Promise, challenges, and paths forward. Proceedings of the National Academy of Sciences 123, 13: e2526192123. [Google Scholar] [CrossRef]
- Laurence, Timothy, Joshua Harris, Leo Loman, Amy Douglas, Yung-Wai Chan, Luke Hounsome, Lesley Larkin, and Michael Borowitz. 2025. Review gide–restaurant review gastrointestinal illness detection and extraction with large language models. arXiv arXiv:2503.09743. [Google Scholar]
- Lewis, Patrick, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, and et al. 2020. Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems 33: 9459–9474. [Google Scholar]
- Li, Chenxiang, Qiqiao Zhang, Yue Zhang, Bowen Zhao, Jule Yang, Li Qi, Jun Ding, and Dechao Tian. 2025. Fine-tuned large language models enhance influenza forecasting. medRxiv, 2025–03. [Google Scholar]
- Li, Hai, Jingyi Huang, Mengmeng Ji, Yuyi Yang, and Ruopeng An. 2025. Use of retrieval-augmented large language model for covid-19 fact-checking: development and usability study. Journal of medical Internet research 27, e66098. [Google Scholar]
- Li, Yiming, Jianfu Li, Jianping He, and Cui Tao. 2024. Ae-gpt: using large language models to extract adverse events from surveillance reports-a use case with influenza vaccine adverse events. Plos one 19, 3: e0300919. [Google Scholar] [PubMed]
- Li, Yiming, Deepthi Viswaroopan, William He, Jianfu Li, Xu Zuo, Hua Xu, and Cui Tao. 2025. Enhancing relation extraction for covid-19 vaccine shot-adverse event associations with large language models. Research Square, rs–3. [Google Scholar]
- Liang, Zongjing, Gongcheng Liang, Yun Kuang, Zhijie Li, and Kuang Yun. 2025. Application and comparative study of generative artificial intelligence for epidemic prediction of coronavirus disease. Cureus 17, 8. [Google Scholar] [CrossRef] [PubMed]
- Liu, Junyu, Qian Niu, Momoko Nagai-Tanima, and Tomoki Aoyama. 2025. Understanding human papillomavirus vaccination hesitancy in japan using social media: content analysis. Journal of Medical Internet Research 27, e68881. [Google Scholar] [CrossRef]
- Liu, Ollie, Sami Jaghouar, Johannes Hagemann, Shangshang Wang, Jason Wiemels, Jeff Kaufman, and Willie Neiswanger. 2025. Metagene-1: Metagenomic foundation model for pandemic monitoring. arXiv arXiv:2501.02045. [Google Scholar]
- Liu, Siying, Shisheng Zhang, and Indu Bala. 2025. Robust or suggestible? exploring non-clinical induction in llm drug-safety decisions. arXiv arXiv:2510.13931. [Google Scholar]
- Liu, Xiaoyu, Lu He, Eman Alanazi, Echu Liu, Arianna Goss, and Lionel Gumireddy. 2025. Assessing the accuracy and explainability of using chatgpt to evaluate the quality of health news. BMC Public Health 25, 1: 2038. [Google Scholar] [CrossRef] [PubMed]
- Liu, Yuqi, Jing Li, Peihan Li, Yehong Yang, Kaiying Wang, Jinhui Li, Lang Yang, Jiangfeng Liu, Leili Jia, Aiping Wu, and et al. 2025. Arnle model identifies prevalence potential of sars-cov-2 variants. Nature Machine Intelligence 7, 1: 18–28. [Google Scholar]
- Liu, Zewen, Juntong Ni, Max SY Lau, and Wei Jin. 2025. Pre-training epidemic time series forecasters with compartmental prototypes. arXiv arXiv:2502.03393. [Google Scholar]
- Liu, Zewen, Guancheng Wan, B Aditya Prakash, Max SY Lau, and Wei Jin. 2024. A review of graph neural networks in epidemic modeling. Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining; pp. 6577–6587. [Google Scholar]
- Lu, Linqi, Yanshu Sybil Wang, Jiawei Liu, and Douglas M McLeod. 2026. Human–generative ai interactions and their effects on beliefs about health issues: Content analysis and experiment. JMIR AI 5, 1: e80270. [Google Scholar] [CrossRef] [PubMed]
- Malek, Samira, Christopher Griffin, Robert D Fraleigh, Robert Lennon, Vishal Monga, and Lijiang Shen. 2026. Intervention in health misinformation using large language models for automated detection, thematic analysis, and inoculation: Case study on covid-19. Journal of medical Internet research 28, e75500. [Google Scholar] [CrossRef]
- Mao, Kangkun, Fang Xu, Jinru Ding, Yidong Jiang, Yujun Yao, Yirong Chen, Junming Liu, Xiaoqin Wu, Qian Wu, Xiaoyan Huang, and et al. 2025. Epiplanagent: Agentic automated epidemic response planning. arXiv arXiv:2512.10313. [Google Scholar]
- Martinson, Sarah, Lingkai Kong, Cheol Woo Kim, Aparna Taneja, and Milind Tambe. 2025. Llm-based agent simulation for maternal health interventions: uncertainty estimation and decision-focused evaluation. arXiv arXiv:2503.22719. [Google Scholar]
- McMurry, Andrew J, Dylan Phelan, Brian E Dixon, Alon Geva, Daniel Gottlieb, James R Jones, Michael Terry, David E Taylor, Hannah Callaway, Sneha Manoharan, and et al. 2025. Large language model symptom identification from clinical text: Multicenter study. Journal of medical Internet research 27, e72984. [Google Scholar] [CrossRef]
- Menon, Vaishnavi, Natnael Shimelash, Samuel Rutunda, Cyprien Nshimiyimana, Lucinda Archer, Mira Emmanuel-Fabula, Derbew Fikadu Berhe, Jaspret Gill, Emery Hezagira, Eric Remera, and et al. 2025. Assessing the potential utility of large language models for assisting community health workers: protocol for a prospective, observational study in rwanda. BMJ open 15, 10: e110927. [Google Scholar] [CrossRef] [PubMed]
- Mittal, Shravika, Hayoung Jung, Mai ElSherief, Tanushree Mitra, and Munmun De Choudhury. 2025. Online myths on opioid use disorder: A comparison of reddit and large language model. Proceedings of the International AAAI Conference on Web and Social Media 19: 1224–1245. [Google Scholar] [CrossRef]
- Modi, Natansh D, Bradley D Menz, Abdulhalim A Awaty, Cyril A Alex, Jessica M Logan, Ross A McKinnon, Andrew Rowland, Stephen Bacchi, Kacper Gradon, Michael J Sorich, and et al. 2025. Assessing the system-instruction vulnerabilities of large language models to malicious conversion into health disinformation chatbots. Annals of internal medicine 178, 8: 1172–1180. [Google Scholar] [CrossRef] [PubMed]
- Moon, Jaeuk, Jonghwa Shim, Eunbeen Kim, and Eenjun Hwang. 2025. Miflu: large language model-based multimodal influenza forecasting scheme. IEEE Journal of Biomedical and Health Informatics. [Google Scholar] [CrossRef] [PubMed]
- Morita, Plinio P, Matheus Lotto, Jasleen Kaur, Dmytro Chumachenko, Arlene Oetomo, Kristopher Dylan Espiritu, and Irfhana Zakir Hussain. 2024. What is the impact of artificial intelligence-based chatbots on infodemic management? Frontiers in public health 12, 1310437. [Google Scholar] [CrossRef]
- Mu, Yida, Mali Jin, Kalina Bontcheva, and Xingyi Song. 2024. Examining temporalities on stance detection towards covid-19 vaccination. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024); pp. 6732–6738. [Google Scholar]
- Muller, Jacob, Daniel Petti, Changying Li, Serap Gorucu, Matthew Pilz, and Bryan P Weichelt. 2025. Large language models for agricultural injury surveillance. Safety 11, 1: 15. [Google Scholar] [CrossRef]
- Nananukul, Navapat, and Mayank Kejriwal. 2024. A large language model-based approach for analyzing covariates of health equity in registered research projects. medRxiv, 2024–09. [Google Scholar]
- Narayan, Aditya, Michael Blasingame, Sam Warmuth, Gabriella Palmeri, India Halm, Ramin Bastani, Whitney Engeran-Cordova, Harold J Phillips, Leandro Mena, and Nirav R Shah. 2026. Ai-augmented communication improves hiv prep initiation and persistence in populations disproportionately impacted by hiv. npj Digital Medicine. [Google Scholar] [PubMed]
- Oh, Yoo Jung, Muhammad Ehab Rasul, Emily McKinley, and Christopher Calabrese. 2025. From digital traces to public vaccination behaviors: leveraging large language models for big data classification. Frontiers in Artificial Intelligence 8, 1602984. [Google Scholar] [CrossRef]
- Olatunji, Tobi, Charles Nimo, Abraham Owodunni, Tassallah Abdullahi, Emmanuel Ayodele, Mardhiyah Sanni, Chinemelu Aka, Folafunmi Omofoye, Foutse Yuehgoh, Timothy Faniran, and et al. 2024. Afrimed-qa: a pan-african, multi-specialty, medical question-answering benchmark dataset. arXiv arXiv:2411.15640. [Google Scholar]
- Omaki, Elise, Felipe Restrepo, Wendy C Shields, and Alan Abrahams. 2025. Natural language processing tool for extracting information about opioid overdoses in the usa from case narratives in the violent death reporting system. Injury Prevention. [Google Scholar] [CrossRef] [PubMed]
- Ong, Jasmine Chiat Ling, Yilin Ning, Rui Yang, Danielle S Bitterman, Xiaoxuan Liu, Yih Chung Tham, Gary S Collins, Michelle María Jiménez de Tavárez, Bilal A Mateen, Kwesi Nyan Amissah-Arthur, and et al. 2026. Large language models in global health. Nature Health 1, 1: 35–47. [Google Scholar] [CrossRef]
- Ong, Jasmine Chiat Ling, Benjamin Jun Jie Seng, Jeren Zheng Feng Law, Lian Leng Low, Andrea Lay Hoon Kwa, Kathleen M Giacomini, and Daniel Shu Wei Ting. 2024. Artificial intelligence, chatgpt, and other large language models for social determinants of health: Current state and future directions. Cell Reports Medicine 5, 1. [Google Scholar] [CrossRef] [PubMed]
- Ouyang, Long, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems 35: 27730–27744. [Google Scholar] [CrossRef]
- Painter, Jeffery L, Venkateswara Rao Chalamalasetti, Raymond Kassekert, and Andrew Bate. 2025. Automating pharmacovigilance evidence generation: using large language models to produce context-aware structured query language. JAMIA open 8, 1: ooaf003. [Google Scholar] [PubMed]
- Pan, Jie, Seungwon Lee, Cheligeer Cheligeer, Elliot A Martin, Kiarash Riazi, Hude Quan, and Na Li. 2025. Integrating large language models with human expertise for disease detection in electronic health records. Computers in Biology and Medicine 191, 110161. [Google Scholar] [CrossRef]
- Panja, Madhurima, Ojas Modak, Grace Younes, and Tanujit Chakraborty. 2025. Zero-shot forecasting of epidemics. In. Recent Advances in Time Series Foundation Models Have We Reached the’BERT Moment’? [Google Scholar]
- Pant, Devesh, Rishi Raj Grandhe, Jatin Agrawal, Jushaan Singh Kalra, Sudhir Kumar, Saransh Khanna, Vipin Samaria, Mukul Paul, Satish V Khalikar, Vipin Garg, and et al. 2025. Health sentinel: An ai pipeline for real-time disease outbreak detection. Proceedings of the Fourth Workshop on NLP for Positive Impact (NLP4PI); pp. 23–42. [Google Scholar]
- Parekh, Tanmay, Jeffrey Kwan, Jiarui Yu, Sparsh Johri, Hyosang Ahn, Sreya Muppalla, Kai-Wei Chang, Wei Wang, and Nanyun Peng. 2024. Speed++: A multilingual event extraction framework for epidemic prediction and preparedness. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing; pp. 12936–12965. [Google Scholar]
- Parker, Susan T. 2025. Supervised natural language processing classification of violent death narratives: Development and assessment of a compact large language model. JMIR AI 4, e68212. [Google Scholar] [CrossRef]
- Perret, Jasmin, and Adrian Schmid. 2024. Application of openai gpt-4 for the retrospective detection of catheter-associated urinary tract infections in a fictitious and curated patient data set. Infection Control & Hospital Epidemiology 45, 1: 96–99. [Google Scholar]
- Pfohl, Stephen R, Heather Cole-Lewis, Rory Sayres, Darlene Neal, Mercy Asiedu, Awa Dieng, Nenad Tomasev, Qazi Mamunur Rashid, Shekoofeh Azizi, Negar Rostamzadeh, and et al. 2024. A toolbox for surfacing health equity harms and biases in large language models. Nature Medicine 30, 12: 3590–3600. [Google Scholar] [CrossRef] [PubMed]
- Pierson, Emma, Divya Shanmugam, Rajiv Movva, Jon Kleinberg, Monica Agrawal, Mark Dredze, Kadija Ferryman, Judy Wawira Gichoya, Dan Jurafsky, Pang Wei Koh, and et al. 2025. Using large language models to promote health equity. [Google Scholar] [CrossRef] [PubMed]
- Proctor, Joshua L, and Guillaume Chabot-Couture. 2024. Democratizing infectious disease modeling: an ai assistant for generating, simulating, and analyzing dynamic models. medRxiv, 2024–07. [Google Scholar]
- Quigley, Ashley, Damian Honeyman, Haley Stone, Rebecca Dawson, and C Raina MacIntyre. 2025. Epiwatch, an artificial intelligence early-warning system as a valuable tool in outbreak surveillance. International Journal of Infectious Diseases 152, 107579. [Google Scholar]
- Raffel, Colin, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research 21, 140: 1–67. [Google Scholar]
- Rahman, Shagoto, Cornelia Pechmann, and Ian G Harris. 2026. Enhancing detection of message intents in a mobile health smoking-cessation intervention using large language model fine-tuning, data downsampling, and error correction: Algorithm development and validation. Journal of Medical Internet Research 28, e83437. [Google Scholar]
- Ramjee, Pragnya, Mehak Chhokar, Bhuvan Sachdeva, Mahendra Meena, Hamid Abdullah, Aditya Vashistha, Ruchit Nagar, and Mohit Jain. 2025. Ashabot: An llm-powered chatbot to support the informational needs of community health workers. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems; pp. 1–22. [Google Scholar]
- Reis, Florian, Lea J Bayer, Claudius Malerczyk, Christian Lenz, and Christof von Eiff. 2026. Leveraging large language models to address common vaccination myths and misconceptions. medRxiv, 2026–02. [Google Scholar]
- Rizzo, Alberto, Enrico Mensa, and Andrea Giacomelli. 2024. The future of large language models in fighting emerging outbreaks: lights and shadows. The Lancet Microbe 5, 11. [Google Scholar] [CrossRef] [PubMed]
- Rosebrock, Tracy R, Zhen Yang, Lauren D’Arco, Tapan Pathak, Rebecca Vislay-Wade, Karen Fowler, John Diaz-Decaro, and Colin Kunzweiler. 2026. Using artificial intelligence methods to evaluate the effect of the national cytomegalovirus awareness month on the content and sentiment of social media posts: Infodemiology study. JMIR infodemiology 6, e80922. [Google Scholar] [CrossRef]
- Rostami, Melika, and Suliman Hawamdeh. 2025. Debunk lists as external knowledge structures for health misinformation detection with generative ai. Systems 13, 10: 882. [Google Scholar] [CrossRef]
- Saeidnia, Hamid Reza, Shamim Jahani, Nasrin Ghiasi, and Hamid Keshavarz. 2026. Generative ai and health misinformation: production, propagation, and mitigation—a systematic review. BMC Public Health. [Google Scholar] [CrossRef] [PubMed]
- Samaei, Mohammad Hosseini, Faryad Darabi Sahneh, Lee W Cohnstaedt, and Caterina M Scoglio. 2026. Epidemiqs: Prompt-to-paper llm agents for epidemic modeling and analysis. In IEEE Transactions on Artificial Intelligence. [Google Scholar]
- Sheridan, Niamh, and et al. 2025. ILI surveillance from Twitter in Wales. Proceedings of the 10th Workshop on Noisy and User-generated Text (W-NUT). [Google Scholar]
- Shi, Ziyi, Xusen Guo, Hongliang Lu, Mingxing Peng, Haotian Wang, Zheng Zhu, Zhenning Li, Yuxuan Liang, Xinhu Zheng, and Hai Yang. 2026. Coordinated pandemic control with large language model agents as policymaking assistants. arXiv arXiv:2601.09264. [Google Scholar]
- Shimelash, Natnael, Samuel Rutunda, Vaishnavi Menon, Mira Emmanuel-Fabula, Angel Uwimbabazi, Crystal Rugege, Cyprien Nshimiyimana, Ivan Rwema, Mouna Kandekwe, Derbew Fikadu Berhe, and et al. 2026. A ‘silent trial’assessing the accuracy of large language models for assisting community health workers in low-resource settings. medRxiv, 2026–02. [Google Scholar]
- Si, Yafei, Yuyi Yang, Xi Wang, Jiaqi Zu, Xi Chen, Xiaojing Fan, Ruopeng An, and Sen Gong. 2024. Quality and accountability of chatgpt in health care in low-and middle-income countries: simulated patient study. Journal of medical Internet research 26, e56121. [Google Scholar]
- Sidorov, Grigori, Muhammad Ahmad, Pierpaolo Basile, Muhammad Waqas, Rita Orji, and Ildar Batyrshin. 2025. Monitoring opioid-related social media chatter using natural language processing and large language models: Temporal analysis. JMIR infodemiology 5, 1: e77279. [Google Scholar] [CrossRef] [PubMed]
- Simmons, Zalaya, Beti Evans, Tamsyn Harris, Harry Woolnough, Lauren Dunn, Jonathon Fuller, Kerry Cella, and Daphne Duval. 2025. Assessing the feasibility and acceptability of a bespoke large language model pipeline to extract data from different study designs for public health evidence reviews. Cochrane Evidence Synthesis and Methods 3, 6: e70061. [Google Scholar] [CrossRef] [PubMed]
- Singh, Nina, Katharine Lawrence, Safiya Richardson, and Devin M Mann. 2023. Centering health equity in large language model deployment. PLOS Digital Health 2, 10: e0000367. [Google Scholar] [CrossRef] [PubMed]
- Song, Xiaoying, Anirban Saha Anik, Dibakar Barua, Pengcheng Luo, Junhua Ding, and Lingzi Hong. 2025. Speaking at the right level: Literacy-controlled counterspeech generation with RAG-RL. In Findings of the Association for Computational Linguistics: EMNLP 2025. Edited by C. Christodoulopoulos, T. Chakraborty, C. Rose and V. Peng. Suzhou, China: Association for Computational Linguistics, November, pp. 2812–2830. [Google Scholar] [CrossRef]
- Sridi, Chayma, and Salem Brigui. 2023. The use of chatgpt in occupational medicine: opportunities and threats. Annals of occupational and environmental medicine 35, e42. [Google Scholar] [CrossRef]
- Stoll, Dragan, Samuel Wehrli, and David Lätsch. 2025. Case reports unlocked: Harnessing large language models to advance research on child maltreatment. Child Abuse & Neglect 160, 107202. [Google Scholar]
- Stureborg, Rickard, Sanxing Chen, Roy Xie, Aayushi Patel, Christopher Li, Chloe Zhu, Tingnan Hu, Jun Yang, and Bhuwan Dhingra. 2024. Tailoring vaccine messaging with common-ground opinions. Findings of the Association for Computational Linguistics: NAACL 2024, 2553–2575. [Google Scholar] [CrossRef]
- Tao, Jun, Ellie Pavlick, Amaris Grondin, Josue D Bustamante, Harrison Martin, Hannah Parent, Natalie Fenn, Alexi Almonte, Amanda Maguire-Wilkerson, Mofan Gu, and et al. 2026. Evaluation of an artificial intelligence conversational chatbot to enhance hiv preexposure prophylaxis uptake: Development and usability internal testing. Journal of Medical Internet Research 28, e79671. [Google Scholar] [CrossRef]
- Testagrose, Conrad, Sakshi Pandey, Mohammadali Serajian, Simone Marini, Mattia Prosperi, and Christina Boucher. 2025. Leveraging large language models to predict antibiotic resistance in mycobacterium tuberculosis. Bioinformatics 41 Supplement_1: i40–i48. [Google Scholar] [CrossRef] [PubMed]
- Tiwari, Anushree, Amit Kumar, Shailesh Jain, Kanika S Dhull, Arunkumar Sajjanar, Rahul Puthenkandathil, Kapil Paiwal, Ramanpal Singh, and Arun Sajjanar. 2023. Implications of chatgpt in public health dentistry: A systematic review. Cureus 15, 6. [Google Scholar] [CrossRef] [PubMed]
- Touvron, Hugo, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, and et al. 2023. Llama: Open and efficient foundation language models. arXiv arXiv:2302.13971. [Google Scholar]
- van Hoek, Albert Jan, Sebastian Funk, Stefan Flasche, Billy J Quilty, Esther van Kleef, Anton Camacho, and Adam J Kucharski. 2024. Importance of investing time and money in integrating large language model-based agents into outbreak analytics pipelines. The Lancet Microbe 5, 8. [Google Scholar] [CrossRef] [PubMed]
- Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems. Curran Associates, Inc: Volume 30. [Google Scholar]
- Verma, Shresth, Alayna Nguyen, Niclas Boehmer, Lingkai Kong, and Milind Tambe. 2025. Priority2reward: Incorporating healthworker preferences for resource allocation planning. Proceedings of the AAAI Conference on Artificial Intelligence 39: 29709–29711. [Google Scholar] [CrossRef]
- Vos, Michiel, Markus Göker, Richard Bendall, and Fabrizio Costa. 2025. Large language model-assisted text mining reveals bacterial pathogen diversity. bioRxiv, 2025–07. [Google Scholar]
- Wang, Song, Yishu Wei, Haotian Ma, Max Lovitt, Kelly Deng, Yuan Meng, Zihan Xu, Jingze Zhang, Yunyu Xiao, Ying Ding, and et al. 2025. A multi-stage large language model framework for extracting suicide-related social determinants of health. Communications Medicine 5, 1: 404. [Google Scholar] [PubMed]
- Wang, Yuanlong, Pengqi Wang, Changchang Yin, and Ping Zhang. 2025. Sathealth: A multimodal public health dataset with satellite-based environmental factors. Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2; pp. 5819–5830. [Google Scholar]
- Wattamwar, Aniket, and Sampson Akwafuo. 2026. Aries: A scalable multi-agent orchestration framework for real-time epidemiological surveillance and outbreak monitoring. arXiv arXiv:2601.01831. [Google Scholar]
- Wei, Jason, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35: 24824–24837. [Google Scholar]
- Wei, Mingyang, Dehai Min, Zewen Liu, Yuzhang Xie, Guanchen Wu, Carl Yang, Max SY Lau, Qi He, Lu Cheng, and Wei Jin. 2026. Epiqal: Benchmarking large language models in epidemiological question answering for enhanced alignment and reasoning. arXiv arXiv:2601.03471. [Google Scholar]
- Wen, Richard, Anteneh Tesfaye Assalif, Andy Sze-Heng Lee, Rajeev Kamadod, Asha Behdinan, Ronald Carshon-Marsh, Catherine Meh, Thomas Kai Sze Ng, Patrick Brown, Prabhat Jha, and et al. 2026. Computer assisted verbal autopsy: comparing large language models to physicians for assigning causes to 6939 deaths in sierra leone from 2019–2022. BMC medicine 24, 1: 49. [Google Scholar] [CrossRef] [PubMed]
- Williams, Ross, Niyousha Hosseinichimeh, Aritra Majumdar, and Navid Ghaffarzadegan. 2023. Epidemic modeling with generative agents. arXiv arXiv:2307.04986. [Google Scholar]
- Wilson, Rory, Ciara M Weets, Amanda Rosner, and Rebecca Katz. 2024. Evaluating generative artificial intelligence’s limitations in health policy identification and interpretation. PLoS One 19, 12: e0312078. [Google Scholar] [CrossRef] [PubMed]
- Wu, Jiaying, Zihang Fu, Haonan Wang, Fanxiao Li, Jiafeng Guo, Preslav Nakov, and Min-Yen Kan. 2025. Beyond the crowd: Llm-augmented community notes for governing health misinformation. arXiv arXiv:2510.11423. [Google Scholar]
- Wu, Julie T, Bradley J Langford, Erica S Shenoy, Evan Carey, and Westyn Branch-Elliman. 2025. Chatting new territory: large language models for infection surveillance from pilot to deployment. Infection Control & Hospital Epidemiology 46, 3: 224–226. [Google Scholar] [CrossRef]
- Xia, Dengke, Mengyao Song, and Tingshao Zhu. 2025. A comparison of the persuasiveness of human and chatgpt generated pro-vaccine messages for hpv. Frontiers in public health 12, 1515871. [Google Scholar]
- Xie, Jiacheng, Ziyang Zhang, Shuai Zeng, Joel Hilliard, Guanghui An, Xiaoting Tang, Lei Jiang, Yang Yu, Xiufeng Wan, Dong Xu, and et al. 2025. Leveraging large language models for infectious disease surveillance—using a web service for monitoring covid-19 patterns from self-reporting tweets: Content analysis. Journal of Medical Internet Research 27, 1: e63190. [Google Scholar] [PubMed]
- Xu, Fengyi, Jun Ma, Nan Li, and Jack CP Cheng. 2025. Large language model applications in disaster management: An interdisciplinary review. International Journal of Disaster Risk Reduction 127, 105642. [Google Scholar]
- Xu, Gelei, Xueyang Li, Yixiong Chen, Yuying Duan, Shuqing Wu, Haoxinran Yu, Ching-Hao Chiu, Juntong Ni, Ningzhi Tang, Toby Jia-Jun Li, Alan Yuille, Wei Jin, and Yiyu Shi. 2026. A comprehensive survey of ai agents in healthcare. Journal of Biomedical Informatics 179, 105045. [Google Scholar] [CrossRef]
- Xu, Jiaxiang, Zhengdong Wu, Lily Wass, Heidi J Larson, and Leesa Lin. 2024. Mapping global public perspectives on mrna vaccines and therapeutics. npj Vaccines 9, 1: 218. [Google Scholar] [CrossRef] [PubMed]
- Xu, Shan, Zhaokun Yan, Chengxiao Dai, and Fan Wu. 2025. Mega-rag: a retrieval-augmented generation framework with multi-evidence guided answer refinement for mitigating hallucinations of llms in public health. Frontiers in Public Health 13, 1635381. [Google Scholar]
- Xu, Xuhai, Bingsheng Yao, Yuanzhe Dong, Saadia Gabriel, Hong Yu, James Hendler, Marzyeh Ghassemi, Anind K Dey, and Dakuo Wang. 2024. Mental-llm: Leveraging large language models for mental health prediction via online text data. Proceedings of the ACM on interactive, mobile, wearable and ubiquitous technologies 8, 1: 1–32. [Google Scholar]
- Yang, Chenghao, Tuhin Chakrabarty, Karli Hochstatter, Melissa Slavin, Nabila El-Bassel, and Smaranda Muresan. 2024. Identifying self-disclosures of use, misuse and addiction in community-based social media posts. Findings of the Association for Computational Linguistics: NAACL 2024, 2507–2521. [Google Scholar]
- Ratnam, Yoga, and K. Kishwen. 2025. Generative artificial intelligence in public health research and scientific communication: A narrative review of real applications and future directions. Digital Health 11, 20552076251362070. [Google Scholar]
- Zhang, Zhihao, Yiran Zhang, Xiyue Zhou, Liting Huang, Imran Razzak, Preslav Nakov, and Usman Naseem. 2025. From generation to detection: A multimodal multi-task dataset for benchmarking health misinformation. In Findings of the Association for Computational Linguistics: EMNLP 2025. Edited by C. Christodoulopoulos, T. Chakraborty, C. Rose and V. Peng. Suzhou, China: Association for Computational Linguistics, November, pp. 24245–24260. [Google Scholar] [CrossRef]
- Zhou, Jiawei, Amy Z Chen, Darshi Shah, Laura M Schwab-Reese, and Munmun De Choudhury. 2025. A risk taxonomy and reflection tool for large language model adoption in public health. Proceedings of the ACM on Human-Computer Interaction 9, 7: 1–32. [Google Scholar] [CrossRef]
- Zhou, Xinyu, Jiaqi Zhou, Chiyu Wang, Qianqian Xie, Kaize Ding, Chengsheng Mao, Yuntian Liu, Zhiyuan Cao, Huangrui Chu, Xi Chen, and et al. 2025. Ph-llm: public health large language models for infoveillance. In medRxiv. [Google Scholar]
- Ziletti, Angelo, and Leonardo DAmbrosi. 2024. Retrieval augmented text-to-sql generation for epidemiological question answering using electronic health records. Proceedings of the 6th Clinical Natural Language Processing Workshop; pp. 47–53. [Google Scholar]
- Zong, Ruohan, Yang Zhang, and Dong Wang. 2025. Empowering llms to synthesize ai and human intelligence for explainable public health misinformation detection on social media. Proceedings of the International AAAI Conference on Web and Social Media 19: 2334–2348. [Google Scholar] [CrossRef]
| Task / Section | N | Dominant LLM Role(s) | Typical Data Modality |
|---|---|---|---|
| T1 Surveillance | 53 | Sensor (), Predictor () | News / outbreak reports; social media / reviews; EHR / clinical text; structured surveillance records |
| T2 Forecasting | 18 | Predictor (), Simulator () | Epidemic time series; simulated epidemic trajectories; policy / genomic covariates |
| T3 Infodemiology | 52 | Sensor (), Communicator (), Auditor () | Social media; news / fact-check content; surveys; synthetic / multimodal misinformation data |
| T4 Health Equity | 22 | Auditor (), Sensor () | EHR / clinical notes; surveys / benchmarks; simulated prompts; literature / policy corpora |
| T5 Intervention | 18 | Communicator (), Sensor (), Simulator () | Surveys / chatbot interactions; policy / guideline documents; social media; EHR / mobile-health data |
| T6 Governance | 25 | Simulator (), Auditor () | Simulated epidemic environments; policy documents; benchmark / literature corpora; public health QA / evaluation data |
| Feature | Public Health | Clinical Medicine / Healthcare |
|---|---|---|
| Target | Population, community, region, or jurisdiction (Burstein et al. 2025,Chen et al. 2025) | Individual patient |
| Primary goal | Prevention, surveillance, preparedness, equity (Quigley et al. 2025,Singh et al. 2023) | Diagnosis, treatment, recovery, continuity of care |
| Typical data | Surveillance counts, news reports, policies, environmental and social signals, mobility (Deiner et al. 2024,Wang et al. 2025) | EHRs, labs, imaging, medications, clinical notes |
| Typical actions | Vaccination, screening, outbreak response, risk communication, resource allocation (Song et al. 2025,Verma et al. 2025) | Prescription, procedure, triage, monitoring, care planning |
| Key stakeholders | Health departments, CDC/WHO, policymakers, communities (Harris et al. 2025,van Hoek et al. 2024) | Clinicians, hospitals, care teams, patients |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).