Submitted:
17 August 2026
Posted:
19 August 2026
You are already at the latest version
Abstract
Synthetic health data enables privacy-safe data sharing, development of code, and mitigation of data scarcity and bias. While most research focuses on generation from electronic health records (EHR), this scoping review focuses on generation with access to only descriptive metadata information (metadata-based methods). Metadata-based methods are suited to privacy sensitive-contexts because they do not require access to individual EHR.This study aimed to identify current methods and research gaps, as well as synthesise relevant evaluation, ethical, and legal frameworks within the context of code development for federated analysis. We searched PubMed, Ovid Emabse, and Web of Science through March 2025 to identify peer-reviewed articles and preprints which focused on metadata-based methods, or which reviewed synthetic data evaluation and/or legal and ethical requirements for health research from all years. After deduplication, abstract and full-text screening were conducted by two researchers for each record. Data extraction and synthesis were conducted by one researcher and validated by a second. Twelve original articles and eleven reviews were included. The twelve articles covered six techniques: four stepwise, one LLM-based, and one syntactical. The methods covered in the primary sources were not benchmarked according to existing evaluation frameworks to identify their strengths and weaknesses; only fidelity metrics were commonly used. Ten of eleven reviews proposed evaluation frameworks and three addressed ethical or legal considerations. We observed inconsistent evaluation definitions and limited exploration of ethical, legal, and governance issues. No reviews specifically discussed code development. In conclusion, in this scoping review we identified several metadata-based methods but limited evaluations for these methods. We found a strong body of literature regarding evaluation but limited specific context-specific guidance. There is a need for consistent evaluation of metadata-based methods, and for specific guidelines aligned with generation method and intended use to ensure quality and compliance with emerging AI requirements.
Keywords:
test data
; generated data
; electronic health record
; EHR
; healthcare data
Author Summary
Electronic health records (EHR) and other routinely collected data are increasingly used within health research, but strict regulations for the sharing of health data can limit researchers’ access. In privacy-sensitive contexts, synthetic data, which mimics the real data but does not contain actual patient records, may be used as an alternative for some applications. Current research often focuses on generating synthetic data directly from individual EHR, but this study focuses on methods which generate synthetic data from descriptive information about the original data: metadata-based methods. In this scoping review, we searched for existing metadata-based methods as well as guidance for the generation and use of synthetic data which may be applicable to these methods. We identified six metadata-based methods but found limited evaluations of these methods. We also found eleven review articles containing relevant guidance. These reviews offered a strong body of literature on evaluation for synthetic data, but with inconsistent definitions and limited discussion of legal and ethical concerns. Our findings show a need for evaluation of metadata-based methods and for guidelines aligned with generation method and intended use. Overall, this review identifies key research gaps and provides researchers with an overview of available metadata-based methods.
Introduction
The analysis of large-scale healthcare data to generate real world evidence (RWE) is becoming increasingly accepted for regulatory decision making, driving a demand for fast and reliable evidence generation pipelines [1,2,3]. RWE generation in the field of pharmacoepidemiology often relies on federated analysis, in which researchers develop central analytical pipelines that can be run locally by data access holders (S1) [1,4]. This type of federated analytics allows researchers to conduct multi-database studies while maintaining restrictions for the sharing of sensitive health data according to the General Data Protection Regulation [1,4]. In order to develop and test scripts without access to real data, researchers may use synthetic data which mimics the structure, distributions, and semantics of real electronic health records (EHR) [5]. Synthetic data used for this purpose may have different data quality and governance standards than in other contexts.
Most of the available research on methodology for the generation of synthetic EHR focuses on deep generative techniques such as generative adversarial networks (GANs) or variational autoencoders (VAEs) [6,7,8]. These black-box approaches generate synthetic data directly from the original EHR without explainability, and with a risk that they may include sensitive data into the resulting datasets [6,7]. Additionally, imputation based methods such as synthpop allow users to generate synthetic data using an R package [9]. However, the need for original EHR input makes these techniques less suitable for federated analysis pipelines in which data access is limited. The common use of black-box methods also reduces transparency, which limits the auditability and human oversight of such methods. While there are methods such as differential privacy metrics which help to mitigate these risks, governmental agencies are hesitant to implement these solutions due to the challenges in practical implementation [10]. In contrast, transparent and human-auditable approaches allow databases and governmental authorities to simply assess the safety of resulting synthetic data [11].
A promising approach to generate synthetic health data without directly using EHR is the use of metadata-based methods, which we define in the background section of this paper. Metadata-based methods require inputs such as public data, expert opinion, and metadata (structured descriptive and quantitative information about original data). These methods reduce privacy risks by avoiding generation from individual level patient records (microdata) and are suitable for the federated analysis context as they do not require access to EHR. Outside the healthcare domain, techniques such as metasyn have been developed to generate privacy-preserving synthetic data from metadata information, but further research would be needed to adapt and apply such tools to EHR data [12,13].
There is little literature available on the application of metadata-based methods in healthcare and on the guardrails around data fidelity, privacy and governance within federated analysis. Guidelines for production, evaluation, and use of synthetic data are still in development [6,7,14]. Notably, major regulatory bodies like European Medicines Agency and U.S. Federal Drug Administration have not yet published concrete guidance on synthetic data generation or use in healthcare research [15].
Existing reviews of synthetic health data generation focus mainly or exclusively on methods which require access to original EHR, and to our knowledge no other reviews have focused primarily on metadata-based approaches or application to federated analysis. Still, reviews focusing on evaluation frameworks or other ethical and legal considerations may have relevant findings which could apply within this context. We therefore focused both on articles describing metadata-based methods for synthetic health data generation, as well as literature reviews on synthetic health data evaluation and more general considerations.
In this scoping review, we identified methods and research gaps on metadata-based synthetic health data generation, as well as synthesised evaluation, legal, and ethical frameworks which could apply to synthetic data generation for code development in federated analysis.
Background
Synthetic data generation methods may be classified into two categories according to the input they require. The first category includes methods which generate directly from information at the individual level, referred to as microdata. In the case of synthetic health data generation, these methods require access to the full EHR including each individual patient record. Examples of this generation method include MDClone, medGAN, and Synthpop [9,16,17].
The second category includes methods which generate synthetic data without requiring access to the microdata. Instead of individual level data, the input for these methods is information about the data of interest. This includes information directly about the data of interest such as its structure, summary statistics, and semantics, as well as public summary statistics related to the data of interest (e.g. data regarding disease rates for a country when the aim is to generate EHRs from a specific hospital in that country), and expert prior knowledge (e.g. knowledge regarding typical patient characteristics for a given disease) [18,19].We characterise this type of information as metadata, and these methods as metadata-based.
While most current literature focuses on the former category, the second is arguably much more suitable for the strict privacy requirements of the healthcare domain, particularly when it comes to federated analysis as no sharing original EHR data is required. To our knowledge, this is the first review focused on metadata-based generation methods.
Results
Selection of Sources of Evidence
A total of 736 records were identified through search across all databases, and 4 records were identified through manual inclusion. After removing duplicates and conducting initial screening, 537 abstracts were screened. Of these, 134 abstracts proceeded to full-text screening. Full-text reports were successfully retrieved and screened for 118 records. After screening for eligibility, 89 reports were screened for relevance, of which 74 described generation methods and 15 were review articles. Among the sources describing generation methods (primary sources of evidence), n=10/74 were considered relevant as they focused on metadata-based methods. Among the reviews (secondary sources of evidence), n=9/15 were considered relevant systematic literature reviews. Four reports identified through manual inclusion (n=4) were considered relevant: 2 primary sources describing generation methods and 2 systematic literature reviews. Non-relevant reports pertained to methods not related to metadata or presented overviews of synthetic data generation without systematic processes. Ultimately, 23 reports (hereafter referred to as articles) were chosen for additional charting and discussion, including 12 articles describing metadata-based generation methods and 11 systematic literature reviews. Figure 1 shows the results of the selection process. Publication year, authors, title, and article type (review or primary article) for all 23 relevant articles are presented in Appendix S3. Charting results for all included articles are available in Appendix S4. Further charting results for literature reviews (search period, purpose, methods evaluated, and evaluation dimensions) are presented in Table 1. Further charting results for metadata-based primary articles (name of generation method, inclusion within review papers, input, output, and evaluations performed) are organised by method in Table 2.
Metadata-Based Methods
The 12 metadata-based generation articles described six generation methods, described below:
GitHub Documentation [39]
Synthea is an open-source community project which generates synthetic patient records from public records and user-inputted specifications. First, Synthea generates a patient population. Demographic information concerning gender, age, race, education level, and income level are derived from US Census Bureau data by default, but users can optionally modify this through an input CSV file.
Each patient in the generated population is then simulated individually throughout their life through a series of condition-based modules. Each module contains a set of states and transitions between these states with probabilities conditional on the patient characteristics (e.g. age, gender, existing conditions) using Monte Carlo simulation. The generic framework processes states one at a time to trigger conditions, encounters, medications, and other clinical events. For example, the module “Ear Infections” begins with the state “No_Infection”, and the likelihood for a patient to transition to state “Gets_Ear_Infection” depends on age. Transitioning to state “Gets_Ear_Infection” results in a SNOMED-CT “Otitis media” (diagnosis) and a SNOMED-CT: Encounter for symptom (procedure) added to the patient record. Patients who do not transition to state “Gets_Ear_Infection” proceed directly to termination of the module. Among patients who do proceed to state “Gets_Ear_Infection”, the likelihood to proceed to different treatment varies based on module specifications. At each timestep, each patient goes through all applicable modules. Modules can happen at the same time and can be conditional on each other. Users can add and edit modules during use. After generation, Synthea exports medical records within an observation period set by the user in FIHR, CCDA, Text, or CSV formats using SNOMED or LOINC coding systems.
- 2)
GitHub Documentation [40]
OSIM was developed by the Observational Health Data Sciences and Informatics (OHDSI) to generate synthetic patient records directly in its common data structure, the Observational Medical Outcomes Partnership Common Data Model (OMOP CDM). However, OSIM is no longer updated through OHDSI. Instead, OHDSI now supports users to convert synthetic data generated through Synthea into the OMOP CDM [41].
To generate with OSIM, users first upload “Simulation Attribute Tables”, which contain descriptive metrics, multinomial distributions, and single-state transition probability distributions about a target dataset. Users may also add a table containing attributes of drug treatment effects if desired. OSIM then generates a population in which each person is assigned gender, age, number of distinct conditions, and observation duration. Gender is assigned based on probability and each following attribute is assigned based on multinomial distributions of the attributes before (e.g. number of distinct conditions is assigned based on gender and age).
Then, OSIM simulates incident conditions for each patient until the number of conditions per patient is reached. Incident conditions and the time intervals between them are simulated based on gender, age, and condition count. Number of subsequent occurrences of each incident condition are simulated based on condition, gender, age, condition count, and time remaining in the observation period. Time intervals between each occurrence are simulated based on condition count and time remaining in the observation period. The number of distinct drug therapies for each patient is assigned based on gender, age and condition count. For each condition occurrence (not distinct condition), a likely number of drugs to associate to the condition is assigned based on the condition, age, drug count, and condition count. That number of draws is made for initial drug therapies until the number of distinct therapies for each person is reached. For each simulated distinct drug therapy, the number of drug exposures, combined sum of all drug exposures, and calendar duration are then simulated.
Then, OSIM modifies the data to include known relationships between drugs and outcomes. Treatment effects defined by drug, condition, effect size, and time-to-event relationship are algorithmically introduced into the generated data. This can add or remove conditions. After generation, OSIM outputs an OMOP dataset with the following CDM tables: person, observation_period, condition_era, drug_era, and procedure_occurrence.
- 3)
- SASC (Simple Approach to Synthetic Cohorts): Khorchani [36]
GitHub Documentation[42]
SASC is designed to generate patient-level datasets for COVID-19 through a web-based Shiny application. For each patient, multiple rows with data under the columns: age, gender, admission time, discharge time, outcome (0 or 1), and lab data (e.g. hemoglobin, serum chloride, and albumin) are generated. Users can specify population size, visit numbers, and “date range max available”. Users upload patient-level data or summary tables. If patient level data is uploaded, SASC first generates a table of grouped probabilities. This table shows the mean and SD of the attributes age, hospitalisation days, and lab values such as albumin, neutrophil count, and lymphocyte percentage based on outcome (death or survival) and gender. SASC uses the summary tables to generate patients with outcome, age, gender, and visit numbers. It then generates other numerical variables for each patient based on the outcome and gender, followed by generating visits per patient and visit times. SASC results in a patient-level excel file.
- 4)
- Davis, et al. [37]
The method by Davis et al. is intended to generate patient-level datasets which account for learning of best practices for treatments over time at the provider and institution level. Users may input summary tables with aggregate data, which include distributions and correlations of patient features, outcome risk factors, the adverse outcome rate, device factors, provider and institution learning factors, and real-world factors. Device factors include prevalence, treatment assignment factors, and safety signal (the difference in risk of adverse outcome associated with the device). Provider and institution learning factors include values such as the number of institutions, distribution of providers within institutions, institutional learning speed, provider learning speed. Real-world factors are missingness, noise, and omitted variables. Patients are then generated based on the distributions and correlations of patient features. Each patient is assigned an institution, provider, and time of treatment. Next, treatments are assigned based on patient features and patients are assigned case orders. Outcomes are then assigned based on patient features, treatment assignment, and learning effects by provider and institution. Finally, the user specifies parameters to introduce noise by randomly sampling patients and flipping their assigned outcome, missingness by randomly dropping a predefined proportion of features across all patients, and omitted variables by dropping all patient features in selected variables. After generation, the output is patient-level synthetic datasets.
- 5)
- Operational Data Model (ODM) Clinical Data Generator: Brix [19]
This web-based application generates patient-level synthetic clinical data in ODM format: an XML-based standardised format used to exchange clinical research data, ranging from clinical trials to routine documentation. ODM files contain metadata outlining a structure for the clinical data as well as the clinical data itself. There are over 23,000 configurations available in ODM format, and the purpose for this generator is to generate syntactically correct data for any ODM file. It is not intended to generate with fidelity to the original data.
Users input an ODM file and an optional configuration file. The configuration file is used for optional parameters including population size as well as ranges and distributions of variables. The generator then randomly generates data based on the structure defined in the ODM file metadata as well as the defined ranges and distributions in the configuration file. Uniform distribution, normal distribution, and fixed probabilities for discrete answer options are possible. However, dependencies between variables are not included. The result is data in ODM format which is syntactically, but not necessarily statistically, accurate to the original metadata.
- 6)
- Barr, et al. [38]
This method uses GPT-4o to generate a patient-level dataset using a single prompt and without pre-training. For the feasibility study by Barr et al., a perioperative dataset concerning patients who underwent routine and emergency cardiac surgery was used to generate input clinical parameters. Thirteen clinical parameters were chosen based on relevance and representation of a range of data formats and variable types. These thirteen parameters included patient demographics, preoperative morbidity factors, intraoperative data, and postoperative outcomes. Patient demographics included case ID, age, sex, height, weight, and BMI. Preoperative morbidity factors included ASA physical status classification, preoperative hypertension, and preoperative diabetes. Intraoperative data included operation type, operation duration, and intraoperative transfusion. Postoperative outcomes included length of stay. For continuous parameters, descriptive statistics including mean, standard deviation, and range were calculated from the real dataset. For categorical and binary parameters, corresponding proportions were calculated from the real dataset.
GPT-4o was then given a list of each parameter and its corresponding descriptive statistics and prompted to generate a table with one column for each parameter for a certain number of patients. The prompt included statements to “ensure that values for each patient make sense in the context of other values in that row”, “all numbers in the table must be positive, except for in columns 2 and 3”, and “Re-iterate until exactly correct. Provide a downloadable excel file of the dataset”. The exact prompt is available in Appendix S5. This method results in a patient-level excel file.
Overall, four of the six methods use a step-wise approach in which patients are first generated and then diagnoses, events, and treatments are subsequently generated [18,34,36,37]. Probabilities and distributions come from public knowledge, user designated variables, and existing database metadata. However, some differences exist in which parameters are generated and what kinds of data format inputs and outputs are used (Table 2). Among the two non-stepwise methods, one uses an LLM for generation [38] and the other solely generates syntactically correct data [19].
Literature Reviews
Table 1 shows an overview of existing literature review articles focused on generation of synthetic EHR. Among these eleven review articles, seven were primarily intended to evaluate generation methods [7,20,21,22,26,27,28], three were primarily intended to review evaluation metrics and/or propose evaluation frameworks [23,24,25], and one was intended to review data governance and legal requirements for synthetic health data generation [14]. The latest review period included articles published through March 2025 [7]. Among the nine articles which reviewed methods, a majority (n=7/9) listed GANs as one of the methods they included, and 2 reported only on GAN methods [27,28]. Other reviewed methods include VAEs, diffusion models (DMs), large language models (LLMs), and Bayesian networks (see Table 1 for all method types in each review). Ten of the eleven literature reviews proposed or defined evaluation dimensions for synthetic health data. Box 1 shows definitions for each evaluation dimension included in these reviews, presented alongside definitions from the UK Synthetic Data Community Group Report (UK-SDCG). Appendix S6 shows a charting of all definitions in each literature review.
Box 1. Definitions in Reviews Alongside UK-SDCG Definitions.
| Fidelity (Data resemblance, Fidelity, Similarity, Resemblance): statistical similarity between synthetic and real data | UK-SDCG: how closely synthetic data mirrors both the structure and statistical behavior of the source data [5] |
|
Utility (Usability, Utility): synthetic data’s performance in tasks compared to real data |
UK-SDCG: 1) Analytical utility: extent to which synthetic data preserves the statistical properties of the original dataset; 2) Functional utility: actual usefulness of synthetic data for specific tasks which may not necessarily require high statistical fidelity [5] |
| Privacy: protection from disclosure of sensitive information | UK-SDCG: Privacy: minimisation of the risk of de-identification or disclosure, ensuring that no individual in the original dataset can be directly or indirectly identified [5] |
|
Computational cost/carbon footprint: resource needs to produce synthetic data |
|
|
Qualitative assessment/qualitative evaluation: expert visual assessment of synthetic data |
|
|
Diversity: how well synthetic data captures variability of the real data |
|
|
Fairness: whether any groups were underrepresented or bias was introduced |
|
| Clinical validation: ability of synthetic data to correctly represent clinical scenarios and effectively support medical decision-making |
Of the ten review articles which named and defined evaluation dimensions, all included some form of a fidelity metric (“Data resemblance”, “Fidelity”, “Similarity”, “Resemblance”), utility metric (“Utility”, “Usability”), and privacy metric (“Privacy”). Four included a computational/carbon cost metric (“Computational cost”, “Carbon footprint”, “Performance dimensions”) [7,22,24,27]. Other included metrics were qualitative assessment/evaluation (n=2/10) [21,28], diversity (n=2/10) [21,22], fairness (n=2/10) [23,24], and clinical validation (n=1/10) [21]. The definition of fidelity was fairly consistent across review articles as the similarity between synthetic and real data. In defining utility, most articles (n=7/10 referred specifically to the usefulness of synthetic data for inference and modelling, while only three referred to the use of synthetic data for other downstream tasks such as data augmentation [21,22,23]. None specifically mentioned utility for code development in federated analysis. Although all articles agreed that privacy involves protection of sensitive information, the description of what constitutes disclosure varied from a narrower definition of “risk of re-identification of individuals from the original dataset used to train the generator” to a more broadly defined “risk of disclosing personal information inadvertently through the generated synthetic data” [20,24]. One review noted “there is no universally accepted standard definition for privacy”, while another identified “22 different ways to discuss privacy” [22,23].
Nine out of eleven review articles included an overview or description of evaluation metric selection, such as a flowchart leading future researchers to the most appropriate metric based on data type and evaluation purpose. Of the two which did not provide this, one instead provided a list of evaluation categories and referred readers to other published literature for details on specific metrics and decision-making process [22].
Regarding ethical and legal considerations, two literature review included fairness as an evaluation category. One defined fairness as related to concerns which emerge when decision making is informed by biased datasets which may unfairly penalise minority groups within a population [24]. The other defined fairness as “how well synthetic data provides equitable treatment across different groups” [23]. However, both reviews noted that despite a growing interest in this area, actual use of fairness-related metrics were scarce in the articles included in their review. One identified six articles which made use of fairness metrics, while the other identified three [43,44,45,46,47,48]. Another review included diversity as an evaluation category as a proxy for fairness [22]. This review identified two primary articles which measured diversity through sample coverage [49,50], and four further papers which highlighted the importance of fairness metrics [24,47,51,52,53], but concluded the area needs more attention.
Only one review article focused on ethical and legal implications of health data re-use [14]. This 2023 scoping review found that health synthetic data generally falls within a grey area of regulation and that generation and sharing have generally been performed on a case-by-case basis with no clear answers to resulting legal and ethical questions. The authors note some have advocated for FAIR principles to be followed, and specifically advocate for CARE principles, which complement the FAIR principles by providing guidance for preserving the rights and interests of Indigenous people in data governance. Other articles acknowledge these issues as areas of relevance in their introductions, but did not explicitly include sections addressing the concerns within the reviews.
Discussion
This scoping review included twelve primary articles on synthetic data generation from metadata and eleven literature review articles on synthetic data evaluation and considerations. Existing reviews emphasised deep learning-based methods and generally did not include metadata-based methods: only Synthea was included within existing literature reviews [7,27,28]. All review articles were published between 2023 and 2025, indicating synthetic health data generation to be a developing field with ongoing research efforts. By contrast, only one of the articles concerning metadata-based method was published after 2023, showing limited recent research on this topic [38].
Evaluation practices for metadata-based methods were limited in scope compared to those described in the literature reviews. Fidelity was reported for four methods, while computation time and utility were each reported for two methods. No other metrics were used. Among the eleven total literature reviews, all but two either directly outlined evaluation metrics or referred readers to existing literature, confirming there is a strong body of existing literature regarding synthetic data evaluation frameworks [7,14,20,21,22,24,25,26,27,28]. However, none provided specific discussion of metrics or standards to use in the context of code development for federated analytical pipelines. The lack of evaluation among metadata-based methods and lack of specific guidance within the reviews means it is difficult to assess methods’ usability within the federated analysis context.
There is also a need for alignment among existing frameworks, both in the evaluation categories named and in how similar terms were defined. As shown in Table 4, only fidelity, utility, and privacy were included in all reviews. There was also disagreement among the meaning of similar terms; even the widely used terms fidelity and utility overlapped across frameworks [24,26]. The UK Synthetic Data Community Group and Duke Margolis have each defined and suggested evaluation categories, neither of which agreed with those identified in this review [5,15]. While some existing literature reviews noted this lack of agreement, there is still a need for standardisation through consensus [15,22,24]. Recent research is beginning to work towards this, including a recent Delphi consensus on a privacy metrics framework for synthetic data [54].
There was also a limited amount of research regarding ethical and legal considerations, as only three reviews and no metadata-based articles covered the topics of fairness or the ethicality and legality of personal health data re-use [14,22,24]. Addressing fairness is of particular concern as dataset bias can lead to the creation of algorithms which unfairly penalise minority groups, and synthetic data brings in unique concerns regarding consent and community engagement [14,24,55]. The EU recognises this within its Artificial Intelligence (AI) Act, which mandates that when it is necessary for ensuring bias detection and correction, providers may “exceptionally process special categories of personal data” when the bias detection and correction “cannot be effectively fulfilled by processing other data”, including synthetic data (among other conditions) [56]. However, there remains a lack of clear governance guidelines within EU regulation and existing research [14].
Some sections of the AI act affect synthetic data generation and use, such as the requirement that AI-generated synthetic data should be detectable as artificially generated or manipulated, and that training, validation, and testing data sets should be relevant and sufficiently representative for the intended purpose [56]. However, there is a need for specific guidelines as issues particular to synthetic data can get lost in broader discussion regarding AI data ethics [57]. Neither the FDA nor the EMA have published guidance on the use of synthetic data for regulatory purposes, highlighting a need for future regulation. Researchers should consider data type, patient consent, and public sharing as factors which affect privacy and data governance requirements [5,15]. Draft guidelines for the European Health Data Space make recommendations but also note open questions regarding privacy criteria and documentation [58].
Implication for Researchers and Future Research
Existing evaluation frameworks should apply to metadata-based generation methods, as evaluation depends on comparison with original datasets without regard for the generation method. Similarly, existing RWE quality frameworks such as those from ISPOR or HARPER may be used for synthetic data quality assurance as they likewise depend on the data itself rather than generation method [15]. However, researchers should be aware that synthetic data may be used across various stages of the research process with different requirements. For example, fidelity requirements for code development may be much lower than for inference [5]. Although Synthea can already be used to generate data in the OMOP common data model [18], most reviews (8/10) only considered the use of synthetic data for inference and Synthea is the only metadata-based method to be included within a benchmarking review [7,14,20,24,25,26,27,28]. There is a need for future research to apply existing evaluation metrics to benchmark metadata-based methods and resulting databases against each other.
While it is important to continue researching methods which produce high-fidelity data, future work should also focus on maximizing accessibility of lower-fidelity data as its application can greatly streamline RWE studies [11]. Additionally, the importance of aligning standards for fidelity, utility, privacy, and other evaluation metrics with focus on the intended use is clear [5]. Because no baseline standards for synthetic data currently exist, future research should define technical and legal requirements for synthetic data within different contexts, taking into account the prevalent use of synthetic data for purposes beyond inference [5].
Strengths of this study include the systematic design, focus on metadata-based methods, and harmonisation of existing evaluation dimensions. Limitations include the possibility of missed studies in private pre-print servers, as well as the lack of abstraction for non-metadata based primary articles which may have led to some missed discussion of legal or ethical considerations within these articles if it was not highlighted in an included literature review.
In conclusion, our scoping review shows that there is an existing body of work regarding metadata-based synthetic health data generation, but a lack of both applied evaluation and guidelines for assessing these methods within a federated analysis context. Existing evaluation metrics beyond fidelity assessment have rarely been applied to metadata-based methods, and use of synthetic data for code development is not discussed within existing literature reviews of synthetic health data generation. There is a need for consensus on a standardised framework for synthetic data evaluation and further exploration of governance, ethical and legal considerations, especially as these concerns are increasingly recognised within AI regulations. As standards may differ depending on intended use, future research should focus on developing context-specific guidelines aligned with both generation methods and use cases.
Methods
We used the Preferred Reporting Items for Systematic Reviews and meta-analysis Protocols for Scoping Reviews (PRISMA-SCR) (7) to guide reporting of the scoping review (Appendix S7) and preregistered the protocol with the Open Science Framework on 12 March 2025 (https://osf.io/9chu6).
Eligibility criteria: We included 1) peer-reviewed articles and pre-prints which reported methods for synthetic data generation, and 2) peer-reviewed articles and pre-prints which reviewed synthetic data evaluation and/or legal and ethical requirements within healthcare research from all years. Both articles reporting generation methods and literature reviews focused on synthetic health data were included in order to identify existing methods as well as potential evaluation frameworks and ethical or legal considerations. We excluded 1) articles regarding generation of non-EHR data, 2) articles describing the use of synthetic data as part of the methods to answer a specific research question, 3) articles not concerning human healthcare research, 4) articles without full-text access, and 5) articles in language other than English.
Relevance criteria: Articles were further screened for relevance. Primary articles were considered relevant if they met all eligibility criteria and focused on metadata-based generation methods. Secondary articles were considered relevant if they were systematic reviews. Both article types were considered relevant as primary articles presented methods and secondary articles presented potential evaluation frameworks and ethical or legal considerations which may be relevant to these methods.
Information Sources: PubMed was used to search PubMed (MEDLINE), Ovid Embase to search Embase, and Web of Science to search pre-print servers. The final search was performed on 18 March 2025 and included all articles through this date. No contact was made with authors to identify additional information sources.
Literature search: The search strategy was refined through initial searches, team discussion, and meeting with a research librarian (NJ). Grey literature was included and searched through Web of Science. The final search strategy included terms related to “synthetic data” alongside either “electronic health record” or “metadata”. The full search strategy for each database can be found in Appendix S2. Articles discovered during the research process through other sources were manually included.
Selection of sources of evidence: All search results were saved and uploaded to Rayyan for abstract and then full-text screening [59]. One reviewer (CR) screened all articles and another three (ACR, VH, CP) collectively screened the same articles. Disagreements were resolved through bilateral discussion or, when necessary, involvement of a third reviewer (CAN). To increase consistency between reviewers, five articles were selected at each stage for test screening to align decisions.
Data charting and data items: One reviewer (CR) performed data charting in Excel through an iterative process based on initial screening and full text review of included article. A second reviewer (CP) confirmed the data extraction results. Articles were first classified as primary articles (describing individual generation methods) or literature reviews. For all articles, characteristics including title, authors, year of publication and type of article (e.g. metadata-based method, non metadata-based method, systematic review) were charted. No contact was made with authors of included records. Metadata-based primary articles were then further charted to identify: input data, an overview of their methodologies, and evaluation criteria applied to the resulting synthetic data. From the literature reviews, the following items were abstracted: review purpose, generation methods included, evaluation dimensions (if any), definitions for each evaluation dimension (if any), and mention of review themes: privacy, fairness, legal requirements, and technical requirements.
Critical appraisal: The AMSTAR tool was used to informally assess included literature reviews based on whether the following items were stated: research question, inclusion/exclusion criteria, search string, PRISMA-ScR screening flow chart, databases searched, and number of reviewers (Appendix S8). No assessment was made for primary sources.
Synthesis of results: Primary articles and literature reviews were synthesised separately to distinguish between articles presenting methods and reviews presenting evaluations frameworks or ethical and legal considerations.
Declaration of generative AI and AI-assisted technologies in the manuscript preparation process
During the preparation of this work, the authors used ChatGPT-4.o to improve the readability of the text. After using this tool/service, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article.
Supplementary Materials
The following supporting information can be downloaded at the website of this paper posted on Preprints.org.
Acknowledgments
We would like to thank the UU Librarian Najoua Ryane for contributions to the search strategy for this project.
References
- EMA. Real-world evidence framework to support EU regulatory decision-making: 3rd report on the experience gained with regulator-led studies from February 2024 to February 2025 2025 [cited European Medicines Agency. Available online: https://www.ema.europa.eu/en/documents/report/real-world-evidence-framework-support-eu-regulatory-decision-making-3rd-report-experience-gained-regulator-led-studies-february-2024-february-2025_en.pdf.
- Natividad, M.; Sestak, I.; Luna, G.; Luijken, J.; Gibson, B.; Delaitre-Bonnin, C.; et al. HTA244 Evaluation of the Acceptance of Real-World Evidence (RWE) in Eunethta Joint Clinical Assessments (JCA) to Support Evidence for Reimbursement Decisions. Value Health 2023, 26(12), S367. [Google Scholar] [CrossRef]
- Heinz, S.; Kumari, C.; Ataide, J.; Bourakkadi, M.; Castellano, G. P22 Current and Future Use of RWE in HTA Decision-Making: Payers View Globally. Value Health 2024, 27(12), S6. [Google Scholar] [CrossRef]
- Gini, R.; Sturkenboom, M.C.J.; Sultana, J.; Cave, A.; Landi, A.; Pacurariu, A.; et al. Different Strategies to Execute Multi-Database Studies for Medicines Surveillance in Real-World Setting: A Reflection on the European Model. Clin. Pharmacol. Ther. 2020, 108(2), 228–35. [Google Scholar] [CrossRef] [PubMed]
- Hotchkiss, L.; Squires, E.; Oliver, E.; McCall, S.; Arora, A.; Magder, C.; Lugg-Widger, F.; Trubey, R.; Moore, S.; Rittman, T.; Gallacher, J.; Thompson, S. Understanding User Requirements of Synthetic Data 2025. Available from. [CrossRef]
- Gonzales, A.; Guruswamy, G.; Smith, S.R. Synthetic data in health care: A narrative review. PLoS Digit Health 2023, 2(1), e0000082. [Google Scholar] [CrossRef] [PubMed]
- Chen, X.; Wu, Z.; Shi, X.; Cho, H.; Mukherjee, B. Generating synthetic electronic health record data: a methodological scoping review with benchmarking on phenotype data and open-source software. J. Am. Med. Inf. Assoc. 2025, 32(7), 1227–40. [Google Scholar] [CrossRef] [PubMed]
- Wang, Z.; Myles, P.; Tucker, A. Generating and evaluating cross-sectional synthetic electronic healthcare data: Preserving data utility and patient privacy. Comput. Intell. 2021, 37(2), 819–51. [Google Scholar] [CrossRef]
- Nowok, B.; Raab, G.M.; Dibben, C. synthpop: Bespoke Creation of Synthetic Data in R. J. Stat. Softw. 2016, 74(11). [Google Scholar] [CrossRef]
- Drechsler, J. Differential Privacy for Government Agencies—Are We There Yet? J. Am. Stat. Assoc. 2023, 118(541), 761–73. [Google Scholar] [CrossRef]
- van Kesteren, E.J. To democratize research with sensitive data, we should make synthetic data more accessible. Patterns (N Y) 2024, 5(9), 101049. [Google Scholar] [CrossRef] [PubMed]
- ODISSEI. Metasyn Documentation 2024. Available online: https://metasyn.readthedocs.io/en/latest/.
- Schram, R.; Spithorst, S.; van Kesteren, E.-J. Metasyn: Transparent Generation of Synthetic Tabular Data with Privacy Guarantees. J. Open Source Softw. 2025, 10(105). [Google Scholar] [CrossRef]
- Tsao, S.F.; Sharma, K.; Noor, H.; Forster, A.; Chen, H. Health Synthetic Data to Enable Health Learning System and Innovation: A Scoping Review. Stud. Health Technol. Inf. 2023, 302, 53–7. [Google Scholar] [CrossRef] [PubMed]
- Hendricks-Sturrup RME, N.; Nafie, M. Synthetic Data Generation Using Generative AI to Support Biomedical Innovation: A Health Policy Perspective: Duke-Margolis Institute for Health Policy. 2025. Available online: https://lnkd.in/e2mjkgrw.
- Foraker, R.E.; Yu, S.C.; Gupta, A.; Michelson, A.P.; Pineda Soto, J.A.; Colvin, R.; et al. Spot the difference: comparing results of analyses from real patient data and synthetic derivatives. JAMIA Open 2020, 3(4), 557–66. [Google Scholar] [CrossRef] [PubMed]
- Choi EaB, Siddharth and Malin, Bradley and Duke, Jon and Stewart, Walter F. and Sun, Jimeng. Generating Multi-label Discrete Patient Records using Generative Adversarial Networks. Proc. 2nd Mach. Learn. Healthc. Conf. 2017, 68, 286–305.
- Walonoski, J.; Kramer, M.; Nichols, J.; Quina, A.; Moesel, C.; Hall, D.; et al. Synthea: An approach, method, and software mechanism for generating synthetic patients and the synthetic electronic health care record. J. Am. Med. Inf. Assoc. 2018, 25(3), 230–8. [Google Scholar] [CrossRef] [PubMed]
- Brix, T.J.; Becker, L.; Harbich, T.; Oehm, J.; Fechner, M.; Dugas, M.; et al. ODM Clinical Data Generator: Syntactically Correct Clinical Data Based on Metadata Definition. Stud. Health Technol. Inf. 2021, 278, 35–40. [Google Scholar] [CrossRef] [PubMed]
- Liu, Y.; Acharya, U.R.; Tan, J.H. Preserving privacy in healthcare: A systematic review of deep learning approaches for synthetic data generation. Comput Methods Programs Biomed. 2025, 260, 108571. [Google Scholar] [CrossRef] [PubMed]
- Ibrahim, M.; Khalil, Y.A.; Amirrajab, S.; Sun, C.; Breeuwer, M.; Pluim, J.; et al. Generative AI for synthetic data across multiple medical modalities: A systematic review of recent developments and challenges. Comput Biol. Med. 2025, 189, 109834. [Google Scholar] [CrossRef] [PubMed]
- Liu, K.; Altman, R.B. Conditional Generative Models for Synthetic Tabular Data: Applications for Precision Medicine and Diverse Representations. Annu Rev. BioMed Data Sci. 2025. [Google Scholar] [CrossRef] [PubMed]
- Kaabachi, B.; Despraz, J.; Meurers, T.; Otte, K.; Halilovic, M.; Kulynych, B.; et al. A scoping review of privacy and utility metrics in medical synthetic data. npj Digit Med. 2025, 8(1), 60. [Google Scholar] [CrossRef] [PubMed]
- Vallevik, V.B.; Babic, A.; Marshall, S.E.; Elvatun, S.; Brøgger, H.M.B.; Alagaratnam, S.; et al. Can I trust my fake data - A comprehensive quality assessment framework for synthetic tabular data in healthcare. Int. J. Med. Inform. 2024, 185, 105413. [Google Scholar] [CrossRef] [PubMed]
- Budu, E.; Etminani, K.; Soliman, A.; Rögnvaldsson, T. Evaluation of synthetic electronic health records: A systematic review and experimental assessment. Neurocomputing 2024, 603. [Google Scholar] [CrossRef]
- Murtaza, H.; Ahmed, M.; Khan, N.F.; Murtaza, G.; Zafar, S.; Bano, A. Synthetic data generation: State of the art in health care domain. Comput. Sci. Rev. 2023, 48. [Google Scholar] [CrossRef]
- Hernandez, M.; Epelde, G.; Alberdi, A.; Cilla, R.; Rankin, D. Synthetic data generation for tabular health records: A systematic review. Neurocomputing 2022, 493, 28–45. [Google Scholar] [CrossRef]
- Ghosheh, G.O.; Li, J.; Zhu, T. A Survey of Generative Adversarial Networks for Synthesizing Structured Electronic Health Records. ACM Comput. Surv. 2024, 56(6), 1–34. [Google Scholar] [CrossRef]
- Walonoski, J.; Klaus, S.; Granger, E.; Hall, D.; Gregorowicz, A.; Neyarapally, G.; et al. Synthea™ Novel coronavirus (COVID-19) model and synthetic data set. Intell. Based Med. 2020, 1, 100007. [Google Scholar] [CrossRef] [PubMed]
- Meeker, D.; Kallem, C.; Heras, Y.; Garcia, S.; Thompson, C. Case report: evaluation of an open-source synthetic data platform for simulation studies. JAMIA Open 2022, 5(3), ooac067. [Google Scholar] [CrossRef] [PubMed]
- Prasanna, A.; Jing, B.; Plopper, G.; Miller, K.K.; Sanjak, J.; Feng, A.; et al. Synthetic Health Data Can Augment Community Research Efforts to Better Inform the Public During Emerging Pandemics. medRxiv 2023. [Google Scholar] [CrossRef] [PubMed]
- Hodges, R.; Tokunaga, K.; LeGrand, J. A novel method to create realistic synthetic medication data. JAMIA Open 2023, 6(3), ooad052. [Google Scholar] [CrossRef] [PubMed]
- Schulz, N.A.; Carus, J.; Wiederhold, A.J.; Johanns, O.; Peters, F.; Rath, N.; et al. Learning debiased graph representations from the OMOP common data model for synthetic data generation. BMC Med. Res. Methodol. 2024, 24(1), 136. [Google Scholar] [CrossRef] [PubMed]
- Murray, R.E.; Ryan, P.B.; Reisinger, S.J. Design and validation of a data simulation model for longitudinal healthcare data. AMIA Annu Symp. Proc. 2011, 2011, 1176–85. [Google Scholar] [PubMed]
- Ayilara, O.F.; Platt, R.W.; Dahl, M.; Coulombe, J.; Ginestet, P.G.; Chateau, D.; et al. Generating synthetic data from administrative health records for drug safety and effectiveness studies. Int. J. Popul Data Sci. 2023, 8(1), 2176. [Google Scholar] [CrossRef] [PubMed]
- Khorchani, T.; Gadiya, Y.; Witt, G.; Lanzillotta, D.; Claussen, C.; Zaliani, A. SASC: A simple approach to synthetic cohorts for generating longitudinal observational patient cohorts from COVID-19 clinical data. Patterns (N Y) 2022, 3(4), 100453. [Google Scholar] [CrossRef] [PubMed]
- Davis, S.E.; Ssemaganda, H.; Koola, J.D.; Mao, J.; Westerman, D.; Speroff, T.; et al. Simulating complex patient populations with hierarchical learning effects to support methods development for post-market surveillance. BMC Med. Res. Methodol. 2023, 23(1), 89. [Google Scholar] [CrossRef] [PubMed]
- Barr, A.A.; Quan, J.; Guo, E.; Sezgin, E. Large language models generating synthetic clinical datasets: a feasibility and comparative analysis with real-world perioperative data. Front Artif. Intell. 2025, 8, 1533508. [Google Scholar] [CrossRef] [PubMed]
- Synthea Patient Generator (Github): Synthea. Available online: https://github.com/synthetichealth/synthea.
- OSIM-v5 (Github): OHDSI. 2018. Available online: https://github.com/OHDSI/OSIM-v5.
- ETL-Synthea: Utility to Load Synthea CSV data to OMOP CDM (Github): OHDSI. 2024. Available online: https://github.com/OHDSI/ETL-Synthea.
- SASC (Github). Fraunhofer ITMP Hamburg. 2021. Available online: https://github.com/Fraunhofer-ITMP/SASC/tree/v1.0.
- Sliman, H.; Megdiche, I.; Alajramy, L.; Taweel, A.; Yangui, S.; Drira, A.; et al. MedWGAN based synthetic dataset generation for Uveitis pathology. Intell. Syst. With Appl. 2023, 18, 200223. [Google Scholar] [CrossRef]
- Hameed, M.A.B.; Alamgir, Z. Improving mortality prediction in Acute Pancreatitis by machine learning and data augmentation. Comput. Biol. Med. 2022, 150, 106077. [Google Scholar] [CrossRef] [PubMed]
- Rodriguez-Almeida, A.J.; Fabelo, H.; Ortega, S.; Deniz, A.; Balea-Fernandez, F.J.; Quevedo, E.; et al. Synthetic patient data generation and evaluation in disease prediction using small and imbalanced datasets. IEEE J. Biomed. Health Inform. 2022, 27(6), 2670–80. [Google Scholar] [CrossRef] [PubMed]
- Ram, P.K.; Kuila, P. GAAE: a novel genetic algorithm based on autoencoder with ensemble classifiers for imbalanced healthcare data. J. Supercomput. 2023, 79(1), 541–72. [Google Scholar] [CrossRef]
- Bhanot, K.; Qi, M.; Erickson, J.S.; Guyon, I.; Bennett, K.P. The Problem of Fairness in Synthetic Healthcare Data; Entropy (Basel), 2021; 9, p. 23. [Google Scholar]
- Yoon, J.; Mizrahi, M.; Ghalaty, N.F.; Jarvinen, T.; Ravi, A.S.; Brune, P.; et al. EHR-Safe: generating high-fidelity and privacy-preserving synthetic electronic health records. npj Digit Med. 2023, 6(1), 141. [Google Scholar] [CrossRef] [PubMed]
- Alaa, A.; Van Breugel, B.; Saveliev, E.S.; Van Der Schaar, M. (Eds.) How faithful is your synthetic data? sample-level metrics for evaluating and auditing generative models. In International conference on machine learning; PMLR, 2022. [Google Scholar]
- Naeem, M.F.; Oh, S.J.; Uh, Y.; Choi, Y.; Yoo, J. (Eds.) Reliable fidelity and diversity metrics for generative models. In International conference on machine learning; PMLR, 2020. [Google Scholar]
- Van Breugel, B.; Van der Schaar, M. Beyond privacy: Navigating the opportunities and challenges of synthetic data. arXiv 2023, arXiv:230403722. [Google Scholar]
- Pereira, M.; Kshirsagar, M.; Mukherjee, S.; Dodhia, R.; Lavista Ferres, J.; De Sousa, R. Assessment of differentially private synthetic data for utility and fairness in end-to-end machine learning pipelines for tabular data. PLoS ONE 2024, 19(2), e0297271. [Google Scholar] [CrossRef] [PubMed]
- Qian, Z.; Davis, R.; van der Schaar, M. Synthcity: a benchmark framework for diverse use cases of tabular synthetic data; Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S., Eds.; 2023; p. 3173--88. [Google Scholar]
- Pilgram, L.; Dankar, F.K.; Drechsler, J.; Elliot, M.; Domingo-Ferrer, J.; Francis, P.; et al. A consensus privacy metrics framework for synthetic data. Patterns 2025. [Google Scholar] [CrossRef] [PubMed]
- Susser, D.; Schiff, D.S.; Gerke, S.; Cabrera, L.Y.; Cohen, I.G.; Doerr, M.; et al. Synthetic Health Data: Real Ethical Promise and Peril. Hastings Cent. Rep. 2024, 54(5), 8–13. [Google Scholar] [CrossRef] [PubMed]
- The EU Artificial Intelligence Act 2025. Available online: https://artificialintelligenceact.eu.
- Shanley, D.; Hogenboom, J.; Lysen, F.; Wee, L.; Lobo Gomes, A.; Dekker, A.; et al. Getting real about synthetic data ethics: Are AI ethics principles a good starting point for synthetic data ethics? EMBO Rep. 2024, 25(5), 2152–5. [Google Scholar] [PubMed]
- Data TSJATtEH, Space. M7.2 Draft guideline on data minimisation,pseudonymisation, anonymisation and synthetic data. 2025. Available online: https://tehdas.eu/wp-content/uploads/2025/09/draft-guideline-on-data-minimisation-pseudonymisation-anonymisation-and-synthetic-data.pdf.
- Ouzzani, M.; Hammady, H.; Fedorowicz, Z.; Elmagarmid, A. Rayyan-a web and mobile app for systematic reviews. Syst. Rev. 2016, 5(1), 210. [Google Scholar] [CrossRef] [PubMed]
Figure 1.
Selection Results.

Table 1.
Charting of Individual Literature Reviews.
| Year | Search Period* | Authors | Title | Purpose | Method types | Evaluation Dimensions |
| 2025 | 2019-2023 | Y. Liu et al [20] | Preserving privacy in healthcare: A systematic review of deep learning approaches for synthetic data generation. | Review generation methods | GAN, VAE, DMs | Data resemblance, Utility, Privacy |
| 2025 | Jan 2021-Nov 2023 | Ibrahim et al [21] | Generative AI for synthetic data across multiple medical modalities: A systematic review of recent developments and challenges. | Review generation methods | GANs, VAEs, DMs, LLMs | Utility, Fidelity, Diversity, Qualitative assessment, Clinical validation, and Privacy |
| 2025 | Jan 2019 - May 2024 | K. Liu et al [22] | Conditional Generative Models for Synthetic Tabular Data: Applications for Precision Medicine and Diverse Representations. | Review generation methods | CGMs | Utility, Resemblance, Privacy, Diversity, Computational Cost |
| 2025 | Jan 2019 – July 2024 | Kaabachi et al [23] | A scoping review of privacy and utility metrics in medical synthetic data | Review privacy and utility metrics | -- | Broad Utility, Narrow Utility, Privacy, Fairness |
| 2024 | 2019-Oct 2023 | Vallevik et al [24] | Can I trust my fake data - A comprehensive quality assessment framework for synthetic tabular data in healthcare. | Review evaluation metrics and propose evaluation framework | -- | Similarity, Usability, Privacy, Fairness, Carbon Footprint |
| 2024 | 2017-2022 | Budu et al [25] | Evaluation of synthetic electronic health records: A systematic review and experimental assessment | Review evaluation metrics | Deep learning: GAN, VAE statistical methods: Bayesian networks |
Fidelity, Utility, Privacy |
| 2025 | Until March 2025 | Chen et al [7] | Generating synthetic electronic health record data: a methodological scoping review with benchmarking on phenotype data and open-source software | Review generation methods and introduce package for benchmarking | GAN, VAE, diffusion, transformer, rule-based | Fidelity, Utility, Privacy, Computational Cost |
| 2023 | 2012-Dec 2022 | Tsao et al [14] | Health Synthetic Data to Enable Health Learning System and Innovation: A Scoping Review | Review data governance and legal requirements | -- | -- |
| 2023 | until Jul 2022 | Murtaza et al [26] | Synthetic data generation: State of the art in health care domain | Review generation methods | Classical: Mixture models, Copulas, Monte Carlo, Bayesian network, Kernel Density Estimation ML: GANs, Neural networks, autoencoders, decision trees |
Realism (Resemblance, Utility), Privacy |
| 2022 | Jan 2016 - May 2021 | Hernandez et al [27] | Synthetic data generation for tabular health records: A systematic review | Review generation methods | GANs | Resemblance, Utility, Privacy, Performance Dimensions |
| 2024 | until Jan 2022 | Ghosheh et al [28] | A Survey of Generative Adversarial Networks for Synthesizing Structured Electronic Health Records | Review generation methods | GANs | Similarity, Privacy preservation, Utility, Qualitative evaluation |
*Note: Search Period is specified with the same degree of precision as specified in each literature review Synthesis of Results.
Table 2.
Synthesised Results of Articles Presenting Metadata-Based Methods.
| Year | Publication Title(s) | Name | Review Papers | Input | Output | Evaluation |
| 1) 2018 2) 2020 3) 2022 4) 2023 5) 2023 6) 2024 |
1) Synthea: An approach, method, and software mechanism for generating synthetic patients and the synthetic electronic health care record [18] 2) Synthea‚ Novel coronavirus (COVID-19) model and synthetic data set [29] 3) Case report: evaluation of an open-source synthetic data platform for simulation studies [30] 4) Synthetic Health Data Can Augment Community Research Efforts to Better Inform the Public During Emerging Pandemics [31] 5) A novel method to create realistic synthetic medication data [32] 6) Learning debiased graph representations from the OMOP common data model for synthetic data generation [33] |
Synthea | Included. (Chen, Hernandez) | Public demographic information and module files based on public data describing a progression of states and transitions between them. Users can modify these as needed. | Patient-level dataset with demographics, clinical encounters (primary care, emergency room, symptom-driven encounters), conditions, medications, vaccinations, observations, labs, procedures, outcomes. Format can be FHIR, CSV, TXT, or CCDA. | 1) Fidelity: comparison of generated vs publicly available statistical and treatment properties 2) Fidelity: comparison of statistics between generated vs primary data sources 3) Fidelity: comparison of distributions between generated vs actual study, Utility: comparison of results between generated study vs actual study which was replicated 4) Fidelity: dataset distributions on age/sex alongside number of total records and other demographics 5) Fidelity: distribution of asthma medications in patients ages 0-5 6) N/A: article evaluates Synthea graph results, not synthetic data |
| 2021, 2023 | 1) Design and Validation of a Data Simulation Model for Longitudinal Healthcare Data [34] 2) Generating synthetic data from administrative health records for drug safety and effectiveness studies [35] |
OSIM | Not included | Source dataset or summary tables. | Patient-level dataset in OMOP format with tables person, observation_period, condition_era, drug_era, and procedure_occurrence. | 1) Computation time; Fidelity: similarity of summary statistics, demographic distributions, number of distinct conditions/condition occurrences, condition (co)-occurrence, prevalence, number of drugs and drug occurrences, drug prevalence, condition/drug co-occurrence, drug/drug co-occurrence, condition timing, drug timing, prevalence agreement with gender 2) Fidelity: similarity of demographic distribution, prescription drug use measures, and health conditions |
| 2021 | ODM Clinical Data Generator: Syntactically Correct Clinical Data Based on Metadata Definition [19] | -- (Brix, et al) |
Not included | ODM file. | Syntactically correct patient-level dataset without fidelity to original data. | Computation time |
| 2022 | SASC: A simple approach to synthetic cohorts for generating longitudinal observational patient cohorts from COVID-19 clinical data [36] | SASC | Not included | Source dataset or summary tables. | CSV or JSON patient-level dataset with patient demographics, visits, and lab results. | Fidelity: tables showing summary statistics for real vs generated cohort. boxplot comparisons for each generated parameter Utility: correlation plot of correlations of variables to gender and hospitalisation time in real vs generated cohort. Comparison of survival analysis of age and gender in both cohorts |
| 2023 | Simulating complex patient populations with hierarchical learning effects to support methods development for post-market surveillance [37] | -- (Davis, et al) |
Not included | Source dataset or summary tables. | Patient-level dataset which also simulates learning effects over time. | Fidelity: boxplots of agreement between distributions of patient features, treatment prevalence, outcome prevalence, treatment effect estimate |
| 2025 | Large language models generating synthetic clinical datasets: a feasibility and comparative analysis with real-world perioperative data. [38] | -- (Barr, et al) |
Not included | Descriptive statistics and conditions for desired variables. | Patient-level dataset. | Fidelity: Compared statistical similarity (means and CIs for parameters) with real dataset. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.