Submitted:
22 September 2026
Posted:
23 September 2026
You are already at the latest version
Abstract
Synthetic data are increasing their popularity in different disciplines. Indeed, in addition to their affordability, synthetic data may play an effective protective role to address privacy issues, which are extremely critical in certain application domains such as healthcare, as well as they represent a relatively simple solution to overcome access restrictions. However, despite synthetic data are not a novelty and their potentialities in the different domains have been largely discussed, their operationalization and utility are still object of research. More recently, the unprecedented AI capabilities have naturally enabled and quickly consolidated a feedback loop, where synthetic data is, at the same time, a valuable resource to further develop and refine technology, as well as an added, often critical, value within the different application domains. By combining hybrid retrieval and analysis techniques, 250+ relevant papers have been selected, classified and analyzed. In a context of discussion about legal, ethical and political implications amid opportunities and challenges, the review has clearly pointed out a need for a more consolidated engineering approach to synthetic data as a response to the intrinsic complexity of applications in the current technological climate.
Keywords:
synthetic data
; machine learning
; LLMs
; data engineering
; requirement modeling
; data governance
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.