Preprint
Review

This version is not peer-reviewed.

AI-Enabled Digital Twins in Healthcare: Epistemic Foundations, Principles, Progress, and Future Directions

Submitted:

16 July 2026

Posted:

21 July 2026

You are already at the latest version

Abstract
Digital twins have rapidly entered healthcare discourse, accompanied by ambitious claims regarding personalization, prediction, and decision support. Originating in engineering domains characterized by well-observed, controllable, and mechanistically understood systems, the concept is increasingly being applied to biological, clinical, and organizational settings whose underlying assumptions differ substantially. As a result, the term now encompasses a heterogeneous collection of models, simulations, and artificial intelligence (AI) systems, often without clear specification of what is being represented, what can be inferred, or what actions can legitimately be justified.This review develops a scale-aware framework for evaluating healthcare digital twins. Through a structured analysis spanning cardiovascular disease, diabetes, Alzheimer's disease, maternal–fetal health, addiction, radiology, and hospital operations, we argue that the scientific validity and practical utility of a digital twin depend fundamentally on scale. As applications move from mechanistic physiological systems to individual patients, coupled human systems, sociotechnical settings, and healthcare organizations, the central challenges shift from model fidelity and state estimation to non-identifiability, behavioral adaptation, contested objectives, and governance.Across domains, recurring failure modes arise less from insufficient computational capability than from misalignment among state definitions, uncertainty representation, causal support, decision authority, and governance. We examine the roles of AI in state estimation, prediction, uncertainty quantification, and decision support while highlighting its limits in addressing poorly defined state spaces, non-identifiability, causal ambiguity, and normative questions regarding acceptable actions. We further argue that governance is not an external constraint but a foundational component of healthcare digital twins, shaping accountability, lifecycle oversight, data stewardship, and the legitimacy of decisions informed by digital twin outputs.This review reframes healthcare digital twins not as a single technology but as a family of computational representations whose validity depends on scale, context, purpose, and governance. Future progress will depend less on increasingly sophisticated models than on maintaining alignment among representation, evidence, intervention, and accountability, and on ensuring digital twin ambitions remain commensurate with the epistemic conditions and institutional constraints of the systems they seek to represent. Building on this analysis, the review offers complementary research roadmaps for AI and governance that emphasize explicit state semantics, uncertainty-aware inference, decision-centered evaluation, adaptive oversight, and governance-by-design.
Keywords: 
;  ;  ;  ;  

1. Introduction

The concept of the digital twin has gained remarkable traction in healthcare over the past decade. Originally developed in aerospace, manufacturing, and infrastructure management, digital twins are computational counterparts of physical systems that are continuously updated with data to forecast behavior, monitor state, detect deviations, diagnose complications, and support decision-making [1,2,3,4,5,6,7]. In healthcare, the concept has been enthusiastically adopted to describe systems promising personalized prediction, real-time monitoring, and optimized intervention across a wide range of clinical and operational domains.
Figure 1. Visual abstract: Digital twins in healthcare as bounded, scale-dependent epistemic instruments. Digital twins are increasingly applied across healthcare domains spanning physiological subsystems, individual patients, coupled human systems, healthcare organizations, and broader sociotechnical contexts. Unlike the engineered systems in which digital twins originated, healthcare applications are characterized by partial observability, limited identifiability, constrained controllability, and complex governance requirements. Across domains, recurring failures arise from epistemic and governance misalignment. Artificial intelligence and governance emerge as complementary foundations of trustworthy healthcare digital twins, motivating parallel research roadmaps focused on state representation, uncertainty handling, mechanistic–AI integration, accountability, lifecycle oversight, and data stewardship.
Figure 1. Visual abstract: Digital twins in healthcare as bounded, scale-dependent epistemic instruments. Digital twins are increasingly applied across healthcare domains spanning physiological subsystems, individual patients, coupled human systems, healthcare organizations, and broader sociotechnical contexts. Unlike the engineered systems in which digital twins originated, healthcare applications are characterized by partial observability, limited identifiability, constrained controllability, and complex governance requirements. Across domains, recurring failures arise from epistemic and governance misalignment. Artificial intelligence and governance emerge as complementary foundations of trustworthy healthcare digital twins, motivating parallel research roadmaps focused on state representation, uncertainty handling, mechanistic–AI integration, accountability, lifecycle oversight, and data stewardship.
Preprints 223586 g001
Despite this rapid adoption, substantial ambiguity remains about what a healthcare digital twin is, what it can legitimately claim to represent, and under what conditions it can responsibly inform action. In the current literature, the term is applied to an eclectic mix of mechanistic simulations, machine learning predictors, hybrid models, and even generative AI systems. In many cases, the presence of an AI model — particularly one operating on large or multimodal datasets — is implicitly treated as sufficient to justify the digital twin label. This conflation obscures the deeper epistemic commitments that distinguish a digital twin from a conventional predictive model.
A digital twin is not simply a predictive model, a simulation, or an AI system operating on large datasets. At its core, it is an epistemic instrument: a computational representation that sustains a claim of correspondence to a specific real-world system over time. The legitimacy of that claim depends on explicit assumptions regarding state, observation, updating, uncertainty, and intervention. When those assumptions are weakly specified or violated, the digital twin metaphor ceases to clarify and begins to obscure.
The central argument of this review is that the scientific validity and practical utility of healthcare digital twins fundamentally depend on scale. As digital twins move from physiological systems to individual patients, coupled human systems, healthcare organizations, and broader sociotechnical settings, the assumptions required for successful twinning degrade in systematic ways. What can be observed, inferred, predicted, and controlled changes across scales, as do the epistemic and governance requirements needed to justify digital twin claims.
At the same time, it is important to recognize a fundamental limitation that is often left implicit: in many healthcare settings, a strict digital twin—in the sense of a continuously synchronized, decision-complete replica of an individual patient or organization—is not attainable even in principle. This conclusion is consistent with foundational assessments of digital twin feasibility, which argue that for complex, weakly observable, and socially embedded systems, the assumptions required for high-fidelity twinning break down irreducibly rather than incrementally [8]. Limits imposed by partial observability, non-identifiability, care-mediated measurement, and irreducible uncertainty constrain what any model—whether mechanistic, statistical, or AI–based—can faithfully represent.
In such settings, what can be constructed is better understood, at best, as a digital cousin: a computational representation that remains systematically coupled to its real-world referent, supports bounded inference or task planning, and is explicitly scoped to particular decisions rather than claiming exhaustive or isomorphic representation. The distinction is important because it shifts attention from aspirations of complete replication to questions of scope, validity, and appropriate use.
This problem is especially acute in healthcare. Biological and clinical systems are only partially observable; observations are often incomplete and inconsistent; interventions are indirect and delayed; and outcomes are shaped by social, behavioral, and institutional factors. Unlike engineered systems, patients and healthcare organizations adapt to monitoring and intervention, measurement processes are confounded by care pathways, and latent states are frequently composite, subjective, or non-identifiable. These features do not merely complicate digital twin construction; they fundamentally alter what digital twins—or their more realistic digital cousins—can represent, infer, predict, and legitimately influence. This conclusion aligns with findings from a recent National Academies consensus study [8], which identifies fundamental limits on observability, validation, uncertainty propagation, and governance as intrinsic constraints on digital twin fidelity rather than problems resolvable through additional data or model complexity.
The consequences of scale extend beyond differences in model complexity or data availability. As digital twins move from cellular and organ-level systems to individual humans, coupled human systems, sociotechnical phenomena, and healthcare organizations, the assumptions required for successful twinning degrade in systematic ways. Observability weakens, identifiability collapses, controllability diminishes, and governance constraints intensify. These transitions are not gradual degradations of fidelity. They are epistemic phase changes that alter what can be known, inferred, predicted, and controlled. Consequently, they require fundamentally different approaches to modeling, validation, and governance, consistent with formal analyses of the scale-dependent breakdown of the assumptions underlying digital twins [8]. As digital twins move from cellular and organ-level systems to individual humans, coupled human systems, sociotechnical phenomena, and healthcare organizations, the assumptions required for successful twinning degrade in systematic ways. Observability weakens, identifiability collapses, controllability diminishes, and governance constraints intensify. These transitions are not gradual degradations of fidelity. They are epistemic phase changes that alter what can be known, inferred, predicted, and controlled. Consequently, they require fundamentally different approaches to modeling, validation, and governance, consistent with formal analyses of scale-dependent breakdown of the assumptions underlying digital twins [8].
To make this argument concrete, we examine digital twins across a diverse set of healthcare domains. Cardiovascular disease and diabetes represent comparatively favorable settings in which mechanistic structure, dense sensing, and actionable interventions support stateful modeling, while still exposing limitations around uncertainty, safety, and decision alignment. Alzheimer’s disease highlights the dominance of non-identifiability at the scale of the individual human, constraining digital twins to roles centered on monitoring and planning rather than optimization. Maternal–fetal health reveals how coupled systems, asymmetric risk, and equity considerations sharply narrow the space of responsible application. Radiology exposes the dangers of equating rich observations with well-defined states, while addiction serves as a boundary case in which subjectivity, agency, and social embedding destabilize the digital twin metaphor itself. Hospital operations, by contrast, represent a domain where digital twins most closely resemble their engineering antecedents, yet where optimization is inseparable from normative and governance choices.
Across these domains, we identify recurring failure modes, including false precision, observation-dominant modeling, behavioral confounding, misaligned optimization, and governance blind spots that arise when digital twins are deployed without scale-sensitive epistemic discipline. We show that artificial intelligence, while indispensable, often amplifies these failures when treated as a substitute for clarity about state, uncertainty, causality, and decision authority.
Building on this analysis, we articulate complementary, scale-aware research roadmaps for artificial intelligence and governance in healthcare digital twins. Rather than advocating ever larger models or broader deployment, these roadmaps emphasize explicit state semantics, principled handling of irreducible uncertainty, causal validity, decision-centered evaluation, and lifecycle governance. Progress, we argue, should be measured not by ambition or coverage, but by the discipline with which digital twin claims are bounded.
The contribution of this review is not a new digital twin architecture, algorithm, or deployment framework. Rather, it is a scale-aware account of what healthcare digital twins can legitimately represent, infer, predict, and influence. We argue that digital twins should be understood as bounded epistemic instruments whose validity depends on the alignment of representation, evidence, intervention, and governance. Recognizing where digital twins can support inference and decision-making, where they should remain descriptive, and where they should be treated as constrained digital cousins — or not deployed at all—is not a retreat from innovation. It is a prerequisite for scientific rigor, clinical trust, and the responsible deployment of healthcare AI.

2. Primer: Digital Twins in Healthcare

The term digital twin has entered healthcare discourse with remarkable speed, often accompanied by expansive claims about personalization, prediction, and clinical decision support. Yet, for many artificial intelligence researchers, biomedical scientists, and clinicians encountering the concept for the first time, it remains unclear whether a digital twin represents a genuinely new scientific construct or a rebranding of existing modeling approaches. This ambiguity is not accidental. The digital twin concept did not originate in medicine, nor was it designed for the epistemic conditions characteristic of biological and clinical systems—a point emphasized across recent conceptual and critical reviews [9,10,11].
Digital twins emerged in engineering domains concerned with complex, safety-critical systems, including aerospace and advanced manufacturing, where direct experimentation on physical systems is costly, risky, or infeasible. In these settings, a digital twin denotes a computational counterpart of a specific physical system that is continuously updated using observational data and is used to forecast future states, diagnose deviations, and support operational decisions. Crucially, the digital twin is not merely a model but part of a closed-loop epistemic system linking sensing, inference, prediction, and action. Analyses of efforts to extend this paradigm to humans emphasize that such systems presuppose levels of observability, controllability, and mechanistic understanding that may not hold in biological systems [9,12,13].
Figure 2. A conceptual architecture of an AI-powered healthcare digital twin. A healthcare digital twin integrates multimodal observations—including clinical, imaging, sensor, contextual, and workflow data—to maintain a continuously updated representation of latent patient state. Mechanistic and AI-based models support state inference, uncertainty quantification, forecasting, scenario analysis, and decision support. Clinical actions and outcomes generate new observations that are assimilated through repeated cycles of updating and learning. Physicians remain central to interpretation and decision-making, while governance, context, uncertainty, and longitudinal adaptation shape the responsible use of digital twins across healthcare domains.
Figure 2. A conceptual architecture of an AI-powered healthcare digital twin. A healthcare digital twin integrates multimodal observations—including clinical, imaging, sensor, contextual, and workflow data—to maintain a continuously updated representation of latent patient state. Mechanistic and AI-based models support state inference, uncertainty quantification, forecasting, scenario analysis, and decision support. Clinical actions and outcomes generate new observations that are assimilated through repeated cycles of updating and learning. Physicians remain central to interpretation and decision-making, while governance, context, uncertainty, and longitudinal adaptation shape the responsible use of digital twins across healthcare domains.
Preprints 223586 g002
When the digital twin metaphor migrated into healthcare, it responded to genuine needs: longitudinal disease risk prediction and management, personalization beyond population averages, and integration of increasingly heterogeneous data sources. Recent surveys of medical and healthcare digital twins emphasize these motivations while acknowledging that most current implementations remain conceptual, preclinical, or limited to narrow use cases [14,15,16,17]. Ethical and sociotechnical analyses further caution that the rhetorical power of the digital twin metaphor risks outpacing its scientific grounding, particularly when systems are framed as holistic representations of patients without adequate epistemic justification or governance safeguards [18,19,20].
For researchers trained in artificial intelligence, a common source of confusion lies at the boundary between a digital twin and a predictive machine learning model. Much contemporary AI in healthcare focuses on estimating outcome probabilities from historical data. Such models are typically trained offline, evaluated retrospectively, and deployed as relatively static decision aids. Even when they operate on longitudinal data, they are rarely embedded in explicit feedback loops linking observations, inferences, actions, and outcomes. Consequently, many systems labeled as digital twins function in practice as retrospective predictors augmented with aspirational framing rather than as dynamically coupled clinical instruments [21,22,23].
A digital twin, by contrast, is intrinsically prospective. Its defining feature is not prediction per se, but the maintenance of an evolving internal representation of a specific system. Predictions arise as conditional forecasts of future trajectories under uncertainty, often with the explicit aim of exploring hypothetical scenarios or informing decisions. This distinction is critical because retrospective predictive performance alone is insufficient to establish clinical utility in dynamic, real-world environments—a point emphasized in recent critical and implementation-focused reviews of healthcare digital twins [15,24].
Digital twins are also frequently conflated with mechanistic or physiological models, such as cardiovascular flow simulations, pharmacokinetic models, or disease progression equations. These models play a foundational role in biomedical research, but they are typically population-level, calibrated infrequently, focused on components of a larger system, and used primarily for explanation or simulation rather than real-time decision support. Digital twin frameworks in cardiovascular medicine explicitly frame their contribution as a shift from static simulation toward patient-specific inference and continuous updating, combining narrow components into larger system representations and enabling continual assimilation of new observations through AI-based surrogates and data assimilation techniques [11,25,26].
This shift from static modeling to dynamic inference has significant epistemic consequences. Mechanistic fidelity alone does not guarantee clinical usefulness if a model cannot be reliably aligned with an individual patient or updated as new evidence, physiological changes, or contextual changes emerge. Conversely, data-driven personalization without mechanistic constraints risks overfitting and overconfidence, particularly when individual-level data are sparse, biased, care-mediated, or collected at different time scales. Hybrid digital twins—combining mechanistic structure with AI-based components—are increasingly proposed as a way to navigate this tension, especially in cardiovascular and maternal–fetal health contexts [27,28,29]. However, hybridization introduces its own challenges, including error propagation, regime failure, and ambiguity about which modeling assumptions are responsible for specific inferences [15,23].
These distinctions ultimately reflect a deeper issue. A digital twin is not defined by the presence of artificial intelligence, mechanistic knowledge, or predictive capability alone. Rather, it is defined by the maintenance of an evolving representation of latent system state under uncertainty. This perspective motivates a minimal formalism that captures the core epistemic commitments shared by most healthcare digital twin architectures.

A Minimal Formalism for Digital Twin Dynamical Systems

The conceptual commitments implicit in these discussions can be made more precise using a light formalism. At an abstract level, a healthcare digital twin can be understood as a stateful dynamical system coupled to a real-world entity. Let x t denote the latent state of the system at time t. In a patient-centered digital twin, this state may represent disease burden, physiological reserve, treatment response capacity, or a composite of such constructs. Crucially, x t is not directly observable.
Clinical measurements provide only partial, noisy, and behaviorally mediated observations of this state. Observations at time t, denoted y t , can be written as
y t = g ( x t , u t , ϵ t ) ,
where u t represents clinical actions or interventions, such as medications, procedures, or diagnostic tests, and ϵ t captures measurement noise and unmodeled influences. In healthcare, the observation process g is often poorly specified and strongly confounded by care pathways and institutional practices, a limitation particularly evident in radiology and neurodegenerative disease modeling [21,22,30].
The evolution of the latent state is governed by a transition process
x t + 1 = f ( x t , u t , η t ) ,
where f encodes disease progression, recovery, or deterioration dynamics, and η t represents stochastic influences and modeling uncertainty. Mechanistic digital twins specify f using physiological or biophysical models, as in cardiovascular digital twin frameworks [25,26]. Data-driven twins approximate f using learned dynamics, while hybrid twins combine mechanistic structures with AI-based surrogates [11,27,29].
Within this framing, a digital twin does not “know” the true state x t . Instead, it maintains a belief distribution over possible states, updated as new observations arrive:
p ( x t y 1 : t , u 1 : t ) .
Inference corresponds to updating this belief distribution rather than producing a single point estimate. This perspective clarifies why uncertainty quantification is not optional but a defining requirement of credible digital twins, a point repeatedly emphasized in recent critical reviews [15,16,24].
This belief-based view also exposes a central limitation of healthcare digital twins: identifiability. Multiple latent trajectories may be consistent with the same observation history. The problem is especially acute in slowly evolving diseases such as Alzheimer’s disease, where biomarkers, cognitive scores, and imaging findings do not uniquely determine disease state or progression [31,32,33]. In such settings, additional observations alone may be insufficient to resolve uncertainty without stronger assumptions or interventional evidence.
Decision-making enters the digital twin framework through a policy that maps beliefs about the current state to actions:
u t = π p ( x t y 1 : t , u 1 : t 1 ) .
Here, π encodes clinical objectives, constraints, and risk preferences. Many proposed healthcare digital twins stop short of specifying such a policy, leaving it unclear whether the twin is intended to support monitoring, recommendations, or optimization. Ethical, legal, and regulatory analyses stress that once a digital twin informs actions rather than merely recommending further observations, it assumes responsibility and must be evaluated accordingly [18,20,34,35].

Implications for Evaluation and Governance

This formal perspective clarifies why the validation of digital twins cannot be reduced to predictive accuracy. A digital twin may generate accurate short-term forecasts of observations while remaining unsafe or ineffective as a decision-support system, particularly over longer horizons, if uncertainty is mischaracterized, identifiability assumptions are violated, or inferred states are poorly aligned with clinically meaningful outcomes. This misalignment is evident in radiology, where high-performing predictive models often fail to demonstrate downstream clinical impact [21,30], and in Alzheimer’s disease, where uncertainty about the disease state undermines intervention planning [22,32].
Taken together, these considerations underscore a central message for AI researchers, biomedical scientists, and clinicians: a digital twin is not simply a more sophisticated model. It is a claim about what can be known, inferred, and acted upon in a clinical setting over time. Such claims demand explicit articulation of assumptions, rigorous validation aligned with clinical intent, and governance mechanisms commensurate with their potential impact [15,24,35].
The central implication is that a healthcare digital twin is fundamentally an epistemic instrument rather than a predictive engine. Its credibility depends not only on computational sophistication but also on alignment among latent state definition, observability, actionability, uncertainty representation, and decision intent. The domain-specific analyses that follow are therefore best understood as stress tests of these epistemic requirements. We begin with cardiovascular medicine, widely regarded as one of the most favorable clinical settings for digital twins, before turning to diabetes, Alzheimer’s disease, adverse maternal–fetal health outcomes, radiology, addiction, hospital operations, and broader healthcare systems.

3. Cardiovascular Digital Twins

Cardiovascular disease is often presented as the prototypical healthcare digital twin application. Among major disease domains, the cardiovascular system combines several characteristics favorable for digital twinning: comparatively well-characterized physiology, established structure–function relationships, actionable interventions, and growing streams of longitudinal data spanning clinical encounters, biomarkers, imaging, electrophysiology, wearable sensors, and environmental context. Consequently, cardiovascular medicine is frequently portrayed as the domain in which digital twins are most likely to achieve clinically meaningful impact [11,25,26].
Yet cardiovascular medicine is instructive precisely because it represents a best-case scenario. Despite decades of mechanistic modeling and rapid advances in AI-driven prediction, the translation of cardiovascular digital twins into routine clinical decision-making remains limited. The domain therefore provides a valuable stress test of the digital twin paradigm itself: if substantial challenges remain in a setting with relatively strong mechanistic understanding and rich observational data, those challenges are likely to be amplified elsewhere [14,15].
Figure 3. A conceptual architecture of a patient-specific cardiovascular digital twin. Multimodal observations from a patient—including clinical measurements, imaging, and wearable or electrophysiological signals—are integrated with mechanistic and AI-based models to infer latent cardiovascular state. The digital twin is continuously updated as new data arrive and can be used to forecast future trajectories, evaluate intervention scenarios, estimate uncertainty, and support clinical decision-making. Outputs from the digital twin can guide acquisition of new observations, enabling an iterative cycle of measurement, inference, prediction, and care.
Figure 3. A conceptual architecture of a patient-specific cardiovascular digital twin. Multimodal observations from a patient—including clinical measurements, imaging, and wearable or electrophysiological signals—are integrated with mechanistic and AI-based models to infer latent cardiovascular state. The digital twin is continuously updated as new data arrive and can be used to forecast future trajectories, evaluate intervention scenarios, estimate uncertainty, and support clinical decision-making. Outputs from the digital twin can guide acquisition of new observations, enabling an iterative cycle of measurement, inference, prediction, and care.
Preprints 223586 g003

3.1. What Is the Cardiovascular Digital Twin a Twin Of?

In cardiovascular medicine, the object of the digital twin is rarely the patient in a holistic sense. Instead, most proposed twins focus on a subset of cardiovascular state variables that are presumed to be both clinically meaningful and inferable from available data. These include latent constructs such as myocardial contractile function, ventricular remodeling, vascular compliance, electrophysiological stability, volume status, and propensity for near-term decompensation or adverse events [25,26].
Formally, the latent state x t represents a composite of physiological and pathological variables that cannot be directly observed but are clinically consequential. These may include myocardial contractile function, ventricular remodeling, vascular compliance, electrophysiological stability, volume status, and propensity for near-term decompensation or adverse events [25,26]. Observations y t —including vital signs, laboratory values, ECG signals, imaging-derived phenotypes, and wearable measurements—provide only partial and irregular views of this state, while actions u t include medications, procedures, device therapies, behavioral interventions, and changes in monitoring intensity.
This framing immediately reveals a key requirement: a cardiovascular digital twin must operate under partial observability, treatment-induced confounding, and heterogeneous data fidelity. The apparent maturity of cardiovascular digital twins therefore reflects not the absence of epistemic challenges but the presence of relatively strong mechanistic priors that partially constrain inference. Design-oriented analyses emphasize that this distinction is essential for understanding both the promise and limits of cardiovascular digital twins [11,27].

3.2. Mechanistic Foundations and the Role of Hybrid Models

The cardiovascular domain has a long history of mechanistic modeling, including biophysical models of cardiac electrophysiology, hemodynamics, and myocardial mechanics. Digital twin frameworks in cardiology explicitly build on this tradition, arguing that patient-specific instantiations of such models—continuously updated with data—can support personalized diagnosis and therapy planning [25].
However, fully mechanistic personalization is computationally expensive and data-intensive. As a result, most contemporary cardiovascular digital twins rely on hybrid architectures, in which mechanistic models provide structural constraints while AI-based components serve as surrogates, estimators, or adapters. Reviews of health and cardiovascular digital twins emphasize this hybridization as both inevitable and desirable [11,26,29]. Design-oriented analyses further argue that digital twins should be treated as engineered systems with explicit component boundaries, update mechanisms, and uncertainty flows [27].
Hybridization mitigates some limitations of purely mechanistic or purely data-driven approaches, but it also introduces new epistemic risks. When AI surrogates approximate mechanistic simulations, approximation error is rarely propagated to downstream clinical decisions. A surrogate that is accurate on average may still be unreliable in clinically salient regions of the state space—precisely where decisions about escalation or intervention are most consequential. This concern is echoed in broader critiques of AI-enabled digital twins that warn against surrogate-driven false precision [15,23].

3.3. State Inference, Updating, and Confounding by Care

A defining feature of cardiovascular digital twins is the promise of continuous updating: as new data arrive, the twin’s estimate of the latent cardiovascular state is refined. In practice, however, this updating process is deeply entangled with care delivery itself. Patients who deteriorate are monitored more closely, imaged more frequently, and treated more aggressively. Consequently, the observation process is confounded by prior actions, and naive updating risks learning patterns of care rather than underlying physiology.
This confounding-by-care problem is not unique to cardiovascular medicine, but the richness of cardiovascular data can obscure it. High-frequency wearable signals and dense imaging do not automatically translate into better state inference if the data-generating process is biased or treatment-dependent. Recent reviews emphasize that failure to model treatment-dependent observation processes is a major barrier to credible digital twin updating and state prediction across healthcare domains [15,24].

3.4. Personalization: Strength and Fragility

Personalization is frequently presented as the defining advantage of cardiovascular digital twins. By conditioning on an individual’s history, imaging, and physiological signals, a digital twin aims to forecast trajectories and tailor interventions more effectively than population-level models [11,26].
Yet personalization is also a source of fragility. Individual-level data are often sparse relative to model complexity, making personalization difficult to distinguish from overfitting. Models calibrated within a particular institution, imaging protocol, or device ecosystem may also fail under distribution shift, limiting transportability and clinical reliability. Reviews of digital twin deployment consistently identify robustness and external validity as major unresolved challenges [23,24].

3.5. Evidence and Validation: From Prediction to Decision Support

Despite current interest, cardiovascular digital twins remain thinly validated relative to their ambitions. Most published work demonstrates feasibility, retrospective predictive performance, or simulation accuracy, but stops short of demonstrating that digital twin outputs improve clinical decisions or outcomes when deployed prospectively. This gap mirrors broader critiques in the digital twin literature, which emphasize that validation must be aligned with declared clinical intent rather than generic accuracy metrics [14,15,24].
The cardiovascular domain, therefore, illustrates a broader lesson. Accurate prediction is not equivalent to clinically useful decision support. Demonstrating that a digital twin forecasts outcomes more accurately than existing models is fundamentally different from demonstrating that its recommendations improve patient outcomes, clinician decision-making, or healthcare efficiency. Bridging this gap remains one of the most important unresolved challenges in healthcare digital twins.

3.6. Domain-Specific Research Gaps

A critical synthesis of the cardiovascular digital twin literature reveals several gaps that are not merely technical but also structural. First, there is no systematic framework for propagating uncertainty from hybrid models to clinical decisions. AI surrogate error, measurement noise, and confounding effects are typically assessed in isolation, if at all [15,23]. Second, personalization lacks robustness guarantees. Few cardiovascular digital twins explicitly evaluate how individualized models behave under distribution shifts, missing data, or device heterogeneity, despite repeated calls for such analyses in digital twin reviews [11,24]. Third, causal reasoning is largely absent. Most twins forecast trajectories under observed care patterns but cannot reliably answer counterfactual questions about alternative interventions, limiting their usefulness for therapy optimization [15]. Finally, evidence pathways remain misaligned with clinical claims. Prospective studies assessing decision impact, safety, clinician trust, and workflow integration remain rare despite repeated calls for such evaluations in both cardiovascular and general healthcare digital twin literatures [14,24].
The cardiovascular case is instructive because it represents one of the most favorable environments for healthcare digital twins. Mechanistic understanding is comparatively strong, clinically meaningful state variables are partially observable, interventions are available, and rich longitudinal data exist. Yet even here, substantial challenges remain in state inference, uncertainty quantification, personalization, causal reasoning, and decision validation. These limitations become even more pronounced in domains where latent states are less observable, interventions are less controllable, or human behavior becomes part of the system itself.

3.7. A Cardiovascular Digital Twin Roadmap

Near-term progress should focus on explicit state representation, uncertainty quantification, and decision-bounded applications. Digital twins should be evaluated within narrowly defined clinical use cases, such as risk stratification, monitoring intensity, or therapy surveillance, with calibrated uncertainty treated as a primary output rather than an auxiliary metric [11,15].
Medium-term priorities include the prospective validation of hybrid digital twins under real-world conditions. Beyond predictive accuracy, studies should assess robustness to distribution shifts, missing data, device heterogeneity, and treatment-dependent observation processes. Equally important is the evaluation of whether digital twin outputs meaningfully alter clinician behavior or improve intermediate clinical outcomes [23,24].
Long-term progress will require intervention-aware digital twins capable of supporting causal reasoning, counterfactual analysis, and adaptive decision support. Achieving this vision will depend not only on advances in AI and mechanistic modeling but also on governance frameworks capable of overseeing continuously updating systems whose outputs influence patient care [15,35].

4. Diabetes Digital Twins: Control, Adaptivity, and Human–AI Co-Evolution

Diabetes mellitus occupies a distinctive position in the landscape of healthcare digital twins. Unlike many chronic diseases, diabetes combines relatively rich longitudinal observations with frequent opportunities for intervention. Continuous glucose monitoring, medication dosing records, wearable sensors, and self-reported behavioral data provide repeated measurements of system state, while therapeutic actions can be adjusted on timescales ranging from hours to weeks. Consequently, diabetes is often regarded as one of the most promising settings for realizing adaptive, patient-specific digital twins in clinical practice [36,37,38].
Yet, diabetes is important for reasons that extend beyond data availability. It is one of the few healthcare domains in which digital twins are explicitly intended to influence ongoing decision-making. The central challenge is therefore not merely state estimation but control under uncertainty. Recommendations alter patient behavior, patient behavior alters future observations, and future observations reshape subsequent recommendations. Diabetes digital twins thus provide a valuable stress test for adaptive healthcare digital twins, in which inference, intervention, and human behavior become tightly coupled.

4.1. What Is a Diabetes Digital Twin a Twin Of?

In diabetes, the objective of the digital twin is typically to serve as a patient-specific metabolic control system rather than a static disease representation. The twin seeks to maintain an evolving representation of the latent metabolic state that is sufficiently accurate to support treatment decisions.
Formally, the latent state x t represents a composite of metabolic and behavioral variables that are only partially observable but directly relevant to treatment decisions. These include insulin sensitivity, hepatic glucose production, peripheral glucose uptake, beta-cell function, medication adherence, dietary behavior, and physical activity [37,39]. Observations y t include glucose trajectories, medication records, dietary logs, activity measures, and laboratory values, while actions u t include medication adjustments, nutritional recommendations, behavioral interventions, and changes in monitoring intensity.
This structure makes diabetes digital twins unusually amenable to frequent belief updating and policy adjustment. However, it also means that the system being modeled is explicitly adaptive: patient behavior responds to recommendations, monitoring influences adherence, and optimization targets alter future data generation. As emphasized in recent reviews, this reflexivity distinguishes diabetes digital twins from more inference-centric clinical twins [38,40].

4.2. Control, Feedback, and Policy Learning

Unlike many healthcare digital twins that are primarily inferential, diabetes twins are fundamentally intervention-oriented. Their purpose is not simply to estimate latent states but to support decisions that modify future system trajectories. A defining feature of diabetes digital twins is their orientation toward control rather than prediction. Many proposed systems aim to support or automate decisions about medication dosing, nutrition, or lifestyle modification, often framed as learning optimal policies and informing clinical decision-making under uncertainty [36,41].
Figure 4. Conceptual architecture of a patient-specific diabetes digital twin. Multimodal observations from a patient—including glucose monitoring, medication records, wearable sensor data, and self-reports—are integrated within a hybrid mechanistic and AI-based modeling framework to infer latent metabolic state. The resulting digital twin is continuously updated as new data arrive and provides uncertainty-aware recommendations for medication dosing, nutrition, physical activity, and monitoring. Because recommendations influence patient behavior and future observations, diabetes digital twins operate as closed-loop systems in which inference, intervention, and behavioral adaptation are tightly coupled. This feedback structure introduces challenges of safety, stability, uncertainty quantification, and governance that extend beyond those encountered in purely predictive models.
Figure 4. Conceptual architecture of a patient-specific diabetes digital twin. Multimodal observations from a patient—including glucose monitoring, medication records, wearable sensor data, and self-reports—are integrated within a hybrid mechanistic and AI-based modeling framework to infer latent metabolic state. The resulting digital twin is continuously updated as new data arrive and provides uncertainty-aware recommendations for medication dosing, nutrition, physical activity, and monitoring. Because recommendations influence patient behavior and future observations, diabetes digital twins operate as closed-loop systems in which inference, intervention, and behavioral adaptation are tightly coupled. This feedback structure introduces challenges of safety, stability, uncertainty quantification, and governance that extend beyond those encountered in purely predictive models.
Preprints 223586 g004
From a formal perspective, diabetes digital twins come closer than most healthcare applications to the classical control-theoretic ideal: frequent observations, short feedback loops, and actionable interventions. This has motivated work on adaptive and reinforcement–learning–inspired approaches, as well as simulation-based evaluation of alternative therapeutic strategies [41,42]. Yet this apparent tractability masks a critical limitation. Unlike engineered systems, the diabetes control loop includes a human agent whose preferences, habits, and responses evolve over time. Consequently, the policy π ( · ) governing actions does not act on a passive system but co-evolves with patient and provider behavior. Reviews of diabetes digital twins repeatedly note that the failure to model this co-adaptation undermines both safety and personalization claims [37,38].

4.3. Behavioral Confounding and Instability

Diabetes digital twins vividly illustrate a challenge that is often less visible in slower-moving disease domains: behavioral confounding. Monitoring intensity, recommendation frequency, interface design, and perceived accountability all influence patient behavior, which in turn shapes future observations and inferred states. As a result, changes in observed outcomes may reflect physiological adaptation, behavioral adaptation, or both.
Empirical studies of digital twin–enabled diabetes programs report improvements in glycemic control and metabolic outcomes [43,44,45]. However, many of these studies are retrospective or lack clear counterfactual baselines, making it difficult to disentangle physiological effects from behavioral and care-delivery effects, a concern echoed in systematic reviews of diabetes digital twins [38,40].
This instability has important implications. A digital twin that optimizes short-term glucose control without accounting for behavioral fatigue, treatment burden, or long-term adherence may degrade outcomes over time. Diabetes thus exemplifies a broader lesson for adaptive digital twins: fast feedback amplifies both learning and error.

4.4. Hybrid Modeling and Metabolic Structure

As in cardiovascular medicine, hybrid modeling that combines mechanistic models with AI-based components plays a central role in diabetes digital twins. Mechanistic glucose–insulin models provide structural constraints, while AI-based components are used to personalize parameters, integrate heterogeneous data, or approximate unmodeled dynamics [37,39,46].
Hybridization has enabled practical advances, including simulation-based evaluation of alternative therapies and data augmentation in low-data regimes [41,42]. However, hybrid models also introduce ambiguity about the sources of error and how uncertainty propagates to decisions—particularly when components are continuously updated through learning and adaptation. Reviews emphasize that surrogate accuracy alone is insufficient if validity domains are poorly characterized [24,38].

4.5. Evidence, Validation, and Decision Alignment

Among healthcare digital twins, diabetes systems have progressed furthest toward real-world deployment. Nonetheless, the evidence base remains uneven and limited. Many studies demonstrate improvements in intermediate outcomes such as hemoglobin A1c or glycemic variability, but fewer establish long-term safety, robustness under distribution shift, or sustained benefit across diverse populations [24,38].
More fundamentally, commonly used evaluation metrics are often poorly aligned with clinical objectives. High predictive accuracy or close simulation fit does not guarantee that resulting control policies are safe or beneficial. A glucose prediction model may achieve low average error while systematically underestimating impending hypoglycemia, leading to clinically harmful dosing decisions. The disconnect between predictive performance and decision quality is a recurring concern in healthcare digital twins but is especially acute in diabetes because interventions are frequent, feedback loops are short, and errors can rapidly propagate through the control system [23,24].

4.6. Domain-Specific Research Gaps

A synthesis of the diabetes digital twin literature highlights several gaps that are structurally distinct from those in other domains.
First, human–AI co-adaptation is rarely modeled explicitly, despite being central to system dynamics [37,38]. Second, uncertainty is often underrepresented in control decisions, increasing the risk of brittle or unsafe policies. Third, causal reasoning about interventions remains limited; many systems learn from observed trajectories without the counterfactual grounding needed to support reliable intervention planning [24,36]. Finally, governance and accountability frameworks for adaptive diabetes digital twins remain underdeveloped despite their increasing real-world impact [24,35].
The diabetes case illustrates a central lesson for healthcare digital twins. Rich data and frequent observations do not eliminate uncertainty; they merely shift the challenge from state estimation to control. Once digital twins begin influencing behavior, the system being modeled becomes partially endogenous to the model itself. This reflexivity creates opportunities for personalization and adaptation, but it also introduces new risks related to instability, safety, and governance.

4.7. A Roadmap for Diabetes Digital Twins

Near-term priorities should focus on uncertainty-aware decision support rather than autonomous optimization. Digital twins should provide transparent recommendations for medication adjustment, nutrition, and monitoring while explicitly communicating uncertainty and preserving clinician and patient oversight [24,38].
Medium-term research should emphasize human–AI co-adaptation. Prospective studies should evaluate not only physiological outcomes but also behavioral responses, adherence patterns, trust, burden, and long-term engagement. Explicit modeling of feedback between recommendations and behavior should become a central design requirement rather than an afterthought [37,40].
Long-term progress will require diabetes digital twins to be treated as governed adaptive control systems. Achieving this vision will depend on advances in causal inference, safety-constrained policy learning, continuous validation, and post-deployment monitoring, together with regulatory and governance frameworks capable of overseeing continuously learning systems that directly influence patient care [23,35].

5. Alzheimer’s Disease Digital Twins

Alzheimer’s disease and related dementias (ADRD) are frequently cited as high-impact targets for healthcare digital twins. The motivation is compelling: Alzheimer’s disease unfolds over decades, exhibits substantial heterogeneity, and imposes enormous personal and societal burdens. Digital twins are therefore proposed as tools for early detection, individualized progression modeling, trial enrichment, and personalized care planning [22,32].
Yet Alzheimer’s disease exposes some of the most fundamental epistemic limits of healthcare digital twins. Unlike diabetes, where dense observations and frequent interventions support adaptive control, or cardiovascular disease, where mechanistic structure partially constrains inference, Alzheimer’s disease is characterized by weak observability, long time horizons, and limited intervention leverage. The central challenge is therefore not prediction or control but identifiability: determining what can legitimately be inferred about the latent disease state from incomplete and indirect observations.
Consequently, Alzheimer’s disease serves as a critical test case for distinguishing predictive ambition from inferential credibility. The domain highlights how uncertainty, non-identifiability, and limited actionability constrain what healthcare digital twins can reasonably claim to know and support [11,15].
Figure 5. Conceptual architecture of an Alzheimer’s disease digital twin. Multimodal observations from a patient—including neuroimaging, cognitive assessments, fluid biomarkers, and longitudinal clinical monitoring—are integrated within mechanistic and AI-based models to maintain a probabilistic belief over latent disease state. The digital twin is continuously updated as new observations become available and supports uncertainty-aware estimation of disease trajectories, prognosis, monitoring strategies, and care planning. Additional measurements may be acquired to reduce uncertainty and refine state estimates. In contrast to control-oriented digital twins, Alzheimer’s disease digital twins primarily serve as epistemic instruments for state estimation and prognosis under partial observability.
Figure 5. Conceptual architecture of an Alzheimer’s disease digital twin. Multimodal observations from a patient—including neuroimaging, cognitive assessments, fluid biomarkers, and longitudinal clinical monitoring—are integrated within mechanistic and AI-based models to maintain a probabilistic belief over latent disease state. The digital twin is continuously updated as new observations become available and supports uncertainty-aware estimation of disease trajectories, prognosis, monitoring strategies, and care planning. Additional measurements may be acquired to reduce uncertainty and refine state estimates. In contrast to control-oriented digital twins, Alzheimer’s disease digital twins primarily serve as epistemic instruments for state estimation and prognosis under partial observability.
Preprints 223586 g005

5.1. What Is an Alzheimer’s Digital Twin a Twin Of?

At first glance, the object of an Alzheimer’s digital twin appears straightforward: an individual patient’s disease state and its future evolution. In practice, however, defining the latent state x t is already contentious.
Formally, the latent state x t is best understood as a composite of partially coupled processes, including neuropathological burden, neurodegeneration, cognitive reserve, functional capacity, and behavioral symptoms [22,31]. Observations y t —including cognitive assessments, neuroimaging, fluid biomarkers, digital biomarkers, and functional evaluations—provide only indirect and noisy views of these processes. Actions u t include monitoring decisions, symptomatic therapies, disease-modifying treatments, and non-pharmacologic interventions.
This structure fundamentally distinguishes Alzheimer’s disease from more favorable digital twin domains. The latent disease state is only weakly constrained by available observations; interventions have limited and heterogeneous effects, and clinically meaningful outcomes often emerge over long time horizons. These characteristics impose intrinsic limits on fidelity, identifiability, and actionability [15,33].

5.2. Conceptual and Translational Motivations

Digital twin approaches are increasingly proposed as mechanisms for integrating imaging, biomarkers, cognitive assessments, and clinical history into individualized disease models that support progression forecasting, trial enrichment, and treatment monitoring [11,22,32].
At the same time, Alzheimer’s disease reflects a broader pattern seen throughout healthcare AI: strong predictive performance on curated datasets often fails to translate into robust clinical impact. Digital twin framing can amplify this tension by implying individualized understanding and control that exceed what current data, models, and interventions can support [24,33].

5.3. Identifiability as the Central Scientific Constraint

The defining methodological challenge for Alzheimer’s digital twins is identifiability. In formal terms, multiple latent trajectories x 1 : t can generate observational histories y 1 : t that are statistically indistinguishable. This problem persists even as more data are collected because the observation process itself is indirect and noisy, and distinct pathological pathways can lead to similar clinical manifestations.
Recent work explicitly analyzing identifiability in Alzheimer’s disease modeling demonstrates that many commonly used progression models are fundamentally under-determined by available data [31]. From a digital twin perspective, this implies that belief distributions over x t may remain broad, multimodal, and weakly informative even under idealized updating schemes. Crucially, this is not merely a technical nuisance but a structural limitation: no amount of algorithmic sophistication can resolve non-identifiability without introducing additional assumptions, constraints, or interventional signals [11,15].
The implications are profound. A digital twin that reports a precise disease stage, progression trajectory, or treatment forecast without explicitly representing irreducible uncertainty risks conveying a level of knowledge that the underlying evidence cannot support. In Alzheimer’s disease, uncertainty is not merely a consequence of limited data; it is often a consequence of the structure of the problem itself [18,19].

5.4. Dynamics, Updating, and Long Time Horizons

Alzheimer’s disease unfolds over years to decades, far exceeding the time scales typically addressed by clinical AI systems. Long horizons exacerbate every challenge faced by digital twins: sparse and irregular observations, evolving diagnostic criteria, cohort effects, and changes in standards of care. Updating mechanisms that function reasonably well over short horizons in cardiovascular or operational settings become increasingly fragile when extrapolated to long-term disease trajectories [15,24].
Long horizons also magnify treatment-dependent confounding. Patients perceived to be at higher risk often receive more intensive monitoring, specialist evaluation, or experimental therapies, altering both observations and outcomes. As a result, digital twins may learn patterns of care as much as patterns of disease progression. These biases become increasingly difficult to detect as prediction horizons lengthen [15,23].

5.5. Decision Interfaces and the Problem of Actionability

Perhaps the most under-examined question in Alzheimer’s digital twin research is: what decisions is the twin meant to support? Unlike cardiovascular disease, where acute interventions and near-term risk stratification are common, many Alzheimer’s-related decisions involve monitoring, care planning, trial participation, and anticipatory support rather than immediate therapeutic optimization.
Ethical analyses emphasize that digital twins in this domain risk conflating prediction with prescription, particularly when uncertainty is high and interventions are limited [18]. Regulatory and ethical perspectives further caution that the clinical utility of personalized forecasts must be weighed against psychological harm, resource misallocation, and inequitable access to care [34,35,47].
These considerations suggest that Alzheimer’s digital twins should adopt a deliberately conservative notion of actionability. Their primary role may be to support monitoring, anticipatory planning, trial enrollment decisions, and uncertainty-aware conversations among patients, caregivers, and clinicians, rather than to generate deterministic treatment recommendations.

5.6. Evidence and Validation Challenges

Validation of Alzheimer’s digital twins faces unique obstacles. Prospective interventional studies are slow and costly, while retrospective evaluations are vulnerable to temporal leakage, cohort effects, and endpoint instability. Reviews of healthcare digital twins repeatedly identify Alzheimer’s disease as an exemplar of the gap between predictive performance and clinical utility [14,15,24].
As a result, claims about individualized progression forecasting or treatment optimization often rest on fragile evidence. Without validation strategies that explicitly account for uncertainty, identifiability, and decision context, such claims remain aspirational rather than decision-grade.

5.7. Domain-Specific Research Gaps

A critical synthesis of the Alzheimer’s digital twin literature reveals several gaps that qualitatively differ from those in more mechanistically constrained conditions.
First, identifiability is rarely treated as a design constraint. Most models implicitly assume that latent disease state can be inferred with sufficient precision despite strong evidence to the contrary [11,31]. Second, uncertainty is underreported and poorly integrated into decision-making. Many systems provide point estimates without credible uncertainty bounds or guidance regarding how uncertainty should influence clinical choices [15,24]. Third, causal reasoning remains largely absent. Most twins forecast observed trajectories but are poorly equipped to evaluate counterfactual interventions, particularly when treatments have modest and heterogeneous effects [23]. Finally, ethical and governance considerations are often treated as peripheral rather than intrinsic despite the high stakes and limited actionability characteristic of Alzheimer’s disease [18,19,35].
The Alzheimer’s case illustrates a complementary lesson to the diabetes case. In diabetes, rich data shift the challenge from inference to control. In Alzheimer’s disease, sparse and indirect observations shift the challenge from control back to inference itself. The central limitation is not insufficient model complexity but persistent uncertainty about latent disease state and future trajectory.

5.8. A Roadmap for Alzheimer’s Digital Twins

Near-term priorities should focus on uncertainty-aware disease modeling. Digital twins should explicitly represent non-identifiability, calibration, and failure modes rather than emphasizing single-point forecasts. Validation efforts should prioritize robustness, transparency, and alignment with realistic clinical decisions [15,24].
Medium-term efforts should focus on decision-bounded applications such as trial enrichment, monitoring strategies, and shared decision-making. Success should be evaluated in terms of decision quality and clinical utility rather than predictive performance alone [32,47].
Long-term progress will require advances in causal inference, richer longitudinal and interventional data, and governance frameworks capable of overseeing adaptive systems operating under persistent uncertainty. The goal is not complete disease reconstruction but trustworthy support for decisions made in the face of incomplete knowledge [23,35].

6. Maternal and Maternal–Fetal Health Digital Twins

Maternal and maternal–fetal health represent uniquely stringent test cases for healthcare digital twins. Pregnancy is characterized by rapid physiological change, time-sensitive clinical decisions, and asymmetric risk, where relatively rare adverse events can have catastrophic consequences for the mother, fetus, or both. Outcomes are further shaped by structural inequities in access to care, social determinants of health, the growing prevalence of maternity care deserts, and increasing rates of preexisting conditions such as obesity, hypertension, and diabetes. Consequently, digital twins in this domain confront technical, ethical, and governance challenges that are inseparable from scientific credibility [18,19,47].
Where Alzheimer’s disease highlights the limits of inference under deep uncertainty, maternal–fetal health introduces a different challenge: decision-making under uncertainty in a coupled biological system. Maternal and fetal states evolve together but not symmetrically, and interventions that reduce risk for one may alter risk for the other. Digital twins in this setting therefore raise fundamental questions about safety, equity, responsibility, and acceptable risk.
Recent discussions of maternal digital twins emphasize their potential for early risk detection, dynamic monitoring, and individualized escalation of care [28]. At the same time, these ambitions expose limitations of generic digital twin frameworks that do not explicitly account for coupled physiology, safety-critical decision-making, and inequitable observation processes [15,24].
Figure 6. Conceptual illustration of a maternal–fetal digital twin. Multimodal observations from pregnancy—including maternal clinical measurements, fetal monitoring, placental biomarkers, and contextual information—are integrated within mechanistic and AI-based models to maintain a probabilistic belief over coupled maternal, placental, and fetal states. The digital twin is continuously updated throughout pregnancy and supports uncertainty-aware estimation of disease trajectories, adverse outcome risks, monitoring strategies, and care planning. Additional observations may be acquired to reduce uncertainty and refine state estimates. Maternal–fetal digital twins therefore function as probabilistic models of a partially observed, dynamically coupled biological system, supporting individualized monitoring and clinical decision-making under uncertainty.
Figure 6. Conceptual illustration of a maternal–fetal digital twin. Multimodal observations from pregnancy—including maternal clinical measurements, fetal monitoring, placental biomarkers, and contextual information—are integrated within mechanistic and AI-based models to maintain a probabilistic belief over coupled maternal, placental, and fetal states. The digital twin is continuously updated throughout pregnancy and supports uncertainty-aware estimation of disease trajectories, adverse outcome risks, monitoring strategies, and care planning. Additional observations may be acquired to reduce uncertainty and refine state estimates. Maternal–fetal digital twins therefore function as probabilistic models of a partially observed, dynamically coupled biological system, supporting individualized monitoring and clinical decision-making under uncertainty.
Preprints 223586 g006

6.1. What Is the Maternal Digital Twin a Twin Of?

Unlike many disease-specific digital twins, maternal digital twins must represent a dyadic system. The latent state x t is naturally decomposed into maternal and fetal components, x t = ( x t ( m ) , x t ( f ) ) , whose dynamics are interdependent but not symmetric. Maternal physiology adapts to support fetal development, while fetal well-being depends on placental function, maternal health, and environmental exposures [11,28].
Clinical observations y t include maternal characteristics, vital signs, laboratory values, imaging findings, fetal monitoring, behavioral factors, and social or environmental determinants. These observations are irregularly sampled and strongly influenced by gestational age, access to care, local practice patterns, and clinical judgment. Actions u t include monitoring strategies, pharmacologic interventions, timing of delivery, and escalation to higher levels of care.
This structure fundamentally distinguishes maternal–fetal digital twins from most patient-level twins. The challenge is not simply uncertainty about a single latent state, but uncertainty about a coupled system in which risks, benefits, and responsibilities are distributed unevenly across multiple stakeholders. Recent conceptual analyses emphasize that such coupled decision contexts fundamentally alter how uncertainty and responsibility must be handled in digital twins [11,18].

6.2. Clinical Motivation and the Appeal of Dynamic Monitoring

The primary motivation for maternal digital twins lies in the early detection, monitoring, and management of high-risk conditions such as excessive gestational weight gain, hypertension, preeclampsia, gestational diabetes, fetal growth restriction, preterm birth, and macrosomia. These conditions often emerge gradually but can escalate rapidly, making static risk scores insufficient. Maternal–fetal digital twins, therefore, emphasize continuous updating and individualized trajectories rather than one-time predictions [15,28].
This motivation aligns naturally with the digital twin paradigm. However, it also raises a central question: how much information is required to support safe decision-making when observations are sparse, unevenly distributed, and systematically influenced by healthcare access? Unlike engineered systems, maternal–fetal digital twins cannot assume continuous sensing or unbiased observation processes [23,24].

6.3. Partial Observability, Measurement Inequity, and Bias

Maternal health illustrates particularly clearly that partial observability is not merely a technical nuisance but a structural feature of healthcare systems. Observation frequency and quality are strongly associated with socioeconomic status, geographic access, institutional practices, and implicit bias. Consequently, the observation process itself is often inequitable.
Within a digital twin framework, this has two important implications. First, updating mechanisms that assume uniform data availability may systematically underestimate risk in under-monitored populations. Second, personalization strategies may inadvertently amplify disparities by producing more reliable predictions for populations that are already better served by healthcare systems. Reviews of digital twins and AI-enabled healthcare systems repeatedly identify this feedback loop as a major source of inequitable deployment [15,20,24].
These dynamics suggest that observation processes must be treated as part of the modeled system rather than as external nuisances. Failure to account for differential monitoring risks embedding existing inequities directly into the digital twin itself.

6.4. Safety, Asymmetric Risk, and Decision Thresholds

Maternal–fetal digital twins operate in a regime of asymmetric and high-stakes risk. False reassurance can lead to catastrophic maternal or fetal outcomes, whereas false alarms may increase anxiety, resource utilization, and unnecessary interventions. Consequently, conventional performance metrics such as accuracy or area under the receiver operating characteristic curve (AUC) are poorly aligned with clinical decision-making in this domain [23,24].
From a formal perspective, the decision policy π ( · ) governing actions must be explicitly risk-sensitive, incorporating asymmetric losses and appropriate safety margins. Yet most proposed maternal–fetal digital twins stop short of specifying such policies, instead reporting predictive performance without clarifying how outputs should influence clinical decisions. This pattern mirrors broader critiques of healthcare digital twins that emphasize the gap between prediction and actionability [11,15].
More fundamentally, predictions alone do not determine action. Safe deployment requires explicit articulation of decision thresholds, acceptable risk tradeoffs, and escalation policies. Without these elements, predictive outputs cannot be translated responsibly into clinical care.

6.5. Hybrid Modeling and Coupled Dynamics

As in cardiovascular medicine, hybrid modeling approaches are frequently proposed for maternal–fetal digital twins. Mechanistic understanding of maternal physiology, placental function, fetal development, and maternal adaptation provides structural constraints, while AI-based components integrate heterogeneous data streams and identify complex temporal patterns [28,29].
However, coupling two evolving physiological systems magnifies the challenges of hybridization. Errors or biases in one component may propagate through the coupled system, and surrogate approximations may fail precisely in regimes where safety margins are smallest. Unlike many chronic disease settings, maternal–fetal health offers little tolerance for such failures, reinforcing concerns raised in broader critiques of hybrid digital twins [15,23].

6.6. Evidence, Validation, and Ethical Constraints

Validation of maternal–fetal digital twins is constrained by both methodological and ethical considerations. Retrospective analyses are vulnerable to confounding by care intensity, while prospective evaluations must prioritize safety, equity, and informed consent. Consequently, evidence standards for maternal–fetal digital twins extend beyond predictive performance to include demonstration of clinical benefit and avoidance of harm [14,24].
Ethical analyses emphasize that digital twins in pregnancy cannot be evaluated solely on predictive accuracy; they must also be assessed according to whether they reduce harm, improve equity, and support informed decision-making without undermining patient autonomy [18,19,47]. Regulatory perspectives similarly stress that adaptive systems influencing perinatal care will require robust lifecycle oversight [35].

6.7. Domain-Specific Research Gaps

Several research gaps emerge as particularly salient in maternal-fetal health.
First, principled frameworks for modeling coupled maternal and fetal states under uncertainty remain limited, especially when observations are uneven, sparse, or systematically biased [11]. Second, fairness is rarely treated as a first-class design requirement. Subgroup calibration, measurement equity, and asymmetric error distributions must be explicitly evaluated rather than assumed [15,24]. Third, decision policies remain underspecified. Without explicit representations of risk tolerance, escalation criteria, and safety constraints, digital twins cannot be responsibly deployed in safety-critical maternal healthcare environments. Finally, prospective validation remains limited, and post-deployment monitoring frameworks are largely absent [14,35].
The maternal-fetal case highlights a distinct limitation of healthcare digital twins. In Alzheimer’s disease, uncertainty constrains what can be inferred. In maternal–fetal health, uncertainty constrains what can safely be done. The central challenge is not merely accurate prediction but responsible decision-making in the presence of asymmetric risk, coupled physiology, and unequal access to observation.

6.8. A Roadmap for Maternal-Fetal Digital Twins

Near-term priorities should focus on uncertainty-aware monitoring and early warning systems, with explicit reporting of subgroup calibration, measurement limitations, and decision thresholds. Digital twins should augment clinical judgment rather than automate high-stakes decisions [15,24].
Medium-term efforts should evaluate whether digital twin–informed monitoring and escalation strategies improve maternal and fetal outcomes without increasing unnecessary interventions or widening disparities. Equity should be treated as a primary evaluation criterion rather than a secondary consideration [14,47].
Long-term progress will require integrated decision-support systems with formal safety cases, continuous validation, post-deployment surveillance, and governance frameworks capable of overseeing adaptive AI systems operating in safety-critical maternal care environments [23,35].

7. Addiction as a Boundary Case for Digital Twins

Maternal–fetal digital twins illustrate how safety, equity, and asymmetric risk constrain the responsible use of prediction, even when the underlying state remains fundamentally physiological. Addiction pushes the digital twin paradigm to its conceptual limits. Unlike cardiovascular disease, diabetes, or pregnancy, addiction cannot be cleanly localized to a physiological subsystem or understood solely through mechanistic or control-oriented frameworks. Rather, it emerges from interactions among neurobiology, cognition, affect, behavior, and social environment, unfolding through processes that are highly context-dependent, adaptive, and shaped by individual agency [19,20].
For this reason, addiction serves as a revealing boundary case for healthcare digital twins. It exposes assumptions about state, observability, actionability, and alignment that often remain implicit in more favorable domains. The resulting limitations are not merely technical; they raise fundamental questions about what can legitimately be represented, inferred, and acted upon using computational models of human behavior [11,18].
Figure 7. Conceptual architecture of an addiction digital twin. Partial and mediated observations—including self-reports, digital traces, biological measurements, and social-contextual information—are integrated within mechanistic and AI-based models to maintain a probabilistic representation of latent addiction-related state. Because neurobiological, psychological, behavioral, and social processes are only partially observable and are altered by measurement and intervention, addiction digital twins operate under conditions of reflexivity, deep uncertainty, and limited predictability. Rather than supporting deterministic prediction or optimization, addiction digital twins are more appropriately viewed as constrained epistemic instruments for understanding, monitoring, support planning, and scenario exploration. These limitations motivate governance-first approaches emphasizing consent, fairness, transparency, and accountability.
Figure 7. Conceptual architecture of an addiction digital twin. Partial and mediated observations—including self-reports, digital traces, biological measurements, and social-contextual information—are integrated within mechanistic and AI-based models to maintain a probabilistic representation of latent addiction-related state. Because neurobiological, psychological, behavioral, and social processes are only partially observable and are altered by measurement and intervention, addiction digital twins operate under conditions of reflexivity, deep uncertainty, and limited predictability. Rather than supporting deterministic prediction or optimization, addiction digital twins are more appropriately viewed as constrained epistemic instruments for understanding, monitoring, support planning, and scenario exploration. These limitations motivate governance-first approaches emphasizing consent, fairness, transparency, and accountability.
Preprints 223586 g007

7.1. What Is the Addiction Digital Twin a Twin Of?

At the core of any digital twin lies a latent state x t that is presumed to represent the system of interest. In the case of addiction, defining such a state is already problematic. Relevant dimensions include physiological dependence, craving intensity, impulse control, affective state, stress, social and cue exposures, access to substances, and learned expectations. Many of these variables are subjective, context-dependent, and difficult to measure directly, challenging the assumption that a coherent, inferable latent state exists [11,19].
Clinical observations y t — including self-reports, toxicology results, digital traces, and clinical encounters — are partial, noisy, and strategically mediated. Individuals with substance use disorders may under-report use, avoid monitoring, or alter behavior in response to surveillance. Actions u t include pharmacologic and behavioral treatments, case management, monitoring strategies, and, in some settings, legally mandated interventions.
Unlike most domains considered in this review, addiction challenges the assumption that a stable latent state exists. Key variables are subjective, reflexive, and socially embedded. This raises a foundational question: when the state itself depends on interpretation, context, and self-understanding, what does it mean to claim that a computational system is a “twin’’ of a person? Similar concerns arise in sociotechnical analyses of digital health surveillance and algorithmic risk assessment [18,20].

7.2. Observation as Intervention

A defining feature of addiction is that observation itself can be interventional. Monitoring substance use, tracking behavior, or assessing relapse risk can alter the very dynamics that the digital twin seeks to model. Surveillance may induce avoidance, stigma, or disengagement, while self-monitoring may either support recovery or exacerbate anxiety and shame [19].
In formal terms, the observation process g ( x t , u t , ϵ t ) cannot be treated as passive. Measurement choices shape behavior, modify incentives, and alter social relationships. This violates a tacit assumption underlying many digital twin frameworks: that observations primarily reveal information about the state rather than act upon it. Reviews of healthcare digital twins increasingly recognize this feedback as a critical but under-modeled failure mode [15,24].
Addiction therefore exposes a deep limitation of conventional digital twin formulations. When observation changes the system behaviorally, socially, and sometimes legally, updating mechanisms that ignore these effects risk generating misleading or harmful inferences.

7.3. Dynamics, Nonstationarity, and Context Dependence

Addiction dynamics are highly nonstationary. Relapse risk is shaped by fluctuating stressors, social environments, access constraints, and episodic life events. Periods of apparent stability may be punctuated by abrupt transitions, while the same individual may respond differently to identical interventions at different times.
These properties challenge the assumption that a stable transition process exists to be learned. The system being modeled is reflexive: beliefs about risk, monitoring, and intervention become part of the state itself [11,19]. Consequently, long-horizon forecasting and control-oriented digital twins are particularly fragile in this domain. Apparent predictive success over short windows may reflect transient regularities rather than stable mechanisms, a concern echoed in broader critiques of AI-enabled healthcare prediction under nonstationarity [23,24].

7.4. Actionability, Agency, and Alignment

Perhaps the most profound challenge for addiction digital twins lies in the question of actionability. In many healthcare domains, the goals of the digital twin are assumed to align with the goals of the patient and clinician. In addiction, this assumption is often violated. Individuals may resist or contest interventions, and their preferences may be unstable or internally conflicted.
From a decision-theoretic perspective, the policy π ( · ) governing actions cannot assume a cooperative agent with stable objectives. Instead, agency is endogenous to the disorder itself. Attempts to “optimize” behavior risk crossing ethical boundaries, particularly when interventions are coercive or tied to legal or social consequences [18,47].
This makes addiction a domain where digital twins cannot be evaluated solely on technical grounds. Questions of consent, autonomy, and acceptable influence are inseparable from claims about prediction or control [19,35].

7.5. Governance, Consent, and Secondary Use

The governance challenges surrounding addiction digital twins extend beyond healthcare. False-positive risk assessments may contribute to stigma, loss of employment, or legal consequences, whereas false negatives may result in overdose or other serious harms [18,47].
More broadly, addiction-related digital twins intersect with systems of surveillance and social control. Risk assessments developed for clinical purposes may be repurposed for punitive or exclusionary uses, raising concerns about mission creep and secondary use [19,20].
These realities imply that transparency, contestability, limits on secondary use, and robust consent mechanisms are not optional safeguards but prerequisites for ethical legitimacy [35].

7.6. What Addiction Teaches Us About Digital Twins

The addiction domain clarifies several general lessons for healthcare digital twins. First, not all clinically relevant phenomena admit well-defined, inferable latent states. In some domains, ambiguity is not a temporary limitation of measurement but a property of the phenomenon itself. Treating subjectivity and social context as noise to be engineered away is, therefore, a category error [11]. Second, observation is not neutral. Digital twins that rely on intensive monitoring must account for the behavioral and ethical consequences of measurement itself [19]. Third, actionability is normative. Decisions about what a digital twin should recommend cannot be separated from values concerning autonomy, coercion, and harm [18,47]. Finally, limits on applicability are scientifically meaningful. Recognizing where digital twins should not be used—or should be used only in constrained, reflective ways—is a mark of maturity rather than failure [11,15].

7.7. A Constrained Roadmap for Addiction Digital Twins

A responsible research roadmap for addiction digital twins must be intentionally conservative.
In the near term, digital twin concepts may be most appropriate as reflective or exploratory tools for clinicians and patients, supporting self-awareness and dialog rather than optimization or, especially, enforcement [19].
In the medium term, research should focus on understanding how digital monitoring and modeling affect behavior, trust, and outcomes, with rigorous safeguards against coercive use [20,47].
In the longer term, progress will depend less on algorithmic sophistication than on governance frameworks that define acceptable uses, ensure informed consent, protect autonomy, and prevent harm [23,35].
Addiction illustrates a boundary beyond which the digital twin metaphor becomes increasingly strained. The challenge is no longer uncertainty about state estimation or intervention effects, but uncertainty about the meaning and stability of the state itself. Radiology presents a contrasting challenge: a domain rich in observations but often poor in state semantics.

8. Radiology Digital Twins: Observation, State Ambiguity, and Workflow Integration

Addiction exposes a boundary where the digital twin paradigm becomes unstable because the latent state is subjective, reflexive, and socially constructed. Radiology presents the opposite challenge. Here, the problem is not insufficient observation but an abundance of observation. Modern imaging generates extraordinarily rich data, creating a persistent temptation to equate observation with understanding.
Radiology, therefore, occupies a distinctive position in the healthcare digital twin landscape. It is frequently portrayed as a natural home for digital twins because imaging plays a central role in diagnosis, monitoring, and treatment planning. Yet radiology also exemplifies a recurring problem in digital twin discourse: the conflation of high-dimensional observations with a well-defined latent state [21,30].
Radiology digital twins thus serve as a test of whether the digital twin paradigm can move beyond pattern recognition toward stateful, clinically meaningful, and decision-relevant systems.
Figure 8. Conceptual architecture of a radiology digital twin. Imaging observations, clinical information, and workflow data are integrated within mechanistic and AI-based models to maintain an evolving representation of latent patient state. The digital twin supports trajectory estimation, risk assessment, uncertainty-aware interpretation, and clinical decision support. Radiologist review and downstream care pathways remain integral components of the system, linking model outputs to clinical action. Because images constitute observations rather than state, meaningful inference requires explicit consideration of acquisition context, longitudinal change, workflow integration, and uncertainty.
Figure 8. Conceptual architecture of a radiology digital twin. Imaging observations, clinical information, and workflow data are integrated within mechanistic and AI-based models to maintain an evolving representation of latent patient state. The digital twin supports trajectory estimation, risk assessment, uncertainty-aware interpretation, and clinical decision support. Radiologist review and downstream care pathways remain integral components of the system, linking model outputs to clinical action. Because images constitute observations rather than state, meaningful inference requires explicit consideration of acquisition context, longitudinal change, workflow integration, and uncertainty.
Preprints 223586 g008

8.1. What Is the Radiology Digital Twin a Twin Of?

In many proposed radiology digital twins, the answer to this question remains implicit. The “twin” is often described as a digital counterpart of an organ, lesion, or patient, yet the underlying latent state x t is rarely specified with precision. Instead, emphasis is placed on images, image-derived features, and learned representations.
Formally, this corresponds to increasingly sophisticated modeling of the observation process g ( x t , u t , ϵ t ) while leaving the latent state itself under-specified. Longitudinal imaging may reveal change, but without a clearly defined state, updating becomes a narrative metaphor rather than a principled inference process. Conceptual analyses argue that this omission weakens claims of individualized, stateful modeling [11,15].
Unlike cardiovascular or diabetes twins, where latent physiological variables are at least conceptually identifiable, radiology faces a persistent risk of mistaking rich observations for the state they are intended to reveal.

8.2. Observation Dominance and the Limits of Imaging-Centric Twins

Radiology-centered digital twins often treat imaging richness as a substitute for mechanistic understanding. Deep learning models trained on large imaging datasets can achieve impressive performance in lesion detection, classification, and short-term outcome prediction. These successes are frequently cited as evidence that digital twins are already becoming feasible in radiology.
However, high predictive accuracy does not establish the existence of a well-posed digital twin [21,30]. Imaging models may capture correlational structures without representing the biological processes that drive disease progression or treatment response. When embedded within a digital twin narrative, such models can create an illusion of individualized understanding while functioning primarily through population-level pattern matching [15,24].
Without explicit state definitions, imaging-centric twins also struggle to support counterfactual reasoning, intervention planning, or long-term forecasting. Their utility, therefore, remains largely confined to detection, classification, and prioritization tasks.

8.3. Longitudinal Imaging and State Drift

Longitudinal imaging appears to provide a natural foundation for digital twins because repeated scans permit continual updating. In principle, serial imaging should support increasingly refined estimates of disease progression and treatment response.
In practice, longitudinal radiology introduces substantial challenges. Scanner hardware, acquisition protocols, reconstruction algorithms, and contrast usage vary across institutions and over time. These changes often induce distribution shifts that can masquerade as biological change [23,30].
Consequently, changes in observations y t cannot be interpreted straightforwardly as changes in latent state x t . Provenance, acquisition context, and measurement processes must be modeled explicitly rather than treated as nuisance variables. Radiology, therefore, illustrates a broader principle for healthcare digital twins: observation systems are part of the system being modeled [15,24].

8.4. Workflow-Centered Digital Twins

A more grounded interpretation of radiology digital twins shifts attention from patients to workflows. In this framing, the physical system is the radiology enterprise itself: scheduling, protocol selection, image acquisition, reporting, follow-up recommendations, and communication with clinical teams.
This formulation aligns more closely with the engineering origins of digital twins because states, actions, and outcomes are explicit and measurable. Queue lengths, report turnaround times, protocol compliance, and follow-up completion rates can be directly observed and optimized. Reviews increasingly suggest that workflow-centered twins may offer more immediate and verifiable benefits than patient-level imaging twins while avoiding unsupported claims of individualized understanding [11,21,48].

8.5. Decision Support and the Utility Gap

Despite the widespread deployment of AI tools in radiology, evidence that these systems improve patient outcomes remains limited. This gap persists even as algorithms achieve expert-level performance on narrowly defined tasks [24].
From a digital twin perspective, the problem is not predictive performance but decision alignment. Many radiology systems focus on detection or classification without specifying how outputs should alter downstream decisions, follow-up pathways, or treatment plans. Improved predictions do not automatically translate into improved care [15,30].
Radiology, therefore, illustrates a recurring theme across healthcare digital twins: prediction without an explicit decision interface rarely delivers meaningful clinical impact.

8.6. Domain-Specific Research Gaps

Several research gaps emerge when radiology is viewed through the digital twin lens. First, latent state definitions remain underdeveloped. Without explicit links between imaging findings and clinically meaningful state variables, digital twins remain observational artifacts rather than inferential systems [11]. Second, longitudinal validation paradigms remain immature. Few evaluations distinguish true disease progression from technical variability, institutional drift, or measurement artifacts [23,24]. Third, integration with clinical workflows remains limited. Most radiology digital twins are evaluated in isolation rather than as components of end-to-end care pathways [21,48]. Finally, governance challenges persist. Imaging-based digital twins raise important questions regarding accountability, bias, transparency, and portability across institutions [18,35].

8.7. A Roadmap for Radiology Digital Twins

The radiology domain suggests a roadmap centered on state semantics, workflow integration, and decision alignment.
In the near term, research should prioritize longitudinal robustness, explicit modeling of acquisition context, calibration under distribution shift, and uncertainty-aware interpretation [23,24].
In the medium term, workflow-centered digital twins offer the clearest path to measurable impact because they provide explicit state variables, auditable outcomes, and well-defined governance structures [35,48].
In the longer term, patient-level radiology digital twins may become viable as components of multimodal systems that integrate imaging with clinical, molecular, physiological, and behavioral data. Achieving this goal will require explicit state definitions, principled uncertainty representation, and decision-centered evaluation rather than continued reliance on imaging performance alone [11,15].
Radiology highlights a failure mode distinct from those encountered in earlier domains. The challenge is not insufficient data but the absence of explicit state semantics. Rich observations alone do not constitute a digital twin. The transition from images to clinically meaningful, decision-relevant representations requires careful attention to state definition, uncertainty, workflow integration, and governance. This lesson becomes even more apparent in healthcare organizations, where state variables and control levers are often more explicit, but optimization itself becomes inseparable from institutional objectives, resource constraints, and normative choices.

9. Hospital Systems and Healthcare Operations Digital Twins

If radiology exposes the limits of observation-dominant digital twins at the patient level, hospital systems and healthcare operations represent the opposite end of the spectrum. Here, state variables are comparatively explicit, observations are frequent, and interventions are direct and repeatable. As a result, hospital operations are often regarded as the most mature and tractable application domain for digital twins in healthcare.
Unlike patient-centered digital twins, which contend with partial observability, weak interventions, and deep biological uncertainty, hospital operations involve systems with comparatively well-defined states, robust data streams, and explicit control levers. Beds are occupied or available, staff are scheduled or absent, patients are admitted or discharged, and resources are allocated according to policies that can, at least in principle, be formalized [49,50].
For this reason, hospital operations digital twins hew more closely to the original engineering conception of a digital twin than most clinical applications. Yet their apparent tractability introduces a different challenge. When inference becomes easier, the central questions shift from state estimation to objective specification, governance, and accountability. Improvements in efficiency may conflict with safety, workforce well-being, equity, or patient-centered outcomes, a tension increasingly emphasized in critical reviews of healthcare digital twins [11,15].
Figure 9. Conceptual architecture of a hospital operations digital twin. Real-time operational data are integrated into a continuously updated representation of hospital state, including demand, capacity, staffing, patient flow, and resource utilization. Simulation, forecasting, and scenario analysis support evaluation of alternative operational strategies and inform decisions regarding capacity, staffing, resource allocation, and service priorities. Decision makers implement actions that modify system behavior and generate new observations, creating a closed loop of monitoring, prediction, intervention, and learning. Unlike many patient-level digital twins, hospital operations digital twins possess relatively explicit state variables and direct control levers. Their responsible use therefore depends not only on predictive accuracy but also on governance mechanisms that ensure alignment with safety, equity, workforce, cost, and policy objectives.
Figure 9. Conceptual architecture of a hospital operations digital twin. Real-time operational data are integrated into a continuously updated representation of hospital state, including demand, capacity, staffing, patient flow, and resource utilization. Simulation, forecasting, and scenario analysis support evaluation of alternative operational strategies and inform decisions regarding capacity, staffing, resource allocation, and service priorities. Decision makers implement actions that modify system behavior and generate new observations, creating a closed loop of monitoring, prediction, intervention, and learning. Unlike many patient-level digital twins, hospital operations digital twins possess relatively explicit state variables and direct control levers. Their responsible use therefore depends not only on predictive accuracy but also on governance mechanisms that ensure alignment with safety, equity, workforce, cost, and policy objectives.
Preprints 223586 g009

9.1. What Is the Hospital Digital Twin a Twin Of?

In healthcare operations, the object of the digital twin is typically a care delivery system rather than an individual patient. The latent state x t may include bed occupancy, patient location and acuity, staffing levels, equipment availability, queue states for diagnostic or therapeutic services, and the progression of patients through care pathways.
Observations y t are derived from administrative systems, electronic health records, scheduling platforms, and sensor-enabled infrastructure. Compared with clinical domains, these observations are relatively frequent and standardized, although they remain subject to institutional practices and documentation artifacts. Actions u t include staffing adjustments, scheduling policies, routing decisions, discharge planning, and escalation protocols.
This framing aligns closely with operations research and systems engineering, making hospital operations one of the few healthcare domains in which a digital twin possesses a relatively clear physical referent and measurable performance objectives [48,49]. The implication is not that evaluation is easy, but that evaluation is at least definable when objectives and constraints are made explicit.

9.2. Existing Approaches and Their Implications

Several strands of work illustrate both the promise and limitations of operational digital twins. Early efforts focused on patient pathway modeling and simulation, using digital replicas of hospital flows to evaluate scheduling, resource allocation, and throughput under alternative scenarios [49]. More recent work extends these ideas to specific care models, such as hospital-at-home programs, where digital twins are used to coordinate resources and monitor system-level risk [48].
Facility-oriented digital twins further broaden the concept by integrating energy usage, equipment maintenance, and environmental controls into operational models of healthcare infrastructure [50]. Unlike patient-level digital twins, these systems can often be validated against objective outcomes such as wait times, length of stay, resource utilization, and cost [14].
However, operational tractability should not be mistaken for neutrality. Choices about what to optimize embed assumptions about value, priority, and acceptable tradeoffs. A digital twin that optimizes throughput may inadvertently increase clinician workload, compromise safety, or exacerbate disparities if clinical and ethical constraints are not explicitly incorporated [18,24]. Optimization itself therefore becomes a governance problem rather than a purely technical exercise.

9.3. Optimization Under Clinical and Ethical Constraints

Hospital digital twins frequently frame their contribution in terms of optimization: reducing congestion, improving efficiency, or maximizing utilization. Formally, this corresponds to defining a policy π ( · ) that maps estimated system states to operational actions. Such policies can often be simulated and stress-tested in ways that are infeasible for patient-level digital twins.
Yet operational decisions remain inseparable from clinical risk. Earlier discharge may improve throughput while increasing readmission risk. Staffing reductions may lower costs while increasing burnout and medical errors. These tradeoffs cannot be resolved through optimization alone; they require explicit constraints, risk modeling, and normative judgments [11,24].
Hospital operations therefore illustrate a broader lesson that recurs throughout this review: even when observability and controllability are relatively high, decision-making remains value-laden. Digital twins must therefore operate within governance frameworks that define acceptable objectives, safety margins, and limits of automation [15,35].

9.4. Interoperability, Semantics, and Transportability Across Settings

Operational feasibility does not guarantee scalability. Hospital digital twins face persistent challenges related to interoperability and transportability. Operational data are recorded using heterogeneous information systems with inconsistent semantics, evolving coding practices, and institution-specific conventions. Variables that appear comparable across organizations may encode subtly different meanings in practice. Stable data standards, provenance tracking, and semantic alignment are therefore prerequisites for credible digital twins at scale [51].
Semantic drift presents an additional challenge. As workflows, policies, and documentation practices evolve, the relationship between recorded variables and operational reality can change over time. Without explicit monitoring of such shifts, digital twins may continue optimizing against outdated representations of the system. Consequently, operational digital twins require ongoing semantic governance to remain valid, interpretable, and transferable across settings [23].

9.5. Human Factors and Sociotechnical Dynamics

Even when states and objectives are formally specified, hospital systems remain fundamentally sociotechnical. Staffing decisions, scheduling policies, and workflow changes interact with professional norms, labor agreements, institutional culture, and informal practices. A technically sophisticated digital twin that ignores these factors is unlikely to be adopted [20].
Moreover, optimization itself changes behavior. Clinicians and administrators may adapt to performance targets by altering documentation practices, reallocating effort, or developing workarounds that improve measured metrics without improving care. Awareness of monitoring and optimization objectives can therefore reshape the very state the twin seeks to represent [11,19].
Hospital digital twins thus operate on adaptive organizations rather than passive systems. Organizational learning, resistance, and adaptation must be treated as part of the modeled environment rather than as external noise.

9.6. Hospital Digital Twins as Algorithmically-Mediated Learning Healthcare Systems

These characteristics position hospital operations as the clearest point of convergence between digital twins and the learning healthcare system (LHS) paradigm. LHS frameworks envision continuous cycles in which data generated through routine care are transformed into actionable knowledge and used to improve practice [52,53,54].
Hospital digital twins operationalize this vision through explicit state representations, predictive or simulation-based inference, and policy-driven actions that influence staffing, scheduling, routing, and discharge decisions. Compared with patient-centered digital twins, operational settings provide comparatively favorable conditions for learning because states are more observable, interventions are frequent, and outcomes are measurable [4,55].
However, the LHS literature also provides an important corrective to optimistic narratives of continuous learning. Learning loops are difficult to sustain and depend on governance, incentives, data quality, and institutional culture [56,57,58,59]. Learning does not emerge automatically from data availability.
Hospital digital twins intensify these challenges. By accelerating the translation of data into operational decisions, they may amplify maladaptive learning, optimizing for throughput, cost, or short-term efficiency at the expense of safety, workforce well-being, or equity. Their legitimacy, therefore, depends not on optimization performance alone but on whether learning objectives are explicitly articulated, ethically grounded, and subject to ongoing oversight [53,54,58,59].

9.7. Domain-Specific Research Gaps

Several research gaps emerge from this analysis.
First, clinical risk is rarely integrated into operational twins in a principled manner. Most systems treat clinical outcomes as downstream consequences rather than as explicit optimization constraints [24]. Second, equity impacts are seldom evaluated. Operational policies may affect patient populations differently; yet, these effects are rarely measured or reported [15]. Third, governance mechanisms for adaptive operational twins remain underdeveloped. As policies evolve over time, questions of accountability, transparency, and oversight become increasingly important [20,35]. Finally, transportability across institutions remains a major barrier to large-scale deployment [23,51].

9.8. A Roadmap for Hospital Operations Digital Twins

A realistic roadmap emphasizes disciplined scope, explicit constraints, and governance aligned with learning objectives rather than operational metrics alone.
In the near term, digital twins should focus on well-defined operational problems with clear performance measures and explicit clinical constraints, such as bed management, staffing, or diagnostic scheduling [14,49].
In the medium term, integration of clinical risk models and equity-aware objectives can enable more responsible optimization, supported by prospective evaluations of system-level impact [15,24].
In the longer term, hospital operations digital twins may evolve into components of broader healthcare delivery twins that coordinate patient flow, resource allocation, and quality of care across organizational boundaries. Achieving this vision will require governance structures that evolve alongside technical capabilities and maintain alignment with the principles of accountable learning healthcare systems [11,35].
Hospital operations digital twins demonstrate that favorable conditions for state estimation and intervention do not eliminate the central challenges of healthcare digital twins. Rather, they shift the locus of difficulty from inference to governance. As observability and controllability increase, questions of objective specification, fairness, accountability, and institutional learning become increasingly central. These themes provide a natural bridge to the final synthesis, where the recurring patterns across domains are brought together into a broader assessment of what healthcare digital twins can realistically achieve.

10. Digital Twins Across Scales

The preceding case studies suggest that healthcare digital twins do not constitute a single technological paradigm. Across applications ranging from cardiovascular physiology and diabetes to Alzheimer’s disease, maternal–fetal health, addiction, radiology, and hospital operations, digital twins rely on fundamentally different assumptions regarding state, observation, intervention, and decision-making. Their feasibility, validity, and usefulness therefore depend critically on scale.
Figure 10. Digital twins across scales of healthcare systems. Different scales imply different assumptions about state, observation, intervention, and governance. Some scales admit strong mechanistic grounding and experimentally tractable interventions, whereas others are dominated by uncertainty, adaptation, agency, or organizational objectives. Diabetes occupies an important intermediate regime at the organ-human interface, where feedback, control, and human-AI co-adaptation become central. At organizational scales, digital twins function primarily as instruments of governed optimization rather than individualized inference.
Figure 10. Digital twins across scales of healthcare systems. Different scales imply different assumptions about state, observation, intervention, and governance. Some scales admit strong mechanistic grounding and experimentally tractable interventions, whereas others are dominated by uncertainty, adaptation, agency, or organizational objectives. Diabetes occupies an important intermediate regime at the organ-human interface, where feedback, control, and human-AI co-adaptation become central. At organizational scales, digital twins function primarily as instruments of governed optimization rather than individualized inference.
Preprints 223586 g010
Although healthcare digital twins are often grouped under a common conceptual umbrella, they should not be viewed as a single technology applied at different levels of biological or organizational complexity. Rather, they comprise a family of computational representations whose scientific validity depends on how latent state, uncertainty, intervention, and decision authority are defined at a particular scale.
Table 1 summarizes this perspective. Different scales are associated with different assumptions regarding state, observability, intervention, and decision authority. As a result, they present different technical, epistemic, and governance challenges. At lower biological scales, the dominant concerns involve mechanistic fidelity, calibration, and uncertainty propagation. At patient and human-social scales, challenges increasingly center on non-identifiability, behavioral adaptation, and causal ambiguity. At organizational and population scales, objective specification, competing stakeholder interests, accountability, and governance become central.
At cellular and organ scales, digital twins benefit from relatively well-defined state variables, mechanistic priors, and experimentally grounded interventions. The primary challenges involve model fidelity, parameter estimation, calibration, and uncertainty quantification. Cardiovascular digital twins illustrate how mechanistic knowledge and physiological constraints can support clinically useful state estimation despite incomplete observability.
At patient scales, latent state becomes increasingly difficult to define and observe directly. Alzheimer’s disease highlights the limits imposed by non-identifiability, heterogeneous disease trajectories, and long forecasting horizons. Multiple latent explanations may remain consistent with the same observations, limiting the degree to which individualized prediction can be justified.
Diabetes occupies an important intermediate regime. Frequent observations and actionable interventions make adaptive control plausible; yet, behavioral adaptation, safety constraints, and feedback between prediction and action introduce challenges that are largely absent from purely physiological systems. Maternal-fetal health further demonstrates how coupled state spaces and asymmetric risks constrain decision-making, even when prediction is feasible.
Addiction pushes these difficulties even further. Observation may alter behavior, agency becomes endogenous to the system being modeled, and the notion of a stable latent state becomes contested. In such settings, prediction, intervention, and governance become tightly intertwined,
At organizational scales, digital twins re-emerge in a different form. Hospital operations involve comparatively explicit state variables, measurable outcomes, and controllable interventions. Yet optimization becomes inseparable from questions of safety, workforce impact, equity, resource allocation, accountability, and governance. At these scales, digital twins function less as individualized predictive models and more as instruments for supporting decisions within complex sociotechnical systems. At population scales, these challenges are compounded by emergent dynamics, heterogeneous stakeholders, and contested policy objectives.

10.1. Implications of Scale

The examples examined throughout this review suggest that the feasibility, validity, and limitations of healthcare digital twins are fundamentally scale-dependent. At some scales, the limiting factor is mechanistic fidelity or uncertainty propagation. At others, it is non-identifiability, behavioral adaptation, competing objectives, or governance. The central challenge is therefore not simply building more sophisticated computational models, but ensuring that the claims a digital twin makes are commensurate with what can actually be observed, inferred, predicted, and acted upon at the scale of interest.
This scale-dependent perspective has two important implications. First, it determines what forms of artificial intelligence can be used reliably for state estimation, prediction, simulation, and decision support. Second, it shapes the governance requirements associated with accountability, oversight, legitimacy, and the responsible exercise of decision authority.
The next two sections examine these implications in turn. We first consider how scale shapes the capabilities, limitations, failure modes, and research priorities associated with artificial intelligence in healthcare digital twins. We then examine the governance challenges that arise as digital twins move from representation and prediction toward intervention, decision support, and organizational deployment.

11. Artificial Intelligence in Healthcare Digital Twins: Capabilities, Limits, and Research Priorities

Artificial intelligence (AI) is a central enabling technology for healthcare digital twins, supporting state estimation, prediction, simulation, uncertainty quantification, and decision support. Yet, AI is often treated in the healthcare digital twin literature as synonymous with digital twinning itself. This conflates two distinct ideas.
A digital twin is defined not by the use of AI, but by a sustained computational representation of a specific real-world system together with explicit assumptions regarding state, observation, updating, and intervention. AI methods can support these functions, but they do not establish them [9,10,11]. As discussed in Section 10, healthcare digital twins operate across scales that differ substantially in observability, identifiability, actionability, and governance requirements. Consequently, there can be no single AI strategy for healthcare digital twins. Many limitations attributed to AI-enabled digital twins arise not from deficiencies in machine learning algorithms but from deeper challenges involving state definition, identifiability, causality, decision alignment, and governance.

11.1. What AI Contributes

The scale framework developed in Section 10 helps explain why AI succeeds in some healthcare digital twins and disappoints in others. Across domains, AI contributes through a relatively small set of recurring functions. What varies is not the existence of these functions, but the degree to which they are supported by identifiable states, informative observations, actionable interventions, and governance structures.
Across healthcare applications, AI contributes through four closely related functions: inference under partial observability, prediction and simulation, uncertainty quantification, and decision support.

Inference under partial observability

Healthcare systems are rarely observed directly. AI methods can assist in estimating latent physiological or operational states from high-dimensional observations, including imaging, physiological measurements, electronic health records, laboratory tests, and wearable sensor streams [29,60]. This role is particularly effective when latent states are conceptually well defined and constrained by mechanistic knowledge.
However, increasingly sophisticated representation learning does not eliminate ambiguity regarding the underlying state itself. In domains such as Alzheimer’s disease, multiple latent trajectories may remain consistent with the same observations despite increasingly expressive models [22,31].

Prediction and simulation

AI is widely used to forecast future trajectories and approximate complex dynamics. Learned surrogate models can accelerate simulation, personalize mechanistic models, and support scenario analysis [25,26,36]. Such models can substantially improve computational efficiency and expand the practical scope of digital twins.
Their usefulness, however, depends critically on the relationship between training data and deployment conditions. Outside the regimes represented in training data, surrogate errors may propagate through the digital twin and generate misleading forecasts or recommendations [15,23].

Uncertainty quantification

Probabilistic AI methods provide mechanisms for representing uncertainty through Bayesian inference, calibration, ensembling, and probabilistic forecasting. In principle, these approaches allow digital twins to communicate uncertainty as well as predictions.
In practice, uncertainty is frequently underrepresented. Confidence scores are often interpreted as evidence of reliability even when structural uncertainty, model misspecification, or non-identifiability remain unresolved. The resulting false precision is one of the most persistent failure modes of healthcare digital twins [10,11].

Decision support

AI can support decisions regarding diagnosis, treatment, monitoring, resource allocation, and operational management. In domains characterized by frequent feedback and relatively clear objectives, such as diabetes management or hospital operations, AI-supported decision-making may improve outcomes [36,49,50].
However, decision support requires more than prediction. Legitimate action depends on causal understanding, safety constraints, ethical considerations, and governance. Predictive accuracy alone does not justify intervention.
Importantly, AI can improve inference and prediction; it cannot determine what should be inferred, what can legitimately be predicted, or what actions are justified. Those questions depend on state semantics, identifiability, causal support, decision context, and governance.

11.2. The Limits of AI

Several challenges identified throughout this review lie beyond the reach of AI alone.
First, increasingly sophisticated models cannot compensate for poorly defined latent states. If the object being modeled lacks a coherent state representation, additional model complexity merely produces more elaborate representations of ambiguity.
Second, additional predictive power cannot overcome fundamental non-identifiability. When multiple latent explanations remain consistent with available observations, the missing information is not recoverable through machine learning alone.
Third, most current digital twin models remain predominantly associational. As a result, they provide limited support for counterfactual reasoning, intervention planning, and policy optimization [15,23]. This limitation becomes increasingly important as digital twins move from prediction toward decision support and optimization.
Finally, questions of acceptable risk, fairness, autonomy, accountability, and legitimacy are fundamentally governance questions rather than prediction problems [20,35]. These considerations become increasingly important as digital twins move toward higher levels of autonomy and decision authority.

11.3. Recurring Failure Modes

The limitations described above are not merely theoretical concerns. Across healthcare digital twin applications, they recur as identifiable failure modes. Although the specific manifestations vary across domains and scales, the underlying pattern is consistent: claims about state, prediction, intervention, or optimization exceed what available observations, causal knowledge, and governance structures can legitimately support.
At physiological scales, failure often results from surrogate error, incomplete personalization, or inadequate uncertainty propagation. At patient scales, non-identifiability and sparse observations create risks of false precision and misleading individualized forecasts. In adaptive systems such as diabetes, behavioral feedback can destabilize policies that appear effective under retrospective evaluation. At organizational scales, optimization objectives may become misaligned with broader goals related to safety, equity, workforce sustainability, or trust.
The specific failure modes differ across scales, but they share a common origin: a mismatch between the claims made by a digital twin and the evidence available to support those claims.
Across domains, four patterns recur repeatedly.

False Precision

AI systems frequently generate stable numerical outputs even when uncertainty is substantial. In weakly observable domains, this can create unwarranted confidence in individualized forecasts or risk estimates [11,31]. Such outputs may appear highly personalized while masking fundamental ambiguity regarding the underlying state of the system.

Observation-Dominant Modeling

AI may learn properties of measurement processes rather than underlying system dynamics. Radiology provides a canonical example, where sophisticated image models can achieve impressive predictive performance while remaining disconnected from clinically meaningful state representations [21,30]. Similar problems arise whenever observational artifacts are easier to learn than the mechanisms that generate them.

Behavioral Confounding

In adaptive settings such as diabetes, addiction, and digital health monitoring, predictions and recommendations alter behavior and thereby change the system being modeled. Models that assume passive observation may therefore become invalid once deployed because interventions modify future observations, decision contexts, and user behavior [19,36].

Optimization Without Justification

AI systems often optimize objectives that are convenient to measure rather than clinically or ethically justified. Throughput, utilization, or short-term outcomes may improve while safety, equity, autonomy, or workforce sustainability deteriorate [24,50]. The resulting systems may be technically effective while remaining misaligned with the goals they are intended to serve.
These failure modes are not isolated technical defects. They arise when claims about prediction, understanding, or control exceed what the available data, causal knowledge, and governance structures can support.

11.4. Research Priorities

The roadmap in Figure 11 organizes research priorities across three horizons: epistemic foundations, adaptive decision systems, and institutionally embedded digital twins. These horizons should not be interpreted as independent stages. State semantics, uncertainty representation, causal validity, decision-centered evaluation, adaptive validation, human–AI co-adaptation, equity-aware optimization, and governance remain important throughout. What changes is the scope at which these challenges must be addressed: from individual models and applications in the near term, to adaptive decision systems in the medium term, and ultimately to multi-scale and institutionally embedded systems in the long term.

Near-Term Priorities: Epistemic Foundations

Near-term work should focus on strengthening the epistemic foundations of healthcare digital twins. Priority areas include explicit state semantics, principled treatment of uncertainty and non-identifiability, validity-bounded surrogate modeling, uncertainty reporting standards, causal validity, and evaluation frameworks centered on decision impact rather than predictive accuracy alone. These priorities correspond directly to the foundational layers of the roadmap: state semantics, uncertainty and identifiability, causal validity, and decision-centered evaluation.
Particular attention should be devoted to developing digital twins that communicate not only what is inferred but also what remains uncertain, unidentifiable, or outside the scope of available evidence. In many healthcare settings, honest representation of uncertainty may be more valuable than increasingly confident predictions.

Medium-Term Priorities: Adaptive and Decision-Centered Systems

As healthcare digital twins become more deeply integrated into clinical and operational workflows, research must move beyond static prediction toward adaptive decision support.
Priority areas include integrating causal reasoning with digital twin architectures, developing methods for adaptive validation as systems evolve over time, modeling human-AI co-adaptation, and incorporating fairness, safety, and workforce sustainability as explicit constraints in optimization. These priorities align with the roadmap’s emphasis on causal reasoning, adaptive validation, human-AI co-adaptation, and equity-aware optimization.
These challenges are particularly important in domains such as diabetes management, maternal–fetal health, addiction, and healthcare operations, where recommendations influence future observations and behavior. In such settings, evaluation must account not only for predictive performance but also for the dynamic interaction between models, users, and institutions.

Long-Term Priorities: Multi-Scale and Institutionally Embedded Systems

Long-term progress will depend on connecting digital twins across scales while preserving uncertainty, provenance, and accountability. Priority areas include composable multi-scale digital twins that support reasoning across molecular, physiological, patient, organizational, and population levels; governance-aware infrastructures that support auditability, interoperability, privacy, security, and lifecycle oversight; accountable learning systems with explicit objectives, constraints, adaptation mechanisms, and oversight structures; and learning systems whose objectives and constraints remain explicit and subject to oversight.
These priorities correspond to the roadmap’s emphasis on multi-scale digital twins, governance infrastructures, and accountable learning systems embedded within healthcare institutions. At organizational and population scales, digital twins should increasingly be viewed as components of learning healthcare systems rather than as isolated predictive tools.
Across all horizons, the central objective remains the same: aligning representations, uncertainty estimates, causal claims, recommendations, and oversight mechanisms with the scale and decision context in which a digital twin is deployed.

11.5. Synthesis

TThe central lesson is that there can be no single AI strategy for healthcare digital twins. Each scale presents distinct challenges of state representation, observability, uncertainty, causality, adaptation, and decision authority. Methods that are effective for molecular or physiological digital twins may be inadequate for patient-level, organizational, or population-scale systems because the underlying epistemic and institutional conditions differ.
Progress should therefore be assessed not by predictive performance alone, but by the extent to which digital twins remain aligned with the systems they represent, the decisions they inform, and the evidence available to support those decisions. From this perspective, the central challenge for AI research is not simply improving prediction, but maintaining alignment between representation and reality, uncertainty and action, prediction and intervention, and ultimately between decision authority and accountability.
These considerations motivate the governance agenda developed in the next section. As digital twins move from representation and prediction toward intervention, optimization, and decision support, questions of legitimacy, accountability, oversight, and institutional responsibility become inseparable from technical performance. The challenge is no longer merely to build more capable models, but to ensure that their claims, recommendations, and actions remain appropriately bounded by evidence, authority, and mechanisms for oversight.

12. Governance of Healthcare Digital Twins: Challenges and Research Priorities

Healthcare digital twins do more than generate predictions; they influence decisions. Governance is therefore not primarily a question of regulatory compliance but of legitimate decision authority. Across recent reviews, the principal barriers to deployment are less often algorithmic than organizational: unclear responsibility, weak evidence pathways, inadequate lifecycle monitoring, interoperability failures, and deficits of trust [14,15,23,24].
Unlike static predictive models, digital twins are intended to evolve over time, integrating new observations, updating internal representations, and influencing downstream actions. Governance must therefore address not only whether a system is accurate, but also who may act on its outputs, under what conditions, with what evidence, and under whose oversight. In this sense, governance is not an external constraint on digital twins but a constitutive component of their legitimacy.

12.1. Core Governance Challenges

Decision Authority

A central governance question is not whether a digital twin can generate predictions, but whether it is legitimate to act on them. Across domains, digital twins are frequently described as decision-support systems without clearly specifying where authority resides, what decisions are in scope, and what degree of automation is acceptable.
At patient scales, uncertainty, limited interventional leverage, and asymmetric risks constrain legitimate action. At organizational scales, governance questions shift toward objective setting, value tradeoffs, and accountability for operational decisions. Across scales, a recurring failure mode is the overextension of actionability: systems influence decisions that exceed their epistemic support or ethical mandate [11,15,24].

Lifecycle Oversight

Digital twins are designed to evolve. They incorporate new observations, update internal representations, and may adapt recommendations over time. These characteristics make traditional “train once, validate once” approaches inadequate. Oversight must therefore encompass continuous validation, performance monitoring, drift detection, auditability, and mechanisms for rollback or intervention when performance degrades. Recent reviews consistently identify lifecycle oversight as one of the least developed aspects of healthcare digital twins, despite its central importance for safe deployment [15,23,24].

Data Governance, Provenance, Consent, and Trust

Digital twins depend on longitudinal integration of clinical, behavioral, imaging, and operational data. Consequently, governance challenges extend beyond privacy to include consent, provenance, stewardship, and secondary use.
A distinctive risk is consent erosion: digital twins can generate inferences that exceed the scope of original data-use expectations, even when data are formally de-identified and access controls are compliant. Without credible governance of data use and secondary inference, trust may erode despite technical adherence to privacy requirements [61,62,63].

Scale-Dependent Governance Requirements

Governance burdens vary across scales. At organ and subsystem levels, governance centers on validation, safety, and clearly delimited decision support. At patient and coupled-human scales, uncertainty communication, consent, equity, and legitimacy become increasingly important because errors can have profound consequences. At organizational scales, governance focuses on transparency of objectives, accountability for value tradeoffs, auditability of adaptive policies, and workforce and equity impacts.
Some domains, such as addiction, demonstrate that governance concerns may dominate technical feasibility. In such settings, questions of surveillance, coercion, and power asymmetries can determine whether deployment is ethically permissible at all [19,20].

12.2. Governance Failure Modes

Several governance failure modes recur across domains and scales (Table 2). These failure modes show that governance breakdowns rarely originate from model inaccuracies alone. More often, they arise from misalignment between epistemic claims, decision authority, and institutional oversight. Addressing these failures requires a governance research agenda that spans system design, deployment, adaptation, and organizational integration.

12.3. Research Priorities

The roadmap in Figure 12 is organized around a set of cross-cutting governance foundations that remain relevant across all time horizons: decision authority, accountability, evidence generation, data stewardship, and interoperability. The priorities described below build on these foundations as healthcare digital twins evolve from individual deployments to adaptive systems and ultimately to institutionally embedded infrastructures.

Near-Term Priorities: Governance by Design

Near-term work should focus on making governance operational. Priorities include standardized governance artifacts documenting intended use, decision authority, validation scope, monitoring plans, and escalation procedures; explicit assignment of oversight responsibilities; safety-case approaches for high-consequence applications; and baseline requirements for privacy, security, consent, and secure deployment.
The objective is to make governance operational, auditable, and enforceable from the outset rather than retrofitted after deployment.

Medium-Term Priorities: Adaptive Oversight and Equity

As healthcare digital twins become adaptive and increasingly integrated into clinical and operational workflows, governance must shift from initial approval to continuous oversight.
Research priorities include lifecycle monitoring for performance, calibration, and semantic drift; governance mechanisms for measuring and mitigating equity impacts over time; frameworks for governed optimization in operational settings; and mechanisms that allow affected stakeholders to understand, contest, and override system outputs when appropriate.
The central challenge at this stage is ensuring that adaptive systems remain aligned with their intended purpose as they evolve.

Long-Term Priorities: Adaptive Regulation and Institutional Infrastructure.

Long-term progress depends on aligning healthcare digital twins with durable institutional and regulatory structures.
Research should focus on regulatory frameworks for continuously learning systems, governance-aware infrastructures that support provenance, interoperability, auditability, and policy enforcement, and participatory oversight mechanisms for socially embedded applications. At organizational scales, digital twins should be governed as components of accountable learning healthcare systems with explicit objectives, constraints, accountability structures, and mechanisms for oversight.
The long-term goal is to create institutional environments in which digital twins can evolve responsibly while remaining subject to transparent accountability, effective oversight, and legitimate decision authority.

12.4. Synthesis

The central governance challenge for healthcare digital twins is maintaining alignment among evidence, authority, accountability, and action throughout the system lifecycle. This challenge grows as digital twins become more adaptive, more autonomous, and more deeply embedded in healthcare delivery.
Across the domains examined in this review, trustworthy deployment depends not simply on technical performance but on disciplined limits to decision authority, honest representation of uncertainty, clear accountability, and oversight that evolves as systems and institutions change. Governance is therefore not a constraint on healthcare digital twins; it is part of the infrastructure that makes them scientifically credible, clinically useful, and socially legitimate.
The future of healthcare digital twins depends less on increasingly sophisticated models than on maintaining alignment among representation, evidence, intervention, and accountability across scales. Digital twins become trustworthy not when they predict more, but when the claims they make, the decisions they support, and the actions they enable remain commensurate with what can actually be observed, inferred, justified, and governed.

13. Conclusions

Healthcare digital twins are often portrayed as a transformative technology capable of delivering personalized prediction, optimization, and decision support across medicine. This review argues that the central challenge is not technological capability, but rather establishing when the claims made by a digital twin are scientifically justified and clinically actionable. Across cardiovascular disease, diabetes, Alzheimer’s disease, maternal–fetal health, addiction, radiology, and hospital operations, the feasibility and credibility of a digital twin fundamentally depend on what can be observed, inferred, predicted, and acted upon within a particular domain.
A central finding is that many limitations of healthcare digital twins arise not from insufficient machine learning capability, but from misalignment among state semantics, uncertainty, causal support, decision authority, and governance. When latent state is poorly specified, when non-identifiability is ignored, when uncertainty is collapsed into apparent precision, or when optimization exceeds legitimate authority, digital twins create an illusion of individualized understanding while offering limited support for trustworthy decision-making. These failures are not isolated implementation errors; they are recurring consequences of extending digital twin concepts beyond the conditions under which their claims can be justified.
Artificial intelligence is a central enabling technology for healthcare digital twins. It supports state estimation, prediction, simulation, uncertainty quantification, and decision support across a wide range of applications. However, the value of these capabilities depends on the broader inferential framework in which they are embedded. Machine learning can improve prediction and representation, but the usefulness of a digital twin ultimately depends on how state is defined, how uncertainty is characterized, how interventions are justified, and how decisions are governed.
Governance is equally fundamental. Digital twins allocate authority, mediate risk, and distribute responsibility across patients, clinicians, institutions, and technology providers. As systems become increasingly adaptive and action-oriented, questions of accountability, oversight, contestability, and legitimacy become inseparable from technical performance. Governance is therefore not an external constraint on digital twins but a constitutive component of their responsible use.
The analyzes presented here suggest that healthcare digital twins are best understood not as a single technology but as a family of computational representations whose usefulness and credibility depend on scale, context, and purpose. At some scales, particularly where state variables are well defined, interventions are understood, and uncertainty can be bounded, digital twins can provide meaningful support for prediction, simulation, and decision-making. At other scales, especially where systems are only partially observable, deeply adaptive, or socially embedded, the assumptions required for a strict engineering-style digital twin may not hold. In such settings, more limited computational representations may provide greater scientific validity and practical utility than claims of full digital twinning.
The central contribution of this review is a scale-aware framework for assessing healthcare digital twins. Across domains, the key questions are not whether increasingly sophisticated models can be built, but whether the modeled state is meaningful, whether uncertainty is adequately characterized, whether proposed actions are supported by causal evidence, and whether the authority to act is matched by appropriate oversight and accountability. Viewed through this lens, progress in healthcare digital twins is fundamentally a problem of alignment: alignment between representations and reality, between uncertainty and action, between predictive claims and causal support, and between decision authority and governance. The future of the field will depend not only on advances in computation and artificial intelligence but also on the development of methods, institutions, and governance structures capable of sustaining that alignment. Only under these conditions can healthcare digital twins mature into trustworthy instruments for biomedical discovery, clinical care, and health-system improvement.

Funding

This work was supported in part by grants from the National Science Foundation (2226025) and the National Center for Advancing Translational Sciences of the National Institutes of Health (UL1 TR002014).

Acknowledgments

The figures included in this review were generated and iteratively refined using AI (specifically, ChatGPT 5.3) based on the textual prompts provided by the authors. All of the textual content was written and typeset using Overleaf, a collaborative editor for typesetting documents with LaTeX, with AI use limited to the spell-checking and grammar correction functions offered by Overleaf. Google Scholar and PubMed Central were used to find the cited articles.

References

  1. Sun, T.; He, X.; Li, Z. Digital twin in healthcare: Recent updates and challenges. Digit. Health 2023, 9, 20552076221149651. [Google Scholar] [CrossRef]
  2. Ferdousi, R.; Laamarti, F.; Hossain, M.A.; Yang, C.; El Saddik, A. Digital twins for well-being: an overview. Digit. Twin 2025, 2, 2530296. [Google Scholar] [CrossRef]
  3. Wang, W.; Zaheer, Q.; Qiu, S.; Wang, W.; Ai, C.; Wang, J.; Wang, S.; Hu, W. Digital twins technologies. In Digital twin technologies in transportation infrastructure management; Springer, 2023; pp. 27–74. [Google Scholar]
  4. Vallée, A. Digital twin for healthcare systems. Front. Digit. Health 2023, 5, 1253050. [Google Scholar] [CrossRef] [PubMed]
  5. Li, J.; Yang, S.X. Digital twins to embodied artificial intelligence: review and perspective. Intell. Robot. 2025, 5, 202–227. [Google Scholar] [CrossRef]
  6. Nadeem, M.; Kostic, S.; Dornhöfer, M.; Weber, C.; Fathi, M. A comprehensive review of digital twin in healthcare in the scope of simulative health-monitoring. Digit. Health 2025, 11, 20552076241304078. [Google Scholar] [CrossRef] [PubMed]
  7. Vohra, M. Digital twin technology: fundamentals and applications; John Wiley & Sons, 2023. [Google Scholar]
  8. Willcox, K.; Bingham, D.; Chung, C.; Chung, J.; Cruz-Neira, C.; Grant, C.; Kinter, J.; Leung, R.; Moin, P.; Ohno-Machado, L.; et al. Foundational research gaps and future directions for digital twins; National Academies Press: Washington, DC, USA, 2023. [Google Scholar]
  9. Shengli, W. Is human digital twin possible? Comput. Methods Programs Biomed. Update 2021, 1, 100014. [Google Scholar] [CrossRef]
  10. Vallée, A. Envisioning the future of personalized medicine: role and realities of digital twins. J. Med. Internet Res. 2024, 26, e50204. [Google Scholar] [CrossRef] [PubMed]
  11. Vallée, A. Digital twins for cardiovascular diseases: towards personalised and sustainable care. Acta Cardiol. 2025, 80, 1055–1062. [Google Scholar] [CrossRef] [PubMed]
  12. De Maeyer, C.; Markopoulos, P. Are Digital Twins Becoming Our Personal (Predictive) Advisors? ‘Our Digital Mirror of Who We Were, Who We Are and Who We Will Become’. In Proceedings of the International Conference on Human-Computer Interaction, 2020; Springer; pp. 250–268. [Google Scholar]
  13. Downs, D.S.; Pauley, A.M.; Rivera, D.E.; Savage, J.S.; Moore, A.M.; Shao, D.; Chow, S.M.; Lagoa, C.; Pauli, J.M.; Khan, O.; et al. Healthy Mom Zone Adaptive Intervention With a Novel Control System and Digital Platform to Manage Gestational Weight Gain in Pregnant Women With Overweight or Obesity: Study Design and Protocol for a Randomized Controlled Trial. JMIR Res. Protoc. 2025, 14, e66637. [Google Scholar] [CrossRef] [PubMed]
  14. Tortora, M.; Pacchiano, F.; Ferraciolli, S.F.; Criscuolo, S.; Gagliardo, C.; Jaber, K.; Angelicchio, M.; Briganti, F.; Caranci, F.; Tortora, F.; et al. Medical Digital Twin: A Review on Technical Principles and Clinical Applications. J. Clin. Med. 2025, 14, 324. [Google Scholar] [CrossRef] [PubMed]
  15. Rudsari, H.K.; Tseng, B.; Zhu, H.; Song, L.; Gu, C.; Roy, A.; Irajizad, E.; Butner, J.; Long, J.; Do, K.A. Digital twins in healthcare: a comprehensive review and future directions. Front. Digit. Health 2025, 7, 1633539. [Google Scholar] [CrossRef]
  16. Elgammal, Z.; Albrijawi, M.T.; Alhajj, R. Digital twins in healthcare: a review of AI-powered practical applications across health domains. J. Big Data 2025, 12, 1–28. [Google Scholar] [CrossRef]
  17. Li, T.; Shen, Y.; Li, Y.; Zhang, Y.; Wu, S. The status quo and future prospects of digital twins for healthcare. EngMedicine 2024, 1, 100042. [Google Scholar] [CrossRef]
  18. Bruynseels, K.; Santoni de Sio, F.; Van den Hoven, J. Digital twins in health care: ethical implications of an emerging engineering paradigm. Front. Genet. 2018, 9, 31. [Google Scholar] [CrossRef] [PubMed]
  19. Lupton, D. Language matters: the ‘digital twin’metaphor in health and medicine. J. Med. Ethics 2021, 47, 409–409. [Google Scholar] [CrossRef]
  20. Jørgensen, C.S.; Shukla, A.; Katt, B. Digital twins in healthcare: security, privacy, trust and safety challenges. In Proceedings of the European Symposium on Research in Computer Security, 2023; Springer; pp. 140–153. [Google Scholar]
  21. Pesapane, F.; Rotili, A.; Penco, S.; Nicosia, L.; Cassano, E. Digital twins in radiology. J. Clin. Med. 2022, 11, 6553. [Google Scholar] [CrossRef] [PubMed]
  22. Amato, L.G.; Lassi, M.; Vergani, A.A.; Carpaneto, J.; Mazzeo, S.; Moschini, V.; Burali, R.; Salvestrini, G.; Fabbiani, C.; Giacomucci, G.; et al. Digital twins and non-invasive recordings enable early diagnosis of Alzheimer’s disease. Alzheimer’s Res. Ther. 2025, 17, 125. [Google Scholar] [CrossRef]
  23. Ringeval, M.; Etindele Sosso, F.A.; Cousineau, M.; Paré, G. Advancing health care with digital twins: meta-review of applications and implementation challenges. J. Med. Internet Res. 2025, 27, e69544. [Google Scholar] [CrossRef] [PubMed]
  24. Shen, M.d.; Chen, S.b.; Ding, X.d. The effectiveness of digital twins in promoting precision health across the entire population: a systematic review. npj Digit. Med. 2024, 7, 145. [Google Scholar] [CrossRef] [PubMed]
  25. Corral-Acero, J.; Margara, F.; Marciniak, M.; Rodero, C.; Loncaric, F.; Feng, Y.; Gilbert, A.; Fernandes, J.F.; Bukhari, H.A.; Wajdan, A.; et al. The ‘Digital Twin’to enable the vision of precision cardiology. Eur. Heart J. 2020, 41, 4556–4564. [Google Scholar] [CrossRef] [PubMed]
  26. Coorey, G.; Figtree, G.A.; Fletcher, D.F.; Redfern, J. The health digital twin: advancing precision cardiovascular medicine. Nat. Rev. Cardiol. 2021, 18, 803–804. [Google Scholar] [CrossRef] [PubMed]
  27. Iyer, A.A.; Umadevi, K. Design and analysis of TwinCardio framework to detect and monitor cardiovascular diseases using digital twin and deep neural network. Sci. Rep. 2025, 15, 24376. [Google Scholar] [CrossRef] [PubMed]
  28. Calcaterra, V.; Pagani, V.; Zuccotti, G. Maternal and fetal health in the digital twin era. Front. Pediatr. 2023, 11, 1251427. [Google Scholar] [CrossRef] [PubMed]
  29. Pawar, B.; Prakash, V.; Garg, L.; Galdies, C.; Buttigieg, S.; Calleja, N. Artificial Intelligence and Digital Health Twin Applications in Healthcare—A Systematic Review. Artif. Intell. Healthc. 2024, 1–25. [Google Scholar] [CrossRef]
  30. Mueller, T.T.; Starck, S.; Llalloshi, R.; Kaissis, G.; Ziller, A.; Rueckert, D.; Braren, R. Medical Images as Biomarkers of Ageing-From Global and Local Patterns to Digital Twins. bioRxiv 2025, 2025–05. [Google Scholar]
  31. Jiang, J.; Petrella, J.R.; Hao, W.; Initiative, A.D.N. Identifiability-Guided Assessment of Digital Twins in Alzheimer’s Disease Clinical Research and Care. bioRxiv 2025, 2025–08.
  32. Dolciotti, C.; Righi, M.; Grecu, E.; Trucas, M.; Maxia, C.; Murtas, D.; Diana, A. The translational power of Alzheimer’s-based organoid models in personalized medicine: an integrated biological and digital approach embodying patient clinical history. Front. Cell. Neurosci. 2025, 19, 1553642. [Google Scholar] [CrossRef] [PubMed]
  33. Koksalmis, G.H.; Soykan, B.; Brattain, L.J.; Huang, H.H. Artificial Intelligence for Personalized Prediction of Alzheimer’s Disease Progression: A Survey of Methods, Data Challenges, and Future Directions. arXiv 2025, arXiv:2504.21189. [Google Scholar]
  34. Huang, P.h.; Kim, K.h.; Schermer, M. Ethical issues of digital twins for personalized health care service: preliminary mapping study. J. Med. Internet Res. 2022, 24, e33081. [Google Scholar] [CrossRef] [PubMed]
  35. Lal, A.; Dang, J.; Nabzdyk, C.; Gajic, O.; Herasevich, V. Regulatory oversight and ethical concerns surrounding software as medical device (SaMD) and digital twin technology in healthcare. Ann. Transl. Med. 2022, 10, 950. [Google Scholar] [CrossRef] [PubMed]
  36. Chu, Y.; Li, S.; Tang, J.; Wu, H. The potential of the Medical Digital Twin in diabetes management: a review. Front. Med. 2023, 10, 1178912. [Google Scholar] [CrossRef]
  37. Mosquera-Lopez, C.; Jacobs, P.G. Digital twins and artificial intelligence in metabolic disease research. Trends Endocrinol. Metab. 2024, 35, 549–557. [Google Scholar] [CrossRef] [PubMed]
  38. Cappon, G.; Facchinetti, A. Digital twins in type 1 diabetes: a systematic review. J. Diabetes Sci. Technol. 2025, 19, 1641–1649. [Google Scholar] [PubMed]
  39. Zhang, Y.; Qin, G.; Aguilar, B.; Rappaport, N.; Yurkovich, J.T.; Pflieger, L.; Huang, S.; Hood, L.; Shmulevich, I. A framework towards digital twins for type 2 diabetes. Front. Digit. Health 2024, 6, 1336050. [Google Scholar] [CrossRef] [PubMed]
  40. Cáceres-Gutiérrez, D.A.; Bonilla-Bonilla, D.M.; Liscano, Y.; Díaz Vallejo, J.A. From Architecture to Outcomes: Mapping the Landscape of Digital Twins for Personalized Diabetes Care—A Scoping Review. J. Pers. Med. 2025, 15, 504. [Google Scholar] [CrossRef] [PubMed]
  41. Cappon, G.; Vettoretti, M.; Sparacino, G.; Del Favero, S.; Facchinetti, A. ReplayBG: a digital twin-based methodology to identify a personalized model from type 1 diabetes data and simulate glucose concentrations to assess alternative therapies. IEEE Trans. Biomed. Eng. 2023, 70, 3227–3238. [Google Scholar] [CrossRef] [PubMed]
  42. Prendin, F.; Facchinetti, A.; Cappon, G. Data Augmentation Via Digital Twins to Develop Personalized Deep Learning Glucose Prediction Algorithms for Type 1 Diabetes in Poor Data Context. IEEE Transactions on Biomedical Engineering, 2025. [Google Scholar]
  43. Shamanna, P.; Joshi, S.; Shah, L.; Dharmalingam, M.; Saboo, B.; Mohammed, J.; Mohamed, M.; Poon, T.; Kleinman, N.; Thajudeen, M.; et al. Type 2 diabetes reversal with digital twin technology-enabled precision nutrition and staging of reversal: a retrospective cohort study. Clin. Diabetes Endocrinol. 2021, 7, 21. [Google Scholar] [CrossRef] [PubMed]
  44. Joshi, S.; Shamanna, P.; Dharmalingam, M.; Vadavi, A.; Keshavamurthy, A.; Shah, L.; Mechanick, J.I. Digital twin-enabled personalized nutrition improves metabolic dysfunction-associated fatty liver disease in type 2 diabetes: results of a 1-year randomized controlled study. Endocr. Pract. 2023, 29, 960–970. [Google Scholar] [CrossRef] [PubMed]
  45. Shamanna, P.; Joshi, S.; Thajudeen, M.; Shah, L.; Poon, T.; Mohamed, M.; Mohammed, J. Personalized nutrition in type 2 diabetes remission: application of digital twin technology for predictive glycemic control. Front. Endocrinol. 2024, 15, 1485464. [Google Scholar] [CrossRef]
  46. Surian, N.U.; Batagov, A.; Wu, A.; Lai, W.B.; Sun, Y.; Bee, Y.M.; Dalan, R. A digital twin model incorporating generalized metabolic fluxes to identify and predict chronic kidney disease in type 2 diabetes mellitus. npj Digit. Med. 2024, 7, 140. [Google Scholar] [CrossRef] [PubMed]
  47. Jabin, M.S.R.; Mirza, A.; Ilodibe, A.; Eldabi, T.; Yaroson, E.V. Ethical and quality of care–related challenges of digital health twins in care settings for older adults: scoping review. JMIR Aging 2025, 8, e73925. [Google Scholar] [CrossRef] [PubMed]
  48. Yahya, F.; Cooper, M.; Saif, W.; Kassem, M.; Nazar, H. Development of a Hospital-at-Home Digital Twin for Patients With Frailty: Scoping Review. J. Med. Internet Res. 2025, 27, e81510. [Google Scholar] [CrossRef] [PubMed]
  49. Karakra, A.; Fontanili, F.; Lamine, E.; Lamothe, J. HospiT’Win: a predictive simulation-based digital twin for patients pathways in hospital. In Proceedings of the 2019 IEEE EMBS international conference on biomedical & health informatics (BHI); IEEE, 2019; pp. 1–4. [Google Scholar]
  50. Song, Y.; Li, Y. Digital twin aided healthcare facility management: a case study of Shanghai tongji hospital. Proc. Constr. Res. Congr. 2022, 2022, 1145–1155. [Google Scholar] [CrossRef]
  51. Memon, U.; Mayer, W.; Selway, M.; Stumptner, M. Interoperability of AI-enhanced digital twins. J. Ind. Inf. Integr. 2025, 100961. [Google Scholar] [CrossRef]
  52. Olsen, L.; Aisner, D.; McGinnis, J.M. The learning healthcare system: workshop summary. 2007. [Google Scholar] [CrossRef] [PubMed]
  53. McGinnis, J.M.; Olsen, L.; Goolsby, W.A.; Grossmann, C. Engineering a learning healthcare system: A look at the future: Workshop summary; National Academies Press, 2011. [Google Scholar]
  54. Bindman, A. Learning healthcare systems: a perspective from the US. Public Health Res. Pract. 2019, 29, e2931920. [Google Scholar] [CrossRef]
  55. Xames, M.D.; Topcu, T.G. A systematic literature review of digital twin research for healthcare systems: Research trends, gaps, and realization challenges. IEEE Access 2024, 12, 4099–4126. [Google Scholar] [CrossRef]
  56. Budrionis, A.; Bellika, J.G. The learning healthcare system: where are we now? A systematic review. J. Biomed. Inform. 2016, 64, 87–92. [Google Scholar] [CrossRef] [PubMed]
  57. Enticott, J.; Johnson, A.; Teede, H. Learning health systems using data to drive healthcare improvement and impact: a systematic review. BMC Health Serv. Res. 2021, 21, 200. [Google Scholar] [CrossRef] [PubMed]
  58. Ellis, L.A.; Sarkies, M.; Churruca, K.; Dammery, G.; Meulenbroeks, I.; Smith, C.L.; Pomare, C.; Mahmoud, Z.; Zurynski, Y.; Braithwaite, J. The science of learning health systems: scoping review of empirical research. JMIR Med. Inform. 2022, 10, e34907. [Google Scholar] [CrossRef] [PubMed]
  59. Golburean, O.; Nordheim, E.S.; Faxvaag, A.; Pedersen, R.; Lintvedt, O.; Marco-Ruiz, L. A systematic review and proposed framework for sustainable learning healthcare systems. Int. J. Med. Inform. 2024, 192, 105652. [Google Scholar] [CrossRef] [PubMed]
  60. Balasubramaniam, S.; Sumina, S.; Kumar, K.S.; Prasanth, A. Machine learning based models for implementing digital twins in healthcare industry. In Metaverse Technologies in Healthcare; Elsevier, 2024; pp. 135–162. [Google Scholar]
  61. Evangeline, S.I. Ethical, Privacy, and Security Implications of Digital Twins. In AI-Powered Digital Twins for Predictive Healthcare: Creating Virtual Replicas of Humans; IGI Global Scientific Publishing, 2025; pp. 397–424. [Google Scholar]
  62. Rani, S.; Hasanpuri, V. Data security and ethical considerations in healthcare digital twins. Digital Twin Technology for Better Health: A Healthcare Odyssey; 2025. [Google Scholar]
  63. Sabri, S.; Aghaabbasi, M.; Atkinson, S.R.; Amon, M.J.; Hancoock, P.; Azevedo, R.; Weidbudsh, M.; Maraj, C.; Mondesire, S.; Soykan, B.; et al. Integrating human–machine systems and digital twin technologies: navigating trust, interoperability, and ethical challenges. Cogn. Syst. Res. 2025, 101414. [Google Scholar] [CrossRef]
Figure 11. A scale-aware AI research roadmap for healthcare digital twins. Research priorities progress from epistemic foundations to adaptive systems and ultimately to institutionally integrated digital twins. Near-term efforts focus on state semantics, uncertainty, validity, and decision-centered evaluation. Medium-term priorities emphasize causal reasoning, adaptive validation, human–AI co-adaptation, and equity-aware optimization. Long-term goals include multi-scale digital twins, governance infrastructures, and accountable learning systems. Across all horizons, progress is measured by alignment between epistemic claims, decision authority, and lifecycle governance rather than predictive performance alone.
Figure 11. A scale-aware AI research roadmap for healthcare digital twins. Research priorities progress from epistemic foundations to adaptive systems and ultimately to institutionally integrated digital twins. Near-term efforts focus on state semantics, uncertainty, validity, and decision-centered evaluation. Medium-term priorities emphasize causal reasoning, adaptive validation, human–AI co-adaptation, and equity-aware optimization. Long-term goals include multi-scale digital twins, governance infrastructures, and accountable learning systems. Across all horizons, progress is measured by alignment between epistemic claims, decision authority, and lifecycle governance rather than predictive performance alone.
Preprints 223586 g011
Figure 12. Governance research roadmap for healthcare digital twins. Governance spans the full lifecycle of healthcare digital twins. Foundational priorities include decision authority, accountability, evidence generation, data stewardship, and interoperability. Near-term efforts focus on governance-by-design, including authority mapping, safety cases, and secure deployment. Medium-term priorities emphasize adaptive oversight through continuous validation, equity monitoring, governed optimization, and contestability. Long-term goals include adaptive regulation, governance-aware infrastructures, participatory oversight, and accountable learning healthcare systems. Across all horizons, governance functions as a design requirement rather than a post hoc compliance layer.
Figure 12. Governance research roadmap for healthcare digital twins. Governance spans the full lifecycle of healthcare digital twins. Foundational priorities include decision authority, accountability, evidence generation, data stewardship, and interoperability. Near-term efforts focus on governance-by-design, including authority mapping, safety cases, and secure deployment. Medium-term priorities emphasize adaptive oversight through continuous validation, equity monitoring, governed optimization, and contestability. Long-term goals include adaptive regulation, governance-aware infrastructures, participatory oversight, and accountable learning healthcare systems. Across all horizons, governance functions as a design requirement rather than a post hoc compliance layer.
Preprints 223586 g012
Table 1. Scale-dependent assumptions and challenges in healthcare digital twins. Different scales imply different notions of state, observability, intervention, and decision-making, giving rise to distinct technical, epistemic, and governance challenges.
Table 1. Scale-dependent assumptions and challenges in healthcare digital twins. Different scales imply different notions of state, observability, intervention, and decision-making, giving rise to distinct technical, epistemic, and governance challenges.
Scale State and Observability Actionability Dominant Challenge
Cellular / Molecular Mechanistically grounded; high observability through assays and omics High in vitro or ex vivo Mechanistic fidelity; transportability
Organ / Physiological Subsystem Partially observable; constrained by mechanistic priors Moderate; therapies, procedures, monitoring Surrogate validity; uncertainty propagation
Organ–Human Interface (e.g., diabetes) Feedback-driven; dense but behavior-dependent observations Frequent and risk-sensitive Feedback-aware learning; adaptive control
Individual Human Composite latent state; sparse, biased, care-mediated observations Limited, delayed, uncertain effects Non-identifiability; false precision
Coupled Humans (e.g., maternal–fetal) Joint asymmetric state; uneven observations High-stakes, safety-critical Coupled-state inference; asymmetric uncertainty
Human–Social Interface (e.g., addiction) Subjective, reflexive, socially embedded state Contested and behavior-modifying Reflexive prediction; behavioral feedback
Organizations / Systems (e.g., hospitals) Explicit operational states; frequent but heterogeneous observations High; policies, staffing, routing Objective specification; adaptive optimization
Societal / Population Emergent state; indirect and contested observations Policy-driven Multi-scale modeling; causal attribution
Table 2. Recurring governance failure modes of healthcare digital twins across domains and scales. These failures arise from misalignment between epistemic claims, decision authority, and lifecycle oversight rather than from deficiencies in model accuracy or computational capacity.
Table 2. Recurring governance failure modes of healthcare digital twins across domains and scales. These failures arise from misalignment between epistemic claims, decision authority, and lifecycle oversight rather than from deficiencies in model accuracy or computational capacity.
Governance Failure Mode Structural or Epistemic Origin Observed or Anticipated Consequences
Diffuse or ambiguous accountability Adaptive systems span developers, vendors, clinicians, and institutions without clear responsibility assignment across the lifecycle Inability to assign liability, delayed response to harm, erosion of trust, and stalled deployment in safety-critical settings
Overextension of decision authority Digital twins are treated as decision-makers or optimizers despite limited epistemic support or ethical mandate Unsafe recommendations, coercive or paternalistic interventions, and legitimacy failures
Validation lag in adaptive systems Models update faster than evaluation, audit, and oversight mechanisms can respond Undetected performance drift, silent failure modes, and accumulation of unvalidated changes over time
Unconstrained or opaque optimization Operational objectives are optimized without explicit constraints reflecting safety, equity, or workforce sustainability Efficiency gains that compromise care quality, clinician well-being, or fairness
Consent erosion and secondary-use drift Longitudinal data integration enables inferences beyond original consent scope without explicit governance controls Loss of patient trust, ethical violations, resistance to participation, and institutional risk
Insufficient contestability and appeal Affected stakeholders lack mechanisms to question, override, or appeal digital twin outputs Automation bias, inappropriate deference, reduced professional judgment, and accountability gaps
Governance retrofitting Oversight mechanisms are added post hoc rather than embedded at design time Fragmented compliance, brittle controls, and failure to scale beyond pilot deployments
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings