Preprint
Review

This version is not peer-reviewed.

Artificial Intelligence and Digital Health for Women's Cardiovascular Health: A Survey of Bias, Equity, and Inclusive Innovation

Submitted:

28 August 2026

Posted:

31 August 2026

You are already at the latest version

Abstract
Cardiovascular disease is the leading cause of death among women worldwide. However, women remain underdiagnosed, undertreated, and underrepresented in the clinical studies, health records, images, and device data used to build artificial intelligence (AI). Machine learning, large language models, and consumer digital health tools can therefore inherit and amplify a century of male-default medicine. When designed and deployed with equity in mind, the same tools can help identify and reduce these disparities.This survey connects three literatures that are usually reviewed separately: women's cardiovascular health, algorithmic fairness in medical AI, and digital health and “FemTech.” We organize the evidence around four questions. Where does inequity enter the AI lifecycle? Which AI applications address women's cardiovascular needs? How do digital health and FemTech change access and risk? Which technical and governance methods can improve fairness?Across imaging, risk prediction, and language models, current studies document disadvantages for women. Imbalanced imaging data reduce diagnostic accuracy. Chest-radiograph classifiers underdiagnose female patients. Language models overgenerate male cardiac cases and can downplay women's needs. We also review sex-specific risk models, electrocardiographic and imaging systems, multimodal foundation models, simulation, telehealth, mobile health, and wearables. The methods literature covers bias measurement, auditing, mitigation, causal and mechanistic analysis, regulation, open science, and World Health Organization ethics guidance. We close with open problems in intersectionality, sex-disaggregated data, the reproductive life course, real-world fairness, and benchmarks designed for women's cardiovascular AI. Overall, equity must be incorporated throughout system design and deployment rather than treated as a post-deployment correction.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

Cardiovascular disease (CVD) is the single largest cause of death among women. It accounts for roughly a third of female deaths globally [1,2]. Despite this burden, the perception of heart disease as a “man’s disease” persists in both public and clinical settings. Women present with a broader and often less “textbook” range of symptoms. They are diagnosed later, referred for invasive procedures less often, and prescribed guideline-directed therapy less reliably. Cardiovascular trials have also enrolled women at rates far below their share of disease burden [3,4,5]. These disparities have measurable consequences. In the Asia Pacific region, women face longer diagnostic delays for heart attacks and systematic gaps across research, education, access, treatment, and investment [6]. These are not isolated features of one disease. Together, they reflect a structural pattern in which the male body has served as the default subject of biomedical knowledge.
Artificial intelligence presents both opportunities and risks in this setting. Machine learning (ML) and deep learning can analyze electrocardiograms (ECGs), echocardiograms, and electronic health records (EHRs) at scale. Large language models (LLMs) and multimodal foundation models extend this capability. These methods may reveal sex-specific risk hidden by pooled models and extend screening to settings where cardiologists are scarce [7,8,9].
However, these systems are trained on data generated by inequitable care processes. Historical records may omit disease that was not recognized in women. Models trained on these records can underestimate women’s risk and reproduce underdiagnosis in groups whose disease has historically been missed [10,11]. AI can therefore either detect inequity or reproduce it (Figure 1). The outcome depends on decisions made throughout the system lifecycle, not only on the algorithm. Automated systems can also encode and scale social inequity outside medicine [12,13]. This survey examines these risks and opportunities in women’s cardiovascular health.
Digital health is also changing care delivery through telehealth, mobile applications, wearables, remote monitoring, and the growing “FemTech” sector. These tools can decentralize cardiovascular prevention and self-management and broaden trial participation. They can also reinforce the digital divide or commercialize unmet needs without improving care [14,15]. Clinical AI and digital health are increasingly interconnected. LLM chatbots are digital health products, while wearables provide data for AI systems. Few surveys examine these areas together, particularly in the context of women’s cardiovascular health.
Three bodies of work inform this survey. Clinical and epidemiological studies examine women’s cardiovascular health and sex and gender differences. A technical literature studies algorithmic fairness in medical AI, especially in imaging and, more recently, LLMs. Digital-health and human–computer-interaction (HCI) research examines telehealth, mobile health (mHealth), and FemTech. General surveys exist within each area. They include fairness benchmarks for medical imaging [16], calls for open and unbiased health data [17], and reviews of sex and gender bias in biomedical AI [18]. What is missing is an integrative survey at their intersection. This survey addresses that gap by treating women’s cardiovascular health as a high-stakes domain in which to examine how AI and digital health create, measure, and may reduce inequity.
We make the following contributions.
  • We propose a unifying organization of the field (Section 2, Table 2) that connects four threads: sources of inequity, AI applications, digital health and FemTech, and fairness methods and governance. These threads are organized around women’s cardiovascular health.
  • We synthesize empirical evidence from medical imaging, tabular risk prediction, and language models. This evidence shows that contemporary cardiovascular and clinical AI systems can disadvantage women. We map each finding to the lifecycle stage at which the bias originates (Section 3, Table 4).
  • We review AI applications for women’s cardiovascular care and the digital health/FemTech ecosystem, drawing out where sex- and gender-specific design is present, absent, or actively contested (Section 4Section 5).
  • We survey methods and governance for fairness, including measurement, mitigation, causal and mechanistic approaches, regulation, and open science. We also examine well-supported negative results that limit the expected benefits of technical interventions (Section 6Section 7).
  • We articulate an agenda of open challenges, arguing for equity as a design constraint and for benchmarks and data infrastructure specific to women’s cardiovascular AI (Section 8).
We assembled the corpus through three complementary search streams. The screening counts appear in Table 1. Appendix A records the databases, query strings, inclusion and exclusion rules, and de-duplication procedure. Stream 1 comprises a curated set of recent clinical and policy articles on women’s cardiovascular health, sex/gender differences, and AI or digital health (sourced from cardiology and digital health venues such as Circulation, JACC: Advances, the Lancet Digital Health, and Current Cardiology Reports).
Stream 2 comprises landmark and topical open-access papers on bias in medical AI, retrieved by targeted queries against open repositories (PubMed Central, PNAS, Nature portfolio, arXiv) for the field’s most-cited results. Stream 3 comprises a broader pool of fairness- and health-AI papers from the machine-learning and natural-language-processing literatures. We retrieved it by running roughly twenty thematic queries against a large indexed corpus and screening the results for relevance. The queries covered sex and gender bias, demographic fairness in imaging and EHRs, clinical-LLM bias, fairness benchmarks, and mitigation methods.
We applied three inclusion criteria. A work had to concern artificial intelligence or digital health. It also had to address women’s cardiovascular health or provide transferable evidence about clinical fairness, equity, or demographic bias. Finally, it had to be peer-reviewed or recognized as a preprint or standard. We emphasized work published between 2019 and 2026 because the clinical and technical literatures have developed rapidly. We also retained seminal earlier results (e.g., [11,19]). We excluded work with no methodological or evidential bearing on the intersection, purely promotional material, and duplicate reports of the same study.
Each retained article was read and condensed into a common structured summary with five coded fields: background, gap, contribution, method, and evidence. We used these fields as comparison axes throughout (e.g., for the bias-evidence map in Table 4 and the method families in Table 5). The searches were run in May–June 2026. Example queries include “women, cardiovascular AI, sex differences”, “underdiagnosis bias, chest radiograph”, and “fairness benchmark, medical imaging”.
After screening, the structured review corpus comprised about eighty papers (Table 1, Figure 2). We also cite the canonical literature on women’s cardiovascular health and fairness in health, including foundational sex/gender-medicine statements, landmark bias studies, and standard fairness methods. Together, these sources account for the 106 references in this survey. This is a structured integrative review. It is not a PRISMA-style systematic review or meta-analysis. We document each search stream and inventory the resulting corpus, but we do not claim exhaustive recall within any single sub-area. All quantitative values are drawn from the cited primary studies. Section 8.1 discusses the resulting scope limitations.
The rest of this paper is organized as follows. Section 2 gives background on women’s cardiovascular health and on the AI and digital health methods involved, and presents our organizing taxonomy. Section 3 examines the sources of inequity and the empirical evidence of bias against women. Section 4 surveys AI applications for women’s cardiovascular care. Section 5 covers digital health and FemTech. Section 6 reviews fairness measurement and mitigation methods. Section 7 addresses governance, ethics, regulation, and open science. Section 8 sets out open challenges and future directions, and Section 9 concludes the paper.

2. Background and Scope

2.1. Women’s Cardiovascular health

Before examining the clinical evidence, we first establish the key terminology used throughout this paper. We use sex to refer to biological attributes, including chromosomal, hormonal, and anatomical characteristics. We use gender to refer to socially constructed roles and identities. These concepts interact, but source datasets frequently conflate them or fail to record either one. Most cardiovascular datasets provide, at best, a binary sex or gender field. Much of the literature also uses “sex” and “gender” interchangeably. When a cited study uses these terms ambiguously, we preserve the terminology of the original study. This inconsistency limits the available evidence and contributes to the measurement bias discussed in Section 3.
Beyond differences in prevalence and presentation, women differ from men in cardiovascular anatomy, physiology, risk-factor weighting, and treatment response [20,21,22]. Women also experience sex-specific and sex-predominant conditions, including pregnancy-associated cardiovascular disease, spontaneous coronary artery dissection, and the coronary microvascular dysfunction characterized by the WISE study [23,24]. These differences have clinical consequences. Women with acute myocardial infarction may present atypically and are diagnosed and treated later than men [25,26]. Several structural deficits recur across the literature. Cardiovascular trials and the datasets derived from them have historically skewed toward white men. Reporting of sex, race and ethnicity, or socioeconomic status (SES) remains incomplete [4,27]. Even when women are enrolled, results are frequently pooled. This practice hides sex-dependent effects [6,28].
Updated prevention guidance increasingly reflects these sex-specific considerations [3,5,29]. Cardiovascular disadvantage also compounds across sex, race, ethnicity, sexual identity, and socioeconomic position. For example, sexual-minority women from several racial and ethnic groups show lower overall cardiovascular health than their heterosexual counterparts on a nationally representative measure [30]. These structural deficits shape the clinical data from which AI systems learn.

2.2. AI and Digital Health: A Primer

We use artificial intelligence broadly to include classical machine learning on tabular data, deep learning on signals and images, and the transformer-based large language and multimodal foundation models that now play a prominent role in the field. Broad reviews describe this trajectory in medicine [31,32]. In cardiovascular medicine, supervised deep networks read ECGs, echocardiograms, and cardiac magnetic resonance imaging (MRI) and computed tomography (CT). Gradient-boosted and random-forest models estimate risk from EHRs. They are increasingly joined by representation-learning approaches over longitudinal records [33,34]. LLMs summarize records, answer questions, and draft documentation [7,8,9].
Digital health refers to the technologies and services used to deliver care. It includes telemedicine, short-message and app-based interventions, wearables and biosensors, remote monitoring, and decentralized trials. FemTech refers to the commercial segment of consumer products targeting women’s health, from cycle and fertility tracking to menopause services. These categories overlap: a chatbot is both an LLM application and a digital health product, and a wearable is both a consumer device and a data source for AI. Table 2 presents the taxonomy that structures this survey. It relates four functional threads to the lifecycle stages at which design decisions and bias arise. The taxonomy organizes the sections that follow and locates individual contributions within the broader field.

2.3. Relation to Prior Surveys

Existing reviews address individual components of this topic. However, to our knowledge, none integrates all three fields around women’s cardiovascular health. Table 3 compares representative prior surveys with the scope of this article. Reviews of fairness in medical imaging provide detailed analyses of one modality but are not specific to women or cardiovascular care [16,42]. Reviews of sex and gender bias in biomedical AI are domain-general and predate the LLM era [18]. Work on open and unbiased health data addresses governance more broadly [17]. Clinical reviews of digital health or AI for women’s cardiovascular care cover applications but engage only lightly with technical fairness methods [3,38].
This survey integrates these strands through the specific setting of women’s cardiovascular health. It examines clinical inequity, technical evidence of AI bias, and digital health delivery within a single framework. The review is focused and integrative rather than exhaustive within any one field. Table 7 summarizes the strength of evidence for each main conclusion.

3. The Double-Edged Sword: How AI Encodes and Amplifies Inequity

3.1. Where Bias Enters the Lifecycle

A consistent theme across recent cardiovascular and biomedical AI reviews is that bias is not only a property of training data. It enters at every stage of the AI lifecycle (Figure 3). The lifecycle framework also identifies opportunities to promote equity (Figure 1) [17,18,38]. This lifecycle view of harm is now standard across health-AI fairness scholarship [45,46,47,48,49].
At problem formulation, the choice of prediction target can encode inequity. A widely deployed population-health algorithm illustrates the problem. It used health costs as a proxy for health needs and systematically underestimated the needs of Black patients, who incur lower costs at equal sickness [19]. At data collection, women, pregnant patients, and gender minorities are undersampled, and access barriers mean their disease is differentially unrecorded [37]. Geographic and institutional skew compounds this problem. A model that generalizes at one hospital can fail at another [50,51,52].
At labeling and measurement, ground-truth itself can be biased when historical diagnoses reflect clinician bias. At modeling and evaluation, aggregate accuracy can mask large subgroup error gaps. At deployment and post-deployment learning, feedback loops and uneven access to the technology can entrench disparities. Cirillo et al. [18] distinguish clinically meaningful sex and gender differences from harmful bias. Models should preserve the former and remove the latter. This distinction guides the analysis that follows.
Several studies adapt this lifecycle view to cardiovascular care. Mihan, Pandey, and Van Spall present an AI health-equity framework for CVD that pairs each bias source with mitigation strategies [38,53]. Their synthesis of primary cardiovascular AI studies found that few studies use representative data, objective labels, or subgroup-aware evaluation. Thamman et al. [37] reframe cardiovascular AI explicitly as a health-equity problem. They show how access-driven data missingness, biased reimbursement, and nonrepresentative validation cohorts can produce models that perform well during development but fail in diverse clinical settings. Amponsah et al. [54] examine how AI can help identify and reduce racial and ethnic cardiovascular disparities. Such benefits require representative data, subgroup fairness testing, and continuous post-deployment monitoring. Patient-facing tools present similar concerns. Healthcare chatbots can produce biased or culturally insensitive guidance, which motivates equality-, diversity-, and inclusion-focused design throughout the chatbot lifecycle [55].

3.2. Empirical Evidence of Bias Against Women

Empirical studies increasingly show that clinical models can disadvantage women. Table 4 summarizes representative findings across modalities. It is useful to distinguish two types of evidence. The first is directly women- and cardiovascular-specific. It includes sex-disparate error in cardiac-disease models [36], sex-specific risk structure [28], over-generation of male cardiac cases [56], and skewed heart-failure datasets [35]. The second type is transferred evidence from the broader clinical-AI fairness literature. In these studies, the mechanism is general but has direct implications for women. Examples include chest-radiograph underdiagnosis, which is most pronounced among female patients [10], and imaging shortcuts that encode demographic attributes [57].
Transferred evidence can operate along different demographic axes. Some findings concern sex directly, such as the lower female performance caused by sex-imbalanced imaging data. Other findings were demonstrated for race rather than sex. Examples include the cost-as-proxy algorithm [19] and pulse-oximeter bias [58]. These findings remain relevant because the underlying mechanisms can harm multiple underrepresented groups. Race-based evidence should not be presented as direct evidence of sex-based bias. Transferred evidence remains necessary because women’s cardiovascular AI remains understudied. General clinical-AI results are often the best available signal, but they are not a substitute for direct audits. This evidence gap is a primary motivation for the research agenda in Section 8.
In medical imaging, a controlled study by Larrazabal et al. [11] trained thoracic-disease classifiers on chest X-rays under varying male/female training ratios and showed that the underrepresented gender suffers consistent accuracy loss. Even a 25/75 imbalance significantly lowered minority-group performance across architectures and diseases. Seyyed-Kalantari et al. [10] demonstrated a clinically important form of underdiagnosis. Chest-radiograph classifiers disproportionately labeled diseased patients as “no finding.” This false-reassurance error concentrated in female patients, younger patients, and racial and ethnic minorities. The largest gaps occurred at their intersections.
Part of the mechanism is that imaging models encode demographic identity directly. Banerjee, Gichoya, and colleagues [57] showed that deep networks can predict self-reported race from radiographs, CT, and mammograms with high accuracy, even after aggressive image corruption and across external datasets. This result indicates a demographic shortcut with no known human-visible correlate [59]. Such findings extend a longer line of demographic-bias-in-vision work into cardiovascular-relevant imaging. Earlier examples include the “Gender Shades” audit of commercial gender classifiers [60] and disparities in dermatology [61]. The findings have spurred dedicated fairness methods, including group-fairness guarantees [62] and equalized-coverage conformal predictors [63].
In tabular risk prediction, sex disparities are equally present in classical ML on structured data. Straw, Rees, and Nachev [36] reproduced cardiac-disease models on open datasets and found significantly higher false-negative rates for women in 13 of 16 experiments. A literature audit in the same study found that only three of 127 papers discussed sex differences, and none examined race. Kwak et al. [28] trained separate sex-specific risk models on a 258,279-person cohort. Their interpretable analysis showed that cholesterol-related factors had greater weight for men, whereas age and waist circumference had greater weight for women. Pooled models can conceal this clinically meaningful heterogeneity. A systematic review identified broader limitations in AI cardiovascular risk models. Across 79 studies and 486 models, none had undergone independent external validation, and all were judged to have a high risk of bias [64].
Recent evidence also concerns LLMs. Ducel et al. [56] fine-tuned seven models to generate French clinical cases. The models systematically overgenerated male patients for cardiac and other conditions, at odds with real prevalence. Rickman’s counterfactual study of long-term-care summarization found that one widely used open model emphasized men’s physical and mental-health needs while downplaying women’s, even though a contemporaneous model showed no measurable gap [65]. The contrast shows that bias is model-specific and must be audited for each system. Broader clinical-LLM studies report a similar pattern. Zhao et al. [66] introduced a diagnosis-bias score over 330,000 records and 193 diseases. They found, for example, stricter myocardial-infarction diagnosis in women. Ahsan et al. [43] localized gender encoding to specific MLP activations. Intervening on those activations could flip the patient’s gender in generated vignettes and shift downstream risk predictions.
Beyond per-system audits, Wu et al. [67] proposed a model-agnostic allocation–deterioration framework that scores the inequality an AI system induces, separating inequality already present in data from inequality the model adds. Applied to ICU cohorts, it found women up to 33% worse off on prognostic markers. All four evaluated allocation models also induced significant additional inequality against non-white patients.

3.3. Underrepresentation and Opacity in Datasets

Underlying many of these failures is the state of the data itself. A systematic review of heart-failure AI datasets found that of 72 datasets covering more than two million people, 85% reported sex but only 29% reported race/ethnicity and 11% reported SES. Among datasets that reported race, 89% of individuals were white. Only 20 datasets were fully accessible [35] (Figure 4). Such opacity makes representativeness impossible to assess and reproducibility difficult. It also interacts with dataset drift. Public medical-imaging datasets can accumulate duplicates, lose metadata, and acquire unclear licenses as they are copied across platforms. Norori et al. [17] characterize bias as a data, measurement, human, and governance problem. Modeling choices alone cannot resolve these sources of bias. For women’s cardiovascular AI, subgroup performance should therefore be measured before a model is considered equitable.

4. AI Applications for Women’s Cardiovascular Health

4.1. Risk Prediction and Screening

Figure 5 situates the applications below along the care pathway. Sex-aware risk stratification is an important application of AI. Its goal is to improve on pooled clinical scores while identifying sex-specific mechanisms. The sex-specific random-forest models of Kwak et al. [28] illustrate this use of ML. They match pooled cohort equations on discrimination while revealing different risk-factor profiles for women and men. AI-enabled ECG analysis is a particularly active screening frontier. Relevant methods include self-supervised and contrastive ECG representations [68], early-classification policy networks [69], and photoplethysmography-to-ECG translation [70]. The last approach could extend rhythm analysis to consumer wearables.
Scientific statements and reviews describe deep models that infer reduced ventricular function or future heart failure from routine ECGs, with particular relevance to peripartum and postpartum women. In this population, early detection can prompt timely echocardiography and treatment [7,71]. Mammography is acquired in women at population scale for cancer screening. It has also been proposed as an opportunistic source of data for cardiovascular risk estimation. This approach uses a population-scale, women-specific data stream for cardiovascular screening [5]. However, a systematic review found that current AI CVD risk models remain largely unvalidated [64].

4.2. Diagnosis and Cardiac Imaging

Deep learning has automated measurement, segmentation, and classification across echocardiography, cardiac CT, MRI, and nuclear imaging. Sengupta et al. [9] survey this shift toward “augmented” rather than “replacement” intelligence. Video-based networks can estimate ejection fraction beat by beat. A blinded randomized trial also found AI echocardiographic assessment non-inferior to sonographer assessment. Efficient video-segmentation models support this work [72]. Data scarcity, weak external validation, and workflow mismatch remain persistent constraints. The same drive toward verifiable automation extends to therapeutic devices, where formal methods have been used to synthesize correct-by-construction cardiac pacemaker controllers [73]. For women, imaging presents both clinical opportunities and documented equity risks. It has produced important diagnostic advances, but it also contains some of the clearest evidence of model bias described in Section 3. The methods literature therefore increasingly examines how to develop imaging models that are both accurate and sex-equitable (Section 6).

4.3. Large Language and Multimodal Foundation Models

LLMs and multimodal foundation models extend cardiovascular AI beyond single-task classifiers. These systems can integrate EHRs, imaging, biosensors, and genomics and communicate in natural language [8]. Quer and Topol argue these models can operate with incomplete annotations and generate clinician- and patient-intelligible output. They also identify confabulation, privacy, and limited prospective evidence as major concerns. Reports from the clinical community similarly emphasize deployment readiness rather than benchmark performance alone. Examples include guideline-tuned chatbots and multi-step clinical agents. They also note that medical-exam benchmarks poorly reflect real workflows [74], a concern that has driven dedicated medical-LLM benchmarks [75]. Because these same models exhibit the gender and demographic biases documented in Section 3, their cardiovascular deployment is inseparable from the fairness agenda.

4.4. Simulation, Digital Twins, and Synthetic Data

Where real sex-balanced data are scarce, simulation offers an alternative. Safdar et al. [39] couple a machine-learning classifier trained on guideline-derived synthetic records to an agent-based model. The system reproduces expected cardiovascular risk dynamics without patient data. This approach may be useful when access to patient records is constrained by privacy, although its validity depends on the fidelity of the guideline encoding. Synthetic data and digital twins may help augment underrepresented subgroups. Recent methods synthesize longitudinal patient records across multiple visits [76]. However, naive generation can preserve or amplify the correlation and representation biases it is meant to fix (Section 6).
The evidence for these applications remains preliminary. Sex-specific modeling and AI-enabled ECG and imaging show what is possible, yet almost none of these tools has prospective, externally validated evidence in women specifically [64]. These applications extend AI beyond generic risk scoring toward tools designed for women’s cardiovascular phenotypes. Their clinical value will depend on appropriate fairness and governance practices.

5. Digital Health and FemTech for Women’s Cardiovascular Care

5.1. Telehealth, Mobile Health, Wearables, and Remote Monitoring

Digital health provides decentralized care that may address barriers faced by women, including caregiving burdens, distance, and cost. Azizi et al. [3] review evidence that text-message and structured telephone support, telehealth, and app- and wearable-based monitoring can improve cardiovascular outcomes and adherence. The reviewed trials and meta-analyses include reductions in adverse cardiovascular events. User-engagement research helps explain when such tools are actually used [77]. Mobile-health interventions have been used extensively in low-resource maternal programs. Field deployments have improved engagement in large maternal-and-child-health cohorts [78,79]. Some deployments use reinforcement-learning schedulers to decide whom to call. This setting is directly relevant to peripartum cardiovascular risk.
In heart failure, Myhre et al. [41] map digital tools across the care continuum from case-finding to chronic monitoring. They emphasize that the largest gaps and opportunities lie in under-resourced settings. These tools are increasingly AI-enabled. Remote photoplethysmography and signal-quality-aware heart-rate estimation can extract cardiac signals from ordinary cameras and wrist-worn sensors [80,81,82]. Digital health and clinical AI are therefore parts of the same system.
Sensor measurement bias affects this delivery layer. The physiological signals on which wearables and remote-monitoring tools depend can be systematically less accurate for some groups. These signals include photoplethysmography for heart rate and rhythm and optical or cuffless blood-pressure estimation. The clearest documented case is the pulse oximeter, which overestimates blood oxygen in patients with darker skin and thereby propagates a biased measurement into biased downstream clinical decisions [58]. The equity of a digital health tool is therefore determined partly upstream of any model, at the point of measurement. Women’s underrepresentation in device-validation cohorts is therefore also relevant [18,38].

5.2. Equity by Design and the Digital Divide

Digital health can narrow access gaps or widen them, depending on whether equity is engineered in. Nair et al. [14] formalize this as an “equity by design” framework that treats equitable access as a design constraint. It combines policy support, broadband and device provision, digital navigators, participatory design, and outcome stratification. In the reported quality-improvement work, targeted support meaningfully increased video-visit uptake among Black and Hispanic patients. The HCI literature indicates that surface-level localization is insufficient. Translation and visual adaptation do not address relational context or culturally specific motivations among diverse groups of women. Equity-by-design therefore emphasizes participatory, culturally grounded methods [14,55]. Editorials on workforce and care diversity add that digital literacy support and inclusive design must accompany the tools themselves [83].

5.3. FemTech, Commercialization, and the Illusion of Empowerment

FemTech also introduces commercial and clinical risks. Nickel et al. [15] analyze the commercialization of women’s health technologies. Their examples include cycle and fertility tracking, direct-to-consumer hormone testing, menopause products, and elective egg freezing. They argue that “empowerment” rhetoric can mask uncertain clinical validity, normalize self-surveillance, monetize intimate data, and shift responsibility from health systems onto individual women. In cardiovascular care, consumer enthusiasm and clinical evidence can diverge sharply. Evidentiary and regulatory standards must therefore keep pace with marketing. This critical analysis also shows that access to a consumer product should not be equated with clinical benefit.

5.4. Human-Centered and Participatory Design

Human-centered research grounds digital tools in the experiences of women and their clinicians. Jacob et al. [40] conducted a structured needs assessment for an AI-enabled women’s heart-health application, finding a dual-value structure: patients prioritize control, awareness of sex-specific symptoms, personalization, and reduced anxiety, while clinicians prioritize engagement, adherence, workflow integration, and validated benefit. Such human-centered methods help ensure that sex-specific content, including pregnancy, menopause, and hormonal context, is present by design rather than omitted. At the research-system level, digital health technologies can also make women’s health research more inclusive and participatory across the study lifecycle [84]. This approach can give participants greater influence over the research that concerns them. Participatory design nevertheless remains limited in cardiovascular applications and wearables.
Digital health equity depends on how technologies are implemented and accessed. The same telehealth service or wearable may extend care or deepen the digital divide. FemTech may expand access, but it may also commercialize unmet needs without improving care. Accordingly, equity in delivery requires direct measurement. Uptake and outcomes should be reported by group. Whether commercial incentives can support such measurement without regulation remains an open question.

6. Methods for Fairness and Bias Mitigation

6.1. Measuring and Auditing Bias

Bias mitigation requires reliable measurement. Common measures include group accuracy and area-under-the-curve (AUC) gaps, true- and false-positive/negative-rate disparities, equalized odds, and calibration within subgroups. Benchmarks can standardize comparisons across these measures. The most common criteria are demographic parity, which requires equal positive-prediction rates, and equal opportunity, which requires equal true-positive rates. Equalized odds requires equal true- and false-positive rates. Calibration within groups requires predicted risk to match observed outcomes in each group. Equalized coverage requires prediction sets to be valid within each subgroup. These criteria are provably incompatible in general [85,86]. A women’s-cardiovascular deployment therefore cannot satisfy all of them at once. It must choose which error to equalize and defend that choice clinically. Equal false-negative rates may deserve priority because underdiagnosis of women is the dominant documented harm (Section 3).
MEDFAIR [16] evaluated eleven bias-mitigation algorithms across ten medical-imaging datasets, multiple sensitive attributes, and model-selection strategies, in both in- and out-of-distribution settings. No mitigation algorithm significantly outperformed well-tuned empirical risk minimization (ERM). Complementary tools target specific failure modes. Uncertainty- and coverage-based methods (e.g., conformal prediction with equalized coverage) attach subgroup-valid confidence to predictions, and the allocation–deterioration framework of Wu et al. [67] scores induced inequality directly. For women’s cardiovascular AI, sex-stratified error reporting should be standard. Aggregate accuracy hid the female false-negative gaps uncovered by Straw et al. [36]. Post-hoc explainability is sometimes presented as evidence of trustworthiness. However, current methods can provide false assurance and do not independently establish fairness [87].

6.2. Mitigation Across the Lifecycle

Mitigation strategies map onto the lifecycle they target. Data-level methods rebalance or augment training data, including fairness-aware synthetic generation that seeks to preserve minority-subgroup densities and break spurious correlations rather than merely match overall statistics [88]. In-processing methods modify training through adversarial removal of protected information, fairness-constrained objectives, parameter-efficient fine-tuning (PEFT) that selects which parameters to update for fairness, mixture-of-experts routing conditioned on attributes, and mutual-information regularizers that push representations toward demographic invariance.
Post-processing methods adjust thresholds or coverage per group after training, building on the classical equal-opportunity construction [89]. Causal approaches examine the mechanisms through which disparities arise. Deconfounding methods correct for unmeasured confounders when learning treatment policies from observational records [90]. Path-specific effect analysis can isolate the part of a treatment disparity mediated by a biased measurement device. Fair survival models address informative censoring that distorts risk across groups [91].
Mechanistic approaches are more recent. Ahsan et al. [43] use activation patching to locate and manipulate demographic encodings inside clinical LLMs. Related work probes whether sparse autoencoders can reveal and steer racial associations in clinical text. Prompt-based debiasing offers a lightweight option for LLMs. Zhao et al. [66] reduce diagnosis bias by masking demographic cues and warning against overdiagnosis, although the effect may not transfer across settings. Table 5 organizes these families by the lifecycle stage they target and pairs each with a representative approach and its principal caveat.

6.3. Limitations of Current Fairness Methods

Several well-supported negative results indicate that no current intervention is universally effective. In-distribution fairness may not transfer. Yang et al. [42] trained thousands of chest-radiograph models. Models that were fair on the training distribution often lost that property under real-world distribution shift. False-negative-rate gaps reached 30%, and in-distribution fairness often failed to predict out-of-distribution fairness. Models that encoded less demographic information sometimes generalized more fairly. The authors therefore propose this measure as a model-selection criterion.
MEDFAIR provides evidence at a larger scale. It compared eleven mitigation algorithms on ten datasets and trained roughly seven thousand models, using about 6,800 GPU-hours. No method statistically outperformed well-tuned ERM in either in- or out-of-distribution settings [16]. Some methods achieve parity by “leveling down.” They reduce performance for an advantaged group instead of improving it for a disadvantaged group. This outcome can be clinically unacceptable. Subgroup performance also varies substantially across tasks. Based on three clinical datasets, Roller et al. [92] argue that routine subgroup reporting and validation provide a more dependable safeguard than a universal fairness algorithm.
These findings support the broader argument that algorithmic fairness methods have ethical and legal limits in health care [93,94]. They also support the use of inherently interpretable models in high-stakes settings [95]. Demographic errors in facial-analysis tools, including BMI estimators, illustrate these limits [96]. For women’s cardiovascular AI, fairness must be treated as an ongoing engineering and governance practice rather than a single post-hoc intervention.

6.4. Benchmarks and the Demographic-Diagnosis Distinction

Medical benchmark design must account for demographic differences that have clinical relevance. DiversityMedQA [97] perturbs patient demographics in medical questions while filtering out items whose correct answer should legitimately change, isolating model bias from genuine clinical dependence. This turns the “desirable vs. undesirable difference” described by Cirillo et al. [18] into an evaluation protocol. Related biomedical-NLP audits similarly probe whether models encode stereotyped disease–demographic associations [66,97]. For cardiovascular care, where sex differences in presentation are both real and historically used to dismiss women, this distinction has practical consequences. A fair system must preserve useful sex-specific signals while discarding harmful stereotypes. The methods literature therefore offers useful but partial tools. Measurement and auditing should be routine. Mitigation remains an open problem, and no single technique dominates.

7. Governance, Ethics, Regulation, and Open Science

7.1. Ethical Frameworks and the WHO Principles

Technical mitigation operates within an ethical and regulatory framework. The World Health Organization’s guidance on the ethics and governance of AI for health articulates six principles: protect autonomy; promote well-being and safety; ensure transparency; foster accountability; ensure inclusiveness and equity; and promote responsiveness and sustainability. The guidance also covers data, public and private organizations, and global coordination [44].
In cardiovascular care, Adedinsewo [71] describes AI health equity as a collective ethical responsibility. A risk–benefit lens also distinguishes decision support from automation. It favors the early use of low-risk, externally validated systems in high-need populations where equity gains are plausible. These frameworks make equity a primary requirement. They build on calls to ensure fairness, avoid harm, and address ethical challenges when machine learning enters care [98,99,100,101]. Two further points recur. Opaque “race correction” in clinical algorithms can entrench disadvantage and requires scrutiny [102]. When developed for this purpose, AI can also reduce unexplained disparities. For example, a learned pain model narrowed the gap between radiologist scores and pain reported by underserved patients [103].

7.2. Dataset Transparency and Open Science

Because so much inequity originates in data, data governance is essential. The heart-failure dataset audit [35] and broader concerns about the quality and provenance of shared medical datasets [17] argue for documentation standards covering demographics, provenance, licensing, and persistent identifiers. These practices support the FAIR principles of findable, accessible, interoperable, and reusable data. Datasheets and model cards can make them concrete [104,105]. Norori et al. [17] present open science as a bias intervention. Their recommendations include participant-centered design, inclusive data sharing, interoperable standards, code sharing, and methods that synthesize underrepresented data. These practices support the assessment, replication, and correction of fairness problems. For women’s cardiovascular AI, sex-disaggregated reporting is the most immediately actionable standard.
Open data creates a tension with privacy, and that tension is not gender-neutral. Reproductive, pregnancy, and hormonal data are sensitive but valuable for closing evidence gaps. Commercial FemTech products often collect these data outside clinical data-protection norms [15]. Federation, differential privacy, and fairness-aware synthetic data can support representative datasets without unrestricted data pooling. The WHO guidance also treats data governance and benefit sharing as core parts of ethical AI for health [44].

7.3. Regulation and Clinical-Trial Reform

Regulatory regimes are adapting: cardiovascular AI deployment increasingly contends with medical-device regulation, post-market surveillance, and emerging AI-specific law, themes prominent in clinical-community reports [74]. Upstream, the evidence base itself is being reformed. Diversity, equity, and inclusion “primers” and ecosystem models for cardiovascular research provide phase-by-phase strategies for recruiting and retaining women and other underrepresented groups [4,27], and scoping-review protocols seek to map how equity is operationalized across the AI lifecycle in healthcare [106]. Methodological work on more inclusive and adaptive trial designs aims to recruit and retain underserved subgroups while preserving regulatory-grade evidence [4,27], pointing toward trials that are both efficient and equitable. More representative trials produce better training and evaluation data, which can support fairer models.

8. Open Challenges and Future Directions

Table 6 summarizes the open challenges. The discussion below examines each challenge in detail.

8.0.0.1. Intersectionality as the default unit of analysis.

The most severe disparities appear at intersections, such as female and Black, female and Hispanic, or female and a sexual minority [10,30]. Yet most audits and mitigations operate one attribute at a time, and intersectional subgroups quickly become too small for reliable estimation. Methods that borrow strength across related subgroups, and reporting norms that make intersectional analysis routine, are therefore needed.

8.0.0.2. Sex-disaggregated data infrastructure.

Many failures trace to missing or unreported sex and the pooling of sexes in analysis [35,36]. The field needs shared, well-documented, sex-disaggregated cardiovascular datasets. Sex-stratified evaluation should also become a publication requirement. This intervention is feasible for most studies and could substantially improve transparency.

8.0.0.3. Pregnancy and the reproductive life course.

Pregnancy-associated cardiovascular disease, the menopausal transition, and hormonal context are simultaneously high-stakes, sex-specific, and poorly covered by both datasets and consumer tools [40,71]. AI that explicitly models the female cardiovascular life course remains underexplored. Such models should not treat male physiology as the default.

8.0.0.4. From in-distribution to real-world fairness.

Given that in-distribution fairness does not reliably transfer [42] and that few cardiovascular AI models are externally validated [64], prospective and multi-site evaluation should assess both performance and fairness. Where feasible, this evaluation should be conducted in randomized or pragmatic trials.

8.0.0.5. Trustworthy foundation models and mechanistic fairness.

As multimodal foundation models enter cardiovascular care [8], the mechanistic tools that localize and steer demographic encodings [43] remain at an early stage. It is not yet clear whether representation-level interventions improve clinical fairness without adverse effects on other model behavior.

8.0.0.6. Evidence standards for FemTech.

Consumer cardiovascular and women’s-health products require independent clinical evaluation [15]. Regulatory requirements should be proportionate to the clinical claims made for each product.

8.0.0.7. Benchmarks for women’s cardiovascular AI.

The field lacks benchmarks built specifically for this intersection. Needed resources include curated, consented, and sex-balanced cardiovascular tasks for ECG, imaging, risk, and LLM-based applications. They should include intersectional subgroup labels and distinguish demographic bias from clinically relevant differences [97]. Such benchmarks would let the community measure progress on the problem this survey describes.

8.1. Limitations of This Survey

This survey has three main limitations. The field is moving quickly and spans clinical, technical, and HCI venues with different publication practices. Despite our broad search, we may have missed relevant work. Our emphasis on 2019–2026 also means that some older foundational results appear only through later citations. This is an integrative, qualitative survey rather than a PRISMA-style systematic review or meta-analysis. We favor breadth across three fields over exhaustive coverage within a narrow query. Coverage of each sub-area is therefore selective. The literature is concentrated in English-language and high-income-country settings. It also tends to encode sex as a binary variable. These properties limit what we can conclude about gender-diverse people and low-resource contexts. They are also substantive gaps in the field.

8.2. An Evidence-Strength Map

This survey draws on both direct and transferred evidence. Table 7 maps the basis of each main conclusion. We label evidence as direct when it comes from women-specific cardiovascular studies and transferred when it comes from the wider clinical-AI fairness literature. We use agenda-setting for positions and research directions rather than settled empirical findings. We also provide a qualitative assessment of the current evidence strength. The map identifies the conclusions for which direct evidence is still missing.

9. Conclusion

This survey integrates three rapidly developing fields: women’s cardiovascular health, medical AI, and digital health. AI systems can inherit and scale male-default assumptions in medicine. They can also support the detection and reduction of cardiovascular inequity. The evidence reviewed here demonstrates both outcomes. Contemporary imaging, tabular, and language models can disadvantage women and intersecting groups. Sex-specific modeling, equitable digital health design, and fairness auditing show how the same technologies can also help.
The evidence also identifies important limitations. Fairness may fail under distribution shift. Mitigation methods may not outperform empirical risk minimization. Some methods obtain parity by lowering performance for an advantaged group. These findings show that no single algorithm can ensure equity. Equity therefore requires a continuous set of practices rather than a single technique. It must be incorporated throughout the lifecycle and monitored after deployment. Concrete next steps include sex-disaggregated data, intersectional and real-world evaluation, and purpose-built benchmarks. With these safeguards, AI and digital health can help build cardiovascular science that serves women.

Funding

This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.

Institutional Review Board Statement

Not applicable. This study is a review of published literature and did not involve human participants, animals, or identifiable personal data.

Data Availability Statement

No new datasets were generated or analysed for this review. All evidence is drawn from publicly available sources cited in the manuscript.

Acknowledgments

Use of AI-assisted tools: During manuscript preparation, OpenAI Codex and Anthropic Claude were used to assist with literature organization, manuscript review, language editing, LaTeX formatting, and quality assurance. The authors reviewed and revised all outputs, verified the cited evidence, and take full responsibility for the content of the manuscript.

Conflicts of Interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Appendix A. Search and Screening Protocol

This appendix documents the procedure used to construct the corpus described in Section 2 and Table 1. The protocol supports a structured integrative review and does not claim the exhaustive recall expected of a PRISMA systematic review.

Appendix A.0.0.8. Sources.

Stream 1 (clinical/policy) drew on cardiology and digital health venues, principally Circulation, JACC: Advances, the Lancet Digital Health, Current Cardiology Reports, Current Cardiovascular Risk Reports, JAMA Network Open, and BMC Medical Informatics and Decision Making. Stream 2 (seminal medical-AI bias) queried open repositories: PubMed Central, PNAS, the Nature portfolio, Patterns, arXiv, and the WHO publications library. Stream 3 (fairness/health-AI methods) queried a large index spanning the ACL Anthology, NeurIPS, ICLR, ICML, CVPR/ECCV, AAAI, IJCAI, and OpenAlex. Searches were run in May–June 2026.

Appendix A.0.0.9. Queries.

Stream 3 used roughly twenty thematic queries. Queries on sex, gender, and representation included “women, cardiovascular heart disease, artificial intelligence, sex differences”; “gender bias, large language models, healthcare, clinical”; “sex bias, machine learning, clinical prediction, fairness”; and “sex-disaggregated data, electronic health records.”
Queries on data and equity included “health equity, artificial intelligence, disparities, underserved” and “underrepresentation, women, clinical trials, data diversity.” Other queries addressed dataset bias, demographic representation, and racial or ethnic disparities in health AI. Application queries covered ECGs, cardiac imaging, digital health, maternal health, wearables, and foundation models. Mitigation queries included “fairness mitigation, reweighting, healthcare audit.”

Appendix A.0.0.10. Inclusion / exclusion.

A record was included if it (a) concerned AI or digital health, (b) bore on women’s cardiovascular health or on transferable clinical fairness/bias, and (c) was peer-reviewed or a recognized preprint/standard. A record was excluded if it had no methodological or evidential bearing on the intersection, was purely promotional, or duplicated a study already retained.

Appendix A.0.0.11. De-duplication and coding.

Duplicates (preprint/published pairs, and works returned by more than one stream) were merged by title and DOI, keeping the version of record.
Each retained record was read and coded on five fields: background, gap, contribution, method, and evidence. These fields served as the comparison axes in Table 4 and Table 5. Screening reduced 226 screened records to a review corpus of about eighty papers (Figure 2). The survey also cites canonical women’s-cardiovascular and fairness-in-health literature, for 106 references in total.

Appendix B. Corpus Ledger by Sub-Area

Table A1 is a reader-verifiable map of the cited corpus, grouped by sub-area, with whether each cluster is directly women-cardiovascular-specific (W-CVD) or transferred from the broader clinical-AI fairness literature (transfer), and the section it supports.
Table A1. Corpus ledger: representative cited works by sub-area, evidence type, and the section each supports. The list is representative, not exhaustive; the full set is the reference list.
Table A1. Corpus ledger: representative cited works by sub-area, evidence type, and the section each supports. The list is representative, not exhaustive; the full set is the reference list.
Sub-area Representative works Type §
Women’s CVD foundations [1,2,20,22,23,24,25,26] W-CVD 2
Imaging bias [10,11,57,59,60,61] transfer 3
Tabular / risk bias [28,36,64] W-CVD 3
LLM clinical bias [43,56,65,66,97] mixed 3
Datasets / representation [17,35,50,51] mixed 3
Lifecycle / equity frameworks [18,37,38,45,46,54] W-CVD 3
AI cardiology applications [7,8,9,68,70,72] W-CVD 4
Digital health & FemTech [3,14,15,40,41,58] W-CVD 5
Fairness methods [16,42,67,88,89,91] transfer 6
Governance / ethics [44,87,98,99,102,103] transfer 7
Trials / DEI reform [4,27,29] W-CVD 7

References

  1. Vogel, B.; Acevedo, M.; Appelman, Y.; et al. The Lancet Women and Cardiovascular Disease Commission: Reducing the Global Burden by 2030. The Lancet 2021, 397, 2385–2438. [Google Scholar] [CrossRef]
  2. Tsao, C.W.; Aday, A.W.; Almarzooq, Z.I.; et al. Heart Disease and Stroke Statistics—2023 Update: A Report From the American Heart Association. Circulation 2023, 147, e93–e621. [Google Scholar] [CrossRef]
  3. Azizi, Z.; Adedinsewo, D.; Rodriguez, F.; et al. Leveraging Digital Health to Improve the Cardiovascular Health of Women. Curr. Cardiovasc. Risk Rep. 2023, 17, 205–214. [Google Scholar] [CrossRef]
  4. Prichard, R.; Maneze, D.; Straiton, N.; Inglis, S.C.; McDonagh, J. Strategies for Improving Diversity, Equity, and Inclusion in Cardiovascular Research: A Primer. Eur. J. Cardiovasc. Nurs. 2024, 23, 313–322. [Google Scholar] [CrossRef]
  5. Mehran, R.; et al. It’s Time for Artificial Intelligence to Go Red For Women. Circulation 2026, 153, 477–479. [Google Scholar] [CrossRef]
  6. Allen, et al. Cardiovascular Health Inequity for Women in the Asia Pacific Region, 2024. Report on women’s cardiovascular health and the gender health gap in the Asia Pacific region.
  7. Armoundas, A.A.; et al. Use of Artificial Intelligence in Improving Outcomes in Heart Disease: A Scientific Statement From the American Heart Association. Circulation 2024, 149, e1028–e1050. [Google Scholar] [CrossRef]
  8. Quer, G.; Topol, E.J. The Potential for Large Language Models to Transform Cardiovascular Medicine. Lancet Digit. Health 2024, 6, e767–e771. [Google Scholar] [CrossRef]
  9. Sengupta, P.P.; Dey, D.; Davies, R.H.; Duchateau, N.; Yanamala, N. Challenges for Augmenting Intelligence in Cardiac Imaging. Lancet Digit. Health 2024, 6, e739–e748. [Google Scholar] [CrossRef]
  10. Seyyed-Kalantari, L.; Zhang, H.; McDermott, M.B.A.; Chen, I.Y.; Ghassemi, M. Underdiagnosis Bias of Artificial Intelligence Algorithms Applied to Chest Radiographs in Under-Served Patient Populations. Nat. Med. 2021, 27, 2176–2182. [Google Scholar] [CrossRef]
  11. Larrazabal, A.J.; Nieto, N.; Peterson, V.; Milone, D.H.; Ferrante, E. Gender Imbalance in Medical Imaging Datasets Produces Biased Classifiers for Computer-Aided Diagnosis. Proc. Natl. Acad. Sci. 2020, 117, 12592–12594. [Google Scholar] [CrossRef]
  12. O’Neil, C. Weapons of Math Destruction: How Big Data Increases Inequality and Threatens Democracy; Crown Publishing Group, 2016. [Google Scholar]
  13. Benjamin, R. Race After Technology: Abolitionist Tools for the New Jim Code; Polity Press, 2019. [Google Scholar]
  14. Nair, P.; Dai, M.; Patel, D.A.; et al. Equity by Design: Using Digital Technology to Overcome Cardiovascular Health Disparities. Curr. Cardiol. Rep. 2025. [Google Scholar] [CrossRef]
  15. Nickel, B.; et al. The Illusion of Empowerment: Commercializing Women’s Health in the Digital Age. Health Promot. Int. 2026, 41, daag074. [Google Scholar] [CrossRef]
  16. Zong, Y.; Yang, Y.; Hospedales, T.M. MEDFAIR: Benchmarking Fairness for Medical Imaging. In Proceedings of the International Conference on Learning Representations (ICLR), 2023. [Google Scholar]
  17. Norori, N.; Hu, Q.; Aellen, F.M.; Faraci, F.D.; Tzovara, A. Addressing Bias in Big Data and AI for Health Care: A Call for Open Science. Patterns 2021, 2, 100347. [Google Scholar] [CrossRef]
  18. Cirillo, D.; Catuara-Solarz, S.; Morey, C.; Guney, E.; Subirats, L.; Mellino, S.; et al. Sex and Gender Differences and Biases in Artificial Intelligence for Biomedicine and Healthcare. npj Digit. Med. 2020, 3, 81. [Google Scholar] [CrossRef]
  19. Obermeyer, Z.; Powers, B.; Vogeli, C.; Mullainathan, S. Dissecting Racial Bias in an Algorithm Used to Manage the Health of Populations. Science 2019, 366, 447–453. [Google Scholar] [CrossRef]
  20. Mauvais-Jarvis, F.; Bairey Merz, N.; Barnes, P.J.; et al. Sex and Gender: Modifiers of Health, Disease, and Medicine. The Lancet 2020, 396, 565–582. [Google Scholar] [CrossRef]
  21. Regitz-Zagrosek, V. Sex and Gender Differences in Health. EMBO Rep. 2012, 13, 596–603. [Google Scholar] [CrossRef]
  22. Garcia, M.; Mulvagh, S.L.; Bairey Merz, C.N.; Buring, J.E.; Manson, J.E. Cardiovascular Disease in Women: Clinical Perspectives. Circ. Res. 2016, 118, 1273–1293. [Google Scholar] [CrossRef]
  23. Bairey Merz, C.N.; Shaw, L.J.; Reis, S.E.; et al. Insights From the NHLBI-Sponsored Women’s Ischemia Syndrome Evaluation (WISE) Study. J. Am. Coll. Cardiol. 2006, 47, S21–S29. [Google Scholar] [CrossRef]
  24. Maas, A.H.E.M.; Appelman, Y.E.A. Gender Differences in Coronary Heart Disease. Neth. Heart J. 2010, 18, 598–602. [Google Scholar] [CrossRef]
  25. Mehta, L.S.; Beckie, T.M.; DeVon, H.A.; et al. Acute Myocardial Infarction in Women: A Scientific Statement From the American Heart Association. Circulation 2016, 133, 916–947. [Google Scholar] [CrossRef]
  26. Woodward, M. Cardiovascular Disease and the Female Disadvantage. Int. J. Environ. Res. Public Health 2019, 16, 1165. [Google Scholar] [CrossRef]
  27. Prichard, R.; et al. Research 4 All: Addressing the Diversity, Equity, and Inclusion Deficit in Cardiovascular Research, 2023. Concept. Res.-Ecosyst. Model Divers. Equity Incl. Cardiovasc. Res. [CrossRef]
  28. Kwak, S.; Lee, H.J.; Kim, S.; Park, J.B.; Lee, S.P.; Kim, H.K.; Kim, Y.J. Machine Learning Reveals Sex-Specific Associations Between Cardiovascular Risk Factors and Incident Atherosclerotic Cardiovascular Disease. Sci. Rep. 2023, 13, 9355. [Google Scholar] [CrossRef]
  29. Cho, L.; Davis, M.; Elgendy, I.; et al. Summary of Updated Recommendations for Primary Prevention of Cardiovascular Disease in Women: JACC State-of-the-Art Review. J. Am. Coll. Cardiol. 2020, 75, 2602–2618. [Google Scholar] [CrossRef]
  30. Rosendale, N.; et al. Differences in Cardiovascular Health at the Intersection of Race, Ethnicity, and Sexual Identity. JAMA Netw. Open 2024, 7, e249060. [Google Scholar] [CrossRef]
  31. Topol, E.J. High-Performance Medicine: The Convergence of Human and Artificial Intelligence. Nat. Med. 2019, 25, 44–56. [Google Scholar] [CrossRef] [PubMed]
  32. Rajpurkar, P.; Chen, E.; Banerjee, O.; Topol, E.J. AI in Health and Medicine. Nat. Med. 2022, 28, 31–38. [Google Scholar] [CrossRef]
  33. Katsuki, T.; et al. Cumulative Stay-time Representation for Electronic Health Records in Medical Event Time Prediction. In Proceedings of the Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, 2022. [Google Scholar] [CrossRef]
  34. Gullapalli, B.T.; et al. Pharmacokinetics-Informed Neural Network for Predicting Opioid Administration Moments with Wearable Sensors. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2024. [Google Scholar] [CrossRef]
  35. Laws, E.; et al. Diversity and Inclusion Within Datasets in Heart Failure: A Systematic Review. JACC Adv. 2025. [Google Scholar] [CrossRef]
  36. Straw, I.; Rees, G.; Nachev, P. Sex-Based Performance Disparities in Machine Learning Algorithms for Cardiac Disease Prediction: Exploratory Study. J. Med. Internet Res. 2024, 26, e46936. [Google Scholar] [CrossRef]
  37. Thamman, R.; Yong, C.M.; et al. Role of Artificial Intelligence in Cardiovascular Health Disparities: The Risk of Greasing the Slippery Slope. JACC Adv. 2023, 2. [Google Scholar] [CrossRef]
  38. Mihan, A.; Pandey, A.; Van Spall, H.G.C. Mitigating the Risk of Artificial Intelligence Bias in Cardiovascular Care. Lancet Digit. Health 2024. [Google Scholar] [CrossRef]
  39. Safdar, et al. Integrating Clinical Assessment Indicators into Cardiovascular Risk Event Simulation Using Machine Learning and Agent-Based Modeling, 2026. Hybrid. Mach.-Learn. Agent-Based Model. Framew. Cardiovasc. Risk Simul. [CrossRef]
  40. Jacob, et al. Bridging Gaps in Women’s Heart Health: A User-Centered Needs Assessment of Patient and Clinician Interviews, 2026. Qualitative human-centered design study of an AI-enabled women’s cardiovascular health application.
  41. Myhre, P.L.; Tromp, J.; Ouwerkerk, W.; Lam, C.S.P.; et al. Digital Tools in Heart Failure: Addressing Unmet Needs. Lancet Digit. Health 2024, 6, e755–e766. [Google Scholar] [CrossRef]
  42. Yang, Y.; Zhang, H.; Gichoya, J.W.; Katabi, D.; Ghassemi, M. The Limits of Fair Medical Imaging AI in Real-World Generalization. Nat. Med. 2024, 30, 2838–2848. [Google Scholar] [CrossRef]
  43. Ahsan, H.; Sen Sharma, A.; Amir, S.; Bau, D.; Wallace, B.C. Elucidating Mechanisms of Demographic Bias in LLMs for Healthcare. Proc. Find. Assoc. Comput. Linguist. EMNLP 2025, 14614–14631. [Google Scholar] [CrossRef]
  44. World Health Organization. Ethics and Governance of Artificial Intelligence for Health: WHO Guidance; Technical report; World Health Organization: Geneva, 2021. [Google Scholar]
  45. Suresh, H.; Guttag, J. A Framework for Understanding Sources of Harm throughout the Machine Learning Life Cycle. In Proceedings of the Equity and Access in Algorithms, Mechanisms, and Optimization (EAAMO), 2021. [Google Scholar] [CrossRef]
  46. Celi, L.A.; Cellini, J.; Charpignon, M.L.; et al. Sources of Bias in Artificial Intelligence that Perpetuate Healthcare Disparities—A Global Review. PLoS Digit. Health 2022, 1, e0000022. [Google Scholar] [CrossRef]
  47. Panch, T.; Mattie, H.; Atun, R. Artificial Intelligence and Algorithmic Bias: Implications for Health Systems. J. Glob. Health 2019, 9, 010318. [Google Scholar] [CrossRef]
  48. Parikh, R.B.; Teeple, S.; Navathe, A.S. Addressing Bias in Artificial Intelligence in Health Care. JAMA 2019, 322, 2377–2378. [Google Scholar] [CrossRef]
  49. Chen, I.Y.; Joshi, S.; Ghassemi, M. Treating Health Disparities with Artificial Intelligence. Nat. Med. 2020, 26, 16–17. [Google Scholar] [CrossRef]
  50. Kaushal, A.; Altman, R.; Langlotz, C. Geographic Distribution of US Cohorts Used to Train Deep Learning Algorithms. JAMA 2020, 324, 1212–1213. [Google Scholar] [CrossRef]
  51. Zech, J.R.; Badgeley, M.A.; Liu, M.; Costa, A.B.; Titano, J.J.; Oermann, E.K. Variable Generalization Performance of a Deep Learning Model to Detect Pneumonia in Chest Radiographs: A Cross-Sectional Study. PLoS Med. 2018, 15, e1002683. [Google Scholar] [CrossRef]
  52. Gianfrancesco, M.A.; Tamang, S.; Yazdany, J.; Schmajuk, G. Potential Biases in Machine Learning Algorithms Using Electronic Health Record Data. JAMA Intern. Med. 2018, 178, 1544–1547. [Google Scholar] [CrossRef]
  53. Mihan, A.; Pandey, A.; Van Spall, H.G.C. Artificial Intelligence Bias in the Prediction and Detection of Cardiovascular Disease. npj Cardiovasc. Health 2024, 1, 31. [Google Scholar] [CrossRef]
  54. Amponsah, D.; Thamman, R.; Brandt, E.J.; Yong, C.M. Artificial Intelligence to Promote Racial and Ethnic Cardiovascular Health Equity. Curr. Cardiovasc. Risk Rep. 2024, 18, 153–162. [Google Scholar] [CrossRef]
  55. Kucukkaya, et al. Equality, Diversity, and Inclusion in Artificial Intelligence-Driven Healthcare Chatbots: Addressing Challenges and Shaping Strategies, 2025. Discussion of equality, diversity and inclusion across the healthcare chatbot lifecycle.
  56. Ducel, F.; Hiebel, N.; Ferret, O.; Fort, K.; Névéol, A. “Women Do Not Have Heart Attacks!” Gender Biases in Automatically Generated Clinical Cases in French. Proc. Find. Assoc. Comput. Linguist. NAACL 2025, 7160–7174. [Google Scholar]
  57. Banerjee, I.; Bhimireddy, A.R.; Burns, J.L.; Gichoya, J.W.; et al. Reading Race: AI Recognises Patient’s Racial Identity in Medical Images. arXiv 2021, arXiv:2107.10356. [Google Scholar] [CrossRef]
  58. Sjoding, M.W.; Dickson, R.P.; Iwashyna, T.J.; Gay, S.E.; Valley, T.S. Racial Bias in Pulse Oximetry Measurement. N. Engl. J. Med. 2020, 383, 2477–2478. [Google Scholar] [CrossRef]
  59. Gichoya, J.W.; Banerjee, I.; Bhimireddy, A.R.; et al. AI Recognition of Patient Race in Medical Imaging: A Modelling Study. Lancet Digit. Health 2022, 4, e406–e414. [Google Scholar] [CrossRef]
  60. Buolamwini, J.; Gebru, T. Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification. In Proceedings of the Proceedings of the Conference on Fairness, Accountability and Transparency (FAccT), 2018; pp. 77–91. [Google Scholar]
  61. Adamson, A.S.; Smith, A. Machine Learning and Health Care Disparities in Dermatology. JAMA Dermatol. 2018, 154, 1247–1248. [Google Scholar] [CrossRef]
  62. Luo, Y.; et al. On Demographic Group Fairness Guarantees in Deep Learning. IEEE Trans. Pattern Anal. Mach. Intell. 2026. [Google Scholar] [CrossRef]
  63. Lu, C.; et al. Fair Conformal Predictors for Applications in Medical Imaging. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2022. [Google Scholar] [CrossRef]
  64. Cai, Y.; Cai, Y.Q.; Tang, L.Y.; Wang, Y.H.; Gong, M.; Jing, T.C.; et al. Artificial Intelligence in the Risk Prediction Models of Cardiovascular Disease and Development of an Independent Validation Screening Tool: A Systematic Review. BMC Med. 2024, 22, 56. [Google Scholar] [CrossRef]
  65. Rickman, S. Evaluating Gender Bias in Large Language Models in Long-Term Care. BMC Med. Inform. Decis. Mak. 2025, 25, 274. [Google Scholar] [CrossRef]
  66. Zhao, Y.; Wang, H.; Liu, Y.; Wu, S.; Wu, X.; Zheng, Y. Can LLMs Replace Clinical Doctors? Exploring Bias in Disease Diagnosis by Large Language Models. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP, 2024; pp. 13914–13935. [Google Scholar]
  67. Wu, H.; Wang, M.; Sylolypavan, A.; Wild, S. Quantifying Health Inequalities Induced by Data and AI Models. In Proceedings of the Proceedings of the 31st International Joint Conference on Artificial Intelligence (IJCAI), 2022; pp. 5192–5198. [Google Scholar]
  68. Wang, N.; et al. Adversarial Spatiotemporal Contrastive Learning for Electrocardiogram Signals. IEEE Trans. Neural Netw. Learn. Syst. 2023. [Google Scholar] [CrossRef]
  69. Huang, Y.; Yen, G.G.; Tseng, V.S. Snippet Policy Network for Multi-class Varied-length ECG Early Classification. IEEE Trans. Knowl. Data Eng. 2022. [Google Scholar] [CrossRef]
  70. Shome, D.; Sarkar, P.; Etemad, A. Region-Disentangled Diffusion Model for High-Fidelity PPG-to-ECG Translation. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2024. [Google Scholar] [CrossRef]
  71. Adedinsewo, D.A. Advancing Cardiovascular Health Equity With Artificial Intelligence: A Collective Ethical Responsibility. Circulation 2024, 150, 174–176. [Google Scholar] [CrossRef]
  72. Wu, H.; Lin, J.; Xie, W.; Qin, J. Super-efficient Echocardiography Video Segmentation via Proxy- and Kernel-Based Semi-supervised Learning. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2023. [Google Scholar] [CrossRef]
  73. Dole, K.; et al. Correct-by-Construction Reinforcement Learning of Cardiac Pacemakers from Duration Calculus Requirements. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2023. [Google Scholar] [CrossRef]
  74. European Society of Cardiology. Igniting the Future of Cardiovascular Care: Highlights from the ESC Digital and AI Summit 2025 in Berlin, 2025. Meeting report; European Society of Cardiology Digital and AI Summit: Berlin.
  75. Cai, Y.; et al. MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2024. [Google Scholar] [CrossRef]
  76. Sun, H.; Lin, H.; Yan, R. Collaborative Synthesis of Patient Records through Multi-Visit Health State Inference. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2024. [Google Scholar] [CrossRef]
  77. Olaniyi, B.Y.; del Río, A.F.; Periáñez, África; Bellhouse, L. User Engagement in Mobile Health Applications. In Proceedings of the Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2022. [Google Scholar] [CrossRef]
  78. Mate, A.; et al. Field Study in Deploying Restless Multi-Armed Bandits: Assisting Non-profits in Improving Maternal and Child Health. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2022. [Google Scholar] [CrossRef]
  79. Verma, S.; et al. Increasing Impact of Mobile Health Programs: SAHELI for Maternal and Child Care. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2023. [Google Scholar] [CrossRef]
  80. Sabour, R.M.; Benezeth, Y. Gated Recurrent Unit-Based RNN for Remote Photoplethysmography Signal Segmentation. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2022. [Google Scholar] [CrossRef]
  81. Shao, H.; et al. Video-Based Multiphysiological Disentanglement and Remote Robust Estimation for Respiration. IEEE Trans. Neural Netw. Learn. Syst. 2024. [Google Scholar] [CrossRef]
  82. Gao, H.; Wu, X.; Geng, J.; Lv, Y. Remote Heart Rate Estimation by Signal Quality Attention Network. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2022. [Google Scholar] [CrossRef]
  83. Thompson, et al. Diversity in Cardiovascular Care: Advancing Health Equity, 2024. Editorial synthesis on workforce diversity and health equity in cardiovascular care.
  84. Grace, et al. Digital Health Technologies to Transform Women’s Health: Innovation and Inclusive Research, 2025. Narrative review on digital health technologies across the women’s health research lifecycle.
  85. Chouldechova, A. Fair Prediction with Disparate Impact: A Study of Bias in Recidivism Prediction Instruments. Big Data 2017, 5, 153–163. [Google Scholar] [CrossRef]
  86. Mehrabi, N.; Morstatter, F.; Saxena, N.; Lerman, K.; Galstyan, A. A Survey on Bias and Fairness in Machine Learning. ACM Comput. Surv. 2021, 54, 1–35. [Google Scholar] [CrossRef]
  87. Ghassemi, M.; Oakden-Rayner, L.; Beam, A.L. The False Hope of Current Approaches to Explainable Artificial Intelligence in Health Care. Lancet Digit. Health 2021, 3, e745–e750. [Google Scholar] [CrossRef]
  88. Ramachandranpillai, R.; Sikder, M.F.; Bergström, D.; Heintz, F. Bt-GAN: Generating Fair Synthetic Healthdata via Bias-transforming Generative Adversarial Networks. J. Artif. Intell. Res. 2024. [Google Scholar] [CrossRef]
  89. Hardt, M.; Price, E.; Srebro, N. Equality of Opportunity in Supervised Learning. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2016. [Google Scholar]
  90. Yin, C.; Liu, R.; Caterino, J.M.; Zhang, P. Deconfounding Actor-Critic Network with Policy Adaptation for Dynamic Treatment Regimes. In Proceedings of the Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2022. [Google Scholar] [CrossRef]
  91. Rahman, M.M.; Purushotham, S. Fair and Interpretable Models for Survival Analysis. In Proceedings of the Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2022. [Google Scholar] [CrossRef]
  92. Roller, R.; Hahn, M.; Madhavan Ravichandran, A.; Osmanodja, B.; Oetke, F.; et al. One Size Fits None: Rethinking Fairness in Medical AI. Proceedings of the Proceedings of the 6th Workshop on Gender Bias in Natural Language Processing (GeBNLP) 2025, ACL, 282–289. [Google Scholar] [CrossRef]
  93. McCradden, M.D.; Joshi, S.; Mazwi, M.; Anderson, J.A. Ethical Limitations of Algorithmic Fairness Solutions in Health Care Machine Learning. Lancet Digit. Health 2020, 2, e221–e223. [Google Scholar] [CrossRef]
  94. Wachter, S.; Mittelstadt, B.; Russell, C. Why Fairness Cannot Be Automated: Bridging the Gap Between EU Non-Discrimination Law and AI. Comput. Law Secur. Rev. 2021, 41, 105567. [Google Scholar] [CrossRef]
  95. Caruana, R.; Nori, H. Why Data Scientists Prefer Glassbox Machine Learning. In Proceedings of the Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2022. [Google Scholar] [CrossRef]
  96. Siddiqui, H.; Rattani, A.; Ricanek, K.; Hill, T.J. An Examination of Bias of Facial Analysis based BMI Prediction Models. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2022. [Google Scholar] [CrossRef]
  97. Rawat, R.; Moon, J.; McBride, H.; Alamuri, D.; Ghosh, R.; O’Brien, S.; Nirmal, D.; Zhu, K. DiversityMedQA: Assessing Demographic Biases in Medical Diagnosis Using Large Language Models. In Proceedings of the Proceedings of the Workshop on NLP for Positive Impact (NLP4PI), 2024; EMNLP. [Google Scholar]
  98. Rajkomar, A.; Hardt, M.; Howell, M.D.; Corrado, G.; Chin, M.H. Ensuring Fairness in Machine Learning to Advance Health Equity. Ann. Intern. Med. 2018, 169, 866–872. [Google Scholar] [CrossRef]
  99. Chen, I.Y.; Pierson, E.; Rose, S.; Joshi, S.; Ferryman, K.; Ghassemi, M. Ethical Machine Learning in Healthcare. Annu. Rev. Biomed. Data Sci. 2021, 4, 123–144. [Google Scholar] [CrossRef]
  100. Wiens, J.; Saria, S.; Sendak, M.; et al. Do No Harm: A Roadmap for Responsible Machine Learning for Health Care. Nat. Med. 2019, 25, 1337–1340. [Google Scholar] [CrossRef]
  101. Char, D.S.; Shah, N.H.; Magnus, D. Implementing Machine Learning in Health Care—Addressing Ethical Challenges. N. Engl. J. Med. 2018, 378, 981–983. [Google Scholar] [CrossRef]
  102. Vyas, D.A.; Eisenstein, L.G.; Jones, D.S. Hidden in Plain Sight—Reconsidering the Use of Race. Correction in Clinical Algorithms. New England Journal of Medicine 2020, 383, 874–882. https://doi.org/10.1056/NEJMms2004740.. [CrossRef]
  103. Pierson, E.; Cutler, D.M.; Leskovec, J.; Mullainathan, S.; Obermeyer, Z. An Algorithmic Approach to Reducing Unexplained Pain Disparities in Underserved Populations. Nat. Med. 2021, 27, 136–140. [Google Scholar] [CrossRef]
  104. Gebru, T.; Morgenstern, J.; Vecchione, B.; et al. Datasheets for Datasets. Proc. Commun. ACM 2021, Vol. 64, 86–92. [Google Scholar] [CrossRef]
  105. Mitchell, M.; Wu, S.; Zaldivar, A.; et al. Model Cards for Model Reporting. In Proceedings of the Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT*), 2019; pp. 220–229. [Google Scholar] [CrossRef]
  106. Nyariro, M.; et al. Integrating Equity, Diversity and Inclusion Throughout the Lifecycle of AI Within Healthcare: A Scoping Review Protocol, 2023. Prism.-ScR. Scoping Rev. Protoc. Equity Divers. Incl. Healthc. AI. [CrossRef]
Figure 1. The survey’s central thesis: the same AI and digital health capability can amplify or close cardiovascular inequity, and which edge prevails turns on how the system is built and deployed. Figure 3 examines the relevant lifecycle decisions. The two branches show representative outcomes and mechanisms.
Figure 1. The survey’s central thesis: the same AI and digital health capability can amplify or close cardiovascular inequity, and which edge prevails turns on how the system is built and deployed. Figure 3 examines the relevant lifecycle decisions. The two branches show representative outcomes and mechanisms.
Preprints 230627 g001
Figure 2. Corpus-construction flow (counts as in Table 1). Three documented search streams feed a single relevance screen. The detailed-review corpus is then supplemented with canonical literature to form the cited set. The full protocol is in Appendix A.
Figure 2. Corpus-construction flow (counts as in Table 1). Three documented search streams feed a single relevance screen. The detailed-review corpus is then supplemented with canonical literature to form the cited set. The full protocol is in Appendix A.
Preprints 230627 g002
Figure 3. Where bias against women enters the cardiovascular-AI lifecycle. Each stage carries a characteristic failure mode (italic), discussed in Section 3, Section 4, Section 5 and Section 6. A deployment feedback loop can re-encode earlier inequities into new training data. Bias should therefore be treated as a system-level property rather than an isolated model defect.
Figure 3. Where bias against women enters the cardiovascular-AI lifecycle. Each stage carries a characteristic failure mode (italic), discussed in Section 3, Section 4, Section 5 and Section 6. A deployment feedback loop can re-encode earlier inequities into new training data. Bias should therefore be treated as a system-level property rather than an isolated model defect.
Preprints 230627 g003
Figure 4. Representation and transparency gaps behind women’s cardiovascular AI, drawn from the surveyed studies: heart-failure (HF) AI datasets seldom report race, ethnicity, or SES and are rarely fully accessible [35]; cardiac machine-learning papers almost never audit sex [36]; and no surveyed cardiovascular AI risk model had been externally validated [64]. Each bar reports a percentage of its own labelled denominator (datasets, papers, or models); the bars are not mutually comparable counts. Overall, the evidence base required to develop equitable models remains incomplete or poorly documented.
Figure 4. Representation and transparency gaps behind women’s cardiovascular AI, drawn from the surveyed studies: heart-failure (HF) AI datasets seldom report race, ethnicity, or SES and are rarely fully accessible [35]; cardiac machine-learning papers almost never audit sex [36]; and no surveyed cardiovascular AI risk model had been externally validated [64]. Each bar reports a percentage of its own labelled denominator (datasets, papers, or models); the bars are not mutually comparable counts. Overall, the evidence base required to develop equitable models remains incomplete or poorly documented.
Preprints 230627 g004
Figure 5. The women’s cardiovascular care pathway with representative AI and digital health touchpoints (above the line) and the equity risks they introduce (below). Applications are surveyed in Section 4Section 5. Every stage offers both an AI opportunity and a matching equity risk.
Figure 5. The women’s cardiovascular care pathway with representative AI and digital health touchpoints (above the line) and the equity risks they introduce (below). Applications are surveyed in Section 4Section 5. Every stage offers both an AI opportunity and a matching equity risk.
Preprints 230627 g005
Table 1. Corpus inventory. Three documented search streams (Appendix A) feed one relevance screen. “Screened” counts the candidates examined, and “retained” counts the papers selected for detailed review and structured coding. Counts are approximate because some works appear in more than one stream. The screen yields a detailed-review corpus of about 80 papers. The survey cites 106 references in total, namely this corpus together with the canonical grounding literature (Section 2).
Table 1. Corpus inventory. Three documented search streams (Appendix A) feed one relevance screen. “Screened” counts the candidates examined, and “retained” counts the papers selected for detailed review and structured coding. Counts are approximate because some works appear in more than one stream. The screen yields a detailed-review corpus of about 80 papers. The survey cites 106 references in total, namely this corpus together with the canonical grounding literature (Section 2).
Stream Focus and sources Screened Retained
1. Clinical Women’s CVD, sex/gender differences, AI/digital health (cardiology and digital health venues) 27 25
2. Seminal Landmark/topical medical-AI bias (PMC, PNAS, Nature, arXiv, WHO) 13 11
3. Methods Fairness- and health-AI (ML/NLP), ∼20 thematic queries over a large index 186 45
Total Detailed-review corpus (de-duplicated) 226 ∼80
Table 2. Organizing taxonomy of AI and digital health for women’s cardiovascular health. Rows are the four threads of this survey; columns sketch where each thread engages the AI and digital health lifecycle. Representative references are illustrative, not exhaustive.
Table 2. Organizing taxonomy of AI and digital health for women’s cardiovascular health. Rows are the four threads of this survey; columns sketch where each thread engages the AI and digital health lifecycle. Representative references are illustrative, not exhaustive.
Thread Data & problem framing Modeling & evaluation Deployment & governance
Sources of inequity (§Section 3) Underrepresentation; cost/outcome proxies; missing sex labels [19,35] Underdiagnosis; sex-stratified error gaps [10,36] Feedback loops; access-driven data missingness [37,38]
AI applications (§Section 4) Sex-specific cohorts; EHR/imaging curation [28] ECG/echo models; LLMs; simulation [8,9,39] Prospective validation; workflow integration [7]
Digital health & FemTech (§Section 5) User needs; participatory design [40] Telehealth, mHealth, wearables [3,41] Digital divide; commercialization [14,15]
Fairness methods & governance (§Section 6Section 7) Fair/synthetic data; dataset transparency [17] Auditing, mitigation, causal/mechanistic methods [16,42,43] WHO principles; regulation; trial reform [4,44]
Table 3. Representative prior surveys/reviews and their coverage relative to this article. : substantial coverage; ½: partial; blank: out of scope. Columns: women’s CVD focus (W-CVD), fairness methods (Fair), LLMs, digital health/FemTech (DH).
Table 3. Representative prior surveys/reviews and their coverage relative to this article. : substantial coverage; ½: partial; blank: out of scope. Columns: women’s CVD focus (W-CVD), fairness methods (Fair), LLMs, digital health/FemTech (DH).
Prior survey / review (focus) W-CVD Fair LLMs DH
Medical-imaging fairness benchmarks [16,42]
Sex/gender bias in biomedical AI [18] ½
Open science vs. bias in health AI [17] ½
Bias in CVD prediction/detection [38] ½
Digital health for women’s CVD [3]
This survey
Table 4. Representative empirical evidence that clinical AI disadvantages women (and intersecting groups), by modality and by the lifecycle stage where the bias chiefly originates. “Type” marks whether the evidence is directly women-cardiovascular-specific (W-CVD) or transferred from the broader clinical-AI fairness literature. FNR: false-negative rate.
Table 4. Representative empirical evidence that clinical AI disadvantages women (and intersecting groups), by modality and by the lifecycle stage where the bias chiefly originates. “Type” marks whether the evidence is directly women-cardiovascular-specific (W-CVD) or transferred from the broader clinical-AI fairness literature. FNR: false-negative rate.
Modality Finding Origin (lifecycle) Type Ref.
Chest X-ray Underrepresented gender loses accuracy under dataset imbalance Data composition transfer [11]
Chest X-ray Underdiagnosis (“no finding”) concentrates in women, youth, minorities Label/threshold + data transfer [10]
Imaging (multi) Models predict self-reported race from pixels (hidden shortcut) Representation transfer [57]
Tabular (cardiac) Higher female FNR in 13/16 reproductions; sex rarely audited Evaluation W-CVD [36]
EHR risk models 0/486 models externally validated; all high risk of bias Modeling, reporting W-CVD [64]
LLM (generation) Overgeneration of male cardiac cases vs. real prevalence Training and decoding W-CVD [56]
LLM (summ.) Women’s needs downplayed (model-specific) Modeling transfer [65]
LLM (diagnosis) Stricter MI diagnosis in women; demographic diagnosis-bias Modeling W-CVD [66]
ICU allocation Models add inequality beyond that present in the data Deployment transfer [67]
Table 5. Fairness and bias-mitigation method families, organized by the lifecycle stage they target (cf. Figure 3), with representative approaches from the surveyed literature and a key caveat for each. The pattern across the column of caveats is the survey’s main methods finding: auditing is mature and should be routine, while every mitigation family carries an open failure mode.
Table 5. Fairness and bias-mitigation method families, organized by the lifecycle stage they target (cf. Figure 3), with representative approaches from the surveyed literature and a key caveat for each. The pattern across the column of caveats is the survey’s main methods finding: auditing is mature and should be routine, while every mitigation family carries an open failure mode.
Family Lifecycle stage Representative approaches Key caveat
Data-level Data / problem Re-balancing; fairness-aware synthetic data [17] Synthetic data can preserve correlation/representation bias
In-processing Modeling Adversarial removal; fairness-constrained PEFT; mixture-of-experts; MI regularizers [16] Rarely beats well-tuned ERM [16]
Post-processing Evaluation Group thresholds; equalized-coverage conformal sets Can “level down” performance
Causal Problem / eval Path-specific effects; fair survival analysis Needs a credible causal model
Mechanistic Modeling Activation patching; sparse-autoencoder steering [43] Immature; clinical effect unproven
Auditing Cross-cutting Subgroup error reporting; induced-inequality scores [36,67] In-distribution fairness need not transfer [42]
Table 6. Open challenges for AI and digital health in women’s cardiovascular care: why each matters and a concrete direction.
Table 6. Open challenges for AI and digital health in women’s cardiovascular care: why each matters and a concrete direction.
Challenge Why it matters Direction
Intersectionality Worst harms at race×sex×SES intersections [10] Borrow strength across subgroups; routine intersectional reporting
Sex-disaggregated data Missing/pooled sex drives measurement bias [35] Documented datasets; sex-stratified evaluation as a norm
Reproductive life course Pregnancy/menopause poorly covered [40] Model the female cardiovascular life course explicitly
Real-world fairness In-distribution fairness does not transfer [42] Prospective, multi-site fairness evaluation
FemTech evidence Commercial claims outpace evidence [15] Independent evaluation; matched regulation
Benchmarks No women’s-CVD-AI benchmark exists Consented, sex-balanced, intersectional tasks
Table 7. Evidence-strength map of the survey’s main conclusions. Basis: direct (women-CVD-specific studies), transferred (general clinical-AI fairness applied to women), or agenda (a position/direction). Strength is our qualitative reading of the current literature.
Table 7. Evidence-strength map of the survey’s main conclusions. Basis: direct (women-CVD-specific studies), transferred (general clinical-AI fairness applied to women), or agenda (a position/direction). Strength is our qualitative reading of the current literature.
Conclusion Basis Strength
Clinical AI underdiagnoses women and intersecting groups transferred + direct strong
Demographic data imbalance degrades subgroup performance transferred strong
HF/CVD datasets under-report sex, race, and SES direct strong
LLMs encode demographic bias in cardiac contexts direct + transferred moderate
Sex-specific modeling exposes distinct risk structure direct moderate
No mitigation reliably beats ERM; in-distribution fairness does not transfer transferred strong
Measurement/sensor bias propagates into decisions transferred (race) moderate
FemTech commercialization outpaces clinical evidence direct (critique) agenda
Sex-disaggregated data and women-CVD benchmarks are missing direct agenda
Equity must be engineered across the lifecycle, not patched on synthesis agenda
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.