Submitted:
13 August 2026
Posted:
14 August 2026
You are already at the latest version
Abstract
Purpose of Review: Heart failure affects over 64 million people worldwide, yet diagnosis remains challenging, with nearly two-thirds of cases estimated to remain undiagnosed overall. This review synthesises current evidence on multimodal artificial intelligence (AI) for heart failure detection, covering fusion strategies, explainability, and implementation challenges. Recent Findings: Unimodal AI achieves good-to-excellent discrimination from routine clinical measurements, 12-lead electrocardiograms, and chest X-rays, while echocardiography and cardiac magnetic resonance enable measurement automation and phenotyping. However, subgroup performance and external validation remain limited. Multimodal AI combining ECG, echocardiography, radiography, and electronic health records improves discrimination and phenotyping compared to single-modality models, with the largest gains observed from foundation models pretrained on large imaging and waveform corpora. Summary: Multimodal AI could capture the multidimensional complexity of heart failure, but prospective validation, calibration, and generalisability remain critical gaps. Multicentre trials with standardised endpoints, co-designed interfaces, and regulatory frameworks are priorities for deployment-centric algorithms.
Keywords:
multimodal artificial intelligence
; heart failure
; electrocardiography
; echocardiography
; radiography
; electronic health records
; explainable AI
Introduction
Heart failure is a complex clinical syndrome that represents the final common pathway of a diverse range of cardiac diseases. It affects more than 64 million people worldwide and continues to rise amid population ageing and a growing cardiovascular disease burden, representing an increasing strain on healthcare systems and patient quality of life [1,2]. Five-year mortality after diagnosis exceeds that of many other conditions, including many cancers, and recurrent hospitalisation generates substantial health economic costs, with combined direct and indirect expenditures estimated in hundreds of billions of dollars annually worldwide [3,4]. Moreover, each heart failure phenotype: heart failure with reduced ejection fraction (HFrEF, LVEF <40%), heart failure with mildly reduced ejection fraction (HFmrEF, LVEF 40–49%), and heart failure with preserved ejection fraction (HFpEF, LVEF ≥50%), reflects distinct underlying mechanisms with different implications for treatment [5,6,7,8,9]. HFpEF now accounts for more than half of all heart failure presentations. However, it remains the most diagnostically elusive phenotype, characterised by near-normal systolic function and a varied comorbidity profile (obesity, hypertension, atrial fibrillation, and chronic kidney disease), closely overlapping with non-cardiac causes of dyspnoea.
Accurate heart failure diagnosis requires the complex integration of information across multiple data flows. The European Society of Cardiology (ESC) and American Heart Association (AHA) guidelines recommend a structured, multi-faceted pathway for the diagnosis and assessment of heart failure [10,11]. Thus, clinical assessment requires examination of multiple data modalities, including clinical history and examination, 12-lead electrocardiogram (ECG), natriuretic peptide measurements (such as BNP or NT-proBNP), chest radiography (CXR), and echocardiography. Each modality contributes to a diagnostic signal yet carries inherent limitations. For instance, natriuretic peptide specificity is reduced in atrial fibrillation, chronic kidney disease, and advanced age, conditions that co-occur with heart failure in many affected patients [12]. Electrocardiographic signs of cardiomyopathy can manifest on the 12-lead ECG, but in many heart failure patients, these are subtle or absent [13]. Beyond individual modality constraints, these pathways demand coordinated specialist input across multiple assessments, creating diagnostic delays that commonly extend to months in community and primary care settings, with real-world consequences for timely initiation of guideline-directed therapy [14,15].
Artificial intelligence (AI) has emerged as a promising approach to improving diagnostic pathways, with its clinical application progressing from algorithmic decision-support toward autonomous diagnostic tools [16]. Early applications of supervised machine learning to structured electronic health record (EHR) data, incorporating demographics, comorbidities, laboratory biomarkers, and prescribing patterns, have established the feasibility of algorithmic risk stratification for heart failure and laid the groundwork for the deep learning approaches that followed [17,18]. One major milestone arrived with the application of deep convolutional neural networks (CNNs) to cardiac signals, demonstrated by Attia et al. 2019 in a study of 44,959 patients from a US academic centre [19]. In this study, a CNN applied to the standard 12-lead ECG could detect left ventricular systolic dysfunction (LVSD) with a receiver operating characteristic area under the curve (AUROC) of 0.93 and identify patients with a fourfold elevated risk of ventricular dysfunction before any clinical presentation. Subsequent AI-ECG models validated in cohorts exceeding 100,000 patients have demonstrated AUROC values of 0.85-0.93 for heart failure detection across independent populations and healthcare systems [20]. However, the performance ceiling of unimodal AI is becoming increasingly evident, with limitations such as collection bias, data missingness, or limited adaptability across health systems hampering clinical utility [21]. These factors have motivated the development of multimodal AI systems that integrate two or more data sources to produce a richer, more clinically faithful patient representation.
In this review, we synthesise the current evidence on the utility of multimodal AI algorithms for the diagnosis and interpretation of heart failure. We examine the fusion strategies and model architectures underpinning contemporary multimodal algorithms, with emphasis on diagnostic gain over unimodal approaches, heart failure phenotype discrimination, explainability, and external validation. We then address the challenges of clinical translation before outlining directions for prospective trials and clinical deployment.
Diagnostic Pathways for Heart Failure
Guideline-directed diagnosis of heart failure follows a structured pathway designed to confirm the syndrome and identify suitable treatment strategies. Recent guideline updates from the ESC and the AHA/ACC/HFSA both recommend an initial assessment that combines clinical history and examination with measurement of natriuretic peptides (BNP or NT-proBNP), a 12-lead ECG and transthoracic echocardiography, where natriuretic peptides are elevated, or clinical suspicion is high [5,9,10,11]. For the substantial proportion of patients presenting with dyspnoea where the left ventricular ejection fraction is normal, a stepwise framework integrating these modalities is required [22]. Table 1 summarises the clinical role and key characteristics for each diagnostic test, along with the opportunities for AI integration to improve the quality of care.
Each modality has well-characterised limitations that constrain diagnostic accuracy when used in isolation. Natriuretic peptides remain the most widely used first-line test for acute heart failure in the emergency department, with NT-proBNP demonstrating excellent diagnostic performance for acute heart failure [24]. An age-independent rule-out threshold of 300 pg/mL has a negative predictive value approaching 99%, whereas age-stratified rule-in thresholds of 450, 900, and 1800 pg/mL (for ages 75 years and older) have been reported to preserve specificity [41]. However, natriuretic peptide specificity is reduced in atrial fibrillation, chronic kidney disease, obesity, and advanced age, which frequently co-exist with heart failure [42]. The 12-lead ECG is rapid, inexpensive, and widely available, but lacks sensitivity for HFpEF, in which electrical abnormalities may be subtle or absent at presentation [43]. Echocardiography remains the cornerstone investigation for confirming structural or functional abnormality and assigning heart failure phenotype, yet it is operator-dependent and not always accessible at the point of first contact, particularly in primary care [44]. Thus, no single modality combines high testing availability, diagnostic accuracy, and phenotype discrimination, a gap that motivates the adoption of multimodal AI-enabled screening strategies.
Beyond the performance of individual tests, the diagnostic pathway as a whole faces substantial real-world bottlenecks, particularly in primary care and community settings where most patients first present with breathlessness. The HFA-ESC clinical consensus statement on heart failure diagnosis in the general community emphasises that diagnostic delays of several months are common, driven by restricted access to natriuretic peptide testing and echocardiography outside specialist clinics [14]. To support earlier recognition by non-specialist clinicians, the HFA of the ESC has recently proposed the FIND-HF acronym (Fatigue, Increased water accumulation, Natriuretic peptide testing, Dyspnoea) as a simple classification to consider a heart failure diagnosis and to initiate NT-proBNP testing in at-risk patients [24]. Pragmatic point-of-care strategies, such as the Handheld-BNP programme [45], in which general practitioners were trained to perform bedside BNP testing and portable echocardiography, have demonstrated the feasibility of bringing these data modalities closer to first contact. However, evidence to date highlights persistent gaps in availability, sensitivity, and specificity across the conventional diagnostic pathway, highlighting the need for more individualised and targeted approaches.
Unimodal AI in Heart Failure: Advances and Challenges
To date, AI models for diagnosing heart failure have typically been trained on a single data modality. Figure 1 illustrates how each data modality can be mapped onto the diagnostic pathway, from low-specificity screening in primary care and community settings to high-specificity advanced imaging in secondary and tertiary care. At each integration point, AI algorithms could be embedded to support triage, risk stratification, and referral decisions for subsequent testing. However, most published models are developed to target only a single diagnostic point along this pathway.
Structured EHR algorithms
Structured EHR data provide a longitudinal view of clinical history and have been widely adopted to predict adverse outcomes and heart failure phenotypes. McGilvray et al. [46] developed a deep learning model trained on structured EHR data to predict death or severe decompensation in patients with established heart failure, supporting its potential use for early identification of patients who may benefit from escalation of care. The Collaboration for the Diagnosis and Evaluation of Heart Failure (CoDE-HF) tool, developed by Lee et al. [18] demonstrated the importance of integration of single NT-proBNP measurements within EHR-driven algorithms using an extreme gradient-boosted machine learning model, externally validated across a pooled international cohort of patients presenting with suspected acute heart failure, providing a confident rule-in and rule-out classification for substantially more patients than NT-proBNP thresholds alone. A comparable approach has since been developed for BNP and MR-proANP [26], and has been shown to provide more accurate rule-in and rule-out classification than age-stratified biomarker thresholds [47]. In genomics, a recent model integrating large-scale genetic and clinical information improved the prediction of incident heart failure cases beyond clinical risk factors alone [48]. Machine learning-based phenotyping combining 107 demographic and clinical variables identified three distinct HFpEF phenogroups with differential treatment responses, including a reduced risk of heart failure rehospitalisation in one phenogroup and a reduced risk of all-cause mortality in another, with 96% consistency and maintained benefits after external validation [49].
AI-ECG algorithms
The 12-lead ECG has become the most intensively studied diagnostic test for unimodal AI in heart failure, owing to its low cost, universal availability, and ability to capture electrophysiological signatures of structural and functional cardiac abnormalities. A demonstration of this potential came from Hannun et al. [50], who trained a deep neural network on more than 91,000 single-lead ECGs to achieve cardiologist-level arrhythmia detection and classification on raw ECG waveforms. Building on this approach, Attia et al. [19] demonstrated that a convolutional neural network (CNN) applied to the 12-lead ECG could identify left ventricular systolic dysfunction (LVSD) with high discrimination. Sangha et al. [51] reported comparable discrimination for LVSD using a CNN trained on photographs of printed ECG traces, broadening applicability to settings without access to raw digital waveforms. Single-lead implementations adapted for wearable and portable devices have similarly demonstrated strong performance for LVSD detection [52]. More recently, AI-ECG models have been extended to predict HFpEF from the 12-lead ECG [53] with additional evidence supporting consistent external validation performance across independent cohorts [54]. Beyond systolic dysfunction, AI-ECG models have been extended to detect specific and less common aetiologies of heart failure. In a real-world prospective evaluation, an AI-ECG algorithm for hypertrophic cardiomyopathy (HCM) analysed more than 100,000 electrocardiograms, achieving a sensitivity of 95% with specificity and accuracy above 98% for identifying HCM in 5% of flagged patients who had no prior diagnosis [55]. AI-ECG has been applied comparably to the detection and monitoring of cardiac amyloidosis [56]. Similar approaches also extend the diagnostic uptake of clinical presentations, enabling broader screening of asymptomatic individuals [24,57], illustrating how at-risk people might be identified before any emerging symptoms develop. Collectively, these studies illustrate a maturing evidence base in which AI-ECG models have begun to move toward prospective, multicentre, and federated settings that demonstrate clinical utility.
Transthoracic echocardiography
Echocardiography is the primary imaging modality for heart failure diagnosis and phenotyping and has attracted substantial AI development efforts to automate and standardise interpretation. The novel EchoNet-Dynamic model, developed by Ouyang et al. [58], used a spatiotemporal CNN to perform beat-to-beat segmentation of the left ventricle across full echocardiographic videos, achieving expert-level estimation of left ventricular ejection fraction (LVEF). Building on this work, He et al. [59] conducted a blinded, randomised trial comparing sonographer-derived and AI-derived measures of cardiac function, providing among the first pieces of evidence that an AI model can match a cardiologist’s assessment in interpreting echocardiographic results. Deep learning applied to echocardiography has also enabled high-throughput precision phenotyping of left ventricular hypertrophy (LVH), a key structural indicator of HFpEF and hypertrophic cardiomyopathy (HCM), across large unselected clinical populations [60]. At the point of care, the PANES-heart failure study demonstrated that AI-enhanced echocardiography performed by briefly trained novice operators could screen for heart failure with an AUROC of 0.88, sensitivity of 84.6% and specificity of 91.4%, outperforming NT-proBNP-based screening by approximately 30% [61]. Taken together, these studies demonstrate that AI-driven echocardiography applications have progressed from automating expert measurements to enabling novice operators to perform screening in lower-resource settings, although questions remain about the adaptability of these algorithms across diverse patient populations and imaging platforms, as well as the practicality of real-world deployment.
Despite this breadth of evidence across modalities, several limitations remain: most models remain derived and validated in single-centre or single-country populations, external validation and prospective evaluation are reported for only a minority of studies, and formal calibration and subgroup performance analyses, by age, sex, and race or ethnicity, are inconsistently performed or not reported. Crucially, this evidence base remains weighted toward reported discriminative performance, with far less attention to whether unimodal models change clinical decisions or improve outcomes under prospective, real-world conditions. These limitations, together with the inherently partial view of heart failure pathophysiology that any single modality can provide, motivate the adoption of multimodal AI approaches, in which complementary data streams are combined within shared model architectures.
Cardiac imaging
Chest radiography (CXR), cardiac magnetic resonance (CMR), and coronary computed tomography angiography (CTCA) are key imaging modalities for heart failure diagnosis, providing tissue and vascular characterisation at varying levels of specificity. For CXR, deep learning frameworks have been developed to simultaneously identify and localise multiple thoracic abnormalities, including cardiomegaly [62], severe LVH and a dilated left ventricle [63] with consistent performance across sex, age, and ethnicity in external validation. CXR-derived deep learning features have also been used prognostically, analysing chest radiographs from patients with established predicted subsequent cardiac events, suggesting a role for AI-CXR algorithms beyond initial diagnosis and into risk stratification [64].
At the higher end of diagnostic specificity, Shad et al. developed a generalisable deep learning algorithm for cardiac MRI to interpret CMR sequences and pathologies relevant to heart failure and structural heart disease [65]. CTCA, increasingly used to exclude coronary artery disease as a cause of new-onset heart failure, has likewise become a surrogate for AI-derived risk stratification beyond visual stenosis assessment. For instance, Oikonomou et al. developed a machine learning-derived radiomic signature of perivascular adipose tissue, capturing coronary inflammation and fibrosis and improving prediction of major adverse cardiovascular events beyond existing risk factors [66]. This radiomic signature was subsequently incorporated into an AI-derived cardiac risk algorithm and externally validated in patients undergoing clinically indicated CTCA across the UK National Health Service [67,68]. While CXR-based models offer the advantage of usability in resource-limited and emergency settings, CMR and CTCA-based models offer comprehensive tissue characterisation. However, all three approaches remain oriented more broadly toward cardiovascular risk assessment, and prospective evaluation of their incremental value within heart failure-specific diagnostic pathways remains comparatively sparse.
Foundational Concepts of Multimodal AI
In this section, we introduce the foundational architectural concepts underlying multimodal AI for heart failure. We first describe the principal strategies for combining information from multiple data modalities before outlining the model architectures most commonly used for development.
Fusion Strategies
By definition, multimodal AI integrates information from two or more data modalities within a single predictive framework, and the primary architectural decision is the stage at which information from each modality is combined. Three fusion paradigms are recognised in the literature: early, intermediate, and late fusion [69,70]. Figure 2 illustrates how these strategies differ with a focused example on combining structured health record measurements with ECG and chest radiograph data.
Early fusion concatenates raw or minimally processed embeddings from each modality into a single input vector at the data level, before a downstream model (typically a deep neural network such as a residual neural net [71] or transformer [72]) is trained to learn joint representations directly from the combined data. This approach is conceptually simple but assumes that modalities can be meaningfully represented within a shared feature space and is sensitive to differences in dimensionality, scale, and missingness across modalities. Intermediate, or joint, fusion instead employs separate encoder networks for each modality to produce modality-specific feature representations, which are subsequently combined at the feature level (commonly via a cross-attention fusion layer) before a shared classification head produces the final prediction. Because each encoder can be tailored to its modality and fusion occurs after feature extraction, intermediate fusion is generally considered the most flexible strategy for task-specific adaptation [73]. Finally, late fusion trains independent models for each modality in isolation and then combines their outputs (typically class probabilities or risk scores) using an ensemble algorithm such as gradient-boosted trees or logistic regression. Because each unimodal model can be developed, validated, and updated independently, late fusion is often the most pragmatic strategy when modalities are collected at different sites or have other domain-specific constraints. A practical challenge across all fusion strategies is the handling of missing modalities, which are common when data are drawn from routine care rather than curated cohorts. Approaches include imputation, modality dropout during training, and architectures that generate predictions from modalities available at inference. Intermediate and late fusion is often more robust in this respect, as developed encoder modules can still contribute to the weight of a predicted risk score, even in the absence of other modalities.
Model Architectures
The fusion strategies described above are developed using a range of underlying model architectures, the choice of which is typically guided by the structure of the input data for each modality. For imaging modalities (echocardiography, CXR, and CMR), convolutional neural networks (CNNs) remain the dominant architecture, using stacked convolutional filters to learn hierarchical spatial patterns, progressing from low-level edges and textures in early layers to clinically meaningful structures such as heart chamber borders or pulmonary congestion in deeper layers. Residual architectures (ResNet [71]), which introduce skip connections to ease the training of very deep networks, are particularly prominent [74,75]. For sequential data such as 12-lead ECG waveforms and longitudinal EHR time series, recurrent architectures, including long short-term memory (LSTM) networks, gated recurrent units (GRUs), and one-dimensional CNNs, have historically been used to capture temporal dependencies [76,77,78].
Transformer architectures [72], which use self-attention mechanisms to dynamically weigh the relevance of different input features to a given prediction, have become increasingly prominent across all data modalities. They are particularly well-suited to intermediate fusion strategies that explicitly model interactions between different representations. For example, Yang et al. developed CaMPNet [79], a transformer-based architecture that fused raw 12-lead ECG waveforms, structured ECG-derived features, and demographic data via cross-attention for disease classification. Variational autoencoders (VAE) [80] are a class of generative models that learn to compress high-dimensional inputs into a lower-dimensional latent representation while preserving clinically relevant variation. These algorithms provide an alternative route to intermediate fusion. For example, Beetz et al. used a VAE to derive shared latent features from CMR and ECG data [81]. For structured EHR and tabular data, including biomarkers, demographics, and comorbidities, gradient-boosted ensemble methods such as XGBoost and CatBoost remain the architectures of choice, owing to their strong performance on tabular data and native handling of missing values [82,83,84]. Collectively, these architectures provide the building blocks from which multimodal architectures that combine these components achieve relative advantages over unimodal AI.
Advances of Multimodal AI in the Diagnosis of Heart Failure
In this section, we examine how multimodal AI approaches compare with their unimodal counterparts for heart failure diagnosis, summarised in Table 2, before considering how these models perform across patient subgroups and how well they are calibrated for clinical use. We then turn to the training paradigms, self-supervised learning, supervised transfer learning, and hybrid approaches, that underpin current multimodal architectures, contrasted in Table 3.
Benchmarks Against Unimodal AI
Multimodal model benchmarks continue to yield meaningful advantages to cardiac diagnostics. Desai et al. [85] provided a recent example of such a comparison, pooling baseline clinical and ECG data from three longitudinal cohorts (N>14,000) to assess whether a composite 12-lead ECG-AI model improved predictions for systolic and diastolic dysfunction. Participants with a positive composite AI-ECG screen had a 10- to 20-fold higher risk of incident heart failure than those with a negative screen, and the addition of the AI-ECG algorithm produced notable net reclassification improvements. This represents a broader pattern across the summarised studies, in which multimodal fusion most reliably improves on unimodal or clinical baselines when the added modality integrates new, meaningful clinical patterns.
Table 2 suggests that the strongest evidence for this improvement comes from imaging, waveform, text-based and tabular foundation models. EchoCLIP [86] and EchoPrime [87] combined large-scale contrastive pre-training on echocardiographic video with downstream fine-tuning, achieving expert-level performance across diagnostic tasks, including estimating ejection fraction, detecting valvular disease, and identifying structural heart failure phenotypes. Soto et al. [88] similarly demonstrated that fusing echocardiographic measurements with ECG-derived features improved detection of left ventricular hypertrophy beyond unimodal benchmarks, while Kolk et al. [89] showed that combining ECG waveform data with structured clinical variables in the DEEP RISK model improved prediction of adverse outcomes after heart failure hospitalisation. Other approaches integrated echocardiography with electronic health record variables for cardiopulmonary exercises [90], chest radiography with electronic health record data for acute presentations [91], ECG with blood biomarkers [92], and ECG with heart rate variability metrics [93], each reporting incremental performance gains. Meanwhile, Oikonomou et al. [94] developed a flexible ECG-Echo foundation model to discriminate 26 phenotypes of structural heart disease, showcasing the broader utilities of AI-guided screening.
Table 2.
Stand-out studies for the application of multimodal artificial intelligence to the diagnosis and interpretation of heart failure.
Table 2.
Stand-out studies for the application of multimodal artificial intelligence to the diagnosis and interpretation of heart failure.
| Study | Modalities |
Fusion level & strategy |
Architecture | Population |
Diagnostic endpoint |
Key results | Benefits over prior work |
|---|---|---|---|---|---|---|---|
|
Christensen et al. 2024 (EchoCLIP) [86] |
Echo (video) + Cardiology report text |
Late: Contrastive image-text (CLIP-based) |
EchoCLIP: ViT encoder + text transformer; Contrastive self-supervised |
Adults; echo archive; N=1,032,975 videos; multi-centre; USA |
LVEF estimation; device identification; clinical transitions |
LVEF MAE 7.1% (external validation); Device AUC 0.84–0.97; Transplant AUC 0.79; Patient reidentification AUC 0.86 |
Contrastive echo–text models surpass task-specific supervised models in diagnostic accuracy |
|
Vukadinovic et al. 2025 (EchoPrime) [87] |
Echo multi-view (video) + Report text |
Late: View-primed anatomic attention; Retrieval- augmented generation |
EchoPrime: Contrastive video-language model; Contrastive self-supervised + supervised fine-tuning |
Adults; 5 international health systems; N=12M+ video- report pairs; USA / Taiwan |
23 cardiac benchmarks (LVEF; diastolic dysfunction; structural heart failure aetiologies incl. HCM) |
SOTA on all 23 benchmarks across 5 international systems; |
Surpasses task- specific unimodal algorithms on all 23 benchmarks |
|
Soto et al. 2022 (LVH-fusion) [88] |
12-lead ECG + Echo (video) |
Intermediate: Simultaneous joint modelling of ECG + echo video |
Joint CNN (ECG + echo); SHAP + saliency maps; Supervised |
N ≈18,000+; USA (multi-centre) |
HCM; occult hypertension; |
F1 0.71 (HCM); F1 0.96 (occult hypertension); |
Outperforms human readers and unimodal algorithms on both tasks |
|
Huang et al. 2026 [90] |
Echocardiography (multi-view video) + EHR (demographics, labs, medications) incl. CPET measurements | Intermediate (multi-instance); CNN (echo) + MLP (EHR) joint model | Multi-instance CNN + MLP; attention weights for view importance | N=1000; multi-centre; USA (Weill Cornell + Columbia + New York Hospitals) | Peak VO2 classification (high-risk advanced heart failure; transplant/LVAD need stratification) | AUROC 0.85 (internal); AUROC 0.87 (external); R2 0.603 (internal) and 0.541 (external) in peak VO2 classification | Surpassing prior benchmarks in peak VO2 classification; Supports decisions for advanced heart failure therapies |
|
Kolk et al. 2024 (DEEP RISK) [89] |
CMR (short-axis) + 12-lead ECG + Clinical data |
Intermediate: Residual variational autoencoder extracts CMR + ECG features, used as inputs into machine learning fusion module with routine clinical data |
Residual VAE+XGBoost; attention + saliency maps; Supervised |
N=289; Netherlands (2 tertiary hospitals) |
Malignant ventricular arrhythmia onset |
AUROC 0.84 (0.71–0.96); Sensitivity 0.98 (0.75–1.00); Specificity 0.73 (0.58–0.97) |
Outperforms unimodal benchmarks |
|
Yang et al. 2026 [91] |
Echo (4 views) + EHR |
Late: Per-view CNN feature vectors concatenated with tabular classifier for clinical measurements |
Inception-v3 (echo) + XGBoost (joint); Supervised; Explainability with Grad-CAM + SHAP |
Adults; N=26,936; Single-centre; China |
Structural heart disease | AUC 0.81 (±0.01); Sensitivity 84.6%; Specificity 72.4%; NPV 98.8% |
Outperforms unimodal benchmarks |
|
Lee CK et al. 2024 [95] |
CXR (image features; lung-heart feature mask) + EHR (vital signs at triage) |
Late: 4 CXR sub- models + EHR vital signs model |
ResNet for CXR with lung-heart mask features + XGBoost classifier for tabular data; Supervised |
ED patients; N=1,432; USA (MIMIC-IV + MIMIC-CXR) |
NT-proBNP surrogate for acute heart failure |
AUROC 0.89 | Outperforms unimodal benchmarks |
|
Botros et al. 2025 [92] |
ECG + Blood tests |
Late: CNN (ECG features) + XGBoost (blood test features) |
Supervised CNN and XGBoost pipeline; Explainability with LIME + SHAP |
N=1,250; heart failure patients and controls; clinical ECG + blood test cohort; USA (MIMIC-IV) |
Classification of coded diagnosis of heart failure | AUROC 0.96; Clinically meaningful LIME/SHAP explanations |
Outperforms unimodal benchmarks |
|
González et al. 2024 [93] |
12-lead ECG (30 s) + Long-term heart rate variability (beat- to-beat samples) |
Intermediate: Raw ECG via ResNet + approximate long-term heart rate variability; TFM-ResNet for temporal dynamics |
XGBoost model with Accelerated Failure Time + ResNet + Transformer-ResNet; Supervised |
Apple Watch wearable cohort; Taiwan (Linkou) |
Heart failure hospitalisation risk prediction (survival modelling) |
AUROC 0.85 (TFM-ResNet; best of 14 models); Apple Watch ECG validation |
Stronger generalisation compared to established tabular and signal models |
|
Oikonomou et al. 2026 (TARGET-AI) [94] |
ECG-echo pairs combined with + EHR foundation model |
Sequential: AI-ECG integrated into EHR clinical decision workflow only if EHR predicts a favourable diagnostic performance |
CNN (AI-ECG) + EHR work- flow integration; Targeted deployment; Supervised |
Yale New Haven Health System (YNHHS); LVSD screening cohort |
LVSD, aortic stenosis, systolic pressure prediction; Actionable echo referral rate; new heart failure diagnosis |
AUROC 0.90 for LVSD, 0.85 for aortic stenosis, 0.82 for right ventricular systolic pressure; improvements in targeted screening in external validation | Increased actionable echo referrals; Targeted deployment strategy; |
|
Woolley et al. 2021 [96] |
Clinical variables + Biomarkers + Echo parameters |
Multivariate feature integration |
K-means + K-medoids + hierarchical clustering; Unsupervised learning |
HFpEF patients; N=2,161; USA |
HFpEF phenogroup identification (4 clusters; differential outcomes and treatment response) |
4 phenotype groups: young/obese/ metabolic; older/hyper- tensive; advanced cardiorenal; COPD; Differential response in biomarker profiles; |
Identification of mutually-exclusive subgroups of patients with HFpEF |
|
Sanchez- Martinez et al. 2018 [97] |
Echo (LV strain; myocardial velocity) + Clinical variables |
Multivariate feature integration |
K-means + LDA clustering; Unsupervised |
HFpEF patients; N=156; multi-centre across Wales, Italy and Norway | HFpEF subtype identification and outcome prediction |
72.6% correlation with HFpEF diagnosis; blinded reinterpretation of imaging revealed important abnormalities not included as clinical diagnostic criteria | Improved understanding of HFpEF pathways and definition of diagnostic criteria |
Abbreviations: CPET, cardiopulmonary exercise testing; CXR, chest radiograph; ECG, electrocardiogram; EHR, electronic health record; HCM, hypertrophic cardiomyopathy; HFpEF, heart failure with preserved ejection fraction; LV, left ventricular; LVEF, left ventricular ejection fraction; MAE, mean absolute error; NPV, negative predictive value; NT-proBNP, N-terminal pro-B-type natriuretic peptide; SHAP, Shapley additive explanations; VO2, oxygen consumption; XGBoost, extreme gradient boosting classifier.
Collectively, these findings support a shift toward multimodal architectures as the standard for diagnostic AI algorithms for heart failure, particularly where modalities capture distinct pathophysiological information such as electrical activity, structural anatomy, and longitudinal clinical trajectory. However, it is worth noting that comparisons are limited by considerable heterogeneity in diagnostic endpoints, model targets, and validation strategies, and few studies report external validation in populations distinct from their development cohort. It remains unclear whether these advantages consistently persist across the full range of patient subgroups encountered in routine clinical practice.
Subgroup Performance, Calibration and Bias
Heart failure outcomes are not distributed equally across patient populations. Despite the increasing adoption of multimodal AI, very few studies report subgroup-stratified discrimination or formal calibration metrics. This gap is particularly consequential given emerging evidence that AI models trained on ECG data can exhibit substantial demographic bias. For example, Kaur et al. [113] evaluated a CNN model trained to predict incident heart failure within five years from 300,000 12-lead ECGs at a single US academic centre, finding that model discrimination declined markedly with age and was considerably worse in Black patients than in other racial groups of the same age [98]. Critically, these disparities were not resolved by incorporating demographic variables into the model architecture, training separate race-specific models, or rebalancing the training set for equal racial representation. Li et al. [99] applied machine learning classifiers to predict prolonged length of stay and in-hospital mortality in over 200,000 patients, and found that the best-performing model selectively under-identified adverse outcomes in female, Black, and socioeconomically disadvantaged patients. Integrating social determinants of health into the feature space improved fairness across these subgroups without compromising overall predictive performance.
By contrast, foundation models illustrate that broad multi-institutional training data can support more consistent cross-site performance, although whether this translates into demographic fairness has not been extensively tested. The choice of training paradigm may itself influence these properties.
Self-Supervised and Supervised Learning
Apart from the model architecture, the choice of training strategy has a direct bearing on how well a multimodal AI model generalises across institutions and patient subgroups, and on how much labelled data is required to reach clinically useful performance. Supervised transfer learning (STL) is used to pretrain a model on a large labelled dataset from a related domain [100,101] and then fine-tune the resulting weights on a smaller, task-specific heart failure dataset. This approach is effective to the extent that the source and target domains share structural features [102]. Self-supervised learning (SSL), by contrast, first pretrains a model on large volumes of unlabelled data using pretext tasks such as contrastive learning, in which the model learns to distinguish matched from mismatched input samples without requiring expert-annotated labels [103]. Weimann and Conrad [104] demonstrated this principle for 12-lead ECG analysis, showing that pretraining convolutional neural networks on a large unlabelled ECG corpus before fine-tuning on a small labelled dataset substantially reduced the number of annotations required to reach a given level of classification performance compared with training from scratch.
This principle underpins several recent foundation models in the domain of heart failure, such as EchoCLIP [86] and EchoPrime [87], which adapts the CLIP (Contrastive Language-Image Pre-training) framework [105] to learn joint representations of echocardiographic videos and their accompanying text reports. Table 3 summarises the key comparisons between SSL, STL, and hybrid approaches that combine self-supervised pre-training with task-specific supervised fine-tuning. In particular, hybrid approaches that combine domain-specific self-supervised pre-training with supervised fine-tuning currently offer the strongest combination of label efficiency and cross-site generalisation, although the evidence base remains concentrated in ECG and chest radiography applications relative to multimodal cardiac imaging [106,107,108].
Table 3.
Comparison of learning strategies for multimodal AI in heart failure diagnosis: self-supervised learning (SSL), supervised transfer learning (STL), hybrid SSL with supervised fine-tuning, and large-scale foundation models contrasted by pre-training clinical data.
Table 3.
Comparison of learning strategies for multimodal AI in heart failure diagnosis: self-supervised learning (SSL), supervised transfer learning (STL), hybrid SSL with supervised fine-tuning, and large-scale foundation models contrasted by pre-training clinical data.
|
Learning strategy |
Pre-training requirements |
Label efficiency |
Generalisation | Diagnostic application |
|---|---|---|---|---|
|
Self-supervised learning (SSL) |
Large unlabelled domain-specific datasets; no manual annotations required for the pretraining phase. 100K-10M samples. Examples: Echocardiogram archives [60,109], public CXR repositories [110]. |
High: competitive performance at 1-10% label availability for LVH, LVSD and aortic stenosis. SSL outperforms supervised baselines in low-prevalence events [103]. The advantage narrows as the amount of labelled data increases. |
Strong zero-shot and few-shot learning for novel tasks. Contrastive objectives (SimCLR [105], BYOL [111], CLIP [112]) yield robust representations that generalise across demographic subgroups and acquisition settings. Performance advantage over supervised approaches is greatest when labels are scarce. |
Risk stratification for heart failure, adapting large unlabelled population-level data; Foundation models achieve state-of-the-art cross-cohort generalization [113,114]. LVH and severe AS detection from echocardiography [115]. CXR few-shot detection of cardiomegaly and pleural effusion [116]. |
|
Supervised transfer learning (STL) |
Large, labelled domain-specific datasets; high-quality annotations required at scale in the source domain. Examples: CheXpert [100], INSPECT-EHR [101]. Layer-freezing and fine-tuning strategy for adapting the source-domain weights to the target task. |
Moderate: requires sufficient labelled fine-tuning data; performance degrades markedly with <1% target labels [117]. Provides minimal benefit when models are trained from scratch, and when medical imaging datasets are of modest size. Early convolutional layers are transferable; later task-specific layers require substantial fine-tuning data [102]. | Strong when the source and target domains share feature and modality structure. Cross-site generalisation can be prominent in federated sites [118]. Risk of negative transfer (source domain data undesirably affects performance in the target domain [117]). | Multimodal models for heart failure prediction on representative labelled datasets [118,119]. Incorporating multi-site federated learning. |
|
Hybrid SSL with supervised fine-tuning* |
Two-stage pipeline: Stage 1 involves SSL on a large unlabelled domain corpus; Stage 2 involves supervised fine-tuning on a smaller labelled heart failure task dataset. Pre-training data scale: 1M-10M unlabelled samples [110,113]; finetuning data: representative labelled events. Recommended pipeline when large unlabelled clinical archives exist alongside task-specific cohorts suitable for domain adaptation. |
Highest: domain-specific SSL representations require fewer fine-tuning labels to reach the performance ceiling [120]. Models converge faster and with fewer labels than training from scratch. | Strong cross-site generalisation [120]. Reduced the impact of model generalisation issues, such as catastrophic forgetting, when model weights are partially frozen. |
Foundation models pre-trained on >1M events across multiple modalities and fine-tuned across multiple cardiac conditions [106,107,108]. Superior label efficiency and generalisability across cardiac prediction tasks. Superior performance in external validation. |
*The hybrid SSL approach is the current recommended paradigm for diagnostic prediction tasks where large unlabelled clinical datasets are available alongside limited labelled annotations. It consistently outperforms both pure SSL and supervised transfer learning when a representative labelled fine-tuning set is available for task adaptation.
Label efficiency is defined as the ability to achieve competitive performance with a smaller proportion of labelled training data than fully supervised models trained from scratch. High label efficiency is critical for rare heart failure phenotypes (e.g., hypertrophic cardiomyopathy, cardiac amyloidosis).
Generalisation refers to cross-site, cross-institution, and cross-demographic performance. External validation on geographically or institutionally distinct cohorts is the minimum acceptable evidence for clinical translation.
Abbreviations: AS, aortic stenosis; CXR, chest radiograph; ECG, electrocardiogram; EHR, electronic health record; heart failure, heart failure; LVH, left ventricular hypertrophy; LVSD, left ventricular systolic dysfunction; SSL, self-supervised learning; STL, supervised transfer learning.
Explainability of Multimodal Algorithms for the Interpretation of Heart Failure
In this section, we examine the explainability methods applied to multimodal AI models for heart failure, with particular attention to whether these techniques yield insights useful for clinical assessment. This broadly includes approaches that highlight which modality, region, or feature drove a given prediction, before turning to feature attribution and counterfactual methods. We then evaluate the gap between explainable AI (XAI) outputs and the requirements of clinical decision-making.
Attention mechanisms, which were popularised by the transformer architecture [72], allow a model to dynamically weight which segments of the input (e.g., ECG intervals, image regions, or clinical variables) contributed most to a prediction. These approaches have become a common explainability layer in multimodal models. For instance, in the DEEP RISK framework for predicting malignant ventricular arrhythmia in non-ischaemic cardiomyopathy, modality-contribution analysis showed that CMR with late gadolinium enhancement carried greater weight than ECG or clinical inputs in the final risk estimate [89]. For imaging modalities, gradient-weighted class activation mapping (Grad-CAM [121]) projects gradients from the final convolutional layer back onto the input image to generate a heatmap of the regions most influential to the prediction, and has been widely used to highlight relevant cardiac structures on echocardiography and chest radiography [122,123]. Saliency methods applied to 12-lead ECG localise waveform segments, such as the QRS complex or ST-T morphology, providing clinicians with a visual anchor for prediction.
Feature-level attribution methods using Shapley Additive exPlanations (SHAP [124]) apply an alternative game-theoretic approach to assign each input feature a contribution value reflecting its marginal effect on the model output. Applied to structured EHR data, SHAP analysis of tree-based models identifies natriuretic peptide levels, renal function, and comorbidity burden among the features most predictive of heart failure-related outcomes, consistent with clinical reasoning [125,126]. In multimodal settings, although promising algorithms, such as MM-SHAP [127] have emerged to estimate modality contributions in vision-language models, but these have not yet been adopted in algorithms for heart failure screening. In general, SHAP methods can serve as a suitable post hoc approach for late-fusion frameworks that combine fundamentally different architectures, such as CNNs, Transformers, and XGBoost models, to confirm that predictions are driven by clinically plausible biomarkers rather than spurious correlations. Beyond attribution-based methods, counterfactual explanations [128], which estimate how a prediction would change under a hypothetical alteration of the input, represent an emerging direction with potential relevance to digital twin applications, although their clinical validation in heart failure remains in its early stages.
Figure 3 illustrates how explainability methods could be integrated into a patient-level dashboard, showing a hypothetical scenario in which a multimodal model predicts HFpEF with 87% confidence. The prediction is decomposed into modality-specific explanations, comprising an ECG saliency map derived from integrated gradients [129], an echocardiographic heatmap of high-activation regions, a Grad-CAM overlay on chest radiography, and a SHAP waterfall plot in which NT-proBNP and the E/e′ ratio emerge as the dominant contributors.
Despite their growing sophistication, these methods carry some important caveats. Attention weights and saliency maps indicate where a model focuses, but this localisation does not necessarily correspond to a causal or mechanistically meaningful explanation, and attribution maps can be unstable under small input perturbations or retraining [130,131]. More broadly, it has been argued that current XAI techniques may offer a false sense of transparency for high-stakes clinical decisions, and that rigorous external validation may be a more direct route to clinical trust [132]. For high-stakes domains such as the development of diagnostic algorithms for heart failure, the consensus is that using inherently interpretable models should be prioritised over the post-hoc explanation of black-box architectures altogether [133]. These considerations underscore the importance of co-designing explainability outputs with cardiology teams, patients, and regulators, so that the methods reported in the literature converge with explanations that are actionable for patient care.
Figure 3.
Mock-up of patient-level dashboard integrating multimodal AI and explainability methods for predicting the diagnosis of heart failure with preserved ejection fraction (HFpEF). In this fictional scenario, a multimodal AI model integrates four data streams to predict HFpEF with 87% confidence. Each panel shows the modality-specific explanation: (1) ECG signal importance via integrated gradients; (2) echocardiography summary of mapped high-activation regions according to pre-defined clinical criteria; (3) chest X-ray GradCAM saliency map focused on regional radiographic findings; (4) SHAP waterfall plot for structured tabular features, showing NT-proBNP and E/e′ ratio as the dominant positive contributors. Overall modality contribution weights affect the final confidence of the model. All data shown are synthetic and illustrative only; they do not correspond to any real patient. Such a dashboard is intended for use in secondary or specialist care, where the multimodal data streams shown here are routinely available. Abbreviations: E/e′, early mitral inflow to annular velocity ratio; eGFR, estimated glomerular filtration rate; GLS, global longitudinal strain; GradCAM, gradient-weighted class activation mapping; LAVi, left atrial volume index; LVEF, left ventricular ejection fraction; SHAP, SHapley Additive exPlanations; TR, tricuspid regurgitation. Created in BioRender. Georgiev, K. (2025).
Figure 3.
Mock-up of patient-level dashboard integrating multimodal AI and explainability methods for predicting the diagnosis of heart failure with preserved ejection fraction (HFpEF). In this fictional scenario, a multimodal AI model integrates four data streams to predict HFpEF with 87% confidence. Each panel shows the modality-specific explanation: (1) ECG signal importance via integrated gradients; (2) echocardiography summary of mapped high-activation regions according to pre-defined clinical criteria; (3) chest X-ray GradCAM saliency map focused on regional radiographic findings; (4) SHAP waterfall plot for structured tabular features, showing NT-proBNP and E/e′ ratio as the dominant positive contributors. Overall modality contribution weights affect the final confidence of the model. All data shown are synthetic and illustrative only; they do not correspond to any real patient. Such a dashboard is intended for use in secondary or specialist care, where the multimodal data streams shown here are routinely available. Abbreviations: E/e′, early mitral inflow to annular velocity ratio; eGFR, estimated glomerular filtration rate; GLS, global longitudinal strain; GradCAM, gradient-weighted class activation mapping; LAVi, left atrial volume index; LVEF, left ventricular ejection fraction; SHAP, SHapley Additive exPlanations; TR, tricuspid regurgitation. Created in BioRender. Georgiev, K. (2025).

Future Directions for Trials and Decision-Support Tools
The evidence showcased in this review demonstrates that multimodal approaches for the diagnosis and assessment of heart failure, combining electrocardiographic, imaging, signal, and structured electronic health record data, meaningfully improve discrimination and phenotype classification beyond the capacity of single-modality models. Novel foundation models pretrained on large imaging and waveform corpora further extend these gains. Explainability methods such as attention-based modality attribution and SHAP-based feature attribution are increasingly adopted to accompany these models, although the relationship between attribution maps and causal relationships with heart failure remains unclear. However, the evidence supporting these approaches remains predominantly retrospective and derived from single-centre cohorts, with external validation, subgroup-stratified performance, and formal calibration reported inconsistently across studies. Translating these findings into clinical benefit will require sustained attention to three areas in particular: the generation of prospective trial evidence, the integration of multimodal models into decision-support workflows, and the regulatory and governance frameworks that enable both.
Some prominent exceptions to retrospective analysis included the blinded randomised trial conducted by He et al. [59], in which an AI model matched cardiologist-level interpretation of echocardiographic function, and the multi-centre external validation of the EchoNext ECG-based screening model by Poterucha et al. [109] to detect structural heart disease across three health systems. Foundation models, such as TARGET-AI of Oikonomou et al. [94] showcased a novel utility of multimodal approaches for stepwise screening for risk of structural heart disease, demonstrating a clear pathway to deployment. These studies illustrate the kind of real-world evidence that researchers evaluating multimodal heart failure algorithms must now focus on generating. Extending this paradigm to the next generation of multimodal architectures that combine 12-lead ECG, echocardiography, CMR, biomarkers, genotyping, and structured EHR data could be used to interrogate the optimal setup for embedding AI outputs as actionable triggers within clinician workflows. This distinction between discrimination and clinical impact is critical as even a well-calibrated algorithm may fail to change outcomes if it does not alter clinician behaviour. For instance, the MARS-ED trial [134] showed that a well-calibrated risk-stratification algorithm had no measurable effect on patient outcomes, while the TRICORDER implementation trial [135] similarly demonstrated that strong diagnostic performance did not translate into increased heart failure detection at scale, largely due to inconsistent uptake in routine primary care. Prospective evaluation of multimodal algorithms should therefore report not only discriminative performance but also their influence on clinical decisions and outcomes.
A clear priority in this field is the design of prospective trials to evaluate these data fusion strategies rather than single-modality components in isolation. Pragmatic trial designs embedded within existing care pathways [59] offer one template for evaluating the level and quality of diagnostic support of multimodal AI without disrupting clinical workflow. Such trials should specify standardised diagnostic endpoints, with predefined subgroup analyses to address demographic disparities [98,99], and report calibration. Recruitment across multiple health systems and geographies will be necessary to generate the external validation evidence that multimodal architectures need to justify model resilience across health services, populations and treatment pathways.
In parallel, translating multimodal models into decision-support tools that operate within routine clinical workflows is a key challenge that requires a different mindset than using single-modality algorithms. To ensure interoperability, we must now pose several questions: (i) which data modality is worth integrating in this care setting?; (ii) is a stepwise integration of modalities at different stages of care preferable to a stacked approach embedding all modalities?; (iii) how does the interplay between data modalities affect patient decisions?; and (iv) which fusion strategies are optimal for predicting patient outcomes? Explainability dashboards, such as the one illustrated in Figure 3, should be evaluated prospectively as part of clinical workflow integration to determine whether these tools improve clinician trust and decision-making, rather than simply making model outputs more interpretable. Federated learning approaches [118], may allow such tools to be trained and updated across health systems without centralising patient data, although continuous-learning systems of this kind will require adaptive regulatory pathways analogous to those proposed for AI-based software as a medical device [136].
Conclusions
Our review synthesises contemporary evidence that multimodal AI, combining electrocardiographic, imaging, signal and structured clinical data within multimodal fusion architectures consistently improves discrimination and phenotype classification for heart failure beyond unimodal approaches. The strongest gains have been observed for foundation models pretrained across large imaging and waveform datasets using novel mechanisms with elements of self-supervised learning. Explainability methods, including attention-based modality attribution and SHAP-based feature attribution, increasingly accompany these models and offer a potential route toward clinically interpretable outputs with some existing caveats regarding attribution instability and the gap between feature localisation and causal understanding.
Despite these advances, some critical gaps persist: most evidence derives from single-centre, retrospective cohorts and subgroup-stratified performance and formal calibration are inconsistently reported or not reported at all. There is no conclusive evidence that sociodemographic biases identified in unimodal AI models have been resolved in multimodal or foundation model architectures. Addressing these gaps should be the near-term priority, through multicentre prospective trials with standardised diagnostic endpoints, development and validation of service-agnostic foundation models and definition of clearer regulatory pathways for multimodal decision-support tools. For instance, federated learning strategies may offer a route to multimodal heart failure algorithms that generalise across health systems while preserving patient privacy. Ultimately, realising the promise of multimodal AI for the diagnosis and assessment of heart failure will depend on sustained collaboration between data scientists, cardiologists, patients, and regulators, so that these tools are validated and embedded within clinical pathways in ways that can demonstrate improved treatment planning and patient outcomes.
Key References
Attia ZI, Kapa S, Lopez-Jimenez F, McKie PM, Ladewig DJ, Satam G, et al. Screening for cardiac contractile dysfunction using an artificial intelligence–enabled electrocardiogram. Nat Med. 2019;25:70–4.
This landmark study demonstrated that a convolutional neural network applied to the standard 12-lead ECG could detect left ventricular systolic dysfunction with an AUROC of 0.93, establishing an evidence base for AI-ECG screening in heart failure.
Ouyang D, He B, Ghorbani A, Yuan N, Ebinger J, Langlotz CP, et al. Video-based AI for beat-to-beat assessment of cardiac function. Nature. 2020;580:252–6.
This study introduced a spatiotemporal convolutional neural network that performs beat-to-beat segmentation of the left ventricle in echocardiographic videos to achieve expert-level ejection fraction estimation, thereby establishing automated echocardiographic interpretation as a benchmark for subsequent multimodal approaches.
Acosta JN, Falcone GJ, Rajpurkar P, Topol EJ. Multimodal biomedical AI. Nat Med. 2022;28:1773–84.
This review articulated the conceptual framework for multimodal biomedical AI, including fusion strategies and model architectures, which underpins the synthesis of multimodal AI approaches for heart failure presented in this review.
He B, Kwan AC, Cho JH, Yuan N, Pollick C, Shiota T, et al. Blinded, randomized trial of sonographer versus AI cardiac function assessment. Nature. 2023;616:520–4.
Blinded randomised trial that provided among the first prospective evidence that an AI model can match cardiologist-level interpretation of echocardiographic function, offering a template for the prospective trial designs identified as a priority for multimodal heart failure algorithms.
Poterucha TJ, Jing L, Ricart RP, Adjei-Mosi M, Finer J, Hartzel D, et al. Detecting structural heart disease from electrocardiograms using AI. Nature. 2025;644:221–30.
Multicentre external validation of an ECG-based screening model across three independent health systems illustrates evidence that multimodal heart failure algorithms have a strong case for supporting clinical deployment.
Supplementary Materials
The following supporting information can be downloaded at the website of this paper posted on Preprints.org.
Author Contributions
K.G., K.K.L. and D.D. conceived the design and scope of the review. K.G. performed the literature extraction, curation and critical evaluation of the literature. K.G. drafted the manuscript. K.K.L. A.J.F.T, A.A. and N.L.M. provided clinical domain expertise to refine the review narrative. D.D. and O.P. provided support in the examination of the eligible methodologies. All authors revised the manuscript critically for important intellectual content. All authors are accountable for the work.
Funding
K.G. and D.D. are supported by the British Heart Foundation through a Project Grant (PG/24/12136). K.K.L. is supported by a British Heart Foundation Intermediate Clinical Research Fellowship (FS/ICRF/25/26134). N.L.M. is supported by British Heart Foundation Chair (CH/F/21/90010), Programme Grant (RG/F/25/110169) and Research Excellence (RE/24/130012) awards from the British Heart Foundation. A.J.F.T. is supported by a Clinical Research Training Fellowship (FS/CRTF/25/24735) from the British Heart Foundation. A.J.F.T. has received honoraria from Roche Diagnostics, outside the submitted work. The funders had no role in the study design; in the collection, analysis, and interpretation of data; in the writing of the literature review; and in the decision to submit the review for publication.
Informed Consent Statement
This article does not contain any studies with human or animal subjects performed by the author. No ethical approval or patient consent was required.
Data Availability Statement
No datasets containing patient or population-level data were processed during this study. The data spreadsheet used to extract and map the key topics for the literature review is made available as supplementary data.
Conflicts of Interest
The authors declare no competing interests.
References
- Khan MS, Shahid I, Bennis A, Rakisheva A, Metra M, Butler J. Global epidemiology of heart failure. Nature Reviews Cardiology. Nature Publishing Group UK London; 2024;21:717–34.
- Conrad N, Judge A, Tran J, Mohseni H, Hedgecott D, Crespillo AP, et al. Temporal trends and patterns in heart failure incidence: a population-based study of 4 million individuals. The Lancet. Elsevier; 2018;391:572–80. [CrossRef]
- Darvish M, Shakoor A, Feyz L, Schaap J, van Mieghem NM, de Boer RA, et al. Heart failure: assessment of the global economic burden. Eur Heart J. 2025;46:3069–78. [CrossRef]
- Averbuch T, Lee SF, Zagorski B, Pandey A, Petrie MC, Biering-Sorensen T, et al. Long-term clinical outcomes and healthcare resource utilization in male and female patients following hospitalization for heart failure. European Journal of Heart Failure. 2025;27:377–87. [CrossRef]
- Bozkurt B, Coats AJS, Tsutsui H, Abdelhamid CM, Adamopoulos S, Albert N, et al. Universal definition and classification of heart failure: a report of the Heart Failure Society of America, Heart Failure Association of the European Society of Cardiology, Japanese Heart Failure Society and Writing Committee of the Universal Definition of Heart Failure. European Journal of Heart Failure. 2021;23:352–80. [CrossRef]
- Redfield MM, Borlaug BA. Heart Failure With Preserved Ejection Fraction: A Review. JAMA. 2023;329:827–38. [CrossRef]
- Wu L, Rizwan A, Rodriguez M, El Hachem K, Hassan Virk HU, Khawaja M, et al. Heart Failure With Mildly Reduced Ejection Fraction. JACC: Advances. American College of Cardiology Foundation; 2026;5:102476. [CrossRef]
- Murphy SP, Ibrahim NE, Januzzi Jr JL. Heart failure with reduced ejection fraction: a review. Jama. 2020;324:488–504.
- Walsh MN, Kober L, Sliwa K, Adamo M, Agarwal A, Banerjee A, et al. AHA/ACC/ESC/WHF Expert Consensus Document: Second Universal Definition of Heart Failure (2026). Circulation [Internet]. American Heart Association; [cited 2026 July 2];0. [CrossRef]
- McDonagh TA, Metra M, Adamo M, Gardner RS, Baumbach A, Böhm M, et al. 2023 Focused Update of the 2021 ESC Guidelines for the diagnosis and treatment of acute and chronic heart failure: Developed by the task force for the diagnosis and treatment of acute and chronic heart failure of the European Society of Cardiology (ESC) With the special contribution of the Heart Failure Association (HFA) of the ESC. Eur Heart J. 2023;44:3627–39. [CrossRef]
- Heidenreich PA, Bozkurt B, Aguilar D, Allen LA, Byun JJ, Colvin MM, et al. 2022 AHA/ACC/HFSA Guideline for the Management of Heart Failure: A Report of the American College of Cardiology/American Heart Association Joint Committee on Clinical Practice Guidelines. Circulation. American Heart Association; 2022;145:e895–1032. [CrossRef]
- Reddy YNV, Tada A, Obokata M, Carter RE, Kaye DM, Handoko ML, et al. Evidence-Based Application of Natriuretic Peptides in the Evaluation of Chronic Heart Failure With Preserved Ejection Fraction in the Ambulatory Outpatient Setting. Circulation. American Heart Association; 2025;151:976–89. [CrossRef]
- Van Ommen A-M, Kessler EL, Valstar G, Onland-Moret NC, Cramer MJ, Rutten F, et al. Electrocardiographic Features of Left Ventricular Diastolic Dysfunction and Heart Failure With Preserved Ejection Fraction: A Systematic Review. Front Cardiovasc Med [Internet]. Frontiers; 2021 [cited 2026 June 10];8. [CrossRef]
- Docherty KF, Lam CSP, Rakisheva A, Coats AJS, Greenhalgh T, Metra M, et al. Heart Failure Diagnosis in the General Community – Who, how and When? A Clinical Consensus Statement of the Heart Failure Association (HFA) of the European Society of Cardiology (ESC). Eur J Heart Fail. 2023;25:1185–98. [CrossRef]
- Cox ZL, Nandkeolyar S, Johnson AJ, Lindenfeld J, Rali AS. In-hospital Initiation and Up-titration of Guideline-directed Medical Therapies for Heart Failure with Reduced Ejection Fraction. Card Fail Rev. 2022;8:e21. [CrossRef]
- Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. Nature Publishing Group; 2019;25:44–56. [CrossRef]
- Rajpurkar P, Chen E, Banerjee O, Topol EJ. AI in health and medicine. Nat Med. Nature Publishing Group; 2022;28:31–8. [CrossRef]
- Lee KK, Doudesis D, Anwar M, Astengo F, Chenevier-Gobeaux C, Claessens Y-E, et al. Development and validation of a decision support tool for the diagnosis of acute heart failure: systematic review, meta-analysis, and modelling study. BMJ. British Medical Journal Publishing Group; 2022;377:e068424. [CrossRef]
- Attia ZI, Kapa S, Lopez-Jimenez F, McKie PM, Ladewig DJ, Satam G, et al. Screening for cardiac contractile dysfunction using an artificial intelligence–enabled electrocardiogram. Nat Med. Nature Publishing Group; 2019;25:70–4. [CrossRef]
- Siontis KC, Noseworthy PA, Attia ZI, Friedman PA. Artificial intelligence-enhanced electrocardiography in cardiovascular disease management. Nat Rev Cardiol. Nature Publishing Group; 2021;18:465–78. [CrossRef]
- Schouten D, Nicoletti G, Dille B, Chia C, Vendittelli P, Schuurmans M, et al. Navigating the landscape of multimodal AI in medicine: A scoping review on technical challenges and clinical applications. Medical Image Analysis. 2025;105:103621. [CrossRef]
- Pieske B, Tschöpe C, de Boer RA, Fraser AG, Anker SD, Donal E, et al. How to diagnose heart failure with preserved ejection fraction: the HFA–PEFF diagnostic algorithm: a consensus recommendation from the Heart Failure Association (HFA) of the European Society of Cardiology (ESC). Eur Heart J. 2019;40:3297–317. [CrossRef]
- Januzzi JL, Chen-Tournoux AA, Christenson RH, Doros G, Hollander JE, Levy PD, et al. N-Terminal Pro-B-Type Natriuretic Peptide in the Emergency Department: The ICON-RELOADED Study. J Am Coll Cardiol. 2018;71:1191–200. [CrossRef]
- Bayes-Genis A, Docherty KF, Petrie MC, Januzzi JL, Mueller C, Anderson L, et al. Practical algorithms for early diagnosis of heart failure and heart stress using NT-proBNP: a clinical consensus statement from the Heart Failure Association of the ESC. European journal of heart failure. Oxford University Press; 2023;25:1891–8.
- Sanders-van Wijk S, Maeder MT, Nietlispach F, Rickli H, Estlinbaum W, Erne P, et al. Long-Term Results of Intensified, N-Terminal-Pro-B-Type Natriuretic Peptide–Guided Versus Symptom-Guided Treatment in Elderly Patients With Heart Failure. Circulation: Heart Failure. American Heart Association; 2014;7:131–9. [CrossRef]
- Doudesis D, Lee KK, Anwar M, Singer AJ, Hollander JE, Chenevier-Gobeaux C, et al. Machine learning to optimize use of natriuretic peptides in the diagnosis of acute heart failure. Eur Heart J Acute Cardiovasc Care. 2025;14:474–88. [CrossRef]
- Neyazi M, Bremer JP, Knorr MS, Gross S, Brederecke J, Schweingruber N, et al. Deep learning-based NT-proBNP prediction from the ECG for risk assessment in the community. Clin Chem Lab Med. 2024;62:740–52. [CrossRef]
- Allen CJ, Guha K, Sharma R. How to Improve Time to Diagnosis in Acute Heart Failure – Clinical Signs and Chest X-ray. 2015 [cited 2026 June 12]; https://www.cfrjournal.com/articles/how-improve-time-diagnosis-acute-heart-failure-clinical-signs-and-chest-x-ray?language_content_entity=en. Accessed 12 June 2026.
- Lee MS, Kim YS, Kim M, Usman M, Byon SS, Kim SH, et al. Evaluation of the feasibility of explainable computer-aided detection of cardiomegaly on chest radiographs using deep learning. Sci Rep. Nature Publishing Group; 2021;11:16885. [CrossRef]
- Horng S, Liao R, Wang X, Dalal S, Golland P, Berkowitz SJ. Deep Learning to Quantify Pulmonary Edema in Chest Radiographs. Radiol Artif Intell. 2021;3:e190228. [CrossRef]
- Conner SM, Husaini M, Fiore M, Ramadan M, Hoemann B, Arnold N, et al. Demonstrating Feasibility of Point of Care Ultrasound (POCUS)-Guided Inpatient Transthoracic Echo Triage Decision Pathway. POCUS J. 10:45–52. [CrossRef]
- Jiménez-Blanco M, Cordero D, Zamorano JL. Left ventricular ejection fraction… What else? Cardiol J. 2020;27:6–7. [CrossRef]
- Xie Y, Zhang L, Sun W, Zhu Y, Zhang Z, Chen L, et al. Artificial Intelligence in Diagnosis of Heart Failure. Journal of the American Heart Association. Wiley; 2025;14:e039511. [CrossRef]
- Wu J, Biswas D, Brown S, Ryan M, Bernstein BS, Tam To B, et al. Artificial intelligence methods to detect heart failure with preserved ejection fraction within electronic health records: an equitable disease detection model. Eur Heart J Digit Health. 2026;7:ztaf107. [CrossRef]
- O’Driscoll JM, Hawkes W, Beqiri A, Mumith A, Parker A, Upton R, et al. Left ventricular assessment with artificial intelligence increases the diagnostic accuracy of stress echocardiography. Eur Heart J Open. 2022;2:oeac059. [CrossRef]
- Betemariam T, Aleka A, Ahmed E, Worku T, Mebrahtu Y, Androulakis E, et al. Barriers to cardiovascular magnetic resonance imaging scan performance and reporting by cardiologists: a systematic literature review. Eur Heart J - Imaging Methods Practice. 2025;3:qyaf010. [CrossRef]
- Bernard O, Lalande A, Zotti C, Cervenansky F, Yang X, Heng P-A, et al. Deep Learning Techniques for Automatic MRI Cardiac Multi-Structures Segmentation and Diagnosis: Is the Problem Solved? IEEE Trans Med Imaging. 2018;37:2514–25. [CrossRef]
- Dhingra LS, Shen M, Mangla A, Khera R. Cardiovascular Care Innovation through Data-Driven Discoveries in the Electronic Health Record. Am J Cardiol. 2023;203:136–48. [CrossRef]
- Liu Y, Tan Z, Zhang Z, Wang S, Guo J, Liu H, et al. A natural language processing-based approach for early detection of heart failure onset using electronic health records. Knowledge-Based Systems. 2025;327:114102. [CrossRef]
- Tang F-SK-B, Verket M, Müller-Wieland D, Brandts J, Jacobsen M, Pütz A, et al. End-to-end pipeline for automated heart failure diagnosis with clinical notes using SNOMED-CT. Sci Rep. Nature Publishing Group; 2026;16:12751. [CrossRef]
- Januzzi JL, Chen-Tournoux AA, Christenson RH, Doros G, Hollander JE, Levy PD, et al. N-Terminal Pro–B-Type Natriuretic Peptide in the Emergency Department. JACC. American College of Cardiology Foundation; 2018;71:1191–200. [CrossRef]
- Tsutsui H, Albert NM, Coats AJS, Anker SD, Bayes-Genis A, Butler J, et al. Natriuretic peptides: role in the diagnosis and management of heart failure: a scientific statement from the Heart Failure Association of the European Society of Cardiology, Heart Failure Society of America and Japanese Heart Failure Society. European Journal of Heart Failure. 2023;25:616–31. [CrossRef]
- Khunti K, Squire I, Abrams KR, Sutton AJ. Accuracy of a 12-lead electrocardiogram in screening patients with suspected heart failure for open access echocardiography: a systematic review and meta-analysis. Eur J Heart Fail. 2004;6:571–6. [CrossRef]
- Chambers J, Shah BN, Garbi M, Campbell B, Vassiliou VS, Schlosshan D. Management of Echocardiography Requests for the Detection and Follow-Up of Heart Valve Disease: A Consensus Statement From the British Heart Valve Society. Clin Cardiol. 2025;48:e70099. [CrossRef]
- Morbach C, Buck T, Rost C, Peter S, Günther S, Störk S, et al. Point-of-care B-type natriuretic peptide and portable echocardiography for assessment of patients with suspected heart failure in primary care: rationale and design of the three-part Handheld-BNP program and results of the training study. Clin Res Cardiol. 2018;107:95–107. [CrossRef]
- McGilvray MMO, Heaton J, Guo A, Masood MF, Cupps BP, Damiano M, et al. Electronic Health Record-Based Deep Learning Prediction of Death or Severe Decompensation in Heart Failure Patients. JACC: Heart Failure. American College of Cardiology Foundation; 2022;10:637–47. [CrossRef]
- Perez Vicencio D, Doudesis D, Thurston AJF, Chenevier-Gobeaux C, Claessens Y-E, Lopez Ayala P, et al. Machine learning to optimize the diagnostic performance of natriuretic peptides for acute heart failure across age groups. ESC Heart Fail. 2026;13:xvaf006. [CrossRef]
- Wu K-HH, Wolford BN, Du J, Yu X, Douville NJ, Mathis MR, et al. Integrating large scale genetic and clinical information to predict cases of heart failure. Commun Med. Nature Publishing Group; 2025;5:493. [CrossRef]
- Li R, Liu Y, Zhao Z, Zhang C, Dong W, Qi Y, et al. Machine learning-based phenotyping and assessment of treatment responses in heart failure with preserved ejection fraction. eClinicalMedicine [Internet]. Elsevier; 2025 [cited 2026 June 3];88. [CrossRef]
- Hannun AY, Rajpurkar P, Haghpanahi M, Tison GH, Bourn C, Turakhia MP, et al. Cardiologist-level arrhythmia detection and classification in ambulatory electrocardiograms using a deep neural network. Nat Med. Nature Publishing Group; 2019;25:65–9. [CrossRef]
- Sangha V, Nargesi AA, Dhingra LS, Khunte A, Mortazavi BJ, Ribeiro AH, et al. Detection of Left Ventricular Systolic Dysfunction From Electrocardiographic Images. Circulation. American Heart Association; 2023;148:765–77. [CrossRef]
- Khunte A, Sangha V, Oikonomou EK, Dhingra LS, Aminorroaya A, Mortazavi BJ, et al. Detection of left ventricular systolic dysfunction from single-lead electrocardiography adapted for portable and wearable devices. NPJ Digit Med. 2023;6:124. [CrossRef]
- Hong D, Song S-H, Shin H, Bak M, Kim J, Kim D, et al. Artificial intelligence-enabled electrocardiogram model for predicting heart failure with preserved ejection fraction: a single-center study. Eur Heart J Digit Health. 2025;6:959–68. [CrossRef]
- Akerman AP, Al-Roub N, Angell-James C, Cassidy MA, Thompson R, Bosque L, et al. External validation of artificial intelligence for detection of heart failure with preserved ejection fraction. Nat Commun. Nature Publishing Group; 2025;16:2915. [CrossRef]
- Desai MY, Jadam S, Abusafia M, Rutkowski K, Ospina S, Gaballa A, et al. Real-World Artificial Intelligence-Based Electrocardiographic Analysis to Diagnose Hypertrophic Cardiomyopathy. JACC Clin Electrophysiol. 2025;11:1324–33. [CrossRef]
- Schlesinger RP, Ferreira Felix I, Harmon DM, Malik AA, Dispenzieri A, Fonseca R, et al. Artificial Intelligence-Enhanced Electrocardiogram. JACC: Case Reports. American College of Cardiology Foundation; 2025;30:102968. [CrossRef]
- Himmelreich JCL, Harskamp RE. Diagnostic accuracy of the PMcardio smartphone application for artificial intelligence-based interpretation of electrocardiograms in primary care (AMSTELHEART-1). Cardiovasc Digit Health J. 2023;4:80–90. [CrossRef]
- Ouyang D, He B, Ghorbani A, Yuan N, Ebinger J, Langlotz CP, et al. Video-based AI for beat-to-beat assessment of cardiac function. Nature. Nature Publishing Group; 2020;580:252–6. [CrossRef]
- He B, Kwan AC, Cho JH, Yuan N, Pollick C, Shiota T, et al. Blinded, randomized trial of sonographer versus AI cardiac function assessment. Nature. Nature Publishing Group; 2023;616:520–4. [CrossRef]
- Duffy G, Cheng PP, Yuan N, He B, Kwan AC, Shun-Shin MJ, et al. High-Throughput Precision Phenotyping of Left Ventricular Hypertrophy With Cardiovascular Deep Learning. JAMA Cardiol. 2022;7:386–95. [CrossRef]
- Huang W, Koh T, Tromp J, Chandramouli C, Ewe SH, Ng CT, et al. Point-of-care AI-enhanced novice echocardiography for screening heart failure (PANES-HF). Sci Rep. Nature Publishing Group; 2024;14:13503. [CrossRef]
- Fan W, Yang Y, Qi J, Zhang Q, Liao C, Wen L, et al. A deep-learning-based framework for identifying and localizing multiple abnormalities and assessing cardiomegaly in chest X-ray. Nat Commun. Nature Publishing Group; 2024;15:1347. [CrossRef]
- Bhave S, Rodriguez V, Poterucha T, Mutasa S, Aberle D, Capaccione KM, et al. Deep learning to detect left ventricular structural abnormalities in chest X-rays. Eur Heart J. 2024;45:2002–12. [CrossRef]
- Kusunose K, Hirata Y, Yamaguchi N, Kosaka Y, Tsuji T, Kotoku J, et al. Deep learning approach for analyzing chest x-rays to predict cardiac events in heart failure. Front Cardiovasc Med. 2023;10:1081628. [CrossRef]
- Shad R, Zakka C, Kaur D, Mathur M, Fong R, Cho J, et al. A generalizable deep learning system for cardiac MRI. Nat Biomed Eng. Nature Publishing Group; 2026;1–16. [CrossRef]
- Oikonomou EK, Williams MC, Kotanidis CP, Desai MY, Marwan M, Antonopoulos AS, et al. A novel machine learning-derived radiotranscriptomic signature of perivascular fat improves cardiac risk prediction using coronary CT angiography. Eur Heart J. 2019;40:3529–43. [CrossRef]
- Chan K, Wahome E, Tsiachristas A, Antonopoulos AS, Patel P, Lyasheva M, et al. Inflammatory risk and cardiovascular events in patients without obstructive coronary artery disease: the ORFAN multicentre, longitudinal cohort study. The Lancet. Elsevier; 2024;403:2606–18. [CrossRef]
- CT coronary angiography in patients with suspected angina due to coronary heart disease (SCOT-HEART): an open-label, parallel-group, multicentre trial. The Lancet. Elsevier; 2015;385:2383–91. [CrossRef]
- Huang S-C, Pareek A, Seyyedi S, Banerjee I, Lungren MP. Fusion of medical imaging and electronic health records using deep learning: a systematic review and implementation guidelines. npj Digit Med. Nature Publishing Group; 2020;3:1–9. [CrossRef]
- Acosta JN, Falcone GJ, Rajpurkar P, Topol EJ. Multimodal biomedical AI. Nat Med. Nature Publishing Group; 2022;28:1773–84. [CrossRef]
- He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. Proceedings of the IEEE conference on computer vision and pattern recognition. 2016. p. 770–8.
- Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is All you Need. Advances in Neural Information Processing Systems [Internet]. Curran Associates, Inc.; 2017 [cited 2026 June 3]. https://papers.nips.cc/paper_files/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html. Accessed 3 June 2026.
- Guarrasi V, Aksu F, Caruso CM, Di Feola F, Rofena A, Ruffini F, et al. A systematic review of intermediate fusion in multimodal deep learning for biomedical applications. Image and Vision Computing. 2025;158:105509. [CrossRef]
- Petmezas G, Stefanopoulos L, Kilintzis V, Tzavelis A, Rogers JA, Katsaggelos AK, et al. State-of-the-Art Deep Learning Methods on Electrocardiogram Data: Systematic Review. JMIR Med Inform. 2022;10:e38454. [CrossRef]
- Huang Y, Yang J, Sun Q, Yuan Y, Li H, Hou Y. Multi-residual 2D network integrating spatial correlation for whole heart segmentation. Comput Biol Med. 2024;172:108261. [CrossRef]
- Guo A, Smith S, Khan YM, Ii JRL, Foraker RE. Application of a time-series deep learning model to predict cardiac dysrhythmias in electronic health records. PLOS ONE. Public Library of Science; 2021;16:e0239007. [CrossRef]
- Kusuma S, Jothi KR. ECG signals-based automated diagnosis of congestive heart failure using Deep CNN and LSTM architecture. Biocybernetics and Biomedical Engineering. 2022;42:247–57. [CrossRef]
- Çınar A, Tuncer SA. Classification of normal sinus rhythm, abnormal arrhythmia and congestive heart failure ECG signals using LSTM and hybrid CNN-SVM deep neural networks. Computer Methods in Biomechanics and Biomedical Engineering. Taylor & Francis; 2021;24:203–14. [CrossRef]
- Yang Z, Wang X, Wang J, Guang Q, Ding X, Liu H, et al. Multimodal Transformer–Based Electrocardiogram Analysis for Cardiovascular Comorbidity Detection: Model Development and Validation Study. JMIR Formative Research. JMIR Publications Inc., Toronto, Canada; 2026;10:e80815. [CrossRef]
- Kingma DP, Welling M. Auto-Encoding Variational Bayes. CoRR [Internet]. 2013 [cited 2026 June 14]; https://www.semanticscholar.org/paper/Auto-Encoding-Variational-Bayes-Kingma-Welling/5f5dc5b9a2ba710937e2c413b37b053cd673df02. Accessed 14 June 2026.
- Beetz M, Banerjee A, Grau V. Multi-Domain Variational Autoencoders for Combined Modeling of MRI-Based Biventricular Anatomy and ECG-Based Cardiac Electrophysiology. Front Physiol [Internet]. Frontiers; 2022 [cited 2026 June 14];13. [CrossRef]
- Lu S, Chen R, Wei W, Belovsky M, Lu X. Understanding Heart Failure Patients EHR Clinical Features via SHAP Interpretation of Tree-Based Machine Learning Model Predictions. AMIA Annu Symp Proc. 2022;2021:813–22.
- Hamid M, Hajjej F, Alluhaidan AS, bin Mannie NW. Fine tuned CatBoost machine learning approach for early detection of cardiovascular disease through predictive modeling. Sci Rep. Nature Publishing Group; 2025;15:31199. [CrossRef]
- Xiong Y-J, Ling S-S, Hu W-T, Wen X-Y, Zheng X-M, Liu W-H, et al. Real-time monitoring of in-hospital mortality risk in intensive care units heart failure patients using an extreme gradient boosting model. Journal of Geriatric Cardiology. 中国人民解放军总医院; 2026;23:184–205. [CrossRef]
- Desai AS, Pandey A, Suratekar R, Ahmad FS, Alger HM, Anto AG, et al. Predicting Heart Failure From 12-Lead ECGs Using AI. JACC. American College of Cardiology Foundation; 2026;87:990–1005. [CrossRef]
- Christensen M, Vukadinovic M, Yuan N, Ouyang D. Vision–language foundation model for echocardiogram interpretation. Nat Med. Nature Publishing Group; 2024;30:1481–8. [CrossRef]
- Vukadinovic M, Chiu I-M, Tang X, Yuan N, Chen T-Y, Cheng P, et al. Comprehensive echocardiogram evaluation with view primed vision language AI. Nature. Nature Publishing Group; 2026;650:970–7. [CrossRef]
- Soto JT, Weston Hughes J, Sanchez PA, Perez M, Ouyang D, Ashley EA. Multimodal deep learning enhances diagnostic precision in left ventricular hypertrophy. Eur Heart J Digit Health. 2022;3:380–9. [CrossRef]
- Kolk MZH, Ruipérez-Campillo S, Allaart CP, Wilde AAM, Knops RE, Narayan SM, et al. Multimodal explainable artificial intelligence identifies patients with non-ischaemic cardiomyopathy at risk of lethal ventricular arrhythmias. Sci Rep. Nature Publishing Group; 2024;14:14889. [CrossRef]
- Huang Z, Pan W, Alishetti S, Beecy AN, Liu Z, Gong A, et al. Multimodal multi-instance learning for cardiopulmonary exercise testing performance prediction. npj Digit Med. Nature Publishing Group; 2026;9:304. [CrossRef]
- Yang B, Qin Y, Li Y, Xie R, Jiang L, He J, et al. Multimodal Fusion of Echocardiogram Images and Electronic Medical Records for Heart Disease Screening: Retrospective Algorithm Development and Validation Study. JMIR Medical Informatics. JMIR Publications Inc., Toronto, Canada; 2026;14:e78949. [CrossRef]
- Botros J, Mourad-Chehade F, Laplanche D. Explainable multimodal data fusion framework for heart failure detection: Integrating CNN and XGBoost. Biomedical Signal Processing and Control. 2025;100:106997. [CrossRef]
- González S, Yi AK-C, Hsieh W-T, Chen W-C, Wang C-L, Wu VC-C, et al. Multi-modal heart failure risk estimation based on short ECG and sampled long-term HRV. Information Fusion. 2024;107:102337. [CrossRef]
- Oikonomou EK, Batinica B, Dhingra LS, Aminorroaya A, Coppi A, Khera R. TARGET-AI: A Foundational Approach for the Targeted Deployment of Artificial Intelligence Electrocardiography in the Electronic Health Record. NEJM AI. Massachusetts Medical Society; 2026;3:AIoa2500588. [CrossRef]
- Lee C-K, Chen T-L, Wu J-E, Liao M-T, Wang C, Wang W, et al. Multimodal deep learning models utilizing chest X-ray and electronic health record data for predictive screening of acute heart failure in emergency department. Computer Methods and Programs in Biomedicine. 2024;255:108357. [CrossRef]
- Woolley RJ, Ceelen D, Ouwerkerk W, Tromp J, Figarska SM, Anker SD, et al. Machine learning based on biomarker profiles identifies distinct subgroups of heart failure with preserved ejection fraction. European Journal of Heart Failure. 2021;23:983–91. [CrossRef]
- Sanchez-Martinez S, Duchateau N, Erdei T, Kunszt G, Aakhus S, Degiovanni A, et al. Machine Learning Analysis of Left Ventricular Function to Characterize Heart Failure With Preserved Ejection Fraction. Circulation: Cardiovascular Imaging. American Heart Association; 2018;11:e007138. [CrossRef]
- Kaur D, Hughes JW, Rogers AJ, Kang G, Narayan SM, Ashley EA, et al. Race, Sex, and Age Disparities in the Performance of ECG Deep Learning Models Predicting Heart Failure. Circulation: Heart Failure. American Heart Association; 2024;17:e010879. [CrossRef]
- Li Y, Wang H, Luo Y. Improving Fairness in the Prediction of Heart Failure Length of Stay and Mortality by Integrating Social Determinants of Health. Circulation: Heart Failure. American Heart Association; 2022;15:e009473. [CrossRef]
- Irvin J, Rajpurkar P, Ko M, Yu Y, Ciurea-Ilcus S, Chute C, et al. CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison. Proceedings of the AAAI Conference on Artificial Intelligence. 2019;33:590–7. [CrossRef]
- Huang S-C, Huo Z, Steinberg E, Chiang C-C, Lungren MP, Langlotz CP, et al. INSPECT: A Multimodal Dataset for Pulmonary Embolism Diagnosis and Prognosis [Internet]. arXiv; 2023 [cited 2025 Feb 11]. [CrossRef]
- Yosinski J, Clune J, Bengio Y, Lipson H. How transferable are features in deep neural networks? Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 2. Cambridge, MA, USA: MIT Press; 2014. p. 3320–8.
- Zhou R, Lu L, Liu Z, Xiang T, Liang Z, Clifton DA, et al. Semi-Supervised Learning for Multi-Label Cardiovascular Diseases Prediction: A Multi-Dataset Study. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2024;46:3305–20. [CrossRef]
- Weimann K, Conrad TOF. Transfer learning for ECG classification. Sci Rep. Nature Publishing Group; 2021;11:5251. [CrossRef]
- Chen T, Kornblith S, Norouzi M, Hinton G. A Simple Framework for Contrastive Learning of Visual Representations [Internet]. arXiv; 2020 [cited 2026 Apr 1]. [CrossRef]
- Mathew G, Barbosa D, Prince J, Venkatraman S. Foundation models for cardiovascular disease detection via biosignals from digital stethoscopes. npj Cardiovasc Health. Nature Publishing Group; 2024;1:25. [CrossRef]
- Gu X, Tang W, Han J, Sangha V, Liu F, Gowda SN, et al. Cardiac health assessment across scenarios and devices using a multimodal foundation model pretrained on data from 1.7 million individuals. Nat Mach Intell. Nature Publishing Group; 2026;8:220–33. [CrossRef]
- Majid MD, Anwar M, Bilal SF, Hussain M, Zubair M, Sultana J, et al. A hybrid learning framework for automated multiclass electrocardiogram classification with SimCardioNet. Sci Rep. Nature Publishing Group; 2026;16:7621. [CrossRef]
- Poterucha TJ, Jing L, Ricart RP, Adjei-Mosi M, Finer J, Hartzel D, et al. Detecting structural heart disease from electrocardiograms using AI. Nature. Nature Publishing Group; 2025;644:221–30. [CrossRef]
- Johnson AEW, Pollard TJ, Berkowitz SJ, Greenbaum NR, Lungren MP, Deng C, et al. MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. Sci Data. 2019;6:317. [CrossRef]
- Grill J-B, Strub F, Altché F, Tallec C, Richemond PH, Buchatskaya E, et al. Bootstrap your own latent a new approach to self-supervised learning. Proceedings of the 34th International Conference on Neural Information Processing Systems [Internet]. Red Hook, NY, USA: Curran Associates Inc.; 2020 [cited 2026 June 3]. p. 21271–84. https://dl.acm.org/doi/10.5555/3495724.3497510. Accessed 3 June 2026.
- Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, et al. Learning Transferable Visual Models From Natural Language Supervision. Proceedings of the 38th International Conference on Machine Learning [Internet]. PMLR; 2021 [cited 2026 June 3]. p. 8748–63. https://proceedings.mlr.press/v139/radford21a.html. Accessed 3 June 2026.
- Li J, Aguirre AD, Junior VM, Jin J, Liu C, Zhong L, et al. An Electrocardiogram Foundation Model Built on over 10 Million Recordings. NEJM AI. Massachusetts Medical Society; 2025;2:AIoa2401033. [CrossRef]
- Zhong H, Wu J, Zhao W, Xu X, Hou R, Zhao L, et al. A Self-supervised Learning Based Framework for Automatic Heart Failure Classification on Cine Cardiac Magnetic Resonance Image. 2021 43rd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC) [Internet]. 2021 [cited 2026 June 7]. p. 2887–90. [CrossRef]
- Holste G, Oikonomou EK, Mortazavi BJ, Wang Z, Khera R. Efficient deep learning-based automated diagnosis from echocardiography with contrastive self-supervised learning. Commun Med (Lond). 2024;4:133. [CrossRef]
- Ma D, Pang J, Gotway MB, Liang J. A fully open AI foundation model applied to chest radiography. Nature. Nature Publishing Group; 2025;643:488–98. [CrossRef]
- Raghu M, Zhang C, Kleinberg J, Bengio S. Transfusion: understanding transfer learning for medical imaging. Proceedings of the 33rd International Conference on Neural Information Processing Systems [Internet]. Red Hook, NY, USA: Curran Associates Inc.; 2019 [cited 2026 June 3]. p. 3347–57. https://dl.acm.org/doi/10.5555/3454287.3454588. Accessed 3 June 2026.
- Goto S, Solanki D, John JE, Yagi R, Homilius M, Ichihara G, et al. Multinational Federated Learning Approach to Train ECG and Echocardiogram Models for Hypertrophic Cardiomyopathy Detection. Circulation. American Heart Association; 2022;146:755–69. [CrossRef]
- Narotamo H, Dias M, Santos R, Carreiro AV, Gamboa H, Silveira M. Deep learning for ECG classification: A comparative study of 1D and 2D representations and multimodal fusion approaches. Biomedical Signal Processing and Control. 2024;93:106141. [CrossRef]
- Nolin-Lapalme A, Sowa A, Delfrate J, Tastet O, Corbin D, Kulbay M, et al. Foundation models for electrocardiogram interpretation: clinical implications. Eur Heart J. 2026;ehaf1119. [CrossRef]
- Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. 2017 IEEE International Conference on Computer Vision (ICCV) [Internet]. 2017 [cited 2025 Feb 12]. p. 618–26. [CrossRef]
- Salih A, Boscolo Galazzo I, Gkontra P, Lee AM, Lekadir K, Raisi-Estabragh Z, et al. Explainable Artificial Intelligence and Cardiac Imaging: Toward More Interpretable Models. Circulation: Cardiovascular Imaging. American Heart Association; 2023;16:e014519. [CrossRef]
- Haupt M, Maurer MH, Thomas RP. Explainable Artificial Intelligence in Radiological Cardiovascular Imaging-A Systematic Review. Diagnostics (Basel). Basel, Switzerland; 2025;15:1399. [CrossRef]
- Lundberg SM, Lee S-I. A Unified Approach to Interpreting Model Predictions. Advances in Neural Information Processing Systems [Internet]. Curran Associates, Inc.; 2017 [cited 2024 Dec 30]. https://papers.nips.cc/paper_files/paper/2017/hash/8a20a8621978632d76c43dfd28b67767-Abstract.html. Accessed 30 Dec 2024.
- Sun Z, Wang Z, Yun Z, Sun X, Lin J, Zhang X, et al. Machine Learning-Based Model for Worsening Heart Failure Risk in Chinese Chronic Heart Failure Patients. ESC Heart Fail. 2025;12:211–28. [CrossRef]
- Luo H, Xiang C, Zeng L, Li S, Mei X, Xiong L, et al. SHAP based predictive modeling for 1 year all-cause readmission risk in elderly heart failure patients: feature selection and model interpretation. Sci Rep. Nature Publishing Group; 2024;14:17728. [CrossRef]
- Parcalabescu L, Frank A. MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) [Internet]. 2023 [cited 2025 Feb 11]. p. 4032–59. [CrossRef]
- Verma S, Boonsanong V, Hoang M, Hines K, Dickerson J, Shah C. Counterfactual Explanations and Algorithmic Recourses for Machine Learning: A Review. ACM Comput Surv. 2024;56:312:1-312:42. [CrossRef]
- Sundararajan M, Taly A, Yan Q. Axiomatic Attribution for Deep Networks. Proceedings of the 34th International Conference on Machine Learning [Internet]. PMLR; 2017 [cited 2026 June 15]. p. 3319–28. https://proceedings.mlr.press/v70/sundararajan17a.html. Accessed 15 June 2026.
- van der Velden BHM, Kuijf HJ, Gilhuijs KGA, Viergever MA. Explainable artificial intelligence (XAI) in deep learning-based medical image analysis. Medical Image Analysis. 2022;79:102470. [CrossRef]
- Arends BKO, van Amsterdam WAC, van der Harst P, van Smeden M, van Es R, van de Leur RR. Signal or noise? Evaluating commonly used attribution methods for explaining deep neural networks in electrocardiogram classification. Eur Heart J Digit Health. 2026;7:ztag038. [CrossRef]
- Ghassemi M, Oakden-Rayner L, Beam AL. The false hope of current approaches to explainable artificial intelligence in health care. Lancet Digit Health. 2021;3:e745–50. [CrossRef]
- Rudin C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat Mach Intell. Nature Publishing Group; 2019;1:206–15. [CrossRef]
- van Dam PMEL, van Doorn WPTM, Sevenich L, Lambriks L, Cals JWL, Bekers O, et al. Machine learning for risk stratification in the emergency department (MARS-ED): a randomized controlled trial. Nat Commun. Nature Publishing Group; 2025;17:242. [CrossRef]
- Kelshiker MA, Bächtiger P, Petri CF, Nakhare S, Mansell J, Chhatwal K, et al. Triple cardiovascular disease detection with an artificial intelligence-enabled stethoscope (TRICORDER) in the UK: a cluster-randomised controlled implementation trial. The Lancet. Elsevier; 2026;407:704–15. [CrossRef]
- Pantanowitz L, Hanna M, Pantanowitz J, Lennerz J, Henricks WH, Shen P, et al. Regulatory Aspects of Artificial Intelligence and Machine Learning. Modern Pathology [Internet]. Elsevier; 2024 [cited 2026 June 15];37. [CrossRef]
Figure 1.
Primary integration points for AI in diagnostic pathways for heart failure, ranging from low-specificity (first presentation) to high-specificity (advanced imaging). AI algorithms can be adopted to guide treatment decisions and provide decision support at each stage of heart failure screening. Additionally, they can be used to guide referral decisions that affect treatment progression based on individual patient requirements, thereby supporting the transfer of risk assessment across layers of heterogeneous data streams. Created in BioRender. Georgiev, K. (2025). Abbreviations: CXR, chest X-ray; CMR, cardiac magnetic resonance; ECG, electrocardiogram; EHR, electronic health record; HFmrEF, heart failure with mildly reduced ejection fraction; HFpEF, heart failure with preserved ejection fraction; HFrEF, heart failure with reduced ejection fraction; LVEF, left ventricular ejection fraction; LVSD, left ventricular systolic dysfunction; SPECT, single-photon emission computed tomography.
Figure 1.
Primary integration points for AI in diagnostic pathways for heart failure, ranging from low-specificity (first presentation) to high-specificity (advanced imaging). AI algorithms can be adopted to guide treatment decisions and provide decision support at each stage of heart failure screening. Additionally, they can be used to guide referral decisions that affect treatment progression based on individual patient requirements, thereby supporting the transfer of risk assessment across layers of heterogeneous data streams. Created in BioRender. Georgiev, K. (2025). Abbreviations: CXR, chest X-ray; CMR, cardiac magnetic resonance; ECG, electrocardiogram; EHR, electronic health record; HFmrEF, heart failure with mildly reduced ejection fraction; HFpEF, heart failure with preserved ejection fraction; HFrEF, heart failure with reduced ejection fraction; LVEF, left ventricular ejection fraction; LVSD, left ventricular systolic dysfunction; SPECT, single-photon emission computed tomography.

Figure 2.
Overview of multimodal fusion paradigms for the diagnosis of heart failure. Three paradigms (early, intermediate, and late fusion) are shown, along with an example data flow for a pipeline that integrates tabular health records with ECG and chest X-ray data. Early fusion concatenates all raw data embeddings before model training (data-level). Intermediate/joint fusion uses separate encoder units per modality with an internal cross-modal fusion layer (feature-level), suitable for task-specific adaptation. Late fusion trains independent models per modality and aggregates their outputs, typically using an ensemble ML model (decision-level). Created in BioRender. Georgiev, K. (2025). Abbreviations: CXR, chest X-ray; ECG, electrocardiogram; EHR, electronic health record.
Figure 2.
Overview of multimodal fusion paradigms for the diagnosis of heart failure. Three paradigms (early, intermediate, and late fusion) are shown, along with an example data flow for a pipeline that integrates tabular health records with ECG and chest X-ray data. Early fusion concatenates all raw data embeddings before model training (data-level). Intermediate/joint fusion uses separate encoder units per modality with an internal cross-modal fusion layer (feature-level), suitable for task-specific adaptation. Late fusion trains independent models per modality and aggregates their outputs, typically using an ensemble ML model (decision-level). Created in BioRender. Georgiev, K. (2025). Abbreviations: CXR, chest X-ray; ECG, electrocardiogram; EHR, electronic health record.

Table 1.
Summary of primary diagnostic modalities for the assessment of heart failure: clinical role, availability, key limitations, and AI potential.
Table 1.
Summary of primary diagnostic modalities for the assessment of heart failure: clinical role, availability, key limitations, and AI potential.
| Modality | Clinical role in heart failure assessment | Availability and care setting | Key limitations | AI potential |
|---|---|---|---|---|
| 12-lead ECG | First-line investigation. Detects LBBB, LVH, ischaemic changes, AF and arrhythmias. High NPV for HFrEF when normal. Essential for rhythm characterisation. | Widely available at all care levels; low cost; point-of-care assessment; applicable in primary care, ED and community [10,14]. | Low sensitivity for HFpEF; cannot directly quantify LVEF; confounded by LBBB, pacemaker rhythms, lead artefacts; limited specificity in isolation. | Deep learning for LVSD detection [19]; population-scale pre-echocardiography triage; multi-lead transformer architectures and waveform foundation models. |
|
Natriuretic peptides (BNP / NT-proBNP) |
Guideline-recommended rule-in/rule-out biomarker for heart failure. Age-stratified NT-proBNP thresholds support heart failure diagnosis [23]. Supports guided therapy in emergency and outpatient care. | Widely available in ED and outpatient settings; point-of-care assays are available; costs vary across healthcare systems. | Elevated in multiple non-heart failure conditions (AF, PE, CKD, sepsis); elevated in renal dysfunction; diagnostic grey zone between 100-400 pg/mL [24]. NT-proBNP-guided therapy alone shows limited benefit [25]. | AI integration for heart failure phenotyping; NT-proBNP trajectory modelling for treatment progression; personalised risk stratification [26,27]. |
| Chest radiograph (CXR) | Assessment of cardiomegaly, pulmonary venous congestion, pleural effusions; rapid triage for non-cardiac causes of dyspnoea. Recommended in acute presentations. | Widely available in hospital/ED settings; portable systems enable bedside assessment; broadly accessible in secondary care. | Low sensitivity for chronic/ambulatory heart failure; up to 20% of patients with acute decompensated heart failure do not show radiographic congestion [28]. | Deep learning models for pulmonary oedema and cardiomegaly detection [29,30]; population-scale pre-echocardiography triage; |
| Transthoracic echocardiography (TTE) | Non-invasive investigation of heart function. Quantifies LVEF, LV dimensions and volumes, wall motion abnormalities, diastolic function, valve disease, and haemodynamic estimates. Essential for heart failure phenotyping and guiding treatment decisions. | Moderate availability; significant waiting times in many healthcare systems; limited in primary care; emerging handheld point-of-care devices [31]. | Operator-dependent; LVEF inter-observer variability; complex grading of diastolic function. Requires specialist referral for timely access [32]. | Automated LVEF measurement; automated wall motion scoring; HFpEF detection; AI-guided decisions for point-of-care screening [33,34,35]. |
| Cardiac magnetic resonance imaging (CMR) | Reference standard for LVEF and biventricular volume; tissue characterisation; Determination of ischaemic cardiomyopathy, myocarditis and other causes of cardiomyopathy. Recommended for indeterminate echocardiography. | Limited to secondary/tertiary centres; long scan times (45-60 min); contraindications; high cost and constrained capacity; not suitable for acute decompensation [36]. | Significant access barriers in most healthcare systems [36]; scan duration affects care use; post-processing variability. | Automated CMR segmentation; deep learning for hypertrophic and dilated cardiomyopathy [37]. |
|
Electronic health records (EHR / structured clinical data) |
Longitudinal clinical history, comorbidity burden, medication records, laboratory, vital signs, and care utilisation data. Enables heart failure risk stratification and supports population-level screening. | Widely available in digitised healthcare systems across all care settings; variable data quality, completeness, and interoperability across providers. | heart failure phenotyping from disease codes has moderate accuracy; structured data misses clinical nuance; missing and heterogeneous data; interoperability [38]; | Diagnostic predictions from EHR trajectories; NLP extraction from clinical notes; personalised risk stratification [18,39,40]. |
Abbreviations: AUROC, area under the receiver operating characteristic curve; AF, atrial fibrillation; CKD, chronic kidney disease; ED, emergency department; LBBB, left bundle branch block; LVH, left ventricular hypertrophy; LVSD, left ventricular systolic dysfunction; NLP, natural language processing; NPV, negative predictive value; PE, pulmonary embolism.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.