Submitted:
03 August 2026
Posted:
05 August 2026
You are already at the latest version
Abstract
Acute myeloid leukemia (AML) produces more molecular, imaging, and clinical data per patient than any hematologist can hold in mind at once, and each revision of the WHO, ICC, and European LeukemiaNet (ELN) frameworks adds to the load. Artificial intelligence (AI) and machine learning (ML) now reach into every stage of AML care. Deep-learning models read therapy-relevant mutations directly from bone-marrow smears; automated flow-cytometry gating reproduces expert calls in under a minute; and the first AI pathology devices for hematology have cleared regulatory review and entered clinical use. Beyond diagnosis, ML captures the age-dependent weight of individual mutations that categorical ELN scoring misses, drug-response prediction for venetoclax–azacitidine has been validated across multiple external cohorts, and large language models are being tested for tumor-board support and trial matching. The next wave, from clonal-architecture modeling and single-cell foundation models to digital twins and reinforcement learning for adaptive dosing, could move AML management from reactive toward predictive, evolution-aware care. This review departs from existing AI-in-hematology surveys in three ways: we (i) restrict scope to AML and organize the field around clinical decision points rather than technology categories, (ii) grade every tool on a five-tier clinical-readiness level (CRL-AML 1–5), which exposes hundreds of models clustered at CRL-AML 1–2 and none yet in prospective clinical evaluation, and (iii) close with a numbered three-year agenda naming the consortia, datasets, and pragmatic trials needed to carry the field from publication to practice.
Keywords:
acute myeloid leukemia
; artificial intelligence
; machine learning
; deep learning
; digital pathology
; flow cytometry
; measurable residual disease
; risk stratification
; drug response prediction
; large language models
1. Introduction
Acute myeloid leukemia is a family of molecularly distinct entities that share clonal expansion of myeloid precursors in the marrow but diverge in biology, prognosis, and response to treatment [1]. Reaching a diagnosis now means reconciling several layers of evidence at once. The 2022 WHO and International Consensus Classification (ICC) schemes, together with the ELN 2022 risk stratification and its 2024 update for less-intensive therapy, ask the clinician to fold morphology, immunophenotype, cytogenetics, and a widening panel of molecular markers into a single coherent assessment [1,2,3]. The quantitative burden is real: a typical AML genome carries roughly five recurrently mutated driver genes [4], on top of karyotypic abnormalities and serial measurable residual disease (MRD) readouts from flow cytometry and molecular assays. All of it must then be weighed against age, fitness, and comorbidity to settle on a treatment plan for one patient.
Even with targeted agents and venetoclax-based combinations, five-year overall survival (OS) in AML still sits near 32%, reaching 40–50% only in younger, fit adults who achieve complete remission. Older patients fare worse, although VIALE-A established venetoclax–azacitidine as a standard for those unfit for intensive therapy, extending median OS to 14.7 months against 9.6 months with azacitidine alone [1,5]. AI and ML offer one route through this complexity, and possibly toward better cure rates. Over five years the field has matured from proof-of-concept studies to multicenter validations and, in a handful of cases, regulatory authorizations and openly available clinical tools [6,7,8].
A 2025 bibliometric analysis counted 376 original AI research articles in hematology, with AML among the most studied diseases [9]. The work now ranges from deep learning that classifies bone-marrow cells and infers mutations from slide images, through ML that resolves ELN risk into continuous estimates, to multicohort-validated drug-response models for venetoclax-based regimens and large language model (LLM) tools for tumor-board support and trial matching [6,7,8,10].
Volume, however, is not the same as clinical readiness, and the gap between a published model and a usable tool stays wide. Of 1,016 AI/ML medical devices authorized by the U.S. Food and Drug Administration (FDA) through late 2024, hematology accounts for about 1.9%, some nineteen devices in all [11]. Prospective validation in AML is, for practical purposes, absent. Most models have been trained and tested at a few Western academic centers, which leaves open real questions of generalizability and equity [10].
Recent reviews have surveyed AI across hematology [6], focused on AML and myelodysplastic syndromes (MDS) [7], or systematically analyzed 426 publications from 2010 to 2024 [8]. This review takes a different stance on three points. First, we hold scope to AML and organize the material around clinical decision points rather than technology categories, so that a hematologist reading at the point of care can find the tool that bears on the question in front of them. Second, we grade every tool on an explicit five-tier translational-readiness scale, which separates regulator-cleared products from internally validated research code. Third, we close with a numbered agenda that names the datasets, trials, and consortia the field needs over the next three years.
To keep the mentioned grading legible, we apply one Clinical Readiness Level scale (CRL-AML) throughout the tables:
CRL-AML 1, Discovery: single-center, internal cross-validation only.
CRL-AML 2, Internal validation: held-out test set within the source cohort.
CRL-AML 3, External validation: independent cohort(s) from different institutions.
CRL-AML 4, Prospective evaluation: real-time clinical use studied prospectively, regardless of randomisation.
CRL-AML 5, Regulatory authorization/clearance: FDA, EMA, or equivalent authorization for clinical use.
The CRL-AML scale is an explicit adaptation of the NASA Technology Readiness Level concept and of the FDA Total Product Lifecycle framework for Software as a Medical Device, recast to grade clinical rather than purely technological maturity; it is intended as an organizing heuristic and has not been formally validated (§8.3).
The review then follows a single patient's trajectory, from clinical suspicion through diagnosis, risk stratification, treatment, response monitoring, and transplant. At each step we ask the same three questions: what AI can contribute today, how strong the evidence behind it is, and what stands between a working model and the clinic. We end by looking forward, to foundation models trained on millions of cells, digital twins that simulate disease progression, and reinforcement-learning systems that could adjust therapy in real time as clonal dynamics shift.
Throughout this review, ML denotes the subset of AI methods that learn from data; deep learning denotes multilayer neural networks, including convolutional networks for images and transformers for sequence and language; and large language models denote transformer-based systems trained on text. Much of the clinical literature still uses "AI" and "ML" interchangeably, and a brief summary table of terms is given in the appendix.
Search strategy and scope. We searched PubMed, bioRxiv, and Crossref for English-language records published between January 2015 and April 2026, pairing disease terms ("acute myeloid leukemia," "AML," "myelodysplastic syndromes," "MDS-AML," "acute promyelocytic leukemia") with method terms ("artificial intelligence," "machine learning," "deep learning," "convolutional neural network," "transformer," "large language model," "foundation model," "graph neural network," "reinforcement learning," "federated learning," "digital twin"). We then hand-screened the reference lists of the three most recent AI-in-hematology and AI-in-AML reviews [6,7,8], the FDA AI/ML medical-device taxonomy [11], and the 2025 ELN-DAVID MRD consensus. We included original primary research, methodological work using AML or MDS-AML data, regulatory filings, and AML-relevant consensus statements; we excluded pure computer-science conference papers without clinical application, non-AML disease foci (except where directly transferable, such as LLMs for oncology tumor boards), and editorials. The evidence is current to mid-April 2026, and given the pace of the field every claim should be read as a snapshot. Clinical-readiness grades for the externally validated and clinical-stage tools are consolidated in Table 4 (Section 7.7).
2. AI in AML Diagnosis
Diagnosis is where AI in AML has advanced furthest. Here the first regulatory clearances have arrived, the training datasets are largest, and the path into routine practice is shortest. Three modalities dominate the work: digital morphology, flow cytometry, and molecular classification.
2.1. Digital Morphology and Automated Blast Detection
Reading a bone-marrow aspirate is among the oldest tasks in hematology, and among the most subjective. Agreement between observers on manual blast counts is moderate at best, with κ values near 0.5 in published comparisons [12]. The consequences are not cosmetic: under ICC 2022 the 10% and 20% blast thresholds separate MDS, MDS/AML, and AML, so a disagreement near a cutoff can change both the diagnosis and the treatment. The WHO 5th edition drops the 20% threshold for AML defined by genetic abnormalities (NPM1, KMT2A-r, core-binding-factor fusions), yet the cutoff still governs the remaining categories. Little wonder that morphology became an early target for automation.
Convolutional neural networks handle blast detection well, and the evidence is now substantial. A systematic review and meta-analysis of AI for AML detection from peripheral-blood images reported pooled accuracy above 0.95 [13], and broader surveys of laboratory hematology and hematopathology document similar gains in cell classification, malignancy detection, and integration with the full MICM (morphology, immunophenotype, cytogenetics, molecular) workup [14,15,16,17]. The more interesting models look past detection to genotype. Eckardt and colleagues trained a fully automated smear pipeline on 1,251 AML patients and 236 healthy controls, reaching AUC 0.97 for AML versus normal and AUC 0.92 (95% CI 0.88–0.96) for NPM1 status inferred from morphology alone, a same-day triage signal that arrives before confirmatory sequencing [18,19]. A later study pushed further with annotation-free multiple-instance learning on whole-slide images, reaching AUCs of 0.90 ± 0.08 for NPM1 and 0.80 ± 0.10 for FLT3-ITD on a held-out set of 116 slides drawn from 572 cases [20]. With prospective validation, such tools could return preliminary molecular information within hours of biopsy; they cannot, however, replace the molecular workup that AML definition and later MRD monitoring require.
Generative modeling adds a complementary capability, exemplified by CytoDiffusion, a diffusion model trained on more than 500,000 peripheral-blood smear images that classifies cells more accurately than trained hematologists and produces synthetic images that experienced observers could not tell from real ones in a blinded Turing test [21]. The same group released the largest public blood-smear dataset to date, now available for benchmarking and model development.
Regulatory approval has arrived in two steps. In April 2024, Scopio Labs received the first De Novo FDA authorization for an all-digital bone-marrow aspirate platform, the X100 Full-Field Bone Marrow Aspirate Application (DEN230034), creating a regulatory category for hematology AI where none had existed [22]. A 510(k) clearance in July 2025 extended it to red-cell morphology across 23 parameters and platelet assessment, and CellaVision holds clearances for AI-assisted peripheral-blood analysis in major markets [22]. Narrow as they are, these authorizations prove that a pathway exists, even if hematology remains a small minority within the broader FDA AI/ML device landscape [11].
One caveat runs under all of these results, namely that most were built on curated single-institution data, with standardized staining and clean imaging. How they behave in a routine laboratory, with different stain batches, older microscopes, and uneven slide preparation, is largely untested. External validation outside European adult cohorts is rare, and pediatric AML is almost absent. The AUCs reported on research cohorts therefore very likely overstate what these tools will achieve on unselected clinical material.
Of the diagnostic modalities, digital morphology is closest to routine use. Two vendors have cleared the regulatory path, and the input, a stained slide, is something every laboratory already produces. Whether mutation prediction from images actually shortens time to treatment will depend on how pathology workflows absorb these tools; for now that is a plausible promise rather than a demonstrated one.
2.2. AI in Flow Cytometry
Multiparametric flow cytometry (MFC) sits at the center of AML diagnosis and MRD assessment, and it is also among the most operator-dependent assays in the laboratory. Manual gating, the work of carving out cell populations in high-dimensional space, demands extensive training, takes 15 minutes or more per case, and gives results that drift between operators, instruments, and centers [23]. The European LeukemiaNet has issued consensus recommendations to standardize MFC-based MRD assessment [24,25].
Even so, residual variability still matters clinically: it limits the comparison of MRD data across trials and constrains the use of flow-based endpoints in regulatory submissions.
AI for flow cytometry divides into two tasks: diagnostic classification and MRD assessment. For classification, Cheng and colleagues trained a deep-learning model on acute-leukemia orientation-tube data, reaching 94.6% sensitivity for AML and 98.2% for B-ALL [26], and a more recent cross-institutional framework classified AML across heterogeneous panel designs at 98.15% accuracy and AUC 0.998 [27]. That second result matters because protocol variability between centers is one of the main obstacles to generalizable flow-cytometry AI, and cross-protocol performance meets it directly. Lewis and colleagues went a step further, showing that a model trained on flow data can both diagnose AML and predict its molecular subtype, a fast proxy for genomic characterization [28].
For MRD quantification, the CCADDAS pipeline, built for B-ALL, cut analysis time from about 15 minutes to 1 minute per case while reaching 100% positive agreement at thresholds of 0.01% or higher [29]. Its AML validation is still pending, since residual AML is immunophenotypically more varied, but the gains in speed and consistency carry over directly. A parallel from CLL makes the broader point: applied to multiparametric flow data, the ALPODS algorithm defined immune-cell populations that predicted outcome better (AUC 0.95) than the standard CLL-IPI (AUC 0.78) [30], extracting signal that conventional gating leaves untouched.
No AI flow-cytometry tool has yet earned FDA authorization, and AML explains much of the delay: its immunophenotype shifts between diagnosis and relapse, its blasts mimic regenerating normal progenitors, and residual cells near MRD-negative thresholds are vanishingly rare. The clinical incentive is nonetheless real, because many community laboratories cannot perform high-quality flow MRD, and a validated automated pipeline would carry that capability beyond the academic centers that now hold it.
2.3. Genomic and Molecular Classification
Molecular classification poses a coordination problem of its own. Under both WHO and ICC 2022, assigning an AML diagnosis means integrating morphologic, immunophenotypic, cytogenetic, and molecular data through a multistep process that spans hematopathology, cytogenetics, and molecular laboratories [1,2].
Several groups have asked whether ML can shoulder this integration. The Unsupervised AML Multi-Omics Classification System (UAMOCS) combines genomic, methylation, and transcriptomic data into three subtypes whose prognostic profiles track the ELN 2022 scheme [31]. Others have used specific mutations and variant allele frequencies to refine the ELN 2022 categories themselves, recalibrating the existing system rather than supplanting it [32].
The most clinically suggestive results come from cheap, universally available inputs. An international validation of an AI algorithm predicting acute-leukemia subtype from the complete blood count and basic biochemistry reached AUC 0.94 for AML, 0.98 for acute promyelocytic leukemia (APL), and 0.84 for ALL across 6,206 patients in 20 centers [33]. These figures, however, hold only at a confidence cutoff that left 71–93% of cases unclassified; without it the real-world accuracy was lower (AUROC 0.84 for AML). Should it hold across varied settings, such a model could act as a rapid triage step, flagging a probable subtype before specialized testing returns. APL is the sharpest case, where any delay carries immediate mortality risk, and graph neural networks have been explored there for fast subtype identification [34].
Anticipating transformation to AML is a related and unsolved problem. In chronic myelomonocytic leukemia, a machine-learning study combined hierarchical clustering with cooperative-mutation analysis to find co-mutation patterns that predict blast transformation, among them NPM1, NRAS+SETBP1, and ASXL1+RUNX1, and validated them in an independent Italian cohort [35]; conventional scoring had missed these interactions. Sequencing has meanwhile widened the molecular target itself, adding fusion transcripts, splicing variants, and non-coding RNA alterations to the features available for ML-based classification [36].
A different route to rapid classification comes from long-read sequencing itself. Because Oxford Nanopore devices call base modifications directly and stream reads in real time, a methylation profile builds during the run, and a classifier trained on array-based atlases can read it as it accrues. The approach was proven in neuro-oncology, where the Sturgeon network classified central-nervous-system tumors intraoperatively from sparse nanopore data in about 40 minutes [37], and unified assays such as ROBIN now add copy number, single-nucleotide variants, and fusions in one same-day workflow [38]. Acute leukemia has followed: MARLIN classified AML, B-ALL, and T-ALL from low-pass nanopore data against a 2,540-sample, 38-class reference, returning subtype calls on prospective cases within about two hours [39], while the Acute Leukemia Methylome Atlas (ALMA) added a 38-CpG model predicting five-year survival in independent pediatric and adult cohorts [40]. Methylation profiling thus offers what morphology and flow cannot, a single assay that assigns subtype and estimates outcome fast enough to inform the first treatment decision. The limits are real: low-pass nanopore gives sparse, largely binary methylation, and rare karyotypes stay undersampled in the training atlases; both classifiers remain externally validated but not yet prospective (CRL-AML 3).
None of these tools has reached the clinic, and AML raises a specific concern. The move from WHO 2016 to WHO/ICC 2022 already reclassified a meaningful share of patients, so layering an AI estimate on top adds a source of variability that current guidelines do not address. The upside is equally concrete: where the full WHO/ICC workup cannot be assembled quickly, in community hospitals, resource-limited settings, or emergencies that force an early decision, an AI approximation from routine data could remove days from the time to first treatment.
Table 1.
Representative AI/ML tools in AML diagnosis, by modality and evidence maturity. CRL-AML column (clinical-readiness class): 5 = regulatory clearance; 3 = external validation in an independent cohort; 2 = held-out or partial internal validation; 1 = internal cross-validation only; n/a = not reported.
Table 1.
Representative AI/ML tools in AML diagnosis, by modality and evidence maturity. CRL-AML column (clinical-readiness class): 5 = regulatory clearance; 3 = external validation in an independent cohort; 2 = held-out or partial internal validation; 1 = internal cross-validation only; n/a = not reported.
| Modality | Tool/Reference | Task | Key metric | Cohort | CRL-AML | Regulatory status |
|---|---|---|---|---|---|---|
| Digital morphology | Eckardt 2022 [18] | AML detection + NPM1 from BM smears | AUC 0.97/0.92 | 1,251 AML + 236 controls (NPM1 subset n=408) | 3 | Research |
| Digital morphology | Kockwelp 2024 [19] | Therapy-relevant genetics from BM smears | — | Single center | 2 | Research |
| Digital morphology | Wei 2025 [20] | NPM1/FLT3-ITD from WSI (annotation-free MIL) | AUC 0.90/0.80 | 572 WSI single cohort | 2 | Research |
| Digital morphology | CytoDiffusion [21] | Generative blood-cell classification | Turing-level vs 10 experts | >500k images | 3 | Research |
| Digital pathology | Scopio Labs FF-BMA [22] | Full-field BM aspirate digital analysis | — | — | 5 | FDA De Novo (Apr 2024) + 510(k) (Jul 2025) |
| Digital pathology | CellaVision [22] | Peripheral blood differential | — | — | 5 | FDA-cleared (multiple) |
| Flow cytometry | Cheng 2024 [26] | AML/B-ALL detection | 94.6 %/98.2 % sens | Single center | 1 | Research |
| Flow cytometry | Wang 2025 [27] | Cross-institute AML classification | 98.15 % acc, AUC 0.998 | Multi-center, heterogeneous panels | 3 | Research |
| Flow cytometry | Lewis 2024 [28] | AML diagnosis + molecular subtypes | — | Single center | 1 | Research |
| Molecular classification | UAMOCS [31] | Multi-omics AML subgroups, ELN-aligned | 3 classes | Pooled datasets | 2 | Research |
| Molecular classification | Tazi 2022 [32] | Unified molecular classification + risk | 16 classes covering 100 % pts | 3,653 pts | 3 | Research |
| Molecular classification | Turki 2026 [33] | Acute leukemia subtype from routine labs | AUC 0.94 AML/0.98 APL/0.84 ALL | 6,206 pts, 20 centers | 3 | Research |
| Molecular classification | MARLIN [39] | Nanopore methylation subtype (AML/B-ALL/T-ALL) | 25/26 concordant; 5/5 prospective | 2,540 reference; prospective real-time | 3 | Research |
| Molecular classification | ALMA [40] | Methylation AML subtype + 5-yr survival | 38-CpG survival model | 3,314 across 11 cohorts | 3 | Research |
3. AI in AML Risk Stratification and Prognosis
3.1. Machine Learning Models Extending the ELN 2024 Framework
The ELN 2022 stratification, updated in 2024 for patients on less-intensive therapy, sorts AML into favorable, intermediate, and adverse groups from cytogenetics and a lengthening list of molecular markers [2,3]. The scheme is useful but coarse. Two patients in the same group can follow very different courses, and three fixed bins capture neither the interactions between mutations nor the fact that risk is continuous.
The first thoroughly validated alternative was the knowledge-bank model of Gerstung and colleagues, who joined multistage Cox modeling to hierarchical Bayesian shrinkage and integrated mutations, cytogenetics, age, and clinical features across 1,540 patients [41]. Its individualized survival estimates outperformed the categorical ELN groups and have anchored the field ever since. Eckardt and colleagues then made the case for context: training age-stratified models on 3,062 patients, they showed that a mutation's prognostic weight shifts markedly with age [42]. A lesion that signals good prognosis in a young adult may be neutral, or adverse, in an older one. Categorical systems cannot express this, because they assign each mutation a single fixed value regardless of the patient who carries it.
Later work has extended this logic to the patients ELN serves least well. Beat-AML 2024 examined 595 adults aged 60 or older on lower-intensity therapy, where IDH2 emerged as an independent favorable factor and KRAS, MLL2, and TP53 as adverse; a combined "Beat-AML mutation score" defined groups with two-year survival of 48%, 33%, and 11% [43]. The DATAML registry applied multilayer perceptrons to 3,687 patients, reporting time-dependent accuracy of 68.5% for survival under intensive chemotherapy and 62.1% under azacitidine [44], figures that are clinically useful yet a reminder of how much individual-level uncertainty remains even with large models on large real-world data. A supervised model trained on 1,383 patients predicted complete remission at AUC 0.77–0.86, with U2AF1, IKZF1, and bZIP-CEBPA carrying weight, and held at AUC 0.71–0.80 on a further 664 patients [45].
Machine-learning clustering of older AML went further still, resolving nine genomic subtypes whose survival depended on whether treatment was intensive or not [46]. The implication is that optimal stratification may have to be treatment-specific rather than universal, a conclusion guidelines have yet to absorb.
A consistent picture emerges across these studies: ML beats categorical ELN stratification, modestly but reliably, with discrimination usually in the 0.7–0.8 AUC or C-index range and at least matching ELN 2022 within each study's own validation cohort [41,42,43,44,45]. Two caveats temper this optimism, however: none of these models has been tested prospectively, and most were trained on intensively treated patients, whereas the growing venetoclax–azacitidine population, older and biologically distinct, is only beginning to enter the training data [47,48]. For now the models supplement existing risk stratification rather than replace it.
3.2. Transcriptomic and Multi-Omic Approaches
Gene expression carries prognostic information that mutation analysis only partly captures. The LSC17 score, a 17-gene signature that approximates leukemia stem-cell burden, was among the first transcriptomic predictors developed and externally validated for AML, and it has held up across multiple cohorts [49]. Where mutation-based models read the genetic lesion, LSC17 reads the leukemia's transcriptional state, and that distinction has teeth: patients with identical mutations can diverge sharply at the expression level, and the divergence predicts outcome.
Newer work builds outward from expression alone in several directions. Silva and colleagues trained RNA-seq models on k-mer features, short nucleotide subsequences, exceeding 90% accuracy in risk prediction [50]; integrative analyses that fuse expression with DNA methylation through network-based feature selection improve on single-omic models [51]; and a multicenter study of epigenetic subtypes resolved two molecularly distinct groups with C-indices of 0.72–0.78 across validation cohorts [52].
Integration now reaches proteomics and metabolomics as well [53,54]. A 2025 spatial-proteomic study profiled AML marrow by high-plex immunofluorescence and mass-cytometry imaging and found tertiary lymphoid-like aggregates whose signatures were prognostic in an independent 480-patient transcriptomic cohort [55]. Methodologically, graph neural networks and tensor factorization are beginning to connect single-cell data to clinical prediction, handling modality prediction, matching, and joint embedding within one framework [53,54,56].
The trajectory is clear: risk stratification is moving from categorical groups toward continuous, individualized probabilities drawn from several data layers. What limits it now is not the algorithms but the assays. Multi-omic profiling needs RNA-seq, methylation arrays, proteomics, and increasingly metabolomics, none of which most clinical laboratories run routinely. The analytic methods exist; the infrastructure to generate the data at the point of care does not. NGS panels became routine once their cost fell far enough over the past decade, and whether multi-omic profiling follows the same curve will decide when these models reach the clinic.
4. AI in AML Treatment Decision Support
4.1. Drug Response Prediction
For most patients the first major treatment decision, between intensive chemotherapy, venetoclax–azacitidine, and a targeted agent, turns on age, fitness, and a few molecular markers. The approach is reasonable on present evidence but blunt: two patients sharing an ELN category and a frontline regimen can still diverge widely, and we have no dependable way to foresee that at the individual level.
The most advanced attempt is the RF8 model of Jin and colleagues for predicting response to venetoclax–azacitidine (VEN/AZA) [57]. Drawing on ex vivo drug-sensitivity testing, transcriptomic profiling, and CRISPR screening, they distilled eight genes whose expression tracks VEN/AZA response, then built a random forest validated across four independent cohorts totaling 498 patients, with scores rising near-monotonically with both response probability and survival. Few AML treatment-prediction models have been externally validated in several datasets, which makes RF8 the most plausible near-term candidate for prospective testing. Three limitations keep it short of the clinic: it rests on only eight genes, it needs RNA-seq input that not every center can provide, and all four validation cohorts were retrospective.
Other approaches widen the lens by representing the links between molecular features and drug sensitivity as networks. These knowledge-graph models flagged expression signatures, among them FGD4-MIR4519, NPC2-GATA2, and BCL2-NFKB2, that predict ex vivo venetoclax response [58]. NetAML scaled the idea up, building 87 sensitivity models across 87 compounds in primary samples from 520 patients to yield a pan-drug atlas for AML [59]. Even mutation patterns alone, without expression, can predict sensitivity in some settings [60].
These predictors do not stand apart from biology; they increasingly fold it in. Monocytic differentiation, marked by loss of BCL2 and a shift to MCL1-dependent survival, is enriched among patients with primary venetoclax resistance, though response is not uniformly lost in this group [61], and post-hoc analyses of VIALE-A show that TP53 and RAS-pathway co-mutations independently shorten responses to venetoclax–azacitidine [5]. ML models are beginning to incorporate such readouts rather than compete with them [52,55,57]. What is still missing is comparison: head-to-head evaluations of VEN/AZA predictors are scarce, and no prospective study has randomized patients to ML-based selection against standard criteria. The logical next step is a pragmatic trial that embeds RF8, or an equally validated predictor, as real-time decision support.
The question hematologists face most often, intensive 7+3 versus venetoclax–azacitidine, has itself been addressed by ML decision-support tools, though these remain exploratory and untested prospectively [62]. At the far end of the treatment arc, MM-AI-AML couples mechanistic modeling of myelopoiesis with AI to predict the severity of chemotherapy-induced myelosuppression before treatment starts, at AUC 0.85 internally and 0.78 externally [63]. Toxicity prediction tends to be neglected next to efficacy, yet for older or frail patients, avoiding severe myelosuppression can matter as much as reaching remission.
Graph neural networks open a different line of inquiry, exemplified by CGMega, an explainable graph-attention network that organized 396 candidate AML genes into functional modules, yielding both predictions and mechanistic hypotheses about drug targets [64]. PDGrapher, a causally inspired framework, predicts combinatorial perturbations, sets of targets able to reverse a disease transcriptional state, with applications in drug repurposing and combination design [65,66].
None of these tools is in clinical use. Ex vivo sensitivity is an unreliable guide to in vivo response, and even RF8 awaits prospective testing. The trials that would settle the matter are overdue.
4.2. Large Language Models and Generative AI
Large language models are being tested for a different kind of work, not pattern recognition in molecular data but the reasoning that surrounds a clinical decision: synthesizing the literature, preparing cases, and matching patients to trials.
Schmutz and colleagues put ChatGPT-4.0 to 20 real-world molecular tumor board (MTB) cases across breast cancer, glioblastoma, colorectal cancer, and rare tumors [67]; the cohort was oncologic rather than AML-specific, but the lesson carries to hematologic boards. The model proposed more therapeutic options per case than the expert panel (median 3 vs 1; p = 0.005) and needed far less preparation time (median 15.2 vs 34.7 minutes; p < 0.001). It also exposed the characteristic weakness of current LLMs: the supporting evidence varied in quality, and some recommendations were poorly sourced, the price of a system that writes plausible and incorrect statements with equal fluency.
Trial matching shows the technology at its most useful. TrialGPT, built by the NIH National Library of Medicine and the National Cancer Institute, recalled more than 90% of relevant trials while screening under 6% of the candidate pool, with 87.3% accuracy at the criterion level and a 42.6% cut in screening time [68]. For AML centers running active trial portfolios, that speaks to a genuine bottleneck: under-enrollment that stems from the sheer difficulty of matching complex patients to the right protocol quickly.
Head-to-head comparisons temper the enthusiasm: tested against one another and against expert judgment for clinical-evidence summarization in hematologic malignancies, Claude, GPT-4, Gemini, and Llama show performance that varies by task and model, with no system uniformly ahead [69,70]. Broader scoping reviews in oncology reach the same verdict: the technology is advancing fast, but validated clinical deployments remain scarce [71,72,73].
The boundary is worth stating plainly: LLMs do not make treatment decisions. A model that offers an option without grasping a patient's comorbidities, prior toxicities, or circumstances is generating suggestions, not judgments. Their value lies in throughput, since a busy academic board reviews dozens of cases a week and halving preparation time compounds quickly [67]. Governance is catching up: early frameworks call for mandatory human review of all AI-generated clinical content, audit trails for LLM-assisted recommendations, and clear disclosure whenever these tools touch the medical record [71,72,73].
Table 2.
Representative ML models for AML risk stratification and drug-response prediction. The CRL-AML column uses the same clinical-readiness classes as Table 1.
Table 2.
Representative ML models for AML risk stratification and drug-response prediction. The CRL-AML column uses the same clinical-readiness classes as Table 1.
| Model/Reference | Task | Features | Cohort | CRL-AML | vs. ELN/standard | Deployment |
|---|---|---|---|---|---|---|
| Gerstung Knowledge Bank [41] | Individualized survival | Mutations + cytogenetics + clinical | n = 1,540 | 3 | Outperforms categorical ELN | Research tool |
| Eckardt age-stratified [42] | Age-specific molecular weighting | 14+ mutations × age bins | n = 3,062 | 3 | Identifies divergent age effects | Research |
| Beat-AML 2024 [43] | ELN refinement for lower-intensity Tx | ELN + IDH2, KRAS, MLL2, TP53 | Beat-AML cohort | 3 | Refines ELN 2024 | Research |
| DATAML [44] | OS under IT vs AZA | Clinical + mutations | n = 3,687 | 1 | Comparable to ELN | Research |
| Eckardt CR-supervised [45] | Complete remission prediction | Mutations + clinical | n = 1,383 | 2 | AUC 0.77–0.86 | Research |
| Silva transcriptomic [50] | Risk from k-mer RNA features | RNA-seq | n = 404 | 2 | > 90 % accuracy | Research |
| RF8 (Jin) [57] | VEN/AZA response | 8-gene random forest | n = 498 pooled | 3 | Outperforms mutation-only | Closest to clinical-ready |
| Knowledge graphs (Qin) [58] | Drug response | Network + expression | — | 2 | n/a | Research |
| NetAML [59] | 87-drug sensitivity atlas | Network + molecular | n = 520 ex vivo | 2 | n/a | Research |
| Qin mutation patterns [60] | Drug sensitivity from mutations | Mutational profile | Clinical cohort | 2 | n/a | Research |
| Islam 7+3 vs VEN/AZA [62] | Treatment selection | Clinical + molecular | Retrospective | 1 | — | Exploratory |
| MM-AI-AML [63] | Myelosuppression severity | 51 clinical features + TabNet | n = 479 + 900 virtual | 3 | No standard comparator | Research |
5. AI in AML Response Assessment and MRD
5.1. AI-Assisted Flow Cytometry MRD
Automated MRD detection is not new and it sets out to make MFC-based MRD faster and more reproducible [74]. Support vector machines were applied to the problem as early as 2016 [75], and by 2018 a clinically validated algorithm matched expert performance in AML and MDS [76]. The more recent MAGIC-DR, an XGBoost pipeline, reached a cross-validation median AUC of 0.97 on its training set (98 diagnostic AML plus 30 MRD-negative samples) and agreed encouragingly with expert gating on a separate 25-sample test set, though that test set is too small to be conclusive [77]. CCADDAS, discussed in Section 2.2, has set the B-ALL benchmark, but its AML validation is not yet complete [29].
Mocking and colleagues, surveying computational immunophenotypic MRD in AML, named the core problem: too many algorithms and no standard [74]. Tools differ in framework, gating strategy, and positivity threshold, and no multicenter study has yet compared them on a common sample set. The parallel is with the early days of molecular MRD, where the absence of an assay standard delayed MRD's entry into trial endpoints by years.
A cloud-based, software-agnostic pipeline could bring reference-quality MRD analysis to the many laboratories that lack the expertise today. In AML trials, where MRD is increasingly used as a surrogate endpoint for approval, standardized automated reading would also strengthen the regulatory case. The 2025 ELN-DAVID consensus made harmonization its priority, issuing 56 recommendations across technical standards and clinical interpretation, and automated gating is one obvious route to implementing them [25].
5.2. Molecular MRD and Clonal Tracking
Sequencing has added a second dimension to AML MRD. Where flow cytometry counts residual blasts by immunophenotype, error-corrected NGS detects persistent somatic mutations down to variant allele frequencies of 0.01–0.1%, depending on the platform (duplex sequencing, unique molecular identifiers). Jongen-Lavrencic and colleagues established that molecular MRD at complete remission independently predicts relapse and survival [78], and Morita and colleagues refined the reading, showing that which mutations persist, and how fast they clear, carries prognostic weight beyond a simple positive-or-negative call [79].
The interpretation of molecular MRD is subtler than it first appears, because not all persistent mutations carry the same weight. Jongen-Lavrencic and colleagues showed that persistent DNMT3A, TET2, and ASXL1 (DTA) mutations at remission mark age-related clonal hematopoiesis and do not predict relapse, whereas persistent non-DTA mutations remain meaningful [78]. Telling true leukemic persistence from background clonal hematopoiesis therefore hinges on which mutations persist, not merely on whether any do, and this is where computation enters. Models trained on mutational patterns reclassify AML past the old primary-versus-secondary divide, reading genomic signatures that betray the disease's clonal origin [80].
A more ambitious line asks whether clonal architecture itself, the relative abundance, order, and subclonal structure of mutations rather than their mere presence, predicts outcome. Analyzing 2,829 patients, Benard and colleagues found branched clonal evolution to carry a better prognosis than linear accumulation, and tied specific subclones to specific signals: subclonal TP53 abundance to white-cell count and blast percentage, subclonal NRAS to ex vivo drug sensitivity, and subclonal IDH1 to both, none of which binary mutation status can see [81]. The clinical weight of hierarchy runs deeper still. Shlush and colleagues traced relapse to pre-leukemic stem cells carrying founder mutations, locating not just who relapses but from which compartment [82], and Hourigan and colleagues showed that conditioning intensity before transplant interacts with genomic MRD to shape outcome, a reminder that monitoring data only mean something in their therapeutic context [83,84].
The computational toolkit for this reconstruction is well developed, beginning with Bayesian methods such as PyClone and its scalable successor PyClone-VI, which infer subclones from bulk variant allele frequencies, while SciClone, PhyloWGS, and newer ML approaches extend the reconstruction to phylogeny and to tumor evolution tracked across sites and timepoints [80,81,85]. At single-cell resolution, Tapestri (used by Morita and colleagues to map AML clonal evolution at high throughput [86]) and GoT (which links genotype to transcriptome in the same cell [87]) supply the ground-truth architecture that bulk methods can only approximate. Combining population genetics with statistical inference, ML has sharpened both the accuracy and the scale of subclonal reconstruction [80,81,85].
These tools matter most against the resistance patterns that modern therapy produces. Venetoclax-based regimens have made the point vividly: response is often clonally heterogeneous, as sensitive clones collapse under selection while pre-existing TP53 or RAS/MAPK subclones, monocytic populations, and emergent FLT3 kinase-domain mutations expand to drive relapse [5,61]. The natural role for AI here is to follow these trajectories forward rather than reconstruct them after relapse.
What does not yet exist is the integration. No tool today takes serial molecular MRD across timepoints, joins it to clonal-evolution modeling, and guides decisions in real time, even though the raw material is in hand. Many patients now have sequential NGS from diagnosis through transplant, flow-based MRD at several timepoints, and, in research cohorts, single-cell clonal maps, and the reconstruction methods are mature. These pieces have simply never been connected into one system that updates individualized risk as each new result arrives.
The resulting clinical dilemma is concrete. A patient in morphological remission with a rising TP53 variant allele frequency and a new NRAS subclone presents a question that guidelines do not answer: is this the start of relapse from a resistant clone about to dominate, or clonal hematopoiesis drifting without consequence? Today the answer comes from judgment and time, from waiting for the next count, the next marrow, the next flow timepoint. A tool that read the patient's full clonal history against treatment context could estimate the probability of imminent relapse and, crucially, name the subclone driving it, because a TP53-driven relapse demands a different response from one driven by FLT3-ITD or IDH2. Building it is the operational goal behind the adaptive-dosing framework in Section 7.6.
5.3. Single-Cell and Spatial Technologies
Single-cell RNA sequencing has mapped the hierarchy of AML clones, picked out rare leukemia stem-cell populations, and resolved the marrow immune microenvironment cell by cell [88,89]. Profiling the same patients through chemotherapy reveals how subpopulations expand or contract under pressure [90], and the GoT approach ties genotype to transcriptional state within a single cell, joining clonal identity to cell behavior [87].
Spatial transcriptomics adds a geographic dimension. In marrow from patients on pembrolizumab and decitabine, single-cell spatial analysis found immune cells clustering near leukemic ones after immunotherapy [91], and spatial proteomics identified tertiary lymphoid structures whose signatures predicted outcome in 480 patients [55,92].
For now these technologies describe rather than decide. No clinical AI tool runs on single-cell or spatial data, and the computational cost rules out routine use. As the assays grow cheaper and foundation models mature (Section 7.6), it should become possible to track not only whether residual disease persists but where it sits in the niche and how the surrounding microenvironment is arranged. That resolution is a long-term ambition, not a near-term deliverable.
Table 3.
AI approaches for MRD assessment, clonal architecture, and single-cell/spatial AML monitoring.
Table 3.
AI approaches for MRD assessment, clonal architecture, and single-cell/spatial AML monitoring.
| Tool/Reference | Assay | Task | Metric | Cohort | Status |
|---|---|---|---|---|---|
| Ni 2016 (SVM MRD) [75] | Flow | First automated AML MRD gating | Expert-level | Single center | Historical |
| Ko 2018 [76] | Flow | Validated AML/MDS MRD | Concordance with experts | Multi-center | Validated |
| MAGIC-DR [77] | Flow | Interpretable AML MRD (XGBoost) | CV AUC 0.97 (n=98 AML + 30 ctrl); separate n=25 test set | Single center | Research |
| CCADDAS (B-ALL) [29] | Flow | MRD pipeline | 15 min → 1 min; 100 % agreement ≥ 0.01 % | Multi-center | Research |
| Mocking review [74] | Flow | Computational MRD landscape & standardisation gap | — | n/a | Review |
| Jongen-Lavrencic NEJM [78] | NGS | Molecular MRD at CR as relapse predictor | Independent OS/EFS predictor | Multicenter | Reference standard |
| Morita mutation clearance [79] | NGS | Kinetics of somatic mutation clearance | — | — | Foundational |
| Awada ML genomics [80] | NGS | ML subclassification beyond primary/secondary AML | — | — | Research |
| Benard clonal architecture [81] | Bulk NGS | Clonal outcomes + drug sensitivity | Branching > linear survival | n = 2,829 | Research |
| Shlush [82] | NGS | Pre-leukemic HSC origins of relapse | — | — | Foundational |
| Morita single-cell Tapestri [86] | scDNA | High-throughput clonal architecture | — | AML pts | Research |
| Nam GoT [87] | scRNA + genotype | Clonal transcriptomes | — | Multiple MPN/AML cohorts | Research |
| Naldini longitudinal [90] | scRNA | Chemotherapy response dynamics | — | AML pts | Research |
| Gui spatial [91] | Spatial | Niche remodeling on immunotherapy | — | AML on pembro + DAC | Research |
| Ly spatial proteomics [55] | Spatial | TLS signatures predictive of outcome | Validated in n = 480 | AML pts | Research |
6. AI in Transplant and Cellular Therapy for AML
6.1. Predicting Transplant Outcomes and Post-HCT Relapse
The transplant decision in AML balances relapse risk against treatment-related mortality, weighing disease biology, donor characteristics, fitness, and MRD status together. The established scores, the Hematopoietic Cell Transplantation–Comorbidity Index (HCT-CI), the Disease Risk Index (DRI), and the EBMT score, were built with conventional statistics and capture only part of that complexity.
Masurekar and colleagues compared random survival forests, elastic-net Cox regression, and standard Cox models in 2,253 patients undergoing allogeneic hematopoietic cell transplantation (allo-HCT), with external validation in 252 patients who had uniform MRD assessment [93]. Age over 60, MRD positivity, and adapted ELN adverse risk were the strongest predictors of poor outcome, yet the gain over conventional models was modest, which the authors read as a sign that pre-transplant variables alone may have hit a predictive ceiling. The contrast with time-updated data is instructive. The DERGA algorithm, applied to 564 consecutive adult allo-HCT recipients (AML among several underlying diseases), reached 93.26% accuracy for survivorship status by combining pre- and post-transplant features [94], and in a pediatric cohort, naïve-Bayes models using peri-transplant data from 30 days before to 30 days after transplant outperformed baseline-only predictors [95], pointing toward dynamic inputs as the way forward.
Relapse remains the most common cause of transplant failure, and several groups have aimed AI at predicting it. Texture and morphology features extracted from myeloblast chromatin in post-HCT aspirates predicted relapse at AUC 0.71, evidence that image-based biomarkers carry post-transplant signal [96]. A dedicated model for relapse after haploidentical HCT in AML has appeared [97], and ML has mapped relapse patterns in pediatric transplant recipients with high specificity, albeit on small samples [98,99]. The most instructive example comes from another disease: Hernández-Boluda and colleagues built a random survival forest for transplant in myelofibrosis and released it as a free web calculator [100]. The biology differs, but the delivery model, a peer-reviewed algorithm paired with a maintained public tool, is exactly the template AML transplant tools will need if they are to change practice.
The pattern is consistent: ML adds incremental value over traditional scores, but reliable individual-level prediction is beyond reach with pre-transplant data alone. The next gains will come from post-transplant dynamics, from chimerism kinetics, serial MRD, early GvHD signals, and immune reconstitution. Transplant programs already collect these data but models trained to learn from them are needed.
6.2. GvHD Prediction
Graft-versus-host disease (GvHD) is the leading cause of non-relapse mortality after allo-HCT, yet risk assessment still rests on donor matching, conditioning, and prophylaxis, all fixed at transplant and blind to the biology of the graft-host interaction as it unfolds. ML is being applied at both ends: before transplant, from clinical and genetic features, and after, from dynamic clinical data.
The pre-transplant work spans clinical, genetic, and molecular features. An ensemble of CatBoost, LightGBM, random forests, and logistic regression predicted acute GvHD at AUC 0.71–0.80 in 270 pediatric recipients, combining 53 clinical factors with 104 candidate single-nucleotide polymorphisms [101]. Targeted transcriptome profiling of pre- and post-transplant marrow yielded a 92-gene signature for acute GvHD and a 20-gene signature for survival, a molecular layer beyond clinical and HLA variables [102]. Convolutional networks have been applied as well, predicting acute GvHD from clinical features [103].
The more promising direction is longitudinal, because GvHD is a process rather than a single event. A model fed daily post-transplant laboratory values, vital signs, medications, and clinical events simply knows more than one limited to the pre-transplant snapshot. Early results bear this out: models built on longitudinal EMR data outperform baseline-only approaches for both acute and chronic GvHD [104].
Still, no GvHD model has been validated prospectively, and none yet guides prophylaxis in a trial [104]. The move from static pre-transplant scoring to continuously updated risk is no longer a technical problem but now a matter of trial-design.
6.3. CAR-T and Novel Cellular Therapies
CAR-T-cell therapy has transformed the treatment of B-cell malignancies but has gained little ground in AML, for one overriding reason: the disease offers no equivalent of CD19 [105]. Its blasts share surface antigens with normal myeloid progenitors, so every candidate target carries an on-target, off-tumor cost. Targeting CD33 produces prolonged cytopenias, as gemtuzumab ozogamicin showed long ago; CD123 constructs cause broad myeloid aplasia; CLL-1 leaves hematopoietic recovery uncertain. In several clinical programs this toxicity has capped both dose and durability [105]. Computational target discovery offers a way through. Single-cell transcriptomic atlases of leukemic and normal marrow, mined with machine-learning surfaceome classifiers, can rank antigen combinations that cover the leukemia while sparing normal progenitors, and can begin to match a patient to the construct most likely to engage their disease [106].
Outcome prediction is further along in lymphoma than in AML. In 416 patients given axicabtagene ciloleucel for diffuse large B-cell lymphoma, Wang and colleagues predicted relapse within six months from seven routine variables: age and six laboratory values (LDH, CRP, ferritin, hematocrit, platelet count, prothrombin time) [107]. Toxicity prediction has drawn similar attention, again by borrowing across diseases. Because severe cytokine release syndrome and COVID-19 share an inflammatory signature, models pretrained on COVID-19 cytokine-storm data have been transferred to predict CRS in CAR-T recipients where dedicated training data are scarce [108]. The same toolkit reaches upstream of the infusion as well, where computational-experimental pipelines nominate selective and synergistic drug combinations for relapsed AML [57] that could debulk resistant subclones before or after cellular therapy.
As CAR-T platforms against CD33, CD123, CLL-1, and dual antigens move through AML trials, the prediction frameworks built in lymphoma will have to be adapted rather than borrowed wholesale. Toxicity models are the priority, because on-target, off-tumor myelosuppression is a distinct and serious hazard with no real counterpart in B-cell disease [109].
7. Challenges, Regulatory Landscape, and Future Directions
7.1. The Validation Gap
Most tools in this review have been validated only internally, trained and tested on data from one institution, often by cross-validation rather than a true external holdout. External validation on independent cohorts is rare, and prospective validation in AML is essentially absent.
This matters because internal cross-validation flatters a model: training and test data share the same biases of demographics, treatment patterns, and laboratory protocols. An AUC of 0.90 at home can fall to 0.75 elsewhere [110]. RF8 [57] and the Gerstung knowledge bank [41] stand out precisely because genuine external validation is so rare.
The trial record makes the imbalance concrete. Against thousands of published models, a ClinicalTrials.gov analysis of AI in oncology found that registered prospective AI trials in hematology number in the single digits [110,111]. The FDA demands prospective evidence for most therapeutic interventions, and a consensus is forming that AI tools shaping clinical decisions should meet a comparable bar. Workable routes exist, adaptive designs that allow continuous model updates without sacrificing rigor, digitized workflows for efficient data capture, and pragmatic trials embedded in routine care [111], but none has been adopted at scale in AML.
The imbalance is structural: the field produces more models than it validates [111]. Until external validation becomes the default expectation for publication rather than the exception, clinicians will rightly hesitate to use these tools.
7.2. The Regulatory Landscape
Within the FDA AI/ML device taxonomy (about 19 hematology devices out of 1,016; Section 1) [11,112], radiology and cardiology dominate the cleared products, the specialties with mature imaging pipelines and large standardized datasets. Scopio Labs and CellaVision (Section 2.1) are the hematology exceptions, and beyond them, few others exist [11,113].
Hematology poses regulatory problems of its own, beginning with its inputs. Flow-cytometry panels, mutational profiles, and expression signatures lack standardized formats, and where radiology has DICOM as a universal image standard, hematology has no equivalent infrastructure. The FDA's Good Machine Learning Practice framework and Predetermined Change Control Plans apply in principle but remain untested here [112,114], and most of the research tools in this review have no regulatory pathway at all.
The field divides along a clear line between two kinds of tool. Those nearest to clinical utility, digital morphology and automated MRD gating, could plausibly win authorization within a few years. Those with the greatest potential impact, drug-response prediction, clonal-evolution tracking, and adaptive treatment optimization, act on data types and decisions for which no regulatory framework yet exists. Closing that gap will take hematologists, regulators, and developers at one table, agreeing on what counts as sufficient evidence, which endpoints fit, and how post-deployment monitoring should work.
7.3. Explainability and Clinical Trust
Accuracy alone does not earn a model a place at the bedside, because a risk score that cannot point to a reason invites skepticism however good its numbers. Explanations are therefore not a courtesy but a condition of clinical use.
SHAP and LIME have become the standard way to decompose a prediction into feature contributions [115], but feature weight is not the same as biological sense. The more persuasive explanations point to something a hematologist already recognizes. In AML, the attention-based SCEMILA framework does exactly that, flagging the diagnostic cells an expert would, atypical hypergranular promyelocytes in PML::RARA, cup-like blasts in NPM1-mutated AML, and myelomonocytic cells in CBFB::MYH11, with 89% concordance between high-attention cells and expert-labeled features [116]. That is explanation that travels to the bedside, because it links model output to morphology clinicians trust. More general solutions, combining several explainability techniques into a trust-centered framework, have also been proposed [117,118].
This distinction is not academic, because a model that leans heavily on CD33 might be reading real biology or merely a staining artifact, and only a hematologist who knows the disease can tell which. The real test of an explanation is whether it survives that scrutiny.
Explainability becomes harder as the underlying models grow more complex. A random forest built on ten clinical features can be explained in terms a clinician recognizes, but a graph neural network over multi-omic data cannot, at least not with today's tools. What would actually build confidence is causal or mechanistic explanation, the kind that ties a prediction back to known AML biology.
7.4. Data Equity, Bias, and Ethics
Most AML models are trained on data from academic centers in the United States and Western Europe. Patients from low- and middle-income countries, minorities within high-income countries, and pediatric populations are underrepresented or missing entirely. A model fit largely to European adults may not transfer to patients with different genetic backgrounds, exposures, and treatment protocols [119,120].
Equity is not only a question of who appears in the training data; it returns at the point of deployment.
Some tools are designed to help settings that lack specialist expertise, yet in practice AI may benefit most the institutions already equipped to adopt it, the large academic centers with digital pathology, comprehensive molecular profiling, and bioinformatics support, widening the gap between specialized and community care rather than closing it. Federated learning (Section 7.5) is a partial remedy, not a complete one.
Ethical questions follow close behind, and several remain unresolved: patient disclosure, institutional governance, error reporting, and liability when AI-guided treatment fails. A recent governance framework for generative AI in oncology offers retrieval-augmented generation (RAG) and human-in-the-loop workflows as practical safeguards, though neither has been tested across diverse real-world settings [121]. In the European Union, the AI Act (Regulation (EU) 2024/1689) requires transparency and human oversight for AI used in clinical decisions, even as implementation timelines for medical applications have slipped and no hematology-specific guidance yet exists [122]. These conversations need to run alongside model development, not wait until after deployment.
7.5. Federated Learning and Multi-Institutional Infrastructure
Pooling data across institutions would address small cohorts, limited diversity, and single-center bias at a stroke. In practice, privacy regulation, institutional governance, and competition make centralized sharing hard to arrange.
Federated learning offers a way around this obstacle by inverting the usual flow. Each site trains the model on its own data and shares only the resulting model updates, never the raw patient records. Early leukemia work has paired federated learning with low-rank adaptation for communication-efficient classification [123], and explainable federated transformers combining vision transformers with ClinicalBERT have been proposed for joint leukemia classification and staging across decentralized devices [124], though both are engineering demonstrations rather than validated clinical tools. The proof of principle comes from imaging, where multi-institutional consortia have shown federated training matching centralized performance in pan-cancer studies; the equivalent hematology evidence is thinner but building within big-data consortia such as the HARMONY Alliance [125].
The practical hurdles remain substantial: homomorphic encryption can slow training roughly fifteenfold, and cross-border model updates raise unsettled questions under GDPR and its counterparts. Most fundamentally, federation distributes computation but does not harmonize data: differences in quality, labeling, and panel composition persist between sites and can degrade a model even when the federation itself is sound [126,127].
Even so, federation is the most realistic route to the large, diverse datasets AML models need. Centralized repositories face legal, logistical, and competitive barriers that have proven hard to overcome even within a single country. For AML in particular, where any one center sees limited numbers and where molecular heterogeneity demands large cohorts to capture rare subtypes, multi-institutional collaboration is not optional. Federated frameworks that tolerate heterogeneous data while preserving privacy are the practical compromise between what the algorithms require and what institutions can give.
7.6. Emerging Technologies
None of the three technologies below has produced a clinical tool for AML but they are moving fast enough that practicing hematologists will meet them within a few years.
Foundation models for single-cell genomics. These models bring the idea behind large language models into cell biology. A network is first trained on tens of millions of single cells with no labels attached, so it picks up features that carry over to many tasks, and is then fine-tuned on the specific problem at hand. scGPT, CellFM, and Nicheformer were pretrained on 33, 100, and 110 million cells [128,129,130]. The hope is that one such model could be reused for AML subtype classification, drug-response prediction, and clonal-state annotation. That has been harder than expected. Used as they are, without fine-tuning on the task in front of them, these models still trail simpler tools built for the job on several benchmarks [131,132].
Digital twins. A patient-specific model that simulates disease progression and treatment response in silico, a "digital twin," has been endorsed by the NCI as a future paradigm for precision oncology [133,134]. In AML it would fold genomic profile, clonal architecture at diagnosis, MRD kinetics, drug pharmacokinetics, and clinical course into one continuously updated prediction. Synthetic-patient generation for trial enrichment has been demonstrated and could supply AI-generated control arms [135,136], but a fully dynamic, continuously learning patient model remains aspirational. Network approaches that jointly model clonal populations, the marrow microenvironment, and the immune system may provide the mathematical backbone, though no working AML prototype has yet been reported.
Reinforcement learning (RL) for adaptive therapy. Here a model learns by trial and error: it takes an action, sees what happens, and adjusts to do better over time, without being shown labeled examples. In treatment, the actions are choices about drug, dose, and timing, and the aim is lasting disease control, which suits a disease whose management has to change as the cancer evolves.
The clearest example is prostate cancer, where Gatenby and colleagues withdrew and reintroduced abiraterone on PSA thresholds instead of dosing to the continuous maximum tolerated dose [137,138]. This evolution-informed schedule delayed resistance in metastatic castration-resistant disease. The goal there is to prolong control, not to eradicate the cancer. The closest AML analogue is therefore not curative-intent induction but the lower-intensity regimens used in older or unfit patients. Azacitidine–venetoclax is the obvious case, where durable control and quality of life matter more than cure. Deep RL dosing has been tested in simulation, recently with Bayesian methods that account for pharmacokinetic variability between patients [139]. RL coupled to population PK/PD models has been tried for propofol anesthesia, warfarin, and several anticancer drugs [140], but none of this has reached AML. The obstacle is not the absence of feedback but its granularity and the cost of exploring. Marrow is already sampled in the first cycles of that regimen, and circulating blasts offer a cheap early signal when present. Yet response at the marrow and MRD level remains intermittent. Prostate adaptive therapy can be switched on and off reversibly. A policy that explored intensification in AML could not, and the harm would be irreversible. A workable system would therefore learn offline, from routinely collected trial and practice data, and stay close to standard care, with that lower-intensity setting the natural place to begin.
Yet AML fits the adaptive-therapy profile almost perfectly: measurable clonal heterogeneity at diagnosis, serial MRD that tracks the disease, clear decision points at induction, consolidation, maintenance, and transplant, and a widening menu of targeted agents to combine or sequence. Adaptive trials that update a model mid-study are technically feasible and already run elsewhere in oncology. Bringing them to AML treatment optimization is among the most underexplored opportunities in the field.
7.7. Consolidated CRL-AML 3+ registry
Most of the tools in this review sit at the lower end of clinical readiness. Table 4 gathers the exceptions, every AML AI tool here that has reached external validation (CRL-AML 3) or higher, in contrast with the large number of CRL-AML 1–2 tools in Table 1, Table 2 and Table 3.
Of the 21 entries, three are FDA-cleared digital-morphology platforms (CRL-AML 5), and none has been evaluated prospectively as real-time decision support in AML (CRL-AML 4).
8. Conclusion and a Three-Year Research Agenda
8.1. Where the Field Stands
No AI tool in AML has reached CRL-AML 4. Until one does, the claim that AI is "transforming AML care" describes a publication trajectory, not a clinical one.
By this scale, the overwhelming majority of tools reviewed here sit at CRL-AML 1–2, and only a handful, RF8, the Gerstung knowledge bank, the Turki international algorithm, and MAGIC-DR, reach CRL-AML 3. The two FDA-cleared digital-pathology platforms, Scopio and CellaVision, are graded CRL-AML 5 on regulatory status, but their predicate- and De Novo-based clearances came without a prospective, AML-specific utility trial, so regulatory status here does not amount to clinical validation. That gap, hundreds of CRL-AML 1–2 models against zero prospective evaluations, frames the next three years.
8.2. From Publication to Practice
The most useful next step is less about building new models than about testing the ones that already exist. Drug-response prediction for venetoclax–azacitidine sits closest to a prospective trial, since its output maps onto a decision hematologists make: how intensively to treat. The clearest example is RF8 [57], so far the only externally validated response predictor of its kind (§4.1), though it rests on a single 2025 study and would need independent confirmation before it could anchor a trial. The broader point holds whichever model is used: a validated response predictor should be carried into a pragmatic randomized comparison with standard ELN 2024 care [3], with survival as the endpoint. That, rather than another retrospective model, is what would move the field toward its first prospectively validated tool.
The larger opportunity lies further out, in adaptive, evolution-aware therapy. AML, with its measurable clonal heterogeneity and serial MRD, is well suited to treatment that is adjusted as the disease changes, yet by 2026 no AML-specific RL trial had been attempted. A sensible first step would be a small safety-and-feasibility study of model-guided dosing in consolidation, with stopping rules set in advance [135,136]. This would provide early evidence without allowing an algorithm to learn on patients in real time. Single-cell foundation models, by contrast, are still not ready for this: without task-specific tuning they still fall short of the supervised tools already in use [131,132], and they belong in the laboratory until benchmarks show otherwise.
None of this will matter without some shared groundwork. External validation in at least one independent cohort should become the minimum bar for any AML model that claims clinical relevance, much as CONSORT-AI and SPIRIT-AI did for trial reporting. Flow-cytometry MRD is the natural place to begin: a shared, federated reference dataset, assembled across countries and harmonized under the 2025 ELN-DAVID consensus [25], would help with both standardization [24,74] and unequal access to data [119,120]. Regulators will need to keep pace, since without an agreed FDA–EMA view of what evidence suffices for AI in MRD, risk stratification, and drug-response prediction [112,114], little beyond digital morphology will reach the clinic. The tool that would matter most to patients still does not exist: a model that reads a patient's clonal trajectory and turns it into an individual relapse risk at the points where treatment is actually decided (§5.2). Building it and testing it prospectively should be the priority for the next few years.
8.3. Limitations of This Review
This is a structured narrative review, not a PRISMA systematic review or meta-analysis: the search strategy is disclosed in §1, but study selection is not exhaustive and no quantitative pooling was performed. English-language sources are heavily weighted, leaving Chinese-, Japanese-, and Spanish-language AI-in-AML work underrepresented. The field moves quickly, so each claim is a snapshot as of mid-April 2026. The proposed CRL-AML 1–5 scale (§1) has not been formally validated against external rubrics such as the FDA's Total Product Lifecycle for SaMD or the WHO AI ethics framework, and we expect it to be refined through use.
8.4. Coda
AML has what it takes to be the disease where AI in hematology earns a clinical role: genomic complexity, measurable clonal dynamics, well-defined decision points, and a strong cooperative-group infrastructure. The next three years will depend less on building new models than on running the trials that test the ones already on the table.
Author Contributions
Conceptualization, F.D. G.C. and M.G.; writing—original draft preparation, F.D.; writing—review and editing, F.D., A.A., G.C., A.S. and M.G.; supervision, M.G. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
Not applicable. No new data were created or analyzed in this study.
Conflicts of Interest
The authors declare no conflicts of interest.
Use of Generative AI in the Preparation of this Manuscript
The authors used large language models (Anthropic Claude) to assist with copy-editing, citation formatting, and consistency checking during manuscript preparation. All scientific claims, citations, interpretations, and conclusions were reviewed, verified against the primary literature, and approved by the authors, who take full responsibility for the content of the work.
References
- Shimony, S.; Stahl, M.; Stone, R.M. Acute Myeloid Leukemia: 2025 Update on Diagnosis, Risk-Stratification, and Management. Am. J. Hematol. 2025, 100, 860–891. [Google Scholar] [CrossRef] [PubMed]
- Döhner, H.; Wei, A.H.; Appelbaum, F.R.; et al. Diagnosis and Management of AML in Adults: 2022 Recommendations from an International Expert Panel on Behalf of the ELN. Blood 2022, 140, 1345–1377. [Google Scholar] [CrossRef] [PubMed]
- Döhner, H.; DiNardo, C.D.; Appelbaum, F.R.; et al. Genetic Risk Classification for Adults with AML Receiving Less-Intensive Therapies: The 2024 ELN Recommendations. Blood 2024, 144, 2169–2173. [Google Scholar] [CrossRef] [PubMed]
- The Cancer Genome Atlas Research Network. Genomic and Epigenomic Landscapes of Adult De Novo Acute Myeloid Leukemia. N. Engl. J. Med. 2013, 368, 2059–2074. [CrossRef] [PubMed]
- DiNardo, C.D.; Jonas, B.A.; Pullarkat, V.; et al. Azacitidine and Venetoclax in Previously Untreated Acute Myeloid Leukemia. N. Engl. J. Med. 2020, 383, 617–629. [Google Scholar] [CrossRef] [PubMed]
- Nazha, A.; Elemento, O.; Ahuja, S.; et al. Artificial Intelligence in Hematology. Blood 2025, 146, 2283–2292. [Google Scholar] [CrossRef] [PubMed]
- Al-Nusair, J.; Lanino, L.; Durmaz, A.; et al. Artificial Intelligence in Myeloid Malignancies: Clinical Applications of Machine Learning in MDS and AML. Blood Rev. 2025, 74, 101340. [Google Scholar] [CrossRef] [PubMed]
- Sabile, J.M.G.; Zhang, P.; Parwani, A.V.; et al. Toward Clinically Actionable Machine Learning and Artificial Intelligence Algorithms in Acute Leukemia: A Systematic Narrative Review. Acta Haematol. 2025, 148, 583–599. [Google Scholar] [CrossRef] [PubMed]
- Kilic Gunes, E.; et al. Artificial Intelligence in Hematology: Current Trends and Application Areas. Ann. Hematol. 2025, 104, 6033–6043. [Google Scholar] [CrossRef] [PubMed]
- Obeagu, E.I.; et al. Big Data Analytics and Machine Learning in Hematology: Transformative Insights, Applications and Challenges. Medicine 2025, 104, e41766. [Google Scholar] [CrossRef] [PubMed]
- Singh, R.; et al. How AI Is Used in FDA-Authorized Medical Devices: A Taxonomy Across 1,016 Authorizations. npj Digit. Med. 2025, 8, 388. [Google Scholar] [CrossRef] [PubMed]
- Zini, G. Hematological Cytomorphology: Where We Are. Int. J. Lab. Hematol. 2024, 46, 789–794. [Google Scholar] [CrossRef] [PubMed]
- Al-Obeidat, F.; et al. Artificial Intelligence for the Detection of Acute Myeloid Leukemia from Microscopic Blood Images: A Systematic Review and Meta-Analysis. Front. Big Data 2025, 7, 1402926. [Google Scholar] [CrossRef] [PubMed]
- Liao, H.; et al. Application of AI in Laboratory Hematology. Acta Pharm. Sin. B 2025, 15, 5702–5733. [Google Scholar] [CrossRef] [PubMed]
- Ghete, T.; et al. Models for the Marrow: A Comprehensive Review of AI-Based Cell Classification Methods and Malignancy Detection in Bone Marrow Aspirate Smears. HemaSphere 2024, 8, e70048. [Google Scholar] [CrossRef] [PubMed]
- Pescia, C.; et al. AI in Haematopathology: Current Perspective and Future Directions. Diagn. Histopathol. 2025, 31, 267–276. [Google Scholar] [CrossRef]
- Czapliński, M.; et al. An Overview of Existing Applications of Artificial Intelligence in Histopathological Diagnostics of Leukemias: A Scoping Review. Electronics 2025, 14, 4144. [Google Scholar] [CrossRef]
- Eckardt, J.-N.; et al. Deep Learning Detects Acute Myeloid Leukemia and Predicts NPM1 Mutation Status from Bone Marrow Smears. Leukemia 2022, 36, 111–118. [Google Scholar] [CrossRef] [PubMed]
- Kockwelp, J.; et al. Deep Learning Predicts Therapy-Relevant Genetics in AML from Pappenheim-Stained Bone Marrow Smears. Blood Adv. 2024, 8, 70. [Google Scholar] [CrossRef] [PubMed]
- Wei, B.-H.; Tsai, X.C.-H.; Sun, K.-J.; et al. Annotation-Free Deep Learning for Predicting Gene Mutations from Whole Slide Images of Acute Myeloid Leukemia. npj Precis. Oncol. 2025, 9, 35. [Google Scholar] [CrossRef] [PubMed]
- Deltadahl, S.; et al. CytoDiffusion. Nat. Mach. Intell. 2025, 7, 1791. [Google Scholar] [CrossRef] [PubMed]
- Kim, H.; Hur, M.; d’Onofrio, G.; Zini, G. Real-World Application of Digital Morphology Analyzers: Practical Issues and Challenges in Clinical Laboratories. Diagnostics 2025, 15, 677. [Google Scholar] [CrossRef] [PubMed]
- Spies, N.C.; Rangel, A.; English, P.; Morrison, M.; O’Fallon, B.; Ng, D.P. Machine Learning Methods in Clinical Flow Cytometry. Cancers 2025, 17, 483. [Google Scholar] [CrossRef] [PubMed]
- Heuser, M.; Freeman, S.D.; Ossenkoppele, G.J.; et al. 2021 Update on Measurable Residual Disease in Acute Myeloid Leukemia: A Consensus Document from the European LeukemiaNet MRD Working Party. Blood 2021, 138, 2753–2767. [Google Scholar] [CrossRef] [PubMed]
- Cloos, J.; et al. 2025 Update on MRD in Acute Myeloid Leukemia: A Consensus Document from the ELN-DAVID MRD Working Party. Blood 2026, 147, 1147–1167. [Google Scholar] [CrossRef] [PubMed]
- Cheng, F.M.; et al. Deep Learning for Acute Leukemia Detection via Flow Cytometry. Sci. Rep. 2024, 14, 8350. [Google Scholar] [CrossRef] [PubMed]
- Wang, Y.-F.; et al. A Machine Learning Framework for Cross-Institute Standardized Analysis of Flow Cytometry in Differentiating AML from Non-Neoplastic Conditions. Comput. Biol. Med. 2025, 193, 110394. [Google Scholar] [CrossRef] [PubMed]
- Lewis, J.E.; et al. Automated Deep Learning-Based Diagnosis and Molecular Characterization of Acute Myeloid Leukemia Using Flow Cytometry. Mod. Pathol. 2024, 37, 100373. [Google Scholar] [CrossRef] [PubMed]
- Seheult, J.N.; Otteson, G.E.; Timm, M.M.; et al. Artificial Intelligence Accelerates the Interpretation of Measurable Residual B Lymphoblastic Leukemia by Flow Cytometry. Blood Adv. 2026, 10, 58–69. [Google Scholar] [CrossRef] [PubMed]
- Hoffmann, J.; Eminovic, S.; Wilhelm, C.; Krause, S.W.; Neubauer, A.; Thrun, M.C.; Ultsch, A.; Brendel, C. Prediction of Clinical Outcomes with Explainable Artificial Intelligence in Patients with Chronic Lymphocytic Leukemia. Curr. Oncol. 2023, 30, 1903–1915. [Google Scholar] [CrossRef] [PubMed]
- Song, Y.; et al. Classification of Acute Myeloid Leukemia Based on Multi-Omics and Prognosis Prediction Value (UAMOCS). Mol. Oncol. 2025, 19, 1836–1854. [Google Scholar] [CrossRef] [PubMed]
- Tazi, Y.; Arango-Ossa, J.E.; Zhou, Y.; et al. Unified Classification and Risk-Stratification in Acute Myeloid Leukaemia. Nat. Commun. 2022, 13, 4622. [Google Scholar] [CrossRef] [PubMed]
- Turki, A.; et al. International Testing and Refinement of AI Algorithms Predicting Acute Leukemia Subtypes from Routine Laboratory Data. Nat. Commun. 2026, 17, 2649. [Google Scholar] [CrossRef] [PubMed]
- Găman, M.A.; et al. Applications of Artificial Intelligence in Acute Promyelocytic Leukemia: A Systematic Review. J. Clin. Med. 2025, 14, 1670. [Google Scholar] [CrossRef] [PubMed]
- Fathima, N.; et al. CMML2AML: Machine-Learning Discovery of Co-Mutations Predictive of Blast Transformation in Chronic Myelomonocytic Leukemia. Blood Cancer J. 2026, 16, 76. [Google Scholar] [CrossRef] [PubMed]
- Ahmed, F.; Zhong, J. Advances in DNA/RNA Sequencing and Their Applications in Acute Myeloid Leukemia. Int. J. Mol. Sci. 2025, 26, 71. [Google Scholar] [CrossRef] [PubMed]
- Vermeulen, C.; Pagès-Gallego, M.; Kester, L.; et al. Ultra-Fast Deep-Learned CNS Tumour Classification during Surgery. Nature 2023, 622, 842–849. [Google Scholar] [CrossRef] [PubMed]
- Deacon, S.; Cahyani, I.; Holmes, N.; et al. ROBIN: A Unified Nanopore-Based Assay Integrating Intraoperative Methylome Classification and Next-Day Comprehensive Profiling for Ultra-Rapid Tumor Diagnosis. Neuro Oncol. 2025, 27, 2035–2046. [Google Scholar] [CrossRef] [PubMed]
- Steinicke, T.L.; Benfatto, S.; Capilla-Guerra, M.R.; et al. Rapid Epigenomic Classification of Acute Leukemia. Nat. Genet. 2025, 57, 2456–2467. [Google Scholar] [CrossRef] [PubMed]
- Marchi, F.; Shastri, V.M.; Marrero, R.J.; et al. Epigenomic Diagnosis and Prognosis of Acute Myeloid Leukemia. Nat. Commun. 2025, 16, 6961. [Google Scholar] [CrossRef] [PubMed]
- Gerstung, M.; et al. Precision Oncology for Acute Myeloid Leukemia Using a Knowledge Bank Approach. Nat. Genet. 2017, 49, 332–340. [Google Scholar] [CrossRef] [PubMed]
- Eckardt, J.-N.; et al. Age-Stratified Machine Learning Identifies Divergent Prognostic Significance of Molecular Alterations in AML. HemaSphere 2025, 9, e70132. [Google Scholar] [CrossRef] [PubMed]
- Hoff, F.W.; et al. Beat-AML 2024 ELN-Refined Risk Stratification for Older Adults with Newly Diagnosed AML Given Lower-Intensity Therapy. Blood Adv. 2024, 8, 5297–5305. [Google Scholar] [CrossRef] [PubMed]
- Didi, I.; et al. Artificial Intelligence-Based Prediction Models for Acute Myeloid Leukemia Using Real-Life Data: A DATAML Registry Study. Leuk. Res. 2024, 136, 107437. [Google Scholar] [CrossRef] [PubMed]
- Eckardt, J.-N.; et al. Prediction of Complete Remission and Survival in Acute Myeloid Leukemia Using Supervised Machine Learning. Haematologica 2023, 108, 690–704. [Google Scholar] [CrossRef] [PubMed]
- Park, S.; et al. Prognostic Value of ELN 2022 Criteria and Genomic Clusters Using Machine Learning in Older Adults with AML. Haematologica 2024, 109, 1095–1106. [Google Scholar] [CrossRef] [PubMed]
- Lachowiez, C.A.; et al. Refined ELN 2024 Risk Stratification Improves Survival Prognostication Following Venetoclax-Based Therapy in AML. Blood 2024, 144, 2788–2792. [Google Scholar] [CrossRef] [PubMed]
- Jia, Z.Y.; et al. A Novel Perspective on Survival Prediction for AML Patients: Integration of Machine Learning in SEER Database Applications. Heliyon 2025, 11, e42030. [Google Scholar] [CrossRef] [PubMed]
- Ng, S.W.K.; et al. A 17-Gene Stemness Score for Rapid Determination of Risk in Acute Leukaemia. Nature 2016, 540, 433–437. [Google Scholar] [CrossRef] [PubMed]
- Silva, R.; et al. AML Risk Stratification Through Transcriptomic Machine Learning. Sci. Rep. 2025, 15, 39821. [Google Scholar] [CrossRef] [PubMed]
- Kosvyra, A.; et al. Machine Learning and Integrative Multi-Omics Network Analysis for Survival Prediction in Acute Myeloid Leukemia. Comput. Biol. Med. 2024, 178, 108735. [Google Scholar] [CrossRef] [PubMed]
- Li, J.; et al. Integrative Analysis of Epigenetic Subtypes in Acute Myeloid Leukemia: A Multi-Center Study Combining Machine Learning for Prognostic and Therapeutic Insights. PLoS ONE 2025, 20, e0324380. [Google Scholar] [CrossRef] [PubMed]
- Soleimani Samarkhazan, H.; et al. Integrating Multi-Omics Approaches in Acute Myeloid Leukemia: Advancements and Clinical Implications. Clin. Exp. Med. 2025, 25, 311. [Google Scholar] [CrossRef] [PubMed]
- Eckardt, J.-N.; et al. Artificial Intelligence for Risk Assessment and Outcome Prediction in Malignant Haematology. Br. J. Haematol. 2025, 208, 25–38. [Google Scholar] [CrossRef] [PubMed]
- Ly, C.; et al. Multimodal Spatial Proteomic Profiling in Acute Myeloid Leukemia. npj Precis. Oncol. 2025, 9, 148. [Google Scholar] [CrossRef] [PubMed]
- Lipkova, J.; et al. Artificial Intelligence for Multimodal Data Integration in Oncology. Cancer Cell 2022, 40, 1095–1110. [Google Scholar] [CrossRef] [PubMed]
- Jin, P.; et al. Precision Prediction of Venetoclax-Azacitidine Treatment Efficacy in Acute Myeloid Leukemia via Integrative Drug Screening and Machine Learning. Cell Rep. Med. 2025, 6, 102461. [Google Scholar] [CrossRef] [PubMed]
- Qin, G.; et al. Knowledge Graphs Facilitate Prediction of Drug Response for Acute Myeloid Leukemia. iScience 2024, 27, 110755. [Google Scholar] [CrossRef] [PubMed]
- Wang, Y.; et al. A Network-Driven Framework for Drug Response Precision Prediction of Acute Myeloid Leukemia (NetAML). Adv. Sci. 2025, 12, 2506447. [Google Scholar] [CrossRef] [PubMed]
- Qin, G.; Dai, J.; Chien, S.; et al. Mutation Patterns Predict Drug Sensitivity in Acute Myeloid Leukemia. Clin. Cancer Res. 2024, 30, 2659–2671. [Google Scholar] [CrossRef] [PubMed]
- Pei, S.; Pollyea, D.A.; Gustafson, A.; et al. Monocytic Subclones Confer Resistance to Venetoclax-Based Therapy in Patients with Acute Myeloid Leukemia. Cancer Discov. 2020, 10, 536–551. [Google Scholar] [CrossRef] [PubMed]
- Islam, N.; et al. Machine Learning-Based Exploratory Clinical Decision Support for Newly Diagnosed Patients with AML Treated with 7+3 Type Chemotherapy or Venetoclax/Azacitidine. JCO Clin. Cancer Inform. 2022, 6, e2200030. [Google Scholar] [CrossRef] [PubMed]
- Zhou, Y.; et al. Improving Severity Grading of Chemotherapy-Induced Myelosuppression in AML via Data-Driven and Model-Based Deep Learning (MM-AI-AML). npj Syst. Biol. Appl. 2026, 12, 76. [Google Scholar] [CrossRef] [PubMed]
- Li, H.; Han, Z.; Sun, Y.; et al. CGMega: Explainable Graph Neural Network Framework with Attention Mechanisms for Cancer Gene Module Dissection. Nat. Commun. 2024, 15, 5997. [Google Scholar] [CrossRef] [PubMed]
- Gonzalez, G.; et al. PDGrapher: Combinatorial Prediction of Therapeutic Perturbations Using Causally Inspired Neural Networks. Nat. Biomed. Eng. 2025. [Google Scholar] [CrossRef] [PubMed]
- Alpsoy, S.; Sezerman, O.U. Transfer Learning with Multiomics Integration and Deep Neural Networks Reveals Drug Resistance Mechanisms in Cancer. Sci. Rep. 2025, 15, 42295. [Google Scholar] [CrossRef] [PubMed]
- Schmutz, M.; Sommer, S.; Sander, J.; et al. Large Language Model Processing Capabilities of ChatGPT 4.0 to Generate Molecular Tumor Board Recommendations—A Critical Evaluation on Real World Data. The Oncologist 2025, 30, oyaf293. [Google Scholar] [CrossRef] [PubMed]
- Jin, Q.; et al. Matching Patients to Clinical Trials with Large Language Models. Nat. Commun. 2024, 15, 9074. [Google Scholar] [CrossRef] [PubMed]
- Rubinstein, S.; et al. Summarizing Clinical Evidence Utilizing Large Language Models for Cancer Treatments: A Blinded Comparative Analysis. Front. Digit. Health 2025, 7, 1569554. [Google Scholar] [CrossRef] [PubMed]
- Yang, J.J.; Hwang, S.-H. Transforming Hematological Research Documentation with Large Language Models: An Approach to Scientific Writing and Data Analysis. Blood Res. 2025, 60, 15. [Google Scholar] [CrossRef] [PubMed]
- Chen, D.; et al. Large Language Model Applications for Health Information Extraction in Oncology: Scoping Review. JMIR Cancer 2025, 11, e65984. [Google Scholar] [CrossRef] [PubMed]
- Mehan, N.; et al. Development and Evaluation of Large-Language Models for Oncology: A Scoping Review. PLoS Digit. Health 2025, 4, e0000980. [Google Scholar] [CrossRef] [PubMed]
- Chase, A.; et al. Large Language Model as Clinical Decision Support System Augments Medication Safety in 16 Clinical Specialties. Front. Pharmacol. 2025. [Google Scholar] [CrossRef] [PubMed]
- Mocking, T.R.; van de Loosdrecht, A.A.; Cloos, J.; Bachas, C. Applications of Machine Learning for Immunophenotypic Measurable Residual Disease Assessment in Acute Myeloid Leukemia. HemaSphere 2025, 9, e70138. [Google Scholar] [CrossRef] [PubMed]
- Ni, W.; et al. Automated Analysis of Acute Myeloid Leukemia Minimal Residual Disease Using a Support Vector Machine. Oncotarget 2016, 7, 71915–71921. [Google Scholar] [CrossRef] [PubMed]
- Ko, B.S.; et al. Clinically Validated Machine Learning Algorithm for Detecting Residual Diseases with Multicolor Flow Cytometry Analysis in Acute Myeloid Leukemia and Myelodysplastic Syndrome. EBioMedicine 2018, 37, 91–100. [Google Scholar] [CrossRef] [PubMed]
- Shopsowitz, K.; et al. MAGIC-DR: An Interpretable Machine-Learning Guided Approach for Acute Myeloid Leukemia Measurable Residual Disease Analysis. Cytom. B Clin. Cytom. 2024, 106, 239–251. [Google Scholar] [CrossRef] [PubMed]
- Jongen-Lavrencic, M.; et al. Molecular Minimal Residual Disease in Acute Myeloid Leukemia. N. Engl. J. Med. 2018, 378, 1189–1199. [Google Scholar] [CrossRef] [PubMed]
- Morita, K.; et al. Clearance of Somatic Mutations at Remission and the Risk of Relapse in Acute Myeloid Leukemia. J. Clin. Oncol. 2018, 36, 1788–1797. [Google Scholar] [CrossRef] [PubMed]
- Awada, H.; et al. Machine Learning Integrates Genomic Signatures for Subclassification Beyond Primary and Secondary Acute Myeloid Leukemia. Blood 2021, 138, 1885–1895. [Google Scholar] [CrossRef] [PubMed]
- Benard, B.A.; et al. Clonal Architecture Predicts Clinical Outcomes and Drug Sensitivity in Acute Myeloid Leukemia. Nat. Commun. 2021, 12, 7244. [Google Scholar] [CrossRef] [PubMed]
- Shlush, L.I.; et al. Tracing the Origins of Relapse in Acute Myeloid Leukaemia to Stem Cells. Nature 2017, 547, 104–108. [Google Scholar] [CrossRef] [PubMed]
- Hourigan, C.S.; et al. Impact of Conditioning Intensity of Allogeneic Transplantation for Acute Myeloid Leukemia with Genomic Evidence of Residual Disease. J. Clin. Oncol. 2020, 38, 1273–1283. [Google Scholar] [CrossRef] [PubMed]
- Yoest, J.M.; et al. Sequencing-Based Measurable Residual Disease Testing in Acute Myeloid Leukemia. Front. Cell Dev. Biol. 2020, 8, 249. [Google Scholar] [CrossRef] [PubMed]
- Caravagna, G.; Heide, T.; Williams, M.J.; Zapata, L.; Nichol, D.; Chkhaidze, K.; Cross, W.; Cresswell, G.D.; Werner, B.; Acar, A.; et al. Subclonal Reconstruction of Tumors by Using Machine Learning and Population Genetics. Nat. Genet. 2020, 52, 898–907. [Google Scholar] [CrossRef] [PubMed]
- Morita, K.; et al. Clonal Evolution of Acute Myeloid Leukemia Revealed by High-Throughput Single-Cell Genomics. Nat. Commun. 2020, 11, 5327. [Google Scholar] [CrossRef] [PubMed]
- Nam, A.S.; Kim, K.-T.; Chaligne, R.; et al. Somatic Mutations and Cell Identity Linked by Genotyping of Transcriptomes (GoT). Nature 2019, 571, 355–360. [Google Scholar] [CrossRef] [PubMed]
- van Galen, P.; et al. Single-Cell RNA-Seq Reveals AML Hierarchies Relevant to Disease Progression and Immunity. Cell 2019, 176, 1265–1281. [Google Scholar] [CrossRef] [PubMed]
- Khosroabadi, Z.; et al. Single Cell RNA Sequencing Improves the Next Generation of Approaches to AML Treatment: Challenges and Perspectives. Mol. Med. 2025, 31, 33. [Google Scholar] [CrossRef] [PubMed]
- Naldini, M.M.; et al. Longitudinal Single-Cell Profiling of Chemotherapy Response in Acute Myeloid Leukemia. Nat. Commun. 2023, 14, 1285. [Google Scholar] [CrossRef] [PubMed]
- Gui, G.; et al. Single-Cell Spatial Transcriptomics Reveals Immunotherapy-Driven Bone Marrow Niche Remodeling in AML. Sci. Adv. 2025, 11, eadw4871. [Google Scholar] [CrossRef] [PubMed]
- Cooper, R.A.; et al. Spatial Transcriptomic Approaches for Characterising the Bone Marrow Landscape: Pitfalls and Potential. Leukemia 2025, 39, 291–295. [Google Scholar] [CrossRef] [PubMed]
- Masurekar, A.N.; et al. Defining the Limits of Pre-Transplant Risk Prediction in AML: Evidence from Machine Learning and Regression Models. Transplant. Cell. Ther. 2026, 32, 575.e1–575.e12. [Google Scholar] [CrossRef] [PubMed]
- Asteris, P.G.; et al. Survival Prediction in Allogeneic Haematopoietic Stem Cell Transplantation Recipients Using Pre- and Post-Transplant Factors and Computational Intelligence (DERGA). J. Cell. Mol. Med. 2025, 29, e70672. [Google Scholar] [CrossRef] [PubMed]
- Zhou, Y.; et al. Longitudinal Clinical Data Improve Survival Prediction After Hematopoietic Cell Transplantation Using Machine Learning. Blood Adv. 2024, 8, 686–698. [Google Scholar] [CrossRef] [PubMed]
- Arabyarmohammadi, S.; et al. Machine Learning to Predict Risk of Relapse Using Cytologic Image Markers in Patients with Acute Myeloid Leukemia Posthematopoietic Cell Transplantation. In JCO Clin. Cancer Inform.; 2022. [Google Scholar] [CrossRef] [PubMed]
- Fan, S.; et al. Artificial Intelligence-Based Predictive Model for Relapse in Acute Myeloid Leukemia Patients Following Haploidentical Hematopoietic Cell Transplantation. J. Transl. Intern. Med. 2025, 13, 253–266. [Google Scholar] [CrossRef] [PubMed]
- Fuse, K.; et al. Patient-Based Prediction Algorithm of Relapse After Allo-HSCT for Acute Leukemia and Its Usefulness in the Decision-Making Process Using a Machine Learning Approach. Cancer Med. 2019, 8, 5058–5067. [Google Scholar] [CrossRef] [PubMed]
- Shyr, D.; et al. Exploring Pattern of Relapse in Pediatric Patients with Acute Lymphocytic Leukemia and Acute Myeloid Leukemia Undergoing Stem Cell Transplant Using Machine Learning Methods. J. Clin. Med. 2024, 13, 4021. [Google Scholar] [CrossRef] [PubMed]
- Hernández-Boluda, J.C.; et al. Use of Machine Learning Techniques to Predict Poor Survival After Hematopoietic Cell Transplantation for Myelofibrosis. Blood 2025, 145, 3139–3152. [Google Scholar] [CrossRef] [PubMed]
- Wang, X.; Chen, Y.; Li, X.; Liu, A.; Qu, Y.; et al. Using Machine Learning to Predict Acute Graft-Versus-Host Disease in Pediatric Patients Undergoing Allogeneic Hematopoietic Stem Cell Transplantation: Integration of Clinical and Genetic Factors. Med. Adv. 2025, 3, 287–299. [Google Scholar] [CrossRef]
- Rowley, S.D.; et al. Using Targeted Transcriptome and Machine Learning of Pre- and Post-Transplant Bone Marrow Samples to Predict Acute Graft-versus-Host Disease and Overall Survival. Cancers 2024, 16, 1357. [Google Scholar] [CrossRef] [PubMed]
- Jo, T.; et al. A Convolutional Neural Network-Based Model That Predicts Acute Graft-Versus-Host Disease After Allogeneic Hematopoietic Stem Cell Transplantation. Commun. Med. 2023, 3, 67. [Google Scholar] [CrossRef] [PubMed]
- Rouzbahani, M.; et al. Predictive Modeling of Outcomes in Acute Leukemia Patients Undergoing Allogeneic Hematopoietic Stem Cell Transplantation Using Machine Learning Techniques. Leuk. Res. 2025, 148, 107619. [Google Scholar] [CrossRef] [PubMed]
- Daver, N.; Alotaibi, A.S.; Bücklein, V.; Subklewe, M. T-Cell-Based Immunotherapy of Acute Myeloid Leukemia: Current Concepts and Future Developments. Leukemia 2021, 35, 1843–1863. [Google Scholar] [CrossRef] [PubMed]
- Gottschlich, A.; Thomas, M.; Grünmeier, R.; et al. Single-Cell Transcriptomic Atlas-Guided Development of CAR-T Cells for the Treatment of Acute Myeloid Leukemia. Nat. Biotechnol. 2023, 41, 1618–1632. [Google Scholar] [CrossRef] [PubMed]
- Wang, M.; et al. ML Model for Early Relapse Prediction Post Axi-cel in DLBCL. Blood Adv. 2025, 9, 5837–5852. [Google Scholar] [CrossRef] [PubMed]
- Wei, Z.; Zhao, C.; Zhang, M.; et al. PrCRS: A Prediction Model of Severe CRS in CAR-T Therapy Based on Transfer Learning. BMC Bioinform. 2024, 25, 197. [Google Scholar] [CrossRef] [PubMed]
- Garuffo, L.; Leoni, A.; Gatta, R.; Bernardi, S. The Applications of Machine Learning in the Management of Patients Undergoing Stem Cell Transplantation: Are We Ready? Cancers 2025, 17, 395. [Google Scholar] [CrossRef] [PubMed]
- Windecker, D.; et al. Generalizability of FDA-Approved AI-Enabled Medical Devices for Clinical Use. JAMA Netw. Open 2025, 8, e258052. [Google Scholar] [CrossRef] [PubMed]
- Tran, D.; et al. Comprehensive Review of Cancer Survival Prediction Using Multi-Omics Integration and Clinical Variables. Brief. Bioinform. 2025, 26, bbaf150. [Google Scholar] [CrossRef] [PubMed]
- Almarie, B.; Gonzalez-Gonzalez, L.F.; dos Santos Barbosa, L.A.; Lutz, A.; Grosse, U.; Fregni, F. Machine Learning-Enabled Medical Devices Authorized by the US Food and Drug Administration in 2024: Regulatory Characteristics, Predicate Lineage, and Transparency Reporting. Biomedicines 2025, 13, 3005. [Google Scholar] [CrossRef] [PubMed]
- Mehta, S.; Komanduri, V.; Bhadouriya, S.; et al. Evaluating Transparency in AI/ML Model Characteristics for FDA-Reviewed Medical Devices. npj Digit. Med. 2025, 8, 673. [Google Scholar] [CrossRef] [PubMed]
- Lin, J.C.; et al. Benefit-Risk Reporting for FDA-Cleared Artificial Intelligence-Enabled Medical Devices. JAMA Health Forum 2025, 6, e253351. [Google Scholar] [CrossRef] [PubMed]
- Salih, A.; Raisi-Estabragh, Z.; Galazzo, I.B.; et al. A Perspective on Explainable Artificial Intelligence Methods: SHAP and LIME. Adv. Intell. Syst. 2025, 7, 2400304. [Google Scholar] [CrossRef]
- Hehr, M.; Sadafi, A.; Matek, C.; Lienemann, P.; Pohlkamp, C.; Haferlach, T.; Spiekermann, K.; Marr, C. Explainable AI Identifies Diagnostic Cells of Genetic AML Subtypes. PLoS Digit. Health 2023, 2, e0000187. [Google Scholar] [CrossRef] [PubMed]
- Thiriveedhi, A.; Ghanta, S.; Biswas, S.; Pradhan, A.K. ALL-Net: Integrating CNN and Explainable-AI for Enhanced Diagnosis and Interpretation of Acute Lymphoblastic Leukemia. PeerJ Comput. Sci. 2025, 11, e2600. [Google Scholar] [CrossRef] [PubMed]
- Parwez, K.; et al. A Trust-Centered Explainable Deep-Learning Framework for Acute Lymphoblastic Leukemia Detection Using Multi-Model Fusion and Interpretability Scoring. Algorithms 2026, 19, 162. [Google Scholar] [CrossRef]
- Hsu, C.-Y.; et al. AI-Driven Multi-Omics Integration in Precision Oncology: Bridging the Data Deluge to Clinical Decisions. Clin. Exp. Med. 2025, 26, 29. [Google Scholar] [CrossRef] [PubMed]
- Diniz, J.M.; et al. Comparing Decentralized Machine Learning and AI Clinical Models to Local and Centralized Alternatives: A Systematic Review. npj Digit. Med. 2026, 9, 174. [Google Scholar] [CrossRef] [PubMed]
- Hamamoto, R.; Koyama, T.; Takahashi, S.; et al. Implementing Generative Artificial Intelligence in Precision Oncology: Safety, Governance, and Significance. J. Hematol. Oncol. 2026, 19, 14. [Google Scholar] [CrossRef] [PubMed]
- European Parliament and Council of the European Union. Off. J. Eur. Union 2024, L 2024/1689; Regulation (EU) 2024/1689 of 13 June 2024 Laying Down Harmonised Rules on Artificial Intelligence (Artificial Intelligence Act). Available online: https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng.
- Suresh, P.; Keerthika, G.; Kumar, K.S.; et al. A Personalized Communication Efficient Federated Learning Framework with Low Rank Adaptation for Intelligent Leukemia Diagnosis. Sci. Rep. 2025, 16, 339. [Google Scholar] [CrossRef] [PubMed]
- Parwez, K.; Sohail, S.I.; Akram, A.; Rashid, J.; Atteia, G.; Sarwar, N. Explainable Federated Transformer Framework for Joint Leukemia Classification and Stage Prediction. Sci. Rep. 2026, 16, 4493. [Google Scholar] [CrossRef] [PubMed]
- Hernández-Sánchez, A.; González, T.; Sobas, M.; et al. Rearrangements Involving 11q23.3/KMT2A in Adult AML: Mutational Landscape and Prognostic Implications—A HARMONY Study. Leukemia 2024, 38, 1929–1937. [Google Scholar] [CrossRef] [PubMed]
- Li, N.; et al. Privacy-Preserving Federated Data Access and Federated Learning: Improved Data Sharing and AI Model Development in Transfusion Medicine. Transfusion 2025, 65, 22–28. [Google Scholar] [CrossRef] [PubMed]
- Valous, N.A.; et al. Graph Machine Learning for Integrated Multi-Omics Analysis. Br. J. Cancer 2024, 131, 205–211. [Google Scholar] [CrossRef] [PubMed]
- Cui, H.; et al. scGPT: Toward Building a Foundation Model for Single-Cell Multi-Omics Using Generative AI. Nat. Methods 2024, 21, 1470–1480. [Google Scholar] [CrossRef] [PubMed]
- Zeng, Y.; et al. CellFM: A Large-Scale Foundation Model Pre-Trained on Transcriptomics of 100 Million Human Cells. Nat. Commun. 2025, 16, 4679. [Google Scholar] [CrossRef] [PubMed]
- Tejada-Lapuerta, A.; Schaar, A.C.; et al. Nicheformer: A Foundation Model for Single-Cell and Spatial Omics. Nat. Methods 2025, 22, 2525–2538. [Google Scholar] [CrossRef] [PubMed]
- Baek, S.; Song, K.; Lee, I. Single-Cell Foundation Models: Bringing Artificial Intelligence into Cell Biology. Exp. Mol. Med. 2025, 57, 2169–2181. [Google Scholar] [CrossRef] [PubMed]
- Kedzierska, K.Z.; Crawford, L.; Amini, A.P.; Lu, A.X. Zero-Shot Evaluation Reveals Limitations of Single-Cell Foundation Models. Genome Biol. 2025, 26, 101. [Google Scholar] [CrossRef] [PubMed]
- Stahlberg, E.A.; et al. Exploring Approaches for Predictive Cancer Patient Digital Twins: Opportunities for Collaboration and Innovation. Front. Digit. Health 2022, 4, 1007784. [Google Scholar] [CrossRef] [PubMed]
- National Cancer Institute. Digital Twins for Cancer — Not If, But When, How, and Why? NCI CBIIT Blog. 2025. Available online: https://www.cancer.gov/about-nci/organization/cbiit/news-events/blog/2025/digital-twins-cancer-not-if-when-how-and-why (accessed on 17 April 2026).
- Bouriga, R.; et al. Advances and Critical Aspects in Cancer Treatment Development Using Digital Twins. Brief. Bioinform. 2025, 26, bbaf237. [Google Scholar] [CrossRef] [PubMed]
- Giansanti, D.; Morelli, S. Exploring the Potential of Digital Twins in Cancer Treatment: A Narrative Review of Reviews. J. Clin. Med. 2025, 14, 3574. [Google Scholar] [CrossRef] [PubMed]
- Gatenby, R.A.; Brown, J.; Vincent, T. Lessons from Applied Ecology: Cancer Control Using an Evolutionary Double Bind. Cancer Res. 2009, 69, 7499–7502. [Google Scholar] [CrossRef] [PubMed]
- West, J.; You, L.; Zhang, J.; Gatenby, R.A.; Brown, J.S.; Newton, P.K.; Anderson, A.R.A. Towards Multidrug Adaptive Therapy. Cancer Res. 2020, 80, 1578–1589. [Google Scholar] [CrossRef] [PubMed]
- Mashayekhi, H.; Nazari, M.; Jafarinejad, F.; Meskin, N. Deep Reinforcement Learning-Based Control of Chemo-Drug Dose in Cancer Treatment. Comput. Methods Programs Biomed. 2024, 243, 107884. [Google Scholar] [CrossRef] [PubMed]
- Tosca, E.M.; De Carlo, A.; Ronchi, D.; Magni, P. Model-Informed Reinforcement Learning for Enabling Precision Dosing via Adaptive Dosing. Clin. Pharmacol. Ther. 2024, 116, 619–636. [Google Scholar] [CrossRef] [PubMed]
Table 4.
CRL-AML 3+ tools across the AML care continuum (externally validated, prospectively evaluated, or regulatory-cleared). All other tools discussed sit at CRL-AML 1–2.
Table 4.
CRL-AML 3+ tools across the AML care continuum (externally validated, prospectively evaluated, or regulatory-cleared). All other tools discussed sit at CRL-AML 1–2.
| Tool/Reference | Modality/Stage | CRL-AML | External validation evidence | Regulatory/Clinical access |
|---|---|---|---|---|
| Scopio Labs FF-BMA [22] | Digital morphology — bone-marrow aspirate | 5 | — | FDA De Novo (DEN230034, Apr 2024); deployed |
| Scopio Labs RBC + platelet morphology [22] | Digital morphology — peripheral blood | 5 | — | FDA 510(k) (Jul 2025); deployed |
| CellaVision DM-/DC- series [22] | Digital morphology — peripheral blood differential | 5 | Multi-site clinical use over >15 years | FDA-cleared (multiple); EU CE-IVDR; widely deployed |
| Eckardt 2022 NPM1-from-BM-smear [18] | Digital morphology — mutation prediction | 3 | Independent test cohort within multi-centre dataset | Research |
| Wang 2025 cross-institute flow [27] | Flow cytometry — AML diagnosis | 3 | Heterogeneous panels across institutions | Research |
| Tazi 2022 unified classification [32] | Genomic — molecular subtyping + risk | 3 | n=3,653 across multiple cohorts; open-access decision-support tool released | Research; web tool |
| Turki 2026 international AI [33] | Routine-lab triage — acute leukemia subtype | 3 | n=6,206 across 20 international centres | Research |
| MARLIN (Steinicke 2025) [39] | Diagnosis — nanopore methylation subtyping | 3 | 25/26 retrospective concordance; 5/5 real-time prospective | Research |
| ALMA (Marchi 2025) [40] | Diagnosis/prognosis — methylation atlas | 3 | Independent pediatric + adult cohorts across 11 harmonized datasets | Research |
| Gerstung knowledge bank [41] | Risk — individualised survival | 3 | Validation across multiple AML cohorts | Research; open-access calculator |
| Eckardt age-stratified [42] | Risk — age × mutation interaction | 3 | n=3,062 pooled multi-cohort | Research |
| Beat-AML 2024 ELN-refined [43] | Risk — older adults on lower-intensity therapy | 3 | Held-out validation within Beat-AML | Research |
| Eckardt CR-supervised [45] | Risk — complete remission prediction | 3 | n=664 external validation cohort | Research |
| RF8 (Jin 2025) [57] | Treatment — VEN/AZA response | 3 | n=498 across four independent cohorts | Research; closest to CRL-AML 4 candidate |
| MM-AI-AML [63] | Toxicity — myelosuppression severity | 3 | Internal + external (AUC 0.78) | Research |
| MAGIC-DR [77] | Flow MRD — interpretable XGBoost | 3 | Concordance with expert gating | Research |
| CCADDAS [29] | Flow MRD — B-ALL (AML pending) | 3 | Multi-centre, cloud-based pipeline | Research |
| Masurekar 2026 transplant [93] | Transplant — pre-HCT survival | 3 | n=252 external cohort with uniform MRD | Research |
| Hernández-Boluda 2025 (myelofibrosis) [100] | Transplant — survival post-HCT | 3 | Random survival forest with public web tool | Research; web calculator |
| Schmutz ChatGPT-4 MTB [67] | LLM — molecular tumor board | 3 | Real-world MTB cases (oncology, not AML-specific) | Research |
| TrialGPT [68] | LLM — patient-trial matching | 3 | Validation across three cohorts; user study | Research/NIH-developed |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.