Preprint
Review

This version is not peer-reviewed.

Artificial Intelligence-Driven Laboratory Technologies for Precision Diagnostics and Global Health Surveillance: Critical Narrative Review

Submitted:

21 August 2026

Posted:

24 August 2026

You are already at the latest version

Abstract
Artificial intelligence (AI) is increasingly applied to laboratory medicine as health systems confront infectious diseases, antimicrobial resistance, chronic disease burdens, and increasingly complex diagnostic requirements. This critical narrative review examines AI-enabled laboratory technologies for precision diagnostics and global health surveillance. A structured literature search was conducted across major biomedical, multidisciplinary, engineering, and health-informatics databases, supported by supplementary citation searching. Included evidence addressed diagnostic accuracy, laboratory workflow, digital pathology, microscopy, molecular diagnostics, biosensors, surveillance, implementation, governance, and equity. Findings indicate that AI can improve pattern recognition, automate selected laboratory processes, support molecular interpretation, strengthen digital pathology, and enhance outbreak detection through integrated clinical, genomic, and epidemiological data. However, strong technical performance does not consistently translate into clinical validity, external generalisability, improved patient outcomes, or equitable implementation. Major limitations include retrospective study designs, weak external validation, algorithmic bias, cybersecurity risks, fragmented interoperability, workforce shortages, infrastructure constraints, and underrepresentation of low-resource settings. AI should therefore augment, rather than replace, qualified laboratory professionals. Sustainable implementation requires representative datasets, rigorous validation, interoperable information systems, cybersecurity safeguards, skilled personnel, regulatory oversight, and context-sensitive infrastructure. When these conditions are met, AI-enabled laboratories can strengthen precision diagnostics, surveillance capacity, and global health decision-making across diverse health system contexts.
Keywords: 
;  ;  ;  ;  ;  ;  

1. Introduction

Science laboratory technology underpins diagnosis, biomedical research, environmental monitoring, and public-health protection by generating reliable evidence for clinical and population-level decisions. Its importance has increased as health systems confront emerging infections, antimicrobial resistance, pandemics, chronic diseases, and increasingly complex molecular and imaging-based diagnostic requirements. However, conventional laboratory systems remain constrained by manual interpretation, delayed workflows, workforce shortages, fragmented information systems, and unequal access. Fleming et al. (2021) estimate that 47% of the global population has little or no access to diagnostics, while only about 19% of people in low-income and lower-middle-income countries can access basic diagnostics at primary-care level.
Artificial intelligence (AI) has emerged as a potential response to some of these limitations through automation, pattern recognition, prediction, image interpretation, and clinical decision support. Machine learning can analyse multidimensional laboratory data, identify subtle patterns, and support faster interpretation across microscopy, pathology, molecular diagnostics, and surveillance. Nevertheless, adoption remains uneven. Cadamuro et al. (2025) found that only 25.6% of surveyed European laboratories reported an ongoing AI project, demonstrating a substantial gap between technological development and routine implementation.
Accordingly, this paper critically examines AI-enabled laboratory technologies for precision diagnostics and global disease surveillance. It evaluates AI applications in modern laboratory systems, diagnostic imaging, laboratory information management systems, molecular diagnostics, outbreak detection, and epidemiological surveillance. It also examines ethical, regulatory, cybersecurity, infrastructural, and equity challenges, while proposing strategies for responsible implementation. The review asks how AI affects diagnostic accuracy, efficiency, reproducibility, surveillance capability, and professional practice, and which risks may limit these benefits. This study is significant for laboratory scientists, clinicians, policymakers, researchers, educators, and global-health institutions because effective AI adoption requires not only technical performance, but also validation, workforce readiness, secure infrastructure, transparent governance, and equitable access. These requirements determine whether innovation becomes useful, sustainable, and relevant rather than technologically impressive alone.

2. Literature Review

2.1. Evolution of AI and Diagnostic Laboratory Applications

Artificial intelligence (AI) marks a shift from predefined laboratory automation towards systems that learn relationships, classify patterns, and support prediction. Hou et al. (2024) argue that AI can operate across pre-analytical, analytical, and post-analytical stages. Baron (2023) similarly notes that machine learning (ML) can identify multidimensional relationships that conventional rules may overlook. However, technological sophistication has not produced equivalent clinical adoption. Cadamuro et al. (2025) found that only 25.6% of surveyed European laboratories reported an ongoing AI project. This implementation gap indicates that technical capability remains constrained by cost, infrastructure, data access, workforce readiness, and regulatory uncertainty. Moreover, Randell et al. (2014) reported 99.5% autoverification using conventional rules-based automation. Therefore, improved laboratory efficiency cannot automatically be attributed to ML.
Laboratory information systems provide another route for applying AI beyond image interpretation. Wang et al. (2021) reported an ML-based autoverification system across 52 biochemistry testing items, achieving an 89.60% passing rate. Seok et al. (2024) found that ML also improved sample-misidentification detection compared with conventional methods during external validation. These results indicate that AI can support error recognition and workflow prioritisation, not only diagnostic classification. However, Seok et al. (2024) partly simulated misidentification errors, which may not reproduce every real pre-analytical failure. Therefore, workflow gains should be judged against existing laboratory controls, prospective error patterns, and the additional complexity introduced by algorithmic systems.
Diagnostic imaging provides some of the strongest evidence for laboratory AI. Zini et al. (2023) evaluated blood-film morphology using 591 samples across two Italian university centres. Sensitivity reached 83.8% for neoplastic cells and 98.8% for immature granulocytes, yet performance varied across morphological categories. Wang et al. (2023) likewise analysed 12,708 images from 380 malaria patients. Their deep-learning model identified 1,116 parasites within 13 seconds, with sensitivity exceeding 96.8% and specificity exceeding 99.3%. Huang et al. (2024) further demonstrated smartphone-assisted leukocyte classification with 98.2% mean accuracy. These findings support automated microscopy, particularly where specialist capacity is limited. Nevertheless, McGenity et al. (2024) caution that controlled image datasets may inadequately capture staining variation, artefacts, equipment differences, and difficult routine specimens. Thus, exceptional internal performance does not establish reliable performance across laboratories.
Digital pathology extends this potential towards tumour detection, grading, molecular prediction, and prognosis. McGenity et al. (2024) synthesised 100 studies involving more than 152,000 whole-slide images. Their meta-analysis reported pooled sensitivity of 96.3% and specificity of 93.3%. Yet 99% of included studies had at least one high or unclear bias or applicability concern. By contrast, Wang et al. (2024) developed CHIEF using 60,530 slides and validated it with 19,491 slides across 24 hospitals. This wider validation strengthens claims of generalisability, although prospective clinical effectiveness remains unproven. Similar caution applies to molecular diagnostics. Cheng et al. (2023) showed AlphaMissense could classify about 89% of possible human missense variants. However, Kong et al. (2025) found lower concordance with clinically curated pathogenic variants. Ardila et al. (2024) also reported variable antimicrobial-resistance prediction across datasets. Accordingly, AI can expand diagnostic reach, but computational performance cannot replace clinical validation, confirmatory testing, or context-sensitive interpretation.

2.2. Precision Diagnostics, Surveillance and Global Health Equity

Artificial intelligence is also reshaping laboratory workflow, point-of-care testing, precision diagnostics, and disease surveillance. Wang et al. (2021) developed an ML autoverification system covering 52 clinical biochemistry tests. The system achieved an 89.60% passing rate and reduced invalid reports requiring review by about 80%. Seok et al. (2024) similarly examined sample-misidentification detection using 397,751 development observations and 215,339 external observations. ML models achieved external balanced accuracy between 0.760 and 0.836, exceeding conventional approaches. However, some errors were simulated rather than naturally observed. This limitation reduces certainty about performance across complex pre-analytical failures. Furthermore, Randell et al. (2014) demonstrated that rules-based autoverification can achieve very high release rates. These comparisons suggest that AI should be adopted where it adds measurable value, rather than replacing effective conventional systems indiscriminately.
Portable technologies illustrate a different benefit. Antonelli et al. (2024) argue that combining ML, biosensors, and microfluidics can improve automated signal interpretation. Huang et al. (2024) demonstrated smartphone-assisted microscopy with strong leukocyte-classification performance and virtual Giemsa-style imaging. Such systems could reduce dependence on large central laboratories. However, proof-of-concept performance does not establish reliability under variable temperatures, maintenance limitations, poor connectivity, or differing operator skills. Precision diagnostics presents a comparable tension. Wang et al. (2024) showed that CHIEF could predict cancer presence, tumour origin, molecular characteristics, and prognosis from pathology data. Hoang et al. (2024) also used deep learning to predict DNA-methylation-defined brain tumour classes from histopathology. Nevertheless, image-based associations do not directly measure molecular alterations. Confirmatory molecular testing therefore remains necessary when treatment depends on validated biomarker status.
At population level, AI may support disease surveillance by combining laboratory, epidemiological, environmental, climatic, wastewater, and digital information. Villanueva-Miranda et al. (2025) report that ML, deep learning, and natural-language processing are increasingly used for infectious-disease early warning. Brunker (2024) further highlights portable sequencing during the 2016 Zika outbreak in Brazil as an operational precedent for field-based genomic surveillance. Yet surveillance algorithms cannot compensate for weak source data, poor interoperability, or incomplete reporting. These limitations become more visible within One Health surveillance. Singh et al. (2024) argue that integrated human, animal, environmental, and food-system information could strengthen cross-sector early warning. However, the supplied 2026 scoping review reports that 89.3% of minimum-dataset surveillance research focused on the human-health interface. Animal and environmental interfaces each appeared in no more than 10.7% of studies. Therefore, One Health remains conceptually persuasive but operationally fragmented.
This fragmentation has major equity implications. Oduoye et al. (2024) argue that AI could partially address shortages of laboratory specialists and diagnostic access in low- and middle-income countries (LMICs). Portable microscopy may be particularly useful where central infrastructure is limited (Huang et al., 2024). Conversely, inadequate datasets, financing, connectivity, and regulatory capacity can restrict safe deployment (Oduoye et al., 2024). Wong et al. (2025) therefore caution against simply exporting algorithms developed in high-income settings. AI may reduce inequality when designed around local constraints, but it may widen inequality when advanced systems remain concentrated in well-resourced institutions. Local validation and workforce support are therefore essential for benefit.

2.3. Validation, Governance, Evidence Limitations and Research Gap

The challenge to laboratory AI concerns governance, validation, safety, and evidence quality. Ethical concerns include privacy, fairness, accountability, transparency, consent, and responsibility for incorrect automated decisions (Spies et al., 2024). Chen et al. (2024) show that bias can emerge from data selection, measurement practices, model development, and deployment conditions. Xu et al. (2024) similarly document subgroup fairness problems in deep-learning medical-image systems. However, bias does not justify rejecting AI outright, because human diagnostic systems also contain variation. The more defensible position is that bias should be identified, quantified, monitored, and corrected before clinical deployment (Spies et al., 2024). Regulatory frameworks increasingly reflect this requirement. The European AI Act introduced enhanced controls for high-risk health AI, including data governance, transparency, human oversight, and risk management (European Commission, 2024).
Cybersecurity creates a related implementation risk. Greater digital interconnection exposes laboratories to ransomware, unauthorised access, service disruption, and manipulation of clinical records (ENISA, 2023; WHO, 2024). ENISA (2024) analysed 487 health-sector incidents, with ransomware representing 45% and data breaches 28%. Importantly, AI is not the sole cause of this exposure. The attack surface arises from connected digital infrastructures. Nevertheless, AI-dependent laboratories may experience greater operational disruption when middleware, electronic records, or decision-support systems become unavailable (WHO, 2024). Safe implementation therefore requires resilient fallback procedures alongside cybersecurity controls.
Validation remains important. Spies et al. (2024) argue that laboratory AI should undergo external validation, prospective evaluation, performance monitoring, and documented change control. Validation should examine generalisability, fairness, interoperability, metric selection, and label quality. This requirement matters because patient populations, instruments, assays, and workflows can change after deployment. Collins et al. (2024) likewise strengthen reporting expectations through TRIPOD+AI, which promotes transparent development and evaluation. The United States Food and Drug Administration (FDA, 2025) further emphasises total-product-life-cycle governance, including testing, bias, modification, transparency, and post-market monitoring. Regulation can strengthen safety, although compliance costs may also delay smaller developers and laboratories.
Viewed critically, the literature still overstates readiness when benchmark accuracy is treated as clinical utility. McGenity et al. (2024) report excellent pathology accuracy, but retrospective designs, heterogeneity, and bias concerns remain widespread. Kong et al. (2025) similarly demonstrate that AlphaMissense benchmark capability does not translate perfectly into patient-level variant interpretation. Moreover, Spies et al. (2024) note that proof-of-concept studies substantially outnumber routine clinical implementations. Cadamuro et al. (2025) reinforce this translation gap through low European laboratory adoption. Geographical inequality further limits generalisability because infrastructure, genomic capacity, connectivity, financing, and trained personnel remain unevenly distributed (Oduoye et al., 2024; Brunker, 2024).
Taken together, the principal research gap is not whether AI can perform laboratory tasks. Existing evidence already supports capability across microscopy, pathology, genomics, workflow management, precision diagnostics, and surveillance (Hou et al., 2024). Rather, evidence remains fragmented across technologies, disciplines, surveillance systems, and geographical contexts. Prospective effectiveness, external validation, equitable implementation, and integrated laboratory-system performance remain less established (McGenity et al., 2024; Spies et al., 2024). Future research should therefore assess whether strong technical performance remains clinically valid, operationally useful, secure, and equitable across diverse health systems.

4. Methodology

4.1. Review Design and Search Strategy

This study adopted a critical narrative review supported by a structured literature search. This design was appropriate because the research question spans laboratory medicine, artificial intelligence (AI), diagnostics, engineering, and public-health surveillance. Snyder (2019) argues that narrative reviews can function as rigorous research methodologies when their procedures are transparent. Baethge et al. (2019) similarly support narrative synthesis for heterogeneous evidence, while Ferrari (2015) argues that systematic procedures reduce selective interpretation. Accordingly, SANRA informed overall review quality, although its validation limitations were recognised (Baethge et al., 2019).
The search covered PubMed/MEDLINE, Scopus, Web of Science Core Collection, IEEE Xplore, and Embase where accessible. Google Scholar and ScienceDirect were used supplementarily for citation tracking and difficult-to-locate studies. Bramer et al. (2017) found that 16% of relevant references appeared uniquely within one database, supporting multi-database searching. However, Gusenbauer and Haddaway (2020) caution that broad search coverage does not guarantee reproducibility. Search concepts therefore combined controlled vocabulary and free-text terms relating to AI, machine learning, laboratory medicine, diagnostics, pathology, microscopy, genomics, biosensors, and disease surveillance. Bramer et al. (2018) informed search-string construction, while PRISMA-S and PRESS principles guided reporting and peer checking (Rethlefsen et al., 2021; McGowan et al., 2016).

4.2. Study Selection, Data Extraction and Synthesis

Eligibility criteria were defined before final selection to reduce post hoc inclusion decisions (Higgins et al., 2024). Eligible studies examined AI within laboratory, diagnostic, or surveillance settings and reported technical, clinical, implementation, or public-health outcomes. Duplicates, unsupported editorials, purely technical studies without healthcare relevance, and inadequately reported conference abstracts were excluded. Records were deduplicated, screened by title and abstract, assessed at full text, and documented using an adapted PRISMA-style flow process (Page et al., 2021). Rayyan supported record organisation where required (Ouzzani et al., 2016).
A standardised extraction framework recorded authorship, year, country, study design, population, laboratory discipline, specimen type, algorithm, training data, validation approach, reference standard, diagnostic metrics, clinical outcomes, implementation conditions, fairness, security, and study limitations. Collins et al. (2024) informed extraction of prediction-model characteristics. Importantly, Andaur Navarro et al. (2021) found high risk of bias in 87% of assessed machine-learning prediction studies, demonstrating why accuracy alone is insufficient. Because substantial methodological heterogeneity was expected, statistical pooling was not assumed. Instead, evidence was organised through thematic synthesis and cross-study comparison following SWiM principles (Campbell et al., 2020).

4.3. Quality and Risk-of-Bias Assessment

Quality appraisal used study-design-specific frameworks rather than one generic instrument. QUADAS-2 assessed diagnostic-accuracy studies (Whiting et al., 2011), while PROBAST+AI assessed prediction-model quality and bias (Moons et al., 2025). STARD-AI informed reporting appraisal for AI diagnostic studies (Sounderajah et al., 2025). TRIPOD+AI supported prediction-model reporting assessment (Collins et al., 2024), CLAIM 2024 guided imaging and microscopy evidence (Tejani et al., 2024), and DECIDE-AI informed evaluation of early real-world implementation (Vasey et al., 2022).
This layered approach was necessary because methodological weaknesses can substantially inflate apparent AI performance. Liu et al. (2019) found that only 25 imaging studies used out-of-sample external validation. Nagendran et al. (2020) reported that only nine of 81 clinician-comparison studies were prospective, while 58 were at high overall risk of bias. Roberts et al. (2021) concluded that none of the reviewed COVID-19 imaging models were clinically ready. Likewise, Wynants et al. (2020) showed that externally validated models performed materially worse than internally assessed models. Therefore, technical accuracy was interpreted separately from clinical validity, external generalisability, reproducibility, and real-world usefulness.
Consequently, synthesis prioritised evidence strength, transparency, and transferability rather than headline performance estimates. Contradictory findings were examined against differences in datasets, validation settings, populations, laboratory workflows, and methodological quality. This approach enabled technical performance to be considered alongside implementation feasibility, equity, regulatory readiness, and relevance across diverse international health systems.

5. Conceptual Foundations of Artificial Intelligence-Enabled Laboratory Systems

Artificial intelligence (AI) should be distinguished carefully from conventional laboratory automation because the two concepts describe different computational functions. The OECD (2024) defines AI around machine-based inference that generates predictions, recommendations, decisions, or related outputs from supplied inputs. By contrast, conventional automation may execute fixed workflows without learning statistical relationships from data. Mencacci et al. (2023) therefore caution against treating every digitally enabled laboratory process as AI. This distinction matters because laboratory efficiency can improve without machine learning (ML). Norgan et al. (2020), for example, reported that radio-frequency identification (RFID) reduced anatomical pathology mislabelling incidents from 24 to six, representing a 75% reduction. The Mayo Clinic system linked RFID with electronic health records and laboratory information systems, yet RFID itself did not constitute ML. Consequently, evidence for automation should not be presented as direct evidence for AI effectiveness. Nevertheless, You et al. (2025), reviewing 144 laboratory ML studies, identify classification, prediction, automation, efficiency, and diagnostic expansion as important areas where AI can add functions beyond deterministic systems.
Machine learning provides the principal computational foundation for many laboratory AI applications. Supervised ML learns relationships between inputs and labelled outcomes, whereas unsupervised ML identifies latent structures without predefined outcome labels (Goodfellow et al., 2016). Semi-supervised learning combines smaller labelled datasets with larger unlabelled datasets, which is particularly relevant where expert medical annotation is expensive (Cheplygina et al., 2019). Reinforcement learning instead learns actions through rewards arising from interaction with an environment (Sutton & Barto, 2018). These categories are conceptually distinct, but they are not equally mature within laboratory medicine. You et al. (2025) show that current laboratory research remains concentrated around supervised classification and prediction. This reflects the availability of diagnostic labels and expert reference standards. However, Cheplygina et al. (2019) note that labelled medical datasets remain limited, creating incentives for semi-supervised and transfer-learning approaches. Therefore, the conceptual breadth of ML should not be confused with equivalent clinical readiness across its different forms.
Deep learning (DL) extends ML through multilayer computational architectures that learn increasingly complex representations from data (Goodfellow et al., 2016). Convolutional neural networks are especially suited to spatial image structures, which explains their importance in microscopy and pathology (Esteva et al., 2019). Matek et al. (2021) trained DL systems using 171,374 expert-annotated bone-marrow cell images from 945 patients, demonstrating the scale achievable in morphological classification. More recent architectures have expanded beyond conventional image models. Chandrashekar et al. (2024) trained Path-BigBird using 2.7 million pathology reports from six United States cancer registries, illustrating transformer-based extraction from unstructured laboratory narratives. Dasdelen et al. (2026) combined foundation-model features with transformer aggregation across more than 3.2 million blood-cell images. Their model achieved an overall malignant-versus-non-malignant area under the curve of 0.97, but lymphoma sensitivity reached only 0.40. This contrast demonstrates that architectural complexity and strong aggregate discrimination do not guarantee reliable performance for individual diagnostic classes.
Natural language processing (NLP) broadens laboratory AI beyond numerical and image-based data by converting unstructured clinical narratives into structured information. Sakai and Lam (2026) reviewed 65 studies on large language models for healthcare text classification after screening 826 records. Twenty-nine studies addressed clinical decision support, indicating growing operational relevance. However, more than 80% of included datasets were English-language, limiting linguistic generalisability. Chandrashekar et al. (2024) similarly demonstrated automated extraction of tumour site, histology, behaviour, laterality, and subsite from pathology reports. Path-BigBird achieved micro-F1 scores of 80.96 for histology and 72.53 for subsite classification. Yet the model was not superior on every task, because Clinical BigBird performed better for tumour site and laterality. Consequently, domain-specific NLP models should be evaluated task by task rather than assumed superior because of larger training corpora or specialised architecture.
Computer vision and predictive analytics further demonstrate the value and limitations of data-driven laboratory inference. Computer vision enables automated recognition, segmentation, classification, and quantification of clinically meaningful structures (Esteva et al., 2019). Evidence spans bone-marrow morphology, culture plates, malaria microscopy, and digital pathology (Matek et al., 2021; Signoroni et al., 2023). Yet field performance may differ materially from controlled validation. Horning et al. (2021) reported 94.3% accuracy on a World Health Organization malaria evaluation set. Das et al. (2022), however, found only 75.6% specificity when the EasyScan GO system was evaluated across 2,250 slides from 11 international sites. Slide quality materially influenced results. Predictive analytics shows a comparable problem. Iscoe et al. (2026) analysed 149,449 emergency-department encounters and achieved an AUROC of about 0.93 for urine-culture prediction. Conversely, Wong et al. (2021) externally validated a widely implemented sepsis model and obtained an AUROC of 0.63, with 67% of sepsis cases missed at a commonly used threshold. These contrasts show why independent validation is essential.
Explainable artificial intelligence (XAI) seeks to make model reasoning or influential inputs more understandable to users. Rosenbacke et al. (2024) found five studies in which XAI increased clinician trust. However, three studies found no significant effect, while two showed that explanations could either increase or reduce trust. Their review included only ten eligible studies, all with moderate or moderate-to-high risk of bias. Therefore, explainability should support calibrated trust rather than simply increase confidence. Ghassemi et al. (2021) further argue that plausible post-hoc explanations can create false reassurance without establishing model safety. XAI should therefore complement external validation, monitoring, and professional judgement rather than substitute for them.

6. Architecture of an Artificial Intelligence-Enabled Laboratory

The architecture of an AI-enabled laboratory begins with a data-acquisition layer capable of integrating analyser outputs, microscopy, sequencing, reports, clinical records, and surveillance data. Brehmer et al. (2024) demonstrated a Fast Healthcare Interoperability Resources (FHIR) framework that linked laboratory results, prescriptions, diagnoses, procedures, and reports across 1,302,988 patient encounters. This illustrates the analytical value of multimodal data integration. However, Edayan et al. (2024) argue that laboratory and hospital systems frequently use incompatible interfaces, data structures, and semantic conventions. Consequently, acquiring more data does not automatically create more reliable AI. The processing layer must therefore address cleaning, harmonisation, annotation, quality assurance, and ground-truth construction (Cheplygina et al., 2019). Yang et al. (2026) further note that laboratory variables can remain method-dependent across instruments and institutions. Data preprocessing should therefore be treated as a clinical governance task, not merely a computational step.
The analytical layer applies models to classification, prediction, segmentation, pattern recognition, and anomaly detection. You et al. (2025) identify convolutional networks, multilayer perceptrons, and tree-based models among recurring approaches in laboratory medicine. Yet no architecture is universally optimal. Model suitability depends on data modality, prevalence, outcome definition, and workflow. Dasdelen et al. (2026), for instance, demonstrate the value of image encoders and transformers in haematological morphology, whereas Iscoe et al. (2026) used XGBoost effectively for structured and textual urinary-tract infection prediction. The decision-support layer then translates analytical outputs into alerts, risk scores, recommendations, or triage decisions. However, Wong et al. (2021) show that poorly performing alerts can increase workload while still missing clinically important cases. Decision support must therefore be evaluated by its workflow consequences, calibration, and clinical usefulness rather than discrimination statistics alone.
Communication and human oversight complete the architecture. AI-generated information must move reliably between laboratory information systems, electronic health records, clinical services, and surveillance infrastructures. Edayan et al. (2024) found that 22 of 28 reviewed integration studies achieved their intended interoperability objectives, with Health Level Seven and FHIR prominent across the literature. Nevertheless, their review also identified weak validation, limited generalisability, incompatible data, security concerns, and evolving standards. Brehmer et al. (2024) show that FHIR can operationalise multimodal information across major clinical use cases, but technical connectivity does not resolve every semantic or governance problem. Human oversight therefore remains essential. The World Health Organization (WHO, 2021) argues that qualified professionals should retain responsibility for consequential AI-supported decisions. Yang et al. (2026) similarly report strong expert support for local verification of externally supplied models. Importantly, oversight does not imply rejection of automation. Bulten et al. (2021) found that AI assistance improved pathologists’ median agreement with expert prostate grading from κ=0.799 to κ=0.872. This evidence supports professional augmentation rather than autonomous replacement.
Taken together, these foundations show that an AI-enabled laboratory is not defined by one algorithm or automated device. Its performance depends on data acquisition, preprocessing, modelling, communication, decision support, and governance. Strong performance at one architectural layer cannot compensate for weaknesses elsewhere. Poor data quality, incompatible systems, unstable models, or weak oversight can undermine algorithms. Therefore, laboratory AI should be evaluated as a socio-technical system whose clinical value depends on coordinated technical and human components.

7. Machine Learning for Diagnostic Image Analysis

Machine learning has become increasingly important for automated microscopy because digitised specimens provide structured visual information suitable for computational classification. Matek et al. (2021) demonstrated this potential using 171,374 expert-labelled bone-marrow cell images from 945 patients. Their deep-learning approach achieved strong differentiation across multiple morphological cell classes. However, high performance on expert-labelled images does not necessarily establish equivalent performance under routine laboratory conditions. Das et al. (2022) show that staining quality, artefacts, microscope variation, and specimen preparation can materially influence field accuracy. Therefore, automated microscopy should be judged against both curated validation and routine operational conditions.
Haematological image analysis further illustrates this tension between technical capability and clinical generalisability. Dasdelen et al. (2026) analysed more than 3.2 million single-cell images from 6,115 patients and 495 healthy controls. Binary malignancy detection achieved sensitivity of 0.94, specificity of 0.83, and an area under the curve (AUC) of 0.97. These results indicate substantial potential for screening abnormal peripheral-blood morphology. Moreover, their proposed threshold reduced potentially unnecessary bone-marrow aspirations from 13.5% to 8.7% while identifying all 46 acute-leukaemia cases. However, performance varied sharply by disease category. Lymphoma sensitivity was only 0.40, despite strong aggregate discrimination. The retrospective, single-centre design also limits transferability. Consequently, strong overall AUC values can conceal weaknesses in clinically important subgroups.
Comparable issues appear in microbiological image analysis. Signoroni et al. (2023) evaluated DeepColony using 5,051 clinical urine-culture plates from a large United States laboratory. Overall agreement with laboratory technologists reached 95.4%, with κ=0.920. Agreement was 99.2% for no-growth cultures and 95.6% for positive cultures. These findings support computer vision for automated culture interpretation. Rajaonison et al. (2022) similarly trained antimicrobial-susceptibility interpretation software using 18,072 antibiogram images and validated the model using 5,100 images. Agreement reached approximately 97%. Nevertheless, validation involved only three commonly isolated bacterial species. Thus, excellent technical agreement within restricted microbiological datasets should not be interpreted as broad diagnostic generalisability.
Malaria microscopy provides an especially strong example of the difference between benchmark and field performance. Horning et al. (2021) evaluated a fully automated system on 55 World Health Organization reference slides. The system achieved 94.3% accuracy, 86.7% sensitivity, and 100% specificity. However, Das et al. (2022) evaluated EasyScan GO using 2,250 slides across 11 international sites. Field sensitivity reached 91.1%, while specificity declined to 75.6%. Sensitivity also fell to 57% where parasite density was below 200 parasites per microlitre. Furthermore, only 23% of density estimates were within ±25% of the reference value. These differences strongly suggest that curated slide sets may overestimate routine performance. Slide preparation, staining, parasite density, and local laboratory conditions therefore remain central determinants of clinical reliability.
Digital pathology provides stronger evidence for human-AI collaboration than complete professional replacement. Bulten et al. (2021) assessed 14 observers grading 160 prostate biopsies before and after AI assistance. Agreement with the expert standard improved from κ=0.799 to κ=0.872, with a statistically significant difference. External validation across 87 cases also showed improvement. However, five of the 14 observers did not improve after receiving AI support. Artefacts and unfamiliar abnormalities also reduced algorithm reliability. Consequently, the evidence does not justify assuming that AI support benefits every professional equally. Instead, it supports augmentation in which AI strengthens consistency while qualified pathologists retain interpretative responsibility.
Multimodal models extend image analysis by combining laboratory biomarkers, physiological signals, clinical variables, and other structured information. Ayoub et al. (2025) analysed 2,258 Mayo Clinic patients receiving immune-checkpoint inhibitor therapy. Their combined electrocardiogram and structured electronic-health-record model achieved an AUROC of 0.717 and a negative predictive value of 0.98. However, the positive predictive value was only 0.08 because adverse events were uncommon. This finding demonstrates that apparently useful discrimination can coexist with poor positive prediction. Collins et al. (2024) therefore argue that sensitivity, specificity, precision, recall, calibration, threshold-dependent measures, and prevalence should accompany AUROC. Likewise, Dasdelen et al. (2026) achieved an AUC of 0.97 while lymphoma sensitivity remained only 0.40. These examples confirm that no single performance statistic establishes clinical usefulness.
The broader limitations of image-based AI arise from annotation error, class imbalance, staining variation, artefacts, equipment differences, dataset shift, and insufficient external validation. Das et al. (2022) observed major performance changes associated with slide quality during multicountry malaria evaluation. Dasdelen et al. (2026) similarly reported acute-leukaemia sensitivity of 0.94 on one external dataset but 0.76 under a different staining background. Wong et al. (2021) provide an additional warning from predictive medicine, where a widely implemented sepsis model achieved an external AUROC of only 0.63. Taken together, these findings support treating external validation as a fundamental development requirement rather than an optional final exercise.

8. Laboratory Information Management Systems and Intelligent Automation

Laboratory information systems (LISs) manage specimen data, analyser interfaces, workflow information, results, and communication with broader health information infrastructures. However, a LIS is not inherently an AI system. Edayan et al. (2024) note that many core functions remain deterministic information-management processes. AI capabilities become relevant when predictive models, anomaly detection, intelligent result review, or adaptive decision support are embedded within these systems. You et al. (2025) argue that such integration can extend LISs from transactional platforms towards predictive environments. Nevertheless, Yang et al. (2026) report that implementation remains uneven. Among surveyed experts, 88% identified inadequate information-technology infrastructure as a barrier, while 78% identified insufficient implementation guidance and inadequate ML expertise. In addition, 66% reported difficulties integrating ML into existing workflows. These findings challenge claims that algorithm availability alone produces laboratory transformation.
Specimen tracking further demonstrates why automation and AI should remain conceptually distinct. Norgan et al. (2020) reported that RFID implementation reduced anatomical pathology mislabelling incidents from 24 to six. The system also helped recover three specimens that might otherwise have been considered lost. Tavares da Souza et al. (2024) reported similar operational benefits in a Washington DC blood bank. Before RFID implementation, 4.0–4.3% of red blood-cell units were discarded annually because of expiry. Following implementation, annual discard fell below 1%, with p<.001. These examples demonstrate substantial improvements in traceability and resource management. However, neither application inherently required ML. Therefore, digital tracking should be recognised as intelligent automation rather than incorrectly presented as evidence of AI performance.
Machine learning may add greater value where laboratories require multivariate error detection or adaptive result review. Seok et al. (2024) trained and internally validated models using 397,751 laboratory records and externally evaluated them using another 215,339 records. Deep neural networks and XGBoost achieved AUCs between approximately 0.834 and 0.903, while conventional delta checks achieved values between 0.705 and 0.816. These results indicate technical improvement in detecting sample misidentification. However, the errors were experimentally simulated rather than derived entirely from naturally occurring incidents. Consequently, the study establishes model discrimination more convincingly than real-world error reduction.
Autoverification provides another important distinction between conventional and AI-enabled systems. Wang et al. (2021) developed an ML autoverification model using 52 laboratory test items and demographic variables. The system achieved an 89.60% passing rate and a false-negative rate of approximately 0.095%. It also produced about 80% fewer invalid reports than the conventional rules-based engine. Nevertheless, the study was conducted within one Chinese centre, limiting generalisability across populations, instruments, and workflows. Mencacci et al. (2023) further note that autoverification predates modern AI and can operate through deterministic laboratory rules. Therefore, the key distinction concerns whether systems merely apply fixed thresholds or learn multivariate decision boundaries from data.
Interoperability ultimately determines whether AI-generated information can move reliably across analysers, LISs, electronic health records, hospitals, and surveillance systems. Edayan et al. (2024) found that 22 of 28 reviewed integration studies achieved their intended interoperability objectives. Health Level Seven and Fast Healthcare Interoperability Resources were important standards within this literature. Brehmer et al. (2024) further demonstrated multimodal FHIR integration across more than 1.3 million clinical encounters. However, technical standards do not eliminate semantic incompatibility, validation weaknesses, cybersecurity risks, or problems created by system modification. Yang et al. (2026) therefore argue for local verification and lifecycle monitoring. The United States Food and Drug Administration (FDA, 2026) similarly requires cybersecurity considerations throughout medical-device design and premarket documentation. Consequently, AI-enabled laboratories require continuous governance, monitoring, professional oversight, and secure interoperability rather than one-time technical validation.

13. Contribution to Global Development

Artificial intelligence (AI)-enabled laboratories could strengthen global health systems by expanding diagnostic capacity, improving surveillance, supporting research, and increasing workforce productivity. Fleming et al. (2021) identify diagnostics as a major health-system bottleneck, estimating that 47% of the global population has little or no diagnostic access. In low-income and lower-middle-income countries, only about 19% of people can access basic diagnostics at primary-care level. AI may partly address this gap through automated interpretation and task-shifting. Wadie et al. (2026) reviewed 20 point-of-care AI studies involving about 78,000 patients across 15 countries and reported median sensitivity of 93.6% and specificity of 90.6%. Thirteen studies demonstrated some task-shifting, suggesting that non-specialists may perform selected diagnostic activities with AI support.
However, AI should be viewed as a health-system multiplier rather than an independent solution. The World Health Organization (WHO) and partners estimate that nearly one billion people in lower-income countries use health facilities without reliable electricity, while only half of hospitals in sub-Saharan Africa have dependable power (WHO et al., 2023). These constraints can prevent implementation before algorithmic performance becomes relevant. Pandemic preparedness illustrates the same dependency. WHO (2022) promotes integrated genomic surveillance and aims for all 194 Member States to access timely pathogen sequencing by 2032. Africa CDC (2025) reports that African countries capable of basic public-health sequencing increased from seven in 2019 to 46 by 2025. Its AGARI platform further supports continental analysis of pathogens such as mpox, cholera, Ebola, and Marburg. Yet earlier detection only improves outcomes when linked to epidemiological investigation, clinical action, governance, and resource mobilisation.
AI may also contribute to research and economic development. Liu et al. (2023) used deep learning to identify abaucin as a candidate antibiotic against multidrug-resistant Acinetobacter baumannii. El Arab et al. (2025) similarly found applications across epitope prediction, vaccine design, molecular docking, and immune-response assessment. Nevertheless, computational candidate generation cannot replace laboratory validation, clinical trials, regulation, or manufacturing. Economic claims require similar caution. El Arab and Al Moosa (2025) found favourable findings across many AI economic evaluations, while Fleming et al. (2021) reported diagnostic benefit-cost ratios reaching about 24.4:1 for drug-susceptible tuberculosis in Bangladesh. However, lifecycle costs, cybersecurity, maintenance, training, and integration may be underestimated. Therefore, developmental value depends on infrastructure, implementation quality, and equitable distribution of benefits.

14. Applications in Low- and Middle-Income Countries

Low- and middle-income countries (LMICs) may gain substantially from AI where diagnostic specialists, laboratory infrastructure, and referral capacity are limited. Wadie et al. (2026) found task-shifting in 65% of reviewed point-of-care imaging studies, with some studies requiring relatively short operator training. This suggests that AI could redistribute specialist expertise towards rural and underserved settings. However, 70% of included studies had high or very high risk of bias, and none demonstrated improved patient-level outcomes. Median sensitivity was also lower in lower-income settings than in high-income settings. Thus, technical feasibility should not be equated with improved population health.
Local validation is particularly important because externally trained models may not represent local populations or laboratory conditions. Chen et al. (2023) show that bias can arise through data acquisition, annotation, biological variation, sampling, and dataset shift. Muralidharan et al. (2024) examined 692 FDA-approved AI-enabled medical devices and found that only 3.6% reported race or ethnicity, while 99.1% omitted socioeconomic information. Such reporting gaps weaken claims of universal applicability. The appropriate response is not to assume that foreign models will fail, but to test transportability empirically within local populations, instruments, workflows, and disease patterns.
Implementation also depends on sustainable capacity. Russell et al. (2023) developed six competency domains and 25 subcompetencies for healthcare professionals using AI, including evidence evaluation, ethics, workflow adaptation, and continuing development. Onywera et al. (2025) similarly argue that African genomic surveillance requires permanent bioinformatics and pathogen-genomics expertise rather than temporary training programmes. Local innovation may strengthen this process by aligning technologies with regional disease burdens and resource constraints. Africa CDC’s AGARI initiative illustrates movement towards locally governed analytical infrastructure. However, local development is not automatically safer or more equitable. Representative data, independent validation, cybersecurity, quality management, and sustainable financing remain necessary. Consequently, successful LMIC implementation requires local participation in problem selection, dataset design, validation, governance, and benefit sharing.

16. Data Security and Cybersecurity Risks

Artificial intelligence (AI)-enabled laboratories create significant cybersecurity responsibilities because laboratory information systems hold sensitive diagnostic, demographic, and genomic information. Security must therefore protect confidentiality, integrity, and availability simultaneously (FDA, 2026). The consequences of failure can extend beyond privacy breaches into direct disruption of patient care. NHS England (2024) reported that the 2024 Synnovis ransomware attack contributed to 10,152 postponed acute outpatient appointments and 1,710 postponed elective procedures across two heavily affected trusts. Cross-matching services were also disrupted, increasing reliance on universal O-type blood products. This case demonstrates that laboratory cybersecurity is fundamentally a patient-safety and continuity issue rather than merely an information-technology concern.
Broader evidence reinforces this risk. ENISA (2023) analysed 215 publicly reported European health-sector incidents and found ransomware represented about 54%, although the dataset does not provide laboratory-specific incidence rates. Genomic information creates additional concerns because biological identifiers cannot simply be replaced after compromise. Erlich et al. (2018) demonstrated that long-range familial searching can enable re-identification from ostensibly de-identified genomic data. The 23andMe investigation similarly found that 18,222 accounts were accessed through credential stuffing, while information concerning almost seven million customers became affected through relational features (OPC & ICO, 2025). Nevertheless, digital laboratories are not inherently unsafe. NIH (2024) and FDA (2026) support controlled access, recognised security standards, strong authentication, encryption, logging, backups, segmentation, and lifecycle monitoring.
AI also introduces adversarial-security risks. Finlayson et al. (2019) argue that deliberately manipulated inputs can alter model predictions without obvious visible changes. However, much of this evidence remains experimental rather than evidence of frequent malicious clinical attacks. The appropriate conclusion is therefore that conventional diagnostic accuracy does not demonstrate adversarial robustness. Cybersecurity governance should consequently combine technical safeguards, staff training, incident exercises, tested recovery procedures, and executive accountability.

17. Validation, Regulation, and Quality Assurance

Strong algorithmic performance cannot establish clinical validity unless both the underlying laboratory measurement and the AI interpretation are rigorously validated. Miller and Valdes (2025) argue that calibration, harmonisation, pre-analytical error, interference, and physiological variation can alter downstream machine-learning reliability. Therefore, validation must assess not only discrimination but also calibration, data quality, explainability, generalisability, and clinically meaningful failure modes. Clinical validation should additionally reproduce intended patient populations, users, disease prevalence, and workflow conditions.
Jung et al. (2022) provide useful evidence from independent validation of prostate-biopsy AI. Their study used 593 whole-slide images, including 463 adenocarcinomas. AI assistance improved weighted grade-group agreement from 0.876 to 0.925 and reduced mean assessment time from 55.7 to 36.8 seconds. However, user validation involved one general pathologist, limiting conclusions about wider clinical effectiveness. Zech et al. (2018) demonstrate why external validation remains essential. Across 158,323 chest radiographs from three United States healthcare systems, models frequently performed worse externally and could identify hospital-specific characteristics. Apparent disease prediction can therefore partly reflect institutional signals rather than transferable pathology.
Prospective evaluation is also necessary because static retrospective datasets cannot reproduce every human-AI interaction. CONSORT-AI formalises reporting expectations for clinical trials involving AI and emphasises AI errors and human interaction (Liu et al., 2020). Regulation increasingly follows the same lifecycle logic. IMDRF (2025) frames Good Machine Learning Practice as a continuous responsibility, while FDA (2025) permits predetermined change-control plans for planned software modifications. This approach attempts to balance adaptive improvement with traceability, revalidation, and safety. Great Britain’s post-market requirements similarly require continued monitoring of in-vitro diagnostic devices (MHRA, 2025).
Post-deployment surveillance is particularly important because models can drift. Ji et al. (2023) examined 1,010 prediction tasks covering 242 outcomes and found population-level temporal shifts in 9.7%, while 93.0% showed shifts affecting at least one subgroup. However, Zhou et al. (2023) show that retraining on only recent data is not always superior to retaining historical information. Drift management should therefore be evidence-driven rather than automatic. AI should ultimately be incorporated into established laboratory quality systems, including ISO 15189:2022, while adding controls for datasets, model versions, updates, monitoring, and revalidation.

19. Discussion

19.1. Summary of Major Findings

The review shows that artificial intelligence (AI) has expanded across diagnostic imaging, digital pathology, molecular diagnostics, laboratory information systems, point-of-care testing, disease surveillance, and workflow management. Reported benefits include faster classification, improved pattern recognition, reduced repetitive work, and better use of complex multimodal data. However, these benefits remain uneven because strong technical performance does not automatically establish clinical validity, generalisability, affordability, or equitable implementation. Cadamuro et al. (2025) found that only 25.6% of surveyed European laboratories reported ongoing AI projects, indicating that implementation remains substantially behind research activity. Similar concerns arise in low-resource settings, where infrastructure, connectivity, workforce capacity, and local validation can determine whether AI produces meaningful benefit (Wadie et al., 2026; Yang et al., 2026).

19.2. Diagnostic Efficiency and Accuracy

AI can improve diagnostic performance in selected settings, but the evidence does not support a universal claim of superior accuracy or productivity. Jung et al. (2022) found that AI assistance improved prostate-biopsy grading agreement and reduced assessment time. Bulten et al. (2021) similarly showed improved pathologist agreement when AI supported grading. However, performance can deteriorate outside development environments. Zech et al. (2018) demonstrated poorer external performance across institutions, while Das et al. (2022) showed that malaria microscopy specificity fell under multicountry field conditions. Therefore, efficiency gains should be evaluated alongside calibration, class-specific performance, workflow effects, and external validation rather than headline accuracy alone (Miller & Valdes, 2025).

19.3. Human–Artificial Intelligence Collaboration

The strongest current evidence supports augmentation rather than autonomous replacement of qualified laboratory professionals. AI can reduce repetitive interpretation and highlight abnormal patterns, but professionals remain responsible for contextual judgement, uncertainty, and clinically consequential exceptions. Bulten et al. (2021) showed that AI-supported pathologists performed better overall, although not every observer improved. This variability demonstrates that human–AI interaction depends on user experience, interface design, and task characteristics. Arvai et al. (2025) also identify concerns about deskilling, automation bias, and professional autonomy. Consequently, laboratories require systems that support calibrated trust, preserve meaningful human oversight, and allow professionals to challenge or override algorithmic recommendations.

19.4. Global Surveillance Impact

Interconnected AI-enabled laboratories could strengthen outbreak detection by combining diagnostic, genomic, epidemiological, environmental, and clinical information. WHO (2022) argues that genomic surveillance should connect sampling, sequencing, analysis, and public-health action. Africa CDC (2025) illustrates expanding regional capacity, with pathogen-sequencing capability increasing across African countries and the AGARI platform supporting continental genomic analysis. Nevertheless, detection alone does not guarantee effective response. Surveillance outputs must connect with epidemiological investigation, logistics, governance, communication, and treatment capacity. Therefore, AI should be assessed across the complete detection-to-action pathway rather than by forecasting performance alone.

19.5. Equity Considerations

AI could reduce inequalities by extending specialist interpretation to underserved settings, yet it may also reproduce existing disparities. Wadie et al. (2026) found evidence of task-shifting in point-of-care imaging, but performance was weaker in lower-income settings and cross-context validation remained inadequate. Muralidharan et al. (2024) also found major demographic reporting gaps among approved AI-enabled medical devices. These findings show that global applicability cannot be assumed from aggregate performance. Equity therefore requires representative datasets, local validation, affordable infrastructure, workforce development, and fair governance of data and benefits.

19.6. Evidence Gaps

The literature remains dominated by retrospective development studies, with fewer prospective evaluations, independent external validations, and patient-outcome studies. Economic evidence is also limited by inconsistent costing methods and incomplete assessment of lifecycle expenditure (El Arab & Al Moosa, 2025). Developing regions remain underrepresented in many datasets, while implementation studies frequently originate from well-resourced institutions. Future evidence should therefore prioritise prospective deployment, subgroup analysis, transparent reporting, real-world clinical outcomes, and comparative cost-effectiveness. Reporting frameworks such as STARD-AI and CONSORT-AI can improve methodological transparency (Liu et al., 2020; Sounderajah et al., 2025).

19.7. Implications for Science Laboratory Technology

AI changes science laboratory technology by expanding professional responsibilities beyond specimen analysis towards informatics, validation, data governance, cybersecurity, and algorithm monitoring. Russell et al. (2023) identify competencies involving AI knowledge, evidence evaluation, ethics, workflow adaptation, and continuing development. Onywera et al. (2025) similarly emphasise sustainable bioinformatics and genomics expertise. Science laboratory technology education should therefore integrate computational literacy without weakening core laboratory science. Professionals must understand both how AI systems operate and when their outputs should be questioned.

20. Recommendations

Governments should develop coordinated policies for laboratory digitalisation, AI regulation, cybersecurity, interoperability, and equitable infrastructure investment. Public investment should prioritise reliable electricity, connectivity, quality laboratory systems, and local validation capacity rather than purchasing algorithms in isolation. Laboratory institutions should establish multidisciplinary AI governance committees involving laboratory scientists, clinicians, informaticians, cybersecurity specialists, quality managers, and patient-safety personnel. These committees should oversee model selection, validation, deployment, monitoring, incident response, and retirement.
Researchers should use diverse and representative datasets, report subgroup performance, conduct independent external validation, and prioritise prospective clinical evaluation. Transparent reporting should follow appropriate frameworks, including STARD-AI, TRIPOD+AI, CONSORT-AI, and related guidance. Technology developers should design systems that are interoperable, secure, affordable, auditable, and suitable for local workflows. Procurement arrangements should protect data portability and reduce unnecessary vendor lock-in. Academic institutions should incorporate AI, bioinformatics, data science, cybersecurity, ethics, and regulatory science into laboratory curricula while preserving training in analytical quality and diagnostic reasoning. International organisations should support shared technical standards, responsible cross-border surveillance, financing, technology transfer, and capacity development, particularly in resource-constrained regions.

21. Future Research Directions

Future research should move from proof-of-concept performance towards resilient, equitable, and clinically integrated AI. Federated learning deserves investigation because it may support multi-institutional model development without centralising sensitive data. Explainable AI should be evaluated for whether explanations improve decisions, not merely whether users prefer them. Multimodal research should examine how laboratory, imaging, genomic, and clinical information can be integrated without creating unnecessary complexity.
Further studies should assess AI-supported antimicrobial-resistance surveillance, edge computing for laboratories with limited connectivity, and synthetic data for rare diseases. Digital twins may support workflow simulation and resource planning, while self-correcting algorithms require careful governance because automatic adaptation can introduce instability. Research is also needed on AI governance in resource-constrained countries, including data sovereignty, regulatory capacity, financing, and local accountability.
Finally, robust economic evaluations should measure total lifecycle costs, including infrastructure, training, maintenance, cybersecurity, revalidation, and opportunity costs. These priorities would shift laboratory AI research from demonstrating technical possibility towards establishing sustainable clinical and public-health value globally.

Conclusion

Artificial intelligence has substantial potential to transform laboratory diagnosis, disease surveillance, workflow efficiency, and global health decision-making. Evidence reviewed across diagnostic imaging, molecular testing, laboratory information systems, and public-health surveillance shows important gains in speed, pattern recognition, and decision support. However, technical performance alone does not guarantee clinical validity, generalisability, affordability, or equitable benefit. Effective implementation requires rigorous external and prospective validation, reliable infrastructure, interoperable data systems, cybersecurity, transparent governance, and continuous performance monitoring. Qualified laboratory professionals must retain meaningful oversight because artificial intelligence should strengthen, rather than replace, scientific judgement. Particular attention is required in low- and middle-income countries, where infrastructure limitations, workforce shortages, and underrepresented populations may restrict benefits or deepen inequalities. Therefore, artificial intelligence-enabled laboratories should be developed as secure, equitable, and clinically governed systems that combine computational capability with professional expertise, local validation, sustained investment in public-health capacity, and responsible international collaboration for future collective resilience.

References

  1. Africa Centres for Disease Control and Prevention. Africa CDC launches AGARI, a continent-wide genomic data platform to strengthen outbreak response. 21 November 2025. Available online: https://africacdc.org/news-item/africa-cdc-launches-agari-a-continent-wide-genomic-data-platform-to-strengthen-outbreak-response/.
  2. Andaur Navarro, C. L.; Damen, J. A. A.; Takada, T.; Nijman, S. W. J.; Dhiman, P.; Ma, J.; Collins, G. S.; Bajpai, R.; Riley, R. D.; Moons, K. G. M.; Hooft, L. Risk of bias in studies on prediction models developed using supervised machine learning techniques: Systematic review. BMJ 375 2021, n2281. [Google Scholar] [CrossRef]
  3. Antonelli, G.; Filippi, J.; D’Orazio, M.; Curci, G.; Casti, P.; Mencattini, A.; Martinelli, E. Integrating machine learning and biosensors in microfluidic devices: A review. Biosens. Bioelectron. 263 2024, 116632. [Google Scholar] [CrossRef]
  4. Ardila, C. M.; Yadalam, P. K.; González-Arroyave, D. Integrating whole genome sequencing and machine learning for predicting antimicrobial resistance in critical pathogens. PeerJ 12 2024, e18213. [Google Scholar] [CrossRef]
  5. Arvai, N.; Katonai, G.; Mesko, B. Health care professionals’ concerns about medical AI and psychological barriers and strategies for successful implementation: Scoping review. J. Med. Internet Res. 27 2025, e66986. [Google Scholar] [CrossRef]
  6. Ayoub, C.; Appari, L.; Pereyra, M.; Farina, J. M.; Chao, C.-J.; Scalia, I. G.; Mahmoud, A. K.; Abbas, M. T.; Baba, N. A.; Jeong, J.; Lester, S. J.; Patel, B. N.; Arsanjani, R.; Banerjee, I. Multimodal fusion artificial intelligence model to predict risk for MACE and myocarditis in cancer patients receiving immune checkpoint inhibitor therapy. JACC Adv. 2025, 4(1), 101435. [Google Scholar] [CrossRef]
  7. Baethge, C.; Goldbeck-Wood, S.; Mertens, S. SANRA—A scale for the quality assessment of narrative review articles. Res. Integr. Peer Rev. 2019, 4, 5. [Google Scholar] [CrossRef] [PubMed]
  8. Baron, J. M. Artificial intelligence in the clinical laboratory: An overview with frequently asked questions. Clin. Lab. Med. 2023, 43(1), 1–16. [Google Scholar] [CrossRef]
  9. Bernhardt, M.; Jones, C.; Glocker, B. Potential sources of dataset bias complicate investigation of underdiagnosis by machine learning algorithms. Nat. Med. 28 2022, 1157–1158. [Google Scholar] [CrossRef]
  10. Bramer, W. M.; de Jonge, G. B.; Rethlefsen, M. L.; Mast, F.; Kleijnen, J. A systematic approach to searching: An efficient and complete method to develop literature searches. J. Med. Libr. Assoc. 2018, 106(4), 531–541. [Google Scholar] [CrossRef]
  11. Bramer, W. M.; Rethlefsen, M. L.; Kleijnen, J.; Franco, O. H. Optimal database combinations for literature searches in systematic reviews: A prospective exploratory study. Syst. Rev. 6 2017, 245. [Google Scholar] [CrossRef]
  12. Brehmer, A.; Sauer, C. M.; Rodríguez, J. S.; Herrmann, K.; Kim, M.; Keyl, J.; Bahnsen, F. H.; Frank, B.; Köhrmann, M.; Rassaf, T.; Mahabadi, A.-A.; Hadaschik, B.; Darr, C.; Herrmann, K.; Tan, S.; Buer, J.; Brenner, T.; Reinhardt, H. C.; Nensa, F.; …; Kleesiek, J. Establishing Medical Intelligence—Leveraging Fast Healthcare Interoperability Resources to improve clinical management: A retrospective cohort and clinical implementation study. J. Med. Internet Res. 26 2024, e55148. [Google Scholar] [CrossRef]
  13. Brunker, K. Rapid pathogen surveillance: Field-ready sequencing solutions. Nat. Rev. Genet. 25 2024, 532. [Google Scholar] [CrossRef]
  14. Bulten, W.; Balkenhol, M.; Awoumou Belinga, J.-J.; Brilhante, A.; Çakır, A.; Farré, X.; Geronatsiou, K.; Molinié, V.; Pereira, G.; Roy, P.; Saile, G.; Salles, P.; Schaafsma, E.; Tschui, J.; Vos, A.-M.; van Boven, H.; Vink, R.; van der Laak, J.; Hulsbergen-van de Kaa, C.; Litjens, G. Artificial intelligence assistance significantly improves Gleason grading of prostate biopsies by pathologists. Mod. Pathol. 34 2021, 660–671. [Google Scholar] [CrossRef]
  15. Cadamuro, J.; Carobene, A.; Cabitza, F.; Debeljak, Z.; De Bruyne, S.; van Doorn, W.; Johannes, E.; Frans, G.; Özdemir, H.; Martin Perez, S.; Rajdl, D.; Tolios, A.; Padoan, A.; European Federation of Clinical Chemistry and Laboratory Medicine Working Group on Artificial Intelligence. A comprehensive survey of artificial intelligence adoption in European laboratory medicine: Current utilization and prospects. Clin. Chem. Lab. Med. 2025, 63(4), 692–703. [Google Scholar] [CrossRef]
  16. Campbell, M.; McKenzie, J. E.; Sowden, A.; Katikireddi, S. V.; Brennan, S. E.; Ellis, S.; Hartmann-Boyce, J.; Ryan, R.; Shepperd, S.; Thomas, J.; Welch, V.; Thomson, H. Synthesis without meta-analysis (SWiM) in systematic reviews: Reporting guideline. BMJ 368 2020, l6890. [Google Scholar] [CrossRef]
  17. Chandrashekar, M.; Lyngaas, I.; Hanson, H. A.; Gao, S.; Wu, X.-C.; Gounley, J. Path-BigBird: An AI-driven transformer approach to classification of cancer pathology reports. JCO Clin. Cancer Inform. 8 2024, e2300148. [Google Scholar] [CrossRef]
  18. Chen, F.; Wang, L.; Hong, J.; Jiang, J.; Zhou, L. Unmasking bias in artificial intelligence. J. Am. Med. Inform. Assoc. 2024, 31(5), 1172–1183. [Google Scholar] [CrossRef]
  19. Chen, R. J.; Ding, T.; Lu, M. Y.; Williamson, D. F. K.; Jaume, G.; Song, A. H.; Chen, B.; Zhang, A.; Shao, D.; Shaban, M.; Williams, M.; Oldenburg, L.; Weishaupt, L. L.; Wang, J. J.; Vaidya, A.; Le, L. P.; Gerber, G.; Sahai, S.; Williams, W.; Mahmood, F. Towards a general-purpose foundation model for computational pathology. Nat. Med. 30 2024, 850–862. [Google Scholar] [CrossRef]
  20. Chen, R. J.; Wang, J. J.; Williamson, D. F. K.; Chen, T. Y.; Lipkova, J.; Lu, M. Y.; Sahai, S.; Mahmood, F. Algorithmic fairness in artificial intelligence for medicine and healthcare. Nat. Biomed. Eng. 2023, 7(6), 719–742. [Google Scholar] [CrossRef]
  21. Cheng, J.; Novati, G.; Pan, J.; Bycroft, C.; Žemgulytė, A.; Applebaum, T.; Pritzel, A.; Wong, L. H.; Zielinski, M.; Sargeant, T.; Schneider, R. G.; Senior, A. W.; Jumper, J.; Hassabis, D.; Kohli, P.; Avsec, Ž. Accurate proteome-wide missense variant effect prediction with AlphaMissense. Science 2023, 381(6664), eadg7492. [Google Scholar] [CrossRef]
  22. Cheplygina, V.; de Bruijne, M.; Pluim, J. P. W. Not-so-supervised: A survey of semi-supervised, multi-instance, and transfer learning in medical image analysis. Med. Image Anal. 54 2019, 280–296. [Google Scholar] [CrossRef]
  23. Collins, G. S.; Moons, K. G. M.; Dhiman, P.; Riley, R. D.; Beam, A. L.; Van Calster, B.; Ghassemi, M.; Liu, X.; Reitsma, J. B.; van Smeden, M.; Boulesteix, A.-L.; Camaradou, J. C.; Celi, L. A.; Denaxas, S.; Denniston, A. K.; Glocker, B.; Golub, R. M.; Harvey, H.; Heinze, G.; …; Logullo, P. TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 385 2024, e078378. [Google Scholar] [CrossRef]
  24. Das, D.; Vongpromek, R.; Assawariyathipat, T.; Srinamon, K.; Kennon, K.; Stepniewska, K.; Ghose, A.; Abu Sayeed, A.; Faiz, M. A.; Abreu Netto, R. L.; Siqueira, A.; Yerbanga, S. R.; Ouédraogo, J. B.; Callery, J. J.; Peto, T. J.; Tripura, R.; Koukouikila-Koussounda, F.; Ntoumi, F.; Ong’echa, J. M.; …; Dhorda, M. Field evaluation of the diagnostic performance of EasyScan GO: A digital malaria microscopy device based on machine-learning. Malar. J. 21 2022, 122. [Google Scholar] [CrossRef]
  25. Dasdelen, M. F.; Kukuljan, I.; Lienemann, P.; Ozlugedik, F.; Sadafi, A.; Hehr, M.; Spiekermann, K.; Pohlkamp, C.; Marr, C. AI-based hematological malignancy prediction from peripheral blood smears in a large diagnostic laboratory cohort. Leuk. 40 2026, 1318–1322. [Google Scholar] [CrossRef]
  26. Edayan, J. M.; Gallemit, A. J.; Sacala, N. E.; Palmer, X.-L.; Potter, L.; Rarugal, J. P.; Velasco, L. C. Integration technologies in laboratory information systems: A systematic review. Inform. Med. Unlocked 50 2024, 101566. [Google Scholar] [CrossRef]
  27. El Arab, R. A.; Al Moosa, O. A. Systematic review of cost effectiveness and budget impact of artificial intelligence in healthcare. npj Digit. Med. 8 2025, 548. [Google Scholar] [CrossRef]
  28. El Arab, R. A.; Alkhunaizi, M.; Alhashem, Y. N.; Al Khatib, A.; Bubsheet, M.; Hassanein, S. Artificial intelligence in vaccine research and development: An umbrella review. Front. Immunol. 16 2025, 1567116. [Google Scholar] [CrossRef]
  29. Erlich, Y.; Shor, T.; Pe'er, I.; Carmi, S. Identity inference of genomic data using long-range familial searches. Science 2018, 362(6415), 690–694. [Google Scholar] [CrossRef]
  30. Esteva, A.; Robicquet, A.; Ramsundar, B.; Kuleshov, V.; DePristo, M.; Chou, K.; Cui, C.; Corrado, G.; Thrun, S.; Dean, J. A guide to deep learning in healthcare. Nat. Med. 25 2019, 24–29. [Google Scholar] [CrossRef]
  31. European Commission. Artificial intelligence in healthcare. 2024. Available online: https://health.ec.europa.eu/ehealth-digital-health-and-care/artificial-intelligence-healthcare_en.
  32. European Commission. In vitro diagnostic medical devices: Overview. 2026. Available online: https://health.ec.europa.eu/medical-devices-vitro-diagnostics/overview_en.
  33. European Union Agency for Cybersecurity. ENISA threat landscape: Health sector. 2023. Available online: https://www.enisa.europa.eu/publications/health-threat-landscape.
  34. European Union Agency for Cybersecurity. ENISA threat landscape 2024. 2024. Available online: https://www.enisa.europa.eu/publications/enisa-threat-landscape-2024.
  35. Ferrari, R. Writing narrative style literature reviews. Med. Writ. 2015, 24(4), 230–235. [Google Scholar] [CrossRef]
  36. Finlayson, S. G.; Bowers, J. D.; Ito, J.; Zittrain, J. L.; Beam, A. L.; Kohane, I. S. Adversarial attacks on medical machine learning. Science 2019, 363(6433), 1287–1289. [Google Scholar] [CrossRef]
  37. Fleming, K. A.; Horton, S.; Wilson, M. L.; Atun, R.; DeStigter, K.; Flanigan, J.; Sayed, S.; Adam, P.; Aguilar, B.; Andronikou, S.; Boehme, C.; Cherniak, W.; Cheung, A. N. Y.; Dahn, B.; Donoso-Bach, L.; Douglas, T.; Garcia, P.; Hussain, S.; Iyer, H. S.; …; Walia, K. The Lancet Commission on diagnostics: Transforming access to diagnostics. The Lancet 2021, 398(10315), 1997–2050. [Google Scholar] [CrossRef]
  38. U.S. Food and Drug Administration; Health Canada; Medicines and Healthcare products Regulatory Agency. Transparency for machine learning-enabled medical devices: Guiding principles. 2024. Available online: https://www.fda.gov/medical-devices/software-medical-device-samd/transparency-machine-learning-enabled-medical-devices-guiding-principles.
  39. Food and Drug Administration. Marketing submission recommendations for a predetermined change control plan for artificial intelligence-enabled device software functions. 2025. Available online: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/marketing-submission-recommendations-predetermined-change-control-plan-artificial-intelligence.
  40. Foroughi, M.; Arzehgar, A.; Seyedhasani, S. N.; Nadali, A.; Zoroufchi Benis, K. Application of machine learning for antibiotic resistance in water and wastewater: A systematic review. Chemosphere 358 2024, 142223. [Google Scholar] [CrossRef]
  41. Ghassemi, M.; Oakden-Rayner, L.; Beam, A. L. The false hope of current approaches to explainable artificial intelligence in health care. Lancet Digit. Health 2021, 3(11), e745–e750. [Google Scholar] [CrossRef]
  42. Goodfellow, I.; Bengio, Y.; Courville, A. Deep learning; MIT Press, 2016; Available online: https://mitpress.mit.edu/9780262035613/deep-learning/.
  43. Gusenbauer, M.; Haddaway, N. R. Which academic search systems are suitable for systematic reviews or meta-analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resources. Res. Synth. Methods 2020, 11(2), 181–217. [Google Scholar] [CrossRef] [PubMed]
  44. Herman, D. S.; Rhoads, D. D.; Schulz, W. L.; Durant, T. J. S. Artificial intelligence and mapping a new direction in laboratory medicine: A review. Clin. Chem. 2021, 67(11), 1466–1482. [Google Scholar] [CrossRef]
  45. Higgins, J. P. T.; Thomas, J.; Chandler, J.; Cumpston, M.; Li, T.; Page, M. J.; Welch, V. A. (Eds.) Cochrane handbook for systematic reviews of interventions (Version 6.5); Cochrane., 2024; Available online: https://www.cochrane.org/authors/handbooks-and-manuals/handbook.
  46. Hoang, D.-T.; Shulman, E. D.; Turakulov, R.; Abdullaev, Z.; Singh, O.; Campagnolo, E. M.; Lalchungnunga, H.; Stone, E. A.; Nasrallah, M. P.; Ruppin, E.; Aldape, K. Prediction of DNA methylation-based tumor types from histopathology in central nervous system tumors with deep learning. Nat. Med. 30 2024, 1952–1961. [Google Scholar] [CrossRef]
  47. Horning, M. P.; Delahunt, C. B.; Bachman, C. M.; Luchavez, J.; Luna, C.; Hu, L.; Jaiswal, M. S.; Thompson, C. M.; Kulhare, S.; Janko, S.; Wilson, B. K.; Ostbye, T.; Mehanian, M.; Gebrehiwot, R.; Yun, G.; Bell, D.; Proux, S.; Carter, J. Y.; Oyibo, W.; …; Mehanian, C. Performance of a fully-automated system on a WHO malaria microscopy evaluation slide set. Malar. J. 20 2021, 110. [Google Scholar] [CrossRef]
  48. Hou, H.; Zhang, R.; Li, J. Artificial intelligence in the clinical laboratory. Clin. Chim. Acta 559 2024, 119724. [Google Scholar] [CrossRef]
  49. Huang, B.; Kang, L.; Tsang, V. T. C.; Lo, C. T. K.; Wong, T. T. W. Deep learning-assisted smartphone-based quantitative microscopy for label-free peripheral blood smear analysis. Biomed. Opt. Express 2024. [Google Scholar] [CrossRef]
  50. Humphries, M. P.; Kaye, D.; Stankeviciute, G.; Halliwell, J.; Wright, A. I.; Bansal, D.; Brettle, D.; Treanor, D. Development of a multi-scanner facility for data acquisition for digital pathology artificial intelligence. J. Pathol. 2024, 264(1), 80–89. [Google Scholar] [CrossRef]
  51. International Medical Device Regulators Forum. Good machine learning practice for medical device development: Guiding principles (IMDRF/AIML WG/N88 FINAL:2025). 2025. Available online: https://www.imdrf.org/documents/good-machine-learning-practice-medical-device-development-guiding-principles.
  52. International Organization for Standardization. ISO 15189:2022 Medical laboratories—Requirements for quality and competence. 2022. Available online: https://www.iso.org/standard/76677.html.
  53. Iscoe, M.; Li, H.; Xue, H.; Socrates, V.; Gilson, A.; Huang, T.; Taylor, R. A. Evaluating the potential impact of AI on urinary tract infection diagnosis in the emergency department across demographic groups: Retrospective cohort study. JMIR AI 5 2026, e91148. [Google Scholar] [CrossRef]
  54. Ji, C. X.; Alaa, A. M.; Sontag, D. Large-scale study of temporal shift in health insurance claims. Proc. Mach. Learn. Res. 209 2023, 243–278. Available online: https://proceedings.mlr.press/v209/ji23a.html.
  55. Jung, M.; Jin, M.-S.; Kim, C.; Lee, C.; Nikas, I. P.; Park, J. H.; Ryu, H. S. Artificial intelligence system shows performance at the level of uropathologists for the detection and grading of prostate cancer in core needle biopsy: An independent external validation study. Mod. Pathol. 35 2022, 1449–1457. [Google Scholar] [CrossRef]
  56. Kherabi, Y.; Thy, M.; Bouzid, D.; Antcliffe, D. B.; Rawson, T. M.; Peiffer-Smadja, N. Machine learning to predict antimicrobial resistance: Future applications in clinical practice? Infect. Dis. Now. 2024, 54(3), 104864. [Google Scholar] [CrossRef]
  57. Kim, S.; Min, W.\.-K. Toward high-quality real-world laboratory data in the era of healthcare big data. Ann. Lab. Med. 2025, 45(1), 1–11. [Google Scholar] [CrossRef]
  58. Kong, S. W.; Lee, I.-H.; Collen, L. V.; Field, M.; Manrai, A. K.; Snapper, S. B.; Mandl, K. D. Discordance between a deep learning model and clinical-grade variant pathogenicity classification in a rare disease cohort. npj Genom. Med. 10 2025, 17. [Google Scholar] [CrossRef]
  59. Krasowski, M. D.; Davis, S. R.; Drees, D.; Morris, C.; Kulhavy, J.; Crone, C.; Bebber, T.; Clark, I.; Nelson, D. L.; Teul, S.; Voss, D.; Aman, D.; Fahnle, J.; Blau, J. L. Autoverification in a core clinical chemistry laboratory at an academic medical center. J. Pathol. Inform. 2014, 5(1), 13. [Google Scholar] [CrossRef]
  60. Leng, H.; Deng, R.; Bao, S.; Fang, D.; Millis, B. A.; Tang, Y.; Yang, H.; Wang, X.; Peng, Y.; Wan, L.; Huo, Y. High-performance data management for whole slide image analysis in digital pathology. Proc. SPIE Med. Imaging 2024 Digit. Comput. Pathol. 12933 2024, 129330Y. [Google Scholar] [CrossRef]
  61. Li, T.; Jia, L.; Qiang, N.; Zheng, J.; Ran, J.; Zhang, X.; Han, L. Progress and challenges in development of minimum essential dataset for disease surveillance through a One Health lens: A scoping review. Infect. Dis. Poverty 15 2026, 45. [Google Scholar] [CrossRef]
  62. Liu, G.; Catacutan, D. B.; Rathod, K.; Swanson, K.; Jin, W.; Mohammed, J. C.; Chiappino-Pepe, A.; Syed, S. A.; Fragis, M.; Rachwalski, K.; Magolan, J.; Surette, M. G.; Coombes, B. K.; Jaakkola, T.; Barzilay, R.; Collins, J. J.; Stokes, J. M. Deep learning-guided discovery of an antibiotic targeting Acinetobacter baumannii. Nat. Chem. Biol. 2023, 19(11), 1342–1350. [Google Scholar] [CrossRef]
  63. Liu, X.; Cruz Rivera, S.; Moher, D.; Calvert, M. J.; Denniston, A. K.; SPIRIT-AI and CONSORT-AI Working Group. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: The CONSORT-AI extension. Nat. Med. 26 2020, 1364–1374. [Google Scholar] [CrossRef]
  64. Liu, X.; Faes, L.; Kale, A. U.; Wagner, S. K.; Fu, D. J.; Bruynseels, A.; Mahendiran, T.; Moraes, G.; Shamdas, M.; Kern, C.; Ledsam, J. R.; Schmid, M. K.; Balaskas, K.; Topol, E. J.; Bachmann, L. M.; Keane, P. A.; Denniston, A. K. A comparison of deep learning performance against health-care professionals in detecting diseases from medical imaging: A systematic review and meta-analysis. Lancet Digit. Health 2019, 1(6), e271–e297. [Google Scholar] [CrossRef]
  65. Mahomed, S.; Ncube, M. V.; Dhai, A.; Janneker, W.; Nodikida, M. Ethical governance of artificial intelligence in healthcare: Critical reflections of the AI Task Team of the South African Medical Association. South Afr. Med. J. 2026, 116(5), e5072. [Google Scholar] [CrossRef]
  66. Matek, C.; Krappe, S.; Münzenmayer, C.; Haferlach, T.; Marr, C. Highly accurate differentiation of bone marrow cell morphologies using deep neural networks on a large image data set. Blood 2021, 138(20), 1917–1927. [Google Scholar] [CrossRef]
  67. McGenity, C.; Clarke, E. L.; Jennings, C.; Matthews, G.; Cartlidge, C.; Freduah-Agyemang, H.; Stocken, D. D.; Treanor, D. Artificial intelligence in digital pathology: A systematic review and meta-analysis of diagnostic test accuracy. npj Digit. Med. 7 2024, 114. [Google Scholar] [CrossRef]
  68. McGowan, J.; Sampson, M.; Salzwedel, D. M.; Cogo, E.; Foerster, V.; Lefebvre, C. PRESS peer review of electronic search strategies: 2015 guideline statement. J. Clin. Epidemiol. 75 2016, 40–46. [Google Scholar] [CrossRef]
  69. Medicines and Healthcare products Regulatory Agency. Medical devices: Post-market surveillance requirements. 2025. Available online: https://www.gov.uk/government/publications/medical-devices-post-market-surveillance-requirements.
  70. Mencacci, A.; De Socio, G. V.; Pirelli, E.; Bondi, P.; Cenci, E. Laboratory automation, informatics, and artificial intelligence: Current and future perspectives in clinical microbiology. Front. Cell. Infect. Microbiol. 13 2023, 1188684. [Google Scholar] [CrossRef]
  71. Miller, H. A.; Valdes, R. Rigorous validation of machine learning in laboratory medicine: Guidance toward quality improvement. Crit. Rev. Clin. Lab. Sci. 2025, 62(5), 327–346. [Google Scholar] [CrossRef]
  72. Moons, K. G. M.; Damen, J. A. A.; Kaul, T.; Hooft, L.; Andaur Navarro, C.; Dhiman, P.; Beam, A. L.; Van Calster, B.; Celi, L. A.; Denaxas, S.; Denniston, A. K.; Ghassemi, M.; Heinze, G.; Kengne, A. P.; Liu, X.; Logullo, P.; Maier-Hein, L.; McCradden, M. D.; Oaken-Rayner, L.; …; van Smeden, M. PROBAST+AI: An updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ 388 2025, e082505. [Google Scholar] [CrossRef]
  73. Munung, N. S.; Royal, C. D.; de Kock, C.; Awandare, G.; Nembaware, V.; Nguefack, S.; Treadwell, M.; Wonkam, A. Genomics and health data governance in Africa: Democratize the use of big data and popularize public engagement. In Hastings Center Report; 2024. [Google Scholar] [CrossRef]
  74. Muralidharan, V.; Adewale, B. A.; Huang, C. J.; Nta, M. T.; Ademiju, P. O.; Pathmarajah, P.; Hang, M. K.; Adesanya, O.; Abdullateef, R. O.; Babatunde, A. O.; Ajibade, A.; Onyeka, S.; Cai, Z. R.; Daneshjou, R.; Olatunji, T. A scoping review of reporting gaps in FDA-approved AI medical devices. npj Digit. Med. 7 2024, 273. [Google Scholar] [CrossRef]
  75. Nagendran, M.; Chen, Y.; Lovejoy, C. A.; Gordon, A. C.; Komorowski, M.; Harvey, H.; Topol, E. J.; Ioannidis, J. P. A.; Collins, G. S.; Maruthappu, M. Artificial intelligence versus clinicians: Systematic review of design, reporting standards, and claims of deep learning studies. BMJ 368 2020, m689. [Google Scholar] [CrossRef]
  76. National Institutes of Health. Implementation update for data management and access practices under the Genomic Data Sharing Policy (NOT-OD-24-157). 2024. Available online: https://grants.nih.gov/grants/guide/notice-files/NOT-OD-24-157.html.
  77. NHS England. Update on cyber incident: Clinical impact in south east London. 26 September 2024. Available online: https://www.england.nhs.uk/london/2024/09/26/update-on-cyber-incident-clinical-impact-in-south-east-london-thursday-26-september-2024/.
  78. Niţulescu, A.; Stoicu-Tivadar, L. Data standardization in the medical field through FHIR and FAIR implementation: A systematic review. Stud. Health Technol. Inform. 316 2024, 1378–1382. [Google Scholar] [CrossRef]
  79. Norgan, A. P.; Simon, K. E.; Feehan, B. A.; Saari, L. L.; Doppler, J. M.; Welder, G. S.; Sedarski, J. A.; Yoch, C. T.; Comfere, N. I.; Martin, J. A.; Bartholmai, B. J.; Reichard, R. R. Radio-frequency identification specimen tracking to improve quality in anatomic pathology. Arch. Pathol. Lab. Med. 2020, 144(2), 189–195. [Google Scholar] [CrossRef]
  80. Oduoye, M. O.; Fatima, E.; Muzammil, M. A.; Dave, T.; Irfan, H.; Fariha, F. N. U.; Marbell, A.; Ubechu, S. C.; Scott, G. Y.; Elebesunu, E. E. Impacts of the advancement in artificial intelligence on laboratory medicine in low- and middle-income countries. Health Sci. Rep. 2024, 7(1), e1794. [Google Scholar] [CrossRef]
  81. OECD. Explanatory memorandum on the updated OECD definition of an AI system; (OECD Artificial Intelligence Papers No. 8); OECD Publishing, 2024. [Google Scholar] [CrossRef]
  82. Office of the Privacy Commissioner of Canada; Information Commissioner’s Office. Joint investigation into a data breach at 23andMe by the Privacy Commissioner of Canada and the UK Information Commissioner. (PIPEDA Findings #2025-001). 2025. Available online: https://www.priv.gc.ca/en/opc-actions-and-decisions/investigations/investigations-into-businesses/2025/pipeda-2025-001/.
  83. Onywera, H.; Mulder, N.; Kebede, Y.; Tessema, S. K. How to sustain a public-health genomics and bioinformatics workforce in Africa. Nat. Med. 2025, 31(8), 2480–2484. [Google Scholar] [CrossRef]
  84. Ouzzani, M.; Hammady, H.; Fedorowicz, Z.; Elmagarmid, A. Rayyan—A web and mobile app for systematic reviews. Syst. Rev. 5 2016, 210. [Google Scholar] [CrossRef]
  85. Page, M. J.; McKenzie, J. E.; Bossuyt, P. M.; Boutron, I.; Hoffmann, T. C.; Mulrow, C. D.; Shamseer, L.; Tetzlaff, J. M.; Akl, E. A.; Brennan, S. E.; Chou, R.; Glanville, J.; Grimshaw, J. M.; Hróbjartsson, A.; Lalu, M. M.; Li, T.; Loder, E. W.; Mayo-Wilson, E.; McDonald, S.; …; Moher, D. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 372 2021, n71. [Google Scholar] [CrossRef]
  86. Paranjape, K.; Schinkel, M.; Hammer, R. D.; Schouten, B.; Nannan Panday, R. S.; Elbers, P. W. G.; Kramer, M. H. H.; Nanayakkara, P. The value of artificial intelligence in laboratory medicine. Am. J. Clin. Pathol. 2021, 155(6), 823–831. [Google Scholar] [CrossRef]
  87. Pearce, G. Data interoperability: Addressing the challenges placing quality healthcare at risk. ISACA J. 2024, 2024(2). Available online: https://www.isaca.org/resources/isaca-journal/issues/2024/volume-2/data-interoperability.
  88. Rajaonison, A.; Le Page, S.; Maurin, T.; Chaudet, H.; Raoult, D.; Baron, S. A.; Rolain, J.-M. Antilogic, a new supervised machine learning software for the automatic interpretation of antibiotic susceptibility testing in clinical microbiology: Proof-of-concept on three frequently isolated bacterial species. Clin. Microbiol. Infect. 2022, 28(9), 1286.e1–1286.e8. [Google Scholar] [CrossRef]
  89. Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence. Official Journal of the European Union. 2024. Available online: https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX%3A32024R1689.
  90. Rethlefsen, M. L.; Kirtley, S.; Waffenschmidt, S.; Ayala, A. P.; Moher, D.; Page, M. J.; Koffel, J. B.; PRISMA-S Group. PRISMA-S: An extension to the PRISMA statement for reporting literature searches in systematic reviews. J. Med. Libr. Assoc. 2021, 109(2), 174–200. [Google Scholar] [CrossRef]
  91. Roberts, M.; Driggs, D.; Thorpe, M.; Gilbey, J.; Yeung, M.; Ursprung, S.; Aviles-Rivero, A. I.; Etmann, C.; McCague, C.; Beer, L.; Weir-McCall, J. R.; Teng, Z.; Gkrania-Klotsas, E.; Rudd, J. H. F.; Sala, E.; Schönlieb, C.-B. Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans. Nat. Mach. Intell. 2021, 3(3), 199–217. [Google Scholar] [CrossRef]
  92. Rosenbacke, R.; Melhus, Å.; McKee, M.; Stuckler, D. How explainable artificial intelligence can increase or decrease clinicians' trust in AI applications in health care: Systematic review. JMIR AI 3 2024, e53207. [Google Scholar] [CrossRef]
  93. Russell, R. G.; Lovett Novak, L.; Patel, M.; Garvey, K. V.; Craig, K. J. T.; Jackson, G. P.; Moore, D.; Miller, B. M. Competencies for the use of artificial intelligence-based tools by health care professionals. Acad. Med. 2023, 98(3), 348–356. [Google Scholar] [CrossRef]
  94. Sakai, H.; Lam, S. S. Large language models for health care text classification: Systematic review. JMIR AI 5 2026, e79202. [Google Scholar] [CrossRef]
  95. Salmi, M.; Atif, D.; Oliva, D.; Abraham, A.; Ventura, S. Handling imbalanced medical datasets: Review of a decade of research. Artif. Intell. Rev. 57 2024, 273. [Google Scholar] [CrossRef]
  96. Schwabe, D.; Becker, K.; Seyferth, M.; Klaß, A.; Schaeffter, T. The METRIC-framework for assessing data quality for trustworthy AI in medicine: A systematic review. npj Digit. Med. 7 2024, 203. [Google Scholar] [CrossRef]
  97. Seok, H. S.; Yu, S.; Shin, K.-H.; Lee, W.; Chun, S.; Kim, S.; Shin, H. Machine learning-based sample misidentification error detection in clinical laboratory tests: A retrospective multicenter study. Clin. Chem. 2024, 70(10), 1256–1267. [Google Scholar] [CrossRef]
  98. Seyyed-Kalantari, L.; Zhang, H.; McDermott, M. B. A.; Chen, I. Y.; Ghassemi, M. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat. Med. 27 2021, 2176–2182. [Google Scholar] [CrossRef]
  99. Shick, A. A.; Webber, C. M.; Kiarashi, N.; Weinberg, J. P.; Deoras, A.; Petrick, N.; Saha, A.; Diamond, M. C. Transparency of artificial intelligence/machine learning-enabled medical devices. npj Digit. Med. 7 2024, 21. [Google Scholar] [CrossRef]
  100. Signoroni, A.; Ferrari, A.; Lombardi, S.; Savardi, M.; Fontana, S.; Culbreath, K. Hierarchical AI enables global interpretation of culture plates in the era of digital microbiology. Nat. Commun. 14 2023, 6874. [Google Scholar] [CrossRef]
  101. Singh, S.; Sharma, P.; Pal, N.; Sarma, D. K.; Tiwari, R.; Kumar, M. Holistic One Health surveillance framework. ACS Infect. Dis. 2024, 10(3), 808–826. [Google Scholar] [CrossRef]
  102. Snyder, H. Literature review as a research methodology: An overview and guidelines. J. Bus. Res. 104 2019, 333–339. [Google Scholar] [CrossRef]
  103. Sounderajah, V.; Guni, A.; Liu, X.; Collins, G. S.; Karthikesalingam, A.; Markar, S. R.; Golub, R. M.; Denniston, A. K.; Shetty, S.; Moher, D.; Bossuyt, P. M.; Darzi, A.; Ashrafian, H.; STARD-AI Steering Committee. The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence. Nat. Med. 2025, 31(10), 3283–3289. [Google Scholar] [CrossRef] [PubMed]
  104. Spies, N. C.; Farnsworth, C. W.; Wheeler, S.; McCudden, C. R. Validating, implementing, and monitoring machine learning solutions in the clinical laboratory safely and effectively. Clin. Chem. 2024, 70(11), 1334–1343. [Google Scholar] [CrossRef]
  105. Sutton, R. S.; Barto, A. G. Reinforcement learning: An introduction, 2nd ed.; MIT Press, 2018; Available online: https://mitpress.mit.edu/9780262039246/reinforcement-learning/.
  106. Syed, R.; Eden, R.; Makasi, T.; Chukwudi, I.; Mamudu, A.; Kamalpour, M.; Kapugama Geeganage, D.; Sadeghianasl, S.; Leemans, S. J. J.; Goel, K.; Andrews, R.; Wynn, M. T.; ter Hofstede, A.; Myers, T. Digital health data quality issues: Systematic review. J. Med. Internet Res. 25 2023, e42615. [Google Scholar] [CrossRef]
  107. Tavares da Souza, A.; Flores, J.; Millendez, L.; Filio, M.; Mo, Y. D.; Jacquot, C.; Delaney, M. Radiofrequency identification tracking system (RFID) significantly improves blood bank inventory management and decreases staff work effort. Transfusion 2024, 64(4), 578–584. [Google Scholar] [CrossRef]
  108. Tejani, A. S.; Klontzas, M. E.; Gatti, A. A.; Mongan, J. T.; Moy, L.; Park, S. H.; Kahn, C. E. CLAIM 2024 Update Panel Checklist for Artificial Intelligence in Medical Imaging (CLAIM): 2024 update. Radiol. Artif. Intell. 2024, 6(4), e240300. [Google Scholar] [CrossRef] [PubMed]
  109. Tun, H. M.; Rahman, H. A.; Naing, L.; Malik, O. A. Trust in artificial intelligence-based clinical decision support systems among health care workers: Systematic review. J. Med. Internet Res. 27 2025, e69678. [Google Scholar] [CrossRef]
  110. U.S. Food and Drug Administration. Good machine learning practice for medical device development: Guiding principles. 2025. Available online: https://www.fda.gov/medical-devices/software-medical-device-samd/good-machine-learning-practice-medical-device-development-guiding-principles.
  111. U.S. Food and Drug Administration. Cybersecurity in medical devices: Quality management system considerations and content of premarket submissions. 2026. Available online: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/cybersecurity-medical-devices-quality-management-system-considerations-and-content-premarket.
  112. United Nations. The 17 goals: Sustainable Development Goals. n.d. Available online: https://sdgs.un.org/goals.
  113. Vasey, B.; Nagendran, M.; Campbell, B.; Clifton, D. A.; Collins, G. S.; Denaxas, S.; Denniston, A. K.; Faes, L.; Geerts, B.; Ibrahim, M.; Liu, X.; Mateen, B. A.; Mathur, P.; McCradden, M. D.; Morgan, L.; Ordish, J.; Rogers, C.; Saria, S.; Ting, D. S. W.; …; McCulloch, P. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat. Med. 2022, 28(5), 924–933. [Google Scholar] [CrossRef] [PubMed]
  114. Villanueva-Miranda, I.; Xiao, G.; Xie, Y. Artificial intelligence in early warning systems for infectious disease surveillance: A systematic review. Front. Public Health 13 2025, 1609615. [Google Scholar] [CrossRef]
  115. Wadie, P.; Zakher, B.; Elgazzar, K.; Alsbakhi, A.; Alhejaily, A.-M. G. AI in point-of-care imaging for clinical decision support: Systematic review of diagnostic accuracy, task-shifting, and explainability. JMIR AI 5 2026, e80928. [Google Scholar] [CrossRef]
  116. Wang, G.; Luo, G.; Lian, H.; Chen, L.; Wu, W.; Liu, H. Application of deep learning in clinical settings for detecting and classifying malaria parasites in thin blood smears. Open Forum Infect. Dis. 2023, 10(11), ofad469. [Google Scholar] [CrossRef]
  117. Wang, H.; Wang, H.; Zhang, J.; Li, X.; Sun, C.; Zhang, Y. Using machine learning to develop an autoverification system in a clinical biochemistry laboratory. Clin. Chem. Lab. Med. 2021, 59(5), 883–891. [Google Scholar] [CrossRef]
  118. Wang, X.; Zhao, J.; Marostica, E.; Yuan, W.; Jin, J.; Zhang, J.; Li, R.; Tang, H.; Wang, K.; Li, Y.; Wang, F.; Peng, Y.; Zhu, J.; Zhang, J.; Jackson, C. R.; Zhang, J.; Dillon, D.; Lin, N. U.; Sholl, L.; …; Yu, K.-H. A pathology foundation model for cancer diagnosis and prognosis prediction. Nat. 634 2024, 970–978. [Google Scholar] [CrossRef]
  119. Whiting, P. F.; Rutjes, A. W. S.; Westwood, M. E.; Mallett, S.; Deeks, J. J.; Reitsma, J. B.; Leeflang, M. M. G.; Sterne, J. A. C.; Bossuyt, P. M. M.; QUADAS-2 Group. QUADAS-2: A revised tool for the quality assessment of diagnostic accuracy studies. Ann. Intern. Med. 2011, 155(8), 529–536. [Google Scholar] [CrossRef]
  120. Wong, A.; Otles, E.; Donnelly, J. P.; Krumm, A.; McCullough, J.; DeTroyer-Cooley, O.; Pestrue, J.; Phillips, M.; Konye, J.; Penoza, C.; Ghous, M.; Singh, K. External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Intern. Med. 2021, 181(8), 1065–1070. [Google Scholar] [CrossRef]
  121. Wong, E.; Bermudez-Cañete, A.; Campbell, M. J.; Rhew, D. C. Bridging the digital divide: A practical roadmap for deploying medical artificial intelligence technologies in low-resource settings. Popul. Health Manag. 2025, 28(2), 105–114. [Google Scholar] [CrossRef]
  122. World Health Organization; Health Level Seven International. WHO and HL7 collaborate to support adoption of open interoperability standards. 2023. Available online: https://www.who.int/news/item/03-07-2023-who-and-hl7-collaborate-to-support-adoption-of-open-interoperability-standards.
  123. World Health Organization; World Bank; International Renewable Energy Agency; Sustainable Energy for All. Energizing health: Accelerating electricity access in health-care facilities; World Health Organization, 2023; Available online: https://www.who.int/publications/i/item/9789240066960.
  124. World Health Organization. Ethics and governance of artificial intelligence for health: WHO guidance. 2021. Available online: https://www.who.int/publications/i/item/9789240029200.
  125. World Health Organization. Global genomic surveillance strategy for pathogens with pandemic and epidemic potential, 2022–2032. 2022. Available online: https://www.who.int/publications-detail-redirect/9789240046979.
  126. World Health Organization. Regulatory considerations on artificial intelligence for health. 2023. Available online: https://www.who.int/publications/i/item/9789240078871.
  127. World Health Organization. Addressing future cybersecurity threats in digital health: Report of the technical consultation, Geneva, Switzerland. 2024. Available online: https://www.who.int/publications/b/75079.
  128. World Health Organization. Guidance for human genome data collection, access, use and sharing. 2024. Available online: https://www.who.int/publications/i/item/9789240102149.
  129. Wynants, L.; Van Calster, B.; Collins, G. S.; Riley, R. D.; Heinze, G.; Schuit, E.; Albu, E.; Arshi, B.; Bellou, V.; Bonten, M. M. J.; Dahly, D. L.; Damen, J. A. A.; Debray, T. P. A.; de Jong, V. M. T.; De Vos, M.; Dhiman, P.; Haller, M. C.; Harhay, M. O.; Henckaerts, L.; …; van Smeden, M. Prediction models for diagnosis and prognosis of COVID-19: Systematic review and critical appraisal. BMJ 369 2020, m1328. [Google Scholar] [CrossRef]
  130. Xu, Z.; Li, J.; Yao, Q.; Li, H.; Zhao, M.; Zhou, S. K. Addressing fairness issues in deep learning-based medical image analysis: A systematic review. npj Digit. Med. 7 2024, 286. [Google Scholar] [CrossRef]
  131. Yang, H. S.; Çubukçu, H. C.; Del Ben, F.; Duan, X.; Feng, J.; Frans, G.; Gruson, D.; King, M.; Nieto-Moragas, J.; Wang, B.; Wang, F. Implementation of AI systems in the clinical laboratory: Insights from an expert survey and recommendations for best practice. Clin. Chem. Lab. Med. 2026, 64(7), 1513–1526. [Google Scholar] [CrossRef]
  132. You, J.; Seok, H. S.; Kim, S.; Shin, H. Advancing laboratory medicine practice with machine learning: Swift yet exact. Ann. Lab. Med. 2025, 45(1), 22–35. [Google Scholar] [CrossRef]
  133. Zech, J. R.; Badgeley, M. A.; Liu, M.; Costa, A. B.; Titano, J. J.; Oermann, E. K. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study. PLoS Med. 2018, 15(11), e1002683. [Google Scholar] [CrossRef] [PubMed]
  134. Zhou, H.; Chen, Y.; Lipton, Z. C. Evaluating model performance in medical datasets over time. Proc. Mach. Learn. Res. 209 2023, 498–508. Available online: https://proceedings.mlr.press/v209/zhou23a.html.
  135. Zini, G.; Mancini, F.; Rossi, E.; Landucci, S.; d’Onofrio, G. Artificial intelligence and the blood film: Performance of the MC-80 digital morphology analyzer. Int. J. Lab. Hematol. 2023, 45(6), 881–889. [Google Scholar] [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.