Submitted:
14 September 2026
Posted:
15 September 2026
You are already at the latest version
Abstract
Artificial intelligence (AI) increasingly interprets vibration and dynamic-response measurements for bridge nondestructive testing (NDT) and structural health monitoring (SHM), yet predictive performance within a dataset does not establish diagnostic or operational maturity. This investigation searched Scopus, Web of Science Core Collection, and IEEE Xplore through 23 August 2026, screening 3,412 records and retaining 727 peer-reviewed studies as the systematic core. Unlike prior reviews organised by algorithm type, each study was coded at the level of the individual validation case for diagnostic capability (D0-D4), validation realism (R1-R4), damage authenticity (A0-A4), generalisation (G0-G3), temporal maturity (T0-T2), engineering-use maturity (M0-M3), and five independent trustworthiness attributes (E, U, X, P, F). The corpus is dominated by enabling inference and detection (D0-D1, 67.4%), whereas prognosis (D4) appears in only 6 studies (0.8%). Numerical validation predominates (R1, 48.1%), confirmed full-scale physical damage occurs in 61 studies (8.4%), and only 8 (1.1%) combine operational monitoring with confirmed damage. Cross-bridge transfer (2.1%), long-term monitoring (2.2%), decision support (0.4%), and explicit trustworthiness (none of the five attributes in 67.0%) are all rare, and only one study combines operational validation, confirmed damage, and verified transfer. The principal gap is therefore not AI capability but the scarcity of evidence where capability, realistic validation, authentic damage, transferability, durability, trustworthiness, and decision relevance converge.
Keywords:
bridge nondestructive testing
; structural health monitoring
; artificial intelligence
; dynamic response
; damage diagnosis
; domain adaptation
; physics-informed machine learning
; uncertainty quantification
1. Introduction
Bridges are long-life transportation assets exposed to ageing, repeated traffic loading, environmental actions, accidental events, and progressive deterioration [1,2,3,4]. Periodic visual inspection and conventional nondestructive testing remain indispensable for infrastructure management, but discrete inspection campaigns cannot continuously characterise changes in structural behaviour between assessment intervals. Structural health monitoring (SHM) therefore provides an important complementary capability by acquiring structural responses under operational or controlled excitation and using those measurements to identify changes that may warrant engineering attention.
In this investigation, AI-enabled bridge nondestructive testing is used as an umbrella term for nondestructive inference of structural condition from vibration and dynamic-response measurements, including continuous SHM implementations where the monitored response is interpreted diagnostically without altering the structure [1,2,3]. This usage is intended to connect the nondestructive-testing and SHM literature without implying that every included study uses identical terminology.
Vibration and other dynamic-response measurements are attractive because structural stiffness, mass, damping, boundary conditions, and load-transfer mechanisms influence observable response characteristics [1,2,3,4]. Acceleration, dynamic strain, displacement, modal properties, and indirect vehicle-bridge interaction measurements can therefore contain information related to structural condition. The inference problem is nevertheless indirect. Temperature, traffic, wind, excitation amplitude, support behaviour, sensor condition, and other environmental and operational factors can change measured response without corresponding physical deterioration. AI-enabled assessment must distinguish damage-sensitive information from normal structural and operational variability, not simply classify differences between datasets.
Artificial intelligence has greatly expanded the analytical strategies available for this problem [1,2,3,4]. Conventional machine-learning methods map engineered statistical, spectral, modal, or time-frequency features to condition states, whereas deep-learning architectures can learn representations directly from raw or minimally processed measurements. More recent developments include healthy-only anomaly detection, self-supervised learning, open-set recognition, transfer learning, domain adaptation, physics-guided models, uncertainty-aware inference, explainable AI, and digital-twin integration. Recent reviews already synthesise broad AI applications in bridge structural-health management, vibration-based intelligent diagnosis, physics-informed machine learning, and machine-learning-based structural-health diagnosis. The present work therefore does not claim novelty from cataloguing AI architectures themselves. Its contribution is instead methodological: to the best of our knowledge, the present investigation applies an ordinal, case-level evidence-maturity appraisal by coding every study in the final corpus against independent axes of diagnostic capability, validation realism, damage authenticity, generalisation, temporal maturity, and engineering-use maturity, rather than synthesising the literature by algorithm type. Where prior reviews describe what methods exist and note gaps narratively, this investigation quantifies, at the level of the individual study, what has actually been demonstrated and under what evidential conditions.
A related published review by Bacha et al. addressed the broader end-to-end domain of dynamic-response-based bridge monitoring and structural assessment, including sensing, damage-sensitive indicators, stiffness- and capacity-related inference, operational and extreme loading, machine learning, and digital-twin integration [1,2,3,4]. That article adopted a structured scoping methodology and explicitly did not claim PRISMA-style comprehensiveness. The present investigation addresses a narrower question. It evaluates how far AI-enabled vibration and dynamic-response methods have progressed in diagnostic capability and how strong the structural evidence supporting those capabilities actually is.
This orientation also distinguishes the present investigation methodologically from the influential surveys that define the field. Widely cited reviews of vibration-based and AI-enabled structural health monitoring are predominantly taxonomic: they organise the literature by method, from traditional techniques through machine learning to deep learning, or from supervised through unsupervised and hybrid learning; describe representative applications; and close with a narrative account of open challenges such as data scarcity, generalisation, field validation, and interpretability [1,2,3,4]. That structure establishes what methods exist and where progress is concentrated. It does not, however, measure how strongly each reported capability is supported by evidence. The present research is therefore complementary rather than competing it takes the gaps those surveys identify qualitatively and resolves them quantitatively, at the level of the individual study, so that a high classification score, a field dataset, a transfer-learning label, or a digital-twin architecture is not mistaken for demonstrated engineering maturity.
This distinction is increasingly necessary because algorithmic sophistication and engineering evidence do not progress in parallel [14]. Duran et al. evaluated a one-dimensional convolutional neural network using vehicle-mounted acceleration measurements from a full-scale reinforced-concrete bridge subjected to progressive controlled physical changes [14]. Such evidence is fundamentally different from a numerically generated damage-classification problem, even if the latter reports a higher predictive metric [14].
Similarly, external transfer and real bridge data require careful interpretation [7,15]. Ferreira et al. transferred healthy-state information between two closely related real concrete bridges using domain adaptation and Bayesian model calibration [7]. Mousavi et al. combined unsupervised learning with multi-domain feature fusion; the Werrington analysis used generated anomaly scenarios on real measurements, whereas the S101 case provided external validation on a separate bridge damage dataset [15]. These studies demonstrate why the phrases field validation, real bridge data, and cross-bridge generalisation cannot be treated as interchangeable descriptors [7,15].
The same problem arises for prediction and digital twins [9,15]. An AI model may predict future acceleration, strain, displacement, or modal properties without forecasting deterioration [9]. A digital twin may provide synchronisation, response estimation, model updating, or anomaly detection without demonstrating prognosis or maintenance optimisation [15]. As a result, a literature synthesis based mainly on reported accuracy or algorithm type risks grouping together studies that answer fundamentally different structural questions.
The present investigation therefore changes the unit of comparison from AI architecture to diagnostic evidence [1,2,3]. Six independent evidence dimensions are used. Diagnostic capability identifies what the AI system actually infers. Validation realism identifies the physical setting in which that capability is tested. Damage authenticity distinguishes numerical, laboratory, synthetically imposed, and physically confirmed structural states. Generalisation records whether inference remains within one structural domain or transfers to unseen conditions and bridges. Temporal maturity identifies exposure to short, repeated, or long-term monitoring. Engineering-use maturity establishes whether the output remains algorithmic, provides structural diagnosis, supports inspection or warning, or contributes explicitly to prognosis and maintenance decisions.
Trustworthiness is evaluated separately through environmental and operational robustness, uncertainty quantification, explainability or physical interpretation, physics integration, and information fusion [1,2,3]. Because these attributes address different failure modes, they are not collapsed into a single score [1,2,3,14]. A physics-informed model may still contain uncertain structural assumptions; an uncertainty-aware model may remain bridge-specific; an explainable model may still perform only modal-property estimation; and a fused representation may use several mathematical domains without combining independent physical sensing modalities.
RQ1. What vibration and dynamic-response information is used by AI-enabled bridge NDT, and how is that information represented for learning?
RQ2. What diagnostic capabilities, from enabling inference to detection, localisation, characterisation, and prognosis, have actually been demonstrated?
RQ3. How does demonstrated diagnostic capability compare with validation realism and physical damage authenticity?
RQ4. How effectively do existing approaches address environmental and operational variability, domain shift, and cross-bridge generalisation?
RQ5. To what extent are physics, uncertainty quantification, explainability, and information fusion incorporated into trustworthy AI-enabled bridge assessment?
RQ6. How far have current systems progressed from algorithmic outputs toward inspection support, deterioration prognosis, maintenance planning, and asset-management decisions?
The principal contribution is an evidence-centred maturity framework comprising diagnostic capability (D0-D4), validation realism (R1-R4), damage authenticity (A0-A4), generalisation (G0-G3), temporal maturity (T0-T2), and engineering-use maturity (M0-M3), together with independent trustworthiness attributes for environmental robustness (E), uncertainty quantification (U), explainability (X), physics integration (P), and information fusion (F) [1,2,3]. Rather than asking only whether AI can detect bridge damage, the investigation asks whether the reported capability has been demonstrated under evidence conditions consistent with credible engineering use.
Applied to the final 727-study core, the framework exposes a pronounced convergence gap. Two-thirds of the corpus (67.4%) reaches only enabling inference or detection (D0-D1), while confirmed prognosis (D4) appears in just 6 studies (0.8%). Operational validation is comparatively common (R4, 200 studies, 27.5%), yet only 8 studies (1.1%) combine operational monitoring with confirmed full-scale physical damage (R4+A4), and genuine cross-bridge transfer (G2-G3) is demonstrated in 15 studies (2.1%). None of the five trustworthiness attributes is demonstrated in 487 studies (67.0%), and a single study simultaneously satisfies operational realism, confirmed damage, and verified cross-bridge transfer. The limiting factor is therefore not AI capability but the scarcity of evidence in which capability, realistic validation, authentic damage, transferability, and engineering relevance converge.
2. Investigation Design, Literature Search, and Evidence-Assessment Framework
2.1. Investigation Design
This study adopts a structured critical investigation supported by a systematic literature search and multidimensional evidence-maturity assessment [5,6]. A statistical meta-analysis of reported predictive performance is not attempted because the bridge-AI literature is heterogeneous in structural typology, sensing configuration, excitation, damage definition, learning task, validation design, and performance metric. Pooling accuracy, F1-score, reconstruction error, or related indicators across such different experimental problems could imply comparability where none exists.
Systematic-search procedures are used to construct the primary evidence corpus, while the synthesis evaluates the structural meaning and evidential strength of each study [5,6]. PRISMA-compatible identification, deduplication, screening, eligibility, and inclusion stages are used for transparent reporting. The formal Scopus, Web of Science Core Collection, and IEEE Xplore searches yielded 3,412 raw records. After database deduplication and targeted recovery, 1,679 records entered title/abstract screening; 839 proceeded to source-level full-text eligibility assessment; and 727 peer-reviewed journal studies satisfied the systematic-core eligibility criteria. These counts were finalised before corpus-level analysis and are the sole basis for the prevalence and maturity statistics reported here; every corpus-level count and percentage in this investigation derives from the final study-level coding in Supplementary Table S1.
2.2. Search Architecture and Eligibility
The systematic core covers English-language peer-reviewed journal literature published between 1 January 2015 and the final database-search timestamp on 23 August 2026 [5,6]. No geographic restriction is applied. The search yielded 3,412 raw database records: 1,969 from Scopus, 1,295 from Web of Science Core Collection, and 148 from IEEE Xplore. The 2026 publication year is therefore partial and should not be compared directly with completed calendar years without this cutoff qualification.
The principal databases are Scopus, Web of Science Core Collection, and IEEE Xplore [5,6]. Backward and forward citation tracing and targeted thematic searches supplement the database queries, particularly for transfer learning, self-supervised learning, physics-informed methods, uncertainty quantification, explainable AI, information fusion, digital twins, and prognosis [5,6,35,43].
A peer-reviewed article is eligible if a publisher-designated citable version with a permanent DOI was publicly available by the search cutoff, including a publisher-posted accepted manuscript [5,6]. Repository-only preprints are not included in the systematic primary corpus.
Search architecture. The core search combines four conceptual blocks: bridge structure × dynamic response × artificial intelligence × structural assessment [5,6]. The bridge block includes bridge* and viaduct*. The dynamic-response block includes vibration*, dynamic response, acceleration, modal, and operational modal analysis. The AI block includes artificial intelligence, machine learning, deep learning, neural networks, CNNs, recurrent networks, transformers, autoencoders, unsupervised learning, self-supervised learning, transfer learning, domain adaptation, physics-informed or physics-guided learning, and Gaussian-process approaches. The assessment block includes structural health monitoring, SHM, damage detection, localisation, classification, condition assessment, diagnosis, prognosis, anomaly detection, and predictive maintenance.
Digital twin is not treated as an AI synonym in the core search [5,6,35,43]. Digital-twin studies are retrieved through a separate supplementary thematic query that combines bridge, dynamic-response, and AI or diagnostic/prognostic terms [5,6]. This avoids importing digital-twin architecture papers that do not contain an AI-enabled structural inference component.
Table 1 summarises the investigation protocol and the final corpus-construction counts [5,6] (Table 1). After removal of 1,743 duplicates, 1,669 database-unique records remained. Targeted recovery added 10 eligible studies identified through supplementary searches, producing a final identification set of 1,679 records. Title/abstract screening excluded 840 records and advanced 839 to source-level full-text eligibility. The final full-text outcome comprised 727 systematic-core studies, 91 supporting primary studies that fell outside one core eligibility boundary, and 21 contextual or excluded records. These full-text eligibility counts are final for PRISMA reporting.
Eligibility. Core primary studies must concern a bridge, bridge component, or experimental/numerical structural model explicitly developed and evaluated by the original authors for bridge monitoring, bridge damage assessment, or bridge-related moving-load response; use vibration, acceleration, dynamic strain, displacement, modal information, vehicle-induced response, or another qualifying dynamic-response quantity; employ an AI or machine-learning component; address structural inference, anomaly detection, damage assessment, condition classification, localisation, prognosis, or a directly enabling SHM function; provide sufficient validation information for evidence coding; and be a peer-reviewed journal publication [5,6].
Visual-only defect detection without qualifying dynamic-response information, sensor-fault classification without structural-condition inference, inspection-database-only deterioration prediction, non-peer-reviewed preprints, and unrelated structural-AI studies are excluded from the systematic core [5,6]. Review articles and broader digital-twin or management studies can be retained as contextual literature but are not counted as primary evidence.
2.3. Evidence Extraction and Case-Level Coding Framework
For each eligible study, information is extracted concerning the structure, sensing configuration, response variable, representation strategy, learning paradigm, diagnostic target, validation environment, damage provenance, environmental and operational variability, cross-structure transfer, monitoring duration, uncertainty treatment, explainability, physics integration, information fusion, and engineering output [5,6].
Coding stability was assessed by comparing the initial and reconciled final codes for all 727 studies. A targeted source-and-consistency reconciliation examined 336 studies and revised at least one code in 277. Agreement was quantified using quadratic-weighted Cohen’s κ for the six ordinal dimensions and unweighted Cohen’s κ for the five binary trustworthiness attributes. Ordinal agreement was temporal maturity κ = 0.98, damage authenticity κ = 0.88, diagnostic capability κ = 0.82, generalisation κ = 0.80, engineering-use maturity κ = 0.80, and validation realism κ = 0.63. Binary agreement was: explainability κ = 1.00, physics integration κ = 0.92, environmental robustness κ = 0.54, information fusion κ = 0.45, and uncertainty quantification κ = 0.33. The lower values for uncertainty, fusion, and environmental robustness identify the most interpretation-sensitive judgements and support the conservative explicit-evidence rules used in final coding. Supplementary Table S2 reports per-dimension percent agreement and confusion between passes. This is a first-pass-to-final stability check within a single coding team rather than fully independent dual coding.
Diagnostic capability (D). Diagnostic capability is coded as D0, enabling inference such as modal identification, virtual sensing, response reconstruction, load estimation, missing-data reconstruction, or response forecasting without direct condition diagnosis; D1, detection of abnormal behaviour, structural change, or damage presence; D2, localisation of the affected region or component; D3, damage-state/type classification, severity estimation, or quantitative condition-related parameter identification; and D4, prediction of future deterioration, cumulative damage, structural-condition evolution, residual performance, or remaining useful life [5,6]. Prediction of future structural response remains D0 unless future condition evolution itself is estimated.
Validation realism and damage authenticity (R, A). Validation realism is coded independently as R1 numerical/synthetic structural validation, R2 laboratory or reduced-scale physical validation, R3 full-scale bridge or established physical benchmark tested in a dedicated campaign, and R4 in-service bridge monitoring under operational conditions [5,6]. Damage authenticity is coded as A0 no confirmed physical damage, A1 numerical/synthetic damage, A2 physical laboratory damage, A3 real bridge measurements with synthetically imposed or numerically transferred damage/anomalies, and A4 confirmed physical structural change or damage in a full-scale bridge.
Where provenance can be established, A4 can be described as A4-controlled, A4-intervention, or A4-confirmed [5,6]. These labels are descriptive subtypes, not additional ordinal levels. The core methodological rule is that R4 does not imply A4.
Generalisation, temporal, and engineering-use maturity (G, T, M). Generalisation is classified as G0 same structure and broadly comparable development-test domain, G1 unseen environmental or operating conditions on the same structure, G2 genuine transfer or independent validation on an unseen bridge, and G3 population/network-level generalisation involving multiple distinct bridges and actual transfer of learned information [5,6]. Target-domain supervision is recorded as ZS zero-shot, UT unsupervised target-domain information, FS few-shot labelled target data, or ST substantial supervised target retraining.
Figure 1.
Systematic corpus construction and evidence-assessment workflow. Database searches returned 3,412 raw records; 1,743 duplicates were removed; 10 eligible studies were recovered through supplementary targeted searches; 1,679 records entered screening; 839 underwent source-level full-text eligibility assessment; and 727 formed the systematic core. Evidence-maturity coding was performed after eligibility was fixed.
Figure 1.
Systematic corpus construction and evidence-assessment workflow. Database searches returned 3,412 raw records; 1,743 duplicates were removed; 10 eligible studies were recovered through supplementary targeted searches; 1,679 records entered screening; 839 underwent source-level full-text eligibility assessment; and 727 formed the systematic core. Evidence-maturity coding was performed after eligibility was fixed.

Temporal maturity is coded as T0 short campaign, T1 repeated or medium-duration monitoring, and T2 long-term operational monitoring, with actual duration recorded where reported [5,6]. Engineering-use maturity is coded as M0 enabling or algorithmic output, M1 structural diagnostic information, M2 warning, inspection, or intervention support, and M3 explicit prognosis and/or maintenance or asset-management decision support.
2.4. Trustworthiness Attributes
Five trustworthiness attributes are recorded independently: E, environmental or operational variability explicitly addressed; U, uncertainty quantitatively assessed; X, explainability or physical interpretation demonstrated; P, structural physics explicitly incorporated into AI/hybrid inference; and F, information fusion demonstrated [5,6]. The framework deliberately avoids a single weighted maturity score because studies can possess very different but equally important strengths.
3. From Bridge Dynamic Response to AI-Ready Information
This section addresses RQ1: what makes bridge dynamic-response data AI-ready, and how representation choices condition every downstream diagnostic claim.
3.1. Structural Observability and Representation
AI performance is fundamentally limited by what the sensing system makes observable [44,45,46,47]. A sophisticated neural architecture cannot infer a structural state that produces no sufficiently identifiable signature in the measured response [48,49,50,51]. Acceleration is widely used because it captures structural dynamics under ambient, traffic-induced, impact, forced, or indirect vehicle excitation [13,14]. Dynamic strain provides more local information concerning deformation and fatigue demand, displacement can reveal lower-frequency or quasi-static changes, and modal quantities compress response measurements into physically interpretable system-level descriptors [19]. Drive-by sensing creates a different observation problem because bridge information is embedded within the coupled vehicle-bridge response [13,14].
Malekjafarian et al. provide an early AI-based example of this indirect formulation using numerical vehicle-bridge interaction data [13,14]. Duran et al. later demonstrated that vehicle-mounted acceleration could support multi-state classification on a full-scale bridge subjected to controlled physical condition changes [13,14]. The progression between these studies illustrates why sensing modality alone does not establish evidence maturity [11,12,13,16].
Representation strategies. AI-ready dynamic information can be organised into three broad representation classes [17,18,19,20]. Engineered features include modal frequencies, damping ratios, mode shapes, statistical descriptors, spectral characteristics, and other damage-sensitive indices [19]. Transformed representations include frequency-domain and time-frequency descriptions [25,26,27,28]. Raw or minimally processed response histories allow deep networks to construct internal latent representations [29,30,31,32]. The three classes are not successive replacements; representation should be selected according to the diagnostic objective, available data volume, environmental variability, required interpretability, and intended transferability [14,33,34,35].
3.2. Modal Inference and Operational Context
Operational modal analysis remains central to vibration-based bridge SHM, but automatic extraction of modal properties is best interpreted as an enabling function unless a subsequent validated step translates those quantities into structural condition [19]. Haghbin et al. developed an attention-based CNN that predicts modal properties and uses explainable-AI methods to interrogate the learned inference [60]. In the present framework this remains D0 because the principal AI output is modal identification rather than damage diagnosis [19].
The same rule applies to virtual sensing, missing-data reconstruction, and future response prediction [17,38]. Such enabling steps can be essential to an SHM pipeline without themselves establishing D1-D4 structural diagnosis [53,54,55,56].
Environmental and operational context. Temperature, traffic, wind, excitation level, support conditions, and sensor state are more than nuisance variables [57,58,59,61]. They can provide contextual information required to determine whether a measured change is structurally meaningful [62,63,64,65]. AI systems can statistically normalise features, condition baselines on operating state, construct latent domains, or identify observations outside the development distribution [66,67,68,69].
Fusion terminology. Information fusion should be described precisely [10,11,60,70]. Multi-sensor fusion combines measurements from different sensors; multi-physical-variable fusion combines quantities such as acceleration and strain; multi-domain fusion combines alternative mathematical representations of the same measurement; and feature- or decision-level fusion combines learned representations or model outputs [12,13,16,17]. Mousavi et al. integrate raw acceleration with time-, frequency-, and wavelet-domain features [15]. Their approach is most precisely described as multi-domain feature fusion rather than multimodal sensing [19].
Representation and generalisation. Representation determines not only within-domain predictive performance but also what can transfer [26,27,28,29]. Raw responses can preserve subtle condition information while encoding sensor and excitation specifics [30,31,32,33]. Modal features have clearer physical interpretation but vary with geometry, boundary conditions, and environmental state [14,19]. Time-frequency representations may isolate transient signatures but depend on scale, window, and normalisation choices [37,38,39,40].
3.3. Answer to RQ1: AI-Ready Bridge Information
AI-enabled bridge NDT draws on acceleration, dynamic strain, displacement, modal information, drive-by response, and other qualifying dynamic-response quantities [13,14,19]. These data can be supplied to learning algorithms as engineered, transformed, raw, or fused representations [46,47,48,49]. No single representation is universally preferable; usefulness depends on structural observability, environmental context, diagnostic target, interpretability, and transfer requirements [50,51,52,53].
Figure 2.
Evidence-centred architecture for AI-enabled bridge NDT. Dynamic-response information and representation support AI inference, but diagnostic capability is interpreted jointly with validation realism, damage authenticity, generalisation, temporal maturity, engineering-use maturity, and the independent trustworthiness attributes E/U/X/P/F.
Figure 2.
Evidence-centred architecture for AI-enabled bridge NDT. Dynamic-response information and representation support AI inference, but diagnostic capability is interpreted jointly with validation realism, damage authenticity, generalisation, temporal maturity, engineering-use maturity, and the independent trustworthiness attributes E/U/X/P/F.

4. AI Paradigms for Vibration and Dynamic-Response Bridge Assessment
This section maps the AI paradigms applied to vibration and dynamic-response data, establishing the modelling vocabulary used in the evidence assessment that follows.
4.1. Supervised, Unsupervised, and Open-Set Learning
Supervised learning provides a direct mapping from response information to predefined structural states [15,71,72,73]. Conventional classifiers operate on engineered features, whereas CNNs, recurrent networks, transformers, and other deep architectures can learn increasingly complex representations from raw or transformed signals [9,16,18,74]. The central limitation is the requirement for representative labelled structural states [19,21,24,75]. Healthy data can often be acquired continuously, but labelled measurements corresponding to several damage mechanisms and severities are difficult to obtain from operational bridges [27,76,77,78]. As a result, many supervised studies rely on numerical simulations, laboratory structures, or controlled physical-damage campaigns [35,37,79,80].
Full-scale controlled classification studies demonstrate that supervised learning can move beyond numerical proof of concept, but high within-study accuracy remains conditional on the structure, excitation, damage sequence, sensor arrangement, and test design [39,81,82,83].
Healthy-only and unsupervised learning. The scarcity of damaged-state labels has motivated one-class and unsupervised approaches that model normal structural behaviour and identify deviations from that learned baseline [44,45,47]. Park and Kim trained a convolutional-autoencoder/Deep-SVDD framework using intact vibration data and evaluated it against a damaged in-situ steel truss bridge [51,54,85,86]. Giglioni et al. similarly used autoencoders to learn healthy behaviour before evaluating novelty against progressive full-scale physical benchmark changes [79].
The limitation is that an anomaly can arise from damage, environmental change, sensor malfunction, maintenance intervention, or a previously unseen operating state [63,89,90,91]. Healthy-only learning aligns naturally with D1 detection, but additional evidence is normally required before an anomalous observation can be interpreted as a particular location, mechanism, or severity [7,69,70]; this builds on earlier machine-learning and novelty-based bridge detection work [100,101,121,122].
Self-supervised and open-set learning. Self-supervised learning exploits structure within unlabelled monitoring data to construct representations without conventional damage labels [8,15,71,73]. Open-set learning addresses the unrealistic assumption that every future observation belongs to a condition represented during training [16,72,73,74]. These methods therefore relax important data assumptions, but their diagnostic capability must still be evaluated independently from validation realism [9,18,19,75]. For example, open-set damage classification has been evaluated on Z24 bridge data, while physics-encoded unsupervised learning has been demonstrated through numerical and laboratory bridge-network cases [16,21,24,62]. Neither methodological sophistication nor label efficiency substitutes for operational field validation [27,35,78,79].
4.2. Transfer Learning, Domain Adaptation, and Population-Based SHM
Transfer learning addresses one of the central obstacles to network-scale monitoring: a model developed on one structure rarely encounters an identical statistical and physical domain on another [7,37,39,70]. Ferreira et al. illustrates a realistic case in which healthy-state information from one bridge is adapted for a related target bridge that lacks a healthy baseline, with Bayesian finite-element calibration used to address model uncertainty [7].
Cross-bridge transfer should be distinguished according to the amount of target information used [7,47,51,54]. Zero-shot transfer, unsupervised target adaptation, few-shot fine-tuning, and substantial supervised retraining are not equivalent deployment conditions [7,57,70]. Above all, use of a transfer-learning algorithm does not by itself establish G2 [7,62,63,70]. If numerical information is transferred to an experimental version of the same structural system, or if a new model is extensively retrained for every bridge, the evidence remains fundamentally different from a locked model tested on a previously unseen bridge [7,69,70].
Population-based SHM. Population-based SHM extends transfer from paired source-target structures toward groups of related assets [7,8,70]. Laboratory populations are useful because geometry, condition state, and environmental effects can be controlled systematically while source-target relationships vary [15,71,72,73]. More recent numerical bridge-network studies investigate population-level similarity requirements for transfer [7,9,16,18]. These approaches provide strong methodological evidence for transferable diagnosis, but their validation realism must remain visible [7,19,21,24].
4.3. Physics-Informed Learning and the Evolution of AI Assumptions
Physics-informed machine learning constrains or augments data-driven inference using structural knowledge [16,27,62]. Physics can enter through governing equations, finite-element models, physically meaningful latent variables, grey-box relationships, hard architectural constraints, or physics-informed loss functions [16,35,37,62]. Examples include a physics-encoded unsupervised bridge-network framework, a physics-informed early-warning and stiffness-identification system on an operational railway bridge, and grey-box thermal-response modelling within operational viaduct monitoring [16,39,62]. The presence of physics should therefore be coded according to where it enters the inference chain and what capability it demonstrably improves [16,44,45,47].
Evolution of AI assumptions. The development of AI-enabled bridge assessment is better interpreted as a progressive relaxation of restrictive data assumptions than as a simple sequence from conventional ML to deep learning: fully labelled bridge-specific classification, healthy-only modelling, self-supervised/open-set learning, domain adaptation, cross-bridge transfer, and population-level or physics-guided inference [7,16,51,54]. The sequence is not a ranking of algorithm quality [57,62,87,88]. More advanced learning paradigms are valuable only to the extent that they remove a practical limitation and retain their capability under increasingly independent and physically representative validation [63,89,90,91].
5. From Enabling Inference to Prognosis: What Can AI Actually Diagnose?
This section addresses RQ2: which diagnostic capabilities, from enabling inference (D0) to prognosis (D4), the corpus actually demonstrates.
5.1. The Diagnostic-Capability Ladder (D0-D4)
D0 comprises AI functions that improve observability, reconstruction, interpretation, or prediction of structural response without directly determining structural condition [12,15,71,72]. It is the largest single enabling category after reconciliation, containing 244 of 727 studies (33.6%) [13,73,94,95]. Examples include AI-assisted operational modal identification, virtual sensing, response reconstruction, missing-data imputation, load or cable-force inference, and strain or displacement forecasting [9,16,74,75]. These functions can be essential to an SHM pipeline, but they remain D0 unless a validated downstream step infers structural condition [21,22,76,77].
D1 - detection of abnormal structural behaviour. D1 is reached when the AI output distinguishes a reference or healthy state from an abnormal, changed, or damaged state [78,79,85]. The final corpus contains 246 D1 studies (33.8%), making detection the most common direct diagnostic endpoint [31,32,33,78]. Healthy-only novelty detection and unsupervised anomaly detection can reach D1 without damaged training labels, but the interpretation remains conditional because an anomaly may arise from structural change, environmental variation, maintenance, or data problems unless those causes are independently distinguished [14,36,78,85].
D2 - localisation. D2 requires identification of the affected structural region or component [16,80,81]. The final corpus contains 102 D2 studies (14.0%) [16,40]. Localisation is demonstrated under numerical, laboratory, controlled full-scale, and selected operational settings, but its evidential meaning depends strongly on whether the damage location is synthetic, physically introduced, or independently confirmed on an in-service bridge [16,41,82,83].
D3 - structural-state characterisation and quantitative diagnosis. D3 includes damage type or state classification, severity estimation, or quantitative damage-related parameter identification [48,65,95]. The final corpus contains 129 D3 studies (17.7%) [65,84,95]. D3 encompasses quantitative stiffness-loss, scour-depth, cable-condition, corrosion-level, and damage-severity inference [52,54,65,85], extending an earlier line of neural-network severity-quantification and vibration-based condition-assessment studies [99,103,123]. D3 therefore reflects a richer structural answer than binary detection, but the final cross-tabulation shows that D3 is still most commonly supported by numerical validation rather than operational physical-damage evidence [65,86,95,110].
Selected operational D3 evidence is also emerging [65,87,88,95]. Khan et al. reconstructed a spatial stiffness field on a railway bridge containing an inspection-confirmed fatigue crack, localising the affected region and quantifying stiffness reduction under operational conditions [8]. This demonstrates that D3 can occur at R4/A4, but it does not imply that all D3 methods have reached comparable field maturity [65,95,110].
D4 - prognosis of future structural conditions. D4 is reserved for prediction of future deterioration, cumulative damage, structural-condition evolution, residual performance, or remaining useful life [9,32,68,90]. Only 6 of 727 studies (0.8%) satisfy this strict definition [7,91,92,124]. The six verified cases comprise predictive maintenance and bridge-deck condition forecasting, field-data-based fatigue-damage prognosis of orthotropic steel decks, bridge-cable fatigue-life prediction, long-term crack-width or structural-condition forecasting linked to maintenance decisions, physics-informed fatigue crack and displacement evolution, and remaining fatigue-life prediction for coastal bridge details [8,9,12,32]. Ordinary prediction of future acceleration, strain, displacement, or modal properties remains D0, even when the forecast horizon is long [15,71,72,94].
5.2. Diagnostic Capability Is Not an Overall Maturity Score
The D0-D4 hierarchy is ordinal only with respect to the type of structural question answered [9,13,16,32]. Corpus-wide counts are D0 244 (33.6%), D1 246 (33.8%), D2 102 (14.0%), D3 129 (17.7%), and D4 6 (0.8%) [9,21,32]. A long-term uncertainty-aware D1 detector on an operational bridge can provide stronger deployment evidence than a numerically validated D4 model [9,22,32]. Diagnostic capability must instead be interpreted jointly with R, A, G, T, M, and E/U/X/P/F rather than treated as an overall maturity score [31,79,96,97].
5.3. Answer to RQ2: Demonstrated Diagnostic Capability
The assessed evidence spans the complete D0-D4 hierarchy, but it is concentrated in enabling inference and detection [9,32,33]. D0 and D1 together account for 490 studies (67.4%), whereas D4 accounts for only 6 (0.8%) [9,14,32,36]. AI-enabled bridge NDT has therefore become broad in diagnostic functionality, yet genuine prediction of future structural condition remains exceptional and must not be inferred from response forecasting alone [40,80,81,102].
6. Beyond Accuracy: Validation Realism and Damage Authenticity
This section addresses RQ3: whether demonstrated diagnostic capability is supported by realistic validation (R) and authentic damage evidence (A), rather than reported accuracy alone.
Predictive performance metrics are meaningful only relative to the experimental problem from which they are obtained [10,11,71,72]. A model can achieve high classification accuracy under numerical damage scenarios, perform less strongly under a noisy laboratory experiment, and fail under real traffic, temperature, sensor drift, and structural uncertainty [13,73,94,95]. Conversely, modest predictive performance obtained under a physically authenticated operational deterioration event may carry greater engineering significance than near-perfect accuracy within a synthetic test set [7,8,16,17]. Validation realism and damage authenticity are evaluated independently from diagnostic capability [20,23,24,26].
6.1. Validation Realism (R1-R4)
R1 numerical or synthetic validation remains the largest realism category, with 350 of 727 studies (48.1%) [28,77,78,79]. Numerical studies allow systematic variation of damage location, magnitude, excitation, noise, environmental effects, and sensor placement, but they cannot reproduce the full joint distribution of modelling error, sensor behaviour, maintenance history, environmental variability, traffic, and structural uncertainty encountered in service [29,30,96,97].
R2 - laboratory and reduced-scale validation. R2 laboratory or reduced-scale physical validation accounts for 91 studies (12.5%) [31,32,34,98]. Physical experiments introduce real sensors, measurement noise, structural imperfections, and physically produced changes that are absent from purely numerical studies [14,35,101,102]. They strengthen physical credibility while remaining limited by scale, controlled boundary conditions, and reduced operational variability [7,8,37,38].
R3 - full-scale controlled and benchmark evidence. R3 full-scale controlled or established physical-benchmark validation accounts for 86 studies (11.8%) [83,104,105,106]. R3 evidence is especially valuable because the structure is full-scale while condition states remain controlled or independently documented [42,44,45,46]. The final authenticity matrix shows that 53 of the 61 A4 studies occur at R3, indicating that most confirmed full-scale physical damage evidence still comes from controlled campaigns and benchmark structures rather than routine operational deterioration [7,8,16,47].
R4 - operational bridge monitoring. R4 reflects monitoring of an in-service bridge under actual operating conditions and accounts for 200 studies (27.5%) [7,8,51,52]. R4 is the strongest realism category, but it does not establish damage authenticity [7,8,78]. Of these 200 operational studies, 138 are A0, 5 are A1, 2 are A2, 47 are A3, and only 8 are A4 [7,8,16]. Thus, most operational studies either monitor healthy or unexplained states, or combine field measurements with synthetic or transferred damage representations [7,8].
6.2. Damage Authenticity (A0-A4)
Damage authenticity is distributed as A0 357 (49.1%), A1 186 (25.6%), A2 76 (10.5%), A3 47 (6.5%), and A4 61 (8.4%) [16,81,90,91]. A0 includes healthy monitoring, maintenance-related changes, unexplained anomalies, or enabling inference without confirmed deterioration [7,69,70,124]. A1 represents numerical or synthetic damage, whereas A2 represents physical laboratory or reduced-scale damage [8,10,81,93]. Together, the A0-A2 categories show that a large fraction of the literature still evaluates algorithms without confirmed full-scale structural damage [11,71,72,94].
A3 - synthetic or injected damage on real bridge measurements. A3 is a critical distinction and accounts for 47 studies (6.5%) [13,16,73,81]. In every A3 case, real bridge measurements coexist with synthetic, injected, or numerically transferred damage effects [17,19,20,75]. Because all 47 A3 studies are operational R4 cases in the final matrix, treating real measurements alone as proof of real damage would substantially overstate field validation maturity [7,8,23,24].
A4 - full-scale physical structural change. A4 includes confirmed physical structural change or damage in a full-scale bridge and is demonstrated by 61 studies (8.4%) [16,28,78,79]. Fifty-three A4 studies occur under R3 controlled or benchmark campaigns, while only 8 combine A4 with R4 operational monitoring [7,8,16]. This 8-study subset, 1.1% of the corpus, is therefore particularly important when judging claims of field-ready diagnosis under physically authenticated deterioration [14,32,34,98].
6.3. Corpus-Level Evidence Profile
Table 3 summarises the final corpus at the dimension level [35,80,101,102] (Table 3). Figure 3 then resolves diagnostic capability against validation realism [37,38,81,104] (Figure 3). The combined view shows why performance claims should be interpreted as evidence profiles rather than as a single maturity ranking: substantial D2-D3 capability exists, but nearly half of the literature remains at R1, confirmed full-scale damage is uncommon, and operational physical-damage evidence is rare [7,8,42,78].
6.4. Answer to RQ3: Diagnostic Capability Versus Evidential Realism
Advanced diagnostic capability and strong validation realism are non-equivalent [44,45,46,47]. D3 appears under R1, R2, R3, and R4, with counts of 68, 22, 15, and 24, respectively [7,8,50,78]. Likewise, D4 includes four R1 studies and only two R4 studies [7,8,51,52]. The principal evidential bottleneck is therefore not whether AI can generate advanced outputs, but whether those outputs survive realistic validation with authentic structural change [86,114,117,118].
7. Environmental Variability, Domain Shift, and Cross-Bridge Generalisation
This section addresses RQ4: how robustly AI-enabled methods withstand environmental and operational variability, and whether they generalise beyond the structure on which they were developed.
A bridge-AI model is operationally useful only if its learned relationship between measurements and structural condition remains valid after the statistical and physical environment changes [10,11,12,72]. Environmental and operational variability, temporal drift, sensor differences, structural geometry, boundary conditions, materials, and deterioration mechanisms all create forms of domain shift [12,13,16,20]. Generalisation must be assessed in relation to what changed between development and evaluation [21,22,23,26].
7.1. Environmental Variability and Within-Structure Robustness
Temperature can alter stiffness, support behaviour, thermal deformation, modal properties, and sensor response [12,27,29]. Traffic changes excitation amplitude and frequency content; wind, moisture, boundary-condition changes, train characteristics, road roughness, and sensor replacement can further shift the measurement distribution [31,33,35,36]. EOV can therefore be larger than the change produced by early deterioration [37,38,39,81]. The assessed literature includes physics-aware thermal compensation under operational monitoring, long-term strain modelling using environmental inputs, drive-by damage detection tested across multiple environmental and operational factors, and temperature-aware drive-by damage detection with road-roughness effects [12,42,47,48]. These approaches address different aspects of environmental and operational domain shift and should not be treated as equivalent forms of generalisation [12,81,84,104].
G0 and G1 - within-structure robustness. G0 represents development and evaluation on the same bridge or structural model under broadly comparable conditions and dominates the corpus with 644 studies (88.6%) [51,52,54,117]. G1, which requires evaluation under unseen environmental, operating, or loading conditions on the same structure, is demonstrated by 68 studies (9.4%) [12,56,57,81]. Thus, even before cross-bridge transfer is considered, explicit within-structure robustness beyond the development domain is relatively uncommon [7,54,62,64].
7.2. Cross-Bridge Transfer and Generalisation (G2-G3)
G2 requires genuine transfer or external validation on a bridge not used as the original development domain [7,54,67,69]. Only 12 studies (1.7%) meet this criterion [7,10,92,93]. Studies were not classified as G2 when transfer-learning terminology referred only to simulation-to-experiment adaptation of the same structure, generic pretraining, or multiple bridge applications without a verified source-to-target transfer test [7,11,12,13].
Target-domain supervision. Target-domain supervision is therefore reported only for verified G2/G3 studies [7,16,20,21]. Among the 15 G2/G3 studies, 4 are zero-shot (ZS), 7 use unsupervised target information (UT), 3 use few-shot labelled target data (FS), none require substantial supervised target retraining (ST) under the final coding, and one G3 study has target supervision recorded as not reported or not verified rather than being inferred [7,22,23,26].
Structural similarity and negative transfer. Cross-bridge transfer is not purely statistical [7,27,29,31]. Two datasets can be mathematically alignable while representing structures with different modal density, span arrangement, boundary conditions, sensors, materials, or deterioration mechanisms [33,35,36,37]. Source-target similarity must be considered physically as well as statistically [27,38,39,70]. Transfer-learning foundations for bridge SHM explicitly frame applicability in terms of structural-domain compatibility, while real twin-bridge transfer with Bayesian calibration illustrates how structural similarity and uncertainty enter a practical source-target problem [7,27,42]. Inappropriate source selection can produce negative transfer, so population-based SHM needs methods that determine not only how to transfer, but whether a particular source bridge should be transferred from at all [7,27,51,54].
G3 - population- and network-level generalisation. G3 requires population- or network-scale generalisation across multiple distinct bridges with actual knowledge transfer [7,52,54,72]. Only 3 studies (0.4%) satisfy this requirement [56,57,62,118]. Population-scale ambition is therefore much more common than population-scale evidence [64,67,89,110]. A study is not G3 merely because it evaluates several bridges or trains a model on a large synthetic inventory; the transfer mechanism and held-out target evaluation must be demonstrated [7,54,69,72].
7.3. Transfer-Learning Terminology and Figure 4 Interpretation
A transfer-learning algorithm can still remain G0 if knowledge is moved only from a numerical model to an experimental version of the same structural system or if the target structure is extensively retrained [7,10,11,54]. Lu et al. provide a useful control case in which domain adaptation supports numerical-to-experimental transfer within one reduced-scale cable-stayed bridge system, but this does not constitute unseen-bridge G2 [82].
Figure 4.
Generalisation maturity across the systematic core. G0 and G1 represent same-structure evidence, whereas G2 and G3 require genuine transfer to an unseen bridge or across a bridge population. Target-supervision qualifiers are reported only for the 15 G2/G3 studies.
Figure 4.
Generalisation maturity across the systematic core. G0 and G1 represent same-structure evidence, whereas G2 and G3 require genuine transfer to an unseen bridge or across a bridge population. Target-supervision qualifiers are reported only for the 15 G2/G3 studies.

7.4. Answer to RQ4: Environmental Robustness and Generalisation
AI-enabled bridge assessment has progressed toward genuine cross-bridge transfer, but the evidence remains sparse: G0 644 (88.6%), G1 68 (9.4%), G2 12 (1.7%), and G3 3 (0.4%) [7,20,21,22]. Only 5 studies (0.7% of the corpus) combine G2/G3 with R4 operational bridge evidence [7,23,26,54]. Generalisation remains one of the clearest barriers between high within-study accuracy and scalable bridge-network deployment [27,29,31,33].
8. Toward Trustworthy AI-Enabled Bridge NDT
This section addresses RQ5: the extent to which the corpus demonstrates the trustworthiness attributes (E, U, X, P, F) required for credible engineering use.
High diagnostic accuracy does not by itself establish that an AI system can be trusted for infrastructure assessment [10,11,12,15]. A bridge model may be accurate but environmentally fragile, physically implausible, poorly calibrated, bridge specific, opaque, or overconfident under previously unseen conditions [9,13,16,74]. Trustworthiness is treated here as a multidimensional evidence problem rather than a label attached to a particular AI architecture [18,22,23,24].
8.1. The Five Trustworthiness Attributes
Environmental and operational robustness (E) is explicitly demonstrated in 148 studies (20.4%) [26,28,29,76]. This category requires the AI or inference pathway to handle or be challenged by variables such as temperature, traffic, wind, humidity, hydraulic conditions, road roughness, or other operational changes [31,34,35,97]. Generic measurement-noise tests alone are not counted as E [36,37,38,81].
Uncertainty quantification (U). Quantitative uncertainty assessment (U) is explicitly demonstrated in only 38 studies (5.2%) [10,23,28,42]. U requires probabilistic or quantitative treatment such as posterior distributions, predictive intervals, uncertainty propagation, heteroscedastic modelling, or comparable calibrated measures [10,23,28,45]. Merely reporting prediction error or robustness to added noise is insufficient [51,52,109,113].
Explainability and physical interpretation (X). Explainability or physically interpretable model reasoning (X) is demonstrated in 37 studies (5.1%) [8,16,53,60]. Qualifying evidence includes explicit explainable-AI analysis, interpretable feature attribution, physically meaningful latent or reconstructed states, or model reasoning that can be related to structural behaviour [8,16,56,58]. Architectural attention mechanisms are not counted as X unless interpretability is actually demonstrated [8,16,60].
Physics integration (P). Explicit structural-physics integration (P) is demonstrated in 42 studies (5.8%) [9,62,67,74]. Qualifying strategies include physics-informed or physics-guided loss functions, governing-equation constraints, mechanics-informed architectures, or similarly explicit hybrid inference [8,9,10,60]. The mere use of finite-element data to generate training samples is not sufficient for P unless physics is incorporated into the inference mechanism itself [9,11,12,13].
Information fusion (F). Explicit information fusion (F) is demonstrated in 28 studies (3.9%) [9,15,16,18]. Fusion can occur across sensors, physical variables, modalities, signal domains, features, latent representations, or decisions, but the combination must be an explicit part of the inference architecture rather than simple concatenation that is not evaluated as fusion [15].
8.2. Trustworthiness Attributes Do Not Substitute for One Another
The trustworthiness attributes are independent and rarely converge [28,29,76,97]. Of the 727 studies, 487 (67.0%) demonstrate none of E/U/X/P/F, 195 (26.8%) demonstrate one, 38 (5.2%) demonstrate two, 6 (0.8%) demonstrate three, and only 1 (0.1%) demonstrates four [31,34,35,36]. No study demonstrates all five [37,38,81,104] (Figure 5). The distribution shows that the literature is far more mature in predictive modelling than in explicit evidence of reliability, interpretability, physical consistency, uncertainty, and information integration [8,10,16,23].
8.3. Integrated Field Evidence and Behaviour Under Domain Shift
Integrated trustworthiness under high validation maturity remains rare [45,48,108,109]. Only 7 studies demonstrate at least three of the five E/U/X/P/F attributes, and just one demonstrates four [51,52,53,113]. These studies are important not because they define a single preferred architecture, but because they illustrate the evidential standard required when multiple failure modes are considered simultaneously [56,115,117,118]. The remaining research needed is to combine these attributes with authenticated physical damage and external transfer rather than demonstrating each property in isolation [58,62,88,119].
Trustworthiness under domain shift. A model intended for deployment across a bridge population should detect when the target domain differs materially from training, represent prediction uncertainty under that shift, identify when source knowledge should not be transferred, preserve physically plausible outputs, and allow engineers to interrogate the basis of warnings [10,23,28]. The requirement moves the problem from conventional model validation toward assurance of AI-supported structural diagnosis [67,90,91,125].
Human and engineering interpretation. Bridge assessment is a safety-relevant engineering process [8,10,11,60]. The strongest near-term role of AI is therefore likely to be prioritising, contextualising, and strengthening human engineering assessment [12,13,15,16]. AI-generated alerts become more useful when they identify the affected region, estimated condition change, confidence in that estimate, environmental context, physical evidence, and whether comparable behaviour has transferred successfully across other structures [9,18,22,74].
8.4. Answer to RQ5: Trustworthiness of AI-Enabled Bridge NDT
Trustworthy AI-enabled bridge NDT remains substantially less mature than the volume of predictive-performance literature suggests [23,24,26,76]. E is demonstrated by 20.4% of studies, whereas U, X, P, and F each occur in fewer than 6% [28,29,31,97]. Two-thirds of the corpus demonstrate none of these five attributes explicitly [34,35,36,37]. Future claims of operational trustworthiness should therefore be based on demonstrated combinations of robustness, uncertainty, explainability, physics, and fusion under realistic validation, not on architectural labels alone [8,9,10].
9. From Digital Monitoring to Prognosis and Engineering Decisions
This section addresses RQ6: how far AI-enabled monitoring translates into prognosis and actionable engineering decisions.
The ultimate value of AI-enabled bridge NDT is not determined solely by whether a model detects an anomaly or predicts a response variable [9,15,32,94]. Infrastructure management requires evidence that can support decisions concerning inspection, intervention, maintenance, prioritisation, and future risk [35,44,107,111]. The requirement creates a transition from structural monitoring to condition diagnosis, from diagnosis to prognosis, and finally from prognosis to engineering decision support [9,32,50,107].
9.1. Digital Twins as Monitoring and Inference Architectures
A bridge digital twin can be understood as a dynamically updated digital representation in which physical measurements, computational models, and analytical algorithms interact to maintain an evolving representation of the asset [8,15,35]. Digital-twin terminology is not itself a maturity category [9,15,32,94]. In the final coding, a digital-twin study remains M0 or M1 unless it demonstrates warning or intervention support and reaches M3 only when explicit future condition or life prediction is linked to maintenance or asset-management decision support [9,32,35,44].
Operational bridge digital twins can support live model calibration and updating without necessarily reaching condition prognosis [9,15,32,35]. AI-enabled digital-twin frameworks can also perform diagnosis, as illustrated by Mousavi et al. [15,59,68]. The architecture therefore does not determine maturity [9,15,32,94]. A digital twin can occupy D0/M0 through response prediction and model updating, D1-D3/M1-M2 through diagnosis and early warning, or D4/M3 when explicit deterioration prognosis and/or maintenance decision support are demonstrated [9,15,32,35].
9.2. Distinguishing Early Warning from Prognosis
An early-warning system identifies a developing abnormal condition earlier than an existing threshold or inspection process [50,112,116,118]. Prognosis predicts the future evolution of the structural condition [8,9,32,59]. Khan et al. provide strong operational M2 warning evidence because their physics-informed system identifies physically interpretable stiffness change associated with a confirmed fatigue crack [8]. This does not automatically constitute D4 because the model does not primarily forecast the future crack state at a specified later time [35,44,107,111].
Response forecasting versus deterioration prognosis. Response forecasting and reconstruction are common enabling functions, but they remain D0 unless future structural conditions or deterioration is predicted [50,112,116,118]. The distinction materially changes the evidence base: 244 studies are D0, while only 6 meet D4 [8,59,68,120]. Long forecast horizons, digital-twin synchronization, or accurate reconstruction of future strain and displacement should therefore not be described as prognosis unless the predicted target is structural condition, damage evolution, residual performance, or remaining life [9,15,32,94].
Mechanism-informed prognosis. Long-horizon prognosis benefits from physical deterioration mechanisms because future conditions can leave the historical data distribution [9,32,35,44]. Xie and Bai demonstrate a physics-informed fatigue-damage prediction framework for orthotropic bridge-deck structures in which crack evolution and displacement degradation are represented over fatigue life [9,32,50,107]. The target is clearly D4, but prognosis maturity must still be evaluated jointly with validation realism, damage authenticity, temporal exposure, uncertainty, and engineering-use maturity [8,9,32,59].
Digital twins and mechanism-specific warning. Sánchez-Haro et al. develop a digital twin of the Espartxo Bridge for early detection of under-foundation scour by linking monitored structural behaviour with calibrated numerical scour scenarios [43]. This provides mechanism-informed warning architecture, but numerical scour states calibrated within a continuously monitored real bridge should not be described as prospective observation of naturally evolving scour [35,44,107,111].
9.3. From Diagnostic Output to Engineering Decisions
Many AI studies end with a probability, novelty score, class label, or estimated condition parameter [50,112,116,118]. This is reflected in the final engineering-use distribution: M0 212 (29.2%), M1 440 (60.5%), M2 72 (9.9%), and M3 3 (0.4%) [8,9,32,59]. M2 is reserved for credible warning, inspection, emergency, or intervention-support outputs rather than a generic statement that a model could support maintenance [9,15,32,94].
M3 and maintenance decision support. Only 3 studies (0.4%) reach M3 [9,32,35,44]. A defensible M3 framework connects estimated structural condition and future evolution to an explicit maintenance, operation, or asset-management decision [9,32,50,107]. All three M3 studies are also D4, while the other three D4 studies remain M1 because they demonstrate prognosis without a sufficiently explicit decision-support layer [8,9,32,59]. Separating the two prevents predictive terminology from being mistaken for demonstrated engineering decision maturity [9,15,32,94].
Human-in-the-loop decision making. Near-term bridge management is more plausibly supported by human-in-the-loop AI than by autonomous maintenance decisions [35,44,107,111]. AI can prioritise structures for inspection, identify unusual response patterns, localise affected regions, estimate physically meaningful condition parameters, quantify confidence, compare behaviour with historical or population information, and support timing of additional inspection or intervention [50,112,116,118]. Engineers then integrate this evidence with visual inspection, conventional NDT, drawings, load history, environmental exposure, structural modelling, and consequence assessment [8,59,68,120].
9.4. Answer to RQ6: From Monitoring to Engineering Decisions
AI-enabled bridge NDT has progressed beyond algorithmic classification toward warning and selected prognosis, but temporal and decision maturity remain limited [9,15,32,94]. Only 16 studies (2.2%) reach T2 long-term monitoring, 6 (0.8%) reach D4 prognosis, and 3 (0.4%) reach M3 maintenance or asset-management decision support [9,11,32,35]. The evidence therefore supports a transition toward decision relevance, not yet a mature population of autonomous or prospectively validated maintenance systems [50,112,116,118].
10. Evidence Gaps and Research Roadmap
The evidence base shows rapid algorithmic expansion but a pronounced convergence gap [1,2,3,4]. Nearly half of the corpus remains R1, only 8 studies combine R4 operational validation with A4 confirmed physical damage, only 15 demonstrate G2/G3 transfer, only 16 reach T2 long-term monitoring, only 6 reach D4 prognosis, and only 3 reach M3 decision support. The research roadmap should therefore prioritise combinations of these dimensions rather than further isolated improvements in within-dataset predictive accuracy.
Table 4 summarises six evidence gaps [1,2,3,4] (Table 4). The prefix EG is used deliberately to avoid confusion with G0-G3 generalisation codes.
10.1. Evidence Gaps EG1-EG3: Realism, Generalisation, and Damage Authenticity
EG1, validation realism: R1 still accounts for 350 studies (48.1%), compared with 200 R4 studies (27.5%) [1,2,3,4]. Advanced D2 and D3 capability is demonstrated across all realism levels, but high diagnostic capability is not evidence of field robustness. Priority should be given to prospective operational validation, time-separated testing, locked models, and independent full-scale evaluation.
EG2 - cross-bridge generalisation. EG2, cross-bridge generalisation: only 15 studies (2.1%) reach G2/G3, and only 5 combine G2/G3 with R4 [1,2,3,4]. Zero-shot, unsupervised-target, and few-shot regimes are all represented, but actual held-out bridge transfer remains rare. Future studies should report target-supervision budgets, structural similarity, source-selection logic, negative transfer, and performance on unseen bridges.
EG3 - damage authenticity. EG3, damage authenticity: 200 studies use R4 operational bridge evidence, but only 8 combine R4 with A4 [1,2,3,4]. Forty-seven R4 studies are A3, meaning real bridge measurements are paired with synthetic, injected, or numerically transferred damage effects. Future work should document inspection or NDT confirmation, interventions, controlled full-scale changes, or prospectively observed deterioration so that field measurements are not mistaken for physical-damage evidence.
10.2. Evidence Gaps EG4-EG6: Trustworthiness, Prognosis, and Decision Integration
EG4, integrated trustworthiness: E is demonstrated by 148 studies, U by 38, X by 37, P by 42, and F by 28, while 487 studies demonstrate none of the five [1,2,3,4]. Only 7 studies demonstrate at least three attributes, and none demonstrate all five. The field therefore needs integrated robustness, uncertainty, explainability, physics, and fusion under authenticated field conditions and external transfer, including mechanisms for abstention or escalation when the target domain is outside validated support.
EG5 - temporal and prognostic maturity. EG5, temporal and prognostic maturity: only 16 studies reach T2, and only 6 reach D4 [1,2,3,4]. Long-term response forecasting remains distinct from future-condition prediction. Stronger evidence requires prospective deterioration forecasts, uncertainty intervals, later physical verification, and enough monitoring duration to evaluate drift in both the structure and the AI model.
EG6 - engineering decision integration. EG6, engineering decision integration: M0 and M1 together account for 652 studies (89.7%), whereas only 72 reach M2 and 3 reach M3 [1,2,3,4]. A technically accurate diagnosis has limited infrastructure value unless its consequences for inspection, warning, maintenance, or asset management are defined. Future studies should report decision thresholds, false-alarm and missed-detection costs, intervention logic, risk reduction, and lifecycle benefit.
10.3. Why These Gaps Persist
The quantified scarcities above are not primarily failures of algorithm design; they follow from structural constraints on how bridge evidence can be obtained. The rarity of confirmed full-scale physical damage (A4 in 8.4% of studies) and of its coincidence with operational monitoring (R4+A4 in 1.1%) reflects that deliberately damaging an in-service bridge is rarely permissible, and that naturally occurring damage on instrumented structures is infrequent and seldom independently verified; researchers therefore substitute numerical or laboratory damage, which is abundant but evidentially weaker. The dominance of same-structure work (G0-G1 in 98.0%) follows from the scarcity of openly shared, multi-bridge datasets with consistent instrumentation, which makes genuine cross-bridge transfer difficult to attempt and harder still to validate. The scarcity of long-term evidence (T2 in 2.2%) reflects the mismatch between multi-year deterioration timescales and the duration of typical research funding and monitoring campaigns. The concentration of reuse on a small set of named benchmark assets (Z24, KW51, Yonghe, and others) is a rational response to these constraints, but it means apparent validation breadth is partly recirculated evidence rather than independent confirmation. Finally, the shortfall in decision integration (M3 in 0.4%) and trustworthiness reporting (67.0% demonstrate none of the five attributes) reflect incentives that reward predictive-accuracy benchmarks over the uncertainty quantification, physical interpretability, and consequence modelling that operational adoption requires. Recognising these causes matters because they identify where the field can realistically intervene shared multi-bridge and physically damaged datasets, provenance and duration reporting, and evaluation protocols that credit evidence maturity rather than leaderboard accuracy would address the root constraints rather than the symptoms.
10.4. Three Research Horizons
Near-term work should standardise D/R/A/G/T/M reporting, damage provenance, monitoring duration, target-domain supervision, and E/U/X/P/F evidence so that studies become comparable [1,2,3,4]. Medium-term work should emphasise prospectively validated cross-bridge transfer under operational variability and confirmed damage. Long-term work should integrate prognosis, calibrated uncertainty, human oversight, intervention consequences, and asset-level decisions. These horizons follow the empirical bottlenecks quantified in Figure 6 rather than a purely algorithmic technology sequence (Figure 5).
10.5. Recommended Benchmark Philosophy
The field would benefit from benchmarks designed around evidence maturity rather than only leaderboard accuracy [1,2,3,4]. The named-asset analysis identified at least 92 publications (12.7% of the corpus) using one of 14 repeatedly named bridge or benchmark assets, including 39 publications explicitly using Z24, 11 Old ADA, 11 KW51, 10 Yonghe, and 6 I-40. The figures are publication counts, not independent validation events. Future benchmarks should therefore report unique structures, unique physical damage events, time-separated campaigns, transfer splits, and target-supervision requirements so that apparent validation breadth is not inflated by repeated reuse of the same assets.
11. Conclusions
AI-enabled vibration and dynamic-response bridge assessment has expanded rapidly, but the 727-study evidence base shows that algorithmic growth has outpaced evidential convergence [6]. The strongest conclusion of this investigation is therefore not that AI lacks diagnostic capability, but that realistic validation, authentic physical damage, transferability, long-term exposure, trustworthiness, prognosis, and engineering decision support rarely occur together.
The investigation evaluates diagnostic capability (D0-D4), validation realism (R1-R4), damage authenticity (A0-A4), generalisation (G0-G3), temporal maturity (T0-T2), engineering-use maturity (M0-M3), and the independent trustworthiness attributes E/U/X/P/F [6]. The multidimensional structure prevents a high classification score, field dataset, transfer-learning label, response forecast, or digital-twin architecture from being interpreted as evidence of maturity that has not actually been demonstrated.
First, diagnostic capability is broad but concentrated at lower maturity endpoints: D0 244 (33.6%), D1 246 (33.8%), D2 102 (14.0%), D3 129 (17.7%), and D4 only 6 (0.8%) [6]. Future response prediction should therefore remain distinct from future structural-condition or damage prognosis [6].
Second, validation of realism and damage authenticity are independent. R1 remains the largest realism category at 350 studies (48.1%), while R4 accounts for 200 (27.5%) [6]. A4 confirmed full-scale physical damage is demonstrated in 61 studies (8.4%), but only 8 (1.1%) combine A4 with R4 [6]. Operational measurements alone should not be described as real-damage validation [6].
Third, cross-bridge generalisation remains a major scalability bottleneck. G0 accounts for 644 studies (88.6%), G1 for 68 (9.4%), G2 for 12 (1.7%), and G3 for 3 (0.4%) [6]. Only 5 studies combine verified G2/G3 transfer with R4 operational evidence [6]. Multiple bridge applications are therefore not equivalent to demonstrated generalisation [6].
Fourth, trustworthy AI should not be represented by a single architectural label. E is explicitly demonstrated in 148 studies (20.4%), U in 38 (5.2%), X in 37 (5.1%), P in 42 (5.8%), and F in 28 (3.9%) [6]. Two-thirds of the corpus, 487 studies, demonstrate none of these five attributes explicitly, and no study demonstrates all five [6].
Fifth, temporal and decision maturity are limited. Only 16 studies (2.2%) reach T2 long-term monitoring, 6 (0.8%) reach D4 prognosis, and 3 (0.4%) reach M3 maintenance or asset-management decision support [6]. Long-term monitoring does not itself establish prognosis, and prognosis does not itself establish decision maturity [6].
Finally, the most important research transition is from isolated algorithm performance toward convergent evidence. Only one study in the corpus combines R4 operational validation, A4 confirmed physical damage, and G2/G3 transfer [6] (Figure 7). Future progress should therefore be judged by whether AI systems remain reliable under realistic variability, diagnose physically authenticated structural states, transfer to unseen bridges, quantify uncertainty, explain or physically constrain their reasoning, remain valid over time, and support defensible engineering actions.
AI-enabled bridge NDT is moving from algorithmic demonstration toward operational diagnosis, but the field is not yet defined by widespread transferable, long-term, trustworthy, prognostic, decision-ready systems. The evidence-centred framework developed here provides a basis for distinguishing genuine advances in engineering maturity from improvements that remain confined to architecture, dataset, or benchmark performance.
Supplementary Materials
The following supporting information can be downloaded at the website of this paper posted on Preprints.org. Table S1: Complete 727-study systematic-core inventory, containing the permanent DOI where available, final D/R/A/G/Z/T/M codes, E/U/X/P/F attributes, source-verification fields, and the study-level data used for every corpus-level count and percentage reported in this investigation. Table S2: First-pass-to-final coding-stability analysis, reporting per-dimension percent agreement and Cohen’s κ (quadratic-weighted for the ordinal dimensions, unweighted for the binary attributes) between the initial and reconciled maturity codes.
Author Contributions
Conceptualization, M.Z.B., M.S. and M.L.P.; methodology, M.Z.B.; formal analysis, M.Z.B.; investigation, M.Z.B.; data curation, M.Z.B.; writing-original draft preparation, M.Z.B.; writing-review and editing, M.S. and M.L.P.; visualisation, M.Z.B.; supervision, M.S. and M.L.P. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
The original contributions presented in this study are included in the article and Supplementary Table S1. All study-level evidence used for the quantitative synthesis is supplied with the article in Supplementary Table S1; no unreported study-level observations are required to reproduce the corpus distributions.
Conflicts of Interest
The authors declare no conflicts of interest.
Use of Generative AI
During the preparation of this manuscript, the authors used ChatGPT (OpenAI) for language and grammar editing, including improvements to sentence structure, clarity, and readability, and for conceptual and visual assistance in developing the graphical abstract. The graphical abstract was subsequently reviewed, revised, and verified by the authors to ensure consistency with the manuscript. All scientific content, data, numerical values, analyses, interpretations, and conclusions were reviewed and verified by the authors, who take full responsibility for the final content of the publication.
Abbreviations
The following abbreviations are used in this manuscript:
| Abbreviation | Definition |
| AI | Artificial intelligence |
| NDT | Nondestructive testing |
| SHM | Structural health monitoring |
| PRISMA | Preferred Reporting Items for Systematic Reviews and Meta-Analyses |
| D0-D4 | Diagnostic capability (enabling inference to prognosis) |
| R1-R4 | Validation realism (numerical to in-service) |
| A0-A4 | Damage authenticity (none to full-scale physical) |
| G0-G3 | Generalisation (same structure to population-level) |
| T0-T2 | Temporal maturity (short to long-term monitoring) |
| M0-M3 | Engineering-use maturity (algorithmic output to decision support) |
| E | Environmental and operational robustness (trustworthiness attribute) |
| U | Uncertainty quantification (trustworthiness attribute) |
| X | Explainability (trustworthiness attribute) |
| P | Physics integration (trustworthiness attribute) |
| F | Information fusion (trustworthiness attribute) |
| ZS | Zero-shot target-domain supervision |
| UT | Unsupervised target-domain information |
| FS | Few-shot labelled target data |
| ST | Supervised target retraining |
References
- Di Mucci, V.M.; Cardellicchio, A.; Ruggieri, S.; Nettis, A.; Renò, V.; Uva, G. Artificial intelligence in structural health management of existing bridges. Autom. Constr. 2024. [Google Scholar] [CrossRef]
- Mammeri, S.; Barros, B.; Conde-Carnero, B.; Riveiro, B. From traditional damage detection methods to Physics-Informed Machine Learning in bridges: A review. Eng. Struct. 2025. [Google Scholar] [CrossRef]
- Niyirora, R.; Ji, W.; Masengesho, E.; Munyaneza, J.; Niyonyungu, F.; Nyirandayisabye, R. Intelligent damage diagnosis in bridges using vibration-based monitoring approaches and machine learning: A systematic review. Results Eng. 2022. [Google Scholar] [CrossRef]
- Bao, Y.; Sun, H.; Xu, Y.; Guan, X.; Pan, Q.; Liu, D. Recent advances in structural health diagnosis: a machine learning perspective. Adv. Bridge Eng. 2025. [Google Scholar] [CrossRef]
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ 2021. [Google Scholar] [CrossRef] [PubMed]
- Bacha, M.Z.; Puppio, M.L.; Zucca, M.; Sassu, M. Dynamic Response-Based Bridge Monitoring and Structural Assessment: A Structured Scoping Review and Evidence Inventory. Infrastructures 2026. [Google Scholar] [CrossRef]
- Ferreira, L.; Yano, M.O.; Souza, L.; Moldovan, I.; da Silva, S.; Lopes, R.; Cimini, C.A.; Costa, J.C.W.A.; Figueiredo, E. Transfer learning and Bayesian calibration addressing data scarcity and uncertainty for structural health monitoring of twin concrete bridges. Mech. Syst. Signal Process. 2025. [Google Scholar] [CrossRef]
- Khan, A.Q.; Raza, A.; Pimanmas, A. Physics-informed neural networks for early-warning fatigue damage detection and localisation in a steel railway bridge using sparse ambient monitoring data. Intell. Transp. Infrastruct. 2026. [Google Scholar] [CrossRef]
- Xie, C.; Bai, Y. A physics-informed neural network for predicting structural fatigue damage of orthotropic bridge deck through updating model uncertainties. Int. J. Fatigue 2026. [Google Scholar] [CrossRef]
- Fan, L.; Guo, Z.; Shi, H. A Bayesian-Boosting Method for Probability Distribution Estimation of Structural Damping. Structural Control and Health Monitoring 2026. [Google Scholar] [CrossRef]
- Park, S.; Chang, M.; Yazdanpanah, O. A condition monitoring of steel box girder bridges based on consecutive system identification and model updating using a long-term monitoring database. Structures 2025. [Google Scholar] [CrossRef]
- Corbally, R.; Malekjafarian, A. A data-driven approach for drive-by damage detection in bridges considering the influence of temperature change. Eng. Struct. 2022. [Google Scholar] [CrossRef]
- Malekjafarian, A.; Golpayegani, F.; Moloney, C.; Clarke, S. A machine learning approach to bridge-damage detection using responses measured on a passing vehicle. In Sensors (Switzerland); 2019. [Google Scholar] [CrossRef] [PubMed]
- Duran, B.; Azam, Y.E.; Linzell, D.G. Bridge health monitoring via drive-by sensing: A novel CNN-based framework with full-scale field validation. Mech. Syst. Signal Process. 2026. [Google Scholar] [CrossRef]
- Mousavi, V.; Rashidi, M.; Ghazimoghadam, S.; Samali, B. A data-driven digital twin model for bridge health monitoring using feature fusion and unsupervised deep learning. Inf. Fusion 2026. [Google Scholar] [CrossRef]
- Wang, Z.; Zhan, J.; Ni, Z.; Xu, X.; Wu, L.; Wang, C. A physics-enhanced deep learning framework for baseline-free damage localisation in railway continuous beam bridge using displacement influence line. Adv. Eng. Inform. 2026. [Google Scholar] [CrossRef]
- Yang, Y.; Feng, D.; Xu, Y. A TCN-PatchTST framework for efficient reconstruction of missing acceleration data in bridge structural health monitoring. Eng. Struct. 2026. [Google Scholar] [CrossRef]
- Tian, H.; Zhu, Q.; Zhang, Q.; Wang, X.; Du, Y. A time-frequency heterogeneous dual-domain joint sampling-reconstruction and uncertainty quantification network for structural monitoring data reconstruction. Struct. Health Monit. 2026. [Google Scholar] [CrossRef]
- Yang, X.-M.; Zhang, Z.-K.; Chen, T.; Wang, Z.; Yao, X.-J. A transformer-based framework for automating frequency domain decomposition in bridge modal analysis. Structures 2026. [Google Scholar] [CrossRef]
- Yan, F.; Geng, Y.; Liu, Y.; Chen, N.; Li, X.; Wang, Y.; Aloisio, A. A vortex-induced vibration warning method based on ensemble-learning-embedded neural network. Comput.-Aided Civ. Infrastruct. Eng. 2025. [Google Scholar] [CrossRef]
- Hassani, S.; Dackermann, U.; Mousavi, M.; Mustapha, S.; Li, J. AAE-CycleWGAN fusion framework for generating fused strain data from sparse to dense domains in bridge monitoring systems. Inf. Fusion 2026. [Google Scholar] [CrossRef]
- Bragança, C.S.C.D.; Souza, E.; Ames, I.; Ribeiro, D.; Bittencourt, T.N. AI-based drive-by methodology for early damage detection in a steel truss railway bridge under multiple environmental and operational factors. Struct. Infrastruct. Eng. 2026. [Google Scholar] [CrossRef]
- Jiang, W.-J.; Kim, C.-W. Ambient vibration-based cable tension monitoring and uncertainty analysis in cable-stayed bridge from a fully probabilistic perspective. Eng. Struct. 2025. [Google Scholar] [CrossRef]
- Wu, S.-H.; Zhu, Y.-C.; Pan, Z.-H.; Di, H. An efficient Expectation-Maximization algorithm for Bayesian operational modal analysis with physics-data fusion model. Mech. Syst. Signal Process. 2025. [Google Scholar] [CrossRef]
- Wan, H.-P.; Fang, T.-L.; Zhu, Y.-K.; Wang, C.; Wang, N.-B. An improved SIFT-based method for non-contact bridge displacement measurement. Adv. Eng. Inform. 2026. [Google Scholar] [CrossRef]
- Ye, X.; Chen, X.; Lei, Y.; Fan, J.; Mei, L. An integrated machine learning algorithm for separating the long-term deflection data of prestressed concrete bridges. Sensors (Switzerland) 2018. [Google Scholar] [CrossRef] [PubMed]
- Giglioni, V.; Eva, A.E.; Worden, K.; Ubertini, F.; Venanzi, I. Assessing similarity requirements for effective transfer learning across a network of rigid frame bridges. Reliab. Eng. Syst. Saf. 2026. [Google Scholar] [CrossRef]
- Sadeghi Eshkevari, S.; Marasco, G.; Sen, D.; Dabbaghchian, I.; Pakzad, S.N. Assessing uncertainty in CNN-attention-based indirect bridge strain estimation from field acceleration data. Struct. Infrastruct. Eng. 2026. [Google Scholar] [CrossRef]
- Pan, C.; Gong, F.; Liu, X.; Yan, B.; Xia, Y. Axle-information-free identification of vehicle moving forces from bridge displacement using U-Net. Adv. Struct. Eng. 2026. [Google Scholar] [CrossRef]
- Nguyen, T.Q.; Nguyen, D.P.; Nguyen, P.T.; Nguyen, T.T. Bayesian Deep Learning for Bridge Structural Damage Assessment Based on Dissipation Coefficient Evaluation. Int. J. Steel Struct. 2026. [Google Scholar] [CrossRef]
- Mu, H.-Q.; Zheng, Z.-J.; Wu, X.-H.; Su, C. Bayesian network-based modal frequency-multiple environmental factors pattern recognition for the Xinguang Bridge using long-term monitoring data. J. Low. Freq. Noise Vib. Act. Control 2020. [Google Scholar] [CrossRef]
- Cao, K.; Ding, Y.; Li, H.; Yang, T.-L.; Liu, G.-G.; Chu, T.-Y. Bridge cables vibration frequency identification and fatigue life prediction based on SHM data with machine learning. Steel Compos. Struct. 2025. [Google Scholar] [CrossRef]
- Huang, S.; Zeng, W.; Deng, X.; Yang, Y.B. Bridge Damage Detection in Vehicle-Bridge Interaction System Under Road Surface Roughness: Convolutional Neural Networks with Baseline-Referenced Residual Transformation. Int. J. Struct. Stab. Dyn. 2026. [Google Scholar] [CrossRef]
- Qu, G.; Song, M.; Xia, Y.; Sun, L. Bridge Girder-End Displacement Reconstruction Using a Novel Hybrid Attention Mechanism Leveraging Multisource Information. Struct. Control Health Monit. 2025. [Google Scholar] [CrossRef]
- Ma, S.; Flanigan, K.A.; Bergés, M. Bridging the Reality Gap in Digital Twins with Context-Aware: Physics-Guided Deep Learning. J. Comput. Civ. Eng. 2026. [Google Scholar] [CrossRef]
- Wu, Y.; Deng, L.; He, W. Bwimnet: A novel method for identifying moving vehicles utilizing a modified encoder-decoder architecture. Sensors (Switzerland) 2020. [Google Scholar] [CrossRef] [PubMed]
- Wu, D.; Li, X.; Xu, N.; Li, Y.; Jin, Y.; He, S.; Xu, W.; Li, N.; chen, W.; Laima, S. Cross-attention based frequency-domain prediction method for buffeting response of long-span bridges. Adv. Bridge Eng. 2026. [Google Scholar] [CrossRef]
- Xiong, J.; Hu, L.; Meng, X.; An, X.; Xie, Y. Cross-Modal Graph Attention for Bridge SHM Data Imputation. Sensors 2026. [Google Scholar] [CrossRef] [PubMed]
- Avcı, M.S.; Ercan, E.; Hızal, Ç.; Malekjafarian, A.; Ozer, E. Cross-structure domain translation with structural state awareness: AT-StarGAN-GP for multi-level synthetic signal generation in structural health monitoring. Struct. Health Monit. 2026. [Google Scholar] [CrossRef]
- Nguyen, D.H.; Nguyen, Q.B.; Bui-Tien, T.; De Roeck, G.; Abdel Wahab, M. Damage detection in girder bridges using modal curvatures gapped smoothing method and Convolutional Neural Network: Application to Bo Nghi bridge. Theor. Appl. Fract. Mech. 2020. [Google Scholar] [CrossRef]
- Zhou, B.; Zhu, J.; Jiao, X.; Yessoufou, F. Damage localisation based on fusion of multi-source vehicle-induced response of bridges. Struct. Infrastruct. Eng. 2026. [Google Scholar] [CrossRef]
- Guo, J.; Zhang, G. Deep learning based identification of vortex-induced vibration in stay cables using multidimensional feature fusion and BiLSTM-MHA. J. Wind Eng. Ind. Aerodyn. 2026. [Google Scholar] [CrossRef]
- Sánchez-Haro, J.; García, M.; Capellán, G.; da Costa, A.; Perez, P.; Añó, J. Digital twin for predictive maintenance on the Espartxo Bridge: Application to early detection of under-foundation scour. Structures 2025. [Google Scholar] [CrossRef]
- Ying, L.Q.; Ying, G.G.; Hu, J.L.; Zhang, W.D. Deep Learning Method for Vehicle Load Identification Using Digital Twin and Transfer Learning Strategies. J. Bridge Eng. 2026. [Google Scholar] [CrossRef]
- Al-Adly, A.I.F.; Kripakaran, P. Developing physics-informed neural networks for virtual sensing in beams with moving loads. Eng. Struct. 2025. [Google Scholar] [CrossRef]
- Calderon Hurtado, A.; Xu, J.; Salleh, R.; Dias-da-Costa, D.; Alamdari, M.M. Development and field validation of a fully customised vehicle scanning system on two full-scale bridges. Struct. Health Monit. 2026. [Google Scholar] [CrossRef]
- Wan, C.; Hou, J.; Zhang, G.; Gao, S.; Ding, Y.; Cao, S.; Hu, H.; Xue, S. Domain adaptation based automatic identification method of vortex induced vibration of long-span bridges without prior information. Eng. Appl. Artif. Intell. 2025. [Google Scholar] [CrossRef]
- Fernandes, T.M.; Minski, L.; De Souza, P.V.G.; Ribeiro, D.R.F.; Miguel, L.F.F.; Lopez, R.H. Early Scour Damage Detection Using Drive-By Monitoring Data through Supervised Learning. J. Struct. Des. Constr. Pract. 2026. [Google Scholar] [CrossRef]
- Nayek, R.; Narasimhan, S. Extraction of contact-point response in indirect bridge health monitoring using an input estimation approach. J. Civ. Struct. Health Monit. 2020. [Google Scholar] [CrossRef]
- Al-Hijazeen, A.; Koris, K. Feedforward Neural Network-Based Digital Twin for SHM of Bridges. Archit. Civ. Eng. Environ. 2025. [Google Scholar] [CrossRef]
- Yang, Y.-L.; Zhu, Y.-C.; Cai, C.S.; Wu, S.-H. Fusion of Bayesian Time Domain and Gaussian Process Models for Modal Identification Under Environmental Variations. Struct. Control Health Monit. 2026. [Google Scholar] [CrossRef]
- Mao, J.-X.; Wang, H.; Spencer, B.F. Gaussian mixture model for automated tracking of modal parameters of long-span bridge. Smart Struct. Syst. 2019. [Google Scholar] [CrossRef]
- Zeng, J.; Cao, Z.-J.; Wan, Q.; Chen, R.; Chen, H. Generative diffusion-aided probabilistic spatiotemporal response reconstruction for structural dynamic system under sparse noisy measurements. Comput.-Aided Civ. Infrastruct. Eng. 2026. [Google Scholar] [CrossRef]
- Liu, J.; Xu, S.; Bergés, M.; Noh, H.Y. HierMUD: Hierarchical multi-task unsupervised domain adaptation between bridges for drive-by damage diagnosis. Structural Health Monitoring 2023. [Google Scholar] [CrossRef]
- Chen, C.; Huang, X.; Liu, Z.; Yang, B.; Zhou, L.; Jiang, Z.; Liu, Y.; Tang, L. Long-term structural health monitoring strain modelling for bridges via data correlation and deep learning. Eng. Res. Express 2025. [Google Scholar] [CrossRef]
- Seon Park, H.; Hong, T.; Lee, D.-E.; Kwan Oh, B.; Glisic, B. Long-term structural response prediction models for concrete structures using weather data, fiber-optic sensing, and convolutional neural network. Expert Syst. With Appl. 2022. [Google Scholar] [CrossRef]
- Zhang, M.; Guo, T.; Zhang, G.; Liu, Z.; Liu, Y. Missing monitoring data reconstruction for cable-stayed bridge using knowledge transfer-based generative pre-trained model. Struct. Health Monit. 2025. [Google Scholar] [CrossRef]
- Yu, H.; Shu, J.; Dong, T.; Yang, H.; Chen, Q.; Wang, J.; Guo, S. Missing structural response recovery under different damage scenarios via a multi-modal diffusion model fusing inspection text and monitoring data. Adv. Eng. Inform. 2026. [Google Scholar] [CrossRef]
- Jasiński, M.; Fawad, M.; Sabzi Khoshraftar, A.; Abbozzo, A.; Tao, Y.; Salamak, M.; Kopeć, B.; Chen, Q. Model-based predictive digital twin for bridge structural health monitoring: the integration of building information modelling, finite element analysis, and machine learning in a Netherlands case study. Eng. Appl. Artif. Intell. 2026. [Google Scholar] [CrossRef]
- Haghbin, Masoud; Tomassini, Elisa; Ubertini, Filippo; García-Macías, Enrique; Chiachío-Ruano, Juan. Explainable AI for Operational Modal Analysis: Field deployment on densely instrumented structures. Eng. Struct. 2026. [Google Scholar] [CrossRef]
- Yang, H.; Yan, W.; He, H. Parameters identification of moving load using ANN and dynamic strain. Shock and Vibration 2016. [Google Scholar] [CrossRef]
- Cianci, E.; Civera, M.; De Biagi, V.; Chiaia, B. Physics-informed machine learning for the structural health monitoring and early warning of a long highway viaduct with displacement transducers. Mech. Syst. Signal Process. 2026. [Google Scholar] [CrossRef]
- Chen, R.; Kim, C.-W.; Wang, J. Physics-informed neural operator for forecasting vehicle-induced bridge vibration trajectories. Eng. Appl. Artif. Intell. 2026. [Google Scholar] [CrossRef]
- Pei, X.-Y.; Zhang, H.-T.; Huang, H.-B.; Liang, D. Probabilistic Machine Learning-Based Frequency Normalization Method for Bridge Damage Detection Considering Environmental Variations. Int. J. Struct. Stab. Dyn. 2026. [Google Scholar] [CrossRef]
- Das, T.; Guchhait, S. Sequential optimisation architecture for bridge damage localisation and severity estimation under dynamic moving loads using an automated machine learning-driven hybrid deep learning model. Eng. Appl. Artif. Intell. 2026. [Google Scholar] [CrossRef]
- Longji, Z.; Zhi, Y.; Jiaqing, L.; Wenhua, L.; Jingchun, M. Spatiotemporal dependency data imputation for long-term health monitoring of concrete arch bridges. Sci. Rep. 2025. [Google Scholar] [CrossRef] [PubMed]
- Li, Z.; Yan, B.; Meng, Q.; Xu, C.; Zhang, F.; Wang, Y.; Domingo, M. Strain prediction in a large-span arch bridge using the TimeXer model considering temperature and traffic loads. Front. Built Environ. 2026. [Google Scholar] [CrossRef]
- Kaewnuratchadasorn, C.; Wang, J.; Kim, C.-W.; Yang, Y. Two-Dimensional Vehicle-Bridge Interaction Neural Operator for Digital Twin of Bridge Structures. Struct. Control Health Monit. 2025. [Google Scholar] [CrossRef]
- Hou, J.; Cao, S.; Hu, H.; Zhou, Z.; Wan, C.; Noori, M.; Li, P.; Luo, Y. Vortex-Induced Vibration Recognition for Long-Span Bridges Based on Transfer Component Analysis. Buildings 2023. [Google Scholar] [CrossRef]
- Yano, M.O.; Figueiredo, E.; da Silva, S.; Cury, A. Foundations and applicability of transfer learning for structural health monitoring of bridges. Mech. Syst. Signal Process. 2023. [Google Scholar] [CrossRef]
- Hadizadeh, A.; Tarighat, A.; Malian, A. A data-driven framework for structural health monitoring using reinforcement learning and deep autoencoders. Sci. Rep. 2026. [Google Scholar] [CrossRef] [PubMed]
- Giglioni, V.; Poole, J.; Venanzi, I.; Ubertini, F.; Worden, K. A domain adaptation approach to damage classification with an application to bridge monitoring. Mech. Syst. Signal Process. 2024. [Google Scholar] [CrossRef]
- Shi, S.; Du, D.; Mercan, O.; Kalkan, E.; Parol, J. A novel decentralized damage detection method for self-powered wireless sensing in structural health monitoring using self-supervised learning. Eng. Struct. 2025. [Google Scholar] [CrossRef]
- Deng, P.; Yang, J.J.; Yee, T.; Oguzmert, M. A Physics-Guided Feature Fusion Network for Vibration-Based Bridge Scour Monitoring. Transp. Res. Rec. 2026. [Google Scholar] [CrossRef]
- Lee, K.; Hwang, J.; Shin, D.H.; Lee, J.-H. A time-interval-based incremental learning paradigm for progressive bridge damage detection using convolutional autoencoders. Struct. Infrastruct. Eng. 2026. [Google Scholar] [CrossRef]
- Giorgi, V.; Tordela, C.; Bernardini, L.; Ramírez Balbiano, P.A.; Somaschini, C.; Strano, S.; Terzo, M. An LSTM Autoencoder-Based Approach for Monitoring Railway Bridges. Applied Sciences (Switzerland) 2026. [Google Scholar] [CrossRef]
- Lu, N.; Xiao, X.; Cui, J.; Liu, Y.; Huang, K.; Yuen, K.-V. An unsupervised cross-domain method for bridge damage detection based on multichannel symmetric dot pattern feature alignment. Comput.-Aided Civ. Infrastruct. Eng. 2025. [Google Scholar] [CrossRef]
- Bayane, I.; Leander, J.; Karoumi, R. An unsupervised machine learning approach for real-time damage detection in bridges. Eng. Struct. 2024. [Google Scholar] [CrossRef]
- Giglioni, V.; Venanzi, I.; Poggioni, V.; Milani, A.; Ubertini, F. Autoencoders for unsupervised real-time bridge health assessment. Computer-Aided Civil and Infrastructure Engineering 2023. [Google Scholar] [CrossRef]
- Shi, S.; Du, D.; Mercan, O.; Kalkan, E.; Parol, J. Contrastive and self-supervised learning for open-set damage classification in structural health monitoring with incomplete and imbalanced vibration data. Expert Syst. With Appl. 2025. [Google Scholar] [CrossRef]
- Ahmad, H.; Matsumoto, Y. Damage detection in a single-span prestressed concrete girder bridge under environmental variations using Gaussian process regression validated by physics-guided surrogate modelling. J. Civ. Struct. Health Monit. 2026. [Google Scholar] [CrossRef]
- Lu, N.; Cui, J.; Zeng, W.; Xiao, X.; Luo, Y. Damage identification of a reduced-scale cable-stayed bridge based on domain adaptation transfer learning. Meas. J. Int. Meas. Confed. 2026. [Google Scholar] [CrossRef]
- Riyahi, A.; Mestari, M.; Bouihi, B. Data-Driven and Physics-Informed Neural Networks for Structural Health Monitoring of the Z24 Bridge. J. Civ. Eng. Forum 2026. [Google Scholar] [CrossRef]
- Hu, Z.; He, W.; Li, H.; Wu, Y. Few-Shot Cross-Bridge Damage Diagnosis from Vibration Sensor Signals via Siamese Contrastive Pretraining with Self-Calibrated Convolution. Sensors 2026. [Google Scholar] [CrossRef] [PubMed]
- Park, S.; Kim, S. Hybrid CAE-DSVDD for unsupervised vibration-based damage detection in in situ steel truss bridge. Struct. Health Monit. 2025. [Google Scholar] [CrossRef]
- Abu Zouriq, M.F.; Linzell, D.G.; Azam, Y.E. Integrating Bi-Bi-LSTM-Based Virtual Sensing and VAE for Sensor-Efficient Unsupervised Structural Damage Detection. J. Eng. Mech. 2026. [Google Scholar] [CrossRef]
- Duran, B.; Eftekhar Azam, S.; Sanayei, M. Leveraging Deep Learning for Robust Structural Damage Detection and Classification: A Transfer Learning Approach via CNN. 2024. [Google Scholar] [CrossRef]
- Liu, J.; Zhang, W.; Sun, L.; Li, Y. Physics-encoded unsupervised deep learning for lightweight structural health monitoring of short and medium-span bridge networks. Eng. Appl. Artif. Intell. 2026. [Google Scholar] [CrossRef]
- Sarwar, M.Z.; Cantero, D. Probabilistic autoencoder-based bridge damage assessment using train-induced responses. Mech. Syst. Signal Process. 2024. [Google Scholar] [CrossRef]
- Nesackon Abraham, J.; Tran, M.Q.; Jayaraj, J.S.; Matos, J.C.; Valluzzi, M.R.; Dang, S.N. Unsupervised Learning-Based Anomaly Detection for Bridge Structural Health Monitoring: Identifying Deviations from Normal Structural Behaviour. Sensors 2026. [Google Scholar] [CrossRef] [PubMed]
- Ge, L.; Wang, C.; Liu, M.; Dan, D.; Jian, F. VIV-SDE-Net: A physics-informed neural framework for long-horizon prediction of bridge vortex-induced vibrations. Adv. Struct. Eng. 2026. [Google Scholar] [CrossRef]
- Giglioni, V.; Poole, J.; Mills, R.; Venanzi, I.; Ubertini, F.; Worden, K. Transfer learning in bridge monitoring: Laboratory study on domain adaptation for population-based SHM of multispan continuous girder bridges. Mech. Syst. Signal Process. 2025. [Google Scholar] [CrossRef]
- Heravi, M.A.; Soleimani-Babakamali, M.H.; Naderpour, H.; Sadhu, A. Zero-shot transfer learning for structural damage detection using target-to-source structure domain data mapping. Mech. Syst. Signal Process. 2026. [Google Scholar] [CrossRef]
- Impraimakis, M.; Palkanoglou, E.N. A generative adversarial network optimisation method for damage detection and digital twinning by deep AI fault learning: Z24 Bridge structural health monitoring benchmark validation. Structural and Multidisciplinary Optimisation 2025. [Google Scholar] [CrossRef]
- Nnamani, N.E.; Matos, J.C.; Komarizadehasl, S.; Nguyen, N.T.T.; Dang, S.N. A Hybrid CNN-LSTM Framework for Vibration-Based Multi-Damage Assessment in Reinforced Concrete Bridges. Applied Sciences (Switzerland) 2026. [Google Scholar] [CrossRef]
- Wu, W.-H.; Chen, C.-C.; Chen, Z.-T.; Lai, G. Automated real-time cable tension warning framework using long-term monitoring data and a convolutional neural network. Struct. Health Monit. 2026. [Google Scholar] [CrossRef]
- Asadollahi, P.; Huang, Y.; Li, J. Bayesian finite element model updating and assessment of cable-stayed bridges using wireless sensor data. Sensors (Switzerland) 2018. [Google Scholar] [CrossRef] [PubMed]
- Weinstein, J.C.; Sanayei, M.; Brenner, B.R. Bridge Damage Identification Using Artificial Neural Networks. J. Bridge Eng. 2018. [Google Scholar] [CrossRef]
- Chun, P.-J.; Yamashita, H.; Furukawa, S. Bridge Damage Severity Quantification Using Multipoint Acceleration Measurement and Artificial Neural Networks. Shock and Vibration 2015. [Google Scholar] [CrossRef]
- Gonzalez, I.; Karoumi, R. BWIM aided damage detection in bridges using machine learning. J. Civ. Struct. Health Monit. 2015. [Google Scholar] [CrossRef]
- Dackermann, U.; Smith, W.A.; Alamdari, M.M.; Li, J.; Randall, R.B. Cepstrum-based damage identification in structures with progressive damage. Struct. Health Monit. 2019. [Google Scholar] [CrossRef]
- Kaspar, K.; Santini-Bell, E.; Petrik, M.; Sanayei, M. Comparison between a Linear Regression and an Artificial Neural Network Model to Detect and Localize Damage in the Powder Mill Bridge. Transp. Res. Rec. 2020. [Google Scholar] [CrossRef]
- Tan, Z.X.; Thambiratnam, D.P.; Chan, T.H.T.; Gordan, M.; Abdul Razak, H. Damage detection in steel-concrete composite bridge using vibration characteristics and artificial neural network. Struct. Infrastruct. Eng. 2020. [Google Scholar] [CrossRef]
- Jin, C.; Jang, S.; Sun, X.; Li, J.; Christenson, R. Damage detection of a highway bridge under severe temperature changes using extended Kalman filter trained neural network. J. Civ. Struct. Health Monit. 2016. [Google Scholar] [CrossRef]
- Xiang, C.; Zhao, H.; Wu, G.; Chen, L.; Yang, Z.; Patelli, E. Damage diagnosis of bridge structures using deep learning strategies: A hybrid neural networks practical tool. Struct. Health Monit. 2026. [Google Scholar] [CrossRef]
- Qiu, Y.; Ahmed, B.; Abueidda, D.W.; El-Sekelly, W.; de Soto, B.G.; Abdoun, T.; Ji, H.; Qiu, J.; Mobasher, M.E. Damage identification for bridges using machine learning: Development and application to KW51 bridge. Digit. Eng. 2026. [Google Scholar] [CrossRef]
- Hakimi, O.; Liu, H.; Abudayyeh, O. Deep learning-driven multi-level data fusion framework for predictive maintenance and structural health monitoring of concrete bridge decks. Autom. Constr. 2025. [Google Scholar] [CrossRef]
- Fallahian, M.; Khoshnoudian, F.; Meruane, V. Ensemble classification method for structural damage assessment under varying temperature. Struct. Health Monit. 2018. [Google Scholar] [CrossRef]
- Ahmad, H.; Matsumoto, Y. Environmental effect compensation and anomaly detection in an ageing prestressed concrete girder bridge using Gaussian process regression. Struct. Infrastruct. Eng. 2026. [Google Scholar] [CrossRef]
- Deng, P.; Yang, J.J.; Yee, T.; Oguzmert, M. Regime-Aware Bridge Scour Depth Estimation Using Mixture of Experts and Accelerometer Data. Transp. Res. Rec. 2026. [Google Scholar] [CrossRef]
- Deng, P.-H.; Cui, C.; Cheng, Z.-Y.; Zhang, Q.-H.; Bu, Y.-Z. Fatigue damage prognosis of orthotropic steel deck based on data-driven LSTM. J. Constr. Steel Res. 2023. [Google Scholar] [CrossRef]
- Lei, Z.; Zhu, L.; Fang, Y.; Niu, C.; Zhao, Y. Fiber Bragg Grating Smart Material and Structural Health Monitoring System Based on Digital Twin Drive. J. Nanomater. 2022. [Google Scholar] [CrossRef]
- Hom, K.L.; Beigi, H.; Betti, R. From laboratory experiments to in-service tests: Learning embeddings for damage identification in structural health monitoring. Mech. Syst. Signal Process. 2026. [Google Scholar] [CrossRef]
- Yang, J.; Liu, D.; Zhao, L.; Yang, X.; Li, R.; Jiang, S.; Li, J. Improved stochastic configuration network for bridge damage and anomaly detection using long-term monitoring data. Inf. Sci. 2025. [Google Scholar] [CrossRef]
- Colacillo, P.; Civera, M.; Surace, C.; Tronci, E.M. Integrating climate projections with Gaussian processes for structural anomaly detection via cointegration. J. Low. Freq. Noise Vib. Act. Control 2026. [Google Scholar] [CrossRef]
- Armijo, A.; Zamora-Sánchez, D. Integration of Railway Bridge Structural Health Monitoring into the Internet of Things with a Digital Twin: A Case Study. Sensors 2024. [Google Scholar] [CrossRef] [PubMed]
- Tyler, G.; Hurtado, A.C.; Hamedani, S.J.; Salleh, R.; Alamdari, M.M. Isolation distributional kernels for indirect bridge health monitoring: A full-scale investigation of anomaly detection. Comput.-Aided Civ. Infrastruct. Eng. 2026. [Google Scholar] [CrossRef]
- Li, S.; Wang, W.; Lu, B.; Du, X.; Dong, M.; Zhang, T.; Bai, Z. Long-term structural health monitoring for bridge based on back propagation neural network and long and short-term memory. Struct. Health Monit. 2023. [Google Scholar] [CrossRef]
- Teng, S.; Chen, G.; Liu, Z.; Cheng, L.; Sun, X. Multi-sensor and decision-level fusion-based structural damage detection using a one-dimensional convolutional neural network. Sensors 2021. [Google Scholar] [CrossRef] [PubMed]
- Lu, Q.; Zhu, J.; Zhang, W. Quantification of Fatigue Damage for Structural Details in Slender Coastal Bridges Using Machine Learning-Based Methods. J. Bridge Eng. 2020. [Google Scholar] [CrossRef]
- Ni, Y.-Q.; Wang, J.; Chan, T.H.T. Structural damage alarming and localisation of cable-supported bridges using multi-novelty indices: A feasibility study. Struct. Eng. Mech. 2015. [Google Scholar] [CrossRef]
- Arangio, S.; Bontempi, F. Structural health monitoring of a cable-stayed bridge with Bayesian neural networks. Struct. Infrastruct. Eng. 2015. [Google Scholar] [CrossRef]
- Khodabandehlou, H.; Pekcan, G.; Fadali, M.S. Vibration-based structural condition assessment using convolution neural networks. Struct. Control Health Monit. 2019. [Google Scholar] [CrossRef]
- Mehrjoo, A.; Hom, K.L.; Beigi, H.; Betti, R. Zero-Shot Bridge Health Monitoring Using Cepstral Features and Streaming LSTM Networks. Infrastructures 2025. [Google Scholar] [CrossRef]
- Gong, F.; Xia, Y.; Ling, Z.; Lozano, F.; He, T. Bayesian deep learning based bridge condition assessment considering uncertainty quantification of missing data. Eng. Struct. 2026. [Google Scholar] [CrossRef]
Figure 3.
Diagnostic capability versus validation realism for the 727-study systematic core. Cell values are publication counts. The matrix demonstrates that higher diagnostic capability does not imply stronger validation realism; even D3 remains concentrated in R1 numerical/synthetic evidence.
Figure 3.
Diagnostic capability versus validation realism for the 727-study systematic core. Cell values are publication counts. The matrix demonstrates that higher diagnostic capability does not imply stronger validation realism; even D3 remains concentrated in R1 numerical/synthetic evidence.

Figure 5.
Co-occurrence of the five trustworthiness attributes (E, environmental/operational robustness; U, uncertainty quantification; X, explainability; P, physics integration; F, information fusion) across the 727-study systematic core. Bar heights give the number of studies in each exclusive attribute intersection; the dot matrix indicates which attributes define each bar; left-hand bars give the total number of studies demonstrating each attribute. Environmental robustness dominates in isolation (116 studies), higher-order combinations are rare, and a single study (highlighted) demonstrates four attributes. 487 of 727 studies (67.0%) demonstrate none of the five, and no study demonstrates all five.
Figure 5.
Co-occurrence of the five trustworthiness attributes (E, environmental/operational robustness; U, uncertainty quantification; X, explainability; P, physics integration; F, information fusion) across the 727-study systematic core. Bar heights give the number of studies in each exclusive attribute intersection; the dot matrix indicates which attributes define each bar; left-hand bars give the total number of studies demonstrating each attribute. Environmental robustness dominates in isolation (116 studies), higher-order combinations are rare, and a single study (highlighted) demonstrates four attributes. 487 of 727 studies (67.0%) demonstrate none of the five, and no study demonstrates all five.

Figure 6.
Evidence-maturity bottlenecks in the systematic core. Operational validation is substantially more common than confirmed physical damage, long-term monitoring, cross-bridge transfer, prognosis, or maintenance/asset-management decision support. The figure therefore represents the evidence-development pathway that must be strengthened for operational AI-enabled bridge NDT.
Figure 6.
Evidence-maturity bottlenecks in the systematic core. Operational validation is substantially more common than confirmed physical damage, long-term monitoring, cross-bridge transfer, prognosis, or maintenance/asset-management decision support. The figure therefore represents the evidence-development pathway that must be strengthened for operational AI-enabled bridge NDT.

Figure 7.
Evidence-convergence funnel for the 727-study systematic core. Successive filtering by validation realism, damage authenticity, and generalisation reduces the corpus from 727 studies to a single study that simultaneously combines R4 operational monitoring, A4 confirmed full-scale physical damage, and G2/G3 verified cross-bridge transfer. The funnel visualises the central finding that convergence of high-maturity evidence, rather than diagnostic capability itself, is the principal bottleneck.
Figure 7.
Evidence-convergence funnel for the 727-study systematic core. Successive filtering by validation realism, damage authenticity, and generalisation reduces the corpus from 727 studies to a single study that simultaneously combines R4 operational monitoring, A4 confirmed full-scale physical damage, and G2/G3 verified cross-bridge transfer. The funnel visualises the central finding that convergence of high-maturity evidence, rather than diagnostic capability itself, is the principal bottleneck.

Table 1.
Investigation design, search strategy, eligibility criteria, and evidence-extraction architecture. [5].
Table 1.
Investigation design, search strategy, eligibility criteria, and evidence-extraction architecture. [5].
| Investigation component | Protocol adopted |
|---|---|
| Investigation type | Structured critical investigation supported by systematic literature search and multidimensional evidence assessment [5] |
| Systematic period | 1 January 2015 to final search timestamp on 23 August 2026 [5] |
| Databases | Scopus; Web of Science Core Collection; IEEE Xplore [5] |
| Language | English-language peer-reviewed journal articles [5] |
| Core concept | bridge × dynamic response × AI × structural assessment [5] |
| Supplementary thematic searches | transfer/domain adaptation; self-supervised/open-set; physics-informed AI; UQ/XAI; fusion; digital twins; prognosis [5] |
| Screening | identification; deduplication; title/abstract screening; full-text eligibility; evidence coding [5] |
| Primary coding | D0-D4; R1-R4; A0-A4; G0-G3; T0-T2; M0-M3 [5] |
| Trustworthiness | E, U, X, P, F [5] |
| Counts | Final: 3,412 raw records; 1,679 screened; 839 full-text eligibility records; 727 systematic-core studies. [5] |
| Search yield | Scopus 1,969; Web of Science 1,295; IEEE Xplore 148; total raw records 3,412 [5] |
| Deduplication and recovery | 1,743 duplicates removed; 1,669 database-unique records; 10 eligible studies recovered through supplementary targeted searches; final identification set 1,679 [5] |
| Title/abstract screening | 840 excluded; 839 advanced to source-level full-text eligibility [5] |
| Full-text outcome | 727 systematic-core studies; 91 supporting primary studies outside one core eligibility boundary; 21 contextual/excluded records [5] |
| Source: search-and-screening protocol and final screening record [5]. | |
Table 2.
Evidence framework used to code diagnostic capability, validation realism, damage authenticity, generalisation, temporal maturity, engineering-use maturity, target supervision, and trustworthiness attributes. [6].
Table 2.
Evidence framework used to code diagnostic capability, validation realism, damage authenticity, generalisation, temporal maturity, engineering-use maturity, target supervision, and trustworthiness attributes. [6].
| Dimension | Code | Definition |
|---|---|---|
| Diagnostic capability | D0 | Enabling inference only [6] |
| D1 | Detection | |
| D2 | Localisation | |
| D3 | Characterisation or quantitative diagnosis [6] | |
| D4 | Future structural-condition/damage prognosis [6] | |
| Validation realism | R1 | Numerical/synthetic |
| R2 | Laboratory/reduced-scale [6] | |
| R3 | Full-scale controlled/physical benchmark [6] | |
| R4 | Operational in-service bridge monitoring [6] | |
| Damage authenticity | A0 | No confirmed physical damage [6] |
| A1 | Numerical/synthetic damage [6] | |
| A2 | Physical laboratory damage [6] | |
| A3 | Synthetic/injected damage on real bridge measurements [6] | |
| A4 | Confirmed full-scale physical structural change; may be described as controlled, intervention, or confirmed [6] | |
| Generalisation | G0 | Same structure/domain [6] |
| G1 | Unseen conditions on same bridge [6] | |
| G2 | Unseen bridge transfer/external validation [6] | |
| G3 | Population/network-level transfer across distinct bridges [6] | |
| Target supervision | ZS | Zero-shot |
| UT | Unsupervised target adaptation [6] | |
| FS | Few-shot labelled target data [6] | |
| ST | Substantial supervised target retraining [6] | |
| Temporal maturity | T0 | Short campaign |
| T1 | Repeated/medium term | |
| T2 | Long-term operational monitoring [6] | |
| Engineering-use maturity | M0 | Enabling/algorithmic output [6] |
| M1 | Structural diagnostic information [6] | |
| M2 | Warning/inspection/intervention support [6] | |
| M3 | Explicit prognosis and/or maintenance/asset-management decision support [6] | |
| Trustworthiness | E | Environmental/operational variability addressed [6] |
| U | Uncertainty quantified [6] | |
| X | Explainability/physical interpretation [6] | |
| P | Physics integrated | |
| F | Information fusion | |
| Source: authors' evidence-maturity framework applied to the final corpus [6]. | ||
Table 3.
Corpus-level maturity distributions for the 727 systematic-core studies.
| Dimension | Final distribution | High-maturity subset | Critical interpretation |
|---|---|---|---|
| Diagnostic capability | D0 244; D1 246; D2 102; D3 129; D4 6 | D4 = 6 (0.8%) | The literature is dominated by enabling inference and detection; prognosis is exceptional. |
| Validation realism | R1 350; R2 91; R3 86; R4 200 | R4 = 200 (27.5%) | Operational evidence is substantial but still smaller than numerical/synthetic validation. |
| Damage authenticity | A0 357; A1 186; A2 76; A3 47; A4 61 | A4 = 61 (8.4%) | Confirmed full-scale physical damage remains uncommon. |
| Realism + authenticity | R4+A4 = 8 | 8 (1.1%) | Operational monitoring rarely coincides with independently confirmed physical damage. |
| Generalisation | G0 644; G1 68; G2 12; G3 3 | G2/G3 = 15 (2.1%) | Most AI remains structure-specific; multiple bridge applications are not equivalent to transfer. |
| Temporal maturity | T0 676; T1 35; T2 16 | T2 = 16 (2.2%) | Long-term operational validation is rare. |
| Engineering-use maturity | M0 212; M1 440; M2 72; M3 3 | M3 = 3 (0.4%) | Decision integration lags algorithmic and diagnostic development. |
| Trustworthiness | E 148; U 38; X 37; P 42; F 28 | 487 studies demonstrate none of E/U/X/P/F | Trustworthiness attributes are much less frequently demonstrated than predictive performance. |
| Source: final study-level maturity coding. | |||
Table 4.
Principal evidence gaps and validation requirements for operational AI-enabled bridge NDT. [1,2].
| Code / evidence gap | Current evidence pattern | Why it matters | Evidence required to substantially narrow the gap |
|---|---|---|---|
| EG1 Validation realism | R1 350 (48.1%); R4 200 (27.5%); advanced D2-D3 exists across realism levels. [1,2,3] | High within-study accuracy can remain model-specific or environmentally fragile. [1,2,3] | Prospective R4 validation, locked models, time-separated tests, independent full-scale bridge evaluation. [1,2,3] |
| EG2 Cross-bridge generalisation | Only 15 studies (2.1%) reach G2/G3; 5 combine transfer with R4. [1,2,3] | Bridge-specific development limits network-scale scalability. [1,2,3] | Held-out target bridges, explicit ZS/UT/FS/ST budgets, structural-similarity reporting, negative-transfer analysis. [1,2,3] |
| EG3 Damage authenticity | A4 occurs in 61 studies, but only 8 combine A4 with R4; 47 R4 studies are A3. [1,2,3] | Real bridge measurements can be mistaken for real-damage evidence. [1,2,3] | Inspection/NDT confirmation, documented interventions, controlled full-scale damage, prospective deterioration observations. [1,2,3] |
| EG4 Integrated trustworthiness | E 148; U 38; X 37; P 42; F 28; 487 studies demonstrate none of the five. [1,2,3] | One trust attribute does not address all operational failure modes. [1,2,3] | Integrated E/U/X/P/F under authenticated field conditions and external transfer, with abstention under domain shift. [1,2,3] |
| EG5 Temporal/prognostic maturity | T2 16 (2.2%); D4 6 (0.8%). [1,2,3] | Long-duration response forecasting is not deterioration prognosis. [1,2,3] | Prospective future-condition forecasts with calibrated uncertainty and later physical verification. [1,2,3] |
| EG6 Engineering decision integration | M2 72 (9.9%); M3 3 (0.4%); M0+M1 652 (89.7%). [1,2,3] | Infrastructure management requires actionable consequences, not only scores or labels. [1,2,3] | Inspection/maintenance triggers, false-alarm and missed-detection costs, risk reduction, lifecycle decision benefit. [1,2,3] |
| Source: corpus-level synthesis and supporting bridge-AI investigations [1,2,3]. | |||
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.