Submitted:
27 September 2026
Posted:
30 September 2026
You are already at the latest version
Abstract
Battery-electric and hybrid ships combine battery energy storage systems with networked control, machine-learning state estimation, digital twins, high-capacity charging, and, in some applications, second-life cells. These configurations introduce risks for which maritime approval instruments may lack verifiable acceptance criteria. This paper defines a Regulatory Readiness Level (RRL) from 0 (unrecognised) to 4 (verifiable acceptance) and applies it to nine frontier risks in three families: cyber-physical, AI-driven, and second life/lifecycle. The assessment covers IMO, EMSA, IACS, IEC, UL, NFPA, and battery requirements of five classification societies, using instruments in force on 31 July 2026. No risk reaches RRL 4 under the study’s target. Median readiness is RRL 2: one risk is at RRL 0, two at RRL 1, four at RRL 2, and two at RRL 3. The two RRL 3 cases illustrate different limitations: state-of-health estimation is verified periodically without qualification between tests, whereas mixed-chemistry and mixed-state-of-health packs are addressed through prohibition rather than qualification. Baseline priority places AI3, auditability of AI safety functions, first, with AI2 and SL1 in the next band; SL1 is sensitive to aggregation. Ordinal dominance excludes CP2, CP3, and SL1 from the lead without assuming equal spacing. Sensitivity analysis shows that uncertainty in readiness classification has a larger effect on ranking than uncertainty in severity and exposure. Adjacent-sector assurance mechanisms provide templates for maritime incorporation, and the paper proposes an acceptance objective, verification method, and regulatory vehicle for each risk.
Keywords:
electric ships
; maritime battery safety
; regulatory readiness
; cyber-physical security
; artificial intelligence
; battery management system
; second-life batteries
; digital twin
; mixed chemistry
; SOLAS alternative design
1. Introduction
1.1. From First-Generation to Next-Generation Shipboard Storage
Battery-electric and hybrid propulsion has moved from demonstration to mainstream within a decade, propelled by the revised IMO greenhouse-gas strategy and EU instruments such as FuelEU Maritime [1,2]. The first wave of shipboard battery energy storage systems (BESS) was, with few exceptions, modest in capacity, built from a single cell chemistry, governed by a self-contained battery management system (BMS), and operated under direct human supervision. The regulatory and standards apparatus that grew up around these systems (EMSA guidance, classification-society rules, and component standards such as IEC 62619) reflects those characteristics [3,4,5,6,7].
The installations now being designed and retrofitted depart from several of these assumptions at once. The ELECTRIC BLUE project, for example, is demonstrating a substantially expanded battery system on an existing RoPax ferry, combined with high-capacity shore charging, an AI-driven energy management system, and digital tools for ship modelling [8]. These documented features illustrate the direction of travel, but they do not imply that every frontier feature considered in this paper is present in one demonstrator. We therefore use the broader term next-generation maritime battery systems for installations exhibiting one or more emerging features that materially change the safety-assurance problem. The defining features used to construct the taxonomy are labelled F1–F7: scale and retrofit (F1); chemistry and state-of-health heterogeneity within one installation (F2); networked, shore-coupled control (F3); machine-learning (ML) state estimation (F4); learning-based or autonomous energy management (F5); digital twins used in safety decisions (F6); and second-life cells (F7). The feature set is intentionally broader than any single project so that the assessment captures plausible combinations that may enter service as maritime battery systems evolve.
1.2. The Readiness Question
Battery installations reach service under SOLAS Reg. II-1/55, which allows machinery and electrical installations to deviate from the prescriptive requirements of parts C, D, E or G provided the alternative design and arrangements meet the intent of those requirements and provide an equivalent level of safety [9,10,11]. The mechanism is important for designs that depart from prescriptive requirements, including emerging battery-electric configurations. As of the assessment cut-off, dedicated IMO safety provisions for batteries as a main source of electrical power were still under development; IMO had initiated work on a safety framework addressing lithium-ion batteries and related technologies [12]. SOLAS Reg. II-1/55 nevertheless provides the existing alternative-design pathway. The engineering analysis under this pathway must determine performance criteria that provide a level of safety not inferior to the displaced prescriptive requirements, and those criteria shall be quantifiable and measurable [9]. The present paper operationalises the stronger idea of a fully checkable acceptance basis as RRL 4. Where a risk remains below that level, an alternative-design case may still be accepted, but the acceptance basis must be developed more case-specifically and is less directly comparable across approvals. Existing scholarship examines how present frameworks treat the risks they do address, and finds that requirements differ enough between them to matter [13]. The problem taken up here is prior to that one and, for next-generation systems, more acute: risks the frameworks were never built to anticipate at all, for which the question is not how requirements differ but whether any requirement exists to differ about. The relevant question is therefore one of readiness, the degree to which the instruments that would govern an approval recognise a frontier risk, impose a control for it, and provide a verifiable basis on which an administration could accept the design.
Readiness is not the same as technical maturity. A hazard can be well understood in the engineering literature yet have no regulatory hook; a control can exist in a general clause yet lack any acceptance criterion against which equivalence could be demonstrated. The contribution of this paper is to make that distinction measurable for the specific frontier risks that next-generation systems introduce.
1.3. Related Work and the Gap Addressed
The component hazards are individually well studied. Thermal-runaway mechanisms and fire behaviour are studied extensively [14,15,16,17]; maritime digital twins and their move toward real-time cyber-physical operation are surveyed [18,19]; data-driven and cybersecurity-enabled approaches are also being applied to maritime operational and environmental analysis [20], while the cybersecurity exposure of shipboard power electronic converters has been examined specifically within cyber-physical architectures [21]; data-driven battery state-of-health estimation is advancing [22]; energy-management strategies for hybrid ships, including learning-based approaches, are mapped [23,24]; the safety readiness of next-generation marine batteries has been flagged from an engineering standpoint, with the supporting literature noted to be scarce [25]; and the prospects and pitfalls of second-life batteries have been critically reviewed [26]. Critically, these contributions establish the technical maturity and hazards of the emerging technologies; they do not assess whether the regulatory instruments are equipped to govern them, and treat regulatory readiness, at most, as a closing remark. On the regulatory side, EMSA’s guidance, the IMO cyber-risk-management instruments and the recent IACS cyber-resilience requirements represent the current state of the maritime instruments [3,27,28,29,30]. Reviews of battery-electric ship safety and of marine lithium-ion risk survey the hazards and the assessment methods, and identify incomplete technical standards as a risk category in their own right [31,32,33]; a systematic overview of the IMO instruments and the IACS unified requirements reaches the complementary conclusion that, while requirements exist for batteries used as an emergency source of power, no specific requirements govern battery systems used as a ship’s main or auxiliary source of power [34]. Work on maritime autonomy has begun to treat certification itself as the object of study, both for the approval pathway [35] and for the identification of AI-specific risks [36], and state-of-safety estimation for maritime battery management has been reviewed in the same spirit [37]. Recent work has also proposed dynamic risk assessment and safety-case frameworks for shipboard lithium-ion batteries, linking hazard evolution to detection, mitigation, and safety-case structure [38].
A further strand of work, outside shipping, has begun to supply exactly the assurance instruments that maritime frameworks lack. ISO/IEC TR 5469 maps the use of AI within safety-related functions onto established functional-safety concepts [39]; ISO/PAS 8800 defines an AI-safety framework for road vehicles that extends ISO 26262 and ISO 21448 [40]; UL 4600 provides an assurance-case standard for autonomous products [41]; EASA’s AI concept paper offers usable certification guidance for machine-learning applications in aviation, including learning assurance and explainability [42]; and DNV has issued cross-industry recommended practices for the assurance of AI-enabled systems and for the qualification of digital twins [43,44]. On the second-life side, the EU Battery Regulation introduces battery passports and state-of-health data obligations [45], and UL 1974 defines a repurposing-evaluation process [46]. None of these is a maritime approval instrument, and none is currently invoked by the maritime instrument set; their existence sharpens the readiness question addressed here, and they return in Section 4.3 as templates for the proposed acceptance objectives.
Two things are therefore absent from the literature. The first is an assessment that takes the emerging risks as the unit of analysis and asks systematically how ready the maritime instruments are for each, distinguishing a risk that is unrecognised from one that is named but uncontrolled, and from one that is controlled but unverifiable: existing work either establishes that a hazard is real, which is an engineering question, or compares how frameworks treat hazards they already cover, which presupposes that coverage exists. The second, and the more consequential, is any treatment of what that absence does to the approval mechanism itself. Regulatory gaps are conventionally discussed as incompleteness, a list of requirements someone ought to write. The argument developed here is different in kind: because alternative-design acceptance is the route by which battery ships are actually approved, and because that route can only certify equivalence against a stated criterion, the absence of such criteria is not a gap in the rules but a limit on what the certification mechanism can be said to establish. Both gaps are addressed here, and the second supplies the reason the first matters.
1.4. Contributions and Organisation
Four things in this paper are new, and it is worth stating them as claims, not as activities.
First, a structural result about the approval pathway. Alternative-design acceptance is shown to carry a precondition that has not previously been made explicit: an equivalence claim is assessable only where some instrument states what would count as acceptable. This reframes regulatory gaps from a completeness question, which requirements are missing, into a question about whether the certification mechanism itself can function, and it yields the paper’s central finding, that for the frontier risks of next-generation systems the precondition mostly fails.
Second, a construct that makes the precondition measurable. Regulatory readiness is distinguished from technical maturity, and the RRL scale separates three states that the literature and practice routinely conflate: a risk no instrument recognises, a risk named but carrying no obligation, and a risk controlled but with no criterion by which compliance could be judged. The distinction is not a matter of degree. Its two thresholds, from acknowledgement to obligation and from obligation to verifiable acceptance, are where readiness is actually lost, and locating a risk relative to them determines what kind of regulatory action would help.
Third, an empirical prioritisation result. Applying the scale to nine frontier risks shows a marked concentration of the largest readiness gaps in the AI-driven family. AI3 has the largest baseline gap and highest baseline priority, while AI2 and SL1 occupy the next priority band; however, the sensitivity and ordinal analyses show that the exact ordering within this group is not invariant to the weighting assumptions. The analysis therefore distinguishes what follows from the ordinal data alone from what depends on the chosen aggregation.
Fourth, a diagnosis with a route out. Adjacent regulated sectors already provide assurance mechanisms and test structures relevant to the identified gaps. Their value here is as evidence and templates for maritime incorporation, not as current maritime requirements. The paper therefore states a specific verifiable acceptance objective, verification method, and candidate regulatory vehicle for each of the nine risks.
Supporting these, the assessment is made contestable, not asserted: an explicit decision-rule rubric (Table 5) fixes the evidence test for each level, and the scoring rationale is published for every instrument–risk cell (Appendix A; Table S1, Supplementary Materials), so a reader who disagrees with a score can locate the exact judgement to challenge. Section 2 defines the method; Section 3 presents the taxonomy, readiness profile, priority ranking and sensitivity results; Section 4 discusses causes, the amplifying role of scale, readiness templates from adjacent regimes, the readiness-raising roadmap, the implication for alternative-design acceptance, a worked retrofit example, and limitations; Section 5 concludes.
2. Materials and Methods
The method has five steps: define the instrument set to be assessed; define a threat-led taxonomy of frontier risks; score the current regulatory readiness of each risk on the RRL scale; rank the residual readiness gap by criticality; and test the robustness of the ranking. The workflow is summarised in Figure 1.
2.1. Instrument Set
The instruments assessed are those an administration or classification society would consult when approving a next-generation installation as an alternative design under SOLAS Reg. II-1/55 with MSC.1/Circ.1455 [9,10]. The set comprises four groups.
The first is the IMO instruments: the framework conventions and the cyber-risk-management resolution with its guidelines [1,9,10,27,47], together with EMSA’s guidance on the safety of BESS on board ships [3]. The second is the IACS unified requirements bearing on battery installations and their control systems, namely the type-approval test specification UR E10 [48], the high-voltage requirements UR E11 [49], the battery-schedule and maintenance-cycle requirement UR E18 [50], the UPS requirements UR E21 [51], the computer-based-systems requirements UR E22 [52], and the cyber-resilience requirements UR E26 and UR E27, in force for ships contracted on or after 1 July 2024 [28,29]. The third is the battery rules and technical references of five classification societies: DNV [5,6,53], ABS [4,54], Bureau Veritas [55], ClassNK [56] and Lloyd’s Register [57]. The fourth is the cross-sector standards drawn upon by the first three: IEC 62619 on industrial cells [7], IEC 60092 on marine electrical installations [58], IEC 61508 on functional safety [59], IEC 62443 on industrial cyber security [60], UL 9540A on thermal-runaway propagation [61] and NFPA 855 on the installation of stationary energy storage [62].
Two boundary decisions deserve to be explicit, because they shape what the scores mean. First, instruments still under development are noted where relevant but are not scored, because a readiness score requires a published text: this excludes the megawatt-charging-system work (the SAE J3271 technical information report and the IEC 63379 project) [63] and the IMO MASS Code. The published shore-connection standard IEC/IEEE 80005-1 [64] was reviewed; because its scope is high-voltage shore power supply and not high-rate charging of propulsion batteries, it informs the assessment of the ship–shore risk (Section 3.2) without governing it. Second, adjacent-sector assurance instruments [39,40,41,42,43,44,45,46] are not part of the scored set, because they are neither maritime approval instruments nor referenced by one. This scoping choice is conservative in one direction only: if an administration chose to invoke such an instrument in an alternative-design case, the effective readiness could exceed the scores reported here, a possibility taken up in Section 4.3.
Selection protocol
The instrument set was assembled by the following procedure, stated so that it can be repeated. The assessment reflects the instruments in force on 31 July 2026, which is the cut-off for all scores reported here; instruments issued or revised after that date are not scored.
Candidates were located in three passes. The first took the instruments named in the approval pathway itself, SOLAS Reg. II-1/55 and MSC.1/Circ.1455, together with those they invoke, and the battery-specific guidance and rules that an administration or society would consult for a BESS installation. The second followed the normative references of each document located in the first pass, one level deep, which is how the cross-sector standards (IEC 61508, IEC 62443, UL 9540A, NFPA 855) entered the set. The third searched the standards catalogues of IMO, IEC, ISO, UL, NFPA, SAE and IACS, and the publication indexes of the classification societies, for battery, energy storage, cyber resilience, artificial intelligence and second-life subject terms, to capture instruments that no located document happened to cite.
Five classification societies were examined in full: DNV, ABS, Bureau Veritas, ClassNK and Lloyd’s Register. They were selected because each publishes dedicated battery-installation requirements instead of relying solely on the IACS unified requirements, and because between them they cover the societies most often engaged for European and Asian battery-electric newbuildings and retrofits. The remaining IACS members were not examined individually. That residual exposure is bounded, not assumed away: UR E10, E11, E18, E21, E22, E26 and E27 bind every IACS member, so an unexamined society could raise a row maximum only by publishing a battery-specific requirement more demanding than all five examined societies. Where the five diverge, the analysis records the strongest, so the scores already reflect the most favourable reading available across a majority of the classed fleet. The marginal return on adding societies is also now visible instead of assumed: Lloyd’s Register was added after the assessment had been completed for the other four, and although it contributes substantive requirements, including a mandatory failure-mode analysis covering the hidden loss of any monitoring, alarm or safety function on which the installation depends [57], it changed no row maximum. Where an instrument exists in several editions, the edition in force at the cut-off was scored and superseded editions were disregarded; where a document is a series published piecemeal (IEC 60092, IEC 62443), the parts relevant to the risk under assessment were scored as one instrument.
Eligibility for the scored set required three conditions to hold together: the document is published (not a draft or a project), it is either a maritime approval instrument or is normatively drawn upon by one, and it is capable in principle of bearing on at least one frontier risk. The second condition is what separates IEC 62443, which IACS UR E27 draws upon for its security capabilities and which therefore reaches a maritime approval through a maritime instrument, from DNV-RP-A204, which is a recommended practice that no maritime approval instrument invokes and which an administration is accordingly under no obligation to apply. The distinction is one of normative incorporation, not of technical quality or of sector of origin: were a maritime instrument to invoke DNV-RP-A204, it would enter the scored set and would raise the readiness of CP3. Table 1 records the decision and its ground for every candidate considered.
2.2. Frontier-Risk Taxonomy
Nine frontier risks were derived by contrasting the defining features F1–F7 of next-generation systems (Section 1.1) with the implicit assumptions of the instrument set, and by reference to the emerging-technology literature [18,23,25,26]. The derivation follows the systems-theoretic view of safety, in which accidents arise from inadequate control of a system’s interactions and not from component failure alone [65,66]: each defining feature is examined for the control loops it introduces or alters (networked and learning-based control, model-mediated sensing, and heterogeneous, ageing energy sources), and a candidate risk is recorded wherever a feature can defeat, bypass or invalidate a safety function that present approvals credit. This control-centred lens is what makes the taxonomy threat-led, and it is well matched to the cyber-physical and autonomous elements of next-generation systems, to which systems-theoretic hazard analysis has increasingly been applied [65]. Formally, a risk was admitted when (a) at least one defining feature creates or materially transforms it relative to a first-generation installation, and (b) it is capable of defeating or bypassing a safety function credited in present approvals, so the taxonomy excludes hazards that are merely scaled versions of well-regulated ones. The risks are grouped into three families (Figure 2, Table 2); the grouping is deliberately threat-led, not domain-led, so that the unit of analysis is the emerging risk and not the framework, and the result is not conditioned on how any existing instrument happens to organise its subject matter. Table 3 records the traceability of each risk to the features that drive it. The nine risks are complete with respect to the defining features F1–F7 in a precise sense: every feature maps to at least one risk it drives (Table 3) and no admitted risk is one a first-generation installation already presents in the same form. Completeness is therefore relative to the feature set, not a claim that no further risk could emerge as the technology evolves, a boundary revisited in Section 4.7.
Two properties of this construction should be made explicit, because the completeness claim is weaker than it may appear. The first concerns the provenance of the features themselves. F1–F7 were abstracted from the design intent of the retrofit configuration of Section 1.1 and then checked against the emerging-technology literature [18,23,25,26] for features that configuration does not exhibit. That procedure establishes internal coverage: every feature is matched to a risk, and no risk lacks a feature that drives it. It does not establish external completeness, because the feature set is itself author-defined and anchored to one motivating configuration, so a feature that neither that configuration nor the surveyed literature displays would be invisible to the method. Candidate omissions can be named: solid-state or sodium-ion chemistries with different abuse signatures, swappable or containerised pack architectures, fleet-level or shore-side aggregation of battery control, and supply-chain integrity of cells and firmware. None is excluded on principle; each would enter as a further feature and would generate its own risks. An independent validation, by structured expert elicitation or a formal horizon scan instead of the authors’ reading, is the appropriate test of the taxonomy and has not been performed here (Section 4.7).
The second concerns the independence of the nine risks. They are not statistically independent, and in two places they share an underlying deficit. AI1, AI2 and AI3 all trace to the absence of a qualification route for non-deterministic functions, and CP3 and AI1 both concern reliance on a model-mediated estimate of physical state. The taxonomy nonetheless separates them because they fail differently and would be closed by different instruments: AI1 is an accuracy-and-uncertainty requirement on an estimator, AI2 an envelope-and-supervision requirement on a controller, and AI3 an auditability requirement on the evidence trail, so a single instrument satisfying one leaves the others open. The consequence for the ranking is that the AI family’s aggregate weight partly reflects one deficit counted through three manifestations, which would matter if priorities were summed within families. They are not: risks are ranked individually and the roadmap of Section 4.4 assigns a separate acceptance objective to each, so no risk’s priority is inflated by its neighbours. A reader who regards AI1, AI2 and AI3 as one risk would obtain a shorter list with the AI deficit still at its head.
2.3. The Regulatory Readiness Level (RRL) Scale
Readiness is scored on a five-point ordinal scale (Table 4, Figure 3) that separates the qualitatively different states a risk can occupy in an instrument. The scale continues the tradition of readiness-level metrics that began with the NASA technology readiness level (TRL) ladder [67,68] and was later extended to system- and integration-level maturity; like those metrics it is an ordinal figure of merit, but its subject is the instrument and not the technology, so that a system can be at high technical maturity while its governing instruments remain at RRL 0. Five levels are used, rather than the nine of the TRL scale, because the regulatory question turns on a small number of qualitative transitions and not on a fine gradation of maturity: the decisive thresholds are between RRL 1 and RRL 2 (from mere acknowledgement to an enforceable control) and between RRL 3 and RRL 4 (from a control to a verifiable acceptance criterion). RRL 4 is the target adopted in this study for a fully verifiable acceptance basis against which an alternative-design equivalence claim can be checked.
The term regulatory readiness level is not new, and the sense used here is deliberately different from the established one. In the technology-transition literature an RRL measures how ready the regulatory environment is to permit a product to reach market, its levels running from a situation where the legal position is unpredictable to one where use and production are approved or unproblematic [69,70,71]. That construct takes the product as its unit of analysis and legalisation as its end state, and it sits alongside market, acceptance and organisational readiness in a balanced assessment of a technology’s prospects [70]. The scale defined here takes the instrument as its unit of analysis and a specific frontier risk as its object, and its end state is not permission to market but the existence of a criterion against which an equivalence claim could be judged. A technology may therefore be at the highest regulatory readiness in the market-entry sense, freely approvable, while the risks it introduces sit at RRL 0 in the sense used here, which is precisely the condition this paper documents.
2.4. Scoring Procedure
To keep scoring rule-based and not impressionistic, the ordinal definitions of Table 4 are operationalised as the explicit decision rules of Table 5: each level is assigned only when a stated evidence test is met, and the test for the next level is not. Two assessors applying these rules to the same instrument set should reach the same score or be able to name the exact clause on which they differ; the rubric is thus the instrument of reproducibility. The need for it is not hypothetical. A multi-industry study of readiness-level practice found the subjectivity of the assessment and the imprecision of the scale to be among the principal difficulties reported by practitioners, with assessors of the same artefact reaching different levels and those favouring a technology tending to score it higher [72]; anchoring each level to a stated evidence test is the available remedy. Appendix A records, for every row maximum, which test was met and why the next was not, and Table S1 of the Supplementary Materials gives the same record for every cell of the matrix, so that the assessment can be reproduced or contested at the level of the individual instrument and not only at the row maximum.
For each frontier risk r, every instrument i in the set was read against the RRL definitions and assigned . The current readiness of the risk is the best any single instrument achieves,
together with a record of the governing instrument(s) attaining that maximum. The maximum (rather than a mean) is used because a single sufficient instrument confers readiness regardless of silence elsewhere; conversely, a risk scores low only when no instrument addresses it, which is the condition of interest. The maximum is the mirror image of the “weakest link” roll-up used in technology-readiness practice, where a system inherits the lowest level among its components, a convention that study of the field has found unsatisfactory because it directs attention to low scores that may be cheap to raise instead of those that matter [72]. Here the direction is reversed for a different reason. The maximum is deliberately a favourable upper-bound rule for the question asked: it credits the regime with the most favourable reading, so a low score cannot be an artefact of overlooking a capable instrument, and it isolates readiness (whether verifiable acceptance is available anywhere) from adoption (whether a given administration invokes it), which is a separate matter taken up in Section 4.4.
Two consequences of the maximum rule need to be stated, because each could otherwise distort a score. The first is that an instrument may be technically capable without being regulatorily applicable. IEC 62443, for example, supplies detailed security capabilities, but a flag administration or society that has not invoked it has no approval basis in it, so crediting a risk with the readiness of an instrument nobody applies would overstate the regime. The assessment therefore admits an instrument to the scored set only where it is either a maritime approval instrument or normatively drawn upon by one (Section 2.1), and the higher levels carry a stricter condition: RRL 3 and RRL 4 are assigned only when the risk has both a maritime regulatory hook, a clause in an instrument an administration or society actually applies, and a technical verification mechanism adequate to the claim. A verifiable criterion sitting in an instrument with no maritime hook raises readiness only once a maritime instrument invokes it, which is why the adjacent-sector instruments of Section 4.3 are treated as templates for future readiness, not as current readiness. In the present assessment this condition binds in one place. SL2 reaches RRL 3 through the ABS prohibition on mixing chemistries in one circuit, which is both a maritime class requirement and determinate, so it satisfies the condition; no risk reaches RRL 3 by a route lacking a maritime hook. The condition nonetheless matters for reassessment, because several of the adjacent-sector instruments of Section 4.3 would otherwise raise scores the moment they were cited. The second is the converse: several instruments might each contribute part of an assurance basis that none supplies alone, in which case the maximum understates the regime. The scored matrix (Table 8) shows this to be a live possibility in two places. For CP1, IMO, IACS and IEC 62443 each reach RRL 2 from a different direction; for AI1, ABS, Bureau Veritas and ClassNK each do so by imposing the same kind of state-of-health obligation. Taken together, however, neither set supplies an acceptance criterion that its members lack individually: three identical obligations to monitor state of health still say nothing about how the estimate is qualified. Neither row maximum is therefore understated. A composite rule would therefore change no score here, but it should be revisited whenever several instruments reach RRL 2 or above on the same risk.
To make the procedure concrete, consider CP1, the compromise of safety-critical battery control. IMO Resolution MSC.428(98) obliges companies to address cyber risks in their safety-management systems, and the associated guidelines elaborate a risk-management process [27,47]; this is a general goal-based clause with no battery-specific content, hence . EMSA’s guidance acknowledges the dependence of safety functions on software and connectivity without attaching a requirement: RRL 1. IACS UR E26/E27 impose dedicated, surveyable cyber-resilience requirements on onboard systems, and IEC 62443 supplies the security capabilities they draw on; because their content is generic to computer-based systems and contains nothing specific to the safety functions of a battery installation, both score RRL 2 under the scale’s definitions and not RRL 3. The cell-level test standards and NFPA 855 are silent: RRL 0. The row maximum is therefore . The corresponding rationale for every row maximum is given in Appendix A, so that the assessment can be reproduced or contested cell by cell.
Scoring was document-based and performed by a single assessor; the scores are therefore provisional, and Section 4.7 states the consequences. The rubric of Table 5 is what makes an independent re-scoring convergent instead of impressionistic: because it fixes the evidence test for each level, a second assessor applying it to the same instrument set can reproduce or contest each score cell by cell, and inter-assessor agreement can be quantified as a linearly weighted Cohen’s on the ordinal scale. Establishing that agreement with an independent assessor is the natural next step and is identified as a required validation in Section 4.7. To avoid resting any conclusion on a single point value in the meantime, results are reported as an ordinal profile and as a priority rank, not as precise measurements, and the stability of the rank is tested in Section 3.4.
2.5. Criticality Weighting and the Readiness Gap
A readiness gap matters in proportion to the harm that would follow if the risk were realised and to how central the risk is to next-generation systems. Each risk was therefore assigned an ordinal severity (worst-credible consequence, informed by the forensic record of large stationary-storage events [73], whose off-gassing and deflagration mechanisms are, if anything, more consequential in the confined and occupied spaces of a ship than in an open-air installation) and an exposure (how intrinsic the risk is to next-generation installations), combined into a criticality
The readiness gap and the resulting priority are
with the target (verifiable acceptance) for all risks. A risk thus ranks high either because it is highly critical and partly addressed, or because it is moderately critical and wholly unaddressed; both are legitimate calls on regulatory attention, and Equation (3) captures both. Both scales are anchored at every level, not at the endpoints alone, so that a rating can be contested against a stated descriptor; the descriptors are given in Table 6 and the rating assigned to each of the nine risks is justified individually in Table 7. The severity and exposure ratings remain expert, ordinal and provisional; the limitations of this assumption are discussed in Section 4.7.
Equations (2) and (3) perform arithmetic on scales that Section 2.3 and Section 2.5 define as ordinal, and the status of that step should be stated plainly instead of assumed. Adding severity to exposure and multiplying criticality by the gap presuppose that successive levels are approximately equally spaced, which the level definitions do not establish: the step from RRL 0 to RRL 1 (from silence to mention) is not self-evidently the same size as the step from RRL 3 to RRL 4 (from a control to a verifiable criterion), and the same reservation applies to the severity and exposure anchors. Three considerations govern how the resulting numbers are used here. First, is treated as a screening device that orders risks into bands, not as a measurement: results are reported as a rank and a band membership, and no claim is attached to a difference of one priority point. Second, the construction is deliberately coarse in a direction that limits the damage, since C is a ceiling of a two-input mean on a five-point scale and takes only the values 4 and 5 across the nine risks, so the ranking is driven mainly by the gap G, which is a difference of readiness levels against a fixed target and not a product of two subjective scales. Third, and decisively, the substantive conclusion is not left resting on the arithmetic: the dominance analysis of Section 2.6 re-derives the leading group using only order comparisons within each scale, so the finding survives any monotone rescaling of the levels. Where the interval assumption does carry weight, in the ordering within the leading group and in the exact priority values, the conclusions are correspondingly hedged.
2.6. Sensitivity Analysis
Because the RRL, severity and exposure inputs are all ordinal, the priority ranking should be robust to reasonable changes in how the inputs are combined; a ranking that depended on the choice of aggregation would not be a safe basis for regulatory prioritisation. Four checks are applied. First, the ceiling-mean aggregation of Equation (2) is replaced by the multiplicative form used in conventional risk-matrix practice,
and the full ranking is recomputed; the agreement of the two aggregations is summarised by their Spearman rank correlation. Second, every one-step perturbation of a single severity or exposure rating ( within the scale bounds) is enumerated exhaustively, all 30 admissible single-input moves across the nine risks, and the resulting rank changes recorded; analytically, a one-step change alters C by at most one and hence P by at most , so the baseline margins between priority bands bound which rank reversals are possible at all. Third, to probe simultaneous rather than isolated rating error, a Monte Carlo experiment resamples every rating at once and recomputes the ranking. Each input is redrawn uniformly from its admissible neighbours, the values within one step of the assigned rating that lie inside the scale bounds, so that a rating of 5 is redrawn from and a rating of 3 from . Sampling the admissible set directly, rather than adding and clipping, avoids the bias that clipping introduces at the scale boundaries, where a value of 5 would otherwise be retained by both and and would therefore be treated as less uncertain than an interior value. Two variants are run over draws each: the first resamples severity and exposure only, isolating uncertainty in the criticality weights; the second additionally resamples on by the same rule, so that uncertainty in the readiness classification itself is propagated alongside it. Each variant records the probability that each risk is among the four highest priorities, that the highest priority is attained by a leading-group member, that the most-ready risks (AI1 and SL2) stay in the lowest third, and the mean Spearman correlation between the perturbed and baseline priority vectors; ties share a rank band throughout.
Fourth, because the objection that ordinal ratings should not be added or multiplied cannot be answered by re-weighting alone, the leading group is also derived without any arithmetic on the ratings at all. Risk a is said to dominate risk b when , and , with at least one of the three strict: a is then at least as severe, at least as exposed, and no better provided for than b, so any prioritisation that respects the order of the three scales must rank a no lower than b. The relation uses only comparisons within each scale and is therefore invariant under any monotone rescaling of severity, exposure or RRL, including rescalings in which successive levels are unequally spaced. The set of non-dominated (Pareto-maximal) risks is the part of the ranking that follows from the ordinal data alone, independently of Equations (2)–(3). The results of all four checks are reported in Section 3.4.
3. Results
3.1. Current Readiness Profile
The instrument-level scores from which the current readiness is derived (Equation (1)) are given in Table 8; the row maximum of that matrix is the current readiness , carried into Table 9 alongside the criticality inputs, the gap and the priority, and plotted against the RRL 4 target in Figure 4. The rationale supporting each row maximum is given in Appendix A.
The overall picture is of low but no longer uniform readiness. Because RRL is ordinal, the profile is reported as a distribution instead of a mean: the median current readiness is RRL 2, and the nine risks divide into one at RRL 0, two at RRL 1, four at RRL 2 and two at RRL 3, with none at RRL 4. No risk, therefore, has a verifiable acceptance criterion, which is the level the alternative-design route requires.
Three features of the profile are worth separating. First, the second-life and lifecycle family, which a reading of the maritime guidance alone would suggest is unaddressed, turns out to be the most provided-for. Mixed-chemistry and mixed-state-of-health packs (SL2) reach RRL 3 because ABS requires that cells of different chemistries and different physical and electrical characteristics are not to be used in the same electrical circuit [4]. Lifecycle revalidation (SL3) reaches RRL 2 on a different basis: IACS UR E18 and Bureau Veritas both require a battery schedule recording maintenance and replacement cycle dates, require replacements to be of an equivalent performance type, and place that schedule in the safety-management system for verification by the surveyor [50,55]. Second, the estimation risks are provided for in a particular and revealing way. ABS, Bureau Veritas and ClassNK each require state of health to be monitored and acted upon [4,55,56] and, more consequentially, independent verification of the state of health calculated by the battery management system is a classification requirement for ships relying on battery power for manoeuvring and propulsion, discharged in practice by annual capacity testing [5,74]. That is a dedicated requirement with a stated verification method, and it lifts AI1 to RRL 3. Bureau Veritas separately requires the type-approval test scope for a battery pack to include tests derived from an FMEA of sensor failures [55], which lifts CP3 to RRL 2. Third, the AI-driven family remains the least ready, and the auditability of AI safety functions (AI3) is the only risk in the set that no instrument names, implies, or can be read to cover.
The highest current readiness, RRL 3, is reached by two risks, AI1 and SL2, and in both cases the manner of reaching it matters for how the score should be read. For AI1 the requirement verifies the estimator’s output at intervals without qualifying the estimator: an annual capacity test establishes that the reported state of health was correct once a year, and says nothing about the accuracy of the estimate between tests, its uncertainty, or its behaviour under operating conditions unlike those it was built for, which is when the estimate is relied upon for a safety decision. Readiness here is real but periodic, and the residual gap is the absence of any bound on the estimator between verifications. For SL2 the position is different again. ABS resolves pack heterogeneity by prohibition instead of qualification: the rule states determinately that the configuration is not permitted, which is a checkable requirement, but it supplies no criteria under which a heterogeneous pack could be accepted. Readiness in the sense measured here, whether an instrument yields a determinate answer, is therefore high; readiness to certify a next-generation retrofit that deliberately combines chemistries and states of health is nil. Section 4.5 returns to this, because it is the clearest instance in the dataset of a rule that forecloses a next-generation configuration instead of governing it.
3.2. Family-Level Findings
Cyber-physical (Family I). IMO Res. MSC.428(98) obliges cyber risks to be addressed in the safety management system, and IACS UR E26/E27 with IEC 62443 establish that networked ship systems must be designed and maintained for cyber resilience, which together lift the compromise-of-control risk (CP1) to RRL 2. UR E27 goes further than a bare obligation: it specifies security capabilities with an approved test procedure stating expected results and acceptance criteria, witnessed by a surveyor, and it requires a computer-based system to set outputs to a predetermined state if normal operation cannot be maintained under attack [29]. UR E22 adds witnessed factory, system and integration testing and requires that a single data-link failure not cause loss of a category III function [52]. The decisive shortfall is specialisation, not rigour: every one of these provisions is written for computer-based systems in general, and none translates into an acceptance criterion for the specific safety functions of a battery installation, so that a deterministic-output requirement tells an approving party what the controller must do on detecting an attack but nothing about whether charge-limit, protection-trip and gas-detection-response functions remain correct and authenticated under a defined threat model. The ship–shore charging interface (CP2) also reaches RRL 2, through the ABS requirements for an offshore charging connection [54], while IEC/IEEE 80005-1 continues to govern the adjacent shore-power case and the megawatt-charging instruments remain unpublished or non-maritime [63,64]. Digital-twin and sensor integrity (CP3) reaches RRL 2 on the strength of the Bureau Veritas sensor-failure FMEA testing and the HAZID coverage of voltage, temperature and gas-sensor failure [55]; the digital-twin half of the risk, however, is untouched, and the term does not appear in any instrument in the scored set.
AI-driven (Family II). This family is now the least ready in absolute terms as well as relative to its criticality, and it contains the whole of the ordinal priority frontier (Section 3.4). The position is not that the instruments are silent about the functions concerned. ABS, Bureau Veritas and ClassNK all require state of charge and state of health to be monitored and used to control charging, discharging and pack connection [4,55,56], so the estimate is not merely permitted but mandated. Nor is the estimate left unchecked: independent verification of the state of health calculated by the BMS is a classification requirement for ships relying on battery power for manoeuvring and propulsion, and is discharged in practice by an annual capacity test in which charge and discharge capacities are measured under rated conditions [5,74]. AI1 therefore reaches RRL 3, the highest readiness observed in this assessment. What the requirement does not do is bound the estimator. It confirms the reported value at one point in the year, by a method the literature itself describes as imperfect, sensitive to depth of discharge, current and temperature, and requiring the ship to be taken out of service [74]; it states no accuracy bound, requires no quantification of uncertainty, and says nothing about behaviour on operating conditions unlike those the estimator was built for. The residual gap is therefore narrow but precisely located: the maritime regime verifies the output periodically and qualifies the estimator not at all. Learning-based energy management in the safety loop (AI2) remains at RRL 1, functional-safety practice presuming deterministic, inspectable logic [59]. Bureau Veritas has published guidance on machine-learning systems, but the rules provide that it may be referred to, so no obligation follows from it [75], which is the acknowledgement-without-obligation threshold in its clearest form. The auditability of AI safety functions (AI3) stands alone at RRL 0 and, combining a maximal gap with moderate criticality, is the single highest-priority readiness deficit in the dataset (Table 9).
A structural point underlies the whole family and is easily missed, because the maritime instruments do formally reach these systems. UR E22 applies to any computer-based system providing control, alarm, monitoring or safety functions, and a battery or energy-management controller is a category III system attracting its fullest treatment. But the verification method UR E22 prescribes is explicitly black-box, confined to correctness, completeness, consistency, intended functionality and intended robustness as observed from outside the system, with no knowledge of its inner workings [52]. Black-box functional testing can establish that a controller behaves correctly on the cases tested; it cannot establish how a learned function generalises beyond them, whether its confidence is calibrated, or how it behaves on inputs unlike its training data. The maritime regime therefore does not simply lack an AI assurance route. The assurance route it does mandate is structurally incapable of qualifying a learned function while formally appearing to cover it, which is a more consequential condition than an acknowledged gap, because it can be satisfied without the underlying question ever being asked.
Second-life and lifecycle (Family III). The recurring theme here is no longer absence but partial provision of an awkward kind. Mixed-chemistry and mixed-SoH packs (SL2) are addressed by prohibition, as described above. Lifecycle revalidation (SL3) reaches RRL 2 through UR E18 and the corresponding Bureau Veritas requirements, which establish a maintenance and replacement regime keyed to cycle dates and equivalent-performance replacement, and place it under survey [50,55]; what they do not do is revalidate the safety margin demonstrated at commissioning against a defined state-of-health or calendar threshold, or state criteria to retire or de-rate. Second-life qualification (SL1) remains at RRL 1, and its evidence is unusually explicit: EMSA’s guidance names second-life batteries, observes that they may pose different safety concerns related to ageing and wear, and states that best practices for their safe reuse are not consolidated and that the guidance accordingly does not cover them [3]. A risk identified, characterised, and expressly left outside scope is the acknowledgement level by definition.
3.3. Priority Ranking
Figure 5 visualizes the relationship between readiness gap and criticality, with bubble area proportional to the resulting baseline priority . The auditability of AI safety functions (AI3) stands apart with the largest gap in the set, followed by learning-based energy management (AI2) and second-life qualification (SL1), each combining a three-level gap with substantial criticality. Below them sit the risks at RRL 2 whose criticality is high but whose gap is halved, the compromise of safety-critical control (CP1) and lifecycle revalidation (SL3), and then the ship–shore interface (CP2) and digital-twin and sensor integrity (CP3). The assurance of ML-based estimation (AI1) and mixed-chemistry pack safety (SL2) fall to the smallest gap in the set, in neither case because the underlying risk is small: the first is verified periodically without being bounded, and the second is foreclosed by rule rather than qualified.
The concentration is therefore sharper than a simple family split would suggest. Regulatory attention has reached the electrochemical and lifecycle properties of battery installations, where the hazards resemble ones the maritime regime already knew how to write about, and has not reached the properties introduced by learning-based and model-mediated control. The practical significance is that the remaining exposure is neither uniform nor diffuse: it clusters in a small number of identifiable places, so a targeted standardisation effort, not a wholesale revision of the instrument set would close most of it.
3.4. Sensitivity of the Ranking
Table 10 compares the baseline ranking with the ranking obtained under the multiplicative criticality of Equation (4). The two aggregations agree closely, at a Spearman correlation of . AI3 and AI2 remain in the leading band under either scheme, whereas SL1 moves from the baseline leading band to fifth place because its lower exposure is weighted differently. AI1 and SL2 remain in the bottom band under both schemes. The result is therefore robust at the level of the main AI-driven priority signal, but the exact position of SL1 is aggregation-dependent.
The perturbation check is both analytic and computational. Enumerating all 30 admissible single-input moves shows that neither AI1 nor SL2 ever enters the leading four and that the largest movement of any risk is three ranks.
Simultaneous rather than isolated error was probed by Monte Carlo, resampling every rating from its admissible neighbours at once (; Table 11). With severity and exposure resampled, the ranking is firmer than in any earlier version of this assessment: the highest priority is attained by a member of the leading group in of draws, AI3 remains in the top four in every draw, AI2 does so in and SL1 in , while AI1 and SL2 never once reach it. The mean Spearman correlation between the perturbed and baseline priority vectors is .
The second variant, which additionally resamples the readiness classification itself, gives a weaker result and is reported to bound what the ranking can bear. One of the three baseline top-ranked risks attains the highest priority in of draws, AI3 stays in the top four in , AI1 and SL2 enter the top four in , and the mean Spearman correlation falls to . The asymmetry between the two variants is structural: priority is the product of a criticality weight and a gap, and the gap is the larger lever, since a one-level error in moves G directly whereas a one-level error in severity or exposure moves C only when it crosses the ceiling in Equation (2). The ranking is therefore robust to the criticality weighting but not to systematic mis-scoring of readiness, which is the quantitative reason why the independent re-scoring described in Section 4.7 is a precondition for treating the ordering, as opposed to the finding of low readiness, as settled. The concern is not peculiar to this scale: the documented difficulties of readiness-level practice in other sectors are dominated by the subjectivity and imprecision of level assignment rather than by the arithmetic applied afterwards [72].
Ordinal dominance check.
The dominance check removes the arithmetic altogether and establishes what follows from the ordering of the three ordinal scales alone. Three risks are dominated: CP2 and CP3 by CP1, and SL1 by AI2. These risks therefore cannot head the ranking under any scheme that respects the ordering of severity, exposure, and readiness. The remaining risks are mutually incomparable under this relation, so their ordering depends on how the scales are weighted against one another. Accordingly, the ordinal analysis excludes CP2, CP3, and SL1 from the lead under any order-respecting rescaling, but it does not establish a unique overall leader. AI3 has the largest baseline gap because it is the only risk at ; its baseline priority also depends on the chosen criticality weighting.
4. Discussion
4.1. Why Readiness Is Low
Three structural features of the instrument set explain the pattern. First, the instruments are component- and point-in-time-oriented: they qualify cells and equipment, predominantly at type approval and commissioning, and have no native mechanism for properties that are emergent (heterogeneous-pack propagation) or time-varying (margin decay). Second, they assume deterministic, local, human-supervised control, which is exactly the assumption that networked and learning-based control violates; cyber resilience has begun to be retrofitted onto this assumption through Res. MSC.428(98) and IACS UR E26/E27, but battery-safety-specific assurance has not. Third, they are reactive by construction: guidance follows demonstrated practice, so genuinely novel configurations (second-life maritime cells, mixed-chemistry retrofits) reach service before any instrument anticipates them. None of these features is a defect of any individual document; they are properties of a standards ecosystem that updates on a slower cycle than the technology it governs.
4.2. Scale and Retrofit as an Amplifier
The order-of-magnitude growth in installed energy characteristic of retrofit programmes [8] does not create new risk types but amplifies the consequences of the nine identified, and stresses the point-in-time approval model hardest. A larger pack means larger off-gas volumes and fault energy for the same architecture; a retrofit means new chemistry and new control are grafted onto an aged installation, activating SL2 and SL3 directly; and a retrofit approved as an alternative design inherits the full equivalency-verification burden of that route, now compounded by frontier risks for which no acceptance criterion exists. Scale is therefore best understood as a multiplier on criticality and not as a tenth risk.
4.3. Readiness Templates from Adjacent Regimes
The readiness gaps identified here are not intrinsic to the subject matter: for every high-priority gap, an adjacent regulated sector has already produced an instrument occupying the RRL 2–3 band that the maritime set lacks. For AI assurance, ISO/IEC TR 5469 classifies AI technologies and usage levels and maps them onto established functional-safety practice [39]; ISO/PAS 8800 extends the automotive functional-safety and safety-of-the-intended-functionality framework to AI elements, including safety-related performance requirements on trained models and their data [40]; UL 4600 structures the whole problem as an assurance case with explicit claims and evidence [41]; EASA’s concept paper sets out a learning-assurance lifecycle (data and model management, performance bounds within a declared operational domain, and explainability proportionate to the function’s criticality) that is already being used in certification projects [42]; and DNV’s recommended practice provides a sector-neutral assurance-case framework for AI-enabled systems [43]. The acceptance objectives proposed below for AI1 and AI3 adapt elements of these approaches to the maritime context, reducing the need to develop every assurance element from first principles.
The same is true of the other families. For digital-twin and sensor integrity (CP3), qualification practice for digital twins (defined validity domain, data-quality and sensor-assurance requirements, and monitoring of model–reality divergence) already exists in recommended-practice form [44]. For second-life qualification (SL1), UL 1974 defines a measurement-based process for sorting, grading and re-rating repurposed cells [46], and the EU Battery Regulation obliges battery passports carrying state-of-health and usage-history data for large batteries placed on the EU market [45]; a maritime acceptance pathway could require a UL-1974-type evaluation anchored to passport data continuity, and the same passport data provide a natural evidence stream for lifecycle revalidation (SL3). For the ship–shore interface (CP2), IEC/IEEE 80005-1 demonstrates what a verifiable interface standard looks like (protection coordination, interlocks, and communication requirements with conformance tests) for shore power [64]; the megawatt-charging work [63] is producing the corresponding connector, control and interoperability content for high-rate charging, and the remaining maritime task is to bind these to the ship-side battery safety case, not to write an interface standard from scratch.
Two caveats temper this otherwise encouraging picture. Adoption is not automatic: an adjacent instrument raises maritime readiness only when a maritime instrument invokes it (by incorporation into class rules, EMSA guidance or an IACS unified requirement) and when its acceptance criteria are re-anchored to maritime consequences (a ship cannot pull over). And the adjacent instruments are themselves mostly at RRL 3 in their home sectors: technical reports and publicly available specifications, not harmonised and testable acceptance criteria. They provide candidate building blocks for a maritime RRL 4 pathway, but they do not remove the need for maritime-specific acceptance criteria.
4.4. A Readiness-Raising Roadmap
For each of the nine risks, Table 12 states the acceptance objective that would satisfy the study’s RRL 4 definition, a verification method, and a candidate standardisation vehicle, drawing on the templates of Section 4.3. The objectives are deliberately framed as verifiable acceptance criteria instead of prescriptive designs, so that they are compatible with the alternative-design route rather than displacing it. The common structure (a defined threat or degradation model, a measurable acceptance threshold, and a stated assurance method) is what distinguishes RRL 4 from the generic clauses that currently leave these risks unverifiable.
4.5. Implication for Alternative-Design Acceptance
The practical consequence is that the SOLAS Reg. II-1/55 equivalence claim becomes harder to sustain exactly where next-generation features are present. An administration accepting an alternative design certifies that the installation provides a level of safety equivalent to the prescriptive requirements it displaces, on the basis of an engineering analysis whose performance criteria must be quantifiable and measurable [9,10,11]. That requirement is the hinge of the present argument. It is not open to an applicant to satisfy Reg. II-1/55 with a qualitative assurance that a design is as safe as a compliant one; the regulation requires criteria that can be measured, and measurement presupposes a stated threshold to measure against. For a conventional pack that claim rests on mature, if divergent, requirements [13]; for a system whose safety depends on an opaque AI estimator, a heterogeneous second-life pack, and a shore-coupled control network, several of the load-bearing risks sit at RRL 0–1, so there is no standardised currency in which equivalence can be expressed. This does not make approval impossible. The performance-based and alternative-design routes exist precisely to allow a case-specific safety argument where prescriptive requirements are absent, and an administration may accept such an argument on its own terms. What the absence of a standardised acceptance criterion removes is the common yardstick: the threshold is then set project by project, so the burden of constructing and defending it falls on the applicant and the approving party, the outcome depends on which administration or society is asked, and neither the applicant nor a third party can check the equivalence claim against anything external to the case itself. The consequence of low readiness is therefore not that equivalence cannot be demonstrated but that it cannot be demonstrated consistently or checkably, which is a weaker claim than the absence of any acceptance basis and is the one the analysis supports. Moving the highest-priority risks toward RRL 4 would provide a more standardised and comparable basis for alternative-design assessments as the technology advances.
A sharper version of the same difficulty appears where an instrument does speak determinately. Pack heterogeneity is the only frontier risk in the set to reach RRL 3, and it does so through a prohibition. For a design that complies, the prohibition is a complete answer; for a design that departs from it, which is what a large mixed-chemistry retrofit necessarily does, the prohibition supplies no criteria at all, because it was written to exclude the configuration instead of qualifying it. Readiness measured as determinacy and readiness measured as capacity to certify a next-generation design therefore come apart, and they come apart exactly where the alternative-design route is invoked. A regime can thus appear more ready than it is: the instrument gives a clear answer to a question the applicant is not asking.
This diagnosis complements, rather than replaces, the IMO’s existing evaluation machinery. A Formal Safety Assessment quantifies the risk of a particular design, and the goal-based and alternative-design routes test whether a design meets the prescriptive intent of the rules [10,76]; all three presuppose that a verifiable acceptance criterion exists against which a design can be judged. RRL measures that precondition directly: it locates the frontier risks for which no such criterion yet exists, and where a Formal Safety Assessment must therefore supply its own benchmark case by case instead of scoring against an agreed one. The claim is not that these methods cannot be applied to a low-readiness risk; it is that their outputs are not comparable between applications until an acceptance criterion is standardised, so each assessment re-litigates the threshold instead of applying it.
4.6. Worked Example: A Retrofit Alternative-Design Case
To show how the readiness profile bears on a concrete approval context, consider the ELECTRIC BLUE retrofit demonstration. The project concerns an existing RoPax ferry with a 5 MWh battery system and aims to substantially increase electric-only operating range through a larger onboard battery system and high-capacity shore charging. It also includes an AI-driven energy management system and digital tools for ship modelling [8]. The project is therefore a useful documented example of the scale-and-retrofit, networked-control, learning-enabled management and model-mediated features that motivate parts of the present taxonomy.
The example should not, however, be read as evidence that the demonstrator already contains all seven features F1–F7 or all nine frontier risks. The public project description does not establish, for example, that the demonstrator uses mixed-chemistry packs, second-life cells, or a machine- learning state estimator. Those features remain part of the broader taxonomy because they are plausible next-generation configurations supported by the literature, not because they are asserted here as features of the ELECTRIC BLUE vessel. This distinction is important for keeping the case study evidence-based.
For the documented project features, the readiness profile identifies several questions that an approval case would need to address explicitly. A networked or AI-driven energy-management function raises the AI2 and CP1/CP2 questions; model-mediated ship or battery decisions motivate the CP3 question; and the large retrofit scale increases the importance of lifecycle margin and degradation assumptions addressed by SL3. If the final design additionally incorporates second-life cells or mixed-chemistry or mixed-state-of-health packs, SL1 and SL2 become directly applicable and the corresponding roadmap objectives provide a structured basis for qualification. The case therefore illustrates how the RRL framework can be used as a screening tool before an alternative-design submission is assembled, while avoiding the stronger claim that every frontier risk is already present in the demonstrator.
The broader approval implication follows from the same logic as in Section 4.5. Where a safety-critical function depends on a frontier feature but the maritime instrument set provides no verifiable acceptance criterion, the applicant and administration must construct a case-specific performance basis. The RRL framework makes that missing basis visible before the approval process begins and identifies the acceptance objectives that could make the resulting safety argument more comparable across projects. A survey of installations already approved under alternative-design provisions would be a valuable empirical complement, but it is outside the scope of the present document-based assessment.
4.7. Limitations
Four limitations bound the conclusions. First, the RRL scores are document-based and single-assessor in origin; they are reported as an ordinal profile and a priority rank, not as precise measurements, and although the full cell-level rationale is published (Appendix A and Table S1 of the Supplementary Materials) so that the assessment can be independently re-scored, an independent second assessment, applying the decision rubric of Table 5, with agreement reported as a linearly weighted Cohen’s , is required before the profile is treated as definitive.
Second, the assessment reflects the instrument set in force on 31 July 2026. The IMO had adopted the non-mandatory MASS Code with effect from 1 July 2026 and had also advanced work on a dedicated safety framework addressing lithium-ion batteries and other new technologies [12,77]. These developments are relevant context but are not included in the scored matrix because the MASS Code is outside the battery-specific instrument scope and the battery safety framework remains under development. Future revisions of the instrument set could therefore change individual scores, which is why the method rather than the snapshot is the durable contribution.
Reg. II-1/55 does require the engineering analysis to be repeated and re-approved if the assumptions or operational restrictions stipulated in it change [9], which is a re-evaluation trigger of a kind; it is tied to changes in the stated assumptions and not to measured degradation of the installation, so it does not supply the threshold-based revalidation that SL3 lacks.
Third, the severity and exposure weights in Equation (2) are expert judgements, so the exact priority values should be read as indicative, not measured. The sensitivity analysis of Section 3.4 shows why the qualitative conclusion is comparatively stable: the bottom rank is shared by the two most-ready risks, AI1 and SL2, because both have a one-level gap, and the ordinal dominance check independently excludes CP2, CP3, and SL1 from the lead. The ordering among the remaining risks, including the position of SL1, depends on how the ordinal scales are combined and should not be over-interpreted.
Fourth, the analysis applies a uniform target of verifiable acceptance (RRL 4) to every risk. This choice is consequential, not neutral: if the target is made criticality-dependent, for instance for and for , the ranking does not simply compress, it reorders. Under that scheme AI3’s priority falls from 16 to 12 while CP1 and SL3 rise to share the leading band, and the Spearman correlation with the baseline ranking is . What survives the change is that the AI-driven risks remain at or near the top and SL2 remains at the bottom. The uniform target is adopted here because verifiable acceptance is what the alternative-design route requires of any load-bearing safety function irrespective of its criticality (Section 4.5), but a regulator who set differentiated targets would obtain a different order of attention, and the ranking should be read subject to that premise. Finally, the Monte Carlo of Section 3.4 probes independent rating error and not mis-scoring correlated across risks, which a single assessor’s systematic bias would produce. None of these limitations affects the central finding that readiness for the frontier risks is low, uneven, and concentrated in identifiable places.
Two further boundaries belong with these. The taxonomy is complete only with respect to an author-defined feature set, so its external validity rests on a reading of the literature instead of on independent validation; a structured expert elicitation or horizon scan could add features, and hence risks, that the present method cannot generate (Section 2.2). And the classification societies were sampled, not enumerated. Five were examined, on the reasoning that the IACS unified requirements bind all members and that an unexamined society could raise a row maximum only by supplying a battery-specific acceptance criterion the five lack. That remains an assumption about the societies not examined, and a single contrary rule would raise the affected score.
5. Conclusions
The guidance, rules and standards that govern shipboard batteries were largely developed around earlier generations of battery installations and are now being applied to configurations with additional digital, lifecycle, and scale-related features. This paper has reframed the resulting problem as one of regulatory readiness and made it measurable. The principal findings are as follows.
- 1.
- Readiness is low, uneven, and absent where it matters most. Across nine frontier risks in three threat families the median current readiness is RRL 2 and the distribution is one risk at RRL 0, two at RRL 1, four at RRL 2 and two at RRL 3. No risk reaches verifiable acceptance (RRL 4), the study’s target for fully verifiable acceptance, and one, the auditability of AI safety functions (AI3), is wholly unrecognised. The full cell-level basis for these scores is published in Appendix A and Table S1 of the Supplementary Materials, so the assessment can be independently re-scored.
- 2.
- The largest readiness gaps concentrate in the AI-driven family. The baseline criticality-weighted ranking places AI3 first, with AI2 and SL1 in the next priority band; the position of SL1 is sensitive to the aggregation rule. The ordinal dominance analysis, which makes no equal-spacing assumption, excludes CP2, CP3, and SL1 from the lead but does not establish a unique overall leader. Sensitivity analysis further shows that uncertainty in RRL classification has a larger effect on the ranking than uncertainty in severity and exposure.
- 3.
- The mandated assurance route cannot qualify a learned function. IACS UR E22 reaches battery and energy-management controllers as category III computer-based systems and subjects them to witnessed factory, system and integration testing. But the verification it prescribes is explicitly black-box, confined to behaviour observed from outside the system, and so cannot establish generalisation, uncertainty calibration or out-of-distribution behaviour. The regime does not merely lack an AI assurance route; the route it mandates can be satisfied without the question being asked. The position on state-of-health estimation is subtler and, in its way, more telling: independent verification of the value the battery management system reports is a classification requirement, discharged by an annual capacity test, so the output is checked once a year while the estimator that produces it the rest of the time is qualified not at all.
- 4.
- Where heterogeneity is addressed, it is foreclosed instead of governed. One of the two risks reaching RRL 3 does so because ABS prohibits cells of different chemistries and characteristics in the same electrical circuit. That is a determinate and checkable rule, but it supplies no criteria under which a mixed-chemistry or mixed-state-of-health pack could be accepted, so precisely the configuration that large retrofits require is excluded by rule instead of qualified, and is pushed onto the alternative-design route.
- 5.
- The missing acceptance criteria already have templates. Aviation (EASA learning assurance), road vehicles (ISO/PAS 8800), cross-sector functional safety (ISO/IEC TR 5469), assurance cases (UL 4600, DNV-RP-0671), digital-twin qualification (DNV-RP-A204) and the second-life instruments (UL 1974, the EU Battery Regulation’s passport and state-of-health data) show that verifiable acceptance of exactly these risk types is practised elsewhere. IACS UR E10 shows that the maritime regime itself knows how to write a verifiable acceptance criterion, pairing defined tests with stated pass criteria and a type-approval certificate. The fastest route to RRL 3–4 is incorporation by reference with maritime-specific acceptance criteria, for which Table 12 states concrete objectives, verification methods and candidate vehicles.
- 6.
- Readiness is a precondition for the alternative-design route. SOLAS Reg. II-1/55 requires the performance criteria supporting an equivalence claim to be quantifiable and measurable. Where load-bearing risks sit at RRL 0–1 such a claim can still be argued case by case, but not against a standardised criterion, so its threshold is set project by project and cannot be checked externally or compared across approvals. Moving the highest-priority risks toward verifiable acceptance would provide a more standardised and comparable basis for alternative-design assessments.
The wider implication concerns the approval mechanism itself. Alternative design is an important approval route when a proposed installation departs from prescriptive requirements. A route that certifies equivalence must be supported by stated performance criteria, and this analysis identifies the frontier risks for which the maritime instrument set does not yet provide the study’s fully verifiable acceptance basis. Closing those gaps would improve the consistency and comparability of alternative- design assessments as battery technology evolves. The method (scale, taxonomy, scoring rule, criticality weighting and sensitivity checks) is reproducible and transfers directly to future standards reviews as the instrument set evolves, and to any sector certifying an emerging technology against instruments written for an earlier generation.
Supplementary Materials
The following supporting information can be downloaded at the website of this paper posted on Preprints.org, Table S1, the complete instrument-level readiness record for all nine frontier risks across the thirteen scored instruments, stating for every one of the 117 instrument×risk cells the level assigned, the basis on which it was assigned, and the reason the next level was not reached.
Author Contributions
Conceptualization, S.R.; methodology, S.R.; formal analysis, S.R.; investigation, S.R.; visualization, S.R.; writing, original draft, S.R.; writing, review and editing, S.R., V.B. and P.K.; supervision, P.K. All authors have read and agreed to the published version of the manuscript.
Funding
This work has been carried out in the framework of the ELECTRIC BLUE project (HORIZON-CL5-2025-04-D5-11), funded by the European Union under the Horizon Europe research and innovation programme, Grant Agreement No. 101270307.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
No new experimental data were generated. The assessment is based on publicly available standards, classification rules, technical guidance documents, and published peer-reviewed literature cited herein. The complete scoring record underlying Table 8, Table 9 and Table 10 is published in full in Appendix A and in the Supplementary Materials. The scripts implementing the sensitivity analysis, the dominance check and the Monte Carlo of Section 3.4 are available from the authors on reasonable request.
Conflicts of Interest
The authors declare no conflict of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| AHJ | Authority having jurisdiction |
| AI | Artificial intelligence |
| BESS | Battery energy storage system |
| BMS | Battery management system |
| EMS | Energy management system |
| EMSA | European Maritime Safety Agency |
| IACS | International Association of Classification Societies |
| IEC | International Electrotechnical Commission |
| IMO | International Maritime Organization |
| LFP | Lithium iron phosphate |
| MASS | Maritime autonomous surface ship |
| MCS | Megawatt charging system |
| ML | Machine learning |
| NFPA | National Fire Protection Association |
| NMC | Nickel manganese cobalt |
| PMS | Power management system |
| RRL | Regulatory Readiness Level |
| SoH | State of health |
| SoX | State of `X’ (charge, health, energy, power, etc.) |
| SOLAS | International Convention for the Safety of Life at Sea |
| UL | Underwriters Laboratories |
| UR | Unified requirement (IACS) |
Appendix A. Scoring Rationale for the Row Maxima
Table A1 records, for each frontier risk, the instrument(s) attaining the row maximum in Table 8, the basis on which that level was assigned, and the reason the next level was not reached. Publishing this rationale is what makes the assessment contestable cell by cell: a reader who disagrees with a score can locate the exact judgement to challenge.
Table A1.
Per-risk scoring rationale for the current readiness (the row maxima of Table 8).
Table A1.
Per-risk scoring rationale for the current readiness (the row maxima of Table 8).
| ID | RRL | Basis of the Maximum | Why Not Higher |
|---|---|---|---|
| CP1 | 2 | IMO: Res. MSC.428(98) obliges cyber risks to be addressed in the safety-management system; MSC-FAL.1/Circ.3/Rev.2 elaborates a risk-management process. IACS: UR E26/E27 impose surveyable cyber-resilience requirements with an approved test procedure stating acceptance criteria, including deterministic output under attack; UR E22 adds witnessed FAT/SAT/SOST and fail-to-safe on data-link failure for category III systems. 62443: The standard specifies security capabilities and zone/conduit requirements, and is drawn upon by UR E27. | Generic to computer-based systems; no clause targets battery safety functions or states an acceptance criterion for them. All content is generic to computer-based systems; nothing targets charge-limit, protection-trip or gas-response functions, so RRL 3 is not reached. Generic to industrial control systems, with nothing specific to a safety-critical battery installation. |
| CP2 | 2 | ABS: The hybrid and all-electric guide sets electrical requirements for an offshore charging connection and for a charging station on the vessel. | Requirements address the electrical interface; the attack surface and the megawatt-rate duty are not covered, and no acceptance criterion is stated for the coupling as a safety-critical interface. |
| CP3 | 2 | BV: The type-approval test scope for a battery pack must include additional tests derived from an FMEA of sensor failures, and the HAZID must cover failure of voltage, temperature and gas sensors. LR: The battery management system is required to monitor the condition of cells, modules and packs continuously and to keep them within a specified safe operating region, with per-sensor alarms and safeguards; the FMEA must additionally consider hidden loss of a monitoring function. | The obligation addresses sensor failure, not drift, spoofing or the validity of a model-mediated estimate; and no instrument in the set addresses a digital twin credited in a safety decision, so RRL 3 is not reached. |
| AI1 | 3 | DNV: Independent verification of the state of health calculated by the BMS is required for ships relying on battery power for manoeuvring and propulsion, discharged in practice by annual capacity testing in which charge and discharge capacities are measured under rated conditions. The requirement is battery-specific, targets the estimate directly, and states a verification method. | RRL 4 is not reached because the verification is periodic and after the fact: it confirms the reported value at one point in the year without stating an accuracy bound, requiring quantification of uncertainty, or governing the estimator’s behaviour between verifications, which is when the estimate is relied upon for a safety decision. |
| AI2 | 1 | EMSA: The guidance acknowledges energy management and its influence on operation. IACS: UR E22 treats an energy-management controller as a category III computer-based system and requires black-box verification of intended functionality and robustness. BV: The rules note that computer-based systems may integrate machine-learning systems and point to the society’s guidance note on them. NK: The rules require EMS functions for monitoring capacity and controlling recharge and discharge. LR: Integration of a battery system into the ship’s electrical power system is governed by the hybrid electrical power systems rules. 61508: The standard governs safety functions carried out by electrical, electronic and programmable electronic systems, and the techniques it prescribes presuppose behaviour that can be specified and verified deterministically. | No requirement attaches to a controller that influences thermal or electrical stress. Black-box verification cannot bound the behaviour of an adaptive controller outside the cases tested; nothing addresses learning-based control. The guidance note may be referred to and not applied, so no obligation follows. Nothing attaches to the adaptive or learning character of the controller. Adaptive or learning-based control is neither addressed nor provided with a qualification route, so no obligation attaches to it. |
| AI3 | 0 | No instrument attains a level above 0. | No instrument in the scored set names, implies or acknowledges the explainability or auditability of AI safety functions; even RRL 1 is not reached. |
| SL1 | 1 | EMSA: The guidance names second-life batteries, states that they may pose different safety concerns related to ageing and wear, and records that best practices for their reuse are not consolidated. IACS: UR E18 requires procedures ensuring that replacement batteries are of an equivalent performance type. DNV: The rules require qualification of the cells used and acknowledge provenance. ABS: The requirements acknowledge cell qualification and documentation of cell origin. BV: Cells and packs require type approval against recognised standards, and replacement batteries must be of equivalent performance. NK: The rules require battery system approval and documentation of cell type. LR: Lithium battery systems must satisfy a stated type-testing specification, and the rules identify the cell chemistries to which they apply. 62619: The standard governs qualification of industrial lithium cells and acknowledges the role of provenance and manufacturing control. | It expressly states that it does not cover second-life batteries, so no requirement attaches. The requirement governs like-for-like replacement, not qualification of cells with a prior service history. Qualification presumes new cells of known provenance; no provision covers repurposed cells. No provision covers repurposed cells. No provision covers cells with a prior service history or uncertain residual state. It presumes new cells of known history and provides no route for repurposed cells. |
| SL2 | 3 | ABS: Battery cells of different chemistries and different physical and electrical characteristics are required not to be used in the same electrical circuit. The requirement is battery-specific and determinate: an independent party can establish compliance by inspection of the design. | RRL 4 is not reached because the requirement resolves the risk by prohibition rather than by stating acceptance criteria: it excludes the configuration instead of defining the evidence under which a heterogeneous pack could be qualified, and states no verification method for that case because the case is disallowed. |
| SL3 | 2 | IACS: UR E18 requires a battery schedule recording maintenance and replacement cycle dates and the dates of last maintenance, requires replacements of equivalent performance type, and places the schedule in the safety-management system for verification by the surveyor. BV: NR467 requires the battery schedule, equivalent-performance replacement and integration into the safety-management system for verification by the surveyor, and requires the BMS to monitor and control state of health. NK: The rules require capacity and state-of-health monitoring functions and periodic survey. | The regime governs maintenance and replacement scheduling, not revalidation of the safety margin: no state-of-health or calendar threshold triggers re-testing, and no criteria to retire or de-rate are stated, so RRL 3 is not reached. The obligations govern maintenance scheduling and monitoring, not revalidation of the demonstrated safety margin against a defined threshold. No threshold-based revalidation of the safety margin is required. |
References
- International Maritime Organization. 2023 IMO Strategy on Reduction of GHG Emissions from Ships; Resolution MEPC.377(80), MEPC 80/17/Add.1, Annex 15; IMO: London, UK, 2023.
- European Parliament and Council of the European Union. Regulation (EU) 2023/1805 of 13 September 2023 on the use of renewable and low-carbon fuels in maritime transport (FuelEU Maritime). Off. J. Eur. Union 2023, L 234, 48–100.
- European Maritime Safety Agency. Guidance on the Safety of Battery Energy Storage Systems (BESS) On-board Ships; EMSA: Lisbon, Portugal, 2023.
- American Bureau of Shipping. Requirements for Use of Lithium-Ion Batteries in the Marine and Offshore Industries; ABS: Spring, TX, USA, 2024.
- DNV. Rules for Classification: Ships, Part 6 Additional Class Notations, Chapter 2 Propulsion, Power Generation and Auxiliary Systems, Section 1 Battery Power; DNV: Høvik, Norway, 2024.
- DNV. Technical Reference for Li-Ion Battery Explosion Risk and Fire Suppression; Report No. 2019-1025, Rev. 4; DNV: Høvik, Norway, 2021.
- International Electrotechnical Commission. IEC 62619:2022 Secondary Cells and Batteries Containing Alkaline or Other Non-Acid Electrolytes, Safety Requirements for Secondary Lithium Cells and Batteries, for Use in Industrial Applications, 2nd ed.; IEC: Geneva, Switzerland, 2022.
- ELECTRIC BLUE Consortium. ELECTRIC BLUE, Demonstration of Battery Energy Storage Systems in Existing and New Vessels via Novel Energy Storage and Ship Design Concepts; Project under Horizon Europe call HORIZON-CL5-2025-04-D5-11 (ZEWT Partnership); European Commission: Brussels, Belgium, 2025.
- International Maritime Organization. SOLAS: International Convention for the Safety of Life at Sea, consolidated ed.; IMO: London, UK, 2022.
- International Maritime Organization. MSC.1/Circ.1455: Guidelines for the Approval of Alternatives and Equivalents as Provided for in Various IMO Instruments; IMO: London, UK, 2013.
- International Maritime Organization. MSC.1/Circ.1212: Guidelines on Alternative Design and Arrangements for SOLAS Chapters II-1 and III; IMO: London, UK, 2006.
- International Maritime Organization. Draft workplan agreed on safety rules for battery, wind and nuclear-powered ships; IMO: London, UK, 29 January 2026.
- Rahimpour, S.; Bratkov, V.; Kujala, P. Compliant but not equivalent: Divergence in lithium-ion battery safety requirements across maritime regulatory frameworks. 2026, manuscript under review.
- Feng, X.; Ouyang, M.; Liu, X.; Lu, L.; Xia, Y.; He, X. Thermal runaway mechanism of lithium ion battery for electric vehicles: A review. Energy Storage Mater. 2018, 10, 246–267. [CrossRef]
- Ren, D.; Feng, X.; Liu, L.; Hsu, H.; Lu, L.; Wang, L.; He, X.; Ouyang, M. Investigating the relationship between internal short circuit and thermal runaway of lithium-ion batteries under thermal abuse condition. Energy Storage Mater. 2021, 34, 563–573. [CrossRef]
- Larsson, F.; Andersson, P.; Blomqvist, P.; Mellander, B.-E. Toxic fluoride gas emissions from lithium-ion battery fires. Sci. Rep. 2017, 7, 10018. [CrossRef]
- Diaz, L.B.; He, X.; Hu, Z.; Restuccia, F.; Marinescu, M.; Barreras, J.V.; Patel, Y.; Offer, G.; Rein, G. Meta-review of fire safety of lithium-ion batteries: Industry challenges and research contributions. J. Electrochem. Soc. 2020, 167, 090559. [CrossRef]
- Madusanka, N.S.; Fan, Y.; Yang, S.; Xiang, X. Digital twin in the maritime domain: A review and emerging trends. J. Mar. Sci. Eng. 2023, 11, 1021. [CrossRef]
- Kabir, M.R.; Halder, D.; Ray, S. Digital twins for IoT-driven energy systems: A survey. IEEE Access 2024, 12, 177123–177143. [CrossRef]
- Rahimpour, S.; Shahin, M.; Gülmez, Y.; Bauk, S. Data mining and cybersecurity-driven solutions for CO2 emissions reduction of different maritime shipping: A multi-faceted analysis. In Proceedings of the Tenth International Congress on Information and Communication Technology (ICICT 2025); Springer: Singapore, 2025; pp. 471–483. [CrossRef]
- Rahimpour, S.; Shahin, M.; Bauk, S. Cybersecurity of power electronic converters in maritime systems: Threat landscape, modeling, and resilient mitigation strategies. In Proceedings of the 10th International Conference on Smart and Sustainable Technologies (SpliTech 2025); IEEE: 2025. [CrossRef]
- Severson, K.A.; Attia, P.M.; Jin, N.; Perkins, N.; Jiang, B.; Yang, Z.; Chen, M.H.; Aykol, M.; Herring, P.K.; Fraggedakis, D.; et al. Data-driven prediction of battery cycle life before capacity degradation. Nat. Energy 2019, 4, 383–391. [CrossRef]
- Guo, X.; Lang, X.; Yuan, Y.; Tong, L.; Shen, B.; Long, T.; Mao, W. Energy management system for hybrid ship: Status and perspectives. Ocean Eng. 2024, 310, 118638. [CrossRef]
- Mylonopoulos, F.; Polinder, H.; Coraddu, A. A comprehensive review of modeling and optimization methods for ship energy systems. IEEE Access 2023, 11, 32697–32707. [CrossRef]
- Swamy, D.; Amin, T.; Kammoun, M.; Raymond, D.; Tomdio, J.; Wang, J.; Gunda, H.; Vaddiraju, S.; Khan, F. Are next-generation batteries ready for marine and offshore electrification? A safety perspective. In Proceedings of the Offshore Technology Conference, Houston, TX, USA, 5–8 May 2025; Paper OTC-35673-MS. [CrossRef]
- Martinez-Laserna, E.; Gandiaga, I.; Sarasketa-Zabala, E.; Badeda, J.; Stroe, D.-I.; Swierczynski, M.; Goikoetxea, A. Battery second life: Hype, hope or reality? A critical review of the state of the art. Renew. Sustain. Energy Rev. 2018, 93, 701–718. [CrossRef]
- International Maritime Organization. Resolution MSC.428(98): Maritime Cyber Risk Management in Safety Management Systems; IMO: London, UK, 2017.
- International Association of Classification Societies. UR E26 Rev.1: Cyber Resilience of Ships; IACS: London, UK, 2023; applicable to ships contracted for construction on or after 1 July 2024.
- International Association of Classification Societies. UR E27 Rev.1: Cyber Resilience of On-board Systems and Equipment; IACS: London, UK, 2023; applicable to ships contracted for construction on or after 1 July 2024.
- Bureau Veritas Marine & Offshore. Maritime Electrification: Maritime Battery Systems and Onshore Power Supply; Technology Report; Bureau Veritas Marine & Offshore: Paris, France, 2025.
- Zhou, R.; Yang, L.; Fan, A.; Liu, Q.; Wang, L.; Yang, J.; Vladimir, N. Systematic review of battery electric ship safety: Risk factors, assessment methods, and preventive measures. Int. J. Nav. Archit. Ocean Eng. 2025, 17, 100710. [CrossRef]
- Yin, R.; Du, M.; Shi, F.; Cao, Z.; Wu, W.; Shi, H.; Zheng, Q. Risk analysis for marine transport and power applications of lithium ion batteries: A review. Process Saf. Environ. Prot. 2024, 181, 266–293. [CrossRef]
- Lucà Trombetta, G.; Leonardi, S.G.; Aloisio, D.; Andaloro, L.; Sergi, F. Lithium-ion batteries on board: A review on their integration for enabling the energy transition in shipping industry. Energies 2024, 17, 1019. [CrossRef]
- Yılmaz, F. Safety of electrical (battery-powered) ships: An overview of IMO’s safety regulations and class rules/requirements. Int. J. New Find. Eng. Sci. Technol. 2025, Special Issue, 30–38. [CrossRef]
- Corsi, P.; Jakovlev, S.; Figari, M.; Djackov, V. Analysis and definition of certification requirements for maritime autonomous surface ship operation. J. Mar. Sci. Eng. 2025, 13, 751. [CrossRef]
- Lee, C.; Lee, S. A risk identification method for ensuring AI-integrated system safety for remotely controlled ships with onboard seafarers. J. Mar. Sci. Eng. 2024, 12, 1778. [CrossRef]
- Aaslund, E.; Wang, S.; Knutsen, K.E. A comprehensive literature review of state of safety (SoS) for maritime battery management systems (BMSs). In Proceedings of the European Conference of the Prognostics and Health Management Society, 2026; Group Research and Development, DNV AS: Oslo, Norway.
- Rahimpour, S.; Kujala, P. Dynamic risk assessment and safety-case framework for shipboard lithium-ion battery systems. Available at SSRN, 2026. [CrossRef]
- International Organization for Standardization; International Electrotechnical Commission. ISO/IEC TR 5469:2024 Artificial Intelligence, Functional Safety and AI Systems; ISO/IEC: Geneva, Switzerland, 2024.
- International Organization for Standardization. ISO/PAS 8800:2024 Road Vehicles, Safety and Artificial Intelligence; ISO: Geneva, Switzerland, 2024.
- Underwriters Laboratories. UL 4600: Standard for Evaluation of Autonomous Products, 3rd ed.; UL Standards & Engagement: Northbrook, IL, USA, 2023.
- European Union Aviation Safety Agency. EASA Artificial Intelligence Concept Paper Issue 02: Guidance for Level 1 & 2 Machine Learning Applications; EASA: Cologne, Germany, 2024.
- DNV. DNV-RP-0671: Assurance of AI-Enabled Systems; DNV: Høvik, Norway, 2023.
- DNV. DNV-RP-A204: Assurance of Digital Twins; DNV: Høvik, Norway, 2020 (rev. 2023).
- European Parliament and Council of the European Union. Regulation (EU) 2023/1542 of 12 July 2023 concerning batteries and waste batteries, amending Directive 2008/98/EC and Regulation (EU) 2019/1020 and repealing Directive 2006/66/EC. Off. J. Eur. Union 2023, L 191, 1–117.
- Underwriters Laboratories. UL 1974: Standard for Evaluation for Repurposing Batteries, 1st ed.; UL: Northbrook, IL, USA, 2018.
- International Maritime Organization. MSC-FAL.1/Circ.3/Rev.2: Guidelines on Maritime Cyber Risk Management; IMO: London, UK, 2022.
- International Association of Classification Societies. UR E10 Rev.10: Test Specification for Type Approval; IACS: London, UK, 2024.
- International Association of Classification Societies. UR E11 Rev.4: Unified Requirements for Systems with Voltages above 1 kV up to 15 kV; IACS: London, UK, 2021.
- International Association of Classification Societies. UR E18 Rev.2: Recording of the Type, Location and Maintenance Cycle of Batteries; IACS: London, UK, 2025; applicable to ships contracted for construction on or after 1 July 2026.
- International Association of Classification Societies. UR E21 Rev.2: Requirements for Uninterruptible Power System (UPS) Units; IACS: London, UK, 2024.
- International Association of Classification Societies. UR E22 Rev.3 Corr.1: Computer-Based Systems; IACS: London, UK, 2025.
- DNV GL. Handbook for Maritime and Offshore Battery Systems; DNV GL: Høvik, Norway, 2016.
- American Bureau of Shipping. Guide for Hybrid and All-Electric Power Systems for Marine and Offshore Applications; ABS: Spring, TX, USA, 2025.
- Bureau Veritas. NR467: Rules for the Classification of Steel Ships, Part C, Chapter 2, Section 7: Battery Energy Storage Systems and Chargers, consolidated ed.; Bureau Veritas: Paris, France, July 2026.
- Nippon Kaiji Kyokai (ClassNK). Rules for the Survey and Construction of Steel Ships, Part D (Machinery Installations), Part H (Electrical Installations) and Part R (Fire Protection, Detection and Extinction); ClassNK: Tokyo, Japan, 2026.
- Lloyd’s Register. LR-RU-001: Rules and Regulations for the Classification of Ships, Part 6, Chapter 2, Section 12: Batteries; Lloyd’s Register: London, UK, July 2026.
- International Electrotechnical Commission. IEC 60092 Series: Electrical Installations in Ships; IEC: Geneva, Switzerland, various dates.
- International Electrotechnical Commission. IEC 61508: Functional Safety of Electrical/Electronic/Programmable Electronic Safety-Related Systems, 2nd ed.; IEC: Geneva, Switzerland, 2010.
- International Electrotechnical Commission. IEC 62443 Series: Security for Industrial Automation and Control Systems; IEC: Geneva, Switzerland, various dates.
- Underwriters Laboratories. UL 9540A: Test Method for Evaluating Thermal Runaway Fire Propagation in Battery Energy Storage Systems, 4th ed.; UL: Northbrook, IL, USA, 2019.
- National Fire Protection Association. NFPA 855: Standard for the Installation of Stationary Energy Storage Systems, 2023 ed.; NFPA: Quincy, MA, USA, 2023.
- SAE International. SAE J3271_202503: Megawatt Charging System for Electric Vehicles; Technical Information Report; SAE International: Warrendale, PA, USA, 2025.
- International Electrotechnical Commission; Institute of Electrical and Electronics Engineers. IEC/IEEE 80005-1:2019 Utility Connections in Port, Part 1: High Voltage Shore Connection (HVSC) Systems, General Requirements; IEC: Geneva, Switzerland, 2019.
- Leveson, N.G. Engineering a Safer World: Systems Thinking Applied to Safety; MIT Press: Cambridge, MA, USA, 2011.
- Leveson, N.G.; Thomas, J.P. STPA Handbook; MIT: Cambridge, MA, USA, 2018.
- Mankins, J.C. Technology Readiness Levels: A White Paper; Advanced Concepts Office, Office of Space Access and Technology, NASA: Washington, DC, USA, 1995.
- Mankins, J.C. Technology readiness assessments: A retrospective. Acta Astronaut. 2009, 65, 1216–1223. [CrossRef]
- Kobos, P.H.; Malczynski, L.A.; Walker, L.T.N.; Borns, D.J.; Klise, G.T. Timing is everything: A technology transition framework for regulatory and market readiness levels. Technol. Forecast. Soc. Change 2018, 137, 211–225. [CrossRef]
- Vik, J.; Melås, A.M.; Stræte, E.P.; Søraa, R.A. Balanced readiness level assessment (BRLa): A tool for exploring new and emerging technologies. Technol. Forecast. Soc. Change 2021, 169, 120854. [CrossRef]
- Lowe, D.C.; Justham, L.; Everitt, M.J. Multi-index analysis with readiness levels for decision support in product design. Technol. Forecast. Soc. Change 2024, 206, 123559. [CrossRef]
- Olechowski, A.; Eppinger, S.D.; Joglekar, N. Technology readiness levels at 40: A study of state-of-the-art use, challenges, and opportunities. In Proceedings of PICMET ’15: Management of the Technology Age, Portland, OR, USA, 2–6 August 2015; pp. 2084–2094.
- DNV GL. McMicken Battery Energy Storage System Event Technical Analysis and Recommendations; Report prepared for Arizona Public Service; DNV GL: Chalfont, PA, USA, 2020.
- Vanem, E.; Liang, Q.; Bruch, M.; Bakdi, A.; Alnes, .Å. Data-informed state of health estimation for maritime lithium-ion battery systems using an ensemble of simple linear models. J. Mar. Eng. Technol. 2026, 25, 73–89. [CrossRef]
- Bureau Veritas. NI692: Guidelines for Machine Learning Systems; Bureau Veritas: Paris, France, 2023.
- International Maritime Organization. MSC-MEPC.2/Circ.12/Rev.2: Revised Guidelines for Formal Safety Assessment (FSA) for Use in the IMO Rule-Making Process; IMO: London, UK, 2018.
- International Maritime Organization. Resolution MSC.595(111): International Code of Safety for Maritime Autonomous Surface Ships (MASS Code); IMO: London, UK, 2026.
Figure 1.
Threat-led readiness assessment workflow. The unit of analysis is the emerging risk, not the framework; the output is a criticality-weighted readiness gap, its robustness, and a set of acceptance objectives.
Figure 1.
Threat-led readiness assessment workflow. The unit of analysis is the emerging risk, not the framework; the output is a criticality-weighted readiness gap, its robustness, and a set of acceptance objectives.

Figure 2.
Frontier-risk taxonomy for next-generation maritime battery systems: three threat families, each with three sub-risks. Scale and retrofit (F1) act as a cross-cutting amplifier on all nine.
Figure 2.
Frontier-risk taxonomy for next-generation maritime battery systems: three threat families, each with three sub-risks. Scale and retrofit (F1) act as a cross-cutting amplifier on all nine.

Figure 3.
The RRL ladder. The two italicised thresholds, acknowledgement to control, and control to verifiable acceptance, are where readiness is most often lost.
Figure 3.
The RRL ladder. The two italicised thresholds, acknowledgement to control, and control to verifiable acceptance, are where readiness is most often lost.

Figure 4.
Current readiness of each frontier risk against the RRL 4 target. The coloured bar is , the grey remainder the readiness gap ; each row is annotated with the criticality-weighted priority . Readiness gap and baseline priority are strongly aligned in this small set: AI3 has the largest gap and highest priority, while the two most-ready risks, AI1 and SL2, share the lowest priority band.
Figure 4.
Current readiness of each frontier risk against the RRL 4 target. The coloured bar is , the grey remainder the readiness gap ; each row is annotated with the criticality-weighted priority . Readiness gap and baseline priority are strongly aligned in this small set: AI3 has the largest gap and highest priority, while the two most-ready risks, AI1 and SL2, share the lowest priority band.

Figure 5.
Readiness gap versus criticality. Bubble area is scaled approximately in proportion to the baseline priority . Small display offsets separate observations with identical coordinates and are used only for visualisation; the underlying readiness-gap and criticality values are unchanged.
Figure 5.
Readiness gap versus criticality. Bubble area is scaled approximately in proportion to the baseline priority . Small display offsets separate observations with identical coordinates and are used only for visualisation; the underlying readiness-gap and criticality values are unchanged.

Table 1.
Inclusion and exclusion decisions for every candidate instrument considered, against the eligibility conditions of Section 2.1.
Table 1.
Inclusion and exclusion decisions for every candidate instrument considered, against the eligibility conditions of Section 2.1.
| Candidate | Decision | Ground |
|---|---|---|
| SOLAS; MSC.1/Circ.1455 | scored | Maritime approval instruments; define the alternative-design route. |
| Res. MSC.428(98); MSC-FAL.1/Circ.3 | scored | Maritime instruments imposing cyber-risk-management obligations. |
| EMSA BESS guidance | scored | Maritime guidance applied by administrations to BESS installations. |
| IACS UR E26 / E27 | scored | Maritime unified requirements, in force for ships contracted from 1 July 2024. |
| IACS UR E10, E11, E18, E21, E22 | scored | Maritime unified requirements governing type-approval testing, high-voltage systems, battery maintenance schedules, UPS units and computer-based systems. |
| DNV battery rules, technical reference and handbook | scored | Maritime class requirements and guidance for battery installations. |
| ABS battery requirements; ABS hybrid/all-electric guide | scored | Maritime class requirements for battery installations and for charging connections. |
| Bureau Veritas NR467 Pt C, Ch 2, Sec 7 | scored | Maritime class requirements dedicated to battery energy storage systems and chargers. |
| ClassNK Rules Parts D, H, R | scored | Maritime class requirements for machinery, electrical installations and fire protection, including accumulator battery systems. |
| Lloyd’s Register LR-RU-001 Pt 6, Ch 2, Sec 12 | scored | Maritime class requirements dedicated to permanently installed batteries, including lithium battery systems. |
| IEC 62619; IEC 60092 | scored | IEC 62619 specifies requirements and tests for the safe operation of secondary lithium cells and batteries in industrial applications; IEC 60092-201 covers the main features of system design of electrical installations in ships. Both are normatively referenced by class battery rules. |
| IEC 61508 | scored | Functional-safety basis drawn upon by maritime control-system requirements. |
| IEC 62443 | scored | Security capabilities drawn upon by IACS UR E27; Part 4-1 specifies process requirements for a secure development life-cycle for products used in industrial automation and control systems. |
| UL 9540A; NFPA 855 | scored | Propagation test method and installation standard referenced in maritime battery guidance and class practice. |
| IEC 63379 (MCS); SAE J3271 | not scored | Unpublished project and non-maritime technical information report respectively; no published text to score. |
| IMO MASS Code | not scored | Adopted as a non-mandatory code in July 2026 and outside the battery-specific scored instrument set; considered as contextual evidence only. |
| IEC/IEEE 80005-1 | reviewed, not scored | Published and maritime, but its scope is high-voltage shore power, not high-rate charging of propulsion batteries; informs CP2 without governing it. |
| ISO/IEC TR 5469; ISO/PAS 8800; UL 4600; EASA AI concept paper | not scored | Not maritime approval instruments and not invoked by one; treated as templates in Section 4.3. |
| DNV-RP-0671; DNV-RP-A204 | not scored | Sector-neutral recommended practices not invoked by any maritime approval instrument. |
| BV NI692 (machine learning systems) | reviewed, not scored | A society guidance note that the rules say may be referred to, so no obligation attaches; it is recorded as evidence at RRL 1 rather than scored as a requirement. |
| UL 1974; Regulation (EU) 2023/1542 | not scored | Repurposing standard and product regulation respectively; neither is a maritime approval instrument. |
Table 2.
Frontier-risk taxonomy. “Why new” states the assumption of first-generation guidance that the risk violates.
Table 2.
Frontier-risk taxonomy. “Why new” states the assumption of first-generation guidance that the risk violates.
| ID | Frontier Risk | Why Current Guidance Is Not Built for It |
|---|---|---|
| Family I: Cyber-physical | ||
| CP1 | Compromise of safety-critical BMS/EMS/PMS | Protection logic, charge limits and fault annunciation are assumed local, trusted and air-gapped. |
| CP2 | Ship–shore charging interface | Megawatt, bidirectional shore coupling creates a fault and attack surface absent in self-contained packs. |
| CP3 | Digital-twin and sensor integrity | Condition-based safety decisions assume faithful sensing and an unmanipulated model; drift and spoofing are not addressed. |
| Family II: AI-driven | ||
| AI1 | Assurance of ML-based SoX/SoH estimation | Safety decisions presuppose deterministic, validated estimators; data-driven estimators have no qualification basis. |
| AI2 | Learning-based EMS/control in the safety loop | Adaptive control that influences thermal and electrical stress is outside functional-safety certification practice. |
| AI3 | Explainability and auditability of AI safety functions | Class approval assumes inspectable, traceable logic; opaque models cannot be audited by present methods. |
| Family III: Second-life and lifecycle | ||
| SL1 | Qualification of second-life / repurposed cells | Cell qualification assumes new cells of known provenance; residual-SoH uncertainty has no maritime pathway. |
| SL2 | Mixed-chemistry and mixed-SoH packs | Rules assume homogeneous packs; differential aging, balancing and heterogeneous propagation are uncharacterised. |
| SL3 | Lifecycle safety-margin revalidation | Type approval is point-in-time at commissioning; degradation of safety margin over service life is not revalidated. |
Table 3.
Traceability of the nine frontier risks to the defining features of next-generation systems (Section 1.1). • = direct driver; ∘ = amplifier. F1: scale and retrofit; F2: chemistry/SoH heterogeneity; F3: networked, shore-coupled control; F4: ML state estimation; F5: learning-based/autonomous EMS; F6: digital twin in safety decisions; F7: second-life cells.
Table 3.
Traceability of the nine frontier risks to the defining features of next-generation systems (Section 1.1). • = direct driver; ∘ = amplifier. F1: scale and retrofit; F2: chemistry/SoH heterogeneity; F3: networked, shore-coupled control; F4: ML state estimation; F5: learning-based/autonomous EMS; F6: digital twin in safety decisions; F7: second-life cells.
| Risk | F1 | F2 | F3 | F4 | F5 | F6 | F7 |
|---|---|---|---|---|---|---|---|
| CP1 | ∘ | – | • | – | • | – | – |
| CP2 | ∘ | – | • | – | – | – | – |
| CP3 | ∘ | – | • | – | – | • | – |
| AI1 | ∘ | – | – | • | – | • | – |
| AI2 | ∘ | – | – | – | • | – | – |
| AI3 | ∘ | – | – | • | • | – | – |
| SL1 | ∘ | – | – | – | – | – | • |
| SL2 | • | • | – | – | – | – | • |
| SL3 | • | • | – | – | – | – | • |
Table 4.
The Regulatory Readiness Level (RRL) scale.
| RRL | State | Definition |
|---|---|---|
| 0 | Unrecognised | No instrument names or implies the risk. |
| 1 | Acknowledged | The risk is mentioned in guidance or referenced literature but no requirement attaches to it. |
| 2 | Generic clause | Covered only by a general goal-/risk-based clause (e.g. “address cyber risks”), with no content specific to safety-critical battery systems. |
| 3 | Specific requirement | A dedicated requirement or referenced standard exists, but acceptance is qualitative or deferred to the authority having jurisdiction. |
| 4 | Verifiable acceptance | A specific requirement plus a defined, testable acceptance criterion and a stated assurance/verification method. |
Table 5.
Operational decision rules for assigning each RRL level. A level is assigned only when its evidence test is satisfied and the test for the next level is not.
Table 5.
Operational decision rules for assigning each RRL level. A level is assigned only when its evidence test is satisfied and the test for the next level is not.
| RRL | State | Evidence Test (assign iff…) |
|---|---|---|
| 0 | Unrecognised | no instrument in the set names, implies, or can be read to cover the risk. |
| 1 | Acknowledged | at least one instrument in the assessed set mentions the risk, but no obligation attaches: no “shall” follows from the mention. Supporting literature may inform interpretation, but does not raise the RRL unless the instrument itself acknowledges the risk. |
| 2 | Generic clause | a goal- or risk-based obligation applies to the risk, but its content is generic: it would read identically for any computer-based or industrial system, with nothing specific to a safety-critical battery installation. |
| 3 | Specific requirement | a dedicated requirement or referenced standard targets the risk, but acceptance is qualitative or deferred to the authority having jurisdiction; no measurable pass/fail criterion is stated. |
| 4 | Verifiable acceptance | a specific requirement is paired with a defined, testable acceptance criterion and a stated assurance/verification method, such that an independent party could determine pass or fail. |
Table 6.
Level descriptors for the severity and exposure scales. Severity is the worst credible consequence of the risk being realised on a next-generation installation; exposure is how intrinsic the risk is to such installations.
Table 6.
Level descriptors for the severity and exposure scales. Severity is the worst credible consequence of the risk being realised on a next-generation installation; exposure is how intrinsic the risk is to such installations.
| Severity S: worst credible consequence | Exposure E: prevalence in next-generation installations | |
|---|---|---|
| 1 | Localised, recoverable fault; no thermal event; normal operation restored by the installed protection. | Peripheral: arises only in unusual configurations. |
| 2 | Equipment damage or loss of the battery system as an asset, without thermal runaway or danger to persons. | Occasional: present in a minority of installations. |
| 3 | Thermal runaway confined to a module, with local fire or off-gassing controlled by the installed barriers; no loss of propulsion. | Common: present in many but not most installations. |
| 4 | Propagation beyond a module, or loss of propulsion or power, or off-gas release into an occupied or confined space; casualty credible but survivable. | Prevalent: present in most installations exhibiting the defining features. |
| 5 | Credible path to loss of the ship or loss of life: uncontrolled propagation, deflagration in a confined space, or failure of a safety function credited in the approval. | Intrinsic: present in essentially every installation exhibiting the defining features. |
Table 7.
Rationale for the severity and exposure rating of each frontier risk, against the descriptors of Table 6.
Table 7.
Rationale for the severity and exposure rating of each frontier risk, against the descriptors of Table 6.
| ID | S | Severity rationale | E | Exposure rationale |
|---|---|---|---|---|
| CP1 | 5 | Defeat of charge-limit, protection-trip or gas-response logic disables a safety function credited in the approval, giving a credible path to uncontrolled propagation. | 4 | Networked, shore-coupled control (F3) is present in most next-generation installations, though not all expose the safety loop equally. |
| CP2 | 4 | A fault or intrusion at a megawatt charging interface can drive abusive charging or disable shore-side protection; casualty credible, but the interface is energised only intermittently. | 4 | Applies to every shore-charged installation (F3), which is most but not all of the class. |
| CP3 | 4 | A drifting or manipulated twin causes a wrong safety decision (missed derating, false healthy state); consequence is mediated by other barriers rather than immediate. | 4 | Digital twins in safety decisions (F6) are prevalent in new designs but not yet universal. |
| AI1 | 5 | SoX/SoH estimates gate charge limits and derating; a confidently wrong estimate removes the margin the approval assumes, with a direct path to runaway. | 5 | ML-based state estimation (F4) is intrinsic to next-generation battery management. |
| AI2 | 4 | An adaptive energy manager influences thermal and electrical stress continuously; consequences accrue through accelerated ageing and stress excursions rather than a single failure. | 4 | Learning-based EMS (F5) is prevalent in hybrid installations but not universal. |
| AI3 | 3 | Non-auditability is an assurance deficit rather than a failure mode: it does not itself cause a thermal event, but it prevents detection of the deficits that do. | 4 | Applies wherever an AI function is credited in a safety role (F4, F5), hence to most such installations. |
| SL1 | 4 | A mis-graded repurposed cell carries unquantified residual defects; propagation is credible, but qualification and screening act as intervening barriers. | 3 | Second-life cells (F7) are proposed and piloted rather than yet common. |
| SL2 | 5 | Cross-chemistry and cross-SoH propagation behaviour is uncharacterised, and differential ageing can defeat balancing, giving a credible path to uncontrolled propagation. | 5 | Chemistry and SoH heterogeneity (F2) is intrinsic to retrofits, which is the dominant next-generation pathway. |
| SL3 | 5 | A safety margin that has silently decayed below its commissioning value leaves the installed barriers under-sized against the hazard they were approved for. | 4 | Margin decay (F1, F2, F7) affects every ageing installation, though its rate varies with duty. |
Table 8.
Instrument-level readiness for each frontier risk (rows) across the scored instruments (columns). The right-hand column is the row maximum (Equation (1)); the documents covered by each column are listed in Section 2.1. Shading: grey = 0, amber = 1, green = 2, dark green = 3.
Table 8.
Instrument-level readiness for each frontier risk (rows) across the scored instruments (columns). The right-hand column is the row maximum (Equation (1)); the documents covered by each column are listed in Section 2.1. Shading: grey = 0, amber = 1, green = 2, dark green = 3.
![]() |
Table 9.
Readiness assessment. S severity, E exposure, C criticality (Equation (2)), current readiness (Equation (1)), G gap, P priority (Equation (3)), Rk rank (ties share a band). Values are from the document-based assessment described in Section 2.4; the single-assessor limitation is discussed in Section 4.7.
Table 9.
Readiness assessment. S severity, E exposure, C criticality (Equation (2)), current readiness (Equation (1)), G gap, P priority (Equation (3)), Rk rank (ties share a band). Values are from the document-based assessment described in Section 2.4; the single-assessor limitation is discussed in Section 4.7.
| ID | Governing Instrument (Current) | S | E | C | RRL | G | P | Rk |
|---|---|---|---|---|---|---|---|---|
| CP1 | MSC.428(98); IACS UR E26/E27; IEC 62443 | 5 | 4 | 5 | 2 | 2 | 10 | 4 |
| CP2 | ABS offshore charging connection | 4 | 4 | 4 | 2 | 2 | 8 | 6 |
| CP3 | BV NR467 (sensor-failure FMEA tests) | 4 | 4 | 4 | 2 | 2 | 8 | 6 |
| AI1 | DNV (independent SOH verification) | 5 | 5 | 5 | 3 | 1 | 5 | 8 |
| AI2 | None (IEC 61508 partial) | 4 | 4 | 4 | 1 | 3 | 12 | 2 |
| AI3 | None | 3 | 4 | 4 | 0 | 4 | 16 | 1 |
| SL1 | None maritime (EMSA expressly excludes) | 4 | 3 | 4 | 1 | 3 | 12 | 2 |
| SL2 | ABS (mixed chemistries prohibited) | 5 | 5 | 5 | 3 | 1 | 5 | 8 |
| SL3 | IACS UR E18; BV NR467 (battery schedule) | 5 | 4 | 5 | 2 | 2 | 10 | 4 |
Table 10.
Sensitivity of the priority ranking to the criticality aggregation. Baseline: , (Equations (2) and (3)); alternative: , (Equation (4)). Tied priorities share a rank band.
| ID | P | Rank | Rank× | Position under both | |
|---|---|---|---|---|---|
| AI3 | 16 | 1 | 48 | 1–2 | leading |
| AI2 | 12 | 2–3 | 48 | 1–2 | leading |
| SL1 | 12 | 2–3 | 36 | 5 | leading / weight-dependent |
| CP1 | 10 | 4–5 | 40 | 3–4 | middle band |
| SL3 | 10 | 4–5 | 40 | 3–4 | middle band |
| CP2 | 8 | 6–7 | 32 | 6–7 | lower band |
| CP3 | 8 | 6–7 | 32 | 6–7 | lower band |
| AI1 | 5 | 8–9 | 25 | 8–9 | bottom (periodic verification) |
| SL2 | 5 | 8–9 | 25 | 8–9 | bottom (prohibition) |
Table 11.
Monte Carlo robustness of the priority ranking ( per variant). is the fraction of draws in which the risk is among the four highest priorities, ties sharing a rank band. Variant (a) resamples severity and exposure from their admissible neighbours; variant (b) additionally resamples .
Table 11.
Monte Carlo robustness of the priority ranking ( per variant). is the fraction of draws in which the risk is among the four highest priorities, ties sharing a rank band. Variant (a) resamples severity and exposure from their admissible neighbours; variant (b) additionally resamples .
| Variant | CP1 | CP2 | CP3 | AI1 | AI2 | AI3 | SL1 | SL2 | SL3 |
|---|---|---|---|---|---|---|---|---|---|
| (a) resampled | 0.61 | 0.42 | 0.43 | 0.00 | 0.95 | 1.00 | 0.83 | 0.00 | 0.61 |
| (b) resampled | 0.49 | 0.42 | 0.42 | 0.16 | 0.79 | 0.90 | 0.68 | 0.16 | 0.49 |
Table 12.
Readiness-raising roadmap. All nine frontier risks are listed in descending baseline priority P, so the table states an acceptance objective for every risk assessed and not for a selected subset; the priority column records where each sits in the ranking of Table 9.
Table 12.
Readiness-raising roadmap. All nine frontier risks are listed in descending baseline priority P, so the table states an acceptance objective for every risk assessed and not for a selected subset; the priority column records where each sits in the ranking of Table 9.
| ID | P | Target Acceptance Objective (RRL 4) | Verification Method | Candidate Vehicle |
|---|---|---|---|---|
| AI3 | 16 | AI safety functions must be auditable: documented training/validation data, performance bounds, and a deterministic fallback that holds the system safe on out-of-distribution input. | Design review against an assurance case; fault-injection of anomalous inputs. | IACS UR; IEC functional-safety extension informed by ISO/IEC TR 5469 and EASA learning assurance |
| AI2 | 12 | A learning-based EMS that influences thermal or electrical stress must operate inside a declared envelope, with a non-learning supervisory limiter that cannot be overridden. | Envelope verification over the declared operating domain; override test of the limiter. | IEC 61508-based class guidance for adaptive control in the safety loop |
| SL1 | 12 | Repurposed cells must be qualified by a measurement-based grading and re-rating process with stated screening thresholds and provenance data. | UL 1974-type evaluation anchored to battery-passport state-of-health data. | Class rule or EMSA guidance invoking UL 1974 and Regulation (EU) 2023/1542 |
| CP1 | 10 | Safety-critical battery functions (charge limit, protection trip, gas-response) must remain correct and authenticated under a defined threat model. | Penetration test of the safety loop; integrity verification. | IACS UR E26/E27 ext.; IEC 62443 zone/conduit profile for battery systems |
| SL3 | 10 | Building on the UR E18 battery schedule, the safety margin must be revalidated over life: defined re-test or inspection at stated SoH or calendar thresholds, with criteria to retire or de-rate. | Periodic survey + capacity/impedance and gas-response checks; battery-passport SoH data as evidence stream. | EMSA guidance update; class survey regime |
| CP2 | 8 | The high-rate shore-charging interface must have defined protection coordination, interlocks and authenticated control communication, with stated conformance tests. | Interface conformance test; fault and intrusion injection at the coupling. | Maritime profile binding the megawatt-charging standards to the ship-side battery safety case |
| CP3 | 8 | A digital twin or model-based estimator credited in safety decisions must be qualified for that role: defined validity domain, sensor-integrity monitoring, and detection and annunciation of model–reality divergence. | Validation against commissioning tests; injected sensor-fault campaign. | Class guidance adapting digital-twin qualification practice; EMSA update |
| AI1 | 5 | ML-based SoX/SoH used for safety decisions must meet a stated accuracy bound with quantified uncertainty, and degrade to a conservative estimate when confidence is low. | Back-test against reserved data; uncertainty-calibration test. | IEC; class guidance note adapting ISO/PAS 8800-style performance requirements |
| SL2 | 5 | Where a mixed-chemistry or mixed-SoH pack is permitted at all, it must be qualified as a system: no propagation across chemistry or SoH boundaries beyond a defined extent under abuse. The present prohibition should be replaced by a qualification route, not merely relaxed. | UL 9540A-type test on the heterogeneous configuration. | IEC 62619 ext.; class rule |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
