Figure 1.
Conceptual motivation for Predictive Operational Safety Engineering (POSE). Left: six established research streams (alarm management, fault detection and diagnosis, digital twins, HAZOP/LOPA, operator support, and prognostics/PHM) each address a partial reactive question but do not collectively produce a continuous estimate of the future operational safety state. Centre: the missing link is future operational safety-state estimation, reframing the operational question from what happened? to what will safety become? Right: POSE addresses four sequential predictive questions (what future safety state is emerging, how much safe operating time remains, which safety margin is being depleted, and which intervention has the highest safety value), implemented through Predictive Process Intelligence, Predictive Safety Intelligence, and Human Safety Intelligence.
Figure 1.
Conceptual motivation for Predictive Operational Safety Engineering (POSE). Left: six established research streams (alarm management, fault detection and diagnosis, digital twins, HAZOP/LOPA, operator support, and prognostics/PHM) each address a partial reactive question but do not collectively produce a continuous estimate of the future operational safety state. Centre: the missing link is future operational safety-state estimation, reframing the operational question from what happened? to what will safety become? Right: POSE addresses four sequential predictive questions (what future safety state is emerging, how much safe operating time remains, which safety margin is being depleted, and which intervention has the highest safety value), implemented through Predictive Process Intelligence, Predictive Safety Intelligence, and Human Safety Intelligence.
Figure 2.
Conceptual analogy between classical process control and industrial process safety. Feedback control (1A) and conventional safety (1B) share a reactive logic in which corrective or protective action follows an observed deviation. Feedforward control (2A) and Predictive Operational Safety Engineering (2B) share a proactive logic in which the system reasons ahead of deviations and intervenes early by estimating RSOT and RSM before safety boundaries are reached. This analogy is conceptual rather than algorithmic; POSE is a safety-reasoning paradigm, not a control algorithm.
Figure 2.
Conceptual analogy between classical process control and industrial process safety. Feedback control (1A) and conventional safety (1B) share a reactive logic in which corrective or protective action follows an observed deviation. Feedforward control (2A) and Predictive Operational Safety Engineering (2B) share a proactive logic in which the system reasons ahead of deviations and intervenes early by estimating RSOT and RSM before safety boundaries are reached. This analogy is conceptual rather than algorithmic; POSE is a safety-reasoning paradigm, not a control algorithm.
Figure 3.
Dynamic HAZOP-informed safety reasoning chain activated by a predicted process deviation. Static HAZOP knowledge (deviations, causes, consequences, safeguards, and recommended actions) feeds a seven-step reasoning sequence: (1) models and data predict a future process deviation before limits are reached; (2) a dynamic HAZOP knowledge engine transforms static documentation into an online reasoning layer; (3) HAZOP reasoning identifies the relevant deviation, infers probable causes, and infers potential consequences; (4) safeguard assessment evaluates the status and effectiveness of existing safeguards; (5) risk prioritization ranks scenarios by OVI, RSM, and RSOT to identify the most critical situation; (6) recommended operator action provides prioritized, actionable guidance before safety boundaries are crossed; and (7) the future safety state is updated to reflect the intervention or continued trajectory, closing the predictive reasoning loop. The three operator-facing output metrics span complementary dimensions: Remaining Safe Operating Time (RSOT, time dimension), Remaining Safety Margin (RSM, margin dimension), and Operational Vulnerability Index (OVI, overall vulnerability dimension).
Figure 3.
Dynamic HAZOP-informed safety reasoning chain activated by a predicted process deviation. Static HAZOP knowledge (deviations, causes, consequences, safeguards, and recommended actions) feeds a seven-step reasoning sequence: (1) models and data predict a future process deviation before limits are reached; (2) a dynamic HAZOP knowledge engine transforms static documentation into an online reasoning layer; (3) HAZOP reasoning identifies the relevant deviation, infers probable causes, and infers potential consequences; (4) safeguard assessment evaluates the status and effectiveness of existing safeguards; (5) risk prioritization ranks scenarios by OVI, RSM, and RSOT to identify the most critical situation; (6) recommended operator action provides prioritized, actionable guidance before safety boundaries are crossed; and (7) the future safety state is updated to reflect the intervention or continued trajectory, closing the predictive reasoning loop. The three operator-facing output metrics span complementary dimensions: Remaining Safe Operating Time (RSOT, time dimension), Remaining Safety Margin (RSM, margin dimension), and Operational Vulnerability Index (OVI, overall vulnerability dimension).

Figure 4.
Conceptual contrast between traditional prognostics and POSE. Left panel (blue): Remaining Useful Life (RUL) is an equipment-centered metric that estimates how long before a single asset’s health index crosses a failure threshold. Right panel (orange): Remaining Safe Operating Time (RSOT) is a process-safety-centered metric that estimates how long before the collective operational safety state crosses a safety boundary. The key distinction is that RSOT depends on predicted process trajectories relative to multi-variable safety boundaries, not on the degradation of any individual component. Two complementary concepts that extend this framework, Safety Horizon (the future interval over which the safety-state forecast remains sufficiently reliable for operational use) and Predictive Safety Confidence (PSC, the probability that a specified intervention can be completed before boundary crossing), are formally defined in
Section 11.
Figure 4.
Conceptual contrast between traditional prognostics and POSE. Left panel (blue): Remaining Useful Life (RUL) is an equipment-centered metric that estimates how long before a single asset’s health index crosses a failure threshold. Right panel (orange): Remaining Safe Operating Time (RSOT) is a process-safety-centered metric that estimates how long before the collective operational safety state crosses a safety boundary. The key distinction is that RSOT depends on predicted process trajectories relative to multi-variable safety boundaries, not on the degradation of any individual component. Two complementary concepts that extend this framework, Safety Horizon (the future interval over which the safety-state forecast remains sufficiently reliable for operational use) and Predictive Safety Confidence (PSC, the probability that a specified intervention can be completed before boundary crossing), are formally defined in
Section 11.
Figure 5.
Synthesis of the six reviewed research streams and the critical research gap they collectively expose. Each stream (Alarm Management, Fault Detection and Diagnosis, Digital Twins, HAZOP/LOPA, Operator Support, Prognostics/RUL) provides a partial view of the operational safety problem. No unified framework presently makes the future operational safety state the central object of continuous estimation. POSE addresses this gap by integrating process prediction, safety-state estimation, safety reasoning, and operator guidance around the POSE metrics: RSOT, RSM, OVI, PSC, and Safety Horizon.
Figure 5.
Synthesis of the six reviewed research streams and the critical research gap they collectively expose. Each stream (Alarm Management, Fault Detection and Diagnosis, Digital Twins, HAZOP/LOPA, Operator Support, Prognostics/RUL) provides a partial view of the operational safety problem. No unified framework presently makes the future operational safety state the central object of continuous estimation. POSE addresses this gap by integrating process prediction, safety-state estimation, safety reasoning, and operator guidance around the POSE metrics: RSOT, RSM, OVI, PSC, and Safety Horizon.
Figure 6.
The three-pillar architecture of Predictive Operational Safety Engineering (POSE). Plant data, process models, and safety knowledge feed into Predictive Process Intelligence, which generates predicted trajectories, hidden-state estimates, and uncertainty bounds. Predictive Safety Intelligence translates these forecasts into safety-state indicators (RSM, RSOT, OVI, PSC, Safety Horizon). Human Safety Intelligence converts these indicators into HAZOP-informed, safeguard-aware, ranked operator guidance. The bottom panel contrasts current practice (alarms, fault labels, reactive response) with the POSE paradigm (safety-state forecast, time-to-unsafe-state, predictive intervention).
Figure 6.
The three-pillar architecture of Predictive Operational Safety Engineering (POSE). Plant data, process models, and safety knowledge feed into Predictive Process Intelligence, which generates predicted trajectories, hidden-state estimates, and uncertainty bounds. Predictive Safety Intelligence translates these forecasts into safety-state indicators (RSM, RSOT, OVI, PSC, Safety Horizon). Human Safety Intelligence converts these indicators into HAZOP-informed, safeguard-aware, ranked operator guidance. The bottom panel contrasts current practice (alarms, fault labels, reactive response) with the POSE paradigm (safety-state forecast, time-to-unsafe-state, predictive intervention).
Figure 7.
Relationship among Remaining Safety Margin (RSM), Remaining Safe Operating Time (RSOT), Operational Vulnerability Index (OVI), and the intervention window. The predicted safety-state trajectory (blue, with uncertainty band) rises toward the safety boundary (red dashed line). RSM is the vertical distance remaining at the current time ; RSOT is the horizontal interval from to the predicted boundary crossing at ; OVI increases as the trajectory accelerates toward the boundary. The shaded green region is the intervention window available to the operator before safe operation is lost.
Figure 7.
Relationship among Remaining Safety Margin (RSM), Remaining Safe Operating Time (RSOT), Operational Vulnerability Index (OVI), and the intervention window. The predicted safety-state trajectory (blue, with uncertainty band) rises toward the safety boundary (red dashed line). RSM is the vertical distance remaining at the current time ; RSOT is the horizontal interval from to the predicted boundary crossing at ; OVI increases as the trajectory accelerates toward the boundary. The shaded green region is the intervention window available to the operator before safe operation is lost.
Figure 8.
Three-layer reference architecture of the Predictive Operational Safety Twin (POST). Real-time plant information (process measurements, historian data, alarm events, equipment state, and operating context) enters Layer I (Predictive Process Intelligence), which performs state synchronization, hybrid process modelling, future trajectory prediction, and uncertainty quantification. Layer II (Predictive Safety Intelligence) translates these trajectories into safety-state metrics: RSM, RSOT, OVI, PSC, Alarm Flood Horizon, and boundary-crossing forecasts. Layer III (Human Safety Intelligence) applies HAZOP/LOPA reasoning, safeguard status assessment, and explainable guidance to produce ranked interventions and operator decision support. Updated plant response closes the feedback loop. POST integrates process prediction, predictive safety metrics, and human-centered safety reasoning into a unified operational safety architecture.
Figure 8.
Three-layer reference architecture of the Predictive Operational Safety Twin (POST). Real-time plant information (process measurements, historian data, alarm events, equipment state, and operating context) enters Layer I (Predictive Process Intelligence), which performs state synchronization, hybrid process modelling, future trajectory prediction, and uncertainty quantification. Layer II (Predictive Safety Intelligence) translates these trajectories into safety-state metrics: RSM, RSOT, OVI, PSC, Alarm Flood Horizon, and boundary-crossing forecasts. Layer III (Human Safety Intelligence) applies HAZOP/LOPA reasoning, safeguard status assessment, and explainable guidance to produce ranked interventions and operator decision support. Updated plant response closes the feedback loop. POST integrates process prediction, predictive safety metrics, and human-centered safety reasoning into a unified operational safety architecture.
Figure 9.
Research roadmap for Predictive Operational Safety Engineering (POSE) across three sequential phases. Phase I (Foundation) establishes metric definitions for RSM, RSOT, OVI, and PSC, develops the POST reference architecture, and benchmarks the framework on representative processes including the Tennessee Eastman Process. Phase II (Validation) conducts industrial case studies, integrates dynamic HAZOP and safeguard reasoning, investigates real-time and edge deployment, and evaluates operator-interface and human-factors dimensions. Phase III (Application) pursues plant-wide safety twins, regulatory and governance acceptance, cross-industry transfer, and semi-autonomous safety intelligence under human oversight. Five cross-cutting themes (uncertainty and confidence, explainability, data quality and governance, human-centered design, and lifecycle model management) span all phases.
Figure 9.
Research roadmap for Predictive Operational Safety Engineering (POSE) across three sequential phases. Phase I (Foundation) establishes metric definitions for RSM, RSOT, OVI, and PSC, develops the POST reference architecture, and benchmarks the framework on representative processes including the Tennessee Eastman Process. Phase II (Validation) conducts industrial case studies, integrates dynamic HAZOP and safeguard reasoning, investigates real-time and edge deployment, and evaluates operator-interface and human-factors dimensions. Phase III (Application) pursues plant-wide safety twins, regulatory and governance acceptance, cross-industry transfer, and semi-autonomous safety intelligence under human oversight. Five cross-cutting themes (uncertainty and confidence, explainability, data quality and governance, human-centered design, and lifecycle model management) span all phases.
Table 1.
Research streams and remaining gaps motivating Predictive Operational Safety Engineering.
Table 1.
Research streams and remaining gaps motivating Predictive Operational Safety Engineering.
| Research stream |
Primary object |
Typical output |
Remaining gap for POSE |
| Alarm management [5,6] |
Alarm events |
Alarm priority, sequences, floods |
Does not estimate future safety-state trajectory |
| Fault diagnosis [40,44] |
Fault class and occurrence |
Fault label, residual, probability |
Does not translate diagnosis into RSOT or RSM |
| Digital twins [13,14] |
Process state |
Predicted variables, virtual representation |
Does not always translate prediction into safety state |
| HAZOP/LOPA [4,60] |
Hazard scenarios/barriers |
Deviations, causes, safeguards |
Usually offline or periodically updated |
| Operator support [73,76] |
Human decision quality |
Displays, explanations, guidance |
Lacks predictive safety-time reasoning |
| Prognostics/PHM [79] |
Asset health |
RUL, failure probability |
Equipment-centered, not process-safety-centered |
Table 2.
What POSE adds beyond existing safety paradigms. Each row identifies what an established discipline provides, where it stops, and the specific layer POSE adds.
Table 2.
What POSE adds beyond existing safety paradigms. Each row identifies what an established discipline provides, where it stops, and the specific layer POSE adds.
| Existing paradigm |
Primary purpose |
Typical outputs |
What it cannot provide |
What POSE adds |
| Alarm management [5,6] |
Detect and prioritize alarm events |
Alarm state, priority rank, flood pattern |
Safety-state forecast before alarm activation |
RSM and RSOT as pre-alarm leading indicators; OVI ranks urgency before alarms fire |
| Fault detection & diagnosis [40,44] |
Identify fault occurrence and type |
Fault label, fault probability, root-cause rank |
How safety margins evolve if no action is taken |
Converts fault evidence into a safety-state trajectory; PSC estimates whether an action can be completed in time |
| Digital twins [13,14,52] |
Replicate and predict process state |
Predicted T, P, flow, composition |
HAZOP-linked safety-boundary assessment; operator intervention guidance |
POST: reorients twin output from process fidelity to RSM, RSOT, OVI evaluated against the HAZOP safety envelope |
| HAZOP / LOPA [4,60,63] |
Identify hazard scenarios and assess safeguard sufficiency |
Deviation lists, safeguard evaluations, SIL requirements |
Continuous online trajectory interpretation during operation |
Activates HAZOP knowledge dynamically; barrier status propagated into OVI in real time |
| Dynamic risk assessment [20,69,86] |
Update event probability from evidence |
Updated event probability |
Time-to-safety-boundary; ranked intervention guidance |
RSOT: translates probability language into intervention-window language; not how likely? but how long? and which action?
|
| Prognostics & PHM [79,80,82] |
Estimate remaining useful life of assets |
RUL distribution; degradation trajectory |
Multi-variable, HAZOP-linked process safety horizon |
RSOT: generalizes RUL from single-asset thresholds to multi-variable, HAZOP-defined safety boundaries |
| Operator decision support [73,76] |
Improve human response to process information |
Alarm displays, XAI explanations |
Proactive safety-state forecasts; intervention-time ranking |
PSC and OVI: time-stamped, safety-ranked intervention guidance before alarms activate (Layer III of POST) |
Table 3.
Why six commonly cited analogues are not POSE: each concept addresses a different primary question and imposes qualitatively different obligations from RSM, RSOT, OVI, and PSC.
Table 3.
Why six commonly cited analogues are not POSE: each concept addresses a different primary question and imposes qualitatively different obligations from RSM, RSOT, OVI, and PSC.
| Concept |
Primary question answered |
Typical output |
Why it is not POSE |
| Time-to-alarm |
When will a single tag hit a preset alarm limit? |
Extrapolated scalar time |
Single-variable; no HAZOP linkage; no multi-variable trajectory; no uncertainty quantification |
| Remaining useful life (RUL) |
When will an asset degrade to failure? |
Asset-level life estimate |
Equipment-centered; does not address safety-boundary proximity, safeguard status, or operator intervention windows |
| Risk Priority Number (RPN) |
How severe, likely, and detectable is a failure mode? |
Static ordinal score |
Assigned offline at design time; not updated from the predicted process trajectory; no time-to-boundary output |
| Dynamic risk index |
How has event likelihood changed given current evidence? |
Updated probability or risk scalar |
Probability-centered; does not produce margin depletion rates, RSOT estimates, or ranked operator-intervention guidance |
| Alarm priority |
Which active alarm deserves the most immediate attention? |
Priority class or rank |
Reactive to alarm activation; does not forecast margin trajectories before alarms fire; no RSOT-equivalent output |
| MPC constraint margins |
How close are process variables to control constraint bounds? |
Constraint distance under active controller |
Serves automatic control-enforcement; not designed for operator safety reasoning or HAZOP-linked intervention-time estimation |
Table 4.
Comparison of early warning / incipient fault detection and POSE.
Table 4.
Comparison of early warning / incipient fault detection and POSE.
| Aspect |
Early warning / incipient fault detection |
POSE |
| Primary question |
Is an abnormality developing before the alarm fires? |
How will safety margins evolve, and how much intervention time remains? |
| Main output to operator |
Warning signal, fault precursor, or predictive alarm |
RSM, RSOT, OVI, PSC, and ranked intervention guidance |
| Time logic |
Before alarm activation or fault classification |
Before safety boundary is approached |
| Safety knowledge |
Detection thresholds or model-based signatures |
HAZOP-informed safety boundaries, safeguard states, consequence severity |
| Operator value |
Alerts operator to a developing situation |
Quantifies urgency, ranks interventions, and estimates action feasibility |
Table 5.
Contrast between current industrial safety practice and the POSE paradigm.
Table 5.
Contrast between current industrial safety practice and the POSE paradigm.
| Dimension |
Current practice |
POSE |
| What is monitored? |
Process variables and alarm threshold crossings |
Operational safety state as a dynamic, predicted trajectory |
| When does the system respond? |
After alarm limits are exceeded |
Before safety boundaries are approached |
| Primary output to operator |
Alarm activation, fault label, risk probability |
RSM, RSOT, OVI, and ranked intervention guidance |
| Time orientation |
Present and near past |
Predicted future under the current operating trajectory |
| Safety knowledge application |
Offline in HAZOP documents |
Connected online to trajectory predictions and deviation reasoning |
| Safety boundary definition |
Single-variable alarm setpoints |
Multi-variable, HAZOP-informed safety envelope |
| Treatment of uncertainty |
Often implicit or separate from alarms and diagnosis |
Explicitly linked to safety horizon and predictive safety confidence |
Table 6.
Functional taxonomy of POSE.
Table 6.
Functional taxonomy of POSE.
| POSE pillar |
Representative inputs |
Core functions |
Representative outputs |
| Predictive Process Intelligence |
Process historian data, current measurements, control actions, equipment state, disturbances |
State estimation, trajectory forecasting, uncertainty quantification |
Predicted trajectories, estimated hidden states, uncertainty bounds |
| Predictive Safety Intelligence |
Process trajectories, alarm limits, safe operating envelopes, safeguard thresholds, topology |
Safety-state classification, margin estimation, time-to-boundary estimation, vulnerability scoring |
RSM, RSOT, OVI, PSC, Safety Horizon, boundary-crossing forecasts |
| Human Safety Intelligence |
Safety-state forecasts, HAZOP knowledge, procedures, safeguard status, operator context |
Explanation, action ranking, consequence interpretation, interface design |
Ranked interventions, HAZOP-linked explanations, safeguard status, operator guidance |
Table 7.
Core safety metrics of POSE and their operational interpretation.
Table 7.
Core safety metrics of POSE and their operational interpretation.
| Metric |
Definition |
Operational question answered |
| Remaining Safety Margin (RSM) |
Normalized remaining distance to the nearest binding safety boundary |
How far is the process from a safety boundary? |
| Remaining Safe Operating Time (RSOT) |
Predicted time until RSM reaches zero under current trajectory |
How much time remains before a boundary is crossed? |
| Operational Vulnerability Index (OVI) |
Bounded weighted composite of proximity, depletion rate, consequence, safeguard unavailability, and uncertainty |
How urgent is intervention for the governing scenario? |
| Predictive Safety Confidence (PSC) |
Probability that first-passage time exceeds the time required for a specified intervention |
What is the probability that action can be completed in time? |
Table 8.
Reference functions of a Predictive Operational Safety Twin.
Table 8.
Reference functions of a Predictive Operational Safety Twin.
| POST layer |
Purpose |
Example output |
| State synchronization |
Estimate the current process and safeguard state from online data |
Current operating state, active constraints, available safeguards |
| Trajectory prediction |
Forecast future process evolution under current or candidate actions |
Predicted pressure, temperature, flow, composition, or inventory trajectories |
| Safety-state translation |
Convert predicted trajectories into safety regimes and margins |
Normal, vulnerable, critical, or unsafe forecast; RSM; OVI |
| Time-to-boundary estimation |
Estimate when relevant safety boundaries may be reached |
RSOT, alarm flood horizon, trip horizon, confidence interval |
| Safety reasoning |
Link predicted deviations to causes, consequences, and safeguards |
HAZOP-linked explanation and consequence pathway |
| Decision support |
Rank interventions according to urgency, feasibility, and safety value |
Prioritized action guidance with rationale and uncertainty |
Table 9.
Ten principles of Predictive Operational Safety Engineering and their meaning for POSE research and practice.
Table 9.
Ten principles of Predictive Operational Safety Engineering and their meaning for POSE research and practice.
| # |
Principle |
Meaning for POSE |
| 1 |
Safety is a dynamic state, not merely an event |
Forecast safety evolution continuously; systems that respond only to discrete events forfeit the prediction window available beforehand |
| 2 |
Future safety-state prediction provides earlier intervention value than present alarm status alone |
Predicting safety degradation before alarm limits are reached provides substantially greater intervention time [6] |
| 3 |
Alarms are symptoms of an evolving safety state, not the safety state itself |
Use alarms as evidence of safety-state change; alarms cannot describe how margins are evolving or when boundaries will be crossed [5,6] |
| 4 |
Fault diagnosis is necessary but insufficient for safety reasoning |
Convert diagnostic outputs into safety-state trajectories; detection accuracy alone does not quantify how rapidly safety margins are depleting [7,40] |
| 5 |
For safety-critical applications, digital twins should predict safety meaning, not only process variables |
Map process forecasts into RSM, RSOT, OVI, and PSC before they become actionable for operational safety [13,14] |
| 6 |
Safety intelligence must integrate physics, data, knowledge, and human expertise |
Combine first-principles models, data-driven methods, HAZOP/LOPA knowledge, safeguard states, and operator expertise; no single source is sufficient [3,22] |
| 7 |
Prediction without interpretation does not improve safety |
Apply HAZOP-informed consequence and safeguard reasoning to translate process forecasts into operator-relevant guidance [27,67] (Sections 6–7) |
| 8 |
Operators remain central to industrial safety decisions |
Design for situation awareness, trust, and action; predictive information must match human cognitive needs to improve rather than impair decision making [76] (Section 7) |
| 9 |
Safety metrics must be predictive |
RSM, RSOT, OVI, and PSC complement fault labels, alarm counts, and diagnostic outputs by translating them into predictive safety-state information that quantifies margin, intervention time, vulnerability, and confidence |
| 10 |
Validation is mandatory |
Benchmark, calibrate, explain, and test human impact; operators must understand and trust predictions for POSE to improve safety-critical decisions (Section 14) |
Table 10.
Candidate benchmark systems and industrial applications for POSE validation.
Table 10.
Candidate benchmark systems and industrial applications for POSE validation.
| System |
Process type |
Key safety challenge |
POSE relevance |
| Tennessee Eastman Process [11] |
CSTR with recycle |
Multiple fault modes, sensor failures |
De facto FDD benchmark with rich alarm and measurement data |
| CO2 capture plant |
Absorption and stripping |
Solvent flooding, column pressure deviation |
Dynamic safety margins under varying load conditions |
| Sulfuric acid converter |
Catalytic SO2 oxidation |
Catalyst deactivation, bed temperature runaway |
RSOT relevant to progressive temperature excursion |
| Distillation column |
Vapor–liquid separation |
Hydraulic flooding, weeping, product deviation |
Multi-variable safety boundary estimation |
| Ammonia synthesis loop |
High-pressure catalytic reaction |
Pressure safety, catalyst activity loss |
RSOT estimation under high-pressure upsets |
| Hydrogen production unit |
Steam reforming |
Reformer tube integrity, H2 composition |
Equipment–process coupled safety horizon |