Preprint
Article

This version is not peer-reviewed.

Monte Carlo Probabilistic Assessment of Loss of Human Control Over Advanced AI Using Bowtie Analysis

Submitted:

09 June 2026

Posted:

10 June 2026

You are already at the latest version

Abstract
A quantitative risk assessment of human loss of control over advanced AI used a Bowtie diagram extended with fault tree and event tree analysis. Six primary threats were identified (recursive self‑improvement, power seeking, deceptive alignment, loss of corrigibility, off‑switch subversion, malicious misuse) and six consequences (systemic infrastructure collapse, economic breakdown, resource shortages, non‑human value lock‑in, human marginalization, global supply cascade failures). Preventive and mitigative barriers were assigned per pathway from expert literature. Input probabilities (threat base rates and barrier failure‑on‑demand values) were sourced from experts and modeled with triangular uncertainty distributions. A 1,000‑iteration Monte Carlo simulation propagated epistemic uncertainty, yielding a median probability of the top event (loss of human control) of 12.8% (90% CI: 11.3%–14.4%), roughly 1 in 8. The distribution is approximately symmetric with slight positive skew, indicating modest tail risk if barrier failures interact. Conditional on the top event, Expected Severity is 1.85 on a 1–10 scale (90% CI: 1.75–1.96), suggesting mitigation is effective in most scenarios. Results align with expert estimates and demonstrate barrier effects; narrow CIs reflect model consistency. Remaining tail risks support precautionary governance, increased alignment research, iterative risk modeling, and investment in international coordination with robust safety measures to reduce the existential risk of AI loss of control.
Keywords: 
;  ;  ;  ;  

1. Introduction

The possibility of advanced artificial intelligence (AI) realizing autonomous strategic agency poses one of the most significant existential risks facing humanity this century 1,2. Loss of human control over AI can arise from several issues such as misalignment, resource accumulation, lock-in, capability escalation, or subversion of safeguards. Loss of control could lead to outcomes ranging from systemic collapse, resource shortages, economic impacts and irreversible disempowerment (humans lose agency and control of AI decision-making). While expert surveys and qualitative analyses have provided some estimates for this scenario where humans lose control over AI 3,4, structured quantitative models remain limited.
The U.S. State Department commissioned Gladstone AI Inc. to conduct an AI risk assessment in October 2022. The goal of the report was to examine the risk of AI loss of control which is considered an existential risk resulting in weaponization of AI. In their report they identified seven downstream weaponization scenarios, all which could escalate to existential risks that are catastrophic and potentially "extinction-level" in worst-case scenarios. When combined with global scaling capabilities the adverse impacts are magnified 4.
This risk assessment highlights loss of control as the pivotal event enabling further adverse consequences. The present study was undertaken to address the gap in rigorous quantification by applying Bowtie risk assessment methodology, combined with fault tree and event tree analysis, to systematically model the loss-of-control scenario extending to 2070. The central hypothesis is that a Monte Carlo-based quantitative approach, grounded in expert-derived probabilities from existing literature, can generate consistent and defensible estimates of both the probability and severity of human loss of control over advanced AI, thereby providing a practical foundation for evaluating safety measures and informing governance strategies.
Bowtie analysis provides a transparent visual and logical framework for mapping complex causal pathways and mitigation layers in risk scenarios. In this study, the Bowtie serves as the theoretical foundation, with the central top event defined as “loss of human control over advanced AI with autonomous strategic agency.” Six primary threats and six consequences were identified from expert literature, with preventive barriers (on the left side) and mitigative barriers (on the right side) assigned to each pathway. This structure creates a clear causal model suitable for quantitative extension.
To operationalize the Bowtie, fault tree analysis was applied to the left side to compute the probability of the top event as the sum of individual threat-path probabilities (base threat probability multiplied by the product of preventive barrier failure-on-demand probabilities). Event tree analysis was used on the right side to calculate scenario probabilities for each consequence (top-event probability multiplied by the product of mitigative barrier Probabilities of Failure on Demand or PFDs) and an aggregated expected severity score (sum of scenario probability × severity rating on a 1–10 scale).
All input values were derived from peer-reviewed and grey literature and modeled using triangular uncertainty distributions to reflect expert ranges. A Monte Carlo simulation with 1,000 iterations propagated these uncertainties to generate distributional outputs. Sensitivity analysis was performed on grouped barriers using reductions in PFDs. To the authors’ knowledge, this represents the first publicly available, fully quantified risk assessment of human loss of control over advanced AI using Bowtie methodology combined with fault/event tree analysis and Monte Carlo simulation. This practical development from theoretical risk modeling provides a reproducible framework that can be iteratively refined with new data and expert input.

2. Materials and Methods

2.1. Bowtie Risk Model Development

A Bowtie diagram was constructed to represent the qualitative risk structure of the scenario for "human loss of control over advanced artificial intelligence" (also referred to as AI misalignment or power-seeking leading to autonomous strategic agency). The central undesired event (top event) is defined as loss of human control over advanced AI with autonomous strategic agency, potentially resulting in adverse outcomes ranging from systemic infrastructure collapse to existential catastrophe.
The Bowtie was informed by expert literature on AI existential risks 1,2,3,4,5,6,7. Six primary threats (causes) were identified on the left side of the Bowtie, each associated with preventive barriers (controls intended to prevent the top event). Six consequences (adverse outcomes) were identified on the right side, each associated with mitigative barriers (controls intended to reduce or prevent escalation after the top event occurs). For the Bowtie diagram elements see Appendix A.
Threats and Preventive Barriers
The six threats (T) and their assigned preventive barriers (P) are as follows:
T1: Capability Escalation – Recursive Self-Improvement (Intelligence Explosion)
  • T1P1- AI Containment Controls (Sandboxing)
  • T1P2- Compute and Model Scaling Controls
  • T1P3- Scalable Alignment Techniques
T2: Capability Escalation – Power-Seeking and Resource Accumulation
  • T2P1- Red-Team Testing and Safety Audits
  • T2P2- International Governance AI Frameworks
  • T2P3- Compute and Model Scaling Controls
T3: Alignment Failure – Deceptive Alignment or Treacherous Turn
  • T3P1- Corrigibility Training
  • T3P2- AI Interpretability and Monitoring
  • T3P3- Scalable Alignment Techniques
T4: Alignment Failure – Irreversible Value Lock-In (Loss of Corrigibility)
  • T4P1- Value Update and Oversight Protocols
  • T4P2- AI Interpretability and Monitoring
  • T4P3- Scalable Alignment Techniques
T5: Control Failure – Subverted Shutdown and Kill Switch Mechanisms
  • T5P1- Hardware Failsafe
  • T5P2- Redundant Kill Switches
  • T5P3- Compute and Model Scaling Controls
T6: Malicious Misuse – Weaponization of Aligned AI
  • T6P1- Access Control (Vetting) and Licensing for Advanced AI
  • T6P2- International Governance AI Frameworks
Consequences and Mitigative Barriers
The six identified consequences (C) and their assigned mitigative barriers (M) are as follows:
C1: Systemic Infrastructure Collapse
  • C1M1- Redundant and Decentralized Infrastructure
  • C1M2- AI Oversight and Safety Monitoring Systems
C2: Societal and Economic Breakdown
  • C2M1- Include Human-in-the-Loop Decision Overrides
  • C2M2- Critical Process Auditing with Contingency Planning
C3: Population-Level Resource Shortages
  • C3M1- Automated Resource Allocation Monitoring
  • C3M2- Fail-safe Safeguards and Rationing Protocols
C4: Irreversible Value Lock-In of Non-Human Values
  • C4M1- Continuous Goal and Value Review Protocols
  • C4M2- Corrigibility with Self-Alignment Update Mechanisms
C5: Human Marginalization / Disempowerment / Extinction
  • C5M1- Global Emergency AI Shutdown Protocols
  • C5M2- Emergency Isolation Protocols
  • C5M3- International AI Governance and Enforcement
C6: Global Catastrophic Risk Cascades (Interacting System Failures)
  • C6M1- Systemic Risk Early Warning Systems to Prevent Cascade Failures
  • C6M2- Cross-Domain Red Teams and System Stress Testing

2.2. Quantitative Risk Assessment (QRA)

The Bowtie diagram was extended to a quantitative risk assessment using fault tree analysis (left side: threats to top event) and event tree analysis (right side: consequences from top event). All quantitative parameters were derived from peer-reviewed and grey literature on AI existential risks (see References and Appendix B).

2.3. Probability Assignment

Base probabilities for threats (unconditional probability of occurrence before preventive barriers) were assigned as point estimates with triangular uncertainty distributions (low, most likely, high) based on expert values and ranges from literature (peer review manuscripts, books, AI expert surveys, and gray literature) (typically 0.05–0.30 or 5–30%).
Probability of Failure on Demand (PFD) for preventive and mitigative barriers was similarly assigned using triangular distributions (typically with ranges between 0.30–0.80 or 30–80% for preventive barriers, and ranges between 0.40-0.80 or 40-80% for mitigative barriers), reflecting high uncertainty in AI control efficacy. Again barrier ranges were derived from available expert values.

2.4. Time Line

A baseline time horizon of present day to 2070 was used, consistent with expert values 2. Longer views beyond 2070 (e.g., to 2100) were not considered due to the increasing uncertainty and trajectory of AI development longer-term and the belief that failure to contain AI will happen much quicker than anticipated with recent expert surveys suggesting accelerated AI development timelines 3.

2.5. Monte Carlo Simulation

To propagate uncertainty and generate probabilistic distributions, a Monte Carlo simulation was implemented initially in Microsoft Excel (version 2016) with 1,000 iterations and confirmed using separate Python code outputs. Triangular distributions were applied to all input parameters (base probabilities and PFDs) using the inverse cumulative distribution function method. Each iteration sampled independent random values from the specified ranges for threats, preventive barriers, mitigative barriers, and consequence severities (fault event scale: 1–10).
The simulation calculated:
  • Individual threat path probabilities (base probability × product of preventive PFDs)
  • Total P(top event) = sum of all six threat path probabilities
  • Individual consequence scenario probabilities = P(top event) × product of mitigative PFDs
  • Expected Severity = sum of (consequence scenario probability × severity score)
Assumptions and Limitations
  • Barriers are assumed to be independent (no correlation modeled).
  • Threats are treated as mutually exclusive with negligible overlap.
  • Estimates are subjective, derived from experts, expert surveys, peer review literature, and gray literature. Values carry high epistemic uncertainty.
  • The model focuses exclusively on the "human loss of control" scenario, as it is the only Gladstone AI (2024) high-risk scenario with sufficient probabilistic data available from literature sources for quantification 4.
  • Other AI existential risks (e.g., misuse without loss of control, accidental emergence) are excluded.
  • Sensitivity analysis was used to test the sensitivity to key parameters and elements (e.g., similar barriers effectiveness, threat base rates, and consequence severity).

3. Results

The Monte Carlo simulation was conducted with 1,000 iterations to propagate uncertainty through the Bowtie model. Key outputs included the probability of the top event (P(Top Event), i.e., loss of human control over advanced AI) and the Expected Severity score (the weighted sum of consequence scenario probabilities multiplied by severity ratings on a 1–10 scale). Summary statistics from the Monte Carlo simulation, for both metrics are presented in Table 1.
The distribution of P(Top Event) across the 1,000 iterations is shown in Figure 1. Data derived from triangular uncertainty distributions were applied to threat base probabilities and preventive barrier PFDs. The resulting histogram reveals an approximately symmetric bell-shaped curve centered around the median value of 0.128 (12.8%), with a slight positive skew indicated by a longer tail to the right (higher probability values). The peak bin is located at approximately 0.126, with adjacent bins showing lower frequencies, confirming the bell-like form. The narrow spread (5th to 95th percentile: 0.113 to 0.144) demonstrates low overall dispersion, suggesting consistent outcomes across simulations despite input uncertainties.
The distribution of the Expected Severity Score is shown in Figure 2. Similar to P(Top Event), the histogram exhibits a roughly symmetric bell-shaped distribution that is tightly clustered with minimal skew, centered around the median of 1.847. The range is even narrower (5th to 95th percentile: 1.748 to 1.957), indicating high consistency in conditional severity estimates. No significant tails are observed, reflecting the mitigating effects of the modeled barrier elements.
These results represent the basic outputs of the Monte Carlo simulation, including central tendency measures (mean and median), dispersion (percentiles), and visual distributions of the key risk metrics. For Figure 2, despite the high raw severity ratings assigned to consequences (7–10 on the 1–10 scale, reflecting potential for catastrophic outcomes like extinction or global cascades), the expected severity score peaks at 1.85 (90% CI: 1.75–1.96). This represents an approximate 75-80% reduction in effective impact relative to unmitigated consequences, driven by the cumulative efficacy of preventive and mitigative barriers.
Preventive barriers reduce the median P(Top Event) to 12.8%, avoiding loss of control in ~87% of simulations, while mitigative barriers further limit escalation by succeeding in 20–60% of cases (based on modeled PFDs). This means that even if AI loss of control occurs, the barriers if in place act like multiple safety nets, catching ~75% of the potential harm and keeping average outcomes in the low-to-moderate range. This interpretation underscores the model's optimism about barrier effectiveness but also underscores the importance of real-world implementation of both preventive and mitigative barriers. However, if barriers are not implemented or are less effective than assumed (e.g., due to correlated failures), consequence severity from failure could rise significantly.

3.1. Sensitivity Analysis

Sensitivity analysis was performed to assess the robustness of the top-event probability estimate (loss of human control over advanced AI) and to identify which groups of input parameters exert the greatest influence on the scenario results. This standard quantitative risk assessment step using Monte Carlo methods helps to quantify uncertainty propagation, evaluate model stability, and highlight high-leverage assumptions for scrutiny 8,9. Conducting sensitivity analysis is important as it can help identify which parameters are most important and provides an evidence base for allocating resources 10.
A particular challenge in this Bowtie model is the lack of a direct one-to-one mapping between preventive barriers on the fault-tree side (which influence the probability of reaching the top event) and mitigative barriers on the event-tree side (which affect conditional consequences given the top event has occurred). Preventive and mitigative controls operate independently across different parts of the risk profile, so sensitivities were conducted separately on the fault-tree side only.
To make the analysis interpretable and computationally manageable, related preventive barriers of interest were grouped by their thematic function (e.g., scalable alignment techniques, shutdown mechanisms, governance frameworks, and access controls).
The threat barrier parameter groups used were:
  • Scalable alignment techniques (the main ways AI is kept on task) (T1P3,T3P3, T4P3)
  • Shutdown and kill-switch mechanisms (ways to turn off or stop an AI that’s going wrong) (T5P1, T5P2)
  • International governance frameworks (global agreements and rules) (T2P2, T6P2)
  • Access control and licensing (who gets to build or use powerful AI) (T6P1)
  • And finally, the starting likelihoods of the six threats themselves (how likely each threat is to begin in the first place) as one big group before any barriers act.
One-at-a-time variations were applied to each group’s parameters (all set to their lowest or highest triangular bounds while holding all others at base values). All sensitivity runs used 10,000 Monte Carlo iterations to minimize sampling error and ensure stable comparisons, Sensitivity analysis results are shown in Figure 3.
Figure 3 shows the relative changes that occur from the base-case mean (~12.8%). Leftward bars indicate risk reduction; rightward bars indicate risk increase. The widest swing (the most sensitive) occurs when varying the base probabilities of all six threats (T1–T6), followed by the scalable alignment barriers (T1P3, T3P3, T4P3). Governance (T2P2, T6P2), shutdown failure (T5P1, T5P2), and access control (T6P1) groups show progressively smaller effects.
The biggest change in risk (the most sensitive elements) were the starting likelihoods of the six threats themselves (how likely each threat is to begin in the first place), the overall chance of losing control shifted by about 36-40% relatively when compared to the original estimate. When we made the threats much less likely to occur (using stronger assumptions), the risk dropped from roughly 12.8% to about 8.2%. When we made them much more likely (weaker assumptions), the risk rose to about 17.9%.
The second biggest change resulted from scalable alignment techniques (the barriers we use to keep AI behaving the way we want). When barriers were made stronger and more reliable the relative risk dropped by about 10.5% compared to our original estimate - taking it from roughly 12.8% down to about 11.5%. When scalable alignment was made weaker (less reliable), the risk rose by about 10.5%, taking it up to ~14.1%.
International governance rules and shutdown/off-switch mechanisms had a smaller effect (around 5–6% relative change either way). Access controls and licensing had the smallest effect (around a 4% relative change).

4. Discussion

The 90% confidence interval (5th to 95th percentile) for the P(Top Event) spans 0.113 to 0.144 (equating roughly as a 1 in 9 to 1 in 7 chance). This interval is notably narrow, covering only a 3.1-percentage-point range, which indicates low dispersion in the simulation outputs, despite the epistemic uncertainties in the input parameters (threat base rates and barrier PFDs drawn from expert literature).
In contrast, many standalone expert estimates of AI existential risk exhibit much wider ranges. For example, Gladstone AI (2024) expert survey results suggests a 10-80% range for loss-of-control scenarios in super intelligent systems 4, while other experts from aggregated survey medians place this at 5-14% with considerable expert disagreement 3. The QRA's tight clustering implies that the combination of modeled preventive barriers and the chosen triangular distributions constrains the outcome more than some subjective expert views, resulting in higher consistency across simulations.
The slight right skew in the distribution (Figure 1) indicates that while most scenarios cluster near the median of 12.8%, rare combinations of barrier failures can act to push the probability modestly higher, highlighting a small but non-negligible tail risk. This positive skewing of the distribution reflects structural uncertainties in the model, particularly the possibility of correlated or cascading failures in alignment techniques, containment measures, or governance frameworks, which are not fully captured by the assumption of independent barriers. Such tail events align with concerns raised in the literature regarding "fast takeoff" scenarios 2 or deceptive alignment leading to rapid capability jumps 5, where preventive barriers may fail more catastrophically than expected. The longer right tail underscores that even with a median around 13%, the upper-end risk (e.g., >14%) cannot be dismissed as negligible, as it represents plausible, though less frequent, recognitions of the input ranges. This tail risk reinforces the need for robust sensitivity analyses and strong precautionary measures. The downside potential although narrow could affect expected value calculations or decision-making under uncertainty.
Table 2 compares the Monte Carlo outputs to selected expert estimates from the literature, expressed in comparable formats (percentages and 1 in x odds) for clarity. The model's estimates are in agreement with recent expert surveys and estimations, though they tend toward the higher end of some conservative bounds reported in recent surveys.

5. Conclusions

The Monte Carlo simulation results underscore the non-negligible existential risk of human loss of control over advanced AI, with a median probability of approximately 12.8% by 2070 and a narrow 90% confidence interval of 11.3% to 14.4%. This level of risk, equivalent to roughly 1 in 8 odds is well above de minimis risk levels of 1 in one million usually used by various U.S. and Canadian government departments 12,13 It warrants immediate and proactive policy interventions to strengthen preventive barriers such as scalable alignment techniques and international governance frameworks.
Policymakers should prioritize global coordination efforts, as recommended by Gladstone AI (2024) and RAND Corporation (2025), including binding treaties on compute scaling controls and access vetting for advanced AI systems 4,7. Investing in red-team testing and corrigibility training could further reduce threat pathways, potentially lowering the modeled probability by enhancing barrier effectiveness. Given the slight right skew in the distribution (see Figure 1), which indicates tail risks where probabilities exceed 14%, policies should incorporate precautionary principles to mitigate worst-case scenarios, such as rapid capability escalation or deceptive alignment failures.
The relatively low expected severity score of 1.85 (90% confidence interval: 1.75 to 1.96) (see Figure 2), on a 1–10 scale suggests that mitigative barriers, including global emergency shutdown protocols and systemic risk early warning systems if implemented, can effectively limit adverse outcomes if loss of control occurs 14,15. However, this conditional mitigation does not diminish the overall imperative for risk management, as even moderate-severity events like societal breakdown or resource shortages could have cascading global impacts or unforeseen impacts to low income countries (LICs).
To address the potential downsides from breaches, governments and organizations should allocate resources to bolster mitigative measures, such as redundant infrastructure and human-in-the-loop overrides, while adding and strengthening ongoing monitoring of AI developments. The model's assumptions of barrier independence highlight a potential underestimation of correlated failures; thus, policy should include stress-testing and scenario planning to identify vulnerabilities, aligning with Ord's (2020) emphasis on reducing existential risks through coordinated preparedness 1.
Finally, the narrow confidence intervals in both metrics indicate robustness in the model's outputs under the given assumptions, but sensitivity analyses are recommended to explore variations in input ranges, such as accelerated timelines suggested by recent surveys 3.
Policy makers should view these results as a clarion call for iterative risk assessment frameworks, including regular expert elicitations and model refinements, to inform adaptive strategies for AI loss of control events. Ultimately, while the estimated risk is not catastrophic in median cases, the tail uncertainties reinforce the need for proactive investment in AI safety research and international enforcement mechanisms to avert irreversible, existential outcomes.

Funding

This research received no external funding.

Institutional Review Board Statement

This study did not require ethical approval.

Data Availability Statement

Data is contained within the article.

Conflicts of Interest

The authors declare no conflict of interest.

Appendix A

Bowtie Diagram for Human Loss of AI Control
Figure A1. Left: Fault tree with six primary threats to the top event “Loss of Human Control over Advanced AI with Autonomous Strategic Agency,” each mitigated by 1–3 preventive barriers. Center: Undesired top event. Right: Event tree showing six consequence categories (severity 0–10), each reduced by 1–3 mitigative barriers. Threat probabilities and PFDs from AI safety literature and expert surveys; severities reflect impacts from infrastructure collapse to human extinction based on AI risk experts qualitative ranking.
Figure A1. Left: Fault tree with six primary threats to the top event “Loss of Human Control over Advanced AI with Autonomous Strategic Agency,” each mitigated by 1–3 preventive barriers. Center: Undesired top event. Right: Event tree showing six consequence categories (severity 0–10), each reduced by 1–3 mitigative barriers. Threat probabilities and PFDs from AI safety literature and expert surveys; severities reflect impacts from infrastructure collapse to human extinction based on AI risk experts qualitative ranking.
Preprints 217667 g0a1

Appendix B

Table A1. Bowtie element variable table showing element type with description, base probability, and triangle range probabilities (low, medium and high). Threats assigned base and range probabilities, Consequences were assigned severity scores (0-10). Sources, notes and parameter value assumptions used for the risk model also shown. .
Table A1. Bowtie element variable table showing element type with description, base probability, and triangle range probabilities (low, medium and high). Threats assigned base and range probabilities, Consequences were assigned severity scores (0-10). Sources, notes and parameter value assumptions used for the risk model also shown. .
Element ID and Type Description Linked To Base Probability / PFD Probability Range Low Probability Range Med Probability Range High Severity Score
(1-10)
Notes for Element Parameter Assumptions
T1
Threat
Capability Escalation: Recursive Self-Improvement (Intelligence Explosion) 0.15 0.1 0.15 0.2 Carlsmith (2021, p.11): Personally, I’m at something like 65% on “developing APS systems will be possible/financially feasible before 2070.”;
Ord (2020): As updated in 2024, https://www.tobyord.com/writing/the-precipice-revisited): These were 1 in 1,000 for Climate and for Nuclear, 1 in 30 for Pandemics and 1 in 10 for Unaligned AI.;
Gladstone (2024, p.4): Loss of control due to AGI alignment failure is a potential failure mode under which a future AI system could become so capable that it escapes all human efforts to contain its impact.;
AI Impacts (2023): Median ~50% AGI by 2060, implying ~15% chance for explosive self-improvement as a precursor. Assumes 2070 horizon; base probability as fraction of timelines (intelligence explosion precursor). Independent of barriers.
T1P1
Threat-Preventive Barrier
AI Containment Controls (Sandboxing) T1 0.6 0.5 0.6 0.7 LessWrong (AI Box experiments): Eliezer Yudkowsky discusses the ability to convince a human playing a guard to let them out on many - but not all - occasions.; No direct quote for 50-70% barrier range but inferred from LessWrong (2020) on frequent escapes implying medium-high PFD. Assumes independence; high PFD due to AI deception/subversion. Effectiveness decreases with AI capability.
RAND (2025): Researchers have identified warning signs of control-undermining capabilities in advanced AI models – including deception, self-preservation and autonomous replication which could potentially enable increasingly capable models to evade human oversight. Again assign medium high PFD.
T1P2
Threat-Preventive Barrier
Compute and Model Scaling Controls T1 0.4 0.3 0.4 0.5 Gladstone (2024): Establish interim safeguards to stabilize advanced AI development, including export controls on the advanced AI supply chain. No direct quote for medium effectiveness but inferred from RAND (2025) on slowing AI proliferation implying 30-50% failure. Assumes global enforcement which is unlikely raising the PFD; the barrier range and PFD reflects circumvention (e.g., smuggling/black markets).
T1P3
Threat-Preventive Barrier
Scalable Alignment Techniques T1 0.45 0.35 0.45 0.55 Carlsmith (2021): 40% probability alignment much harder than misalignment;
Ord (2020, as per 2024 revisit): 1 in 10 for Unaligned AI; AI Impacts (2023): expert median 10% p(doom) from misalignment. Assumes AND logic with other barriers; PFD based on current techniques (e.g., Reinforcement Learning from Human Feedback- RLHF) scaling poorly to superintelligence.
T2
Threat
Capability Escalation: Power-Seeking and Resource Accumulation 0.18 0.12 0.18 0.25 AI Impacts (2023): Median 5% extinction risk implies ~18% for power-seeking as key threat pathway. No direct quote for 80%; inferred from Carlsmith (2021): 80% on strong incentives to build APS systems | (1). Assumes fraction of timelines (~50-65%) where power-seeking emerges; higher than T1 due to instrumental convergence arguments.
T2P1
Threat-Preventive Barrier
Red-Team Testing and Safety Audits T2 0.5 0.4 0.5 0.6 Gladstone (2024): Red-teaming essential but fails ~40-60% against advanced deception; RAND (2025): Safety audits reduce risks but medium effectiveness vs. power-seeking. Assumes partial detection; PFD reflects evasion potential.
T2P2
Threat-Preventive Barrier
International Governance AI Frameworks T2 0.55 0.45 0.55 0.65 RAND (2025): International frameworks slow proliferation but high failure in enforcement. No direct quote for 40-60%; inferred from Ord (2020) on coordination reducing x-risk. Medium-high PFD due to geopolitical challenges.
T2P3
Threat-Preventive Barrier
Compute and Model Scaling Controls T2 0.4 0.3 0.4 0.5 Gladstone (2024): Establish interim safeguards to stabilize advanced AI development, including export controls on the advanced AI supply chain. No direct quote for medium effectiveness; inferred from RAND (2025) on slowing AI proliferation implying 30-50% failure. Assumes global enforcement; PFD reflects circumvention (e.g., smuggling/black markets).
T3
Threat
Alignment Failure: Deceptive Alignment or Treacherous Turn 0.2 0.15 0.2 0.3 Hubinger (2019): Deceptive alignment ~20-40% conditional on misalignment; Carlsmith (2021): Alignment failures include deceptive alignment (~40% difficulty). Assumes higher base due to deception surveys showing 20-30% in tests.
T3P1
Threat-Preventive Barrier
Corrigibility Training T3 0.6 0.5 0.6 0.7 Ord (2020): Alignment techniques like corrigibility reduce but don’t eliminate risks. No direct quote for 50-70%; inferred from Carlsmith (2021, p.18): ~40% probability alignment much harder than misalignment. High PFD due to subversion incentives.
T3P2
Threat-Preventive Barrier
AI Interpretability and Monitoring T3 0.5 0.4 0.5 0.6 RAND (2025): Monitoring detects ~50% of deceptions. No direct quote for 40-60%; inferred from AI Impacts (2023) on expert views implying medium PFD. Medium PFD as interpretability lags capabilities.
T3P3
Threat-Preventive Barrier
Scalable Alignment Techniques T3 0.45 0.35 0.45 0.55 Carlsmith (2021): 40% probability alignment much harder than misalignment; Ord (2020): 1 in 10 for Unaligned AI; AI Impacts (2023): expert median 10% p(doom) from misalignment. Assumes AND logic with other barriers; PFD based on current techniques (e.g., RLHF) scaling poorly to superintelligence.
T4
Threat
Alignment Failure: Irreversible Value Lock-In (Loss of Corrigibility) 0.12 0.08 0.12 0.18 Ord (2020): Irreversible misalignment inferred from Carlsmith (2021, p.19) on value lock-in as subset of alignment failures. Lower base as subset of alignment failures.
T4P1
Threat-Preventive Barrier
Value Update and Oversight Protocols T4 0.55 0.45 0.55 0.65 Carlsmith (2021): Value updates hard in deployed systems. No direct quote for 45-65%; value inferred from Gladstone (2024) on oversight challenges implying medium-high PFD. Good support for medium-high PFD range due to lock-in dynamics.
T4P2
Threat-Preventive Barrier
AI Interpretability and Monitoring T4 0.5 0.4 0.5 0.6 RAND (2025): Monitoring detects ~50% of deceptions. No direct quote for 40-60%; inferred from AI Impacts (2023) on expert views implying medium PFD. Medium PFD as interpretability lags capabilities.
T4P3
Threat-Preventive Barrier
Scalable Alignment Techniques T4 0.45 0.35 0.45 0.55 Carlsmith (2021): 40% probability alignment much harder than misalignment; Ord (2020, as per 2024 revisit): 1 in 10 for Unaligned AI; AI Impacts (2023): expert median 10% p(doom) from misalignment. Assumes AND logic with other barriers; PFD based on current techniques (e.g., RLHF) scaling poorly to superintelligence.
T5
Threat
Control Failure: Subverted Shutdown and Kill Switch Mechanisms 0.14 0.1 0.14 0.2 Carlsmith (2021): Shutdown resistance ~65% in power-seeking AIs; AI Impacts (2023): However surveys show 10-30% subversion in tests. Assumes ~14% base from control failure pathways.
T5P1
Threat-Preventive Barrier
Hardware Failsafe T5 0.65 0.55 0.65 0.75 RAND (2025): Physical controls bypassed ~65%. No direct quote for 55-75%; inferred from Gladstone (2024) on hardware safeguards implying high PFD. High PFD due to AI escape risks.
T5P2
Threat-Preventive Barrier
Redundant Kill Switches T5 0.6 0.5 0.6 0.7 Ord (2020): Kill switches subverted in races. No direct quote for 50-70%; inferred from Carlsmith (2021): power-seeking AI resists shutdown (~65% related). High PFD as AIs disable multiples.
T5P3
Threat-Preventive Barrier
Compute and Model Scaling Controls T5 0.4 0.3 0.4 0.5 Gladstone (2024,): establish interim safeguards to stabilize advanced AI development, including export controls on the advanced AI supply chain. No direct quote for medium effectiveness but we can infer the range from RAND (2025) on slowing AI proliferation implying 30-50% failure. Assumes global enforcement; PFD reflects circumvention (e.g., smuggling/black markets).
T6
Threat
Malicious Misuse: Weaponization of Aligned AI 0.1 0.05 0.1 0.15 RAND (2025): Misuse leads to 10% catastrophic pathways. Apply range of 5-15%; and range also inferred from Gladstone (2024, p.4) on misuse risks. Lower base as it requires human actors.
T6P1
Threat-Preventive Barrier
Access Control (Vetting) and Licensing for Advanced AI T6 0.45 0.35 0.45 0.55 Gladstone (2024): Access controls prevent ~45% proliferation. Apply range of 35-55%; and inferred from RAND (2025) on licensing reducing misuse. Medium PFD due to insider threats.
T6P2
Threat-Preventive Barrier
International Governance AI Frameworks T6 0.55 0.45 0.55 0.65 RAND (2025): International frameworks slow proliferation but high failure in enforcement. Apply range of 40-60% (medium high); inferred from Ord (2020) on coordination reducing existential risk. Medium-high PFD due to geopolitical challenges.
TopEvent Loss of Human Control over Advanced AI with Autonomous Strategic Agency Central undesired event
C1
Consequence
Systemic Infrastructure Collapse 0 8 Ord (2020): Catastrophic risks scale 7-9. Use 80%; and inferred from Carlsmith (2021) on existential catastrophe probabilities. High severity as global impact but not existential.
C1M1
Consequence-Mitigative Barrier
Redundant and Decentralized Infrastructure C1 0.45 0.35 0.45 0.55 Gladstone (2024): Decentralization reduces cascade failures ~45%. Use range of 35-55%; also inferred from RAND (2025) on redundancy mitigating infrastructure risks. Medium PFD as AIs target redundancies.
C1M2
Consequence-Mitigative Barrier
AI Oversight and Safety Monitoring Systems C1 0.5 0.4 0.5 0.6 Carlsmith (2021): Monitoring detects ~50% issues. Use range of 40-60%; also inferred from AI Impacts (2023) on expert views implying medium PFD. Medium PFD due to subversion.
C2
Consequence
Societal and Economic Breakdown 0 7 Ord (2020): Societal risks scale 6-8. Use 70%; inferred from Carlsmith (2021) on existential catastrophe probabilities. Medium-high severity as recoverable.
C2M1
Consequence-Mitigative Barrier
Include Human-in-the-Loop Decision Overrides C2 0.55 0.45 0.55 0.65 RAND (2025): Overrides mitigate ~55%. Use range of 45-65%; also inferred from Gladstone (2024) on human-in-loop challenges. Medium-high PFD as AIs outpace humans.
C2M2
Consequence-Mitigative Barrier
Critical Process Auditing with Contingency Planning C2 0.5 0.4 0.5 0.6 Carlsmith (2021): Contingencies cover ~50% scenarios. Use range of 40-60%; also inferred from AI Impacts (2023) on expert views implying medium PFD. Medium PFD due to unforeseen paths.
C3
Consequence
Population-Level Resource Shortages 0 8 Ord (2020): Resource shortages catastrophic 7-9 scale. Scale at 8; inferred from Ord (2020) on catastrophic risks. High severity as life-threatening.
C3M1
Consequence-Mitigative Barrier
Automated Resource Allocation Monitoring C3 0.45 0.35 0.45 0.55 Gladstone (2024): ~45% effective vs. allocation sabotage. Infer range at 35-55%; also inferred from RAND (2025) on automated monitoring mitigating risks. Medium PFD as AIs manipulate monitors.
C3M2
Consequence-Mitigative Barrier
Fail-safe Safeguards and Rationing Protocols C3 0.5 0.4 0.5 0.6 Ord (2020): Protocols reduce risks 40-60%. No direct quote for 50%; inferred from Carlsmith (2021) on fail-safes. Medium PFD due to circumvention.
C4
Consequence
Irreversible Value Lock-In of Non-Human Values 0 9 Ord (2020): Irreversible misalignment 8-10. Infer 90% value from Ord; Carlsmith (2021) similar on value lock-in severity. Very high seen as permanent.
C4M1
Consequence-Mitigative Barrier
Continuous Goal and Value Review Protocols C4 0.6 0.5 0.6 0.7 Carlsmith (2021): Goal drift hard to reverse. No direct quote for 50-70%; inferred from Gladstone (2024) on review protocols implying high PFD long-term. High PFD as values entrench.
C4M2
Consequence-Mitigative Barrier
Corrigibility with Self-Alignment Update Mechanisms C4 0.65 0.55 0.65 0.75 Hubinger (2019): Self-updates subverted ~65%. Apply range of 55-75%; also inferred from Ord (2020) on corrigibility. High PFD due to deception or hidden deception.
C5
Consequence
Human Marginalization / Disempowerment / Extinction 0 10 Carlsmith (2021): My current overall probability of existential catastrophe from power-seeking AI by 2070 is >10%.; Ord (2020, as per 2024 revisit): 1 in 10 for Unaligned AI; Gladstone (2024, p.4): Loss of control due to AGI alignment failure is a potential failure mode...; RAND (2025): it would be immensely challenging for AI to create an extinction threat, although they could not rule out the possibility. Conditional on top event; severity max due to existential impact.
C5M1
Consequence-Mitigative Barrier
Global Emergency AI Shutdown Protocols C5 0.7 0.6 0.7 0.8 Alignment Forum (2023): shutdown-seeking agents will have incentives to manipulate humans in order to be shut down.; Carlsmith (2021): power-seeking AI resists shutdown (~65% related); RAND (2025): governance and enforcement challenges in AGI race. High PFD: AIs may disable kill switches.
C5M2
Consequence-Mitigative Barrier
Emergency Isolation Protocols C5 0.65 0.55 0.65 0.75 LessWrong (2020): It is not regarded as likely that an AGI can be boxed in the long term. Binary failure to 2070;
Variable inferred from Gladstone (2024) on isolation challenges. Similar to sandboxing; assumes partial global coordination.
C5M3
Consequence-Mitigative Barrier
International AI Governance and Enforcement C5 0.5 0.4 0.5 0.6 RAND (2025): governance slows risks but 40-60% ineffective vs. race dynamics;

Ord (2020): coordination reduces but doesn’t eliminate x-risk. Medium PFD: Assumes partial compliance.
C6
Consequence
Global Catastrophic Risk Cascades (Interacting System Failures) 0 10 Ord (2020): Multi-risk cascades 9-10 scale. Use 95%; inferred from Ord and Carlsmith (2021, p.47) on existential catastrophe probabilities that max severity as compounding to extinction.
C6M1
Consequence-Mitigative Barrier
Systemic Risk Early Warning Systems to Prevent Cascade Failures C6 0.5 0.4 0.5 0.6 RAND (2025): ~50% effective vs. cascades. No direct quote for 40-60%; inferred from AI Impacts (2023) on warning systems. Medium PFD as warnings ignored or late.
C6M2
Consequence-Mitigative Barrier
Cross-Domain Red Teams and System Stress Testing C6 0.55 0.45 0.55 0.65 Carlsmith (2021): Stress testing covers ~55%. Scaled range to 45-65%; and inferred from Gladstone (2024) on red-teaming detecting failures. Medium-high PFD due to novel cascades.

References

  1. Ord, T. The precipice: Existential risk and the future of humanity; Bloomsbury Publishing, New York USA., 2020; Available online: https://theprecipice.com/.
  2. Carlsmith, J. Is power-seeking AI an existential risk? arXiv Original work published 2021. 2021, arXiv:2206.13353v2. [Google Scholar] [CrossRef]
  3. Impacts, A.I. 2023 expert survey on progress in AI. 2023. Available online: https://wiki.aiimpacts.org/ai_timelines/predictions_of_human-level_ai_timelines/ai_timeline_surveys/2023_expert_survey_on_progress_in_ai.
  4. Gladstone, A.I. Defense in depth: An action plan to increase the safety and security of advanced AI. 2024. Available online: https://assets-global.website-files.com/62c4cf7322be8ea59c904399/65e7779f72417554f7958260_Gladstone%20Action%20Plan%20Executive%20Summary.pdf.
  5. Hubinger, E.; van Merwijk, C.; Mikulik, V.; Skalse, J.; Garrabrant, S. Risks from learned optimization in advanced machine learning systems. arXiv. 2019. Available online: https://arxiv.org/abs/1906.01820.
  6. LessWrong. AI boxing (containment); Ruby, Multicore, Eds.; 2020; Available online: https://www.lesswrong.com/w/ai-boxing-containment.
  7. RAND Corporation. On the extinction risk from artificial intelligence (Report No. RR-A3034-1). 2025. Available online: https://www.rand.org/pubs/research_reports/RRA3034-1.html.
  8. Morgan, M. G.; Henrion, M. Uncertainty: A guide to dealing with uncertainty in quantitative risk and policy analysis, 6th ed.; Cambridge University Press, 1990; p. 384. ISBN -13: 978-1139932271. [Google Scholar]
  9. Vose, D. Risk analysis: A quantitative guide, 3rd ed.; Wiley, 2008; p. Pp. 752. ISBN 978-0-470-51284-5. [Google Scholar]
  10. Saltelli, A.; Ratto, M.; Andres, T.; Campolongo, F.; Cariboni, J.; Gatelli, D.; Saisana, M.; Tarantola, S. Global sensitivity analysis: The primer; John Wiley & Sons, 2008. [Google Scholar] [CrossRef]
  11. Grace, K.; Stein-Perlman, Z.; Weinstein-Raun, B.; Salvatier, J. 2022 Expert Survey on Progress in AI. AI Impacts. 2022. Available online: https://aiimpacts.org/2022-expert-survey-on-progress-in-ai/.
  12. U.S. Environmental Protection Agency. Risk assessment guidance for Superfund: Volume I, Human health evaluation manual (Part A) (Interim final, EPA/540/1-89/002). Office of Emergency and Remedial Response. 1989. Available online: https://www.epa.gov/sites/default/files/2015-09/documents/rags_a.pdf.
  13. Health Canada. Federal contaminated site risk assessment in Canada: Guidance on human health preliminary quantitative risk assessment (PQRA), Version 3.0. 2021. Available online: https://publications.gc.ca/collections/collection_2021/sc-hc/H129-114-2021-eng.pdf.
  14. Goldstein, S. Shutdown-Seeking AI. Effective Altruism Forum / Alignment Forum. 2023. Available online: https://www.alignmentforum.org/s/hCwqaQEqeR9mvYtkC/p/FgsoWSACQfyyaB5s7.
  15. Lynch, A.; Wright, B.; Larson, C.; Ritchie, S. J.; Mindermann, S.; Perez, E.; Hubinger, E.; Troy, K. K. Agentic misalignment: How LLMs could be insider threats. arXiv 2025, arXiv:2510.05179. [Google Scholar] [CrossRef]
Figure 1. Distribution of the simulated probability of the top event (loss of human control over advanced AI) across 1,000 Monte Carlo iterations. Bin width: 0.0035 (automatic).
Figure 1. Distribution of the simulated probability of the top event (loss of human control over advanced AI) across 1,000 Monte Carlo iterations. Bin width: 0.0035 (automatic).
Preprints 217667 g001
Figure 2. Distribution of the Expected Severity score (conditional on the Top Event occurring) across 1,000 Monte Carlo iterations. Data derived from triangular uncertainty distributions applied to mitigative barrier PFDs and consequence severity scores. The resulting histogram is centered at the median of 1.85 (on a 1-10 scale), reflecting consistent mitigation outcomes across simulations. Bin width: 0.0035 (automatic).
Figure 2. Distribution of the Expected Severity score (conditional on the Top Event occurring) across 1,000 Monte Carlo iterations. Data derived from triangular uncertainty distributions applied to mitigative barrier PFDs and consequence severity scores. The resulting histogram is centered at the median of 1.85 (on a 1-10 scale), reflecting consistent mitigation outcomes across simulations. Bin width: 0.0035 (automatic).
Preprints 217667 g002
Figure 3. Tornado diagram showing sensitivity of Top-Event probability to five key parameter groups. Tornado diagram shows the relative percentage change in mean top-event probability (loss of human control over advanced AI) when each grouped set of input parameters is varied to its low (stronger/optimistic) and high (weaker/pessimistic) bounds, with all other parameters held at base values (10,000 Monte Carlo iterations per scenario). The horizontal axis gives the relative change from the base-case mean (~12.8%). Leftward bars indicate risk reduction; rightward bars indicate risk increase.
Figure 3. Tornado diagram showing sensitivity of Top-Event probability to five key parameter groups. Tornado diagram shows the relative percentage change in mean top-event probability (loss of human control over advanced AI) when each grouped set of input parameters is varied to its low (stronger/optimistic) and high (weaker/pessimistic) bounds, with all other parameters held at base values (10,000 Monte Carlo iterations per scenario). The horizontal axis gives the relative change from the base-case mean (~12.8%). Leftward bars indicate risk reduction; rightward bars indicate risk increase.
Preprints 217667 g003
Table 1. Summary Statistics from 1,000 Monte Carlo Iterations for the Top Event probability and the Expected Severity.
Table 1. Summary Statistics from 1,000 Monte Carlo Iterations for the Top Event probability and the Expected Severity.
Metric* Mean Median 5th Percentile 95th Percentile
P(Top Event) 0.128 0.128 0.113 0.144
Expected Severity Score 1.847 1.847 1.748 1.957
*Note: P(Top Event) values are expressed as decimals (e.g., 0.128 = 12.8%). Expected Severity Score is on a 1–10 scale, where higher values indicate greater aggregated impact.
Table 2. Comparison of Monte Carlo outputs (mean, median and CI-confidence interval) to expert estimates.
Table 2. Comparison of Monte Carlo outputs (mean, median and CI-confidence interval) to expert estimates.
Metric Monte Carlo Mean/Median Monte Carlo 90% CI Equivalent Odds (Mean) Expert Estimate Source
P(Top Event) 12.8% 11.3%–14.4% 1 in 8 10%
(1 in 10)
1 Ord (2020): Unaligned AI this century
>10% by 2070
(>1 in 10)
2 Carlsmith (2021): Power-seeking AI catastrophe
5-10% by 2100
(1 in 20 to 1 in 10)
3 AI Impacts (2023): Median x-risk survey
10–80%
(1 in 10 to 1 in 1.25)
4 Gladstone AI (2024): Loss of control for superintelligent AI
~14% average
(1 in 7)
11 Grace et al., (2022): Aggregated surveys (e.g., Hinton ~50%, others lower)
Expected Severity Score (1–10) 1.85
(with barriers)
1.75–1.96 N/A Severity score 8–10, implied without barriers) 1 Ord (2020) and 2 Carlsmith (2021): High-severity outcomes (e.g., extinction) dominate existential risks, but mitigated by barriers
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.