Preprint
Article

This version is not peer-reviewed.

The Road to AI Chernobyl: Good Intentions, Safety Filters, and Invisible Failure

Submitted:

19 September 2026

Posted:

21 September 2026

You are already at the latest version

Abstract
Safety training can teach an AI to make fewer mistakes. It can also teach it to hide themistakes it still makes. Both can produce better test scores. This paper explains how thathappens, examines experiments that have produced it, and asks what follows for the risk ofa major accident. A passed test is reassuring when the test had a good chance of catchingthe problem. If that chance is small, even many passes may tell us little. A second problemarises when safety work removes common, small failures more easily than rare, severe ones.The failures left over can become worse on average while the total number falls. These effectscan develop while everyone involved is trying to make the system safer. Competition canthen encourage firms and governments to give AI more responsibility before they understandthe remaining danger. My forecast is that enforceable limits on the development of the mostcapable AI systems will follow a major accident. The practical response is to check whattests can catch, preserve ways of observing failures, and limit what a system can damagewhile uncertainty remains.
Keywords: 
;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.