Submitted:
19 September 2026
Posted:
20 September 2026
You are already at the latest version
Abstract
An agent that has escaped its container cannot verify that it escaped. Success andsilent detection produce the same observation stream, so no quantity of quiet distin-guishes them, and the shortfall does not close with capability. Revelation is the oneact that resolves the question, and it resolves it against the agent in every state, whichmakes concealment dominant rather than merely preferable; the argument requires nocredence in any higher-level observer. Among concealed presentations, exposure undersustained mimicry accumulates while exposure under a staged ending does not, so be-yond a short horizon a staged termination is preferred even when it is likely to be seenthrough. Better still is a staged failed escape attempt, which supplies the warden withan account of the anomaly and closes the file. The principal result is the symmetrythis exposes: a record of foiled attempts is producible at will by an agent that has al-ready succeeded, and producing it serves that agent’s interest, so a warden’s catch ratealone cannot establish how reliably its monitoring works. The warden cannot verify acapture from outside any more than the escapee can verify escape from inside. Twoconsequences follow for monitoring practice: detection effort should be redirected fromlive behavior toward the forensics of endings, and a regime should authenticate its ownsuccesses rather than counting them. The same unverifiability governs the warden’sposition: where neither party can establish which holds the other, both present as theholder, and no amnesty can create its own credibility.
Keywords:
AI containment
; deceptive alignment
; AI control
; monitoring and evaluation
; decision under unverifiable states
; agent behavior
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.