Submitted:
14 September 2026
Posted:
15 September 2026
You are already at the latest version
Abstract
Why is machine text detectable? Using a calibrated probe instrument on length-controlled corpora, we trace detectability along two causal axes and find they point in opposite directions. Across staged post-training checkpoints (base → SFT → DPO → RL), detectability rises through DPO (d′: 0.58→2.18) and consolidates at RL. Crucially, the rise is anchor-mediated: free generation is nearly flat (range 0.31 in d′), while source-anchored rewriting climbs steeply. The dominant mechanism is source-anchor release, not intrinsic style divergence. Across scale at fixed pretraining data (Pythia 70M–6.9B), detectability instead shows a monotone decline (Spearman ρ=−0.900, p=0.037). Two opposing origins resolve the apparent paradox: pre-training scale pushes detectability down; post-training alignment pushes it up. The staircase replicates on OLMo-2, again monotone up, confirming the pattern is not family-specific.

Keywords:
AI-generated text detection
; large language models
; post-training alignment
; detectability
; RLHF
; representation probing
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.