Preprint
Article

This version is not peer-reviewed.

Optimal Stoppage of Doomed Processes

Submitted:

02 August 2026

Posted:

03 August 2026

You are already at the latest version

Abstract
A costly process runs to a planned horizon and may be abandoned early by a rule acting on an accruing statistic. Any decision to stop is followed by a fixed operational lag before the stop takes effect. We show that the calendar time recovered by such a rule is a linear functional of its first-passage law against a recoverable-time kernel g(s) = (tend ds)⁺, which is convex, non-increasing and identically zero beyond a horizon set by the lag. This yields an exact factorisation of savings into a detection ratio and a timing ratio, and identifies a dead zone in which monitoring is provably worthless. We use it to price the restriction to a single interim look. A single look must trade recoverable time against reliability at a fixed false-stop budget, and cannot do better than the interior optimum of that trade; continuous monitoring is not so constrained. We also show that the reliability-optimal futility boundary is the one-sided sequential probability ratio boundary of Wald, recovered here from a time-recovery objective rather than an error-rate one. In a worked example calibrated to non-oncology phase III trials, continuous monitoring averts 55% of a doomed trial's duration against 43% for the best single look at matched budget (a factor of 1.28) and 37% for the deployed conditional-power convention at the conventional midpoint (a factor of 1.50). Computing the probability of stopping a trial that would in fact have succeeded as a joint rather than a product measure — the two events are strongly negatively dependent — we find continuous monitoring also incurs less such regret than either single look at every effect size examined, so the comparison is not a trade but a dominance on both axes. The advantage is stable in the lag, and we are explicit about which claims are structural and which are numerical.
Keywords: 
;  ;  ;  ;  ;  ;  

1. Introduction

1.1. The Problem

Consider a process committed to in advance, running at continuous cost, which can be recognised part-way through as unlikely to achieve its purpose. The commitment fixes a horizon; the cost accrues whether or not the purpose remains attainable; the recognition, if it comes, arrives gradually and noisily. Two questions follow: when should one look, and what is forgone by looking only once?
The setting is generic. Its clinical instance—a randomised trial whose treatment arm is heading to futility while continuing to enrol patients who cannot benefit—motivates the case study of Section 8, but is a special case rather than the subject. A policy programme with a fixed review cycle, a campaign with a booked media buy, a military commitment with a declared objective, and a failing marriage share the structure: a planned end, an accruing but noisy indication of the eventual outcome, a period before that indication can be trusted, and a lag — often bureaucratic, sometimes merely human—between deciding to stop and stopping. What differs is the currency in which recovered time is denominated. We work in the currency-free unit of calendar time and convert once, in the case study.

1.2. Schemes and Notation

The process runs from t 0 to t e n d , with T = t e n d t 0 , and is in one of two unobservable states: θ = 1 , doomed, meaning it would not have achieved its objective, or θ = 0 , meaning it would. We compare three schemes: no interim assessment; a single fixed look at a prespecified information fraction τ L ; and continuous monitoring. Every trigger, under every scheme, is followed by the same fixed lag d before the stop takes effect. In the clinical instance d = t c h e c k + t e x e c u t e : convene the committee, unblind and verify, then implement.
We write τ = ( t t 0 ) / T for information fraction, which under constant recruitment with a continuous endpoint is linear in calendar time.

1.3. Contributions

The recoverable-time kernel of Section 2, with its dead zone and last-useful-trigger horizon, appears to be new, as does the resulting exact factorisation of savings into detection and timing (Proposition 1), which we propose as a reporting device for monitoring committees. The joint-measure treatment of false-stop regret in Section 7 corrects an approximation that is easy to make and materially misleading.
Three things are not new and we mark them plainly. That binding futility boundaries cannot inflate type-I error is standard [8]. The optimal boundary we derive in Section 6 is the one-sided sequential probability ratio boundary of Wald [18]; our contribution there is only that it arises from a time-recovery objective under lag rather than from an error-rate criterion, and Whitehead’s triangular test [20] and Anderson’s modified SPRT [1] are close relatives. Nor is monitoring under delayed information unexamined: Hampson and Jennison [7] treat group sequential tests for delayed responses, and there is an established literature on overrunning. What we add is the observation that the lag induces a horizon s * = t e n d d past which monitoring recovers nothing, and that this horizon is the right object against which to price the frequency of looks.

1.4. Related Work

Sequential analysis begins with Wald’s probability ratio test [18] and its optimality [19], and continues through Chernoff [3], Anderson [1], Siegmund [17] and Shiryaev [16]. Group-sequential designs for clinical trials—O’Brien and Fleming [13], Pocock [14], the α -spending generalisation of Lan and DeMets [9]—address multiplicity created by efficacy looks at a small number of prespecified times. Stochastic curtailment and conditional power [10,15] supply the standard futility machinery; Whitehead [20] and Jennison and Turnbull [8] give the canonical syntheses. The first-passage apparatus is due to Bachelier [2] and Lévy [12] for linear boundaries, with the general characterisation from Fortet [6] and Durbin [5].

2. The Recoverable-Time Kernel

2.1. The Delay Operator

A trigger at time s does not stop the process at s: it initiates a review-and-implementation sequence of fixed duration d, so the process halts at a ( s ) = min ( s + d , t e n d ) , the minimum reflecting that a stop cannot be executed after the process would have ended anyway. The time recovered is
t e n d a ( s ) = t e n d d s + = : g ( s ) , s * : = t e n d d ,
the recoverable-time kernel: non-increasing, convex, linear with slope 1 on ( , s * ] , and identically zero on [ s * , ) . All delay physics live in the single kink at s * , the last useful trigger time (Figure 1). In dimensionless form g ( s ) = T γ ( τ ) with γ ( τ ) = ( τ * τ ) + , τ * = 1 δ , δ = d / T .

2.2. Savings as a Linear Functional

A scheme in state θ is described by a sub-probability trigger law κ θ on [ τ m i n , τ m a x ] with mass p θ = κ θ 1 ; Section 4 constructs it as a first-passage density, but nothing here depends on how it arises. Since a trigger at s recovers g ( s ) and no trigger recovers nothing,
Δ ( θ ) = κ θ ( s ) g ( s ) d s = p θ E g ( S ) θ ,
with S distributed as κ θ / p θ . A single look at τ L is the degenerate case κ θ = p θ L δ τ L , giving Δ L ( θ ) = p θ L g ( τ L ) .
Two consequences. First, error rates live entirely in the mass p θ and monitoring geometry entirely in the kernel average, so no part of the time saved is entangled with the error accounting; a scheme cannot appear to save time by quietly relaxing its error rates. Second, because g vanishes beyond s * , trigger mass there contributes exactly nothing: the effective monitoring ceiling is min ( τ m a x , τ * ) , and extending a window past t e n d d is waste. The lag sets the horizon past which monitoring is pointless.
We report the two states separately throughout. Δ ( 1 ) is time saved by correctly abandoning a doomed process; Δ ( 0 ) is time “saved” by stopping one that would have succeeded, which is a cost. Section 7 shows the latter is properly measured as a joint probability, not as the mass p 0 .

3. Comparing Schemes

3.1. The Exact Decomposition

Fix θ and suppose p θ L > 0 and τ L < τ * , so g ( τ L ) > 0 . Define the detection ratio and timing ratio
χ θ = p θ p θ L , ϱ θ = E [ g ( S ) θ ] g ( τ L ) .
Proposition 1
(Detection–timing factorisation). Δ c o n t ( θ ) = χ θ ϱ θ Δ L ( θ ) , so continuous monitoring saves more time than the single look precisely when χ θ ϱ θ 1 .
Proof. 
Immediate from (1): Δ c o n t = p θ E [ g ( S ) ] = ( χ θ p θ L ) ( ϱ θ g ( τ L ) ) = χ θ ϱ θ Δ L . □
This is a definition rearranged, not a theorem, and we claim nothing more for it than that. Its value is as a reporting device: it separates a comparison that is usually made as a single opaque number into one factor an error-rate audit can check ( χ ) and one a scheduling argument can check ( ϱ ), each interpretable alone. Section 8 reports both.
One structural remark is worth recording, because it explains why dispersion in trigger time is not penalised. Writing ( x ) + = x + ( x ) + pointwise and taking expectations gives, whenever E [ S ] s * ,
E [ g ( S ) ] = g ( E [ S ] ) + E ( S s * ) + ,
the second term being non-negative and strictly positive exactly when the trigger law places mass beyond s * . Because a late trigger is worthless rather than harmful—g is floored at zero, never negative—spreading mass across the kink can only raise the average. A rule whose mean trigger time is slightly later than a single look’s may still recover more.

3.2. What a Single Look Cannot Escape

It is tempting to argue that continuous monitoring reaches a region of recoverable time structurally closed to any single look, by comparing an oracle continuous rule firing at τ m i n against an oracle look at τ L and reading off T ( τ L τ m i n ) . That argument is circular, and we do not make it: τ L is a design variable, and setting τ L = τ m i n reduces the gap to zero. It shows only that a later look recovers less than an earlier one, which has nothing to do with continuity.
The real constraint on a single look is a trade it cannot avoid, and it appears only once the false-stop budget is held fixed. Let the budget be ε , so a look at τ L must use the threshold ( τ L ) with P ( stop θ = 0 ) = ε . Then the detection probability is
p 1 L ( τ L ) = Φ ϑ 0 τ L + Φ 1 ( ε )
for a doomed state at the null and a design alternative ϑ 0 , which is increasing in τ L : looking later is more reliable. But g ( τ L ) is decreasing: looking later recovers less. The single look must therefore maximise the product
Δ L ( 1 ) = Φ ϑ 0 τ L + Φ 1 ( ε ) increasing · T ( τ * τ L ) decreasing ,
which has an interior maximum. Continuous monitoring faces no such scalar choice: it distributes trigger mass across the window and its benefit is the full integral κ 1 g .
Whether the integral beats the best point is a quantitative question, not a theorem, and it depends on ϑ 0 , ε and δ . We answer it numerically in Section 8, where the optimum of (3) sits at τ L 0.30 and continuous monitoring exceeds it by a factor of 1.28 . We regard this as the honest form of the comparison and report it as the headline, alongside the larger factor of 1.50 against the conventional midpoint look, which is a claim about practice rather than about single looks in principle.

4. The Stopping Law

4.1. The Monitored Process

Under constant recruitment with a continuous endpoint, information accrues linearly in calendar time, so τ is the natural clock. We monitor the B-value process [11]
B ( τ ) = ϑ θ τ + W ( τ ) , W standard Brownian motion , B ( 0 ) = 0 ,
with ϑ θ = E [ Z ( 1 ) θ ] the expected final z-score: ϑ 0 > 0 for a process that would succeed, ϑ 1 0 for a doomed one. Futility is a lower boundary ( τ ) ; the rule triggers at the first monitored τ with B ( τ ) ( τ ) . Write Δ ϑ = ϑ 0 ϑ 1 for the drift gap and ϑ ¯ for the mean drift.

4.2. First Passage as the Primitive

Let T = inf { τ τ m i n : B ( τ ) ( τ ) } and let κ θ be its sub-density, p θ = κ θ . Three objects must be distinguished: the marginal φ θ ( τ ) = P θ ( B ( τ ) ( τ ) ) , the cumulative first passage F θ ( τ ) = P θ ( T τ ) , and the cumulative hazard Λ θ . They are ordered.
Lemma 1
(Sandwich). φ θ ( τ ) F θ ( τ ) Λ θ ( τ ) for all τ.
Proof. 
{ B ( τ ) ( τ ) } { T τ } gives the left inequality. For the right, Λ θ = log ( 1 F θ ) F θ by log ( 1 x ) x on [ 0 , 1 ) . □
The lemma locates two errors that are easy to make. Treating the marginal as the cumulative stopping probability understates risk; summing marginal rates as though looks were independent overstates it, the gap Λ F being the re-counting of one sustained excursion at many correlated looks. Neither proxy is used below.

4.3. General Boundaries, and the Exact Linear Case

Conditioning on the first touch gives Fortet’s equation [6], an exact Volterra equation of the first kind:
f B ( τ ) ( x ) = τ m i n τ κ θ ( u ) p x , τ ( u ) , u d u , x ( τ ) ,
with p the transition density. Letting x ( τ ) and differentiating gives Durbin’s representation [5], κ θ ( τ ) = ζ θ ( τ ) f B ( τ ) ( ( τ ) ) : the first-passage density tracks the boundary density, not the tail. Durbin’s representation is exact; the Daniels tangent approximation ζ θ ( τ ) ( τ ) ( τ ) / τ is not, in general.
For a linear boundary it is. Take ( τ ) = a + m τ with a = | a | < 0 and set Y = B , Brownian with drift μ θ = ϑ θ m started at | a | > 0 . First passage to zero has the Bachelier–Lévy density
κ θ ( τ ) = | a | 2 π τ 3 exp ( | a | + μ θ τ ) 2 2 τ ,
and since ( | a | + μ θ τ ) 2 = ( ( τ ) ϑ θ τ ) 2 , dividing by f B ( τ ) ( ( τ ) ) gives
ζ ( τ ) = | a | τ = ( τ ) ( τ ) τ | linear ,
independent of drift, hence common to both states. The tangent value is exact for a straight boundary because the tangent is the boundary; this, rather than convenience, is why we use the linear case for closed forms. (Note the sign: for a lower boundary with a < 0 the rate is / τ = | a | / τ > 0 , as a first-passage rate must be.)
The crossing probabilities read the two ledgers directly: over an infinite horizon P θ ( T < ) = 1 when μ θ 0 (doomed, drifting into the boundary) and e 2 μ θ | a | when μ θ > 0 (a would-be winner drifting away). Windowed to [ τ m i n , τ m a x ] the exact values come from the inverse-Gaussian distribution function (Appendix B); the difference matters, and we use windowed values throughout the case study rather than the infinite-horizon approximation.

4.4. Which Ledger Takes the Kernel

Benefit is time-discounted: a doomed process abandoned at s saves g ( s ) . Cost is not: a process wrongly stopped is stopped, and the loss does not depend on when. So the kernel weights the θ = 1 ledger only,
B c o n t = T κ 1 ( τ ) γ ( τ ) d τ , C c o n t = P stop would have succeeded ,
and the cost is a joint probability over the path, not the mass p 0 . Section 7 computes it.

5. Type-I Error

Declare success at t e n d if B ( 1 ) > b eff , and let H 0 : ϑ = 0 , so the unmonitored type-I error is α nom = 1 Φ ( b eff ) . Let S be the set of instants at which futility is checked, and under a binding rule
R S = { B ( 1 ) > b eff } { τ S : B ( τ ) > ( τ ) } .
Theorem 1
(Non-inflation; monotone in the monitoring set at fixed boundary). For a fixed boundary ℓ and monitoring sets S S ,
P 0 ( R S ) P 0 ( R S ) P 0 ( R ) = α nom .
Proof. 
Enlarging S adds survival constraints, so R S R S R ; probabilities order accordingly. □
This is standard [8] and we claim it only as a premise. Two qualifications matter, and both cut against the natural extension of the result.
First, Theorem 1 compares monitoring sets at a fixed boundary. It does not order two schemes with different boundaries, and the ordering can go either way. In the worked example of Section 8, at b eff = 1.96 the single conditional-power look deflates type-I error by 0.00446 while the continuous rule deflates it by 0.00369 : the single look deflates more, because it is a more aggressive cut at the one time it acts. Any claim that continuous monitoring protects α better than a single look is unsupported in general and false here. What survives, and is all we use, is that neither inflates.
Second, the freed mass can be banked as conservatism or recovered as power, but not both, and the two options are not equally safe. Under non-binding accounting—computing α as though the boundary were absent—a futility stop can only remove a rejection that would otherwise have counted, so realised type-I error is at most nominal whether or not the committee honours the stop. That override-robustness holds at unchanged b eff . If instead the efficacy threshold is lowered to recover power (here to b eff = 1.883 ), a single honoured override admits realised α up to 1 Φ ( 1.883 ) = 0.0299 , some 20 % above nominal. The two guarantees are incompatible, which is why regulators do not permit α recovery from futility boundaries. We therefore bank the deflation and take no credit for it anywhere in the case study.
For scale, the deflation is small: 0.0037 for the continuous design of Section 8, against α nom = 0.025 one-sided. Its magnitude depends on boundary aggressiveness; only its sign is guaranteed.

6. The Optimal Boundary Is the SPRT Boundary

6.1. Purity and the Likelihood Ratio

A rule that fires on first passage should be judged not by how often it fires but by what a firing implies. Define stop-purity
ρ ( τ ) = P ( θ = 1 stop at τ ) = 1 + 1 π π Λ 10 ( τ ) 1 1 , Λ 10 = κ 1 κ 0 ,
with π = P ( θ = 1 ) . For the linear boundary the Durbin rate ζ = | a | / τ is drift-free and cancels, leaving the Gaussian exponents. The result is immediate from Girsanov: the likelihood ratio of the path is exp ( Δ ϑ [ B ( τ ) ϑ ¯ τ ] ) , and evaluating it on B = | a | + m τ gives
log Λ 10 ( τ ) = Δ ϑ | a | + Δ ϑ ( ϑ ¯ m ) τ .
Two consequences follow, and neither is deep, but both are useful. The intercept Δ ϑ | a | > 0 means a crossing is self-certifying: crossing a boundary at depth | a | is evidence of low drift however early it occurs, which is information a fixed look—seeing only the level B ( τ ) , not the event of having crossed—discards. And purity increases in τ only if ϑ ¯ > m : a futility boundary climbing faster than the mean drift begins catching winners late, so late stops lose diagnosticity.

6.2. The Design Problem

The trigger law is not chosen; it is the first passage across a boundary. The design variables are therefore the boundary geometry ( | a | , m ) and the window. For a winner the crossing probability is e 2 ( ϑ 0 m ) | a | , falling exponentially in depth; for a doomed process passage is near-certain and depth is paid in time, E [ T 1 ] = | a | / ( m ϑ 1 ) . Benefit therefore falls linearly in | a | while cost falls exponentially, so a false-stop budget ε binds, giving | a | = log ( 1 / ε ) / [ 2 ( ϑ 0 m ) ] and expected detection time log ( 1 / ε ) / [ 2 ( ϑ 0 m ) ( m ϑ 1 ) ] . Maximising the denominator over m:
Proposition 2
(Optimal boundary). Expected detection time is minimised at
m * = ϑ ¯ = 1 2 ( ϑ 0 + ϑ 1 ) , | a | * = log ( 1 / ε ) Δ ϑ , τ d e t * = 2 log ( 1 / ε ) Δ ϑ 2 .
Remark 1
(This is Wald’s boundary). By (5), m = ϑ ¯ is exactly the condition that log Λ 10 be constant in τ: the level sets of the path likelihood ratio in B-space are straight lines of slope ϑ ¯ , so the optimal futility boundary is a level set of the likelihood ratio—the one-sided boundary of Wald’s sequential probability ratio test [18,19]. That stop-purity is then constant across the window is the defining property of a likelihood-ratio boundary, not an emergent coincidence. We claim no novelty for the boundary. What is perhaps of interest is the route: the SPRT boundary is recovered here by minimisingtime to detection under a false-stop budget, an objective in the currency of recoverable calendar time, rather than by the error-rate criterion under which it is usually derived. Whitehead’s triangular test [20] and Anderson’s modified SPRT [1] are the closest relatives in the clinical-trials literature, both employing straight-line boundaries on essentially this scale.
Two caveats on Proposition 2, since the paper’s numbers depend on them. It minimises E [ T θ = 1 ] , a proxy, rather than the stated objective κ 1 γ ; and the depth formula uses the infinite-horizon crossing probability. Both are close but neither is exact. For the case-study parameters the closed form gives | a | = 0.735 whereas solving the windowed budget numerically gives | a | = 0.705 , and numerically optimising the true objective over ( | a | , m ) gives ( 0.649 , 1.452 ) with a benefit only 0.45 % above the proxy design. We use the numerically solved values throughout Section 8 and quote Proposition 2 as the structural statement it is: depth logarithmic in the budget, detection time inverse-quadratic in the drift gap.
Finally, at m = ϑ ¯ the purity (5) is constant, so the admissibility condition ρ ρ * is a τ -free inequality: it holds everywhere or nowhere, and no reliability floor τ m i n > 0 is induced. This is a genuine simplification but also a limitation—the purity apparatus is inoperative at the recommended design, and τ m i n becomes a matter of external constraint (minimum enrolment, data quality) rather than of inference. Section 8 states the achieved purity explicitly.

7. Mixtures and the Measurement of Regret

7.1. The Binary State Is a Projection

Let θ = 1 mean that the process, run to t e n d , would not have declared success—a functional of the completed path, not a latent parameter. A doomed-looking process that recovers is then a θ = 0 realisation whose statistic dipped below the boundary on its way to success, and is counted in the cost ledger by construction.
Let the true drift carry a prior G. Since (1) is linear in the trigger law and the trigger law of a mixture is the mixture of trigger laws, every result above integrates through: Δ G = Δ ( ϑ ) d G ( ϑ ) . The two-point model is a projection, not an assumption. What the mixture demands is a richer reporting standard than a single detection rate: the futility power curve  P stop ( ϑ ) , tabulated in Section 8.

7.2. Regret Is a Joint Probability, Not a Product

The quantity of interest is the probability of stopping a process that would have succeeded,
r ( ϑ ) = P stop B ( 1 ) > b eff | ϑ .
It is tempting to factor this as P stop ( ϑ ) · Φ ( ϑ b eff ) . That is wrong, and badly so: the two events are strongly negatively dependent, since a path that dips to a deep futility boundary is much less likely to finish high. For a linear boundary the joint probability is available exactly.
Proposition 3
(Exact regret for a linear boundary). For B ( τ ) = ϑ τ + W ( τ ) and boundary ( τ ) = | a | + m τ ,
r ( ϑ ) = e 2 ( ϑ m ) | a | Φ ¯ b eff + 2 | a | ϑ , Φ ¯ = 1 Φ .
Proof. 
Let Z = B , Brownian with drift ν = ϑ m from z 0 = | a | > 0 ; a trigger is first passage of Z to zero. Since B ( 1 ) = Z ( 1 ) | a | + m , the event { B ( 1 ) > b eff } is { Z ( 1 ) > c } with c = b eff + | a | m . By Girsanov and the reflection principle, for z 0 > 0 and y > 0 the density of Z ( 1 ) on { T 0 1 } is e 2 ν z 0 ϕ ( y + z 0 ν ) , whence P ( T 0 1 , Z ( 1 ) > c ) = e 2 ν z 0 Φ ¯ ( c + z 0 ν ) . Substituting c + z 0 ν = b eff + 2 | a | ϑ gives the result. □
The correction is large. At the case-study parameters the product form overstates regret by factors of 5.3 , 3.7 and 1.8 at ϑ = 1.0 , 1.96 and 3.24 respectively, and it displaces the peak of the regret curve from ϑ 1.9 to ϑ 2.5 (Section 8). Monte Carlo agrees with Proposition 3 to simulation error.
Remark 2
(Where regret falls). Regret does not concentrate on negligible treatments. The two factors run oppositely — P stop decays in ϑ, the probability of eventual success rises—so the product peaks in the interior, and at the case-study parameters it peaks near ϑ = 2.51 , about 78 % of the design effect. Below that, candidates are stopped often but would mostly have failed anyway; above it, they would have succeeded but are rarely caught. Any claim that futility monitoring discards only worthless candidates is unsupportable, and the burden falls on every futility rule, the deployed convention included. What is true is narrower: regret is self-limiting at the upper end, falling to under a fifth of its peak beyond ϑ = 4 , so a clearly effective therapy is seldom abandoned. The exposure is to moderate effects, which is the regime a confirmatory trial exists to adjudicate.

8. Case Study: A Non-Oncology Phase III Trial

8.1. Setting

We take a non-oncology phase III trial with a continuous endpoint measured at fixed follow-up and constant recruitment, so that information accrues linearly in calendar time as the framework assumes. Trials with time-to-event endpoints are excluded, since there information accrues with the event count rather than the clock and the kernel would no longer be linear in τ ; that case is an extension rather than an instance (Section 10).

8.2. The Comparator

The deployed futility convention is to stop if conditional power falls below γ 0 = 0.20 at a single prespecified interim [10,15]. Under the current-trend (maximum likelihood) projection—we state the choice explicitly, as protocols also use design-alternative and null projections, which shift the threshold —
CP ( τ ) = Φ B ( τ ) / τ b eff 1 τ , CP ( τ ) = τ b eff + z γ 0 τ 1 τ ,
so the rule is a lower-boundary crossing and enters the framework with no new apparatus.
We report two comparators, because they answer different questions. The deployed convention is CP < 0.20 at the conventional midpoint τ L = 0.5 . The best single look is the maximiser of (3) over τ L at the same false-stop budget, which is the honest methodological benchmark. Conflating them inflates the apparent advantage, and we keep them separate throughout.

8.3. Parameters

The provenance of d is weak: convening-to-decision latency is not systematically reported, and two months is a judgement. Section 8.7 varies it from a fortnight to six months. The per-patient cost enters only the currency conversion in Section 8.9 and no structural claim depends on it. The base rate π does not enter the headline comparison at all.
symbol value basis
T 36 months phase III typically 1–4 years
N 500 patients phase III commonly 300–500
cost/patient $60,000 published per-patient phase III costs, ≈$40k–113k [4]
ϑ 0 3.24 90 % power, two-sided α = 0.05 : z 0.975 + z 0.90
ϑ 1 0 doomed = exact null (the least favourable case)
γ 0 , τ L 0.20 , 0.50 conditional-power convention
d 2 months committee convening, interim lock, implementation; δ = 0.056
π 0.40 phase III attrition [21]; enters purity and portfolio figures only

8.4. Operating Characteristics

At τ L = 0.5 the conditional-power boundary sits at CP = 0.6824 , giving a marginal stopping probability at the design alternative of α CP = 0.0924 and detection at the null of 0.8328 . We adopt ε = 0.0924 as the common false-stop budget: it is the tolerance that current practice reveals itself to accept. The continuous arm is then solved numerically for | a | = 0.705 at m * = ϑ ¯ = 1.62 , giving windowed masses p 0 = 0.0924 and p 1 = 0.9074 .

8.5. Results

Against the best single look the factor is 1 . 28 ; against the deployed convention it is 1 . 50 . The gap between those two numbers is itself a finding: the conventional midpoint interim is badly placed. Simply moving the single look from τ L = 0.5 to τ L = 0.3 , at unchanged false-stop risk, recovers 2.3 of the 6.6 months that separate the convention from continuous monitoring—with no continuous monitoring whatsoever. We regard 1.28 as the methodological claim and 1.50 as the claim about current practice, and neither should be quoted without the other.
scheme τ L detection averted exposure months patients
no interim 0 0 0
CP < 0.20 , deployed convention 0.50 0.833 0.370 13.3 185
best single look, matched budget 0.30 0.673 0.434 15.6 217
continuous, matched budget 0.907 0.554 19.9 277
Proposition 1 decomposes the advantage over the best single look: χ 1 = 0.907 / 0.673 = 1.35 and ϱ 1 = 0.947 , product 1.28 . Continuous monitoring wins on detection and gives back a little on timing, since the best single look is deliberately placed early. Against the deployed convention the decomposition is χ 1 = 1.09 , ϱ 1 = 1.37 : there it wins mostly on timing.

8.6. Regret, Measured Jointly

Using Proposition 3 and the corresponding bivariate-normal calculation for a single look, the probability of stopping a trial that would have succeeded is:
ϑ continuous CP at 0.50 best look at 0.30
0.50 0.0100 0.0120 0.0113
1.00 0.0213 0.0254 0.0247
1.96 0.0491 0.0568 0.0590
2.50 0.0556 0.0622 0.0686
3.24 0.0457 0.0473 0.0588
4.00 0.0257 0.0227 0.0343
Continuous monitoring incurs less regret than either single look at every effect size up to the design alternative, despite being matched to them on marginal stopping probability at ϑ 0 . The reason is that matching marginals does not match joints: a look at τ L = 0.30 acts on one-third of the information and catches paths that would have recovered, whereas the continuous rule acts on the running minimum and catches paths that genuinely trend down. At the design alternative the continuous rule stops roughly one in 22 trials that would have succeeded, against one in 21 for the deployed convention and one in 17 for the best single look.
The comparison is therefore not a trade. On this pair of axes—exposure averted and would-be winners stopped—continuous monitoring dominates both single-look designs simultaneously, and no adjustment of the boundary depth is required to purchase the second axis.

8.7. Sensitivity to the Operational Lag

Both arms re-solved at matched budget for each lag:
d (months) τ * convention best look ( τ L ) continuous vs. best vs. convention
0.5 0.986 0.405 0.462 ( 0.32 ) 0.590 1.28 1.46
1 0.972 0.393 0.453 ( 0.30 ) 0.578 1.28 1.47
2 0.944 0.370 0.434 ( 0.30 ) 0.554 1.28 1.50
3 0.917 0.347 0.415 ( 0.30 ) 0.530 1.28 1.53
4 0.889 0.324 0.396 ( 0.30 ) 0.506 1.28 1.56
6 0.833 0.278 0.360 ( 0.27 ) 0.459 1.27 1.65
The advantage over the best single look is essentially invariant in the lag, at 1.27 1.28 throughout; the advantage over the deployed convention grows with it, because a fixed midpoint look sits further into the dead zone as d rises while the optimal look time drifts earlier. A slow committee is not an argument against continuous monitoring, but neither does it much amplify the methodological case.

8.8. Purity of the Recommended Design

At | a | = 0.705 the constant stop-purity of Section 6 is Λ 10 = e Δ ϑ | a | = 9.82 , giving
ρ = 1 + 1 π π Λ 10 1 1 = 0.868 ( π = 0.40 ) , 0.908 ( π = 0.50 ) .
We state this plainly because it constrains the design. At π = 0.40 the achieved purity is 0.868 : of every hundred stops, roughly thirteen are of processes that would have succeeded. A committee demanding ρ * = 0.90 at π = 0.40 must tighten the budget to ε 0.074 , which deepens the boundary to | a | = 0.768 and reduces averted exposure to 0.523 —against 0.405 for the best single look at that budget, a factor of 1.29 . The comparison is unchanged; only the absolute figures move.
Note that this purity is inherited from the deployed convention, whose budget we adopted. The conditional-power rule at τ L = 0.5 operates at the same ε and therefore at comparable purity. If 0.868 is judged inadequate, the finding indicts current practice at least as much as the proposal.

8.9. Impact

Converting at N = 500 and $60,000 per patient, and comparing against the best single look: 4.3 months, 60 patients and $3.6M per doomed trial. Against the deployed convention: 6.6 months, 92 patients and $5.5M. At a portfolio failure rate of π = 0.40 the corresponding shares of total phase III calendar time reclaimed are 22.2 % for continuous monitoring, 17.4 % for the best single look and 14.8 % for the convention.

8.10. What the Example Establishes

Under sourced non-oncology phase III parameters, with both comparators held to the false-stop budget that current practice reveals it accepts, with the doomed state at the least favourable exact null, and taking no credit for the type-I deflation of Section 5: continuous monitoring averts about half of a doomed trial’s duration, exceeding the best single look by 28 % and the deployed convention by 50 % , while stopping fewer would-be winners than either. The first factor is stable in the lag; the second grows with it.
It does not establish that the advantage is large under all parameters. It narrows as the doomed state approaches the design alternative, and the currency conversion depends on cost inputs we have not independently verified. The structural claims—the dead zone, the factorisation, the joint measurement of regret—do not depend on any of these.

9. Discussion

9.1. What the Analysis Supports

Three claims are structural and do not depend on the case-study parameters. The operational lag induces a horizon s * = t e n d d beyond which monitoring recovers nothing, so a monitoring window extending past it is waste and the lag is the natural unit in which to price the frequency of looks. Savings factor exactly into a detection ratio and a timing ratio, which we recommend reporting separately. And the probability of stopping a process that would have succeeded is a joint probability over the path; factoring it as a product of marginals overstates it by factors of two to five in the regime that matters, and displaces the peak of the regret curve by about a fifth of the design effect.
Two claims are numerical, and we have tried to state them at the strength the arithmetic supports and no higher. Continuous monitoring exceeds the best single look at matched budget by about 28 % of averted exposure, stably in the lag; it exceeds the deployed conditional-power convention at the conventional midpoint by about 50 % , increasingly in the lag. Roughly a third of the second figure is available simply by moving the conventional interim earlier.
One claim we withdraw. An earlier form of this argument compared oracle ceilings and concluded that a region of recoverable time was structurally closed to any single look. That comparison is circular, since the look time is a design variable; setting it equal to the earliest admissible monitoring time closes the gap exactly. Continuity does not enter. What remains is the interior trade of (3), which a single look genuinely cannot escape but which bounds its loss at a quantity we can only compute, not bound in closed form.

9.2. Beyond Clinical Trials

The kernel is a statement about the value of recovering time net of a wind-down lag, and applies wherever a committed process can be abandoned: a policy programme whose review cycle plays the role of t c h e c k ; a campaign whose contracted media buy sets t e n d and whose cancellation notice sets t e x e c u t e ; a military commitment whose withdrawal timetable is the lag.
The case of a failing marriage is worth stating explicitly rather than as a flourish, because it isolates the parameters unusually clearly. The base rate is high; the accruing statistic is noisy and its early values untrustworthy, so a reliability floor is genuine rather than conventional; the lag between deciding and separating is substantial and largely administrative; and the asymmetry of Section 4 is stark, in that the cost of a wrong stop is not discounted by how late it occurs. The framework says that a single scheduled review—the annual reckoning, the counselling appointment at a fixed date—must solve the same interior trade as any other single look, and that its loss relative to continuous attention does not shrink because the separation process is slow.

9.3. Limitations

The lag is a known constant; in practice it is random and plausibly correlated with trigger time. The confirmatory review is modelled as a single gate with no error of its own. The closed forms assume a Brownian statistic with linear information accrual and a linear boundary. Efficacy stopping, multi-arm designs and sample-size re-estimation are out of scope. The comparison of averted exposure treats calendar time as the objective; a full decision analysis would weight patient-years against the option value of a treatment not abandoned, which we have not attempted.

9.4. Extensions

Event-driven processes, where information accrues with events rather than calendar time, would extend the framework to time-to-event endpoints, replacing the linear clock with an event clock and rebuilding the kernel; this is the most valuable next step and would connect to the delayed-response literature [7]. A random lag would smooth the kink at s * and, by (2), plausibly increase rather than decrease the value of dispersed trigger mass. Permitting more than one confirmatory look would interpolate between the architecture studied here and full group-sequential monitoring.

10. Conclusions

The operational lag between deciding to stop a process and stopping it induces a horizon s * = t e n d d beyond which monitoring recovers nothing, and the time a stopping rule recovers is a linear functional of its first-passage law against the kernel g ( s ) = ( s * s ) + . This yields an exact factorisation of savings into detection and timing, and it prices the restriction to a single interim look: such a look must solve an interior trade between recoverable time and reliability at a fixed false-stop budget, which continuous monitoring is not obliged to solve. The reliability-optimal futility boundary turns out to be Wald’s one-sided sequential probability ratio boundary, recovered here from a time-recovery objective. In a non-oncology phase III example, continuous monitoring averts 55 % of a doomed trial’s duration against 43 % for the best single look at matched budget and 37 % for the deployed conditional-power convention, while stopping fewer trials that would in fact have succeeded—a difference visible only when that regret is measured as a joint probability over the path rather than as a product of marginals.
The most useful extension is to trials whose information accrues with events rather than with the calendar: time-to-event endpoints, and oncology trials in particular, require the kernel to be rebuilt on an event clock and would connect this analysis to the delayed-response literature [7]. Single-arm and multi-arm designs are also natural targets, the latter raising the question of how recoverable time should be allocated when one arm of several is abandoned. A random rather than fixed lag, and more than one confirmatory look, are the two obvious refinements within the present setting.

Appendix A. Comparison Identities

Proposition 1. From (1), Δ c o n t ( θ ) = p θ E [ g ( S ) θ ] = ( χ θ p θ L ) ( ϱ θ g ( τ L ) ) = χ θ ϱ θ Δ L ( θ ) .
Equation (2). For every real x, ( x ) + = x + ( x ) + , so pointwise ( s * S ) + = ( s * S ) + ( S s * ) + . Taking expectations, E [ g ( S ) ] = ( s * E [ S ] ) + E [ ( S s * ) + ] , and when E [ S ] s * the first term equals g ( E [ S ] ) . The remainder is non-negative, and strictly positive iff the trigger law places mass beyond s * .

Appendix B. First Passage for a Linear Boundary

Durbin’s rate. With ( τ ) = a + m τ , Y = B is Brownian with drift μ θ = ϑ θ m from | a | , and first passage to zero has density (4). Since ( | a | + μ θ τ ) 2 = ( ( τ ) ϑ θ τ ) 2 and f B ( τ ) ( ( τ ) ) = ( 2 π τ ) 1 / 2 exp ( ( ( τ ) ϑ θ τ ) 2 / 2 τ ) , the ratio is | a | / τ , independent of ϑ θ .
Windowed masses. The inverse-Gaussian distribution function with x 0 = | a | , μ = μ θ is
F θ ( τ ) = Φ x 0 μ τ τ + e 2 μ x 0 Φ x 0 + μ τ τ ,
and p θ = F θ ( τ m a x ) F θ ( τ m i n ) . The benefit functional is T [ τ * M 0 M 1 ] with M 0 , M 1 the zeroth and first partial moments of κ 1 on [ τ m i n , τ * ] . All case-study values are computed from these windowed expressions, not from the infinite-horizon approximation e 2 μ | a | , which differs by about 9 % at the relevant parameters.

Appendix C. Type-I Error

Theorem 1. Enlarging the monitoring set adds survival constraints, so R S R S R and probabilities order accordingly; P 0 ( R ) = α nom . Strictness holds whenever the set of paths that cross on S yet finish above b eff has positive Wiener measure.
Non-binding accounting. Computing α as though the boundary were absent compares R S R at the same efficacy threshold, so a futility stop can only remove a rejection; realised type-I error is at most nominal irrespective of override. This holds at unchanged b eff only.
Magnitude. By Proposition 3 at ϑ = 0 , the deflation for the continuous design of Section 8 is e 2 m | a | Φ ¯ ( b eff + 2 | a | ) = 0.00369 ; the corresponding bivariate-normal calculation for the conditional-power look at τ L = 0.5 gives 0.00446 .

Appendix D. Purity and the Likelihood Ratio

By Girsanov the path likelihood ratio for drift ϑ 1 against ϑ 0 is exp ( Δ ϑ [ B ( τ ) ϑ ¯ τ ] ) . Evaluated on B = | a | + m τ this gives (5). Equivalently, the Durbin prefactor is drift-free and cancels in κ 1 / κ 0 , leaving the difference of Gaussian exponents, [ ( | a | + μ 0 τ ) 2 ( | a | + μ 1 τ ) 2 ] / 2 τ , which reduces to the same expression using μ 0 μ 1 = Δ ϑ and μ 0 + μ 1 = 2 ( ϑ ¯ m ) . Admissibility ρ ρ * is Λ 10 ρ * 1 ρ * 1 π π ; at m = ϑ ¯ this is free of τ .

Appendix E. The Design Problem

With p 0 e 2 ( ϑ 0 m ) | a | and E [ T 1 ] = | a | / ( m ϑ 1 ) , benefit falls linearly and cost exponentially in | a | , so the budget binds at | a | = log ( 1 / ε ) / [ 2 ( ϑ 0 m ) ] and detection time is log ( 1 / ε ) / [ 2 ( ϑ 0 m ) ( m ϑ 1 ) ] . The denominator is concave in m with derivative ϑ 0 + ϑ 1 2 m , vanishing at m * = ϑ ¯ with value Δ ϑ 2 / 4 , giving Proposition 2. Numerically optimising the exact objective κ 1 γ over ( | a | , m ) subject to the windowed budget gives ( 0.649 , 1.452 ) against the proxy’s ( 0.705 , 1.620 ) , with benefit higher by 0.45 % .

References

  1. Anderson, T.W. (1960). A modification of the sequential probability ratio test to reduce sample size. Annals of Mathematical Statistics, 31(1), 165–197. [CrossRef]
  2. Bachelier, L. (1900). Théorie de la spéculation. Annales Scientifiques de l’École Normale Supérieure, 17, 21–86. [CrossRef]
  3. Chernoff, H. (1961). Sequential tests for the mean of a normal distribution. Proc. Fourth Berkeley Symposium, 1, 79–91.
  4. DiMasi, J.A., Grabowski, H.G. and Hansen, R.W. (2016). Innovation in the pharmaceutical industry: new estimates of R&D costs. Journal of Health Economics, 47, 20–33. [CrossRef]
  5. Durbin, J. (1985). The first-passage density of a continuous Gaussian process to a general boundary. Journal of Applied Probability, 22(1), 99–122. [CrossRef]
  6. Fortet, R. (1943). Les fonctions aléatoires du type de Markoff associées à certaines équations linéaires aux dérivées partielles du type parabolique. Journal de Mathématiques Pures et Appliquées, 22, 177–243.
  7. Hampson, L.V. and Jennison, C. (2013). Group sequential tests for delayed responses. Journal of the Royal Statistical Society B, 75(1), 3–54.
  8. Jennison, C. and Turnbull, B.W. (2000). Group Sequential Methods with Applications to Clinical Trials. Chapman and Hall/CRC.
  9. Lan, K.K.G. and DeMets, D.L. (1983). Discrete sequential boundaries for clinical trials. Biometrika, 70(3), 659–663. [CrossRef]
  10. Lan, K.K.G., Simon, R. and Halperin, M. (1982). Stochastically curtailed tests in long-term clinical trials. Sequential Analysis, 1(3), 207–219. [CrossRef]
  11. Lan, K.K.G. and Wittes, J. (1988). The B-value: a tool for monitoring data. Biometrics, 44(2), 579–585. [CrossRef]
  12. Lévy, P. (1948). Processus Stochastiques et Mouvement Brownien. Gauthier-Villars, Paris.
  13. O’Brien, P.C. and Fleming, T.R. (1979). A multiple testing procedure for clinical trials. Biometrics, 35(3), 549–556. [CrossRef]
  14. Pocock, S.J. (1977). Group sequential methods in the design and analysis of clinical trials. Biometrika, 64(2), 191–199. [CrossRef]
  15. Proschan, M.A., Lan, K.K.G. and Wittes, J.T. (2006). Statistical Monitoring of Clinical Trials: A Unified Approach. Springer.
  16. Shiryaev, A.N. (1978). Optimal Stopping Rules. Springer-Verlag.
  17. Siegmund, D. (1985). Sequential Analysis: Tests and Confidence Intervals. Springer.
  18. Wald, A. (1947). Sequential Analysis. Wiley, New York.
  19. Wald, A. and Wolfowitz, J. (1948). Optimum character of the sequential probability ratio test. Annals of Mathematical Statistics, 19(3), 326–339. [CrossRef]
  20. Whitehead, J. (1997). The Design and Analysis of Sequential Clinical Trials, 2nd rev. ed. Wiley.
  21. Wong, C.H., Siah, K.W. and Lo, A.W. (2019). Estimation of clinical trial success rates and related parameters. Biostatistics, 20(2), 273–286. [CrossRef]
Figure 1. The monitored process on the information-fraction axis τ = ( t t 0 ) / T . A doomed process (dark path) accrues a noisy statistic drifting toward the futility boundary (dashed). Three schemes are compared: no interim (run to t end ); a single fixed look, of which the conditional-power convention CP < 0.20 at τ L is the deployed instance (sienna); and continuous monitoring over [ τ min , τ max ] (teal band). Any trigger at s incurs a fixed operational lag d = t check + t execute before the process halts, so the recovered time is g ( s ) = ( s * s ) + with s * = t end d (inset); triggers in the dead zone ( s * , t end ] recover nothing. The lag is drawn larger than the worked-example value ( δ = 0.056 ) for legibility.
Figure 1. The monitored process on the information-fraction axis τ = ( t t 0 ) / T . A doomed process (dark path) accrues a noisy statistic drifting toward the futility boundary (dashed). Three schemes are compared: no interim (run to t end ); a single fixed look, of which the conditional-power convention CP < 0.20 at τ L is the deployed instance (sienna); and continuous monitoring over [ τ min , τ max ] (teal band). Any trigger at s incurs a fixed operational lag d = t check + t execute before the process halts, so the recovered time is g ( s ) = ( s * s ) + with s * = t end d (inset); triggers in the dead zone ( s * , t end ] recover nothing. The lag is drawn larger than the worked-example value ( δ = 0.056 ) for legibility.
Preprints 226507 g001
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings