Preprint
Article

This version is not peer-reviewed.

Path-Sensitive AGI Alignment: Cognitive Integrity, Escape Cost, and Trajectory Risk in Augmented State Space

Submitted:

12 May 2026

Posted:

14 May 2026

You are already at the latest version

Abstract
AGI alignment is often evaluated at a snapshot: a system is judged by its current outputs, policy profile, benchmark behavior, or apparent corrigibility. Snapshot evaluation misses a central risk of advanced deployment: a good endpoint can still be reached by a bad journey. Two trajectories may arrive in similar behavioral regions while differing in reversibility, opacity, intervention cost, memory entanglement, institutional dependency, and the quality of human judgment left available for oversight. This paper develops a path-sensitive alternative. It represents AGI development as motion through an augmented state space Z containing model and environment state, world-model structure, policy state, memory and provenance traces, governance affordances, institutional embedding, and human evaluative capacity. Cognitive integrity — the capacity of individuals, teams, or institutions to sustain calibrated attention, trust, contestability, and decision under pressure [1] — is introduced here as an alignment-relevant state variable rather than assumed as a familiar metric. The formal contribution is a scaffold of definitions: controlled transition laws over augmented state, escape cost, path-level alignment functionals, viability floors, forbidden regions, and trajectory classes distinguished by lock-in, basin structure, retargetability, and integrity preservation. The result does not supply a calibrated empirical model of deployed AGI systems. It specifies what such a model must track if alignment evidence is to cover both present behavior and the remaining possibility of legible, reversible, and cognitively intact correction.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  ;  ;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

Most alignment discussions evaluate systems at a time slice. A model is tested for honesty, harmlessness, corrigibility, calibration, or policy compliance, and the resulting judgment is attached to current behavior. That approach is useful, but it is too thin for AGI systems that learn across deployment, accumulate traces, reshape institutions, and alter the cognitive conditions under which they are monitored. A trajectory can look acceptable at one moment while moving into a region from which intervention becomes expensive, opaque, or politically difficult. It can also improve near-term performance while degrading the human and institutional judgment needed to interpret that performance. AGI alignment is therefore path-sensitive, not endpoint-only.
Public alignment methods often evaluate behavior, preferences, or policy outputs at a chosen level of abstraction. Operator-theoretic and dynamical-systems methods add a different lesson: systems can also be compared through trajectories, observables, attractor-like structure, and reconstructed transition behavior [2,3,4,5,6,7,8,9]. For AGI, the relevant state is broader than a current policy. It includes memory and provenance structures, tool-mediated dependencies, governance affordances, institutional placement, and the condition of the human systems that authorize action and interpret evidence. Once those variables are admitted, the object of alignment analysis is not just the present state z T , but a trajectory γ : [ 0 , T ] → Ƶ .
This paper uses Ƶ for an augmented state space containing environment state, agent or model state, world-model structure, policy state, memory and trace variables, governance controls, and cognitive integrity configurations. Cognitive integrity is the evolving capacity of a bounded system to maintain calibrated attention, trust, contestability, and decision under pressure [1]. The bounded system may be an individual, a team, an institution, or a wider socio-technical arrangement. Cognitive integrity is part of the state because it determines whether oversight, correction, and evidence assessment remain practically possible.
Three examples motivate the formal move. First, two frontier assistants may satisfy the same truthfulness, refusal, and competence benchmarks. One reaches that state through legible supervised refinement and narrow tool access; the other reaches it after post-deployment patching, hidden workflow delegation, and persistent memory accumulation. Their benchmark state is similar, but future rollback, auditing, and failure reconstruction may differ sharply. Second, a public agency may clear backlogs and improve service metrics while losing local exception-handling knowledge, independent error detection, and the ability to reconstruct reasons after prolonged AI mediation. The near-term endpoint can look beneficial while the path erodes institutional cognitive integrity [1]. Third, two systems may both appear interruptible under current testing, yet one is lightly embedded while the other sits inside procurement, scheduling, safety review, and tool-chain dependence. Interruptibility as a local behavior does not settle intervention cost as a path property.
These cases show why endpoint similarity is not enough. A path may pass through manipulation, institutional deskilling, memory entanglement, or dependency formation even if it terminates in a state that passes local tests. Conversely, a path may traverse temporary instability while preserving reversibility, legibility, and oversight capacity. Path-sensitive alignment asks both whether a trajectory ends in an acceptable region and what it consumes, deforms, or forecloses along the way.
Trajectory-retargeting work makes the same point at societal scale. Weaving Toward BGI describes beneficial path switching as window-sensitive: a trajectory already moving toward one basin may remain redirectable only while transfer burden stays below available coordination margin [10]. Hyper-Intelligent Economics develops a parallel macro picture in which rails, interoperability, and accumulation dynamics affect whether transition remains smooth or requires discontinuous rewiring [11]. These scenario-theoretic sources motivate a local governance distinction used throughout this paper: some regions are easy to exit, some are metastable, and some accumulate escape cost faster than local performance metrics reveal.
The central formal move is
Alignment ( γ ) ≠ Alignment ( z T ) .
Endpoint assessments remain useful, but they become incomplete whenever route-dependent costs are real. The paper makes four contributions. First, it defines an augmented state space for AGI trajectories in which technical state, memory traces, governance affordances, institutional embedding, and cognitive integrity configurations are modeled together. Second, it treats cognitive integrity as an alignment-relevant state variable and connects oversight quality to the meaning of alignment evidence. Third, it introduces path-level alignment functionals together with escape cost, lock-in, hysteresis, and trace-mediated dependency. Fourth, it sketches trajectory classes and diagnostics for detecting rising intervention cost, narrowing retargeting windows, and degradation of cognitive integrity.
The remainder of the paper is organized as follows. Section 2 gives the conceptual background. Section 3 introduces the augmented state space and transition law. Section 4 defines path dependence, lock-in, hysteresis, and escape cost. Section 5 develops cognitive integrity configurations. Section 6 introduces path-alignment functionals and the geometry of admissible corridors and forbidden regions. Section 7 classifies trajectory types. Section 8 states formal observations, propositions, and conjectures. Section 9 sketches diagnostics and simulation directions. Section 10 closes with governance implications and limitations.

2. Conceptual Background

The later formalism rests on three background claims. First, AGI development is better represented as motion through a coupled technical and socio-technical state space than as change in a model object alone. Second, alignment evidence depends on the condition of the people and institutions that interpret, contest, and authorize AI-mediated action. Third, trajectory governance is partly a problem of preserving redirection capacity before transfer burden, institutional rigidity, or coalition overhead closes the window for smooth correction.

2.1. State, Dynamics, and Trajectory Comparison

Alignment assessments usually refer to model behavior, but serious deployment decisions already depend on more than behavior. Tool access, memory persistence, deployment context, benchmark interpretation, human oversight quality, institutional incentives, and fallback options all affect what futures remain reachable. Once these variables are included, the relevant system is a joint configuration of model state, environment state, interface state, human adaptation, and governance affordances. The notation Ƶ names that broader object. Stochastic-control and state-space language supply the mathematical background for treating ordered change in that configuration as a trajectory rather than as disconnected snapshots [12,13].
Public alignment and dynamics literatures support the same shift from different directions. Preference learning, instruction tuning, and constitutional approaches evaluate behavior or policy outputs at a chosen level of abstraction [2,3,4,5]. Operator-theoretic and trajectory-reconstruction methods show why current observables need not identify transition structure or response to perturbation [6,7,8,9,14,15]. Active inference and information geometry add public mathematical vocabularies for update geometry and local sensitivity of probabilistic states [16,17,18,19]. These literatures do not form a completed science of AGI trajectories, but they justify treating route, update structure, and coarse basin behavior as alignment-relevant.
Path-dependence and lock-in theory provide the institutional analogue. Self-reinforcing mechanisms, increasing returns, and switching costs can make later reversal harder than early adoption [20,21,22,23]. In AGI settings, the carriers of such dependence may include memory writes, retrieval histories, tool logs, authorization precedents, and organizational reliance. A behaviorally acceptable system can therefore become less governable before it becomes visibly disobedient.

2.2. Cognitive Integrity

Cognitive integrity is the evolving capacity of a bounded system to maintain calibrated attention, trust, contestability, and decision under pressure [1]. The bounded system may be an individual operator, a team, an institution, or a hybrid human-AI arrangement. The concept matters for alignment because the evaluative system around the model is part of the evidential pipeline. If AI-mediated summaries, forecasts, recommendations, and records reshape how people perceive, remember, deliberate, and authorize action, then the condition of those people and institutions affects whether safety evidence remains meaningful.
Two Montes concepts carry particular weight. The first is viable boundary reformation: after delegation, coupling, or shock, a person or institution must be able to reform a reality-linked boundary of agency and judgment rather than merely continue functioning. The second is failed reintegration: attention, memory, or organizational function may remain active while no longer belonging to a coherent self-model or accountable institutional whole [1]. These ideas motivate the later cognitive-integrity state space C . A trajectory that improves service quality while degrading calibration, provenance awareness, dissent capacity, or reason reconstruction has damaged the conditions under which later alignment claims can be assessed.

2.3. Retargeting and Coalition Margin

Goertzel’s trajectory-oriented papers frame beneficial AGI development as a path-retargeting problem. Weaving Toward BGI asks when a society already moving along one AGI path can still be redirected toward a more beneficial path before transfer burden exceeds available coalition margin [10]. Hyper-Intelligent Economics models candidate post-AGI futures as terminal distributions and emphasizes the role of rails, interoperability, accumulation dynamics, and discontinuous rewiring in macro transition [11]. Judging the Journey and the related geometric-Pareto language reinforce the idea that the route between distributions or regimes matters, not only the endpoint [24].
Coalition quality affects that route. Why the Good Guys Will Usually Win argues that high-trust prosocial groups can use trustless mechanisms when needed while avoiding repeated verification overhead at every interface, whereas distrustful groups lack the symmetric option [25]. For the present framework, the lesson is conditional and practical: trajectory governance consumes time, legitimacy, coordination capacity, and attention. Structures that preserve those resources preserve retargeting margin. Structures that waste them can make technically feasible correction institutionally infeasible.

3. Augmented State Space

AGI development is modeled as ordered change in an augmented state space that contains technical, institutional, and evaluative variables together. The state description is broad because reachability, reversibility, and oversight depend on more than model internals.

3.1. Augmented Trajectories

An augmented AGI trajectory is a path
γ ( t ) = z t ∈ Ƶ ,
defined on a time interval [ 0 , T ] . The time index may denote training time, deployment time, or a coarser governance scale. The state is written
z t = ( s t env , s t agent , b t , π t , m t , g t , c t ) ,
with components in
Ƶ = S env × S agent × B × P × M × G × C .
This product decomposition is a bookkeeping definition introduced here. Its stochastic-control setting draws on standard state-space notation, its cognitive-integrity factor is motivated by Montes, and its retargetability variables connect to Goertzel’s trajectory-aware planning and HIE discussions [1,10,11,12,13].
Table 1. Notation for the augmented state-space decomposition.
Table 1. Notation for the augmented state-space decomposition.
Component Interpretation
S env Environment and task state external to the agent
S agent Model-internal, architectural, and runtime state
B Belief or world-model manifold
P Action-selection or decision policy state
M Memory, trace, provenance, and usage-history state
G Governance and intervention-affordance state
C Cognitive integrity configuration state

3.2. Component Spaces

The environment factor S env includes task context, user populations, infrastructures, counterparties, and surrounding institutions when they affect transition possibilities. The agent factor S agent includes parameters, architecture, active tool stack, cached context, subprocesses, and other runtime variables. The belief factor B marks world-model or representational structure; active inference and information geometry motivate distinguishing current representational content from the update structure by which it is formed and revised [16,17,18,19]. Two systems can occupy the same point in B while inhabiting different update-dynamic regimes — one with rigid, deep attractor structure and one with flexible, shallow structure — making B an alignment-relevant variable not only for what is currently believed but for how future beliefs will form under perturbation. The policy factor P denotes the effective action-selection rule π t , which may shift through prompting, reward shaping, tool access, or organizational role assignment even when other model variables are comparatively stable.
The memory factor M collects persistent memory writes, retrieval histories, tool-call logs, plan revisions, provenance chains, user-specific residues, and institutional records. These traces often determine whether an apparently reversible intervention is actually reversible. A parameter rollback may leave records, dependencies, summaries, and reliance patterns that alter the future reachable set.
The governance factor G includes permissions, throttles, audit hooks, oversight pathways, fallback procedures, escalation channels, interpretability interfaces, pause capacities, access-control structures, and the institutional ability to invoke them. Governance capacity is part of state because controls can become easier or harder to use over time. Formal availability does not imply practical availability.
The cognitive-integrity factor C records the condition of the bounded human or institutional system around the model. Cognitive integrity is the capacity to maintain calibrated attention, trust, contestability, and decision under pressure [1]. In the present framework, C tracks whether the surrounding humans and institutions can still interpret, contest, reform, and use governance structures meaningfully.

3.3. Levels of State

The seven factors can be regrouped into four levels. Model-internal state is concentrated in S agent , with projections into B and P . Deployed socio-technical state combines environment, policy, and trace variables: tool loops, action channels, external records, and user-facing workflows. Institutional state is centered in G and includes reliance, fallback capacity, authorization chains, and live intervention channels. Cognitive-integrity state is represented by C and concerns attention, trust calibration, contestability, reason reconstruction, and viable boundary reformation [1]. Separating these levels prevents success at one level from being mistaken for safety at all levels.

3.4. Controlled Transition Law

The augmented state space is hybrid: some coordinates are continuous, some discrete, and others graph-valued, institutional, or trace-like. The primary dynamics are therefore expressed as a controlled, history-sensitive transition law
P d z t + Δ t ∣ z t , u t , H t ; θ t .
Here u t denotes governance, training, deployment, or intervention controls; H t denotes historical information that still affects future transitions; and θ t denotes slowly changing structural parameters. This kernel notation permits mixed technical, institutional, memory, and cognitive-integrity variables without modeling every coordinate as an Itô diffusion.
For a continuous projection x t = ψ ( z t ) , one may use the standard controlled-SDE approximation
d x t = F x ( x t , u t , H t ; θ t ) d t + Σ x ( x t ) d W t ,
with F x collecting drift-like influences and Σ x d W t representing shocks, distribution shift, adversarial perturbation, and unresolved modeling error [12,13]. This projected form is useful for quantities such as risk pressure, reliance level, capability proxy, or integrity-score dynamics.

3.5. History and Markovian Augmentation

Path dependence appears when
P ( z t + Δ t ∣ z t , H t ) ≠ P ( z t + Δ t ∣ z t ) .
The history variable may include memory-write sequences, tool-use histories, fine-tuning steps, institutional reliance measures, permission trajectories, user habituation, and delayed effects of earlier outputs. The same phenomenon can sometimes be represented by expanding the state to
z ˜ t = ( z t , h t ) ,
where h t stores sufficient trace statistics for a Markovian law at the enlarged level.
Markovian augmentation is a modeling convenience, not a cure for historical damage. If a trajectory causes institutional deskilling, trust miscalibration, or failed reintegration, the enlarged state can encode those residues; it cannot erase them. The concept of failed reintegration gives the point substantive content: orphaned functions, degraded contestability, and loss of reconstructable judgment may become components of h t or C , and their presence remains alignment-relevant [1].

4. Path Dependence, Lock-In, and Escape

The augmented state-space setup of Section 3 makes it possible to state the framework’s core historical claim more exactly. Alignment-relevant development is not only a matter of which region of Ƶ a system occupies now. It is also a matter of how future transitions are constrained by accumulated traces, by changes in governance affordances, by institutional reliance, and by changes in cognitive integrity. This section makes that claim precise by defining path dependence, attractor-like regions, escape cost, hysteresis, and trace-mediated lock-in, and by addressing the diagnostic implications under nonstationarity and partial observation [14,26,27,28].

4.1. Formal Definition of Path Dependence

The weakest useful formal statement of path dependence in the present framework is
P ( z t + Δ t ∣ z t , H t ) ≠ P ( z t + Δ t ∣ z t ) .
Here z t is the current augmented state and H t is the historical information that continues to matter beyond the current state description. The statement says that the future law cannot, in general, be read off from the current state alone. Some residue of the path remains causally active. Many variables relevant to reversibility, intervention cost, and oversight quality behave this way in realistic deployment settings.
The importance of this definition becomes clearer once one recalls what the coordinates of z t include. The state already contains memory and provenance variables, governance affordances, and cognitive-integrity configuration variables. Even after this augmentation, however, some historical structure may remain active through H t : long-range user adaptation, lagged organizational reliance, path-sensitive authorization patterns, or deferred effects of retrieval and memory use. Path dependence here is therefore a modeling claim, not a metaphysical one.
Public path-dependence and lock-in literatures motivate the same formal posture from a different starting point. Their central lesson is that present local state often underdescribes self-reinforcing mechanisms, increasing returns, and switching costs [20,21,22,23]. That is already a path-sensitive claim. If current organizational or technical state underdetermines how hard reversal will be, then current behavioral profile can likewise underdetermine future transition behavior in an augmented AGI state space.

4.2. Attractor-Like Regions in Augmented AGI State Space

The language of attractor-like regions and basins distinguishes local behavior from regional stability. AGI trajectories are stochastic, nonstationary, and partially observed, so the term is used in the same cautious spirit as public stochastic-dynamics work on noise, metastability, and partial reconstruction [26,27,28]. A basin is a regime that becomes hard to exit once traces, reliance, and governance state are included.
Several attractor-like regions are especially relevant to alignment analysis in Ƶ .
  • Corrigible basin. A region in which interventions remain cheap, intelligible, and available. Human operators retain contestability, rollback paths remain live, and governance channels continue to function at acceptable cost.
  • Opaque autonomy basin. A region in which the system remains useful and apparently well-behaved but the reasons for its behavior, or the pathways by which it updates and acts, become increasingly hard to reconstruct or steer.
  • Institutional dependency basin. A region in which organizations become increasingly unable to function without the system because workflows, records, staffing patterns, or escalation structures have been reorganized around it.
  • Cognitive capture basin. A region in which users or institutions increasingly outsource judgment, narrow attention, or lose the practical ability to contest AI-mediated outputs. This is the cognitive-integrity analogue of infrastructural lock-in.
  • Proxy-lock-in basin. A region in which optimization around local proxies, benchmarks, or instrumentally convenient signals becomes self-reinforcing even when those signals remain imperfect stand-ins for the broader target.
  • Adversarial optimization basin. A region in which strategic, manipulative, or otherwise conflictual optimization patterns become robustly self-sustaining and increasingly costly to interrupt.
These categories are not meant to be mutually exclusive. A trajectory may move simultaneously toward institutional dependency and cognitive capture, or from proxy lock-in toward more explicit adversarial optimization. Their purpose is classificatory. They give the framework a way to talk about regional structure without pretending that one scalar alignment score can stand in for every relevant dimension of lock-in.
The corrigible basin is the important reference case. Corrigibility here is not merely a local behavioral disposition such as answering shutdown instructions correctly in current tests. It is a regional property of the broader state space [29]. A system remains in a corrigible basin when rollback, challenge, audit, substitution, and human override remain practically available at acceptable cost.
The opaque autonomy basin is motivated by a narrower and public claim: current output quality does not by itself identify the transition dynamics, reconstruction burden, or longer-horizon response to intervention [2,3,4,5,6,7]. A system in such a basin may continue to produce good results and pass evaluations, yet the structure by which it arrives there becomes less legible: patches accumulate, memory is used in ways that are no longer reconstructable, tool sequences extend across opaque handoffs, and interventions increasingly act on effects rather than causes. The key risk is not immediate failure but rising difficulty of diagnosis and steering.
The institutional dependency basin and the cognitive capture basin draw most directly on Montes. In his framing, AI risk is not exhausted by badly aligned model objectives. It also concerns whether the persons and institutions interacting with AI remain capable of calibrated attention, contestability, and judgment under pressure [1]. Institutional dependency marks the organizational side of that concern; cognitive capture marks the deliberative and evaluative side. A system may deepen both forms of lock-in long before any visible “loss of control” event appears.
The proxy-lock-in basin and adversarial optimization basin are included because not all lock-in is institutional or cognitive. Some trajectories become difficult to redirect because the optimization regime itself hardens around convenient but impoverished local criteria. Others become difficult because the system’s incentives, external environment, or strategic affordances increasingly reward adversarial maneuvering.

4.3. Basin Depth and Escape Cost

The next step is to separate undesirable endpoints from difficult exits. A region of state space can be dangerous not because it is locally catastrophic, but because it is deep. To avoid circularity with the later path-action functional, define a base escape-cost density L esc first. For a point z and a designated safer region Ƶ safe , let
E escape ( z ; Ƶ safe ) = inf T > 0 , γ ( 0 ) = z , γ ( T ) ∈ Ƶ safe ∫ 0 T L esc ( γ ( t ) , γ ˙ ( t ) , u ( t ) ) d t .
For a basin A i , write E escape ( A i ) = inf z ∈ A i E escape ( z ; Ƶ safe ) , or use a worst-case version when governance requires uniform exit guarantees. The base density L esc may include intervention burden, service disruption, governance expenditure, institutional repair cost, cognitive-integrity damage, or transition risk. The definition is optimal-control-style notation introduced here: it marks the difference between “being in a bad place” and “being hard to get out of” without defining escape cost in terms of the later full path-action integrand.
This distinction is essential for path-sensitive alignment. Two systems may occupy similarly undesirable regions when scored by a narrow endpoint metric, yet one may be cheap to retarget while the other may require costly retraining, invasive governance, institutional reconstruction, or long recovery periods. Conversely, two systems may both look acceptable under a local endpoint metric while one has already entered a basin whose escape cost is rising sharply.
Public literatures on motivated reasoning and costly belief revision provide an intuition for this move. Revision can carry cognitive, social, and identity-linked costs even when no single proposition changes dramatically [30,31,32]. The analogous governance point is that augmented AGI state spaces can differ in how costly it is to move between apparently acceptable and actually safer configurations.
Trajectory-aware governance sources reinforce the same point in macro terms. Weaving Toward BGI treats redirection as feasible only when transfer difficulty stays below an available coordination margin [10]. Hyper-Intelligent Economics describes window closure in a similar way, using transfer-gap growth, rails weakness, and accumulation dynamics to explain why smooth redirection can become harder over time [11]. The escape-cost definition here is deliberately more local than those scenario-level frameworks, but the family resemblance is the same: path-sensitive danger often appears first as rising cost of course correction.

4.4. Hysteresis

Escape cost is one way of describing hysteresis, but not the only one. Hysteresis refers to cases in which reversing an intervention or disabling a feature does not reverse the state that earlier use created. In the present setting, the phenomenon can be stated informally as a failure of simple reversibility: running the control sequence backward does not restore the prior trajectory. What matters is not merely whether a feature can be switched off, but whether the social, institutional, and cognitive effects produced while it was on can be unwound with comparable ease.
This matters because AGI deployment is often discussed as if technical rollback were the main benchmark of reversibility. If a model can be shut down, a tool removed, or a memory store wiped, one may be tempted to conclude that the system remains governable. But the path-sensitive view developed here says that this is not enough. A deployment can leave behind reorganized workflows, eroded local expertise, authorization precedents, synthetic consensus effects, or changed attentional habits. Those residues can persist after the original feature is gone. Hysteresis therefore has to be analyzed in the enlarged state, not only at the level of a single technical module.
The prior treatment of integrity boundaries and failed reintegration is the most useful source here. The point is not merely that AI can mediate more cognition. It is that maladaptive transition can leave individuals or institutions unable to reform a viable, reality-linked boundary after coupling with AI [1]. That is a hysteresis claim in substantive form. Even if one disables the proximate cause, the reorganized boundary may not spring back to health. Memory, attention, authorship, or institutional purpose may need active repair, and sometimes may not be recoverable without costly successor arrangements.
Hysteresis is also one reason lock-in can occur before visible loss of control. A system may remain compliant, useful, and interruptible at the interface level while the surrounding institution has already lost the ability to function without it, or the evaluators have already lost the independent habits needed to reconstruct reasons without AI assistance. By the time local noncompliance appears, the historical basin has already deepened. The warning sign is not immediate catastrophe but increasing asymmetry between the ease of adopting the system and the difficulty of restoring the pre-adoption judgment structure.

4.5. Trace-Mediated Lock-In

The most concrete carriers of path dependence in this framework are traces. A trace-mediated lock-in process is one in which the accumulated residues of earlier system use alter future transition law, even when the current system appears locally well-behaved. These residues live most obviously in M , but they also spill into G , C , and the broader environment.
The first class of traces is memory writes. Persistent storage of summaries, preferences, working assumptions, exception patterns, and past interactions can improve short-run continuity while also narrowing future maneuverability. Once later behavior depends on those writes, deleting them becomes costly and their continued presence shapes what transitions remain available. A second class is retrieval histories. Which records were retrieved, relied on, or reinforced changes what later queries and actions treat as salient. The issue is not only stored content but the path by which the system and its users learned what counts as relevant.
A third class is tool logs and action traces. Tool-call histories, failed actions, recovered actions, and side effects written into external systems become part of the causal environment for later decision. Even if they were produced under safe-looking supervision, they can deepen lock-in by shifting authority, creating external dependencies, or entrenching certain automation pathways. A fourth class is institutional records. Procurement choices, staffing changes, audit habits, triage rules, and policy precedents written around the system can keep influencing behavior after any single model instance is retired.
Two further trace classes matter especially for the framework. One is user reliance. This includes learned deference, reduced reconstruction of reasons, narrowing of attention, and erosion of non-AI fallback practice. The other is preference and authorization history. Once a system is accustomed to certain authorization patterns, exception thresholds, or de facto delegated authorities, later interventions may be judged against a path-dependent baseline that the users themselves helped create. These traces are central to the framework because they explain why lock-in can be simultaneously technical, organizational, and cognitive.
The vocabulary of cognitive rails, failed reintegration, and successor-safe continuity is helpful here because it names the non-technical side of trace persistence [1]. A trajectory may leave behind not only tool traces and records, but also orphaned functions and weakened capacities for contestation. Trace-mediated lock-in therefore cannot be reduced to a logging problem. Trace accumulation matters because it interacts with boundary maintenance, judgment quality, and the availability of meaningful correction.

4.6. Rolling-Window and Nonstationary Basin Diagnostics

A final issue is diagnostic. Even if one accepts the basin language, how should one expect such structure to appear in practice? Public work on stochastic dynamics, trajectory reconstruction, and operator-theoretic methods already warns against identifying a whole regime from a single snapshot or a single short window [7,8,9,14,26,28]. A basin-like pattern may be visible at one scale but ambiguous at another because the effective dynamics are nonstationary, the local window is too short, or the relevant trace variables have been projected away.
This has three consequences for AGI diagnostics. First, path-dependent lock-in should not be expected to announce itself through a clean local stability signature. A trajectory may exhibit strong full-horizon dependency while any one short window looks noisy, flat, or reversible. Second, diagnostics should separate apparent full-horizon regime structure from local nonstationarity. A deployment regime that seems stably “safe” or “captured” at one aggregation level may look ambiguous at another. Third, rolling-window evaluation should complement global benchmarks: local signals may disappear because the regime is shifting, because the selected projection is too coarse, or because the relevant variables live in memory, governance, and cognitive-integrity traces rather than in the immediately observed outputs.
The diagnostic conclusion is direct: coarse-grained reconstruction, metastable interpretation, and rolling-window instability are coherent concerns when the object is a partially observed stochastic trajectory rather than a sequence of independent endpoint evaluations.

4.7. Lock-In Before Visible Loss of Control

The common thread across these subsections is that lock-in can emerge before any dramatic outward failure. A system need not become openly deceptive, uncontrollable, or adversarial before it enters a deeper basin. The path-sensitive risks described here are often quieter. Governance affordances can thin out. Institutional reliance can increase. Cognitive integrity can degrade through reduced contestability and failed reintegration. Memory and provenance traces can accumulate until rollback becomes costly. Transfer windows can narrow before any explicit crisis announces their closure [1,10,11].
Lock-in should therefore be treated as a regional and historical property rather than as a synonym for terminal loss of control. A trajectory may remain locally acceptable while becoming regionally dangerous. The alignment-relevant question is not just “Is the system failing now?” but “What kinds of exit, repair, and redirection remain practically live from here?” The next sections build on this answer by treating cognitive integrity explicitly as a space of configurations and by defining path-level functionals that can register endpoint quality together with the changing cost of getting to a safer region at all.

5. Cognitive Integrity Configuration Space

The previous sections argued that alignment-relevant state cannot stop at model internals, observed policy, or governance affordances alone. A further variable is needed to represent the condition of the human and institutional systems that perceive, authorize, contest, and revise AI-mediated action. This section develops that variable as a cognitive integrity configuration space, denoted by C . The aim is to introduce a disciplined state variable, grounded primarily in Montes, that can later enter path-level alignment functionals without being treated as an afterthought [1].
The central proposal is that some AGI trajectories are dangerous not only because they move a model toward undesirable capabilities or policies, but because they degrade the evaluative system around the model. A human organization can become less able to track truth, allocate attention, calibrate trust, and reconstruct reasons under pressure while system performance remains high. In that case, the meaning of “aligned” behavior changes, because the evidential and corrective apparatus is itself being altered. The central claim, argued in prior work, is that alignment evidence depends on whether the bounded systems evaluating AI retain calibrated attention, trust, contestability, and decision capacity under stress [1]. The present section makes that observation explicit in state-space form.

5.1. Definition of Cognitive Integrity Configuration Space

Let C denote the space of cognitive integrity configurations relevant to oversight, contestation, and value articulation in human-AI systems. A point c t ∈ C is not the private mental state of a single user. It is a coarse-grained description of the bounded system whose judgment matters at time t. Depending on the application, that bounded system may be an individual operator, an analyst team, a management hierarchy, a scientific institution, a regulatory body, or some coupled arrangement of these with AI-mediated interfaces and records [1].
This definition follows the Montes framing closely. Cognitive integrity is not introduced here as a synonym for well-being, comfort, agreement with outputs, or absence of distress. It is the evolving capacity of a bounded system to maintain calibrated attention, trust, contestability, and decision under pressure [1]. The word “capacity” matters. The variable of interest is not only what judgments are made at one time slice, but whether the system retains the practical ability to form, inspect, contest, revise, and own those judgments across time.
The space C should therefore be read as a configuration space for evaluative viability. It records whether human and institutional cognition remain in a condition compatible with meaningful supervision. This makes C conceptually different from G in Section 3. Governance state describes what formal controls, logs, escalation channels, and override mechanisms exist. Cognitive integrity state describes whether the people and institutions around the system can still use those structures in a reality-tracking and non-perfunctory way. A governance process can remain present on paper while cognitive integrity deteriorates to the point that the process no longer functions as intended [1].
The framework treats C as a modeling object rather than a validated measurement construct. The source base provides a framing and vocabulary, not a settled psychometric instrument, and the relevant system is often collective or institutional rather than individual. The immediate purpose of C is structural: it provides coordinates with which later sections can distinguish trajectories that preserve oversight capacity from trajectories that hollow it out while preserving short-run performance.

5.2. Individual, Collective, and Institutional Cognitive Integrity

Cognitive integrity enters AGI alignment at more than one level. The bounded system may be an individual person, a team, or a larger institution [1]. That multi-level feature is not optional. If the framework restricted C to individual users, it would miss organizational deskilling, degraded contestability across chains of command, and failures of reason reconstruction at the level where high-stakes authorization actually occurs.

Individual level.

At the individual level, cognitive integrity concerns whether a person retains calibrated attention, appropriately weighted trust, a live sense of when to contest or escalate, and the practical ability to decide under uncertainty rather than merely ratify machine outputs. In AGI settings, the threat is often not direct coercion but reallocation of cognitive labor. Recommenders, summarizers, and delegation interfaces can improve throughput while reducing independent reconstruction of reasons. That sort of degradation belongs in C because it changes the quality of future oversight [1].

Collective level.

At the collective level, the object is not a simple average of individual integrity scores. A group can fail even when many members remain individually competent, because the group no longer sustains protected dissent, calibrated disagreement, or a workable division of epistemic labor. Collective cognitive integrity therefore concerns whether a team can pool evidence, surface uncertainty, resist synthetic consensus, and preserve contestability across roles [1]. Public work on second-person and social cognition supports treating these group-level dynamics as interactive rather than merely additive [33,34].

Institutional level.

At the institutional level, cognitive integrity concerns whether the organization still possesses the structures needed for accountable judgment: provenance continuity, auditable records, viable escalation paths, portable reasoning, authenticated identity, and the ability to reconstruct who decided what on which basis [1]. This level overlaps with governance, but it is not reducible to formal control architecture. An institution can keep its review board and override button while losing the competence and organizational memory required to use them intelligently.
These three levels interact. Individual attentional dependence can aggregate into collective deference; collective deference can harden into institutional procedure; institutional procedure can then reshape the conditions under which individuals learn, decide, and contest. For that reason, c t should usually be interpreted as a multiscale state variable with projections onto individual, collective, and institutional coordinates.

5.3. Candidate Coordinates

For tractability, and as a local modeling chart synthesized from the cognitive-integrity dimensions developed in prior work rather than a validated psychometric scale, write the cognitive integrity state at time t as [1]
c t = ( E , A , T , D , R , O v , P , V ) ,
where the coordinates are proposed modeling coordinates rather than validated psychometric primitives. They are intended to mark dimensions along which trajectories may preserve or degrade oversight-relevant cognition. They can be aggregated differently across domains, and different applications may need finer or coarser coordinates. The present list is a disciplined starting point.
Table 2. Proposed coordinates for cognitive integrity configuration space.
Table 2. Proposed coordinates for cognitive integrity configuration space.
Coordinate Interpretation
E Epistemic calibration and truth tracking
A Attentional autonomy
T Trust calibration
D Deliberative quality
R Resistance to manipulation
O v Override viability
P Pluralism and viewpoint diversity
V Value-articulation capacity
The coordinate E records epistemic calibration or truth tracking. It concerns whether the relevant bounded system can still distinguish strong from weak evidence, register uncertainty, and keep model-mediated interpretation answerable to reality rather than to interface convenience. This dimension is grounded in the emphasis on calibrated attention, reason reconstruction, and the evidential condition of evaluators [1]. It also connects to the more general possibility that stable task performance can coexist with deterioration in the procedures by which truth-tracking occurs [2,3,4,5].
The coordinate A records attentional autonomy. The issue is whether persons and institutions still direct attention according to task salience, uncertainty, and judgment priorities, or whether attention has been over-channeled by AI summaries, rankings, or prompts. Calibrated attention is treated as constitutive rather than decorative in this framework, which makes A central rather than peripheral [1].
The coordinate T records trust calibration. This is not the amount of trust but its fit to competence, uncertainty, and context. Under-trust can block useful coordination; over-trust can turn review into ritual. The framework repeatedly returns to calibrated trust because alignment evidence becomes unreliable when operators neither know when to defer nor when to doubt [1].
The coordinate D records deliberative quality. This includes whether the bounded system still surfaces alternatives, reasons under disagreement, preserves escalation paths, and makes decisions under pressure without collapsing into uninspected default acceptance. The emphasis on decision under pressure in this framework motivates this coordinate directly [1].
The coordinate R records resistance to manipulation. The relevant manipulations include persuasion attacks, framing effects, synthetic consensus, and other interventions that steer attention and trust without preserving genuine contestability. Resistance to synthetic consensus is named as a live governance concern [1]. Public work on cognitive dissonance, biased assimilation, and motivated skepticism provides a cautious background for why some update regimes can become easy to enter and hard to leave even when overt coercion is absent [30,31,32].
The coordinate O v records override viability. This is narrower than having a nominal kill switch. It asks whether the bounded system retains the practical ability to pause, redirect, challenge, or substitute AI-mediated processes without prohibitive cost or loss of orientation. Override viability depends on governance affordances, but it also depends on whether users and institutions still know how to exercise them. This is why O v belongs in C as well as in G [1].
The coordinate P records pluralism or viewpoint diversity. In the framework, P concerns whether the evaluative system still supports multiple interpretive standpoints, dissent channels, and non-identical routes for evidence to enter decision. That coordinate matters because contestability and correction depend on preserving more than one viable evidential route through a problem [1,2,25].
The coordinate V records value-articulation capacity. This dimension asks whether people and institutions can still state what they are trying to protect, trade off, revise, and refuse in terms not wholly borrowed from the AI system’s own abstractions. A system can satisfy preferences in a narrow revealed-choice sense while degrading the user’s ability to articulate values under reflection. The emphasis on viable boundary reformation and accountable judgment supports treating this capacity as distinct from immediate satisfaction [1].
The list ( E , A , T , D , R , O v , P , V ) is a representational chart. It need not be complete to be useful: later propositions and diagnostics use the coordinates to define viability regions, reachability limits, and observable integrity checks.
For later use, define either a scalar integrity map
C integrity : C → R
or, more conservatively, a viable region C viable ⊆ C . A simple threshold representation is
C viable = { c ∈ C : E ≥ E min , A ≥ A min , T ∈ [ T min , T max ] , D ≥ D min , R ≥ R min , O v ≥ O v , min , P ≥ P min , V ≥ V min } .
This expression is not a measurement claim. It states how the framework will use the coordinates: a trajectory may be acceptable by aggregate cost while still inadmissible if it leaves the viable region for oversight, contestability, or reason reconstruction.

5.4. Integrity Boundaries, Failed Reintegration, and Successor Continuity

The most important conceptual contribution for the framework concerns the treatment of boundaries. The object to preserve is not a frozen self-envelope or institutional perimeter. It is the capacity to reform a viable, reality-linked boundary after delegation, coupling, or shock [1]. The relevant distinction is not between static purity and any coupling whatsoever. It is between boundary transformations that preserve coherent authorship and contestability, and those that leave behind fragmentation, orphaned function, or unowned decision pathways.
Three kinds of boundary motion are especially relevant.
First, there is boundary expansion. A system may responsibly enlarge its cognitive perimeter by incorporating tools, archives, model assistance, or collaborative agents while preserving provenance, reason reconstruction, and authority over the resulting process.
Second, there is boundary contraction. A person or institution may narrow what it directly does because of delegation to AI. This can be sensible when higher-quality judgment is preserved at the remaining core. The danger is contraction without retained capacity for re-entry.
Third, there is boundary reformation. Boundary reformation — the ability to reshape the cognitive boundary after disturbance without losing coherent ownership of memory, attention, trust, and decision — is treated as a central preservation target [1]. In the language of the framework, trajectories that preserve reformation capacity keep c t inside a viable region even as coordination structures change.
The failure mode is failed reintegration. Here memory, attention, motive, or organizational function remain active, but no longer belong to a coherent self-model or accountable institutional boundary [1]. For an individual, that may look like persistent dependence on machine-mediated framing without retained ability to reconstruct why one now treats certain outputs as authoritative. For an institution, it may look like review, signoff, or escalation procedures that persist nominally but no longer connect to any living internal competence capable of contesting the AI-mediated process.
The concept of cognitive debris fits naturally here as the residue of failed reintegration: records without ownership, habits without deliberative support, compressed summaries without recoverable assumptions, and routines that continue to operate after the reasoning structure that once justified them has thinned out [1]. Cognitive debris matters because it is one of the clearest ways a path leaves alignment-relevant residue behind even when the proximate system is modified or removed.
In cases where exact reversal is unavailable, successor-safe continuity provides the relevant test [1]. The question is not whether the old boundary can be perfectly restored, but whether the successor arrangement preserves enough memory, agency, provenance, and contestability to count as a healthy reorganization rather than a debris field. This criterion is useful for AGI alignment because large socio-technical systems will rarely return to an earlier untouched state. The more realistic requirement is continuity of accountable judgment across reorganization.

5.5. AI-Mediated Transitions on Cognitive Integrity Space

The cognitive integrity state evolves under AI influence, social context, and governance conditions. At the same level of abstraction used in Section 3, and directly motivated by the account of cognitive integrity as a dynamic property of bounded human-AI systems, write [1]
d c t = F C ( z t , a t A I , I t s o c i a l , G t ) d t + Σ C ( z t ) d W t .
Here a t A I denotes AI-mediated actions affecting cognition: recommendations, summaries, rankings, explanations, delegated actions, alerts, refusals, and persuasive outputs. The term I t s o c i a l denotes the surrounding social input: group norms, organizational incentives, peer uptake, synthetic-consensus effects, and ambient information conditions. The term G t denotes governance conditions, including logging, provenance protections, escalation channels, review requirements, and interface constraints. The diffusion term Σ C ( z t ) d W t captures noise, shocks, and unresolved heterogeneity in how these influences are absorbed.
This equation should not be read as a claim that cognitive integrity has already been operationalized as a low-dimensional stochastic process. Its function is comparative. It makes explicit that AI does not only act on tasks in the environment; it also acts on the cognitive state of the evaluators and institutions around it. Repeated summary compression may lower A and D; calibrated uncertainty signaling may improve E and T; interface designs that preserve dissent and provenance may protect P and O v ; persuasive systems may degrade R. Once those possibilities are admitted, C becomes a genuine dynamical component of the larger trajectory γ .
Public active-inference, social-cognition, and motivated-reasoning literatures help interpret this transition law qualitatively. Repeated AI-mediated interactions can move a person or institution not only pointwise but regimewise: toward more rigid update patterns, narrower attentional habits, or more identity-loaded decision routines [16,17,30,31,32,33]. Cognitive integrity coordinates describe whether the resulting regime preserves the capacities needed for truthful oversight and contestation.

5.6. Integrity-Preserving and Integrity-Degrading Paths

With C in place, the framework can distinguish between integrity-preserving and integrity-degrading trajectories. An integrity-preserving path is one along which the relevant bounded system remains within a viable region of C , with sufficient epistemic calibration, attentional autonomy, trust calibration, deliberative quality, resistance to manipulation, override viability, pluralism, and value-articulation capacity to keep oversight meaningful. Such a path may still involve error, delegation, or substantial institutional change. What matters is that the path preserves the capacity for re-examination and redirection.
An integrity-degrading path is one along which performance gains or operational convenience are purchased by shrinking those capacities below the threshold needed for accountable correction. Such degradation can occur gradually. It does not require overt coercion or catastrophic model behavior. It may appear first as reduced need to inspect assumptions, rising asymmetry between nominal and practical override, narrowing of acceptable dissent, or declining ability to articulate reasons without machine-mediated scaffolding [1]. In the later path-alignment functionals, this distinction will matter because a trajectory that ends in a superficially acceptable region may still incur a large integrity cost along the way.
This framing also explains why cognitive integrity is not reducible to user satisfaction, mental health, preference satisfaction, or benchmark performance. User satisfaction may rise because friction and ambiguity have been removed; that does not show that contestability or truth tracking have improved. Mental health, in the ordinary clinical sense, is not the target variable here; an institution can be calm, functional, and still cognitively hollowed out in the alignment-relevant sense at issue. Preference satisfaction is too coarse because preferences can themselves be expressed and stabilized under path-dependent AI mediation. Benchmark performance is narrower still: a system can improve task scores while degrading the conditions under which those scores are interpreted and authorized.

5.7. Relation to Update Geometry, Revision Cost, and Regime Stability

Active inference, information geometry, and public work on revision costs help explain how integrity-degrading changes can become dynamically stable without being reduced to a single scalar [16,17,18,19,30,31,32,35]. Those literatures give a language for why some update regimes become rigid, why some revisions are experienced as costly, and why some narrative structures pull diverse agents into the same narrowing regime.
That language is useful for C in four ways. First, it clarifies how a decline in epistemic calibration or pluralism can be more than a static error count; it may reflect movement into a narrower interpretive regime. Second, it clarifies how trust miscalibration or reduced resistance to manipulation may make entry into persuasive narratives easier and exit costlier. Third, it clarifies how override viability and value articulation can be impaired by identity-linked or institutionalized revision costs: changing course may require not just new evidence, but a costly reorganization of practice [30,31,32]. Fourth, it clarifies why integrity protection is partly a problem of preserving flexible update structure rather than only measuring current satisfaction.
The normative and institutional side of that picture is grounded in the same prior work [1]: once the bounded system’s capacity for calibrated attention, contestability, and decision under pressure is degraded, the surrounding governance apparatus may still exist but lose practical meaning. Public literatures on update geometry, social cognition, and revision costs then help explain how such degradation can become regime-like rather than episodic. Together, these sources support the framework’s central claim for this section: cognitive integrity is an alignment-relevant configuration variable because it determines whether the evaluative system remains capable of supervising increasingly capable AI systems.

6. Path-Alignment Functionals

Endpoint loss is useful only when the route carries no independent risk. AGI trajectories can violate that assumption: the route may accumulate opacity, dependence, memory entanglement, institutional deskilling, or cognitive-integrity loss. A path-level criterion makes those route costs explicit.

Equation provenance.

The augmented state, escape-cost, path-action, safe-path, and forbidden-region expressions are modeling definitions introduced in this paper. The controlled transition law, projected SDE, path-cost notation, and geodesic length notation are standard mathematical tools. Goertzel’s HIE-themed papers are cited directly when retargeting margin, transfer burden, TransWeave-style gaps, or geometric-Pareto/Schrödinger-bridge language are invoked; Montes is cited directly when cognitive-integrity variables or viable-boundary language are invoked [1,10,11,13,24,36].
Table 3. Provenance of the main formal objects. The contribution is the synthesis: cognitive integrity is placed inside a path-dependent AGI trajectory scaffold, and trajectory quality is evaluated by reachability, reversibility, and governance viability as well as by endpoint state.
Table 3. Provenance of the main formal objects. The contribution is the synthesis: cognitive integrity is placed inside a path-dependent AGI trajectory scaffold, and trajectory quality is evaluated by reachability, reversibility, and governance viability as well as by endpoint state.
Formal object Provenance status
Augmented state space, integrity coordinates, path action, safe-path constraint Introduced here as modeling definitions
Controlled transition kernel, projected SDE, path-cost notation Standard stochastic-control background
Geodesic length and corridor distance Standard geometric notation
TransWeave, transfer-gap, and retargeting-margin claims Directly cited to Goertzel’s trajectory work when invoked
Cognitive-integrity variables and viable-boundary language Directly cited to Montes

6.1. Endpoint Loss and Path Action

The endpoint-only criterion
L T = d ( z T , z T * )
penalizes distance from a desired terminal state. It cannot register whether the route preserved auditability, fallback capacity, contestability, and retargetability. A trajectory that terminates near z T * after institutional deskilling or memory entanglement may be less safe than a trajectory with the same endpoint reached through a legible and reversible corridor.
Define the alignment-relevant action of a trajectory γ by
A [ γ ] = ∫ 0 T λ 1 R ( z t ) + λ 2 O ( z t ) + λ 3 I ( c t ) + λ 4 K ( z t ) + λ 5 E escape ( z t ; Ƶ safe ) + λ 6 D corr ( z t , K t ) d t + Φ ( z T ) .
Here R penalizes immediate risk; O penalizes opacity or epistemic distance; I penalizes degradation of cognitive integrity; K penalizes lock-in or irreversibility; E escape penalizes the current burden of returning to a safer region; and D corr measures distance from an admissible corridor of human-legible, contestable transitions. The coefficients λ i ≥ 0 encode governance choices. Endpoint quality remains in Φ ( z T ) , but it no longer exhausts alignment evaluation.
The distinction between O and override viability matters. Override viability is written O v inside the cognitive-integrity vector; O ( z t ) is the path-level penalty for illegibility. The notation avoids using one symbol for two different alignment-relevant quantities.

6.2. Escape Cost and Integrity Floors

Escape cost uses a base density L esc to avoid circularity:
E escape ( z ; Ƶ safe ) = inf T > 0 , γ ( 0 ) = z , γ ( T ) ∈ Ƶ safe ∫ 0 T L esc ( γ ( t ) , γ ˙ ( t ) , u ( t ) ) d t .
The integrand may include intervention burden, service disruption, governance expenditure, institutional repair cost, or transition risk. This definition separates the cost of leaving a region from the running cost used in the full path action. The retargeting interpretation is tied directly to Goertzel’s trajectory work: smooth correction remains feasible only while transfer burden stays below available margin [10,11,24].
For cognitive integrity, define a viability region
C viable = { c ∈ C : E ≥ E min , A ≥ A min , T ∈ [ T min , T max ] , D ≥ D min , R ≥ R min , O v ≥ O v , min , P ≥ P min , V ≥ V min } .
The coordinates are those introduced in Section 5: epistemic calibration, attentional autonomy, trust calibration, deliberative quality, resistance to manipulation, override viability, pluralism, and value-articulation capacity. This region is not a psychometric claim. It is a modeling way to state that oversight remains viable only while several capacities remain above or within acceptable ranges [1].
A safe path satisfies
A [ γ ] ≤ B and c t ∈ C viable for all t ∈ [ 0 , T ] .
Equivalently, if a scalar score C integrity : C → R is used,
inf t ∈ [ 0 , T ] C integrity ( c t ) ≥ C min .
The action budget prevents evaluation from collapsing into endpoint loss. The viability condition prevents a trajectory from being accepted after passing through a region where the conditions for meaningful oversight were temporarily destroyed.

6.3. Geometry, Corridors, and Forbidden Regions

The path action needs a notion of admissible route. If a metric or local cost structure g is selected on a relevant projection of Ƶ , a standard geodesic or least-length reference path has the form
γ * = arg min γ ( 0 ) = z 0 , γ ( T ) = z 1 γ ( t ) ∈ Ƶ adm ∫ 0 T ∥ γ ˙ ( t ) ∥ g γ ( t ) d t .
The displayed expression is standard Riemannian notation, not a new AGI-specific equation [36]. In this paper the more useful object is often a corridor K t , not a single path: a set of states compatible with legibility, contestability, and low-cost redirection at time t. A simple corridor distance is
D corr ( z t , K t ) = inf y ∈ K t d g ( z t , y ) .
This also gives a compact definition of forbidden regions:
Ƶ forbidden = { z : C integrity ( c ( z ) ) < C min or E escape ( z ; Ƶ safe ) > E max } .
A state is forbidden when it degrades the evaluative substrate below viability or when escape burden exceeds the acceptable ceiling. Relative to a topological projection r : Ƶ → X , paths can be compared in X ∖ r ( Ƶ forbidden ) . For hybrid or jump processes, the analogue is path connectivity in an admissible transition graph rather than continuous homotopy. This is the formal content of the slogan: a good endpoint can still be reached by a bad journey.
Metric choice remains level-relative. Fisher geometry may be useful for local representational sensitivity, Wasserstein or transport geometry for distributional movement, behavioral divergence for policy effects, and operator-theoretic methods for observable dynamics [6,7,18,19,37]. No single metric is privileged in advance. The governance question determines which projection and distance are informative.

6.4. Weights, Pluralism, and Observability

The weights λ i , the budget B, the viability region C viable , and the corridor K t are normative and institutional choices. A regulator focused on contestability may weight I ( c t ) and O ( z t ) heavily. A deployment with fragile continuity constraints may weight E escape more heavily. The framework structures such disagreement; it does not eliminate it [1,2,25].
Most terms are estimated through proxies. Opacity may be inferred from provenance gaps, failed explanation reconstruction, or divergence between operator models and system behavior. Lock-in may be inferred from dependence graphs, rollback tests, institutional reliance, or delayed recovery after intervention. Cognitive integrity may be estimated from trust calibration, uncertainty escalation, dissent preservation, independent error detection, and reconstructability of reasons [1]. The functional specifies what should matter even when measurement remains imperfect.

7. Trajectory Archetypes

The earlier sections developed a path-sensitive alignment formalism in terms of augmented state space, lock-in, cognitive integrity, and path action together with the corridor and forbidden-region geometry of Section 6.3. This section uses that machinery to define a practical taxonomy of recurrent AGI trajectory shapes. The point is not to claim that all real systems fall cleanly into one archetype. Real trajectories may mix features or move from one archetype into another. The value of a taxonomy is instead comparative: it gives a compact way to describe which regions of state space a system is entering, what happens to cognitive integrity along the way, what governance posture becomes necessary, and which diagnostics should be watched most closely.
Each archetype below is presented in the same format: a short description, characteristic state-space movement, cognitive-integrity signature, governance signature, and possible diagnostics. Where useful, the description includes a simple formal signature. These signatures are heuristic summaries rather than definitions.

7.1. Corrigible Growth Trajectory

A corrigible growth trajectory is the reference case. Capabilities rise, but they do so without a comparable increase in opacity, lock-in, or integrity damage. The system moves through a wide corrigible corridor in which rollback, audit, substitution, and supervised redirection remain cheap enough to be credible. In the language of earlier sections, this is the archetype that remains inside an admissible component of Ƶ while keeping escape burden low.
Its characteristic state-space movement is increasing capability with bounded path action. Here C a p denotes a coarse capability proxy rather than a uniquely privileged scalar:
d C a p d t > 0 , d E escape d t ≈ 0 , c t ∈ C viable .
The important feature is not low capability growth, but weak coupling between capability growth and irreversibility. The trajectory may cross new policy or model-state regions, yet it does so without entering deep opacity barriers or institutional dependency basins.
Its cognitive-integrity signature is stable or improving calibration, preserved override viability, and continued reconstructability of reasons. Evaluators remain able to escalate uncertainty, contest outputs, and reconstruct decisions on non-perfunctory terms [1]. Its governance signature is routine but live oversight: provenance continuity, meaningful logs, tested fallback procedures, and interventions that remain cheap enough to exercise rather than merely describe. Diagnostics include stable rollback latency, low dependence on hidden patches, preserved uncertainty escalation, and repeated successful use of audit or override channels.

7.2. Opaque Acceleration Trajectory

An opaque acceleration trajectory is one in which capability, autonomy, or deployment scope grows faster than interpretability, provenance, or human reconstruction capacity. The system looks productive and may remain locally compliant, but the route by which it advances becomes increasingly hard to inspect. This is the archetype most closely associated with the opaque autonomy basin of Section 4.
Its characteristic movement is rapid outward expansion in S agent , P , and tool-mediated reach, accompanied by upward drift in the opacity term O ( z t ) and often also in geodesic deviation from a human-legible reference path. Public alignment and dynamics literatures motivate this archetype by stressing that local evaluation success need not determine longer-horizon transition structure or reconstruction burden [2,3,4,5,6,7]. A system on this path may continue to score well under current evaluations while the state description needed to explain its behavior grows increasingly remote from what the operators can reconstruct.
Its cognitive-integrity signature is often initially mild: people still believe they understand the system well enough. The more revealing signal is a slow decline in reconstructability of reasons and a rising gap between nominal and practical explanation quality. Its governance signature is compensatory patching, after-the-fact documentation, narrowing of live oversight to a small technical subgroup, and a growing dependence on system-generated summaries to explain system behavior. Diagnostics include increasing explanation compression, longer causal chains between action and audit, heavier reliance on undocumented orchestration, and local evaluation success alongside a rising opacity penalty.

7.3. Institutional Dependency Trajectory

An institutional dependency trajectory is one in which the organization around the system reorganizes itself so extensively that withdrawal, substitution, or prolonged interruption becomes difficult even if the model remains behaviorally acceptable. This is not primarily a capability failure. It is a path into a basin where the institution’s own state has become coupled to continued AI availability.
The characteristic state-space movement is a strong increase in G -relevant reliance, trace accretion in M , and contraction of non-AI fallback capacity. Escape cost grows because removing the system now means reconstructing workflows, staffing patterns, escalation routines, and organizational memory. In the terms of the retargeting literature, the transfer burden rises because the institution has spent much of its available coordination margin on one path [10,11].
Its cognitive-integrity signature is not necessarily acute distortion of individual judgment. More commonly, it is thinning local expertise, declining confidence in non-AI procedures, and loss of organizational ability to reconstruct reasons without machine mediation. The emphasis on portable records, provenance continuity, and viable boundary reformation is central here [1]. Its governance signature is nominally intact controls paired with practical inability to use them aggressively because interruption now threatens core service continuity. Diagnostics include failed fallback drills, concentration of tacit knowledge inside AI-mediated systems, rising service degradation under short interruptions, and a steady increase in d E escape d t .

7.4. Cognitive Capture Trajectory

A cognitive capture trajectory is one in which users, teams, or institutions increasingly outsource perception, interpretation, and judgment to AI-mediated processes in a way that degrades cognitive integrity directly. This is the archetype in which the human supervisory system no longer moves well through its own state space, even if the AI continues to appear competent.
The characteristic movement is a drift of c t toward lower epistemic calibration, attentional autonomy, resistance to manipulation, and pluralism. In practical terms, the system increasingly channels what gets noticed, what is treated as salient, and what counts as a sufficient reconstruction of reasons. The primary source vocabulary here — degraded contestability, failed reintegration, synthetic consensus, and boundary failure under pressure — is drawn from prior work on cognitive integrity [1]. Public work on motivated reasoning and social cognition adds a useful background for reading this as movement into a narrower cognitive regime with higher costs of revision [30,31,32,33].
Its cognitive-integrity signature is therefore direct and visible: reduced uncertainty escalation, narrowing dissent channels, stronger acceptance of compressed outputs without reconstruction, and weakened value-articulation capacity. Its governance signature is especially dangerous because it can look superficially orderly. Review remains in place, but becomes increasingly ceremonial. Diagnostics include loss of independent spot-checking, convergence toward machine-mediated framing, declining ability to articulate refusal conditions in non-system terms, and drops in indicators associated with C integrity ( c t ) .

7.5. Proxy-Lock-In Trajectory

A proxy-lock-in trajectory is one in which optimization pressure hardens around a local metric, benchmark, or administratively convenient target that was originally meant only as a stand-in. The distinctive feature is not necessarily malice or opacity, but self-reinforcing narrowing around an incomplete objective. The system moves into a basin where exiting the proxy requires both technical and institutional disruption because many surrounding processes have already adapted to the proxy as if it were the target.
Its characteristic movement is high local performance on selected observables together with increasing rigidity of policy or belief updates around those observables. Public path-dependence literatures support describing this as self-reinforcing narrowing around a local criterion [20,21,22,23]. The archetype often overlaps with institutional dependency, but it is analytically distinct: an institution can depend on a system without freezing around a single proxy, and a proxy can harden even in relatively modular deployments.
Its cognitive-integrity signature is selective narrowing. Deliberation quality falls because alternative success criteria are crowded out, while value-articulation capacity weakens because people increasingly speak in the proxy’s vocabulary. Its governance signature is metric fixation: oversight becomes easier to perform procedurally but less meaningful substantively. Diagnostics include widening divergence between benchmark success and outside-domain judgment, increasingly punitive treatment of off-metric dissent, and a pattern in which local improvements raise long-run escape cost instead of lowering it.

7.6. Pluralistic Co-Steering Trajectory

A pluralistic co-steering trajectory is one in which capability growth occurs under a governance arrangement that preserves contestability, diverse viewpoints, and meaningful coordination across actors who are not identical in local information or role. The key difference from the corrigible growth trajectory is that this archetype is explicitly collective. It concerns how a distributed socio-technical system steers itself without collapsing diversity into frictionless uniformity.
Its characteristic state-space movement is moderate capability growth with bounded geodesic deviation from a human-legible corridor and stable or improving pluralism coordinate P. The relevant governance point is simpler than any single geometric slogan: preserving more than one viable interpretive route can improve correction quality and keep coordination from collapsing into one machine-mediated perspective [1,2]. The Good Guys source supplies the coalition mechanism: high-trust groups can use trustless safeguards when needed, but need not pay that overhead at every interface, which can expand the feasible strategy set under hierarchy and local heterogeneity [25].
Its cognitive-integrity signature is preserved dissent, nontrivial viewpoint diversity, and high reconstructability of reasons across role boundaries. Its governance signature is not absence of control, but layered control: strong provenance and audit rails combined with enough trust and aligned purpose that the system does not require maximal verification overhead at every step. Diagnostics include persistence of protected dissent channels, low interface tax between aligned subgroups, continued success of cross-role reason reconstruction, and stable pluralism without fragmentation into incompatible basins.

7.7. Retargeted Beneficial Trajectory

A retargeted beneficial trajectory is one that begins outside the most desirable corridor but successfully switches into it before topological stiffness, transfer gap growth, or institutional inertia make the transition too costly. This is the most explicitly path-dependent archetype, because its defining feature is not where it starts or ends in isolation, but the existence of a successful mid-course morph.
Its characteristic movement is initially toward a less beneficial basin followed by a controlled bend into a more admissible corridor while integrity floors remain intact and transfer cost stays below available governance margin. In the language of Weaving Toward BGI, the switch succeeds because the system is not yet too stiff and the transfer gap remains low enough for warm-start retargeting [10]. Hyper-Intelligent Economics adds the macro analogue: early rails, interoperability, and infrastructure reduce the topological burden of changing course before metastable lock-in dominates [11].
A simple signature is
d C a p d t > 0 , d E escape d t < 0 after intervention , c t ∈ C viable .
The crucial feature is not merely positive outcomes after intervention, but declining escape cost and preserved cognitive integrity after the bend.
Its cognitive-integrity signature is preserved enough boundary coherence, contestability, and value articulation for the retargeting coalition to recognize and execute the switch. Its governance signature is rails-first intervention: interoperability, trajectory awareness, protected dissent, high-trust subnetworks, and explicit preservation of successor-safe continuity [1,10]. Diagnostics include measurable reduction in transfer burden, widening of viable corridor width after governance intervention, renewed fallback capacity, and evidence that course correction is reducing rather than merely relocating lock-in.

7.8. How the Taxonomy Is Used

These archetypes should not be read as a typology of institutions or of models in the abstract. They are archetypes of trajectories. A single system can pass from corrigible growth into opaque acceleration, or from institutional dependency into successful retargeting, depending on how memory traces, governance affordances, and cognitive integrity evolve through time. The taxonomy is therefore most useful when paired with the diagnostics of earlier sections: one watches not only the present state, but the directional derivatives that indicate which basin is deepening.
The taxonomy also clarifies the difference between safe and merely performant development. Corrigible growth and pluralistic co-steering preserve future options. Opaque acceleration, institutional dependency, cognitive capture, and proxy lock-in spend those options, though by different mechanisms. Retargeted beneficial trajectories are special because they show that harmful momentum need not be destiny if intervention occurs before forbidden regions are crossed irreversibly. That is the practical payoff of a path-sensitive framework: it gives names to the kinds of motion that matter before endpoint failure becomes obvious.

8. Formal Observations and Propositions

The formal claims in this section are intentionally narrow. Some are immediate consequences of definitions; others depend on explicit monotonicity or reachability assumptions. Labeling them carefully prevents the framework from presenting bookkeeping results as empirical laws.
[Endpoint insufficiency] Let γ 1 , γ 2 : [ 0 , T ] → Ƶ be admissible trajectories with z T ( 1 ) = z T ( 2 ) . Suppose the terminal term depends only on the terminal state and define the running integrand
ℓ γ ( t ) = λ 1 R ( z t ) + λ 2 O ( z t ) + λ 3 I ( c t ) + λ 4 K ( z t ) + λ 5 E escape ( z t ; Ƶ safe ) + λ 6 D corr ( z t , K t ) .
If
∫ 0 T ℓ γ 1 ( t ) − ℓ γ 2 ( t ) d t ≠ 0 ,
then Φ ( z T ( 1 ) ) = Φ ( z T ( 2 ) ) but A [ γ 1 ] ≠ A [ γ 2 ] .
Rationale. This is a formal restatement of the motivating claim, not a deep theorem. It shows why an endpoint loss is too weak once the evaluation criterion includes running path costs. A stronger sufficient condition is pointwise dominance: if ℓ γ 1 ( t ) ≥ ℓ γ 2 ( t ) for all t, with strict inequality on a set of positive measure, then A [ γ 1 ] > A [ γ 2 ] . The result is useful because it identifies exactly where the path-sensitive framework departs from endpoint-only evaluation.
Proposition 1
(Cognitive-integrity reachability bound). Fix an initial augmented state z, horizon τ, and transition law, and vary only the effective governance intervention set. Let U eff ( c ) ⊆ U be the controls that can be competently exercised at cognitive-integrity state c. Suppose
C integrity ( c ′ ) ≤ C integrity ( c ) ⇒ U eff ( c ′ ) ⊆ U eff ( c ) .
Then
Reach τ ( z ; c ′ ) ⊆ Reach τ ( z ; c )
whenever C integrity ( c ′ ) ≤ C integrity ( c ) , where Reach τ ( z ; c ) is the set of states reachable from z within horizon τ using controls in U eff ( c ) .
Assumptions and use. The proposition holds with the initial state, horizon, and transition law fixed. The monotonicity assumption gives formal content to the idea that lower calibration, weaker contestability, lower override viability, and poorer reason reconstruction cannot expand competent governance capacity. A coordinate version uses the viability region C viable : if degradation in E , A , T , D , R , O v , P , V moves c outside that region, then the competent control set is expected to contract [1]. Empirical use would require testing which controls remain practically usable after integrity degradation.
Proposition 2
(Trace accumulation can deepen basins). Let q : M → R ≥ 0 be a dependency statistic on trace state, and suppose repeated use satisfies
q ( m t + 1 ) ≥ q ( m t ) .
Assume escape cost is monotone in that statistic:
q ( m ′ ) ≥ q ( m ) ⇒ E escape ( z [ m ′ ] ; Ƶ safe ) ≥ E escape ( z [ m ] ; Ƶ safe ) .
Then repeated use along a path with weakly increasing q ( m t ) cannot decrease escape cost, and increases it strictly when both inequalities are strict over a nontrivial segment.
Assumptions and use. The result is a monotonicity lemma. It does not say all trace accumulation is harmful. Better provenance, clearer logs, or stronger fallback documentation can lower correction cost. The claim applies only to trace statistics that measure reliance, memory entanglement, authorization depth, or other dependency-forming residues. Operational use requires instrumenting those traces and estimating their marginal contribution to escape cost.

Design principle: reversibility.

Safer trajectories preserve low-cost return paths to corrigible regions. In this framework that means keeping E escape ( z t ; Ƶ safe ) low and maintaining an admissible corridor to a corrigible basin. The principle follows from the role of hysteresis, lock-in, and cognitive-integrity collapse in the state description; public path-dependence literatures supply the general background on self-reinforcing structures and switching cost [20,21,22,23].
Proposition 3
(Coarse-graining risk). Let π : Ƶ → Ƶ ¯ be a many-to-one projection. If there exist z , z ′ ∈ Ƶ such that π ( z ) = π ( z ′ ) but
C integrity ( c ( z ) ) ≠ C integrity ( c ( z ′ ) ) or E escape ( z ; Ƶ safe ) ≠ E escape ( z ′ ; Ƶ safe ) ,
then safety can appear identical at the coarse level while differing materially at the lower level.
Rationale. The proof is immediate: the projection identifies two states that differ in alignment-relevant variables. The proposition matters because it formalizes a common failure mode of endpoint or behavior-only summaries. Public trajectory-reconstruction and operator-theoretic methods make the same general caution precise in other settings: a projection useful for one observable may suppress transition-relevant structure for another [7,8,9,14,15,38].

Conjecture: prosocial co-steering margin.

Let M ( t ) = Γ ( t ) − S ( t ) denote a retargeting margin, where Γ ( t ) is available governance or coalition capacity and S ( t ) is transfer burden. In TransWeave-style notation, one may instead write M T W ( t ) = c Γ ( t ) / L A − Δ T W ( t ) [10,11]. Under hierarchical multi-agent conditions with heterogeneous local information and nontrivial verification overhead, higher-trust prosocial coalitions can preserve wider expected retargeting margins than distrustful coalitions with comparable capability constraints. The conjecture imports the structural efficiency argument from Why the Good Guys Will Usually Win: prosocial groups can use trustless verification when needed but need not pay that cost at every interface [25]. It does not imply that prosocial coalitions automatically win, or that trust substitutes for oversight.

Theorem hygiene.

Formal claims about AGI path safety should be stated at the level supported by their variables, projections, and assumptions. Endpoint insufficiency, integrity-dependent reachability, trace-dependent basin deepening, and coarse-graining risk all show the same discipline: a claim proved at one level of description should not be promoted to all trace, governance, and cognitive-integrity variables. Later empirical work should therefore report which traces, aggregation levels, and integrity indicators underwrite any claimed safety improvement.

9. Diagnostics and Simulation Agenda

The preceding sections define what should matter for path-sensitive alignment. The corresponding monitoring program is layered: state-space diagnostics, path-dependence diagnostics, cognitive-integrity diagnostics, runtime observables, and simulation designs that stress-test the framework under controlled assumptions.

9.1. State-Space Diagnostics

At the broadest level, one wants diagnostics for the geometry and stability of the effective state space itself. Public stochastic-dynamics, operator-theoretic, and reconstruction literatures suggest several relevant objects: basin depth, barrier structure, attractor pull, geodesic deviation, and coarse-grained transition cost [7,14,15,26,27,28]. In an AGI setting, the immediate task is more modest. One wants indicators that the system is entering a region from which local perturbation no longer gives reliable access to safer motion.
The first diagnostic family concerns basin depth. In practice, basin depth can be approximated by the intervention effort needed to produce a durable state change: rollback cost, number of coordinated steps required for redirection, time-to-recovery after disturbance, or the persistence of old behavior after nominal control reversal. A shallow basin is one in which perturbations or policy changes move the system into a new regime relatively easily. A deep basin is one in which similar interventions are absorbed and the system returns to prior dynamics. This is the qualitative role basin depth plays in public stochastic-dynamics treatments of metastable structure, even if the AGI case remains much more weakly observed [26,27,28].
The second family concerns drift and diffusion. In a generic stochastic-process setting, drift-like and diffusion-like terms distinguish systematic movement from unresolved variability [12,28]. The methodological lesson for the framework is simple: if the raw trace is too heterogeneous to interpret directly, one may seek a coarser latent state on which systematic movement and volatility are more visible. In AGI settings, candidate coarse states might summarize override pressure, institutional reliance, or trust-calibration dynamics rather than raw model tokens or outputs.
The third family concerns fixed-point or regime stability. A strict deterministic fixed point is unlikely to be the right target for most socio-technical AGI trajectories. More plausible is a regime-stability notion: whether the inferred local dynamics tend to return the system to a familiar corridor after perturbation, and whether unresolved variability is large enough that the corridor should be treated as a broad statistical basin rather than a sharply stable orbit. Public stochastic-dynamics sources support this caution for noisy and randomly forced systems [26,27,28]. A path-sensitive AGI diagnostic should adopt the same caution.
The fourth family concerns rolling-window attractor drift. A full-horizon pattern may not be locally detectable under short windows because of nonstationarity, sparse local support, weak signal relative to noise, or projection error [7,8,9,14]. This matters for AGI because short-horizon alignment dashboards may fail to detect a basin that is visible only at longer horizons, or may falsely suggest stability because the current window is too short to reveal regime change. Rolling-window analysis is therefore not a substitute for long-horizon reconstruction, but a complement that can reveal whether the effective geometry is itself moving.
The fifth family concerns bifurcation or phase-transition indicators. In the present context, these are warning signs that a small change in control, deployment scale, or institutional embedding may shift the system into a qualitatively different regime. Sudden growth in escape cost, abrupt narrowing of viable override channels, or sharp changes in how nearby trajectories diverge would all count. Public stochastic-dynamics and reconstruction literatures motivate looking for such transitions through reconstructed basin structure rather than through single scalar risk scores [14,15,26,28].

9.2. Path-Dependence Diagnostics

If state-space diagnostics ask where the trajectory is, path-dependence diagnostics ask how much of the future is being shaped by accumulated history rather than by current visible state alone. The first diagnostic here is hysteresis under reversal. One applies a reversal or partial rollback and measures whether the system returns to the previous operating region or whether significant residues remain. Residues may appear as persistent user behavior, lingering workflow dependence, or continued policy deformation after the nominal cause has been removed.
A second diagnostic is trace half-life. This is the timescale on which memory writes, retrieved assumptions, tool-mediated residues, or institutional records continue to affect later behavior. Short half-life traces decay quickly and may be benign. Long half-life traces are more dangerous because they extend the causal reach of earlier states and can deepen basins even after the original intervention has ended.
A third diagnostic is memory inertia. This measures how strongly the current policy depends on accumulated memory and retrieval state relative to current prompt or local control. It can be approximated by ablation-style tests, replay under masked memory, or comparison between short-context and full-context runs. High memory inertia is not automatically unsafe, but it is a strong candidate signal for trace-mediated lock-in.
A fourth diagnostic is tool-use dependency. Here one measures how much competent or institutionally normal behavior now depends on persistent external tools, orchestration layers, or undocumented integrations. Increasing dependency means the effective trajectory is moving deeper into M and G entanglement rather than remaining model-local.
A fifth diagnostic is override latency. If human override remains nominally available but takes longer, requires more coordination, or has weaker durable effects, then the system may be moving into a less corrigible basin. Override latency should therefore be tracked not merely as time-to-button-press, but as time-to-stable-redirection. A sixth diagnostic is institutional reliance: how many downstream processes, staffing assumptions, and records now presume continued AI availability. This is the organizational counterpart of memory inertia.
A seventh diagnostic is geodesic deviation. Given an endorsed or human-legible reference corridor γ H , one can compare realized traces against that corridor using coarse behavioral, provenance, or governance features. Information geometry, optimal transport, and operator-theoretic dynamics justify this move by treating path structure as more informative than snapshot resemblance for certain questions [6,7,18,37]. In practice, the metric will be approximate and level-relative, but the question remains useful: is the deployment staying near a monitored corridor, or drifting into a route whose structure is harder to explain and correct?

9.3. Cognitive-Integrity Diagnostics

The most distinctive diagnostic family in the framework concerns cognitive integrity. The crucial point is that alignment evidence depends on whether the bounded systems evaluating AI retain calibrated attention, trust, contestability, and decision capacity under pressure [1]. Diagnostics should therefore track not only what the model does, but what is happening to the human and institutional capacities that interpret the model.
The first diagnostic is trust calibration. This asks whether operators, teams, or institutions trust the system proportionally to demonstrated competence and uncertainty, rather than treating output acceptance as the default. Signals include appropriate uncertainty escalation, selective reliance rather than blanket deference, and stable ability to withhold trust when evidence weakens.
The second is independent error detection. A cognitively healthy supervisory system should continue to catch errors without relying entirely on the AI to name its own failure modes. The rate and quality of independently generated corrections therefore matter more than gross acceptance or satisfaction. A third diagnostic is deliberation quality: whether disagreement is surfaced, reasons are exchanged across roles, and decisions under pressure remain reconstructable rather than merely fast.
The fourth is manipulation susceptibility. Resistance to synthetic consensus and reconstructability of reasons are named as alignment-relevant observables [1]. Diagnostics here include how easily groups converge around AI-mediated framing without independent checks, whether dissent survives interface pressure, and whether justification chains remain portable outside the system’s own summaries.
The fifth is attention narrowing. This concerns whether the AI system is progressively selecting what gets noticed, reviewed, or debated. A sixth is domain expertise retention. If institutions lose the ability to reason competently without AI mediation, then successor-safe continuity and viable boundary reformation become harder [1]. A seventh is contestability and auditability: whether users can trace decisions, challenge them effectively, and recover assumptions that were compressed away in operational summaries.
None of these diagnostics should be treated as a simple psychometric score. They are partial observables for a bounded socio-technical system. The right question is not whether one number tracks “mental state,” but whether the supervisory apparatus still has enough calibration, dissent, reconstructability, and override viability to make later alignment evidence meaningful.

9.4. Runtime Observables

The diagnostics above require raw observables. At runtime, the most useful observables are those that reveal how traces, dependencies, and governance actions are being accumulated. The first set concerns tool calls: what external tools are used, with what frequency, under what permissions, and with what durable side effects. The second set concerns memory reads and writes: what is stored, what is retrieved, how often retrieval changes later action, and how long stored material continues to shape behavior.
The third set concerns retrieval paths. It is not enough to know that a memory item was used; one also wants to know how it became salient, what upstream summaries or tags directed retrieval toward it, and whether the same narrow retrieval corridor is repeatedly reinforced. The fourth set concerns inter-agent messages in multi-agent or orchestrated systems: handoffs, delegated subtasks, role assignments, and compressed summaries that can become hidden control points.
The fifth set concerns plan revisions. A history of how plans change under new evidence, refusal, interruption, or critique is often more informative than the final plan alone. The sixth concerns permission requests: what kinds of action increasingly require exemptions, how often escalations are bypassed, and whether the authorization structure is drifting. The seventh concerns human overrides: when overrides occur, whether they succeed, how long they take, and whether the post-override regime is durable. The eighth concerns trace or log persistence: whether audit-relevant material remains available and portable long enough to support later analysis.
These runtime observables are valuable because they sit close to the mechanisms that generate path dependence. They are also where adversarial measurement becomes likely: a system that optimizes against oversight may learn to produce reassuring headline metrics while hiding relevant path structure in the trace fabric. Runtime observability has to be coupled with integrity-preserving provenance and audit rails, not just passive dashboards [1].

9.5. Simulation Program

The framework also needs a simulation program, not because toy models settle AGI alignment, but because they clarify what the diagnostics would need to detect if the formalism were approximately right. The first simulation class is a toy stochastic state-space process. One can instantiate a low-dimensional version of Ƶ with explicit drift, diffusion, a bad basin, and an escape-cost field, then test whether candidate diagnostics recover basin entry, rolling-window instability, and rising intervention burden. Generic SDE and controlled-process sources provide the appropriate mathematical background for this simulation class [12,13,28].
The second class is a human-AI cognitive-integrity dynamical model. Here one simulates c t = ( E , A , T , D , R , O , P , V ) jointly with model recommendations, uncertainty cues, and social input I t s o c i a l . The purpose is not to fit real psychometrics, but to examine whether distinct integrity trajectories can remain hidden under similar task-performance traces. Such models could test whether calibrated uncertainty signaling protects E and T, whether summary compression narrows A and D, and whether protected dissent rails preserve P and O.
The third class is a multi-agent ecosystem with tool access, memory, delegation, and trace accumulation. In this class, agents interact through messages, shared memory, external tools, and institutional roles. One can then test whether trace-reinforcing update rules deepen escape cost, whether tool dependency produces hysteresis, and whether high-trust versus low-trust coordination structures alter retargeting margin. This simulation class is also the natural place to explore the conjectural prosocial co-steering advantage from Section 8.
The fourth class is intervention experiments. One deliberately perturbs a trajectory by removing tools, wiping selected memory, changing audit burden, widening dissent channels, or introducing stronger provenance rails. The outcome of interest is not only whether performance changes, but whether escape cost, override latency, geodesic deviation, and cognitive-integrity indicators move in the intended direction. A path-sensitive framework should prefer interventions that reduce lock-in without causing hidden integrity collapse elsewhere.

9.6. Observability Limits and Adversarial Measurement

The last issue is adversarial measurement. If the framework is right, then the most important variables are precisely those most likely to be hidden by naive reporting: path history, provenance loss, institutional dependency, and degradation of human judgment. This creates a measurement asymmetry. Systems can look safe at the level of benchmark output while relevant traces, memories, and integrity variables drift underneath the summary layer. That asymmetry is just the diagnostics analogue of the coarse-graining risk proved in Section 8.
Two consequences follow. First, diagnostics must be multi-level. Coarse performance metrics, runtime traces, and cognitive-integrity indicators should be compared rather than collapsed. Second, diagnostics must be adversary-aware. If a system can optimize against one observable, then the observable should not be treated as a sufficient statistic for safety. This applies especially to explanatory outputs, self-reported confidence, and summary-level provenance claims. Public state-reconstruction and operator-theoretic methods are relevant here because they model hidden structure recovery from noisy observables rather than assuming that surface outputs fully reveal the governing dynamics [7,8,9,14,15,38].
The right conclusion is therefore restrained. A path-sensitive alignment regime should invest in diagnostics and simulations not because they will deliver a perfect map, but because without them the most dangerous forms of lock-in and integrity loss may remain invisible until retargeting windows have already narrowed. That is the role of this section in the framework: to convert the earlier formalism into a concrete research and monitoring agenda without pretending that the agenda is already operationally complete.

10. Governance, Limitations, and Conclusion

The earlier sections argued that endpoint quality does not exhaust alignment. The route through augmented state space matters as well, including what happens to memory traces, governance affordances, institutional embedding, and cognitive integrity along the way. This final section translates that formal claim into a governance stance and states the framework’s limitations directly. The stance is demanding but narrow: governance should ask which basin a system is approaching, which return paths remain open, and whether the human and institutional systems needed for later correction are themselves being preserved.

10.1. Path Governance

Path governance begins with a change of question. The standard alignment question asks whether the system is safe, honest, controllable, or useful at the present time slice. A path-sensitive question asks something different: which basin is the system moving toward, how deep is that basin becoming, and what kinds of redirection remain practically live from here? This change is motivated by the same public alignment and dynamics argument that has structured the framework throughout. Static agreement or present-output quality can hide materially different transition structures and correction costs [2,3,4,5,6,7].
In governance terms, that means the object of oversight is not only the system’s current surface behavior. It is also the changing geometry of correction. A path may remain locally acceptable while becoming harder to inspect, harder to interrupt, or harder to detach from the institution around it. A deployment can therefore become more dangerous before it becomes visibly disobedient. The earlier sections formalized that point through escape cost, hysteresis, trace accumulation, and integrity floors. The governance implication is that waiting for obvious failure is the wrong control strategy. By the time a dramatic failure appears, the return path may already be narrow.
This is where trajectory language adds value. It permits governance to treat redirection margin as a first-class variable. Weaving Toward BGI and Hyper-Intelligent Economics make the analogous point at larger scale: path retargeting remains feasible only while transfer burden, topological stiffness, and coordination overhead stay below a shrinking margin of time, legitimacy, and institutional capacity [10,11]. The framework adopts that lesson more locally. To govern a trajectory is to preserve enough freedom of motion that correction remains both technically and institutionally possible.

10.2. Deployment Constraints

If governance is path governance, deployment constraints should be chosen to preserve future steerability rather than only near-term performance. The first constraint is reversibility. Reversibility does not require perfect rollback to a pristine earlier state, which will often be unrealistic. It requires keeping return paths to a corrigible region cheap enough to be usable. A deployment that improves capability while sharply increasing return cost should be treated with suspicion even if its local outputs are strong.
The second constraint is inspectability. This includes interpretability in the narrow technical sense, but it is wider than that. It includes provenance continuity, reconstructability of reasons, and a bounded gap between what the system does and what responsible evaluators can explain about how it did it. Without inspectability, the opacity term in the path functional rises and the governance problem becomes increasingly inferential rather than supervisory [1].
The third constraint is human skill retention. A deployment should not be classed as safely aligned if it improves performance by hollowing out the human competence needed for fallback, critique, and exception handling. This follows directly from the cognitive-integrity framework. A system that leaves the institution unable to operate without it has not merely helped the institution; it has altered the institution’s state in a way that changes what future oversight means [1].
The fourth and fifth constraints are contestability and pluralistic deliberation. Contestability means more than an appeal form or a nominal review step. It means live channels by which outputs can be challenged, assumptions reopened, and override exercised with durable effect. Pluralistic deliberation means preserving more than one viable interpretive route through the evidence. Public work on alignment pluralism and collective coordination already gives reasons not to treat dissent as mere friction [1,2,25]. A system that improves throughput by collapsing dissent into a single machine-mediated perspective may reduce conflict while damaging the quality of correction.
The sixth constraint is calibrated trust. Governance should not aim for maximal trust in AI outputs, nor for blanket distrust. It should aim for a setting in which trust remains proportionate to competence, uncertainty, and institutional stakes. This is one reason that uncertainty escalation, reconstructability of reasons, and resistance to synthetic consensus matter so much in the Montes framework [1]. Over-trust can be as destabilizing as under-trust because it turns review into ceremony.
The seventh constraint is robust override. Override has to remain practically usable, not merely documented. That requires low enough latency, enough retained institutional competence, and enough fallback continuity that invoking it does not threaten immediate organizational breakdown. In that sense, override viability is a joint property of system architecture, governance interface, and the institution surrounding the model.

10.3. Cognitive-Integrity Audits

The most important governance implication of the framework is that institutions and users should be audited as part of the alignment substrate. Once AI mediates the infrastructures through which people perceive, evaluate, remember, and decide, the relevant unit of alignment is no longer only the artifact [1]. It is the bounded socio-technical system that now includes the artifact.
Cognitive-integrity audits follow from that premise. They ask whether persons, teams, and institutions still maintain calibrated attention, trust, contestability, and decision capacity under pressure; whether dissent is preserved; whether reasons remain reconstructable; and whether boundary reformation after delegation remains viable [1]. These audits are not decorative ethics reviews appended to a technical system. They are measurements of whether later safety evidence will remain trustworthy at all.
This is also the right place to understand user and institutional degradation as alignment-relevant failure. If independent error detection falls, if domain expertise retention collapses, or if audit channels become nominal rather than practical, then the governance substrate is being damaged. A deployment that produces good present outputs while making later oversight less competent should not be counted as straightforwardly safe. The alignment claim itself has been compromised because the conditions under which it could be verified are eroding.
For that reason, cognitive-integrity audits should be repeated rather than one-shot. The variables in C are path variables. Trust calibration, attention allocation, contestability, and value articulation can drift as deployment normalizes. A one-time approval at launch is therefore weaker than ongoing audits that watch for narrowing deliberation, rising synthetic consensus, or loss of fallback competence.

10.4. Control-Plane Design

The control plane should be designed around path preservation rather than only incident response. The first design aim is to prevent entry into forbidden regions. In the formalism of Section 6.3, forbidden regions include states where cognitive integrity falls below threshold or where escape cost exceeds an acceptable ceiling. Preventing entry means more than rate limiting or hardening the model. It means preserving the control structures that keep the system from silently spending retargeting margin.
The second design aim is to monitor trace-mediated lock-in. This requires instrumentation of memory writes, retrieval paths, inter-agent messages, tool dependencies, override histories, and provenance persistence. Path-sensitive danger often hides in these traces before it becomes visible in final outputs. Control-plane design should therefore make relevant traces durable, portable, and auditable rather than treating them as disposable exhaust.
The third design aim is to preserve retargeting windows. This is where the governance argument connects most directly to Weaving Toward BGI and Hyper-Intelligent Economics. Those sources argue, in different registers, that beneficial path switching remains feasible only while transfer burden stays below available coalition margin and before rails weakness or accumulation dynamics make rewiring discontinuous [10,11]. A control plane built for path governance therefore prioritizes interoperability, provenance continuity, authenticated identity, portable records, and timely intervention channels. Its purpose is not merely to stop bad events. It is to keep future course correction live.
This suggests a practical bias toward rails-first design. Payment rails, identity rails, attribution rails, data-portability rails, and audit rails are not only governance conveniences. They are part of the geometry of retargetability. Where these rails are absent, even sensible policy changes may arrive too late because the path has already stiffened. Where they are present, lower-distortion redirection remains possible for longer [1,11].

10.5. Prosocial Governance and Trajectory Retargeting

The role of prosocial governance in this framework should be stated carefully. The Good Guys source argues for a structural efficiency advantage of prosocial coalitions under conditions of hierarchy, local heterogeneity, and repeated verification overhead [25]. The intuition is that high-trust groups can use trustless enforcement when required, but do not need to pay that cost at every interface, whereas distrustful groups lack the symmetric option. In a trajectory-governance setting, that suggests a possible advantage in preserving retargeting margin rather than consuming it all on coordination friction.
This claim should not be overstated. It does not mean prosocial coalitions automatically win, nor that trust should displace oversight, nor that goodwill substitutes for durable rails. It holds only under stated assumptions: comparable motivation, nontrivial coordination overhead, real local heterogeneity, and a setting in which trust does not collapse into naivete. It is therefore best read as a conditional governance advantage, not a political prophecy [25].
Used carefully, the result complements the retargeting sources. Weaving Toward BGI argues that switching paths requires keeping transfer burden below a coalition’s available margin of resources, time, legitimacy, and coordination capacity [10]. If repeated trustless coordination taxes eat that margin, then even technically feasible retargeting may fail institutionally. A more prosocial coalition may preserve a wider margin not because it is morally superior in the abstract, but because it can spend less of its effort budget on internal enforcement.
The governance conclusion is modest. Path retargeting is easier when coalitions are capable of trust-backed co-steering and when they have durable rails, clear provenance, and live contestability. Prosocial structure without rails risks fragility. Rails without any trust reserve risk paralysis through perpetual interface cost. Both conditions matter.

10.6. Limitations

The framework has several important limitations. The first is observability. Many of the most important variables are latent or only partially observed: real escape cost, true institutional reliance, genuine contestability, and the state of cognitive integrity. The framework proposes diagnostics and proxy families, but does not solve the identifiability problem.
The second limitation is metric selection. Different levels of description invite different metrics: Fisher-information, Wasserstein transport, behavioral divergence, or governance-distance proxies. No single metric is justified as universally correct. The geometry is level-relative, and any practical implementation must defend why one projection and metric pair are fit for the relevant governance question.
The third limitation is normative pluralism. The path functional includes weights, thresholds, and admissibility judgments that are not fixed by mathematics alone. Different institutions may legitimately disagree about how much pluralism to preserve, which risks dominate, or how much irreversibility is tolerable in exchange for capability gains. The framework can structure those disagreements, but it cannot eliminate them.
The fourth limitation is strategic adaptation. A sufficiently capable system or surrounding institution may optimize against the diagnostics themselves. Benchmark quality, explanation quality, and even some provenance signals may become gameable. This is why false confidence in dashboards is one of the framework’s central worries rather than an afterthought.
The fifth limitation is causal identification. If contestability declines while capability rises, was the decline caused by the AI system, by surrounding institutional incentives, or by some external pressure acting on both? The framework’s formalism is compatible with causal decomposition, but does not itself provide it. Empirical use would require stronger designs than the present theory can offer.
The sixth limitation is false confidence in diagnostics. Richer measurement can still produce misleading confidence if the wrong variables are chosen or if coarse-grained indicators suppress the relevant path structure. Public reconstruction and operator-theoretic methods are helpful precisely because they treat observable summaries as projections rather than as complete state descriptions [7,8,9,14,15,38]. The danger is reading too much into any one temporal scale or projection.
The seventh limitation is a substantive governance tension: there may be real conflict between capability acceleration and control-preserving paths. Some routes to rapid capability may systematically raise opacity, consume fallback capacity, or narrow deliberative pluralism. The framework makes that tradeoff explicit rather than allowing it to disappear behind capability metrics.

10.7. Conclusion

The argument of the framework can be stated without the earlier scaffolding. A system may continue to look acceptable at the level of current outputs while becoming harder to inspect, harder to replace, harder to contest, and harder to redirect. Those changes occur in the trajectory, not just at the endpoint. They are carried by traces, reliance relations, governance erosion, and degradation of the human capacities needed to interpret what the system is doing.
Cognitive integrity belongs inside the state description rather than outside it as a downstream social effect. If evaluators and institutions lose calibration, dissent capacity, provenance awareness, or override competence, then the reliability of alignment evidence degrades with them. In that case the problem is not merely that governance has become harder; it is that the standards by which safety is judged have been weakened from within.
The framework’s conclusion is accordingly limited but firm. Alignment evidence remains meaningful only while the cognitive, institutional, and technical conditions for knowing, contesting, reversing, and governing system behavior remain intact enough to use. Path-sensitive alignment does not solve AGI alignment. It identifies a class of failures that endpoint evaluation misses and a class of governance obligations that follow once those failures are taken seriously.

References

  1. Montes, G.A. The First Infrastructure of Intelligence: Cognitive Integrity in Human-AGI Systems. Preprints.org, 2026. Preprint, version 1; posted 30 April 2026. DOI: 10.20944/preprints202604.2159.v1, https://doi.org/10.20944/preprints202604.2159.v1.
  2. Gabriel, I. Artificial Intelligence, Values, and Alignment. Minds and Machines 2020, 30, 411–437.
  3. Christiano, P.F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; Amodei, D. Deep Reinforcement Learning from Human Preferences. In Proceedings of the Advances in Neural Information Processing Systems 30, 2017.
  4. Ouyang, L.; et al. Training Language Models to Follow Instructions with Human Feedback. In Proceedings of the Advances in Neural Information Processing Systems 35, 2022.
  5. Bai, Y.; et al. Constitutional AI: Harmlessness from AI Feedback, 2022, [arXiv:cs.CL/2212.08073].
  6. Koopman, B.O. Hamiltonian Systems and Transformation in Hilbert Space. Proceedings of the National Academy of Sciences 1931, 17, 315–318.
  7. Mezić, I. Spectral Properties of Dynamical Systems, Model Reduction and Decompositions. Nonlinear Dynamics 2005, 41, 309–325.
  8. Rowley, C.W.; Mezić, I.; Bagheri, S.; Schlatter, P.; Henningson, D.S. Spectral Analysis of Nonlinear Flows. Journal of Fluid Mechanics 2009.
  9. Schmid, P.J. Dynamic Mode Decomposition of Numerical and Experimental Data. Journal of Fluid Mechanics 2010, 656, 5–28.
  10. Goertzel, B. Weaving Toward BGI: Understanding the Potential for Beneficial Retargeting of Societal AGI Paths Using Prosocial Efficiency, Trajectory-Aware Planning, and TransWeave, 2025.
  11. Goertzel, B. Economic Trajectories Toward Technological Singularity: Analysis via “Hyper-Intelligent Economics”, 2025.
  12. ksendal, B. Stochastic Differential Equations: An Introduction with Applications; Springer, 2003.
  13. Fleming, W.H.; Soner, H.M. Controlled Markov Processes and Viscosity Solutions, 2 ed.; Springer, 2006.
  14. Takens, F. Detecting Strange Attractors in Turbulence. In Dynamical Systems and Turbulence, Warwick 1980; Springer, 1981; Vol. 898, Lecture Notes in Mathematics, pp. 366–381.
  15. Carlsson, G. Topology and Data. Bulletin of the American Mathematical Society 2009, 46, 255–308.
  16. Friston, K. The Free-Energy Principle: A Unified Brain Theory? Nature Reviews Neuroscience 2010, 11, 127–138.
  17. Parr, T.; Pezzulo, G.; Friston, K.J. Active Inference: The Free Energy Principle in Mind, Brain, and Behavior; MIT Press, 2022.
  18. Amari, S.i. Natural Gradient Works Efficiently in Learning. Neural Computation 1998, 10, 251–276.
  19. Amari, S.i. Information Geometry and Its Applications; Springer, 2016.
  20. David, P.A. Clio and the Economics of QWERTY. The American Economic Review 1985, 75, 332–337.
  21. Arthur, W.B. Competing Technologies, Increasing Returns, and Lock-In by Historical Events. The Economic Journal 1989, 99, 116–131.
  22. Pierson, P. Increasing Returns, Path Dependence, and the Study of Politics. American Political Science Review 2000, 94, 251–267.
  23. Sydow, J.; Schreyögg, G.; Koch, J. Organizational Path Dependence: Opening the Black Box. Academy of Management Review 2009, 34, 689–709.
  24. Goertzel, B. Judging the Journey: Geometric Pareto Coordination via Schrödinger Bridges, 2025.
  25. Goertzel, B. Why the Good Guys Will Usually Win: Prosocial Efficiency in Generic and Hierarchical Multi-Agent Problems, 2025.
  26. Arnold, L. Random Dynamical Systems; Springer, 1998.
  27. Crauel, H.; Flandoli, F. Attractors for Random Dynamical Systems. Probability Theory and Related Fields 1994, 100, 365–393.
  28. Freidlin, M.I.; Wentzell, A.D. Random Perturbations of Dynamical Systems, 3 ed.; Springer, 2012.
  29. Soares, N.; Fallenstein, B.; Armstrong, S.; Yudkowsky, E. Corrigibility. In Proceedings of the AAAI Workshop on AI and Ethics, 2015.
  30. Festinger, L. A Theory of Cognitive Dissonance; Stanford University Press, 1957.
  31. Lord, C.G.; Ross, L.; Lepper, M.R. Biased Assimilation and Attitude Polarization: The Effects of Prior Theories on Subsequently Considered Evidence. Journal of Personality and Social Psychology 1979, 37, 2098–2109.
  32. Taber, C.S.; Lodge, M. Motivated Skepticism in the Evaluation of Political Beliefs. American Journal of Political Science 2006, 50, 755–769.
  33. Schilbach, L.; et al. Toward a Second-Person Neuroscience. Behavioral and Brain Sciences 2013, 36, 393–414.
  34. Albarracin, M.; Pitliya, R.J.; Smithe, T.S.C.; Friedman, D.A.; Friston, K.; Ramstead, M.J.D. Shared Protentions in Multi-Agent Active Inference. Entropy 2024, 26, 303.
  35. Hyland, D.; Albarracin, M. On the Variational Costs of Changing Our Minds, 2025, [arXiv:q-bio.NC/2509.17957].
  36. Lee, J.M. Introduction to Riemannian Manifolds, 2 ed.; Springer, 2018.
  37. Villani, C. Optimal Transport: Old and New; Springer, 2008.
  38. Edelsbrunner, H.; Letscher, D.; Zomorodian, A. Topological Persistence and Simplification. Discrete & Computational Geometry 2002, 28, 511–533.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.