Preprint
Article

This version is not peer-reviewed.

GLHS: A Co-Versioned Disclosure-to-Commit Governance Contract for Longitudinal Health AI

Submitted:

23 September 2026

Posted:

24 September 2026

You are already at the latest version

Abstract
Purpose. Persistent health AI can read an authorized patient snapshot and return a write proposal minutes or hours later. By then, the record, consent state, or policy may no longer be the same. We examine how that read-to-write interval can be governed. Methods. GLHS was implemented in a reference platform. THSS records the governed snapshot supplied to the AI, and GST checks the relevant state and governance again before a proposal is written. Contract enforcement was evaluated separately from context utility. The model study enrolled 64 prospectively frozen synthetic subjects evaluated with Claude and Gemini; the 1,152 solver cells were model-condition evaluations, not independent subjects. Additional experiments covered contract conformance and PostgreSQL state-version concurrency. Results. Under Strict THSS, Claude was exact on all four axes for 63/64 subjects (98.44%) and Gemini for 64/64 (100%). Six of ten planned paired contrasts remained Holmsignificant, although only 9–21 discordant subjects informed those significant tests and the planned power target was not reached. In the tested PostgreSQL state-version races, stale writes were rejected. With unrelated writes, the false-stale pattern followed the expected one-winner consequence of a profile-global version counter. Conclusions. GLHS keeps the snapshot shown to an AI connected to any later proposal that seeks to change persistent state. The two-model cohort provides controlled synthetic evidence that Strict THSS can support longitudinal-state reasoning, while the concurrency experiments show the cost of coarse versioning. These results concern software behavior, not clinical effectiveness or regulatory compliance.
Keywords: 
;  ;  ;  ;  ;  

1. Introduction

Longitudinal records are difficult for reasons that go beyond retrieval. A fact may be clinically valid at one time and only become known later. Another source may disagree with it. The same information may be appropriate for one task or actor but not for another. A system can therefore retrieve something relevant and still show the AI the wrong historical state, or more of the record than the current use permits. The difficulty does not end when the model finishes reasoning: the record, consent state, or policy may change before the model’s proposal is written back.
The components needed to handle these problems are well established. Medical database work separates valid time from transaction or knowledge time [1,2,3]. FHIR Provenance and Consent represent lineage and policy-relevant information [4,5], and FHIR HTTP supports version-aware updates [6]. openEHR maintains versioned EHR content and reconstructable contributions [7], while W3C PROV provides a general provenance model [8]. Zhao et al. [9] already describe bitemporal, supersession-aware, conflict-preserving state arbitration for longitudinal health AI, and related systems address persistent patient state, reconciliation, and provenance-linked timelines [10,11,12].
Recent agent-memory work covers much of the same ground from a different direction. MemTX couples transactional commit with evidence, permissions, provenance, and validity [13]. MemTxn moves update validation, temporal conflict handling, and recovery outside the answer model [14]. MemClaw addresses scoped retrieval and policy-governed propagation [15], while HealthClaw applies privacy-aware memory selection to longitudinal personal health management [16]. These papers make temporal state, governed memory, and transactional write control part of the prior-art baseline for this study.
Our question starts one step later. Suppose an AI has already received an authorized disclosure and then returns a proposal that would alter persistent state. A matching base version alone does not tell us which governed disclosure supplied the evidence and policy context for that proposal. GLHS keeps that relationship explicit across the two operations.
In GLHS, the disclosed snapshot remains part of the proposal’s admission context. Before commit, the service also checks whether the relevant state, authorization, consent, policy, and expiry conditions still hold. The work therefore concerns the way existing mechanisms are joined across the read and write boundary; it does not present bitemporality, provenance, consent enforcement, conflict preservation, task-bounded disclosure, or optimistic concurrency as separate inventions. For the snapshot-bound mode studied here, the operative invariant is that proposal admissibility requires an exact match to the persisted disclosure identity and digest together with governance that remains valid at commit.
Table 1 summarizes the relationship between these established mechanisms and the GLHS contract evaluated here.

2. Materials and Methods

2.1. Design Scope and Governed Longitudinal Substrate

GLHS represents a system view that is evidence-backed and permitted for the current use. It is not meant to stand in for an unquestionable biological truth. Evidence, events, assertions, state, transitions, conflicts, and projections remain separate so that provenance and uncertainty are not flattened into the current view. State can be reconstructed at clinical valid time t v and semantic knowledge time t k . Database write time stays in the audit trail; it is not substituted for knowledge time.
For subject u, we write the governed state as S u ( t v , t k ) . Assertions retain their provenance, epistemic status, and lifecycle. Conflicting assertions can remain visible together until a resolution becomes valid and known. This is the state from which a governed disclosure is compiled.
S u ( t v , t k ) = GovernedView u ( t v , t k ) .

2.2. Task-Bounded Health State Snapshot

THSS is the governed object produced on the read path. Besides the selected content, it stores the coordinates needed to reproduce what the AI was allowed to see. We write a snapshot H as:
H = 〈 u , a , r , ϕ , τ , v s , v π , v c , t v , t k , id H , d H , E H 〉 .
In Equation (2), u identifies the subject or profile, while a and r identify the actor and role; ϕ and τ denote purpose and task. The snapshot also records state version v s , policy version v π , consent/governance version v c , and the valid- and knowledge-time cutoffs t v and t k . id H names the persisted snapshot, d H is its consistency digest, and E H is the disclosed content. In the implementation, canonical bytes are produced using a project-defined canonicalization profile and hashed with SHA-256. The profile sorts object keys, preserves array order, emits compact UTF-8 JSON, normalizes timezone-aware datetimes to UTC and negative zero to zero, and rejects unsupported or non-finite values; it is not presented as RFC 8785.
Canonicalization is applied before digest persistence. The manifest records schema glhs.snapshot.v3, payload schema glhs.snapshot.payload.v3, the digest algorithm and canonicalization profile, and validation fails closed if these values or the recomputed digests disagree. Strict model-facing disclosure additionally records the consent basis, expiry, assertion identifiers and hashes, evidence identifiers, and an opaque profile scope.

2.3. Snapshot-Bound Persistent Proposal and GST

The model is not given unrestricted write access. Instead, it returns a persistent proposal P:
P = 〈 u , a , r , ϕ , τ , v s , v π , v c , id H , d H , Δ , ρ 〉 .
Δ is the proposed change and ρ records proposal provenance, including a model-manifest reference when applicable. The gateway supports a base-version-only path as well as snapshot-bound proposals. This paper evaluates the stronger snapshot-bound path. At the manuscript freeze, however, the service does not yet enforce a provenance-sensitive rule that forces every model proposal derived from Strict THSS to remain snapshot-bound after review. We therefore do not claim downgrade prevention: the exact-binding claims in Equations (2)–(4) apply to proposals submitted through snapshot-bound mode.
Commit ( P ) ⟺ Bound ( P , H ) ∧ StateCurrent ( P ) ∧ GovernanceCurrent ( P ) ∧ SnapshotValid ( P ) .
GovernanceCurrent ( P ) groups the authorization, consent, and policy checks. Bound ( P , H ) fails when the proposal does not agree with the snapshot on profile, actor, role, purpose, task, base state version, snapshot identity or digest, or the declared evidence set. During commit, the gateway obtains a FOR UPDATE lock on the profile row, rechecks the base state, rereads current consent, compares the frozen policy version, and then writes the transition and resulting state in the same database session. This serializes the tested state-version path. We did not establish that concurrent consent or authorization writers take the same lock, and no concurrent governance-revocation race is reported; cross-transaction governance TOCTOU atomicity therefore remains unproven.
The complete disclosure-to-commit path is illustrated in Figure 1.

2.4. Implementation and Reproducibility Boundary

The reference implementation was frozen for this manuscript. Snapshot-bound proposals are validated in the API-owned commitment gateway. Strict disclosures read the persisted manifest identity and digests after canonicalization and fail if required binding fields are missing. The persisted manifest carries the schema, payload-schema, SHA-256 digest-algorithm, and canonicalization-profile fields described above, and the validator rejects metadata or digest mismatch. State replay exposes separate valid-time and semantic knowledge-time cutoffs. Historical conflict visibility is reconstructed from conflict-creation and conflict-resolution transitions rather than from mutable current status. Generic and commitment THSS share the same executable five-stage trace. Each reported run keeps its own implementation SHA and checksum lineage; older results are not reassigned to this manuscript-level freeze.
Model-mediated evaluation used the completed 64-subject solver-v5 cohort on 11 August 2026. The run was tied to a sealed reproducibility record containing the implementation identifier, an internal preregistration SHA-256, and a pre-execution seal. Provider calls used an OpenAI-compatible /chat/completions endpoint with requested IDs antigravity/claude-sonnet-4-6 and antigravity/gemini-3.6-flash-high; returned IDs had to match claude-sonnet-4-6 and gemini-3.6-flash-high. Requests used temperature 0, max_tokens 1024, stream=false, and JSON-schema response mode; top_p and a provider API-version parameter were not set. The client timeout was 60 s and allowed at most two retries for declared transient HTTP failures, with 0.25-s initial backoff. The completed run had zero errors and zero retries. Prompt, schema, cohort, model-mapping and analysis inputs were part of the sealed transitive-input inventory rather than being changed after execution.

2.5. Comparator and Contract-Clause Isolation

The contract experiments and the model comparators answer different questions. The standards-composed mechanism comparator includes bitemporal resolution, version-aware writes, current authorization, provenance, and audit. We did not verify a single comparator that matches GLHS on every current consent, policy, purpose, and expiry check while differing only in exact snapshot binding, so it is used as a mechanism comparator rather than as a binding-only causal ablation. In the solver-v5 cohort, Strict THSS was compared with full authorized history, long chronological context, Naive RAG, LWW, and BTSA. These contrasts test the usefulness of disclosure context for model reasoning, not the later GST persistence boundary.
To exercise individual contract checks, we applied seven progressively stronger variants to the same 16 developer-authored cases, giving 112 deterministic conformance executions. They are not 112 independent samples. The variants add, in order: temporal/provenance resolution; base-version checking; authorization at disclosure; provenance and audit; snapshot identity; full context/digest plus disclosed-evidence membership; and current reauthorization with policy, consent, and expiry checks. Case order, clause order, runner digest, and engine digest were frozen. Separate property-based, adversarial, leakage, and boundary tests exercise the implementation outside this matrix; those tests are not pooled with the 112 cells, and mutation testing was not part of this study.

2.6. PostgreSQL Concurrency and Service-Layer Evaluation

We exercised state-version concurrency on PostgreSQL 16.14. A four-writer test covered both a shared semantic slot and four unrelated slots under the current profile-global version counter. A second contention grid used 1, 2, 4, 8, and 16 writers, with five independent profile races at each level, for 50 races and 310 writer attempts. Losses on the same semantic dependency were classified as true stale and rejections on unrelated slots as false stale. These runs did not introduce simultaneous authorization loss, consent revocation, or policy changes, so they do not test governance TOCTOU races.
Service-layer overhead was measured in a clean, one-worker, in-process run on PostgreSQL 16.14 using read-committed isolation, synchronous commit, a history depth of 50, and 30 repetitions per operation. The artifact was produced from the frozen implementation and its associated checksum inventory. HTTP transport, worker scaling, and provider inference were not included. Because only 30 repetitions were collected, we report p50 and p95 but omit p99; the rate column is a single-worker sequential operation rate, not concurrent throughput.

2.7. PII/PHI and Sensitive-Data Boundary

PII, PHI, and other sensitive health attributes are treated as governed data rather than ordinary model context. The canonical record may retain identifiers or source material when permitted, but THSS decides what is actually disclosed to the AI for the current actor, purpose, task, consent state, and policy. Direct identifiers or sensitive classes can therefore be omitted, minimized, or pseudonymized before model access while their provenance remains in the governed record.
Evaluation artifacts use synthetic or de-identified inputs where applicable, opaque study-local subject tokens, and content hashes. Credentials, unrestricted provider payloads, and raw licensed patient resources are excluded from tracked evidence bundles. Public or sealed benchmark copies also replace synthetic FHIR Patient/… references with an opaque study URN while keeping a separately verified source-cohort hash and redaction provenance. For a stored governed decision, reconstruction can recover the manifest and snapshot identifiers and digests, disclosed evidence/assertion set, state/policy/consent versions, actor/role/purpose/task, temporal cutoffs, transition identifier, base/result state versions, review status, and recorded reason/model references. The clause-ablation artifacts separately record the first rejecting clause for each conformance case; this should not be read as a claim that every production rejection is persisted at clause-level granularity. Hashes are not treated as anonymization guarantees.

2.8. Evidence Classification and Analysis

The preregistered model endpoint was subject-level all-axes exact match. Strict THSS was compared with five alternatives for each model, giving ten paired tests in one Holm family. Missing or malformed outputs counted as errors. For each contrast, paired subject-level differences were bootstrapped 1000 times with seed 20260811; the reported 95% interval is the percentile interval over those resamples. Exact two-sided sign tests used only discordant subjects and were Holm-adjusted across all ten tests. The frozen 64-subject plan targeted at least 51 non-tied pairs and approximately 81.5% power for a directional win probability of 0.75 under a conservative 0.005 per-test threshold. That target was an informative-pair requirement, not a guarantee about comparator discordance. The significant contrasts ultimately contained only 9–21 non-tied subjects, so inference is underpowered relative to the plan even where adjusted p-values are significant.

3. Results

3.1. Exact Disclosure-Context Binding and Clause Isolation

The focused tests exercised Equations (2)–(4) directly. Snapshot-bound proposals were rejected for mismatches in profile, actor, role, purpose, task, base state version, snapshot identity, or digest. The write path also rechecked current policy, consent, and expiry. Incomplete binding metadata was rejected, as were THSS traces with missing or reordered stages. A separate historical regression confirmed that a conflict resolved in August still appears in a July snapshot compiled later, because conflict visibility is replayed from transition history rather than read from the current conflict projection. These tests concern software conformance rather than clinical or legal correctness.
All 112 cells in the 16-case × 7-variant matrix produced the expected conformance result and identified the first clause responsible for rejection. The incremental matrix helps localize which added check changes each case outcome. Because the standards-composed comparator is not fully matched on every current governance semantic, however, these executions should not be read as a causal estimate of exact snapshot binding alone.

3.2. Atomic Write Safety and Version-Granularity Trade-Off

In the N = 4 PostgreSQL test, both the same-slot race and the unrelated-slot race produced one atomic winner and three stale rejections, with no database error. Under these specific concurrent-write conditions, the profile-global contract therefore failed closed.
The larger contention grid confirmed a direct consequence of profile-wide versioning. With one winner per simultaneous profile race, unrelated-slot false-stale rejection followed the expected pattern: 0 with one writer, 0.50 with two, 0.75 with four, 0.875 with eight, and 0.9375 with sixteen. Losses on the same dependency were true stale. We therefore report these values as a design consequence reproduced by the benchmark, not as estimated population rates. Resource-partitioned and dependency-aware alternatives were not run on PostgreSQL.

3.3. Service-Layer Overhead

The descriptive service-layer measurements are reported in Table 2.
Peak process RSS across these operations was approximately 109 MiB. The measurements describe one local, in-process service-layer environment. They should not be read as public-HTTP latency, deployment capacity, or production throughput.

3.4. Prospective Two-Model Solver-v5 Evaluation

The solver-v5 benchmark completed all 1152 solver cells with zero errors and zero retries. Across those cells, accuracy was 93.06% for lifecycle, 96.79% for evidence, 99.22% for timeliness, and 99.83% for escalation; 1027/1152 cells (89.15%) were exact on all four axes. On the subject-level primary endpoint, Strict THSS reached 63/64 (98.44%) with Claude and 64/64 (100%) with Gemini. The single remaining Strict-THSS Claude error—a cancelled history predicted as satisfied—was retained without post-hoc repair. The 128 construction-review requests served a separate, bounded purpose: two model reviews per subject checked the deterministic anchor construction, while code retained final authority. All 64 constructions were accepted; candidate-slot F1 and due-window exact accuracy were both 1.0.
The ten preregistered paired contrasts are reported in Table 3.
Six contrasts remained Holm-significant, but only 9–21 discordant subjects informed those tests, well below the planned 51 informative pairs. Claude versus BTSA was not significant (+4.69 percentage points, Holm p = 1.0 ), and Gemini tied BTSA, full history, and long context exactly. The exact sign tests remain valid for the observed discordances, but the high tie rate limits effect-size precision and confirmatory strength.

3.5. Claim Boundary

The model cohort and the contract tests should not be read as the same kind of evidence. The former asks whether the THSS context helps with a frozen synthetic reasoning task; the latter asks whether the software keeps that disclosure tied to a later write and rechecks it at commit. In the model cohort, Strict THSS performed better than several context constructions for both models, but not every comparator. None of these results measures clinical accuracy or clinical safety. Independent clinical adjudication, real-world policy completeness, deployed-boundary security testing, and prospective patient outcomes were not part of this study.

4. Discussion

4.1. What Is Novel—And What Is Not

Bitemporal health state, provenance, consent representation, conflict preservation, and version-aware writes all have established antecedents [1,2,3,4,5,6,7,8,9]. Transactional commit, governed propagation, and privacy-aware longitudinal memory also appear in recent agent-memory systems [13,14,15,16]. What GLHS adds is the connection across the system boundary: the disclosure supplied to the AI remains part of the admission check when a later proposal tries to change persistent state.
A version check alone can reject stale writes, but it does not identify the disclosure from which a proposal arose. Conversely, a disclosure policy can restrict what the model sees and still disappear from the decision path once the model returns an output. GLHS carries the same coordinates across both steps. THSS records the governed read context; the proposal cites that context and its base state; GST checks them again before persistence.

4.2. Relation to Transactional and Governed Agent Memory

MemTX is the closest transactional neighbor because it separates staged memory writes from committed beliefs and includes permissions, provenance, and validity in the commit discipline [13]. MemTxn likewise keeps validation and recovery outside the answer model [14]. MemClaw treats scope, temporal contradiction, provenance, and policy propagation as service concerns [15], and HealthClaw studies privacy-aware memory selection in longitudinal health assistance [16]. GLHS places its boundary at a different point: the purpose- and consent-bounded health disclosure is still consulted when the later proposal is admitted for persistence.

4.3. Safety, Privacy and Operational Implications

The same boundary is also useful when the world changes while a model is working. If state, consent, actor permission, policy, or snapshot validity changes during inference, the proposal can be rejected or reconciled instead of being written silently. The strict representation carries the persisted manifest and its temporal scope. THSS also reconstructs historical conflicts from transition lineage and keeps a fixed order from authorization through minimization. This exposes the basis of the disclosure without adding answer labels or a scoring shortcut.
THSS can also reduce what reaches the model. PII, PHI, and task-irrelevant health attributes may be removed before disclosure while the persisted snapshot records what was actually shown. The benchmark copies further replace synthetic FHIR subject references with an opaque study URN and record the redaction in artifact provenance. These safeguards do not amount to legal “minimum necessary” compliance, anonymization, sender authentication, or a general privacy guarantee. Their value here is that the disclosure decision remains auditable.
The concurrency experiment gives the design a less comfortable result. A profile-global version counter is simple to reason about, but it also treats unrelated writes as competitors. In our grid, the false-stale rate for unrelated slots reached 0.9375 at 16 writers. Finer-grained versioning could reduce that contention, although it would need its own safety evaluation because lost updates and cross-dependency errors can reappear when the version boundary is split.

4.4. Limitations

Several limits remain. The prospective model cohort contains only 64 synthetic subjects, and frequent ties left 9–21 informative pairs in the significant comparisons rather than the planned 51. The exact paired tests remain valid for those discordances, but effect estimates are less precise and the cohort is underpowered relative to plan. Strict THSS was not significantly better than BTSA for Claude; Gemini tied BTSA, full history, and long context. The current gateway also retains a base-version-only proposal path and does not yet enforce provenance-sensitive downgrade prevention for every Strict-THSS-derived model proposal after review. State-version commit uses a profile-row lock and rereads consent in the same service transaction, but atomicity against concurrent authorization, consent, or policy writers has not been demonstrated with race tests. No single baseline was verified to match every governance semantic while differing only in exact snapshot binding. Service/API tests do not include a deployed-boundary attack matrix across HTTP, cache, and retrieval; historical guarantees are strongest for conflicts with complete transition lineage; the performance run used one in-process worker; and there is no independent human adjudication or real-world clinical validation.

5. Conclusions

GLHS treats persistent longitudinal health AI as a governed read-write process. THSS records the state and governance context shown to the AI. A snapshot-bound proposal carries those source coordinates forward, and GST checks that they are still valid before persistence. Bitemporality, provenance, consent, conflict preservation, and optimistic concurrency serve as components of this process rather than as separate novelty claims.
The implementation rejected the tested context mismatches, the clause ablation showed where invalid proposals failed, and the PostgreSQL experiments prevented stale commits in the tested state-version races. In the prospective two-model cohort, Strict THSS reached 98.44% all-axes exact with Claude and 100% with Gemini; six prespecified comparisons remained significant after Holm correction, although only 9–21 discordant subjects informed those significant tests and the planned power target was not reached. Under the tested conditions, the evidence supports a traceable software path from governed disclosure to persistent write. Clinical benefit, downgrade-resistant policy for every model-derived path, governance-race atomicity, and independent validation of institutional policy require separate work.

Funding

The author declares that no funds, grants, or other support were received during the preparation of this manuscript.

Institutional Review Board Statement

Not applicable. The experiments reported in this manuscript were software evaluations conducted on synthetic data and did not involve human participants, human biological material, or identifiable patient data.

Data Availability Statement

The implementation and reproducibility materials are maintained in the public project repository at https://github.com/Project-CLARA-HBT/CLARA-Care. The public release should be accompanied by the exact implementation identifier, checksum inventory, and frozen evaluation artifacts so that the reported run can be independently audited.

Acknowledgments

Generative AI was used only to help restructure and edit author-supplied text. It did not generate the reported data or run the analyses. The author checked the scientific content and remains responsible for the manuscript.

Conflicts of Interest

The author has no relevant financial or non-financial interests to disclose.

References

  1. Gabrieli, E.R. Longitudinal patient record (LPR). J. Clin. Comput. 1992, 21, 1–16. [Google Scholar]
  2. Kouramajian, V.; Fowler, J. Modeling past, current, and future time in medical databases. In Proceedings of the Annual Symposium on Computer Applications in Medical Care, Washington, DC, USA, 5–9 November 1994; pp. 315–319. [Google Scholar]
  3. Kumar, A.; Tsotras, V.J.; Faloutsos, C. Designing access methods for bitemporal databases. IEEE Trans. Knowl. Data Eng. 1998, 10(1), 1–20. [Google Scholar] [CrossRef]
  4. HL7 International. FHIR R4: Provenance. 2019. Available online: https://hl7.org/fhir/R4/provenance.html (accessed on 12 August 2026).
  5. HL7 International. FHIR R4: Consent. 2019. Available online: https://hl7.org/fhir/R4/consent.html (accessed on 12 August 2026).
  6. HL7 International. FHIR R4: RESTful API/HTTP. 2019. Available online: https://hl7.org/fhir/R4/http.html (accessed on 12 August 2026).
  7. openEHR Foundation. openEHR Reference Model: EHR Information Model. 2026. Available online: https://specifications.openehr.org/ (accessed on 12 August 2026).
  8. World Wide Web Consortium. PROV-O: The PROV Ontology. W3C Recommendation. 2013. Available online: https://www.w3.org/TR/prov-o/.
  9. Zhao, J.; Zhi, X.; Yu, X. Beyond retrieval: Bi-temporal state arbitration for longitudinal healthcare agents. In Proceedings of the 4th Workshop on Towards Knowledgeable Foundation Models (KnowFM 2026); Association for Computational Linguistics: Stroudsburg, PA, USA, 2026; pp. 129–137. Available online: https://aclanthology.org/2026.knowfm-1.10/.
  10. Qu, Z.; Farber, M. Vital Trace: Protocol-Constrained Patient-State Reasoning for Longitudinal Clinical Trajectories. arXiv 2026, arXiv:2602.12833. [Google Scholar]
  11. Pugh, S.L.; Yang, E.; Sutherland, A.M.; Breschi, A. Detecting clinical discrepancies in health coaching agents: A dual-stream memory and reconciliation architecture. arXiv 2026, arXiv:2604.27045. [Google Scholar]
  12. Kiiskinen, T.; Fries, J.; Adamson, P.; et al. VISTA Architect: A graph database-oriented health AI system demonstrated in multidisciplinary tumor boards. arXiv 2026, arXiv:2606.22692. [Google Scholar]
  13. Li, X.; Wang, Y.; Lu, H.; Chen, Z.; Li, M.; Song, P.; Cai, T. MemTX: Transactional Belief Commit for Stateful Agent Memory. arXiv 2026, arXiv:2607.23929. [Google Scholar]
  14. Cui, H.; Tang, Z.; Yao, Z.; Meng, F.; Ma, Q.; Jia, W. MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory. arXiv 2026, arXiv:2607.27834. [Google Scholar]
  15. Margalit, Y.; Cohen-Inger, N.; Avram, E.; Taig, R.; Margalit, O. Governed Shared Memory for Multi-Agent LLM Systems. arXiv 2026, arXiv:2606.24535. [Google Scholar]
  16. Li, H.; Deng, J.; Jin, T.; et al. A Self-Evolving Agent for Longitudinal Personal Health Management. arXiv 2026, arXiv:2607.13940. [Google Scholar]
  17. Lewis, P.; Perez, E.; Piktus, A.; et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. Adv. Neural Inf. Process. Syst. 2020, 33, 9459–9474. [Google Scholar]
  18. Doyle, J. A truth maintenance system. Artif. Intell. 1979, 12(3), 231–272. [Google Scholar] [CrossRef]
  19. Kraljevic, Z.; Bean, D.; Shek, A.; et al. Foresight—a generative pretrained transformer for modelling of patient timelines using electronic health records: A retrospective modelling study. Lancet Digit. Health 2024, 6(4), e281–e290. [Google Scholar] [CrossRef]
  20. Wu, K.; Nagori, A.; Kamaleswaran, R. Planner-Auditor Twin: Agentic discharge planning with FHIR-based LLM planning, guideline recall, optional caching and self-improvement. arXiv 2026, arXiv:2601.21113. [Google Scholar]
Figure 1. Core GLHS disclosure-to-commit contract. THSS determines the context that can be shown to the AI. The returned proposal carries the source snapshot coordinates, and GST either writes it against the still-valid state and governance context or rejects it for reconciliation.
Figure 1. Core GLHS disclosure-to-commit contract. THSS determines the context that can be shown to the AI. The returned proposal carries the source snapshot coordinates, and GST either writes it against the still-valid state and governance context or rejects it for reconciliation.
Preprints 234813 g001
Table 1. Related mechanisms and scope. Each entry summarizes the emphasis of the cited specification or study in relation to the question examined here.
Table 1. Related mechanisms and scope. Each entry summarizes the emphasis of the cited specification or study in relation to the question examined here.
Work/Standard Relevant Mechanism Relation to This Study
FHIR + openEHR [4,5,6,7] Provenance/consent representation, record versioning, history, and version-aware update semantics. The cited specifications provide much of the infrastructure used here; they do not evaluate the AI disclosure-to-proposal contract studied in this paper.
Zhao et al. [9] Bitemporal, conflict-preserving longitudinal state arbitration. The cited work is closest on temporal state arbitration; its evaluation is not organized around a purpose/consent-bounded disclosure carried into a later write.
MemTX [13] Transactional belief commit with evidence, permissions, provenance, validity, and snapshot isolation. The cited work is the closest transactional analogue, but evaluates general agent memory rather than the health-specific read-to-write binding studied here.
MemTxn [14] Source-supported transactional updates, temporal conflict handling, and recovery. The cited work studies update boundaries and recovery; task/purpose/consent-scoped health disclosure is not its evaluation focus.
Governed Shared Memory/MemClaw [15] Scoped retrieval, temporal supersession, provenance, and policy-governed memory propagation. The cited work evaluates governed memory propagation across agents; this study follows one disclosed health-state snapshot into a later proposal.
HealthClaw [16] Health-specific private longitudinal memory and privacy-aware context reduction. The cited work is health-specific and privacy-oriented, but centers on memory selection and assistance rather than write admission after a particular disclosure.
GLHS THSS disclosure record, snapshot-bound proposal, and GST revalidation. Evaluated here as a software governance contract; clinical effectiveness is outside the claim.
Table 2. Descriptive PostgreSQL service-layer benchmark; one worker, history depth 50, 30 repetitions per operation. p99 is omitted because 30 repetitions do not support a stable tail estimate. The rate column is sequential single-worker operations per second, not concurrent throughput. Values include transaction finalization but exclude HTTP and model inference.
Table 2. Descriptive PostgreSQL service-layer benchmark; one worker, history depth 50, 30 repetitions per operation. p99 is omitted because 30 repetitions do not support a stable tail estimate. The rate column is sequential single-worker operations per second, not concurrent throughput. Values include transaction finalization but exclude HTTP and model inference.
Operation p50 (ms) p95 (ms) Sequential Rate (ops/s)
GST transition 30.967 54.685 29.952
State reconstruction 7.298 12.299 129.688
THSS compilation 61.806 68.889 16.937
Governed-decision reconstruction 10.537 17.011 95.350
Audit lookup 2.122 5.501 399.149
Invalidate + rebuild 52.996 74.336 18.295
Enter-in-error + rebuild 69.314 87.309 13.818
Table 3. All ten preregistered subject-level contrasts for Strict THSS. Positive effects favor Strict THSS; Holm adjustment spans the complete ten-test family.
Table 3. All ten preregistered subject-level contrasts for Strict THSS. Positive effects favor Strict THSS; Holm adjustment spans the complete ten-test family.
Model Comparator Effect (pp) 95% Bootstrap CI Holm-Adjusted Result Discordance
Claude Naive RAG +32.81 +21.88 to +43.75 p = 0.00000954 21
Claude LWW +25.00 +15.63 to +35.94 p = 0.00027466 16
Claude Full history +18.75 +7.81 to +29.69 p = 0.01098633 14
Claude Long context +14.06 +6.25 to +23.44 p = 0.01953125 9
Claude BTSA +4.69 not reported p = 1.0 ; n.s. not reported
Gemini Naive RAG +25.00 +14.06 to +35.94 p = 0.00027466 16
Gemini LWW +25.00 +15.63 to +37.50 p = 0.00027466 16
Gemini BTSA 0.00 exact tie n.s. exact tie
Gemini Full history 0.00 exact tie n.s. exact tie
Gemini Long context 0.00 exact tie n.s. exact tie
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.