Preprint
Article

This version is not peer-reviewed.

The Reinforcement Learning Contour: From Document Production to Governed Epistemic Transformation in AI-Assisted Scientific Inquiry

Submitted:

05 August 2026

Posted:

10 August 2026

You are already at the latest version

Abstract
Generative artificial intelligence has reduced the technical cost of producing scientific text, but increased output does not by itself produce cumulative knowledge, conceptual novelty, or epistemic progress. This article develops the Reinforcement Learning Contour (RLC) as a conceptual and architectural framework for converting AI-assisted scientific work from repeated document generation into governed recursive learning. The framework distinguishes an epistemic contour, Cₜ, from the governed transformation operator, Φₜ, through which a research episode may produce a successor contour. The contour records the time-indexed configuration of knowledge, ontology, relations, retrieval, evaluative criteria, inquiry policy, provenance, memory, tools, and governance constraints. The operator organizes evidence retrieval, source verification, adversarial evaluation, human approval, and reintegration. Changes in the answer-producing function within a contour are analytically separated from changes to the transformation operator itself; the latter are treated as Learning III-like events that cannot be autonomously committed and require an explicit human meta-decision. The central proposition is a differentiation principle: the value of a research cycle is proportional not to the quantity of text it produces but to the beneficial, traceable differentiation it creates between the preceding contour and the contours governing subsequent perception, interpretation, decision, and action. The framework is positioned relative to Popperian criticism, Lakatosian research programmes, Bateson’s orders of learning, double-loop learning, socially situated objectivity, provenance, temporal knowledge representation, and agentic science. It specifies a role-based architecture in which generation authority is strictly weaker than human acceptance authority; a twelve-step governed research cycle; a Structured Epistemic Change Record that documents change without collapsing quality into a scalar reward; novelty and anti-recursion controls; failure modes; seven falsifiable hypotheses; and a comparative pilot design. The proposal is explicitly unvalidated. Its claim is not that automation guarantees scientific progress, but that claims of recursive learning can be made more inspectable, contestable, and governable.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

Large language models can retrieve and summarize literature, reorganize arguments, generate hypotheses, propose experimental designs, write code, and draft scientific manuscripts. Agentic systems extend these capabilities toward increasingly complete research workflows. The resulting increase in production capacity is real. The inference that science therefore learns faster is not. A system can generate fluent documents while preserving the same assumptions, conceptual vocabulary, blind spots, and citation errors. Apparent breadth can substitute for understanding, and recursive dependence on generated outputs can narrow rather than expand the represented space of inquiry [22,23].
This article addresses a narrower and more demanding question: under what conditions does an AI-assisted research episode change the epistemic mechanism from which subsequent inquiry begins? The relevant object is not the manuscript alone. A document is an artifact produced by a research configuration. The deeper unit of analysis is the change, if any, in the configuration that determines what evidence can be retrieved, how claims are related, which contradictions remain visible, what questions are considered worth asking, and who is authorized to accept a proposed update.
The Reinforcement Learning Contour (RLC) is proposed as a conceptual and architectural response. The term does not denote a conventional reinforcement-learning algorithm. The framework does not assume a closed Markov state space, a scalar reward, a value function, or automated policy optimization. Instead, it uses the recursive logic of evaluated action to describe a governed human-machine environment in which accepted outcomes alter the conditions and policies governing future inquiry. The word contour names the bounded but evolving configuration within which knowledge, memory, ontology, tools, evaluation, and authority interact.
The revised formulation makes two analytical distinctions explicit. First, an epistemic contour Cₜ is a time-indexed state representation, whereas the RLC is the complete governed architecture that connects such states. Second, the transition is produced through a governed transformation operator Φₜ. An update to the answer-producing function inside a contour is not equivalent to an update to Φₜ itself. The latter is a higher-order event and cannot be treated as an automatically authorized consequence of the operator’s own output.
The article makes six contributions. It (1) defines the epistemic contour as a bounded, inspectable state representation; (2) identifies the governed transformation operator that produces candidate transitions; (3) states a differentiation principle for the value of research experience; (4) specifies a role-based architecture in which provenance, temporality, and human acceptance authority are structural; (5) provides a categorical change record and anti-recursion controls that make claims of learning auditable; and (6) exposes the framework to falsification through seven hypotheses and a comparative pilot design. The proposal remains conceptual and untested. Its purpose is to make recursive epistemic claims governable, not to declare them validated.

2. The Problem: Output Growth Without Epistemic Progress

The technical unit optimized by many AI-assisted workflows is the artifact: a report, literature review, analysis notebook, or manuscript. This creates a category error. Artifacts can multiply while the mechanism that produced them remains unchanged. New titles, rearranged sections, and altered vocabulary may create textual variation without changing central claims, evidence quality, uncertainty, ontology, or the policy by which future questions are selected.
Four recurrent failure conditions motivate the RLC. First, textual novelty may be mistaken for epistemic novelty. Second, a closed internal corpus may become progressively more coherent and less corrigible. Third, unrestricted external search may replace internal confirmation bias with unverified novelty. Fourth, the document may be treated as the terminal product of research rather than as an epistemic event that should modify future inquiry.
These conditions interact with known weaknesses of generative systems. Unsupported assertions may remain fluent; bibliographic references may be fabricated or mismatched to claims [20,21]; repeated ingestion of generated material can amplify representational degradation [22]; and the scale and fluency of models can produce an illusion of understanding [23]. In a recursive system, a local error is not merely repeated. Once accepted into memory, ontology, retrieval, or policy, it can alter the starting conditions of later cycles.
A meaningful recursive research cycle therefore requires more than feedback. It requires a recorded starting state, an explicit proposal for change, independent source verification, adversarial evaluation, a decision by an identified acceptance authority, governed reintegration, and a comparison between the preceding and successor states. Without these elements, recursion may be computationally sophisticated while remaining epistemically opaque.

3. Philosophical and Technical Foundations

3.1. Popper: Criticism and Exposure to Refutation

Popper’s account of scientific knowledge emphasizes conjecture, criticism, and exposure to possible refutation [1,2]. The RLC operationalizes this not as a universal score but as an adversarial obligation. A research cycle must seek disconfirming evidence, state conditions under which a central claim would fail, and distinguish an explanatory proposal from an empirically validated conclusion. A system that only retrieves support has not completed the critical stage of the contour.

3.2. Lakatos: Progressive and Degenerating Research Programmes

Lakatos provides the closest philosophical predecessor to the idea of beneficial contour differentiation. A research programme is progressive when theoretical change produces novel empirical content and degenerating when adjustments mainly protect an established core after difficulties arise [3]. The RLC does not claim to rediscover that distinction. Its contribution is operational: it represents a programme as a sequence of inspectable state transitions with claim-level lineage, temporal provenance, contradiction records, ontology changes, and explicit acceptance decisions.
Lakatos asks when a programme is progressing. The RLC asks how a human-machine research environment can record, govern, and compare the transitions through which progress or degeneration may occur. The former is primarily an appraisal of programmes over time; the latter attempts to make the relevant changes visible at the time they are proposed and committed.

3.3. Kuhn, Laudan, and Socially Situated Objectivity

Kuhn’s paradigms, Laudan’s problem-solving model, and Longino’s account of socially mediated objectivity caution against reducing scientific change to a single metric [4,5,6]. Progress may involve empirical novelty, conceptual reclassification, improved problem solving, exposure of background assumptions, or stronger critical interaction. For this reason, the RLC treats epistemic transformation as multidimensional and contestable rather than as a scalar reward automatically emitted by the system.
Longino is especially important for the framework’s limitation. Objectivity is not merely an internal property of a well-documented individual reasoner; it is produced through critical interaction among differently situated inquirers [6]. A single-researcher recursive configuration lacks that plurality by construction. Internal role separation can reduce conflicts, but it cannot manufacture independent social perspectives.

3.4. Orders of Learning: Bateson, Argyris, and Schön

Bateson distinguished Learning I, in which errors of choice are corrected within an existing set of alternatives, from Learning II, in which the process of Learning I changes through revision of the alternatives or of how experience is punctuated into contexts. He described Learning III as change in the process of Learning II. He also emphasized that premises acquired through Learning II may become self-validating and that their higher-order transformation is difficult and rare [7]. The RLC does not claim this hierarchy of learning as novel. Bateson provides a direct conceptual predecessor for separating correction inside an epistemic configuration from change in the conditions through which future correction occurs.
Argyris and Schön later expressed a related distinction in an organizational register. Single-loop learning corrects action within existing governing variables; double-loop learning revises the governing variables themselves [8]. The RLC’s difference between changing an answer and changing the function that produces future answers belongs to this lineage. Its proposed contribution is not the hierarchy itself, but the instrumentation required to make a claim of higher-order change auditable: state snapshots, provenance, difference records, authority separation, and explicit reintegration.

3.5. Provenance, Reproducibility, and Temporal Knowledge Representation

Modern scientific workflows depend on traceable data, methods, software, claims, and transformations. W3C PROV formalizes the entities, activities, and agents involved in producing information [9]. FAIR principles emphasize findability, accessibility, interoperability, and reuse [10]. Reproducible research requires computational results to remain connected to executable procedures and inspectable evidence [11,12,13], while foundational work on data provenance distinguishes why and where lineage and surveys its representation in e-science [14,15].
The RLC extends this concern from reproducibility of an output to provenance of epistemic change. A workflow may perfectly reproduce a document and still fail to record why its research programme now asks different questions, rejects a former claim, or uses a different ontology. Because belief revision is the phenomenon of interest, the representation must be temporal. Static knowledge graphs can encode support relations [16], but dynamic and temporal graph models are needed to represent qualification, contradiction, and supersession [17,18].
Retrieval-augmented generation can condition output on retrieved material [19], but retrieval alone does not establish that a source supports a claim, nor does it prevent fabricated references and unsupported synthesis [20,21]. In a recursive environment, unverified output that is written back into memory becomes structural corruption. Source verification and claim-to-source provenance are therefore architectural requirements rather than editorial preferences.

3.6. Agentic Science and the Remaining Gap

Autonomous and semi-autonomous research systems demonstrate increasingly complete cycles of planning, execution, analysis, and manuscript production [25,26,27]. In parallel, scholarship on large models in science warns that scale, fluency, and apparent breadth can be mistaken for understanding and that scientific use requires explicit epistemic governance [23,28,29].
Neither line by itself answers the question addressed here. The issue is not whether an agent can produce a scientific artifact, but whether a sequence of artifacts constitutes a progressive, traceable, and corrigible programme. The RLC is positioned as a complementary appraisal and governance layer for human-AI research systems, including systems that are otherwise highly autonomous.

4. Formal Definition of the Reinforcement Learning Contour

4.1. The Epistemic Contour

An epistemic contour is the time-indexed state Cₜ of a bounded human-machine inquiry configuration. It includes the knowledge available to the system, the ontology through which that knowledge is organized, the relations among claims and evidence, retrieval mechanisms, inquiry policies, evaluative criteria, memory and provenance, available tools, and human or institutional governance constraints.
Cₜ = ⟨Kₜ, Oₜ, Gₜ, Rₜ, Pₜ, Eₜ, Mₜ, Hₜ⟩
K denotes available knowledge; O the active ontology; G the graph of entities, claims, evidence, and relations; R retrieval mechanisms; P inquiry and action policies; E evaluative criteria; M memory and provenance; and H the recorded state of human and institutional governance. The contour determines how evidence is perceived, which questions can be formulated, which actions are available, and how outcomes are evaluated at a given time.
The contour is bounded analytically rather than assumed to be closed. Its boundary identifies the knowledge, tools, rules, and authorities that are active for a specified research programme at time t. External sources can enter only through a governed discovery and verification process. This boundedness is necessary for state comparison: without a defined scope, any claim that the system changed remains impressionistic.

4.2. The Governed Transformation Operator

The Reinforcement Learning Contour is the complete governed recursive architecture that links epistemic states. A research episode Xₜ may produce a candidate successor contour through a time-indexed transformation operator Φₜ:
C*ₜ₊₁ = Φₜ(Cₜ, Xₜ)
The asterisk marks a proposed rather than committed state. Φₜ comprises the governed research cycle: question selection, internal retrieval, source-grounded synthesis, external and adversarial discovery, primary-source verification, change proposal, drafting, evaluation, approval routing, reintegration, and state comparison. The operator describes how a successor contour is produced; it does not possess authority to accept its own output.
Commit requires an explicit approval event by the Human Acceptance Authority, denoted A_H:
A_H : approve(C*ₜ₊₁) ⟹ Cₜ₊₁
The recorded governance state Hₜ is a component of the contour. The acceptance authority A_H is the accountable subject or institutional role that can authorize a transition. This distinction prevents the stored description of governance from being confused with the actor who exercises authority.

4.3. Orders of Change: Answers, Contours, and Operators

Three analytically distinct orders of change must be kept separate. First, an individual answer or action may change while the function that produces answers remains stable. Second, the answer-producing function may change as part of a broader transition in the epistemic contour. Third, the governed operator that produces contour transitions may itself be revised.
Table 1. Analytical orders of change in the RLC.
Table 1. Analytical orders of change in the RLC.
Order Object changed Notation Interpretation Authority
Output correction A specific answer, judgment, or action Aₜ → Aₜ₊₁ Correction inside the currently active answer-producing function Execution under the already approved policy; no contour commit
Contour transformation Knowledge, ontology, relations, retrieval, policy, evaluation, memory, or the answer-producing function fₜ(·) → fₜ₊₁(·) as part of Cₜ → Cₜ₊₁ Learning II-like change within the bounded epistemic configuration Commit by A_H
Operator transformation The rules by which candidate contour transitions are produced and governed Φₜ → Φₜ₊₁ Distinct higher-order, Learning III-like event Explicit human meta-decision only
Strong learning as used in this article denotes change in the answer-producing function fₜ within the epistemic contour and is represented as part of the transition Cₜ → Cₜ₊₁. Change in Φₜ modifies the governed operator through which such transitions are produced. It is therefore a distinct higher-order event and is not covered by the definition of strong learning used in Section 5.
An operator change is admissible only through an explicit human meta-decision:
A_H : approve(Φₜ → Φₜ₊₁)
The operator may expose evidence that its rules require revision, but it may not treat its own output as an authorized change to those rules. This is a design boundary rather than a missing capability. The RLC does not automate Learning III-like transformation; it reserves modification of the learning rules to accountable human authority.
Figure 1. Separation of contour transition, human commit authority, and operator revision. The operator may produce a candidate successor contour but cannot commit it or autonomously rewrite itself.
Figure 1. Separation of contour transition, human commit authority, and operator revision. The operator may produce a candidate successor contour but cannot commit it or autonomously rewrite itself.
Preprints 226936 g001

4.4. Relation to Technical Reinforcement Learning

The name Reinforcement Learning Contour is retained deliberately, but the relationship to technical reinforcement learning is structural and analogical rather than formally algorithmic. Conventional reinforcement learning presupposes a state space, action space, reward signal, and a policy updated to maximize expected return [30]. The RLC has an open and partly unformalized state space; evaluative criteria are multidimensional, contested, and often delayed; and updates include human judgments that cannot be reduced to optimization of a scalar objective.
The word reinforcement marks the recursive principle that evaluated outcomes modify the conditions and policies governing future action. The word contour marks an evolving boundary and internal organization rather than a single loop. Technical RL concepts apply directly mainly as warnings. In particular, any attempt to optimize a novelty or transformation score creates a reward-hacking surface [24]. The framework therefore refuses a single objective function for epistemic quality.

5. The Law of Experiential Value

The central proposition of the framework is the following:
The greater the beneficial differentiation between a preceding epistemic contour and the contours governing future perception, decision, and action, the greater the value of the lived or computationally mediated experience.
V(Xₜ) ∝ D⁺(Cₜ, Cₜ₊₁)
D⁺ denotes beneficial differentiation. It is not a validated numerical metric. The qualifier beneficial is indispensable because any episode can alter a system while producing confusion, false confidence, epistemic closure, degraded calibration, or stronger attachment to unsupported claims. Change alone is not learning. The judgment must be supported by evidence, provenance, and an identifiable approval record.
The proposition relocates value from output quantity to mechanism change. A cycle has low RLC value when it merely produces another formulation of an existing answer. It has greater potential value when it changes what the system can subsequently perceive, retrieve, question, test, reject, or understand. This yields a distinction between weak and strong learning:
Weak learning: Aₜ → Aₜ₊₁ Strong learning: fₜ(·) → fₜ₊₁(·)
The architectural claim attached to strong learning is narrow and testable. A defensible claim requires, at minimum, a snapshot of the preceding contour, a provenance-bearing record of the difference, and separation between generation and acceptance. These conditions do not prove that the change is beneficial; they make the claim inspectable.

6. Architecture by Functional Roles

The architecture is defined by roles rather than by short-lived products. Each role has a function and an authority boundary. The authority column is load-bearing: it specifies what a component may commit, not merely what it may compute or propose.
Figure 2. The Reinforcement Learning Contour as a governed recursive architecture. Generative roles may propose changes; only the human governance role may commit them. The return path represents the transition from Cₜ to Cₜ₊₁, not merely the production of a document.
Figure 2. The Reinforcement Learning Contour as a governed recursive architecture. Generative roles may propose changes; only the human governance role may commit them. The return path represents the transition from Cₜ to Cₜ₊₁, not merely the production of a document.
Preprints 226936 g002
Table 2. Functional roles and authority allocation.
Table 2. Functional roles and authority allocation.
Architectural role Function Authority
Prior Research Corpus Publications, manuscripts, notes, reviewer feedback, rejected hypotheses, methods, and unresolved questions. Epistemic memory; no acceptance authority
Locally Governed Knowledge Store Human-readable and editable representation of concepts, claims, evidence, decisions, and research questions. Working representation
Temporal Provenance and Relation Layer Versioned entities and relations, validity intervals, source lineage, contradiction, qualification, and supersession records. System of record for accepted change
Source-Grounded Synthesis Layer Analysis restricted to approved evidence, cross-document comparison, claim extraction, and structured synthesis. May propose; may not commit
External Discovery and Adversarial Retrieval Layer Search for new literature, competing theories, negative evidence, replications, criticism, and adjacent domains. May propose; may not commit
Research Orchestration and Drafting Layer Workflow execution, retrieval coordination, validation scripts, structured drafting, and change reports. May propose; may not commit
Human Epistemic Governance Approval of questions, evidence, ontology changes, interpretations, operator changes, authorship, and publication. Sole acceptance authority; veto
Generation authority is strictly weaker than acceptance authority. Synthesis, discovery, and orchestration layers may propose changes to claims, relations, ontology, uncertainty, policy, or governance, but none may commit them. A configuration in which the drafting layer writes accepted claims directly into the provenance layer without an approval event has abandoned the architecture, regardless of the sophistication of its tooling.

6.1. Requirements by Role

The Prior Research Corpus must retain version, status, domain, extracted claims, evidence type, confidence, citations, known contradictions, and relation to later work. Rejected and falsified material must remain available; a corpus containing only successes cannot support calibration.
The Locally Governed Knowledge Store requires a typed vocabulary that includes at least Concept, Claim, Hypothesis, Evidence, Method, Dataset, Contradiction, Limitation, Research Question, Publication, External Source, Decision, and Contour Update. Links must be typed and directional, and the representation must remain editable by the human without automation.
The Temporal Provenance and Relation Layer must represent not only what is believed, but when it came to be believed, on which source, what later changed, and whether an earlier claim remains active. Supersession is a relation rather than a deletion. A minimal claim record is ⟨content, status, confidence, source, time, verification⟩.
The Source-Grounded Synthesis Layer derives value from controllable scope, not authority. It may analyze the approved corpus but must not silently expand it. The External Discovery and Adversarial Retrieval Layer has the opposite obligation: it must introduce challenge rather than merely more support. A pass returning only confirmatory material should be treated as failed and rerun under adversarial framing.
The Research Orchestration and Drafting Layer must produce a machine-readable change record for each cycle and may not write accepted claims to the system of record without an approval event. Human Epistemic Governance may reject a source, ontology update, causal inference, manuscript, direction of inquiry, or proposed change to Φₜ. No automated output overrides rejection.

6.2. One Replaceable Local Instantiation

The author’s current local implementation uses the following products as one replaceable instantiation of the functional architecture. These products are not part of the RLC definition. Any substitution that preserves provenance, role separation, and acceptance authority is architecturally equivalent; any configuration that collapses synthesis, discovery, orchestration, and commitment into a single unaudited agent is not.
Table 3. One local implementation of the functional roles.
Table 3. One local implementation of the functional roles.
Role Current implementation
Locally governed knowledge store Obsidian [33]
Temporal provenance and relation layer Graphiti [34]
Source-grounded synthesis layer NotebookLM [35]
External discovery and adversarial retrieval Perplexity [36]
Research orchestration and drafting Codex [37]
Final epistemic authority Human researcher

7. The Governed Research Cycle

A governed cycle consists of twelve steps. The sequence is not a claim that inquiry is mechanically linear; it is an accountability structure that identifies what must be recorded before a proposed transition can be accepted.
1. Contour snapshot. Record the current knowledge, ontology, relation graph, retrieval mechanisms, inquiry policies, evaluative criteria, provenance state, tools, and governance constraints.
2. Question selection. Select questions from unresolved contradictions, weak conceptual links, failed hypotheses, external developments, or explicit human priorities.
3. Internal retrieval. Retrieve relevant prior publications, claims, evidence, methods, decisions, and limitations from the governed corpus.
4. Source-grounded synthesis. Analyze the approved internal corpus without silently expanding beyond it.
5. External and adversarial discovery. Search for current evidence, competing explanations, criticism, negative results, replication failures, and adjacent concepts.
6. Primary-source verification. Verify material against original sources; search summaries remain discovery objects rather than evidence objects.
7. Change proposal. Specify what would change in the contour: claim, definition, relation, evidence weight, uncertainty, policy, ontology, or governance rule, and on what basis.
8. Drafting. Produce a manuscript or another scientific artifact while separating evidence, interpretation, hypothesis, and speculation.
9. Adversarial evaluation. Attempt to refute the central claim, detect prior art, identify overgeneralization, and test citation integrity.
10. Human decision. Approve, revise, defer, or reject the proposed artifact and contour changes. Rejection is a valid and often correct outcome.
11. Governed reintegration. Write accepted claims, evidence, contradictions, limitations, and questions back into the governed stores with the approval record attached.
12. State comparison. Document how Cₜ₊₁ differs from Cₜ and how subsequent inquiry should change as a result.
Steps 1 and 12 are the load-bearing pair. Without a recorded starting state and an explicit comparison with the successor state, a workflow may produce useful writing but cannot support an auditable claim of contour differentiation.

8. The Structured Epistemic Change Record

A composite differentiation score would create false precision unless its dimensions, coders, scales, and reliability were empirically validated. It would also create an optimization target: any dimension rewarding novelty or transformation could be satisfied by extreme or weakly supported claims [24]. The framework therefore uses a structured categorical record. Each dimension is documented separately; no total score is calculated.
Table 4. Structured Epistemic Change Record.
Table 4. Structured Epistemic Change Record.
Dimension Governance question Permitted classification
Claim change Which central claim changed? none / refined / rejected / introduced
Evidence change Did the quality or relevance of evidence change? lower / unchanged / higher / contested
Contradiction Was a serious counterargument discovered? none / unresolved / partially resolved / resolved
Ontology Did concepts or relations change? none / local / structural
Uncertainty Was confidence recalibrated? increased / unchanged / reduced with justification
Inquiry policy Will future questions or methods be selected differently? no / possible / demonstrable
Provenance Can the change be traced to evidence and decisions? absent / partial / complete
External review Was the claimed change independently examined? no / internal only / external

8.1. The Uncertainty Dimension

Justification is required for reduced uncertainty, not for increased uncertainty. Contact with disconfirming evidence should often increase uncertainty; treating this as a defect would reward premature closure. A reduction in uncertainty is itself an epistemic claim and must be supported by evidence and reasoning. The asymmetry is intentional.

8.2. Conditions Under Which the Record Is Invalid

The record is categorical, programme-relative, and rater-dependent. Its classifications are not comparable across programmes, across raters without an agreement study, or across time if the anchors change. It documents a decision; it does not validate the decision. Provenance: absent and External review: no are statements about the limits of what has been established, not evidence that change occurred.
The record must not be aggregated into a productivity score, used to rank researchers, or presented as a validated quality metric. In evaluative studies, at least two independent raters should assess central dimensions and report inter-rater agreement. Most importantly, the record must not become an optimization target. Once a measure becomes a target, novelty, contradiction handling, and ontology change become reward-hacking surfaces.

9. Novelty and Anti-Recursion Controls

9.1. Textual Versus Epistemic Novelty

A new title, section order, metaphor, or vocabulary does not establish scientific novelty. The system must compare claims, mechanisms, ontology, evidence, and conclusions with the prior corpus. Fluency is not evidence of difference at the epistemic level.

9.2. Claim-Level Novelty Classification

Each central claim should be classified as previously established, reformulated, extended, contradicted, integrated across domains, newly formalized, or genuinely new. A manuscript in which no central claim exceeds reformulation should ordinarily not be submitted as a new contribution.

9.3. Prior-Art Search

Before a construct is presented as new, it must be searched under alternative names and in adjacent literatures. This includes disciplinary vocabularies that may not share the author’s terminology. Prior-art search is not a ceremonial bibliography step; it is an adversarial test of novelty and one of the most consequential controls before journal review. The present revision documents one such operation: application of this control identified Bateson’s orders of learning [7] as an omitted conceptual predecessor, which narrowed the novelty claim, produced a local clarification of the framework’s ontology, and increased recorded uncertainty. This episode is a documented execution trace of the control, not a pilot study or empirical validation.

9.4. Counterfactual Novelty Test

The system asks: if this work had not been produced, which future question, decision, concept, method, or evidence relationship would remain unavailable? If the answer is none, RLC value is low regardless of prose quality.

9.5. Mechanism-Change Test

The strongest criterion asks whether the episode changes the mechanism that will generate future research rather than merely adding another answer. Content that does not alter question selection, evidence weighting, ontology, retrieval, evaluation, or uncertainty represents weak differentiation by definition.

10. Failure Modes and Controls

10.1. Source Hierarchy

Admissibility runs in descending order: primary peer-reviewed evidence; official documentation and authoritative datasets; systematic reviews and meta-analyses; scholarly preprints; institutional reports; reputable secondary analysis; search summaries; and unverified generated content. The last two categories are never terminal evidence. They may function only as discovery objects.
Table 5. Principal failure modes and architectural controls.
Table 5. Principal failure modes and architectural controls.
Failure mode Risk Control
Automated paper inflation The workflow optimizes document count rather than epistemic value. Count governed cycles, including justified non-publication, rather than papers alone.
Conceptual echo chamber The system repeatedly confirms the author’s framework; recursive self-consumption narrows representation. Mandate adversarial retrieval and at least one serious competing explanation per cycle.
False novelty Existing ideas are renamed and presented as original constructs. Search alternative terminology, compare prior art, and classify novelty at claim level.
Citation laundering A summary or generated statement is treated as primary evidence; fabricated or mismatched references enter memory. Retrieve the original source, verify claim support, and preserve claim-to-source provenance.
Graph corruption Incorrect entity resolution creates false relations. Use typed schemas, confidence thresholds, provenance, and human review of high-impact edges.
Temporal inconsistency New findings contradict old relations without qualifying or superseding them. Store validity intervals and explicit qualification or supersession links.
Ontology proliferation New labels accumulate without consolidation. Apply synonym mapping, concept-merger review, and ontology governance.
Reward hacking Novelty incentives encourage extreme or weakly supported claims. Use no composite score; balance novelty with evidence, calibration, falsifiability, and relevance.
Automation bias Technical sophistication is mistaken for scientific reliability. Expose uncertainty, failed checks, rejected hypotheses, and unresolved disagreement.
Illusion of understanding Breadth substitutes for comprehension and the question space narrows unnoticed. Track question diversity and require explanation of conceptual dependencies.
Loss of intellectual authorship Human and machine contributions become indistinguishable. Record contribution provenance at claim, decision, and paragraph level.
Reflexive self-validation The system produces, detects, and evaluates the change it calls progress. Separate generation and evaluation; use preregistered rubrics, blinded comparison, and external audits.

10.2. Citation Verification Standard

Before acceptance, the verifier must establish that the source exists; that authors, title, year, and publication details are correct; that a DOI resolves where applicable; that the cited passage supports the claim as stated; that the work has not been retracted; and that the wording does not overstate the evidence. When a claim is tied to a page number, the edition actually consulted must be identified and the cited pagination must be verified in that edition; page references must not be transferred across editions or reprints without direct confirmation. Because generated bibliographies contain documented fabrication and error rates [21], this check cannot be delegated to the system that generated the citation.

10.3. Uncertainty Representation

The architecture distinguishes fact, supported inference, weak inference, hypothesis, speculation, and unresolved contradiction. Each status is recorded explicitly on the claim rather than inferred from rhetorical hedging. This prevents fluent prose from concealing epistemic status.

11. Reflexive Validation and the Human Authority Boundary

The RLC contains a structural circularity. The same human-machine environment may generate an epistemic change, define the categories through which the change is represented, detect the change, and judge whether it is beneficial. Human veto does not eliminate the problem because the human may also be the architect of the system, the author of the source corpus, the beneficiary of a novelty claim, and the final evaluator.
This creates risks of metric self-confirmation, endogenous novelty criteria, author-system collusion, and confirmation through ontology design. A system may appear to have learned because it has redefined progress in terms that favor its own outputs. Longino’s account explains why this is not merely a procedural inconvenience: objectivity depends on critical interaction among differently situated inquirers, and a single-researcher recursive system lacks that social structure [6].
The problem should therefore be managed rather than declared solved. Generation and evaluation should use different prompts, agents, or reviewers; central claims should be compared blindly with earlier work; source verification should be independent of drafting; rejected changes should remain visible; and external audits of ontology and provenance should occur periodically. For empirical studies, evaluation rubrics should be preregistered before a cycle begins.
The architecture intentionally reserves modification of Φₜ to explicit human meta-decisions and therefore does not automate Learning III-like change. This boundary prevents endogenous self-modification from being treated as authorized learning, but it does not make the human independent of the system. The same person may define the contour, design the operator, authorize its revision, and evaluate the result. Internal role separation can reduce, but cannot eliminate, this limitation.
A binding consequence follows: any favorable assessment of differentiation that lacks independent and blinded evaluation should be discounted, including assessments made by the present author. Documented execution traces may demonstrate that the architecture was used; they do not constitute a pilot study or empirical validation.

12. Falsifiable Hypotheses

  • H1. RLC-supported workflows will produce greater claim-level novelty than non-recursive LLM drafting when assessed by blinded reviewers.
  • H2. RLC-supported manuscripts will contain fewer unsupported claims and fewer source-to-claim mismatches.
  • H3. Temporal provenance representations will improve detection of superseded, contradicted, or qualified claims compared with static document retrieval.
  • H4. Mandatory Structured Epistemic Change Records will reduce semantic and claim-level repetition across successive manuscripts.
  • H5. Human-governed RLC workflows will show better confidence calibration than fully automated research-generation workflows, at a measurable cost in throughput.
  • H6. The value assigned to a research cycle will correlate more strongly with the quality of subsequent research questions than with manuscript length or generation speed.
  • H7. Contradictory external evidence will produce greater beneficial contour differentiation than confirmatory evidence when governance controls are active, and greater harmful differentiation when they are absent.
H7 is the framework’s sharpest exposure. If contradictory evidence produces no measurable advantage under governance, or if the architecture cannot distinguish beneficial from harmful differentiation, the central mechanism is not doing the work claimed for it.

13. Comparative Pilot Study Design

A pilot of approximately 30 to 90 research cycles can compare three conditions: conventional manual research; AI-assisted drafting without recursive state and provenance; and the full RLC architecture. The conditions should address comparable research questions and be assessed by reviewers blinded to condition.
Primary outcomes should include claim-level novelty, citation accuracy, unsupported-claim rate, contradiction-detection rate, useful ontology change, quality of subsequent research questions, reviewer-rated contribution, and traceability of epistemic change. Secondary outcomes may include researcher cognitive load, graph error rate, source diversity, question diversity across cycles, the number of rejected hypotheses, and the proportion of cycles that correctly end without publication.
At least two independent raters should assess novelty and contribution, with inter-rater agreement reported. Rubrics should be preregistered before the first cycle. Without an agreement study, scored comparisons between conditions should not be reported.
A high pre-publication rejection rate should not automatically count as failure. A governed system may create value by preventing weak, redundant, or unsupported manuscripts from entering the literature. This must be specified as a hypothesis-consistent outcome in advance rather than rationalized after the result is known.

14. Discussion

The RLC changes the unit of analysis from the scientific document to the transformation of the epistemic mechanism that produced it. This shift has practical and philosophical consequences. A cycle may be successful without producing a publishable manuscript if it rejects a false assumption, exposes a contradiction, narrows an unjustified claim, reorganizes an ontology, improves citation integrity, or generates a more productive future question.
Prior publications consequently function as evolving scientific memory rather than a static archive. Rejected hypotheses and failed approaches are retained because they are useful for calibration and future question selection. External search acquires an adversarial obligation: its purpose is not to enlarge the corpus indiscriminately but to challenge the contour. A discovery pass returning only support is incomplete.
The operator formalism clarifies what recursive scientific learning can and cannot mean. A change in a particular answer is not automatically a change in the contour. A change in the contour is not automatically a change in the rules that produce contour transitions. Treating these orders as equivalent would allow ordinary revision to be redescribed as meta-learning. The distinction among Aₜ, fₜ, Cₜ, and Φₜ is therefore not mathematical decoration; it is a safeguard against inflated claims.
The human authority boundary creates a deliberate asymmetry. Automated components may discover, synthesize, challenge, and propose. They may not determine the epistemic status of their own outputs or autonomously rewrite the governance rules under which future outputs are accepted. This restriction reduces speed and may increase cognitive burden, but it preserves accountable authorship and creates a visible point at which responsibility cannot be delegated.
The framework also distinguishes automation of science from recursive scientific learning. Agentic systems may optimize literature review, coding, experimentation, analysis, and drafting. The RLC asks whether an episode changes the epistemic conditions governing the next episode and whether that change can be traced and contested. The system should therefore not be obligated to publish daily. It may complete a governed cycle whose correct result is a contradiction report, negative result, ontology revision, rejected hypothesis, or reasoned decision not to write a paper.
Institutional extension would require further work. Laboratories and organizations need distributed acceptance authority, role-based access control, authorship attribution, conflict resolution, audit policy, and procedures for disagreements among human authorities. These are not implementation details; they determine whether the architecture remains governed when it scales beyond a single researcher.

15. Limitations

1. The RLC is a conceptual construct that has not been empirically validated. The comparative pilot described in Section 13 has not been conducted.
2. Beneficial differentiation is multidimensional, contestable, and often visible only over time. The Structured Epistemic Change Record documents decisions; it does not validate them.
3. The architecture intentionally reserves operator modification to explicit human meta-decisions, but this design boundary does not eliminate reflexive validation. A single human may design, authorize, and evaluate the system, while the critical plurality required by Longino’s account is absent [6].
4. The term reinforcement learning may create technical expectations despite the explicit distinction from conventional reinforcement-learning algorithms.
5. The architectural components rely on probabilistic systems that cannot guarantee truth, novelty, comprehension, or citation accuracy [20,21,23].
6. Temporal graphs remain vulnerable to extraction errors, entity-resolution failures, misleading provenance, and schema decisions that encode the author’s assumptions.
7. A corpus centered on one researcher can preserve that researcher’s biases even under adversarial search because the search is framed from within the contour.
8. Daily or rapid iteration may privilege conceptual production over slow empirical validation, reintroducing the imbalance the framework is designed to oppose.
9. The proposed local implementation has not been compared experimentally with alternative workflows and should not be interpreted as an endorsed product stack.
10. Extension to laboratories or institutions requires formal mechanisms for authorship, access control, disagreement, conflict of interest, and distributed governance that are not specified here.

16. Conclusion

The Reinforcement Learning Contour is a governed epistemic feedback architecture for AI-assisted scientific inquiry. It does not describe a conventional reinforcement-learning algorithm. It describes the evolving configuration of knowledge, ontology, relations, retrieval, evaluation, policy, memory, tools, provenance, and human authority through which research experience shapes future inquiry.
Its central proposition is that the value of an experience is proportional not to the output it produces but to the beneficial and traceable differentiation it creates between the preceding contour and the contours governing subsequent perception, interpretation, decision, and action. The framework makes that proposition operational through state snapshots, a governed transformation operator, source verification, adversarial evaluation, human acceptance authority, structured change records, and explicit reintegration.
A recursive research system becomes scientifically meaningful only when accepted outcomes change not merely what the system says, but what it can subsequently perceive, question, test, reject, and understand—and when those changes remain open to inspection, criticism, and refusal.

Competing: interests

The author declares no competing financial interests arising from this work. The author develops governance-oriented AI systems commercially and discloses this as a potential intellectual conflict with respect to the framework advocated here.

Use: of generative AI

Generative AI systems were used for drafting, structural revision, and literature discovery. The author reviewed and accepted the final text and remains responsible for all conceptual claims, citations, interpretations, the differentiation principle, the architecture, the Structured Epistemic Change Record, and the hypotheses. Bibliographic claims were checked against original or authoritative sources in accordance with Section 10.2. No AI system is listed as an author, consistent with prevailing publisher and editorial policy [31,32].

Ethics: approval

Not applicable; the work involved no human participants or animal subjects.

Data: availability

Not applicable; no datasets were generated or analysed.

Author Contributions

Sole author: conceptualization, methodology, formalization, writing, revision, and final approval.

Funding

None.

References

  1. Popper, K. R. (1959). The Logic of Scientific Discovery. Hutchinson.
  2. Popper, K. R. (1963). Conjectures and Refutations: The Growth of Scientific Knowledge. Routledge.
  3. Lakatos, I. (1970). Falsification and the methodology of scientific research programmes. In I. Lakatos & A. Musgrave (Eds.), Criticism and the Growth of Knowledge (pp. 91-196). Cambridge University Press.
  4. Kuhn, T. S. (1962). The Structure of Scientific Revolutions. University of Chicago Press.
  5. Laudan, L. (1977). Progress and Its Problems: Toward a Theory of Scientific Growth. University of California Press.
  6. Longino, H. E. (1990). Science as Social Knowledge: Values and Objectivity in Scientific Inquiry. Princeton University Press.
  7. Bateson, G. (1987). Steps to an Ecology of Mind: Collected Essays in Anthropology, Psychiatry, Evolution, and Epistemology. Jason Aronson Inc. (Original work published 1972.).
  8. Argyris, C., & Schön, D. A. (1978). Organizational Learning: A Theory of Action Perspective. Addison-Wesley.
  9. Moreau, L., & Missier, P. (Eds.). (2013). PROV-DM: The PROV Data Model. W3C Recommendation, 30 April 2013. https://www.w3.org/TR/prov-dm/.
  10. Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., et al. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3, 160018. [CrossRef]
  11. Peng, R. D. (2011). Reproducible research in computational science. Science, 334(6060), 1226-1227. [CrossRef]
  12. Sandve, G. K., Nekrutenko, A., Taylor, J., & Hovig, E. (2013). Ten simple rules for reproducible computational research. PLoS Computational Biology, 9(10), e1003285. [CrossRef]
  13. Gil, Y., David, C. H., Demir, I., et al. (2016). Toward the Geoscience Paper of the Future: Best practices for documenting and sharing research from data to software to provenance. Earth and Space Science, 3(10), 388-415. [CrossRef]
  14. Buneman, P., Khanna, S., & Tan, W.-C. (2001). Why and where: A characterization of data provenance. In J. Van den Bussche & V. Vianu (Eds.), Database Theory - ICDT 2001, Lecture Notes in Computer Science, vol. 1973 (pp. 316-330). Springer.
  15. Simmhan, Y. L., Plale, B., & Gannon, D. (2005). A survey of data provenance in e-science. SIGMOD Record, 34(3), 31-36. [CrossRef]
  16. Hogan, A., Blomqvist, E., Cochez, M., et al. (2021). Knowledge graphs. ACM Computing Surveys, 54(4), Article 71, 1-37. [CrossRef]
  17. Kazemi, S. M., Goel, R., Jain, K., Kobyzev, I., Sethi, A., Forsyth, P., & Poupart, P. (2020). Representation learning for dynamic graphs: A survey. Journal of Machine Learning Research, 21(70), 1-73. https://jmlr.org/papers/v21/19-447.html.
  18. Cai, L., Mao, X., Zhou, Y., Long, Z., Wu, C., & Lan, M. (2024). A survey on temporal knowledge graph: Representation learning and applications. arXiv:2403.04782. [CrossRef]
  19. Gao, Y., Xiong, Y., Gao, X., et al. (2023). Retrieval-augmented generation for large language models: A survey. arXiv:2312.10997. [CrossRef]
  20. Ji, Z., Lee, N., Frieske, R., et al. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248, 1-38. [CrossRef]
  21. Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, 14045. [CrossRef]
  22. Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631(8022), 755-759. (See also Author Correction, Nature, 2025, https://doi.org/10.1038/s41586-025-08905-3.). [CrossRef]
  23. Messeri, L., & Crockett, M. J. (2024). Artificial intelligence and illusions of understanding in scientific research. Nature, 627(8002), 49-58. [CrossRef]
  24. Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). Concrete problems in AI safety. arXiv:1606.06565. [CrossRef]
  25. Wang, L., Ma, C., Feng, X., et al. (2024). A survey on large language model based autonomous agents. Frontiers of Computer Science, 18(6), 186345. [CrossRef]
  26. Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J., & Ha, D. (2024). The AI Scientist: Towards fully automated open-ended scientific discovery. arXiv:2408.06292. [CrossRef]
  27. Boiko, D. A., MacKnight, R., Kline, B., & Gomes, G. (2023). Autonomous chemical research with large language models. Nature, 624(7992), 570-578. [CrossRef]
  28. Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT 2021) (pp. 610-623). [CrossRef]
  29. Birhane, A., Kasirzadeh, A., Leslie, D., & Wachter, S. (2023). Science in the age of large language models. Nature Reviews Physics, 5(5), 277-280. [CrossRef]
  30. Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.
  31. COPE (Committee on Publication Ethics). (2023). Authorship and AI tools: COPE position statement. https://publicationethics.org/cope-position-statements/ai-author.
  32. Nature Editorial. (2023). Tools such as ChatGPT threaten transparent science; here are our ground rules for their use. Nature, 613(7945), 612. [CrossRef]
  33. Obsidian. Obsidian Help: Graph view and local knowledge management. Product documentation. https://help.obsidian.md.
  34. Zep. Graphiti: Temporal knowledge graph framework. Technical documentation. https://github.com/getzep/graphiti.
  35. Google. NotebookLM: Source-grounded research and thinking tool. Product documentation. https://notebooklm.google.
  36. Perplexity AI. Search and academic discovery documentation. Product documentation. https://docs.perplexity.ai.
  37. OpenAI. Codex documentation. Product documentation. https://developers.openai.com/codex.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.