Generative artificial intelligence has reduced the technical cost of producing scientific text, but increased output does not by itself produce cumulative knowledge, conceptual novelty, or epistemic progress. This article develops the Reinforcement Learning Contour (RLC) as a conceptual and architectural framework for converting AI-assisted scientific work from repeated document generation into governed recursive learning. The framework distinguishes an epistemic contour, Cₜ, from the governed transformation operator, Φₜ, through which a research episode may produce a successor contour. The contour records the time-indexed configuration of knowledge, ontology, relations, retrieval, evaluative criteria, inquiry policy, provenance, memory, tools, and governance constraints. The operator organizes evidence retrieval, source verification, adversarial evaluation, human approval, and reintegration. Changes in the answer-producing function within a contour are analytically separated from changes to the transformation operator itself; the latter are treated as Learning III-like events that cannot be autonomously committed and require an explicit human meta-decision. The central proposition is a differentiation principle: the value of a research cycle is proportional not to the quantity of text it produces but to the beneficial, traceable differentiation it creates between the preceding contour and the contours governing subsequent perception, interpretation, decision, and action. The framework is positioned relative to Popperian criticism, Lakatosian research programmes, Bateson’s orders of learning, double-loop learning, socially situated objectivity, provenance, temporal knowledge representation, and agentic science. It specifies a role-based architecture in which generation authority is strictly weaker than human acceptance authority; a twelve-step governed research cycle; a Structured Epistemic Change Record that documents change without collapsing quality into a scalar reward; novelty and anti-recursion controls; failure modes; seven falsifiable hypotheses; and a comparative pilot design. The proposal is explicitly unvalidated. Its claim is not that automation guarantees scientific progress, but that claims of recursive learning can be made more inspectable, contestable, and governable.