Submitted:
13 August 2026
Posted:
14 August 2026
You are already at the latest version
Abstract
We propose a context-first framework for future AI agents engaged in scientific discovery. Instead of assuming objects, variables, or coordinates, an agent starts from histories of contexts and observations and compares finitely describable predictive programs. The comparison uses one prequential Minimum Description Length (MDL) code: prediction and description costs share a common coding unit, task frequencies are fixed by an external evaluation protocol, inherited paradigms enter through conditional code length, and covariance tests are coded as held-out predictive channels rather than attached with arbitrary penalty coefficients. Predictive quotients, studied objects, attributes, and coarse generative world sketches thereby become model-class choices rather than priors. An open-world context-action dictionary can reuse, compose, split, merge, or introduce primitives under held-out compression. We also formulate local observation generally as a possibly stochastic, noninvertible channel with restricted support; pointwise warps exist only on shared latent support, while the remaining region requires posterior prediction. A finite implementation alternates continuous parameter fitting with discrete structure proposals, evaluates them on ordered held-out blocks, and requests observations that best discriminate the surviving programs. A Solar-System telescope specializes this construction to bounded circular fields, rotation, zoom, parallax, point-spread matching, and active follow-up. The framework is a theoretical and computational blueprint, not an empirical validation, and its selections are conditional on the supplied candidate programs, coding protocol, intervention coverage, and optimization. Overall, this work is intended to serve as a comprehensive operational handbook and practical guide for future autonomous AI agents conducting scientific discovery.
Keywords:
artificial intelligence
; scientific discovery
; minimum description length
; observer covariance
; local observation
; open-world context factorization
1. Introduction
Artificial intelligence is now deeply involved in scientific practice, from materials screening and chemical planning to automated experiment and data analysis [1,2,3,4,5]. Yet many successful AI-for-science pipelines begin after the hardest representational work has already been done. The objects have been segmented, the variables named, the units calibrated, the relevant scale chosen, and the data table prepared. A theory of scientific discovery for AI should begin earlier: with an agent that must decide what in its observations deserves to become an object, a variable, a reference, a state, or a law.
This upstream problem is only partly addressed by existing frameworks. World models, AIXI-like agents, and active-inference approaches formalize compact prediction and action under uncertainty [6,7,8,9,10]. Sparse regression, symbolic regression, AI Feynman, AI Physicist, and Physics-Informed Neural Networks search for concise laws once a representation, or at least a modeling language, has largely been chosen [11,12,13,14,15]. Operational reconstructions, contextuality theory, and renormalization theory explain important failures of classical variables and fixed microscopic descriptions [16,17,18,19,20,21]. These lines of work are indispensable, but they do not yet give a single loss-level account of how an AI should move from raw active observation to the scientific language in which laws are stated.
This paper proposes such an account expressly as an operating framework for future autonomous scientific AI. Its primary addressee is the AI agent itself, not a human reader: statements of the form “the agent should compare”, “retain”, or “observe” are intended as machine-executable instructions. Human readers are secondary designers and auditors who specify the candidate languages, safety limits, and evaluation protocol and who inspect the resulting evidence. The primitive data are histories of observations and contexts. A candidate theory is a predictive program, and it is selected by a Minimum Description Length (MDL) code that combines held-out prediction with the conditional code length of the program [22,23]. The predictive comparison is deliberately purpose-relative. It may evaluate raw future images or trajectories, externally specified coarse-grainings, prompt-selected objects or features, and even textual scientific claims judged by a semantic verifier. This makes explicit a fact that an operating agent must track: a scientific representation is selected relative to what is being asked, at what scale, and through which admissible observations.
The resulting framework is not meant to be an assumption-free derivation of all of physics. It requires candidate model classes, calibrated evaluators, admissible context changes, and real experiments. Optimization may be hard, and many claims are conditional on idealized accessibility and regularity assumptions. The claim is more modest but, we believe, useful: once these ingredients are specified, the same objective can decide when histories should be quotiented, when raw context effects should be reused, composed, split, merged, or represented by a new primitive, when studied objects and object hierarchies should be retained, when attributes should be introduced, and when a classical description should be compared with a richer one. The familiar nuisance, measurement, anchor, purpose, and scale roles are diagnostic signatures of the learned effects, not an exhaustive ontology.
Several ingredients are important for this unification. First, observer covariance turns transformations between observers into constraints on the learned representation rather than into after-the-fact coordinate conventions. Second, local observation is represented as a restricted and often information-losing channel, so a future agent knows when a coordinate warp is valid, where it is undefined, and what must instead be predicted generatively. Third, object discovery is treated as a variational factorization problem: an object is retained only when it reduces the total predictive code. Fourth, attributes are classified by their transformation behavior, with standard Lie-orbit, discrete-component, conservation, and quotient-descent facts used as diagnostics rather than claimed as new theorems. Fifth, the theory allows coarse generative world sketches: the agent need not explain every background pixel if the scientific question evaluates only the coarse, prompt-relevant future.
Finally, we upgrade the initial context taxonomy to an open-world action factorization and work through active celestial observation as a feasibility case. The article contains no empirical demonstration; every proposed selection must ultimately be tested in controlled environments and real experiments. Its intended contribution is an agent-facing workflow: a future scientific AI can use it to decide what to observe next, which distinctions to retain, which local observations can be compared, and which externally supplied representation class currently gives the shortest calibrated predictive code.
2. A General Theory of Scientific Discovery for AI
We now reframe scientific discovery from the strongest possible starting point. The primitive data are only contexts and observations; the goal is to predict future observations as accurately as possible while using as few rules as possible, under the constraint that the theory must transform consistently across admissible observer changes. In this view, “state”, “law”, “concept”, “attribute”, “gluing”, and even the possible use of quantum, second-quantized, or field-theoretic language are not starting assumptions. They are structures to be compared under one selection principle:
A scientific theory is the shortest observer-covariant predictive program for future observations under active context choice.
This formulation deliberately strips the theory down to a single basic objective. The role of the rest of this section is to show how quotienting, context splitting, studied objects, object hierarchies, concepts, attributes, coarse generative sketches, gluing failure, attribute death under isolation, and comparisons among operational, Hilbert, occupation-number, and scale-indexed descriptions can all be tested by that objective rather than by independent patch rules.
2.1. One Basic Principle: the Shortest Observer-Covariant Predictive Program
Let
denote the history of past observations, , and raw contexts, , up to time t. Here may be an image, a sensor vector, a table of measurements, a text record, or any multimodal observation. The raw context records the conditions under which is obtained—for example, the observer, apparatus setting, reference frame, intervention, prompt, or resolution—before any roles are assigned to these conditions. Let be the measurable history space and its time-indexed union. Let be an admissible, possibly stochastic and history-dependent policy for choosing future contexts and interventions. Thus is a rule for selecting what is done or observed next, rather than a realized future context sequence, and is the set of policies allowed by the experimental and operational constraints. Write for the future-observation random variable whose realization is , and let
be the random future observation block over a finite horizon H, with denoting a realized block. Let be the measurable space of such blocks and let denote the set of probability measures on a measurable space . All conditional laws below are understood up to almost-sure equality under an evaluation measure over histories, admissible policies, contexts, and realized futures; empirical implementations replace by held-out sample averages.
A candidate scientific theory is a finitely describable executable predictive program, where is the externally specified class of programs or models available to the discovery procedure. Concretely, T specifies an internal measurable state space , an encoder , controlled state-update kernels, channel prediction kernels, and any learned observer-action, decoder, and nuisance-model parameters needed to execute those kernels. For every admissible history–policy pair and every future quantity specified by the scientific task, where , the program supplies a predictive law
Here is the distribution predicted by T, not the unknown data-generating distribution p used below to define predictive equivalence. If T models the raw future, it supplies and the channel prediction may be the pushforward , where maps a distribution on raw future blocks to the distribution of . The framework also permits T to predict directly, without encoding fine details discarded by C. The program T includes all learned rules and finite-precision parameters needed to construct an internal predictive state from , propagate predictions under , produce the required channel outputs, and implement admissible observer transformations; its full executable specification is what the description-length term below charges. By contrast, the observed data, the admissible-policy class, and fixed task specifications or evaluators are external to T unless they are themselves learned.
The predictive target need not always be the raw future observation itself. Sometimes the agent must predict a future image or trajectory; sometimes it must judge whether a linguistic claim is true; sometimes the scientific question intentionally ignores microscopic detail. We therefore introduce an evaluation family
where J indexes the evaluated channels and is a measurable map from a future observation block to the target of channel j. The declared frequency is the probability that this channel is requested by the evaluation protocol; all channel and auxiliary-event frequencies are normalized together in the complete protocol. It is not a fitted penalty coefficient. For brevity, write for the channel law in Eq. (3). Unless explicitly subscripted otherwise, expectations below are with respect to . The identity channel recovers ordinary image-level or trajectory-level prediction. An externally supplied coarse-graining map
expresses that the theory is required to predict only a chosen resolution of the world, for example the orbit of Jupiter rather than the weather inside Jupiter. The map is part of the question being asked, not something the theory is allowed to choose in order to hide errors.
A prompt or purpose description , such as “in the solar system, focus on the motion of the Moon”, induces another evaluation channel, where is the space of admissible purpose specifications. Let
be a fixed or separately validated prompt-conditioned feature extractor, mask, query operator, or semantic parser. Write for the corresponding prompt-conditioned prediction supplied by T. The interest-conditioned prediction loss is then
This term makes the active policy and the representation pay attention to what the scientific question asks about. A prompt may also generate focused probes or counterfactual test cases through a generator ; those probes simply add further channels of the same form to .
Textual predictions are handled similarly. Suppose the theory emits a claim , for example “tomorrow at 4 p.m. the Sun will be in the east”, where is an externally specified admissible claim language. A calibrated semantic judge or verifier
compares the claim with the realized future and any relevant prompt. The text-claim loss can be written as
If the judge is implemented as a differentiable language model, visual-language model, or distilled verifier, gradients can pass through this term. If the judge is only available as a black-box evaluator, Eq. (9) remains a valid outer objective and can be optimized by reinforcement-style, evolutionary, or other gradient-free search. Thus differentiability is an algorithmic convenience, not a conceptual requirement of the theory.
For an evaluation family , define the multimodal predictive loss
where the first term includes raw, coarse-grained, and prompt-conditioned channels as special cases. The nonnegative frequencies, including when present, are normalized over the complete event family when this pedagogical channel sum is embedded in the prequential protocol below. A distinguished prompt prediction is either one of those events or a separately listed event type; it is never counted twice.
When T predicts several channels directly, the predictions must also be jointly coherent. In particular, whenever for a known measurable map , admissibility requires
More generally, every finite family of jointly evaluated channel laws must admit a common joint extension. This condition prevents a program from winning separate channel scores with mutually contradictory forecasts. To make unlike terms commensurable, fix an external evaluation protocol before comparing theories. It specifies how often each task, prompt, intervention, observer-change test, and held-out data block is encoded. The canonical selection score is log loss, because it is already a predictive codelength in nats. Other proper scores may be reported as diagnostics or external decision utilities, but they are not added to the MDL code; a domain-specific non-log score can enter only through a separately justified probabilistic coding model declared before the comparison. Define
where an evaluation event e includes its channel or test, history, policy, prompt, and target. Covariance is included by drawing paired observer-change events and coding one observer’s target under the law transported from the other; a separately predicted prompt target is handled by its own event frequency. Hard scientific requirements, such as channel coherence, forbidden interventions, or exact symmetries in an idealized model class, define admissibility rather than a soft cost.
The basic objective is therefore the single codelength
where N is the declared number or effective mass of future evaluation events and is a fixed prefix-code length in bits. Multiplication by converts the model code to nats, so no free conversion coefficient remains. Equivalently, one may use total rather than mean held-out log loss and omit N. Here is an optional library of inherited paradigms, background laws, units, and trusted approximation schemes. The relative description length rewards a theory that can be stated compactly as a modification of existing science. In the principled default used here, scientific inheritance is expressed only by this conditional code: a theory may borrow old structure when doing so shortens its message, without an additional freely weighted paradigm penalty. If a user insists on a preference not expressible as data, feasibility, or code length, it must be reported as an external utility and explored by a Pareto or sensitivity analysis rather than hidden inside Eq. (13). The symbol denotes the full external problem specification: , , the active channels and their spaces, the event-frequency protocol , the probabilistic coding conventions, the prompt space and prompt extractor, the claim language and calibrated verifier, admissible observer changes, hard constraints, the evaluation mass N, and the optional inherited library . Thus no learned component can be hidden outside the description-length charge merely by omitting it from a short tuple.
All later gains inherit this convention. For two externally declared, matched candidate classes and , define
A positive gap says only that compresses the declared evaluation stream better than . Finite-data acceptance requires a predeclared uncertainty margin, for example a one-sided confidence bound or sequential coding regret, and stability across scientifically reasonable alternative protocols . The event frequencies are not learned from the same held-out outcomes; they encode the scientific purpose and must be published with every reported gain. For a deployment task they are fixed by the target occurrence distribution; for a designed experiment they are fixed by the experimental design. They may be varied in a disclosed sensitivity analysis, but not tuned post hoc to make a preferred theory win. For the remainder of the paper N and are fixed, and abbreviates .
To write a covariance diagnostic and its associated evaluation events, let be an admissible observer change acting on histories, future policies, future observations, and, when appropriate, the evaluation channels. For each channel let denote its transformed channel and let be the corresponding measurable output transformation. Let be a fixed, normalized distributional diagnostic, such as a held-out Kullback–Leibler divergence when it is estimable. A natural normalized discrepancy is
In the idealized limit this residual vanishes for every admissible observer change and every admissible evaluation channel. In Eq. (13) it is not multiplied by a free : either exact covariance defines the admissible class, or observer-paired cases occur in and their transported predictive log scores contribute nats to . is then a reported diagnostic, or a calibrated surrogate used during optimization, rather than an independently weighted scientific objective. Observer covariance therefore means that the same scientific content can be expressed in different observer descriptions without changing what is predicted at the chosen level of evaluation.
This role is analogous to transformation-first derivations of kinematics and relativity: one does not freely assign distinct observer descriptions, but checks them against identity, reciprocity, composition, and continuity [24]. Whenever reversible, globally supported observer changes between local observer contexts are available, covariance induces a groupoid action on the internal representation. Writing its arrows as , one requires
Let denote the candidate theory’s local state in chart , and let be a metric on the target chart . A convenient empirical residual is
In an exact groupoid, the two-observer inverse law and the three-observer cocycle generate all higher path-independence conditions. Four-observer and longer loops are therefore not independent structural equations in the noise-free limit; they are valuable as overdetermined empirical tests when transformations are learned approximately. Eq. (16) is not yet a claim about what the correct scientific object is. It is only the minimal consistency condition that any admissible object must satisfy if it is to deserve observer-independent meaning.
Most real experiments are more general because each apparatus reveals only part of a latent phenomenon and may destroy information. Let denote a candidate latent world state and let a context c define a support set together with a local observation kernel
The kernel covers deterministic projections, noisy sensors, censoring, saturation, finite temporal windows, spatial crops, limited spectral bands, destructive assays, and coarse resolution. If two deterministic local charts and are one-to-one on their shared latent support, their pointwise correspondence is only the partial map
Outside there is no observation-to-observation warp. The theory must instead push its posterior over through :
For stochastic or many-to-one kernels, Eq. (22) is primary and Eq. (21) may not exist at all. Thus global invertible groupoid actions are a special case of a broader category of local Markov kernels and partially defined correspondences. Every comparison must carry its support or exposure mask, and scores across different supports must be normalized or included through the declared event protocol . This general formulation applies equally to microscope tiles, cropped video, limited-angle tomography, frequency windows, missing detector channels, destructive biological measurements, and telescope fields.
2.2. Predictive Equivalence Selects Quotient States
The first major consequence of Eq. (13) is that quotienting is not an added principle. It is selected by predictive compression, but the quotient is always relative to every positive-frequency evaluation event in the problem specification, not merely to a list of named channels. To make this precise, let index all positive-weight evaluation experiments. For each , let be the admissible report space and let be the realized logarithmic code, or a declared decision diagnostic used only to refine the quotient for that task. For an ordinary channel, a report is a law and ; the prompt and text-claim codes are included in the same way. Define the conditional risk functional
Then define
for every admissible purpose value having positive evaluation mass. For strictly proper channel scores, equality of all conditional risk functionals is equivalent to equality of the corresponding conditional target laws. If the identity channel is included, this is a fine predictive equivalence over raw future observations. If only is evaluated, the quotient intentionally discards differences below the chosen coarse-graining. If an interest prompt is active, histories are distinguished only insofar as they affect the prompt-conditioned future features. Thus the same physical data can support different legitimate scientific quotients depending on whether the question concerns the motion of Jupiter as a whole, the weather inside Jupiter, or the Moon’s trajectory in a solar-system scene.
Let with its quotient sigma-algebra. The induced equivalence class
constitutes the minimal predictive state for the specified scientific purpose. It retains exactly the information in history that matters for future predictions under the admissible policies and evaluation channels, and nothing else. This is the active-observation analogue of causal-state constructions in computational mechanics, predictive-state representations in controlled dynamical systems, and their input-output generalizations [25,26,27].
The reason Eq. (25) is unavoidable is straightforward. If a candidate theory assigns different internal states to two histories that are equivalent under Eq. (24), merging those states cannot worsen any evaluated Bayes risk. Provided the code for a representation does not reward redundant state distinctions, quotienting also cannot increase description length or covariance cost. Hence every admissible predictor has a quotient representative with no larger objective. If the representation code is strictly larger whenever an operationally redundant distinction is retained, every minimum-code representative identifies predictive-equivalent histories almost surely; without that strictness assumption, only existence of an equally good quotient minimizer is guaranteed. In finite data, exact equality is replaced by a held-out two-sample or conditional-risk tolerance, and identifiability additionally requires that the evaluated policy distribution cover the context changes being compared. Quotienting is thus a consequence of the objective under explicit coding and coverage conditions, not a separate postulate. In the special classical case where a hidden Markov microstate exists and global classical gluing remains valid, the predictive quotient may be represented by an ordinary belief state. In that regime, belief states introduce no new ontology; they are simply a classical realization of minimal predictive sufficiency under partial observation.
2.3. Raw Context Splits Itself by How It Acts on the Predictive Quotient
The second major consequence is that context classes need not be assumed in advance. Start from a raw context variable and ask how primitive context changes act on the purpose-relative predictive quotient . The resulting action types induce the familiar scientific split, with one additional class needed for explicit scientific purpose.
- 1.
- Nuisance or gauge contextg: a change is nuisance if it leaves the predictive quotient unchanged,so distinguishing it would increase description length without improving the evaluated prediction.
- 2.
-
Measurement contextm: a primitive change is of measurement type if it changes the readout applied to the current predictive state, rather than redefining the predictive quotient itself. There exists a Markov readout kernel , equivalently , such that
- 3.
-
Anchor or reference contextr: a primitive change is anchoring if removing it changes the quotient map itself. Let erase the reference information supplied by , and let be the quotient recomputed after that erasure. The context is anchoring if there exists a surjective but non-injective map such thatEquivalently, some histories that were predictively distinct become indistinguishable after the anchor is erased. A camera identity, a laboratory frame, a clock synchronization convention, or an external comparison object can therefore be an anchor: it is not itself a coordinate attribute, but a condition that makes certain predictive distinctions available.
- 4.
- Purpose or probe contextp: a change in prompt or evaluation channel changes what the agent is asked to predict, even if the physical apparatus remains fixed. If the prompt only changes the scoring projection , it is a purpose context. If implementing the prompt requires moving a telescope, choosing a sensor, or intervening on the world, that physical action belongs to the policy or measurement context. This separation prevents a scientific interest, such as “track the Moon”, from being confused with an intrinsic property of the world.
- 5.
- Scale context: a change in resolution is scale if it acts as a generally noninvertible coarse-graining semigroup on future observations and may therefore alter which predictive quotient is shortest. The externally supplied in Eq. (10) is one way to specify the scale of the question; a learned renormalization map is another way to describe how descriptions change when varies.
Thus the familiar diagnostic decomposition becomes
The entries are roles, not necessarily mutually exclusive labels: one physical control may instantiate several of them. This is still a post hoc diagnostic vocabulary rather than a proof that five names exhaust all possible context effects. Section 3 removes that remaining restriction by learning an open dictionary of context actions and its factorization directly from the objective.
2.4. Studied Objects Emerge as Variational Predictive Factors
The theory should not assume in advance which things in the visual field are the objects of study. It should allow objects, multiple objects, and hierarchical objects to emerge when they shorten the predictive program. A candidate object decomposition at time t is written as
where is a latent state, is a vector of candidate attributes, is a soft region, slot, graph node, or semantic support in the observation, and records type or sector information. For hierarchical objects, the slots are arranged in a tree or directed acyclic graph ; a planet may be a single object for orbital mechanics and a system of atmospheric objects for climate. Below, and denote the complete finitely coded proposal and update rules that generate these time-indexed slot and hierarchy families, not merely a decomposition of one frame.
A decomposition deserves the name “object” only if it produces a shorter observer-covariant predictor than an unfactored world model. Let be a class of predictors that encode the future without persistent objects, and let be a class that predicts through persistent object factors, their relations, and their hierarchy. Define the object gain
If , the decomposition is not merely a convenient visualization; it is selected by the same variational principle that selects the theory. The pressures that make an object emerge are predictive persistence, transformation covariance, sparse interaction, intervention separability, and code reuse. In an explicit implementation it is useful to report the diagnostic vector
where every is a normalized empirical residual under . Here persistence penalizes arbitrary slot swapping through time; object covariance requires the same object to be recoverable across observers; separability rewards interventions that affect one object without rewriting the whole scene; sparsity rewards local interactions; and hierarchy rewards part-whole descriptions when they reduce total code length. These quantities are not summed with free multipliers in the scientific objective. They are either encoded as probabilistic evaluation events, imposed as declared admissibility tolerances, or used as training surrogates whose final model is still selected by Eq. (13).
The studied object is then selected either endogenously or by a prompt. Given a prompt , the marginal importance of object i can be estimated by an ablation gap
where is the same predictor with object i masked, marginalized, or compressed away. Objects with large become the prompted studied objects. Without a prompt, objects with large expected information gain, large object gain, or large contribution to future prediction are selected automatically. Thus attention to studied objects can itself be variational: it is the lowest-cost way to explain the future at the resolution and purpose specified by the loss.
2.5. Concepts, Attributes, and Lie-structured Transformation Behavior
Once the predictive quotient and possible object factors have been formed, the next question is whether they contain reusable internal structure. A concept is a reusable predictive factor. Let
be a candidate subrepresentation extracted from the predictive quotient, an object slot, or a hierarchy node. Let be a measurable context–policy family with the evaluation measure restricted and renormalized to that family; write for the corresponding restricted objective. We call a concept over if every active evaluated channel can be routed through that shared factor:
The point of introducing is code reuse rather than metaphysical naming. Let denote models that fit separate local predictors for each , and let denote models in which those predictors must factor through . The concept gain is
If , the shared factor is justified as a concept.
An attribute is a low-dimensional coordinate, label, invariant, or readout on a concept or object that remains predictively meaningful across a nontrivial family of contexts. A candidate attribute is a map
such that the dependence of future predictions on can be mediated through :
The attribute gain is
where the two matched model classes share the same encoder budget, decoder family, optimization protocol, and all capacities except explicit access to A; is obtained by deleting or marginalizing that coordinate, while requires predictions to use it. If , retaining A makes the overall predictive program shorter even after paying its complexity cost; if , the coordinate should be deleted.
The next question is not only whether an attribute is useful, but how it should be represented. When the relevant arrows of the observer groupoid restrict to a Lie-group action G on the attribute manifold, write for the identity component and for the stabilizer of a representative in component . When the learned value space is a smooth manifold and the connected observer action is transitive on each connected component, the standard homogeneous-space orbit theorem motivates the representation
Here is a continuous orbit coordinate within one connected component, while is a discrete sector label. Orbit coordinates should be handled by equivariant maps, Lie generators, and smooth chart transformations; sector labels should be handled by classification, counts, occupation numbers, or discrete transition events. This is a conditional routing heuristic based on standard geometry, not a new dichotomy theorem: singular, stratified, fractal, or nonhomogeneous learned spaces must remain available as competing representations. Appendix A.2 states the standard homogeneous-space result under the precise smoothness assumptions used by this routing rule.
Physical attributes are not exhausted by only these two elementary templates. Orbit-like coordinates and sector-like labels are the topological backbone, but real scientific attributes are more finely classified by their transformation behavior under observer changes:
- 1.
- Invariant scalars satisfy under the admissible observer groupoid. Rest mass in a fixed regime, charge after calibration, and some dimensionless constants are examples.
- 2.
- Equivariant orbit coordinates live on a continuous orbit or in a representation of a continuous group. Position, orientation, phase, and many coordinate components are of this type.
- 3.
- Reciprocal relational quantities are attached to ordered observer pairs and obey , with . In one-dimensional inertial kinematics, relative velocity locally has .
- 4.
- Appearance or readout attributes depend on intrinsic attributes and on the observer, anchor, or measurement map. Apparent angular size, brightness in a camera, and projected shape are not the same as true size; they are readouts such as
- 5.
- Thermodynamic, field, and material attributes such as pressure, temperature, density, elastic moduli, and local fields are tied to a coarse-graining scale and a subsystem convention. They are often neither pure coordinates nor pure invariants; they are effective attributes that transform under changes of scale, material chart, or observer readout.
- 6.
- Sector labels and counts distinguish disconnected components of the value space, such as particle number, species, phase sector, or occupation number.
- 7.
- Scale-flow attributes and couplings are parameters whose meaning changes under coarse-graining. Their correct transformation law is an RG flow rather than an ordinary coordinate change.
This classification explains why Lie theory is useful but not sufficient by itself. Continuous observer changes generate Lie groups or Lie groupoids; their orbits, stabilizers, invariants, and representations organize coordinate-like attributes. Discrete sectors, scale flows, and semantic readouts require additional structure, but they still enter the same variational objective.
When the relevant observer changes are continuous, the learned transformations should be locally generated by infinitesimal operators. For a one-parameter observer change acting on an attribute manifold , define
The generators and their commutators reveal the Lie algebra of the observer action. A coordinate is orbit-like when moving along these generators changes its value while preserving predictive form; a scalar invariant is a function constant along all such generators; and an equivariant tensor is an attribute whose components transform under a finite-dimensional representation. In practice these structures can be discovered by minimizing an equivariance residual
together with the inverse and cocycle residuals in Eq. (16). Thus coordinates, invariants, reciprocal quantities, and apparent readouts are not inserted by hand; they are the attribute maps whose predictive gain survives compression and whose transformation behavior closes under the learned observer groupoid.
Even when a candidate coordinate has positive attribute gain locally, this still does not imply that there exists a single sharp, context-independent value of that attribute. What an accessible observer context directly provides is, in general, only a local chart . For overlapping accessible contexts , the local charts must satisfy
Only under this compatibility condition is it meaningful to regard the various as different coordinate realizations of one and the same attribute. Equivalently, there exists a context-independent latent attribute and chart maps such that
If Eq. (45) holds, the attribute is globally sharp. If it fails, the correct conclusion is not that the attribute has a sharp global value that is merely unknown; rather, the supposed sharp attribute is only locally defined over the accessible context family. This is the attribute-level form of gluing failure.
A sharper test is descent to the accessible quotient. If a richer state z and another richer state are merged by the accessible quotient q, then a classical attribute A can survive at the accessible level only if it is constant on such fibres. Let be a specified pair-sampling law on supported on pairs satisfying . A practical residual for a candidate fine-state attribute is
A small residual supports pulling the attribute back from the accessible state; a persistent large residual is evidence against that descent model, not merely a parameter estimate with large uncertainty. The exact case tests constancy on quotient fibres. The elementary descent fact is that a map A factors as exactly when it is constant on every fibre of the quotient q. Consequently, no jointly faithful frozen chart family that still presupposes the same non-descending attribute can repair the failure; this says nothing against a different attribute, a coarser readout, or an enlarged state representation. Appendix A.3 records this standard fact in the notation used here.
Once local charts for the same candidate attribute have been identified, the transformation law between observer contexts becomes a scientific object to be discovered. One should not count equations by the number of observers in a chain. The independent structural constraints are identity, two-observer reciprocity, and three-observer composition; longer loops are implied in the exact case and serve mainly as robustness tests. Unknown constants or functions are fixed only to the extent that these functional constraints, together with regularity, symmetry, and data, have sufficient rank. The one-dimensional Lorentz derivation is the simplest example: reciprocity and composition reduce the transformation family to a one-constant form, and experiment fixes the remaining constant.
A complementary diagnostic applies when a globally sharp orbit-type attribute evolves in a closed homogeneous regime. The standard discrete Noether calculation says that a stationary local path code with a continuous symmetry has a conserved conjugate statistic. In a learned model this is useful as a test, not as a new theorem: discover a candidate continuous generator, compute its conjugate statistic, and measure held-out drift. Large drift can indicate symmetry breaking, a missing context, a wrong observer map, failed stationarity, or an insufficient quotient. Translation-like and rotation-like symmetries may later be interpreted as momentum-like and angular-momentum-like only after these hypotheses are tested; camera rotation alone is merely an observer change. Retention is a separate matched code comparison under Eq. (14): the agent retains the predictive information in the statistic only when the model class that can encode it achieves a statistically supported codelength advantage. This standard diagnostic can also guide active experiments that perturb origin, orientation, time window, or apparatus calibration. Appendix A.1 gives a self-contained discrete statement and approximate residual bound in the present notation.
2.6. Coarse Generative World Sketches Rather than Literal Background Subtraction
Human scientists usually do not construct theories by first subtracting every background pixel and every noise source. More often, they form a rough generative picture of the physical world and judge whether the generated picture matches reality at the relevant level of resolution. This can be made precise inside Eq. (13). Let be a coarse world sketch, consisting for example of object slots, relations, approximate geometry, fields, and a small number of dynamical variables. Let denote nuisance background and sensor noise. A superscript + denotes the corresponding future block over the same horizon as . A generative observation model has the form
Here the decoder parameter and every learned nuisance law are part of T and are charged by . The scientific theory is not required to predict all of and in detail. Instead it may marginalize or cheaply code them while being judged through the coarse and prompt-conditioned channels. The induced coarse-channel law is
Under log score, the resulting sketch loss is
This is an implementation of the g-channel summand in Eq. (10), not an additional loss to be counted twice. Equivalently, a generated rough future may be compared to the true future only after applying and . A theory wins when the code for the sketch dynamics plus the cheap nuisance model is shorter and more predictive than a direct pixel-level memorizer.
This is not the same as ignoring the background. If the background contains predictive structure relevant to the chosen evaluation channel, then compressing it away increases and the theory loses. If the background is irrelevant to the current scientific question, then forcing the main law to explain it wastes description length and can hide the true low-dimensional structure. The generative-sketch formulation therefore explains why rough physical pictures are robust to noise: the objective compares the world at the level at which the scientific claim is made, while nuisance variation is integrated out, separately modeled, or assigned to gauge context.
2.7. Gluing Is a Variational Comparison between One Global State and Many Local Atlases
The previous subsection formulated gluing at the level of a single candidate attribute. There the question was whether local coordinates assigned in different contexts could all be understood as chart representations of one latent attribute . That is the simplest form of gluing. However, the failure of a particular attribute to glue does not yet imply that the scientific description as a whole is only local. It may happen that one chosen coordinate is not globally sharp, while a richer predictive state still exists from which all local descriptions can be obtained. The next question is therefore strictly stronger: can the entire family of local predictive descriptions be compressed into one shared global state, or is the shortest accurate theory genuinely atlas-like at the level of the state itself? This distinction is crucial. Attribute-level gluing asks whether one particular coordinate is globally well defined. State-level gluing asks whether there exists some common predictive object—possibly richer than any one chosen attribute—from which all local coordinates and local readouts may be derived. Thus state-level gluing failure is not merely the statement that one attribute is ill behaved; it is the statement that no single global predictive state can replace the local atlas without worsening the basic objective. We formulate this as a variational comparison between two nested model classes. In the first, , there exists one shared predictive state together with context-dependent charts or readouts such that all local predictions factor through that common state:
Here the local state is only a charted representation of one underlying predictive object. In the second class, , each accessible context carries its own local predictive state and local prediction rule,
with overlap maps between local states required only where the corresponding contexts are jointly accessible. The global class is recovered as the special case in which all local states are merely charts on one common latent state. The atlas class is more general because it allows the possibility that no such common latent state exists. Define the gluing defect
If up to statistical tolerance, then any apparent locality of description can be absorbed into context-dependent charts on one shared predictive state, and the scientific object is globally glueable. If , then the locality is irreducible: forcing all contexts through a single global state makes the theory longer or less accurate than allowing a family of overlapping local states. In that case the right conclusion is not merely that one coordinate system was chosen badly, but that the predictive organization itself is only locally chartable. Sheaf-theoretic obstruction to a global section is the geometric expression of exactly this variational fact [16]. The most restrictive globally glueable subfamily is the classical simplex or point-state branch,
where is a preparation distribution over hidden states and are response functions. If this branch remains competitive under Eq. (52), then science can stay classical: local charts and local attributes are merely different descriptions of one globally valid hidden-state structure. In the partially observed case, the same globally glueable branch is expressed through ordinary belief dynamics rather than through a new nonclassical ontology. Only when even the best global branch loses to the atlas class does one need to ask what richer predictive language should replace it.
2.8. When Local Atlases Remain Irreducible: Operational Lifts First, Hilbert Lifts Only When Coherence Wins
If , the first conclusion is only that one global simplex or point-state description is not the shortest accurate theory. The minimal safe lift is therefore to a more general operational state space,
with a convex state space and a family of affine effects [28,29]. This is the correct immediate lift because failed global gluing rules out the simplex, not yet every nonclassical operational theory. To decide whether a further lift to amplitudes is necessary, compare two additional model classes: , in which alternatives combine only by convex mixing, and , in which alternatives may recombine coherently. The associated coherence lift gap is
Only when does the data demand more than mixture-like operational structure. In two-slit-like situations, Eq. (55) is estimated by interference or nonadditivity witnesses, but those witnesses should be understood as experimental estimators of a model-class gap, not as separate first principles. If and one additionally supplies assumptions such as continuous reversible transformations, sufficient transitivity on pure states, composition, and local tomography, established operational reconstructions motivate including Hilbert-space quantum models in the candidate menu [17,18,30,31,32,33]. A useful comparison between real and complex candidates comes from composition. Let denote the Hilbert-space dimension of subsystem A and let denote the real dimension of its normalized state space. Composition gives , while local tomography requires . For real symmetric density matrices one has , which does not obey this multiplicative rule in general; for complex Hermitian density matrices one has , which does. Thus, within a menu that imposes coherent recombination and local tomography, the complex density-matrix class is a natural competitor, with pure states represented by rank-one projectors [17,18,29,30]. This is a conditional model-class argument, not a derivation of quantum theory from Eq. (13). In that branch the predictive rule takes the standard form
with density operators and POVMs . The Born rule is then not inserted at the start of this framework, but it still requires independent quantum-probability assumptions. Envariance is one proposed derivation among competing approaches and has attracted debate about which probabilistic premises are already used [34,35]; Gleason-type representation theorems provide an independent route from noncontextual probability measures on Hilbert subspaces or generalized effects to the trace rule, under their own dimensional and regularity assumptions [36,37]. Accordingly, the candidate-class comparison may be summarized without a derivation arrow:
global simplex, operational atlas, and coherent Hilbert models are compared as externally supplied classes.
This avoids the premature identification of all contextuality with textbook quantum mechanics.
2.9. Isolation as Context Restriction: Unifying Superposition and Divergence as Attribute Death
The physical concept of “isolation” can now be stated cleanly within this informational framework. Isolation is not a special metaphysical condition; it is simply a progressive contraction of the accessible context family. Let denote the family of admissible intervention policies and denote the accessible observational contexts. We mathematically write this restriction as:
Crucially, this contraction can occur in two fundamental directions: (1) Spatial/Relational isolation: removing external reference anchors (r) or comparison systems. (2) Scale isolation (Coarse-graining): restricting the microscopic resolution context (), essentially isolating the macroscopic observer from the ultraviolet details.
When a value space is a smooth componentwise homogeneous manifold under a connected Lie-group action, standard orbit geometry separates continuous homogeneous-orbit coordinates within components from discrete component labels. This observation applies below only to orbit/sector organization. Couplings that run with scale, singular spaces, and nonhomogeneous spaces are handled by separate candidate classes.
When the accessible context shrinks in either direction, the predictive quotient may coarsen. If a previously supported sharp attribute then ceases to descend to the new quotient or ceases to have positive gain, it should no longer be treated as a globally sharp coordinate. No initial sharp assignment is required by the framework.
This loss of definability—the “death” of an attribute—motivates two distinct conditional model comparisons relevant to quantum and field-theoretic descriptions:
2.9.0.1. 1. Orbit-type attributes and Superposition (Spatial Isolation).
If the space of admissible values forms a connected symmetry orbit, (where is the anchor’s stabilizer), then spatial isolation removes the anchor . Removing the anchor can make a sharp global coordinate (e.g., position, orientation, or phase) non-descending on the restricted predictive quotient. When distinct pre-isolation coordinate values are in fact merged by that quotient, the primary event is not gluing failure but attribute death. The elementary fibre-constancy criterion then obstructs only a jointly faithful frozen classical atlas that still presupposes this same coordinate. If, in addition, a coherent candidate class achieves a statistically supported code advantage and the operational-reconstruction assumptions stated above hold, a superposition-like or Hilbert representation becomes a candidate repair. Non-descent alone does not imply quantum superposition.
2.9.0.2. 2. Coupling-type attributes and Divergence (Scale Isolation).
If the attribute is a coupling parameter, bare mass, or charge interacting across scales, its definability is threatened by scale isolation (coarse-graining, ). Unlike positions, couplings are not generally coordinates on the same compact spatial orbit; they parameterize how an effective description absorbs the influence of eliminated degrees of freedom. Under scale restriction a fixed-parameter class may lose predictive compression to a scale-flow class. The fitted coupling may run to a finite fixed point, vary without divergence, or diverge, depending on the model and regime. Divergence is therefore one possible symptom of failure of a fixed-law parameterization, not the exact counterpart of gluing failure.
A unified diagnostic for both forms of attribute death is the accessible Fisher information:
If , the parameter becomes locally unidentifiable from the accessible experiment family. A diverging fitted value may also diagnose misspecification or a singular parameterization, but neither observation by itself identifies the correct repair; the scale-gap comparison below is still required.
A useful refinement of the gluing logic is the elementary descent test. The primary failure is not gluing itself, but the loss of definability of a fixed candidate classical attribute on the actually accessible predictive quotient. Let denote the passage from a richer (actual or hypothetical) descriptive regime to the actually accessible theory. If a putative classical attribute A is not constant on the fibres of q, then it cannot descend to any well-defined function on . Any jointly faithful frozen atlas that still presupposes that same attribute is therefore non-glueable on the accessible theory. This diagnoses ordinary attribute death and also a proposed attribute non-birth only after a concrete refinement q, candidate A, and fibre violation have been supplied; it does not rule out a different classical attribute, a coarser unfaithful readout, or an enlarged state representation.
2.10. Two Additional Lifts from the Same Objective: Sector Closure and Scale Closure
The unified model-selection objective can now compare structural lifts for these distinct failures. The following sector and scale classes are candidates, not consequences of the three appendix theorems alone.
2.10.0.3. Sector lift.
When measurement isolation prevents the continuous tracking of individual particle identities, fixing the particle sector yields severe predictive defects. We define the sector-closure defect:
If , the tested fixed-sector class is suboptimal to the specified variable-sector class. A natural representation then tracks occupancies rather than persistent individual labels. In a Hilbert model with the usual indistinguishability and composition assumptions, the corresponding standard representation is Fock space,
2.10.0.4. Scale lift (Renormalization).
When an AI attempts to maintain a fixed parameter vector across scale isolation, the parameters diverge, causing a scale-closure defect:
If , the tested fixed-law class is suboptimal to the specified scale-flow class over that scale transition. Within this comparison, the selected repair promotes static parameters to scale-indexed functions and may use an effective action
whose couplings flow under coarse-graining [19]. This casts renormalization as a testable representation change under coarse-graining rather than as a response that follows from divergence alone.
The agent should therefore compare, without treating the sequence as a derivation,
anchor-rich vs. anchor-poor coordinate models;
fixed-identity vs. variable-sector models;
fixed-law vs. scale-flow models.
Selecting all three steps still does not derive a unique quantum field theory; the candidate classes must additionally supply locality, composition, field content, dynamics, and empirical adequacy.
2.11. A Variational Derivation Scheme from the Loss to Scientific Structure
The preceding subsections can be read as an operational model-comparison program rather than as a list of separate philosophical claims. Let
collect the learnable quotient map, object decomposition, hierarchy, attribute maps, observer transformations, transformation-composition rule, coarse world sketch, decoder or readout model, and inherited library. The discovery problem remains the single-codelength comparison
where contains the externally supplied candidate programs satisfying hard coherence and safety constraints. Object persistence, equivariance, and cocycle residuals are either encoded through declared evaluation events or used as optimization surrogates calibrated back to that same code; they are not extra uncalibrated terms in the scientific objective.
The stationarity and model-class comparisons of Eq. (64) generate the later structures in a definite order. Variation over encoders identifies histories that cannot be distinguished by any evaluated future prediction, yielding the predictive quotient in Eq. (25). Variation over raw context actions tests reusable factors whose post hoc signatures may be nuisance-, measurement-, anchor-, purpose-, or scale-like. Variation over object slots and hierarchies compares matched classes through the codelength gap in Eqs. (14) and (31). Variation over attribute maps compares matched classes through Eqs. (14) and (39). Variation over observer transformations imposes the normal equations associated with equivariance, inverse consistency, and the cocycle condition; in continuous regimes the infinitesimal version of this variation identifies Lie generators as in Eq. (42). Variation over the coarse sketch and nuisance decoder decides whether a rough physical world picture is cheaper than direct prediction of raw data. Finally, discrete variation over model classes produces the gluing, coherence, sector, and scale gaps.
This is the precise, limited sense in which the later scientific language is selected by the loss from supplied candidate classes. One should not expect every mature physical theory to appear from a single closed-form Euler–Lagrange equation. The construction is variational in a broader MDL sense: a structure becomes part of the scientific description when removing it worsens the optimized objective within a matched comparison, and a representational lift is selected when the tested simpler admissible class loses that same comparison. An executable demonstration can therefore be organized as follows: train candidate predictors with and without object slots, attributes, transformation laws, sketch decoders, and richer atlas classes; compute the corresponding optimized losses and code lengths; and accept a structure only when its gain remains positive under held-out contexts, observer changes, and prompt-conditioned probes.
2.12. A Finite, Anytime Approximation That Can Be Computed
The minimization over all finitely describable programs is not computable in general. For an operating agent, the global minimization implicit in Eq. (13) must therefore be approximated by a finite, budgeted search whose candidate grammar and search history are explicit. A practical implementation needs only four components: a typed library of executable modules; continuous optimizers for the parameters inside a fixed program graph; discrete edit operators on that graph; and an ordered prequential evaluator.
Partition the available experiments by time, apparatus, or intervention into an optional initialization block and ordered evaluation blocks , while reserving an audit set that is not consulted during proposal generation. Any nonempty initialization block is encoded by a fixed start-up code common to all candidates; this candidate-independent term is omitted below. For a candidate structure T, fit its parameters using only and blocks preceding and compute the MDL score with sequential held-out prediction
The last term codes the program graph, module identities, and parameters at their declared numerical precision. If a proposal generator or learned verifier is not fixed externally, its executable description is charged as well. Differentiable modules may be fitted by gradient descent; latent discrete assignments by EM, variational inference, particle methods, or reversible-jump moves; and black-box modules by evolutionary or other derivative-free search. These choices affect search efficiency, not the final comparison rule.
The discrete neighborhood of T is generated by typed edits: merge or split predictive states; add, delete, or hierarchically group object slots; add or remove an attribute bottleneck; reuse, compose, split, or invent a context primitive; replace a global chart by an atlas; or exchange one supplied representation family for another. Parameters shared by parent and child proposals are warm-started. Candidates that violate channel coherence, support masks, intervention rules, or safety constraints are rejected before scoring. The remaining candidates are ranked by Eq. (65), and a beam of the K shortest programs is retained. During search, an edit is retained only when a predeclared lower confidence bound for its paired codelength gain is positive across fresh sequential validation blocks. After the search terminates, the reported winner and a small preregistered set of alternatives are evaluated once on the untouched ; only an audit-stable gain is reported as accepted. Repeated adaptive reuse of one validation or audit set is forbidden: the agent must use fresh sequential blocks, a reusable-holdout mechanism, or a nested outer audit.
| Algorithm 1 Budgeted context-to-theory search |
|
Active choice is also finite. Here is the search-round budget, m is a predeclared no-improvement patience, and K is the beam width. For the retained beam, assign code weights . For each feasible next context c, approximate the information criterion in Eq. (69) by the predictive disagreement
estimated analytically or by simulation, and choose the admissible c with largest . The new observation becomes a fresh prequential block, so active sensing and structure revision alternate. The procedure is anytime: a larger beam, richer edit grammar, more restarts, and more interventions enlarge the explored menu, while every intermediate output remains an auditable comparison among programs actually tested. It does not solve unrestricted program synthesis or guarantee the global optimum; it converts the conceptual framework into a concrete search protocol with measurable computational and statistical failure modes.
2.13. Unified Selection Rule and Active Experimental Design
All branches, objects, prompts, coarse evaluations, and representational lifts can now be written as one nested model-selection rule. Let be an externally supplied menu containing, where scientifically meaningful, unfactored and object-factorized predictors, fine and coarse sketches, global and atlas states, mixture and coherent operational models, fixed and variable sectors, and fixed-law and scale-flow programs. The selected scientific program is simply
Equation (67) deliberately contains no composition of “lift” operators and no implication arrows. It says only that the future AI compares supplied executable classes under one declared code. Pairwise gaps explain which additional capacity pays for itself, but do not prove that nature or the agent must traverse a fixed representational ladder.
Because contexts are chosen actively, scientific discovery is also an experimental-design problem. The next intervention should reduce uncertainty not only about parameters inside a fixed theory, but also about the quotient, object decomposition, attribute chart, observer transformation law, and representation class. With a prompt and evaluation family , define
A unit-free information criterion is
where denotes the collection of evaluated future channels. imposes the declared resource budget B and safety bound R; alternatively, a user-provided utility may define a Pareto frontier, but an uncalibrated physical cost is not subtracted from information in nats. This is the active-observation counterpart of classical Bayesian experimental design [38,39]. The agent should choose contexts that most strongly discriminate among candidate quotients, objects, transformation laws, and lifts at the resolution specified by the scientific question.
2.14. Summary of the General Principles
The framework in this section gives a model-comparison workflow from raw observations to scientific structure. The core principles are as follows.
- 1.
- Multimodal predictive grounding: The primitive data are histories of contexts and observations. The evaluated future may be a raw image, a trajectory, a coarse-grained state, a prompt-conditioned feature, or the truth of a textual scientific claim.
- 2.
- The core objective: Theories are selected by minimizing the prequential codelength in Eq. (13). Predictive events and model descriptions are measured in nats, task frequencies are declared externally, and covariance is either an admissibility condition or an observer-paired predictive event.
- 3.
- Purpose-relative state emergence: Under the stated coding and coverage conditions, predictive equivalence selects minimal quotient states, Eq. (25). A different coarse-graining or prompt can legitimately induce a different quotient.
- 4.
- Open-world context factorization: Raw context changes are encoded by a learned dictionary whose primitives can be reused, composed, split, merged, or created under held-out predictive compression. Nuisance, measurement, anchor, purpose, and scale are overlapping post hoc signatures, not an exhaustive list.
- 5.
- 6.
- Attribute extraction with standard diagnostics: Attributes emerge as reusable coordinates, invariants, sector labels, readouts, or scale-flow parameters when their matched class saves held-out code. Standard orbit geometry, discrete Noether calculations, and fibre-constancy tests diagnose candidate structure but are not claimed here as new mathematical theorems.
- 7.
- Generative sketch robustness: Scientific theories need not match every background pixel. They may generate a coarse world sketch and compare it with reality only through the chosen coarse and prompt-conditioned channels, while nuisance background and noise are marginalized, separately coded, or treated as gauge.
- 8.
- Calculable paradigm shifts: Gluing failure, non-descent of attributes, coherence requirements, sector defects, and scale dependence are computable model-class comparisons under the same objective, not independent philosophical assumptions.
- 9.
- Explicit candidate-class comparisons: Operational, Hilbert, Fock, and scale-indexed programs are externally supplied competitors. A lower code selects one for the declared task; it does not establish a logical derivation or a compulsory sequence of lifts.
- 10.
- Finite computation: A budgeted anytime implementation alternates parameter fitting, typed structure edits, prequential scoring, independent audit, and model-discriminating observation. It returns the shortest program actually tested, not an unearned claim of a global optimum.
This is the sense in which the present theory is genuinely general. It does not begin by assuming classical particles, quantum waves, effective fields, or even the objects of study. It begins with active observation, scientific purpose, predictive parsimony, observer covariance, and optional inheritance from existing paradigms; the appropriate variables, objects, transformation laws, and representational regime are then selected by the informational topology of the world as exposed through the loss.
3. Open-World Automatic Factorization of Context Effects
Section 2 deliberately retained the familiar words nuisance, measurement, anchor, purpose, and scale because they connect the abstract objective to scientific practice. Those names, however, must not become a closed ontology. A genuinely general discovery principle should be able to reuse several known effects in one raw control, separate two effects previously conflated under one name, and introduce an unanticipated context role when held-out interventions demand it. This section upgrades the diagnostic split in Eq. (29) into an open-world variational factorization.
3.1. Context Changes as Experimentally Identified Operators
Let denote an admissible raw context intervention. Its empirical signature is not its human name but the change it induces in the controlled family of predictive experiments. For a fixed T and history h, collect the active channel laws, admissible policy set, state update, and scoring specification into
Here denotes the controlled internal update kernel and the context-dependent part of the evaluation specification. Two raw interventions are operationally equivalent when they induce the same transformation of this predictive bundle on every covered history, up to the statistical tolerance of the experiment. The learnable object is therefore an operator
which may be deterministic or stochastic, invertible or information-losing, and may act on several entries of the bundle at once.
Let be a finite, learned dictionary of elementary context-action programs. For this section write the full executable program as , where contains the encoder, dynamics, channel kernels, and all other learned components, while the dictionary and intervention codes are factored out solely so that their code lengths are visible and disjoint. For each intervention , let be a finite ordered sparse code, where is a discrete or continuous parameter for primitive . The reconstructed action is
The ordered word is essential because context actions need not commute. A stochastic mixture of such words is allowed when the intervention has unresolved random effects. Several primitives may be active in one raw control, so cross-role contexts arise without special treatment.
3.2. One Objective Chooses Reuse, Composition, Splitting, and Novelty
Let be a train–validation distribution over single interventions and composable intervention sequences. For each sampled context-action event , execute the reconstructed action and let be its predictive law for the held-out target in the transformed context. Define the context-action prequential code
This is a genuine codelength in nats, not a norm multiplied by a tunable conversion coefficient. Typed discrepancies between predicted bundles may still be reported for diagnosis or used as training surrogates, but final selection is based on the held-out code above. Sequential intervention data additionally impose the empirical composition constraint
with identity and inverse tests wherever those interventions are available. Let be the declared effective number of context-action evaluation events. Let be the total prequential code for all other declared events, evaluated with context-dependent predictions generated through Eq. (72), and excluding every model-description charge. The open-world discovery objective is
where K is itself optimized. The arguments of the first two terms are those given immediately above, and a hierarchical prefix code factorizes as
The last term charges the complete finite coding rule and the intervention-to-word assignments required by the training design; if words are transmitted event by event, it is the corresponding total rather than a mean length. The displayed prefix-code terms charge the base program, the reusable dictionary, and the intervention codes exactly once. The dictionary charge penalizes inventing synonymous primitives, and the intervention-code charge favors sparse reuse and short compositions. All terms are nats under a published event protocol and prefix code; there are no independent conversion weights. To prevent a merely correlational factorization, the residual is evaluated on the same interventions applied to new states and on independently composed sequences.
The fitted open-world theory is
This one comparison implements four outcomes. An unseen context intervention is represented by an existing primitive when that is shortest; by a composition when several known effects occur together; by splitting an old primitive when the shorter separate factors generalize better; or by a new primitive when its held-out predictive improvement exceeds its description cost. For example, the gain of a proposed new primitive is
It is retained only if a predeclared one-sided uncertainty bound for is positive and the conclusion is stable across the reported protocol sensitivity set. Conversely, two primitives are merged whenever replacing them by one does not worsen held-out fit enough to pay their separate codes.
The factorization is generally identifiable only up to renaming, reparameterization, and transformations among equally short observationally equivalent dictionaries. That is not a defect: no experiment can distinguish two context roles that induce the same accessible predictive action. Recovery of a more refined split additionally requires intervention coverage, counterfactual variation across states, an expressive program class, and a coding convention that charges redundant distinctions. The defensible universality claim is therefore conditional: under those conditions, Eq. (77) selects a shortest reusable factorization of every context effect that is operationally distinguishable from the available data, including effects not named in advance.
3.3. The Familiar Context Names Reappear as Action Signatures
After optimization, scientific names may be attached to recurrent signatures of the learned primitives rather than imposed beforehand. A primitive is nuisance-like when it changes raw observations but leaves every active conditional risk functional invariant; measurement-like when it changes a readout kernel while preserving the pre-measurement quotient and autonomous update; anchor-like when adding or erasing it refines or coarsens the quotient; purpose-like when it changes the evaluated reports, scores, or weights rather than physical dynamics; and scale-like when it maps between nested descriptions by a directed, generally noninvertible composition law. These signatures overlap. One primitive may have several signatures, and one raw intervention may activate several primitives. A stable learned action that matches none of them remains as an unnamed new role rather than being forced into the five-way vocabulary.
This resolves the cross-category issue directly. For example, optical zoom may compose (i) a projective readout action that changes focal length and point-spread function, (ii) a scale action that changes which angular structures are resolvable, and (iii) an anchor-revealing action when a previously unresolved reference star becomes available. The split is accepted only if applying these factors separately in held-out experiments predicts their recombination more compactly than a single indivisible “zoom” token.
3.4. Solar-System Active Observation as a Feasibility Case
Consider a virtual telescope whose raw context contains time, site, pointing, field of view or zoom, exposure, spectral band, and optional short revisit cadence. The observation is an image; the AI is not given labels “planet”, “star”, “background”, or “camera rotation”. This is a specialization of Eqs. (20)–(22): the latent domain is the celestial sphere plus finite-distance bodies, the local support is the visible spherical cap, and the kernel contains projection, optics, sampling, masks, and noise. The finite footprint of the image must therefore be part of the observation model. For an ideal circular detector let
where a soft detector visibility may instead describe vignetting, occultation, dead pixels, or an arbitrary detector mask. The continuum notation is an optical idealization; on a digital detector the same expressions are restricted to the pixel lattice and all image integrals become weighted sums. Let
be the calibrated projection of the visible spherical cap , with the attitude and the focal-length or zoom parameter. For an ideal pinhole camera the angular half-width is ; the pupil aperture primarily enters the point-spread kernel below, whereas the field stop and detector footprint determine . On a field smaller than a hemisphere an ideal pinhole is one-to-one; a more general optical model may be a calibrated kernel rather than a bijection.
A finite-distance body with inertial position and observer position has apparent inertial direction
Thus site and time enter the direction before camera attitude and projection, making parallax representable rather than merely listed as a diagnostic. A compact candidate generator may use a latent radiance measure
containing diffuse sky, nearly fixed distant sources, and possibly moving finite-distance sources. The finite-frame observation is then
Here is detector-fixed background, is the point-spread/readout kernel, and is residual noise. The sky field, source and observer positions, fluxes, attitude, projection, aperture, and detector terms are hypotheses selected by compression, not labels supplied to the learner.
Repeated pointings and camera rotations can identify a shared action on directions, while multi-zoom pairs constrain both and the angular point-spread function. Crucially, this does not give a global image-to-image map; it instantiates Eq. (21). For , the stationary-sky correspondence back to frame t is the partial map
defined only on the effective overlap
Even for two circular fields, is generally a clipped lens-shaped set after a repointing or zoom; pixels in show newly exposed sky and have no antecedent in . Let be zero where interpolation or either common-resolution kernel lacks valid boundary support and one elsewhere, with a soft transition allowed. The usable overlap weight, including masks in both frames, is
Let and denote calibrated photometric operators that map both frames to a common angular resolution no finer than either available point-spread function. They may smooth a sharp frame, but must not manufacture frequencies destroyed by the coarser frame. The counter-warped residual between two observations can then be written on the overlap as
where and correct learned multiplicative and additive photometry and is the predictive noise scale. A comparable scalar score must normalize by usable support. For , for example,
so that a small overlap is not rewarded merely for containing fewer pixels. Equation (87) is therefore not literal global background subtraction. It is the overlap-supported diagnostic derived from posterior prediction under Eq. (83); the full likelihood still predicts from the learned sky prior and earlier frames. Stationary sky and repeatable detector structure are cheaply explained by the projection, mask, and background terms, whereas sources with coherent residual motion require additional persistent factors. The direction action may be reversible on , but cropping, masking, sampling, and loss of resolution make the induced image action partial or noninvertible. It consequently belongs to the open context-action framework of Section 3, not automatically to the invertible observer groupoid of Eq. (16).
A planet candidate is retained when a moving-source slot plus a low-dimensional trajectory law reduces the total code relative to a star-only explanation. Useful diagnostic evidence includes nonzero motion after attitude compensation, observer-dependent parallax, a point-spread profile consistent across zoom, and brightness or phase behavior that is reproducible across bands and epochs. None is individually a definition of “planet”. Together they enter the matched object-gain comparison in Eq. (31). A distant star is the limiting factor whose direction is approximately fixed over the observing baseline and whose residual is explained by shared camera motion; an artifact fails cross-frame geometric and photometric consistency; diffuse backgrounds are marginalized unless their structure improves an active channel.
The next context is selected to distinguish competing explanations, not merely to maximize immediate image quality. Let denote the posterior model index and latent structure, including star, moving body, artifact, trajectory family, context dictionary, and calibration parameters. An acquisition rule specialized from Eq. (69) is
Wide fields are valuable for discovery and attitude calibration; narrow fields for centroid and morphology; short revisits for distinguishing apparent motion from noise; separated observer locations or times for parallax; and multiple bands for separating achromatic geometric motion from wavelength-dependent artifacts. The choice of zoom is therefore learned through its expected reduction of model uncertainty, its role in context-action composition, and the overlap–resolution trade-off made explicit by Eqs. (85)–(88).
If the current candidate class contains finite-field projective camera models, flexible sky and detector backgrounds, persistent object slots, and sufficiently rich trajectory programs, this case is representable by Eqs. (75)–(89). The framework does not guarantee that finite images uniquely reveal the real Solar System. Exact recovery can fail through insufficient angular resolution, short baselines, occlusion, saturation, indistinguishable trajectories, unmodeled optical distortions, poor intervention coverage, or an inadequate program language. What the theory supplies is a complete selection principle over the supplied program class: infer aperture and projection, calibrate partial observer correspondences, explain reusable background, invent persistent moving factors, choose informative views, and introduce a new context primitive whenever doing so wins held-out predictive compression. The celestial construction is a feasibility argument, not empirical validation: it shows that finite fields, nonoverlap, resolution loss, background prediction, moving-source hypotheses, and active zoom can be expressed inside one objective, while leaving their practical learnability to future experiments.
4. Conclusion
We have proposed a context-first operating framework for future autonomous AI agents whose task is to make scientific discoveries. The starting point is not a prepared list of variables, objects, or coordinates, but active histories of observations and contexts. Candidate theories are externally supplied executable programs and are compared by one prequential MDL code: predictive log losses and model descriptions are both expressed in nats, event frequencies are fixed by a published evaluation protocol, inherited science enters through conditional prefix-code length, and covariance is tested through admissibility or paired predictive events. The resulting selection is therefore conditional rather than deductive, but its codelength gaps have an unambiguous unit and a finite-data interpretation.
The open-world extension sharpens the meaning of automatic context discovery. Raw interventions are represented by a learned dictionary of predictive actions, and one held-out MDL comparison chooses whether to reuse, compose, split, merge, or add primitives. Nuisance, measurement, anchor, purpose, and scale survive only as overlapping signatures that may be assigned afterward; an unseen effect need not fit any predeclared category. This does not give unconditional semantic uniqueness. The shortest factorization is identifiable only up to operational equivalence and coding symmetries, and finer splits require intervention coverage, counterfactual variation, adequate program expressivity, and an honest code for redundant distinctions.
The local-observation formalism extends beyond finite telescope frames. An experiment is a possibly stochastic and information-losing observation kernel supported on only part of a candidate latent phenomenon. A pointwise warp exists only when two deterministic charts are invertible on their shared support; otherwise the agent must compare conditional laws and generate what cannot be transported. Cropping, missing channels, finite time or frequency windows, censoring, saturation, resolution loss, and destructive measurement are all instances of this same structure. In the Solar-System specialization, bounded circular fields, projection, point-spread matching, rotation, zoom, parallax, background prediction, moving-object hypotheses, and active follow-up therefore fit one formalism rather than requiring an image-specific patch.
Computationally, the unrestricted program minimum is replaced by a finite anytime search. The agent alternates continuous fitting with typed structural edits, scores every surviving program on ordered prequential blocks, protects a fresh outer audit from adaptive proposal generation, and chooses new experiments by disagreement among the retained programs. This makes the proposal executable as a search protocol while keeping its limitation visible: it can select only among structures that its grammar, budget, data, and optimizer actually reach.
The limitations are substantial. This article contains no empirical validation, and the Solar-System construction is only a feasibility argument. The framework does not guarantee a global optimum, the correct candidate language, a calibrated semantic verifier, sufficient resolution, adequate intervention coverage, or identifiable context factors from finite data. Nor does it derive quantum mechanics, field theory, or renormalization from an optimization symbol: operational, Hilbert, sector, and scale-dependent descriptions remain candidate classes whose physical assumptions and empirical adequacy must be supplied and tested. Envariance is only one debated route to the Born rule, while Gleason-type results offer an independent route under different assumptions.
The intended reader and user is therefore primarily a future scientific-discovery AI agent, not merely a human seeking a retrospective philosophy of science. Such an agent should learn not only laws on prepared variables, but also what to observe next, which distinctions to retain, where observations can be transported, and which variables, objects, context factors, states, and representation classes are warranted by controlled prediction. Overall, this work is intended to serve as a comprehensive operational handbook and practical guide for future autonomous AI agents conducting scientific discovery.
4.0.0.5. Acknowledgement
We gratefully acknowledge financial support from the National Natural Science Foundation of China (52394272), the National Key Research and Development Program of China (2023YFA0915300), and the Shanghai Science and Technology Innovation Action Plan (24JD1400700).
Appendix A. Supporting Mathematical Statements
The results collected here are retained for precision and convenient reference. They are self-contained specializations of standard mathematics—the discrete Noether calculation, the homogeneous-space orbit theorem, and the elementary criterion for descent through a quotient—rather than the principal mathematical innovations of this paper. The paper’s main contribution is the operational synthesis and agent-facing model-comparison framework in the main text. Accordingly, none of the statements below should be read as deriving a physical representation without the additional candidate-class and empirical assumptions stated there.
Appendix A.1. AI–Noether Theorem and Retention Criterion
The AI–Noether result used in the main text is a specialization of the standard discrete Noether argument and is not intended to rederive physical energy or momentum by name. It gives a representation-learning statement: if the local predictive code has a continuous symmetry, then a stationary predictive program carries a conjugate statistic that cannot drift along the corresponding trajectory. A second, model-selection statement then says when this statistic is worth naming and retaining inside the learned theory. In implementation, this theorem should be read as an algorithmic test rather than as a decorative analogy with mechanics: search the learned attribute manifold for continuous directions that leave the predictive code invariant; compute the corresponding conjugate statistic; measure its drift under new contexts and interventions; and retain it only if the reduction in predictive loss exceeds the cost of adding the statistic to the code. Large drift is informative rather than merely a failure, because it points to hidden forcing, an unmodeled environment, a wrong observer transformation, or an overly coarse predictive quotient. Together, the two statements make symmetry operational for AI discovery: symmetry proposes a conserved latent, experiments try to break it, and the global objective decides whether that latent belongs in the shortest scientific program.
Theorem A1
(AI–Noether theorem: continuous code symmetry implies a conserved conjugate statistic). Let be an orbit-type attribute extracted from the predictive quotient , where Q is a finite-dimensional smooth attribute manifold. Let be a one-parameter transformation group with infinitesimal generator
For a trajectory of length H, define the path objective
where denotes all retained latent variables other than the orbit coordinate. Assume:
- 1.
- Regularity.The local code is in its two orbit-coordinate arguments, and the variations below are taken in local charts, equivalently as arbitrary tangent variations of the interior .
- 2.
- Attribute retention.On the accessible context family, has positive attribute gain and is therefore retained as a predictive coordinate by the objective in Eq. (13).
- 3.
-
Continuous code symmetry.The local predictive code is invariant under the simultaneous transformation of consecutive orbit coordinates:for all sufficiently small s.
- 4.
- Closed homogeneous regime.No explicit anchor, measurement, purpose, or scale term breaks this symmetry along the considered trajectory segment.
- 5.
- Stationarity.With , , and the endpoints fixed, the path is a stationary point of under arbitrary compactly supported interior variations.
Let and denote the corresponding cotangent derivatives of . Define the conjugate code statistic
Then
Thus a continuous symmetry direction of the predictive code generates a time-conserved conjugate code statistic.
Proof.
Corollary A1
(Approximate AI–Noether diagnostic). In a learned or noisy implementation, define the stationarity residual and symmetry residual by
Equip the tangent bundle with a norm and the cotangent bundle with its dual norm. If , , and , then
Proof.
By the definition of and the residuals,
Taking absolute values gives Eq. (A10). □
This corollary is the practical reason to keep the theorem in a discovery framework. A proposed symmetry is testable: it should produce a low-drift latent statistic. If the drift is large, the agent has evidence for symmetry breaking, an unmodeled context, a wrong observer transformation, or an insufficient predictive quotient.
Proposition A1
(AI–Noether retention criterion). Assume the global objective remains Eq. (13), and consider a target evaluated by log loss with declared event frequency among the N evaluation events. Assume all relevant conditional laws are dominated and the conditional entropies below are finite. Compare two matched, sufficiently expressive model classes whose infima attain the corresponding true conditional laws and that differ only in whether they may retain a statistic in addition to another representation ; apart from their predictive code and their prefix-code lengths, all evaluation events and model capacities coincide. Suppose:
- 1.
- Trajectory constancy:
- 2.
- Predictive sufficiency:
- 3.
- Predictive relevance:on the accessible context family,
Then, at the population or oracle optimum under log loss, the increase in this target’s contribution to the total predictive code caused by deleting C is exactly
Let and be the shortest members of the with-C and without-C classes, respectively, among those attaining the corresponding oracle predictive risks assumed above, and define
as their prefix-code difference in bits. If
then
Hence the shortest oracle-risk achiever that deletes all predictive information supplied by C is worse than its matched retaining alternative. This is an exact population comparison of the stated pair; extending it to unrestricted classes requires checking their actual optimized objectives through Eq. (39) and the finite-data rule following Eq. (14). Since C is trajectory-constant, one may choose a representative latent update of the form
Proof.
Under log loss, the Bayes risk conditioned on a representation R is the corresponding conditional entropy; its total contribution to Eq. (13) is multiplied by the declared event count . If the representation keeps both and , the optimal risk is
If the representation deletes , the optimal risk becomes
Therefore the excess risk from removing is
The conditioning includes averaging over the accessible context–policy distribution, yielding Eq. (A14). If Eq. (A16) holds, then the predictive codelength benefit of retaining C exceeds its additional prefix-code cost in the same unit. Substitution into Eq. (13) gives Eq. (A17). Finally, because C is constant along trajectories, a retained representative explicitly equal to C may be implemented with the trivial update law . □
Appendix A.1.1.6. Relation between the theorem and the retention criterion.
Theorem A1 is the Noether-type statement: it converts a continuous symmetry of the local predictive code into a conserved conjugate statistic. Corollary A1 makes the statement experimentally useful for imperfect learned models. Proposition A1 answers a different question: once such a statistic exists, when does an exact matched oracle-risk comparison favor retention of its predictive information, possibly through a sufficient reparameterization? Within the present framework they motivate the following diagnostic sequence, not a new derivation of Noether’s theorem:
Appendix A.2. A Bifurcation Theorem for Elementary Attribute Organization
The following statement records the standard homogeneous-space fact used as a conditional routing test in the main text. After an AI has learned a smooth candidate attribute manifold and a Lie-group observer action, componentwise transitivity identifies each connected component with a homogeneous orbit while the component index is discrete. It does not claim that all scientific attributes fall into only two ontological kinds; singular, stratified, or nonhomogeneous learned spaces are explicitly outside its hypotheses.
Definition A1
(Componentwise homogeneous attribute manifold). Let G be a finite-dimensional Lie group, let denote its identity component, and let be a second-countable Hausdorff smooth manifold on which G acts smoothly. It is called componentwise homogeneousif:
- 1.
- every connected component of is invariant under ;
- 2.
- the restricted action of on every connected component is transitive.
Theorem A2
(Attribute bifurcation theorem). Let be a componentwise homogeneous attribute manifold under the smooth action of a Lie group G, and let be the identity component of G. Then:
- 1.
- for each connected component and each chosen representative , the stabilizer is a closed Lie subgroup and the orbit map induces a -equivariant diffeomorphism ;
- 2.
-
consequently, there is a decompositionwhere is the set of connected components of ;
- 3.
- if, in addition, the component space is treated with the quotient topology, it is discrete because connected components of a manifold are open. The attribute therefore separates into smooth within-component orbit variation and a discrete component label.
This is a routing result conditional on smooth componentwise homogeneity; it is not a classification of arbitrary attribute spaces.
Proof.
Connected components of a manifold are open, closed, and themselves smooth manifolds, so the unique component decomposition is a topological disjoint union
Fix a point , and let be its connected component. Because is connected and the orbit map
is continuous, the orbit is connected. Hence
By componentwise homogeneity, the restricted action of on is transitive, so in fact
Now let
be the stabilizer subgroup of q. Since is Hausdorff and the action is continuous, is closed in , hence is an embedded Lie subgroup by the closed-subgroup theorem. The standard homogeneous-space orbit theorem for smooth transitive Lie-group actions states that the orbit map factors through a unique -equivariant diffeomorphism
Choosing one representative point for each component gives the decomposition in Eq. (A19). Since the components are open, their images under the quotient map are open singletons; hence the component space is discrete. This proves the stated smooth-orbit/discrete-label routing. □
Appendix A.2.2.7. Interpretation for the main text.
Theorem A2 is used in Section 2.5 as a conditional attribute-routing criterion. When the learned value space is a smooth manifold and the connected observer action is transitive on each component, orbit-type attributes correspond to continuous coordinates inside one homogeneous component and should be modeled by equivariant charts, Lie generators, and invariants. Sector-type attributes correspond to discrete component labels and should be modeled by sector variables, counts, or transition events. Attribute spaces that are singular, stratified, fractal, or nonhomogeneous fall outside the theorem and require richer routing rather than being forced into its conclusion. If a learned continuous observer action appears to change a valid sector label, the agent should suspect a missing event, a wrong quotient, or an over-smoothed representation. The additional phenomena discussed in the main text—superposition under anchor loss, sector lift under measurement isolation, and running couplings under scale restriction—concern how these elementary organizations fail, survive, or are repaired once accessible contexts are contracted.
Appendix A.3. Classical-atlas Obstruction from Attribute Undefinability
The following theorem isolates the logical point that is often hidden inside informal phrases such as “isolation causes non-gluability”. The primary obstruction appears earlier: a candidate classical attribute may fail to be well defined on the actually accessible predictive quotient. Operationally, this is a descent test: before an AI tries to glue local classical charts, it should ask whether the attribute is constant on fibres of the accessible quotient. Non-gluability then follows for any classical atlas that still presupposes a non-descending attribute.
Theorem A3
(Classical-atlas obstruction from attribute undefinability). Let
be a surjective coarsening map from a richer descriptive regime to the actually accessible predictive quotient. Let
be a candidate classical attribute, and assume that there exist such that
Then the following two conclusions hold.
(i) Failure of descent.There does not exist any map
for which
Equivalently, A is not a well-defined attribute of the accessible theory.
(ii) Classical-atlas obstruction.Let be a jointly faithful family of readout maps, meaning that
Define the frozen pre-classical chart family on by
Then there do not exist a global attribute and accessible charts
for which
Hence no atlas of this jointly faithful frozen form that still requires the attribute A can be globally glued on the accessible theory. The conclusion does not rule out a different classical attribute, an unfaithful coarse readout, or another enlarged state representation.
Proof.
For part (i), suppose by contradiction that there exists satisfying Eq. (A21). Then whenever , one has
This contradicts Eq. (A20). Therefore no such exists.
For part (ii), assume that a glued accessible atlas of the form Eqs. (A24)–(A25) does exist. Then for every and every ,
Because the family is jointly faithful, Eq. (A22) implies
Hence , contradicting part (i). Therefore no such glued accessible atlas can exist. □
Appendix A.3.3.8. Interpretation.
The theorem unifies two cases. In the first, a richer leakage or anchor regime once made A definable, but after information suppression the coarsening map q merges states with different values of A; this is attribute death. In the second, A is a merely hypothetical classical refinement—for instance, a shape, size or internal form assigned to an object with no lower-level structure that could support it; this is better viewed as attribute non-birth. In both cases the logical order is the same: attribute undefinability ⟹ failure of descent ⟹ non-gluability of any classical atlas that presupposes the attribute.
Thus the theorem does not say that isolation by itself is already quantum. It says that once a putative classical attribute fails to descend to the actual accessible quotient, any jointly faithful frozen chart family that still requires that same attribute is obstructed; deleting the attribute, choosing another classical representation, or making a representation lift are all logically possible next comparisons. In practice, a persistent descent residual such as Eq. (46) is therefore not just a fitting error. It is a diagnostic telling the AI that the candidate variable should be deleted, marginalized, or replaced by a richer operational or scale-dependent description.
References
- Himanen, L.; Geurts, A.; Foster, A.S.; Rinke, P. Data-Driven Materials Science: Status, Challenges, and Perspectives. Adv. Sci. 2019, 6, 1900808. [Google Scholar] [CrossRef] [PubMed]
- Wang, H.; Fu, T.; Du, Y.; Gao, W.; Huang, K.; Liu, Z.; Zitnik, M.; et al. Scientific discovery in the age of artificial intelligence. Nature 2023, 620, 47–60. [Google Scholar] [CrossRef] [PubMed]
- Bran, A.M.; Cox, S.; Schilter, O.; Baldassari, C.; White, A.D.; Schwaller, P.; et al. Augmenting large language models with chemistry tools. Nat. Mach. Intell. 2024, 6, 525–535. [Google Scholar] [CrossRef] [PubMed]
- Liu, Z.; Chai, Y.; Li, J. Toward Automated Simulation Research Workflow through LLM Prompt Engineering Design. J. Chem. Inf. Model. 2025, 65, 114–124. [Google Scholar] [CrossRef] [PubMed]
- Ma, Q.; Zhou, Y.; Li, J. Automated Retrosynthesis Planning of Macromolecules Using Large Language Models and Knowledge Graphs. Macromol. Rapid Commun. 2025, e2500065. [Google Scholar] [CrossRef] [PubMed]
- Hutter, M. Universal artificial intelligence: Sequential decisions based on algorithmic probability; Springer Science & Business Media, 2004. [Google Scholar]
- Friston, K. The free-energy principle: a unified brain theory? Nat. Rev. Neurosci. 2010, 11, 127–138. [Google Scholar] [CrossRef] [PubMed]
- Ha, D.; Schmidhuber, J. World Models, 2018. arXiv arXiv:1803.10122.
- LeCun, Y. A Path Towards Autonomous Machine Intelligence, 2022; OpenReview; version 0.9.2.
- Hafner, D.; Pasukonis, J.; Ba, J.; Lillicrap, T. Mastering diverse control tasks through world models. Nature 2025, 640, 647–653. [Google Scholar] [CrossRef] [PubMed]
- Brunton, S.L.; Proctor, J.L.; Kutz, J.N. Discovering governing equations from data by sparse identification of nonlinear dynamical systems. Proc. Natl. Acad. Sci. 2016, 113, 3932–3937. [Google Scholar] [CrossRef] [PubMed]
- Udrescu, S.M.; Tegmark, M. AI Feynman: A physics-inspired method for symbolic regression. Sci. Adv. 2020, 6, eaay2631. [Google Scholar] [CrossRef] [PubMed]
- Udrescu, S.M.; Tan, A.; Feng, J.; Neto, O.; Wu, T.; Tegmark, M. AI Feynman 2.0: Pareto-optimal symbolic regression exploiting graph modularity. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2020; Vol. 33. [Google Scholar]
- Wu, T.; Tegmark, M. Toward an artificial intelligence physicist for unsupervised learning. Phys. Rev. E 2019, 100, 033311. [Google Scholar] [CrossRef] [PubMed]
- Raissi, M.; Perdikaris, P.; Karniadakis, G.E. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. J. Comput. Phys. 2019, 378, 686–707. [Google Scholar] [CrossRef]
- Abramsky, S.; Brandenburger, A. The Sheaf-Theoretic Structure of Non-Locality and Contextuality. New J. Phys. 2011, 13, 113036. [Google Scholar] [CrossRef]
- Hardy, L. Quantum Theory From Five Reasonable Axioms, 2001. arXiv arXiv:quant.
- Chiribella, G.; D’Ariano, G.M.; Perinotti, P. Informational Derivation of Quantum Theory. Phys. Rev. A 2011, 84, 012311. [Google Scholar] [CrossRef]
- Wilson, K.G.; Kogut, J. The Renormalization Group and the ϵ Expansion. Phys. Rep. 1974, 12, 75–199. [Google Scholar] [CrossRef]
- Polchinski, J. Renormalization and Effective Lagrangians. Nucl. Phys. B 1984, 231, 269–295. [Google Scholar] [CrossRef]
- Penco, R. An Introduction to Effective Field Theories, 2020. lecture notes for the 2nd Joint ICTP-Trieste/ICTP-SAIFR School on Particle Physics.
- Rissanen, J. Modeling by shortest data description. Automatica 1978, 14, 465–471. [Google Scholar] [CrossRef]
- Grünwald, P.D. The Minimum Description Length Principle; MIT Press, 2007. [Google Scholar]
- Dragan, A.; Ekert, A. Quantum principle of relativity. New J. Phys. 2020, 22, 033038. [Google Scholar] [CrossRef]
- Shalizi, C.R.; Crutchfield, J.P. Computational Mechanics: Pattern and Prediction, Structure and Simplicity. J. Stat. Phys. 2001, 104, 817–879. [Google Scholar] [CrossRef]
- Littman, M.L.; Sutton, R.S.; Singh, S. Predictive Representations of State. In Proceedings of the Advances in Neural Information Processing Systems 14; Dietterich, T.G., Becker, S., Ghahramani, Z., Eds.; MIT Press, 2002; pp. 1555–1561. [Google Scholar]
- Barnett, N.; Crutchfield, J.P. Computational Mechanics of Input-Output Processes: Structured Transformations and the ϵ-Transducer. J. Stat. Phys. 2015, 161, 404–451. [Google Scholar] [CrossRef]
- Barrett, J. Information processing in generalized probabilistic theories. Phys. Rev. A 2007, 75, 032304. [Google Scholar] [CrossRef]
- Janotta, P.; Hinrichsen, H. Generalized Probability Theories: What determines the structure of quantum theory? J. Phys. A Math. Theor. 2014, 47, 323001. [Google Scholar] [CrossRef]
- Masanes, L.; Müller, M.P. A derivation of quantum theory from physical requirements. New J. Phys. 2011, 13, 063001. [Google Scholar] [CrossRef]
- Barnum, H.; Müller, M.P.; Ududec, C. Higher-order interference and single-system postulates characterizing quantum theory. New J. Phys. 2014, 16, 123029. [Google Scholar] [CrossRef]
- Wigner, E.P. Group Theory and Its Application to the Quantum Mechanics of Atomic Spectra; Academic Press: New York, 1959. [Google Scholar]
- Chiribella, G.; Aurell, E.; Życzkowski, K. Symmetries of quantum evolutions. Phys. Rev. Res. 2021, 3, 033028. [Google Scholar] [CrossRef]
- Zurek, W.H. Environment-Assisted Invariance, Entanglement, and Probabilities in Quantum Physics. Phys. Rev. Lett. 2003, 90, 120404. [Google Scholar] [CrossRef] [PubMed]
- Zurek, W.H. Probabilities from entanglement, Born’s rule. Phys. Rev. A 2005, 71, 052105. [Google Scholar] [CrossRef]
- Gleason, A.M. Measures on the Closed Subspaces of a Hilbert Space. J. Math. Mech. 1957, 6, 885–893. [Google Scholar] [CrossRef]
- Busch, P. Quantum States and Generalized Observables: A Simple Proof of Gleason’s Theorem. Phys. Rev. Lett. 2003, 91, 120403. [Google Scholar] [CrossRef] [PubMed]
- MacKay, D.J.C. Information-Based Objective Functions for Active Data Selection. Neural Comput. 1992, 4, 590–604. [Google Scholar] [CrossRef]
- Chaloner, K.; Verdinelli, I. Bayesian Experimental Design: A Review. Stat. Sci. 1995, 10, 273–304. [Google Scholar] [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.