Submitted:
29 September 2026
Posted:
01 October 2026
You are already at the latest version
Abstract
Scientific discovery seeks regularities that support explanation and prediction. Compression makes their reuse explicit: shared structure is described once, while parameters specify individual cases. A model can therefore pay for itself through repeated use even when its description is not minimal, provided it captures reusable structure in the data. Because regularities are often easier to identify in simpler systems, a reductionist approach is often employed: discover the laws of the parts and treat them as fundamental. Yet knowing those laws does not by itself provide useful coarse-grained models of the larger systems they compose. Here we formulate Anderson’s distinction between reduction and construction for finite algorithmic observers and prove three barriers to such construction. First, an observer’s coarse-grained record may retain so much information about initial or boundary conditions that no substantially shorter description exists, even when the underlying laws are simple and known. Second, when a shorter description does exist, open-ended search can eventually find one, but there is no computable bound on how long this may take, and no algorithm that always halts with a valid description can succeed in every case. Third, every such algorithm has blind spots at all sufficiently large lengths: records it leaves unshortened even though they admit descriptions of only logarithmic length. Yet observers do discover useful models. We call the event in which an observer acquires a representation relevant to its objective and a reusable model that reveals previously unavailable regularity algorithmic emergence. We give a sufficient certificate in explicit code lengths: the pair passes when, with its own description cost included, it compresses the retained data relative to an agreed baseline and yields further savings on later observations. Even for a stream that repeats one block, no computable procedure that always halts with valid codes can guarantee finding a passing pair with the observations available whenever one exists, nor remain within a fixed number of bits of the best qualifying complete code. Favorable structure, including symmetry and restricted model classes, can nevertheless make discovery feasible.

Keywords:
emergence
; regularity
; compression
; algorithmic information theory
; Kolmogorov complexity
; model discovery
; cognition
; coarse-graining
; uncomputability
; algorithmic statistics
MSC: 68Q30; 94A15
1. Introduction: Reduction Is Not Construction
Scientific discovery seeks regularities that support succinct explanation and prediction across cases—compression and generalization. A model captures what the cases share; its parameters specify what changes from one case to another. The parameter sequence can itself have regularity, such as a recurring movement. Together, the model and jointly encoded parameters generate the modeled observations. They provide a compressed description when the cost of the model, the parameters, and whatever the model leaves unexplained is less than that of describing each case separately. The model’s fixed cost is paid once, so even a large implementation can yield savings through reuse. If the regularity persists in new observations, the same model supports prediction.
This role of compression in science and cognition is explicit in the Kolmogorov Manifesto, in the minimum message length and minimum description length principles, and in Li and Vitányi’s treatment of induction as compression [1,2,3,4,5]. Schmidhuber connects discovery to improvements in an agent’s ability to compress its observations, and recent work on language models develops the relation between prediction and lossless coding [6,7]. Prediction and compression are two uses of one regularity: under logarithmic loss a predictor’s loss is an ideal code length, and a deterministic predictor compresses a binary record by transmitting only the positions of its mistakes, provided these are few (Appendix B). Compression gives evidence that a model captures structure; reuse on new records tests whether that structure generalizes.
There is also a reason to expect short descriptions to generalize. Solomonoff’s universal predictor weights each program p by . For any computable data-generating source, this predictor’s expected cumulative excess logarithmic loss relative to that source is bounded by the source’s description complexity, up to fixed coding constants [8,9,10]. This is a predictive bound within the class of computable sources, not a proof that Nature is simple. Earlier KT work considered a simplicity bias in the structures encountered in the world [11], and related results show how such a bias arises in broad classes of low-complexity computable input–output maps [12].
Models guide the actions of agents in pursuit of their objectives. We write for the objective function that determines what matters to the observer; Figure 2 locates the observer, its prior resources, and the representation–model pair within this setup. One important class is telehomeostatic agency, in which action is organized around regulation that supports continued persistence—of the agent, its lineage, or a larger system to which it contributes [13]. Modeling is instrumental rather than an end in itself: successful regulation requires the agent to capture reusable regularities in its world and in its own dynamics [1].
Among models that are comparably adequate, shorter effective descriptions have two advantages. They provide stronger inductive grounds for expecting the captured regularity to generalize, and they reduce the cost of storing and transmitting the model. But description length is not the only resource that matters. A longer implementation of the same functional regularity may be preferable if it produces predictions faster, uses less working memory, or requires less energy to run. Thus compression is evidence that reusable regularity has been captured, not the ultimate objective of the agent. None of these considerations, however, supplies a general procedure for discovering such regularities.
Furthermore, while compression is necessary for modeling, it is not sufficient. A model that captures reusable regularity can compress the observations it applies to through reuse, even when it is imperfect or far from minimal; its description cost need not yet have been repaid. The converse fails. A compressor of past data does not hand over a model: a short program for one record need not expose which part is shared structure and which part is the case, how the case is parametrized, or what to send when the next observation arrives. This is why the paper measures acquisition by compression and tests it by reuse. Reuse on new observations, with the same coordinates in the same roles, is what the certificate requires beyond a shorter code for the past.
Regularities are often easier to find in simpler systems. This is one motivation for reductionism: isolate smaller parts or fewer degrees of freedom, discover the laws that govern them, and treat those laws as the foundation for understanding composite systems. A microscopic theory is itself a compression success, often obtained in such a restricted setting. The question is whether that success transfers upward. Given the microscopic laws, the scientist would still need a procedure for discovering the theories of the larger systems those parts compose. In More Is Different, Anderson denied that the laws supply one:
“The ability to reduce everything to simple fundamental laws does not imply the ability to start from those laws and reconstruct the universe.” [14]
Here, we take “reconstructing” to mean finding the higher-level variables, organizing principles, and effective theories that make the behavior of a larger system intelligible—in other words, compressible. Compression is the part of intelligibility this paper measures; explanatory or causal adequacy requires further conditions (Section 4.2).
Given a microscopic law and the information specifying a particular case, direct simulation can generate its observed history. We call this generative reduction. The stronger constructionist converse would provide a useful higher-level representation and reusable model. A simulated trajectory does not by itself supply these. A tornado, for example, is understood through variables such as location, rotation, intensity, and expected evolution, rather than a list of every molecule. Whether the resulting macromodel also provides a computational shortcut is a separate question.
Throughout, “microscopic” and “macroscopic” are relative terms. A microscopic description is a set of variables together with a law that updates them; a macroscopic description is another set of variables obtained from the first by a projection. Nothing here requires the microscopic level to be fundamental physics, and nothing requires the projection to be an average. A cell’s chemistry is microscopic for its physiology, and a phase label or a single halting bit can be a macroscopic variable. Any level can play the microscopic role for the level built on it.
The choice of projection also fixes what the observer must then model. Under reversible microscopic dynamics the complete-state algorithmic information is conserved once the law and the elapsed time are supplied, but information held initially in the omitted detail can later reach the observed variables. A projection that makes one snapshot simple can therefore leave a complicated history, or discard the distinctions needed to predict how the retained variables evolve. Section 6 gives this accounting and its consequences for choosing .
Within Kolmogorov Theory (KT), we take algorithmic emergence to be an event in an observer’s descriptive history: the acquisition of a useful, reusable model under which the observer’s retained observations admit a shorter description than it could give before. A pattern is what such a model captures; emergence lies in its discovery, without changing the complexity of the fixed record.
This view aligns with Bédard and Bergeron’s account of emergence in theory space, where new partial models yield discrete gains in description [15]. Our emphasis is on the observer’s acquisition of a usable model and on whether a procedure can guarantee such a transition. The representation may already be supplied; the acquisition concerns a newly available description.
We study a finite algorithmic version of this construction problem:
- Q1.
- Existence. Does the specified observed history admit a substantial lossless shortening relative to its literal code?
- Q2.
- Discovery. If any shortening exists, can one general terminating procedure always find a valid one?
- Q3.
- Near-optimality. Can such a procedure always stay within a fixed number of bits of the shortest description?
Three barriers answer these questions. Simple known dynamics can preserve almost all the initial-condition information in the observed history, so no substantial shortening exists (the residual-information barrier, Theorem 1). No algorithm that always halts with a valid description finds a shortening whenever one exists (the discovery barrier, Theorem 2). No such algorithm stays within a fixed number of bits of the shortest description (the optimality barrier, Theorem 3), and each leaves records with logarithmic-length descriptions unshortened at every sufficiently large length (Proposition 1).
The same barriers reach model discovery. Suppose the observations simply repeat one block, and a constructor must return a representation and a reusable model from the first copy. Every such constructor leaves some block unshortened; that block nevertheless has a short description, one that names it through the constructor’s own failure and also predicts its repetitions. A qualifying pair therefore exists that the constructor, with the observations available, cannot return (Proposition 2).
The existence claim concerns the lossless shortening of the specified history; it does not settle whether a model can capture some regularity or support a narrower predictive task. Figure 1 summarizes the three questions and their answers.
The barriers limit what any one method can guarantee, yet agents do discover reusable models. Such a discovery can carry the surprise often associated with emergence: the observer recognizes coherent organization that was not apparent in the microscopic description. A Game of Life glider crossing a cellular automaton and the coordinated motion of a flock are familiar examples [16,17]; each invites a description in terms of collective entities and their behavior. Surprise is an explicit criterion in the emergence test of Ronald, Sipper, and Capcarrère [18].
For an isolated glider, a recurring shape and a translation rule describe many successive configurations through position, orientation, and phase; Section 5.3 works this case through. The model need not be optimal, and its description cost may be repaid only through repeated use. This connects the recognition of emergent organization to the compression progress associated with discovery [1,6]. Scientific discovery and perceptual learning are instances; evolution can also build useful modeling capacity into organisms. Surprise motivates the definition without entering it. Acquisition changes the descriptions available to the observer, not the complexity of a fixed record under a fixed machine.
Definition 3 gives a sufficient certificate using explicit code lengths. A representation and model pass the test when, with their own description cost counted, they describe the data they keep in fewer bits than a baseline fixed in advance, and do so again on observations obtained later. Savings may arise across observations even when individual chunks cannot be shortened. Passing certifies reusable regularity. Failing does not rule it out, because a model can be useful before its description cost has been repaid.
The classical noncomputability results for Kolmogorov complexity are known. The contribution here is to locate them within the finite micro-to-macro problem and to carry them to the acquisition of a representation and reusable model. Problem 1 asks which effective restrictions allow one terminating method to find a qualifying pair whenever one exists, and Section 7 shows how Curie–Weiss symmetry, renormalization, hydrodynamic assumptions, and the Montbrió reduction constrain that search in physical systems. The record-level reductions are machine-checked in Lean 4; the model-acquisition proposition is proved at paper level (Appendix F).
Section 2 fixes the observer-relative setup and Section 3 proves the three barriers. Section 4 casts representation choice as a retained–residual split, and Section 5 gives the acquisition certificate and transfers the barriers to it. Section 6 shows how conserved information moves between the retained and the omitted coordinates, and what that costs a badly chosen representation. Section 7 examines the physical structures that make discovery feasible, and the Discussion separates these limits from answer-side undecidability. Appendix A collects the notation, and Appendix B develops the relation between compression and prediction.
2. Observer-Relative Setup
Fix a universal prefix-free Turing machine U, a finite observer-agent , and a shared description frame containing its coding conventions and admissibility criteria. Write for the fixed observer/description frame and reserve for a computable projection or coarse-graining. “Finite” means that the observer has finite physical state and computational resources at every finite time; no theorem depends on a particular architecture. Unless shown explicitly, complexities are relative to the fixed U and C, with their cost absorbed into additive constants. Tuples inside K use a fixed computable self-delimiting pairing in . Conditioning on this frame is analogous to specifying the information available to Jaynes’s idealized reasoning robot [19,20].
We use observer and agent for the same system: agent when its actions are in view, observer when only its modeling is. It is the algorithmic agent of Kolmogorov Theory [21,22], which has a modeling engine that builds its world-model, an objective function, and a planning engine that selects actions; only the first two enter the results below.
The observer’s objective function specifies what it seeks to minimize. In a parametrized model it may depend on the model and its changing parameters; more generally it evaluates the program M in a declared context. The objective and its interpretation belong to the fixed setup. A relevance requirement translates that objective into distinctions the representation must preserve, at a declared accuracy, to evaluate or pursue it. A concrete task specifies the required inputs and outputs; equal objective values need not mean equal task behavior.
We model the observer’s construction procedures as Turing-computable, with resources finite at every finite time but no fixed memory bound across instances. Finiteness alone does not justify the computability restriction. Gandy’s analysis of discrete deterministic mechanisms shows that a device assembled from bounded-size parts within a bounded structural hierarchy, whose transitions depend only on causal neighborhoods of bounded size, computes only recursive functions [23]. Its physical assumptions are a lower bound on the size of a part and an upper bound on the speed of propagation of changes. An observer with a halting oracle, or an infinite-time machine [24], falls outside this class and can compute K. Theorems 2 and 3 place no restriction on it, while Theorem 1 applies unchanged because counting involves no computation. For ordinary oracle Turing machines, the computational barriers recur for description complexity defined relative to the same fixed oracle.
Let . Let denote the set of finite binary strings and its nonempty subset. For ,
where p is a self-delimiting program, and denote bit lengths, and denotes execution with y supplied as auxiliary input. Equation (1) is the standard classical numerical notation; none of the operational claims requires a finite observer to possess or compute the exact value of K, and Appendix C gives a formulation directly in terms of description relations. For finite sets, denotes cardinality. All logarithms are base two. The mutual algorithmic information of two strings is , with the complexity of their fixed computable pairing. When common side information z is displayed, we use . These symmetric forms agree with the usual directional definitions up to additive terms.
Unless stated otherwise, and terms may depend on U, , and fixed coding conventions, but not on the varying instance.
A procedure is total computable if it halts on every admissible input, and uniform when the same procedure must work at every input size. A predicate is semidecidable when positive instances are eventually found although negative searches may run forever; dovetailing interleaves candidate computations so every halting one is eventually seen. We call x compressible relative to its raw code when . When a procedure returns a standalone program p for a record x, its additive description regret is . This is a static gap in bits. The benchmark is the bit length of the data itself; the proofs below compare program length against it.
Let denote the substrate state at time t, and let denote its finite history from a fixed starting time. The observer sees through a computable projection acting at each time. Its observation history is . Write for the Modeling Engine implemented by observer . It constructs the world-model
A candidate macroscopic submodel , indexed by j, is a component of this structured world-model. Figure 2 shows these levels. We reserve calligraphic for the full world-model and use M for a candidate parametrized macromodel.
A projection is admissible when it belongs to a class specified under frame C and preserves the distinctions required by the objective. It cannot erase a target distinction that the declared task needs. Stored patterns and lookup tables are allowed, provided their acquired information is charged; unseen observations are not available as free side information. Admissibility is a modeling restriction, not a hypothesis needed by the record-level barrier theorems: those results require only that the displayed be computable. In joint acquisition, may be inherited, supplied, or learned together with M; Section 4.2 explains the relevance constraint. Here specifies retained distinctions, not whether the observer acts on the system. A specified computable observer–world interaction can be included in the dynamics; any additional action information supplied to the decoder must enter the coding account.
A note on “coarse-graining.” We use the term in an observer-relative informational sense, broader than spatial averaging or the elimination of short-wavelength degrees of freedom. A coarse-graining is any computable representation that treats some distinctions in the substrate as irrelevant to the observer’s stated task while retaining others. A block average, an order parameter, a phase label, an object category, or a binary property such as “the device has halted” can all be coarse descriptions in this sense. What makes a variable macroscopic here is therefore not necessarily spatial scale: it is that the observer uses it as a higher-level description that ignores lower-level distinctions. Relevance is essential, since erasing all distinctions would produce maximal compression and no useful model.
To separate microscopic generation from macromodel construction, one must first specify the finite object that contains everything needed to simulate a single observed history.
Definition 1
(Finite micro-experiment). Afinite micro-experimentcomprises an update rule, an initial state, an observation map, and a finite horizon. We write it as . Let Σ be a finite microstate alphabet and N the number of microscopic degrees of freedom. The computable map updates the initial state , with . The computable map gives observations in a finite readout alphabet Ξ, and is the observation horizon, so the record contains observations. Itsmacrohistoryis
encoded as a binary string using the fixed computable convention in ; for a fixed-width binary readout alphabet , this convention writes each r-bit block literally in temporal order (block boundaries are implicit in μ). In particular, for the macrohistory is literally the readout bit sequence, and a single full-state readout from is encoded by its n bits.
The notation distinguishes states from records. The finite experiment encodes the conceptual substrate state as , and its observed history as the binary record . In the general modeling problem, D denotes a record before representation and the retained record; there may act on the whole record. Section 6 uses for a complete encoded state and for its projected state. These are instantaneous states, rather than history encodings.
The tuple is a convenient finite initial-value formal class for the barrier results, not a claim that fundamental physics must provide freely specifiable initial conditions. SubSection 3.5 rewrites the same information accounting in terms of boundary or global prescriptions and residual realization identifiers.
Complete knowledge of the experiment yields its macrohistory by direct simulation, giving the upper bound
and, with the common mechanism supplied as side information,
These are the generation statements announced in the introduction: microscopic determination supplies a generative route and an upper bound on description length. The theorems of Section 3 ask what stronger conclusion follows.
A description of x is a standalone program p with . “Standalone” means that the auxiliary tape is empty: any side information must either be encoded in p or be made explicit through a conditional complexity . Two baselines are distinguished. The residual-information theorem conditions on the common law, observation map, and horizon and asks how much record-specific information remains. The discovery and optimality theorems give the procedure the complete experiment but require it to return a standalone program for . Appendix D shows separately that conditioning on the known laws does not restore computability.
Each experiment has a finite state space and horizon, and produces a finite record. The computational theorems concern a single procedure required to work as these bounds grow without a fixed maximum. On a fixed finite domain, answers could in principle be tabulated. The theorems concern lossless descriptions in an unrestricted program class; restricted predictive, causal, or intervention-based classes require their own analysis.
Simulation generates a record from the experiment. Compression seeks a shorter description of that record. Reusable model construction also selects a relevant representation and captures structure shared across records. The following sections develop these requirements in turn.
3. The Algorithmic Construction Barrier
3.1. Simple Microlaws Need Not Reduce Record-Specific Information
The first barrier does not rely on universal dynamics. It shows that simple microscopic access does not guarantee that the observer-visible history has a substantial compression gap relative to its raw code. Case-specific information can survive the observation process so completely that little of the raw length can be saved, even when the update rule and the instantaneous projection are both simple.
Theorem 1
(Residual-information barrier). For each integer , let be the cyclic left shift on the n-cell binary register and let read the leading cell, so that . Write for the micro-experiment with initial state and horizon ; then . For every integer d with ,
Hence at least a fraction of the n-bit initial microstates produce macrohistories satisfying .
Proof.
Writing and indexing readout steps by , under the register content shifts one cell per step and outputs the cell now leading. The readout at step t is therefore and the n-step record is exactly y. Fix the conditioning string . Conditional descriptions are programs for a fixed machine with z on the auxiliary tape, and there are fewer than programs of length less than m; taking gives fewer than strings y with . Since the map is the identity, (6) follows, and the complement has size at least . □
Interpretation. The shift register provides a minimal witness. For , is pointwise many-to-one: each observation retains only one bit of an n-bit state. Yet the omitted bits are not lost from the history; the reversible shift feeds them into later observations until the complete initial state has been exposed. The example therefore separates pointwise lossiness from history-level lossiness. It establishes a non-guarantee, not a model of a typical natural system.
Under the uniform counting measure, at least a fraction of records retain bits of conditional complexity despite the known law, projection, and horizon; for this exceeds . Relative to the n-bit raw record, the ideal compression gap is therefore at most d bits. The ensemble is simple, but identifying the realized member still costs approximately n bits.
Theorem 1 concerns the lossless shortening of the realized history. Its counting measure is not an assumption that physical initial conditions are uniformly distributed. Restricting the initial or boundary information, or discarding information from the history, can reduce the description burden; compression requires that the retained history become shorter to describe than its raw code. Microlaws may still expose useful symmetries or predictive constraints. In particular, a simple stochastic law can describe an ensemble while identifying its realized member remains costly.
3.2. No Total Procedure Finds a Shortening Whenever One Exists
The second barrier is computational rather than informational. It asks only for improvement, not optimality: if the completed macrohistory has a lossless description shorter than its literal code, can one total procedure always find some such shortening?
No shift-register mechanism is needed. For each , let be the identity update on and let be the full-state readout. For an arbitrary nonempty string y, define the one-observation experiment
Under the fixed-width readout convention of Definition 1, . Thus is a computable embedding of arbitrary finite records into the formal experiment class. The computational barriers below require only this representational richness, not the information-leak mechanism used in Theorem 1.
A universal discovery rule would therefore have to exploit every compression gap in this embedded family while halting and returning a valid description on every input. The discovery barrier rules out that guarantee.
Theorem 2
(Discovery barrier). There is no total computable procedure A that, for every finite micro-experiment μ, returns a program satisfying
and is guaranteed to satisfy
Proof.
Suppose such an A exists. Given any nonempty string y, construct by Eq. (7) and compute . By the unconditional validity condition (8), , so implies . Conversely, if , the guarantee (9) gives . Comparing with would therefore decide whether , contradicting undecidability of the compressible strings (Appendix C). The empty-string case can be hard-coded. □
Interpretation. For every total procedure that always returns a valid lossless description, there exists at least one compressible record on which that procedure fails to shorten. A single such record defeats the guarantee. The theorem therefore does not say that no procedure ever compresses, that a given compressor cannot shorten most records of interest, or that all shorter programs must be produced; nor does it exclude open-ended search, which succeeds on positive instances without a stopping rule.
Even complete access to the finite micro-experiment does not provide a universal shortcut. “Complete” includes the microscopic rule, initial state, observation map, and finite horizon, so the record is straightforwardly computable by simulation. The reduction uses only that the formal experiment class can represent an arbitrary finite record through the trivial identity experiment. Classical undecidability of raw compressibility therefore transfers without invoking the shift-register witness or any assumption about incompressible initial states. Thus a supplied program that reproduces the data gives an upper bound on its description length; finding a shorter description is a further computational task.
Finding a shortening, and knowing when to stop. With the reference machine fixed, a finite record x presents three distinct tasks: exhibit a program shorter than x; establish that no such program exists; or determine the shortest program. For a threshold b, dovetail all programs shorter than b and return the first that halts with output x. If , this search eventually succeeds and the program is a finite witness. There is no computable bound on its waiting time valid for all positive instances, even when the bound may depend on x and b (Proposition A2). At any stipulated execution rate, some successful searches therefore outlast any fixed time budget, including the age of the universe. The same absence of a computable bound holds for the runtime of a fastest shortest program (Appendix C.2).
Certifying that a candidate is shortest instead requires excluding every shorter description. Continued simulation supplies no general stopping rule: short programs can halt extremely late, as measured by runtime Busy Beaver functions [25]. Shortness of the record alone gives no assurance of feasible certification. Particular instances and finite ranges can nevertheless be settled. A complete halting catalog through b program bits decides compressibility for every record of length at most , and gives exact complexity for every output represented in it. Appendix C.2 gives the precise implication and a published example; the cutoff depends on the machine. An unresolved candidate need not obstruct a particular record if one can establish that it cannot output that record.
3.3. Optimality and Unbounded Additive Regret
The third barrier asks a different computational question. A procedure may find useful shortenings on many instances while still returning descriptions arbitrarily far from shortest. For a total computable procedure A on micro-experiments that always returns a valid program, define its additive description regret by
Every fixed total constructor incurs arbitrarily large regret on some compressible records. The quantitative bound below will locate such blind spots at every sufficiently large length.
Theorem 3
(Optimality barrier: unbounded regret on compressible records). Let A be a total computable procedure that returns a valid description for every finite micro-experiment μ. Then for every there is a finite micro-experiment μ whose macrohistory is compressible,
Equivalently, the additive regret is unbounded for every such procedure already on the compressible instances.
Proof.
Proposition 1 below gives witnesses with and . Since , choose n so that . Then is compressible and its regret exceeds c. □
The following elementary counting-and-selection argument makes the failure quantitative.
Proposition 1
(Blind spots at every length). Under the hypotheses of Theorem 3, there is a constant such that, for every , some record in the identity experiment family satisfies
Here uses the fixed integer encoding of the observer frame.
Proof.
Validity makes injective: one program cannot output two distinct records. There are records of length n but only binary programs shorter than n. Let be the lexicographically first length-n record with . It exists by counting and is computable from n by evaluating the total procedure on the finite ordered list. A fixed program implementing this search gives . Subtracting yields the regret bound. The identity embedding preserves the record exactly. □
Exact answers with abstention. Even a partial computable procedure that returns correctly whenever it answers can answer on only finitely many distinct records (Theorem A2, Appendix C.2; cf. [26]). This concerns guaranteed exactness on every answered instance, not a compressor’s incidental success in returning a shortest program without recognizing its optimality.
Interpretation. Every fixed total method has records it fails to shorten at each sufficiently large length, although those records have descriptions of only bits. The bound implies Theorem 3 and Corollary 1, and supplies witnesses for Theorem 2 at every sufficiently large length. It describes worst-case blind spots, not typical performance. The constant depends on the method, reference machine, and fixed encodings; constructing may be extremely expensive. Thus the bound gives no numerical hardness guarantee for a particular short record. Appendix C.1 gives a witness-based version of unbounded regret that does not take exact numerical complexity as primitive.
Section 5 will translate between programs and the representation–model pairs defined there, and each translation costs a fixed number of bits. The hard instances must therefore stay hard after a fixed compression margin is required, so that they can pay that cost. The following strengthening lets the compression margin and the regret tolerance be set independently.
Corollary 1
(Regret with a fixed compression margin). Under the hypotheses of Theorem 3, for every there is a finite micro-experiment μ such that
Proof.
Choose n with in Proposition 1. Its witness satisfies both inequalities. □
The discovery and optimality barriers are logically distinct. Discovery asks for a guarantee of some shortening whenever one exists. Bounded regret asks that every returned description stay within a fixed distance of K. A fixed-c approximation need not shorten records lying within c bits of their literal length, while a procedure that successfully shortens many records may still lie arbitrarily far above K. Both are computational obstructions, but they exclude different guarantees.
Appendix D records further technical variants: exact conditioning on the common microlaw does not make the residual complexity computable, and the open-ended below-threshold search has no computable universal schedule.
A separate limitation concerns proof rather than discovery: Chaitin incompleteness prevents a fixed sound formal theory from proving arbitrarily strong lower bounds on K (Appendix C.3).
3.4. Restricted Construction Problems
The computational reductions use the ability to represent arbitrary finite records, with microscopic simulation itself kept easy. Restricting the physical systems or observations may remove this embedding and permit a constructor on the restricted class. A positive result must identify both the source of the compression gain and a procedure that can find it. We first examine where the record-specific information can reside.
3.5. Source Accounting: Where Realization Information Can Reside
Theorem 1 raises an accounting question: where do the record-specific bits reside? For a computable deterministic evolution
with computable projection , simulation gives
Theorem 1 shows that the right-hand side can scale with record length. The same inequality also identifies one route to making compression possible: if physics constrains the information selecting the realized trajectory, it constrains the residual burden. This is not yet a compression theorem, because the remaining burden must still fall substantially below the raw cost of the retained macrohistory.
From an initial condition to a boundary rule. Suppose first that an additional computable prescription selects the initial state from the dynamics,
Then
and, if is included in the known physical background,
The information has not disappeared by notation: it has moved from the free initial datum into the selecting prescription . If itself requires an arbitrarily long description, no algorithmic gain has been obtained. If instead is a fixed, low-complexity physical principle, the particular shift-register witness of Theorem 1 is excluded on that restricted class. For sufficiently long records this can create a large compression gap relative to the physical background; for short finite records, fixed coding overheads can still consume the practical gain.
The formulation can be made independent of the language of initial conditions. Let denote the histories admitted by the dynamical law, and let a computable global prescription restrict them to a family
If this family contains more than one history, let denote whatever additional finite information is sufficient, together with , to select the realized history relevant to the finite record. Whenever this selection is computable, the information accounting has the schematic form
Equation (20) generalizes Eq. (15). Description cost can reside in the dynamical law, in a boundary or selection principle, or in the information needed to identify one realization among those that principle permits. A short description of a finite admissible family does not make its typical members short: conditioned on , fewer than members can have complexity below m. Thus a unique computably selected history makes the term , whereas a simple large family can still require close to bits to identify the realized member. This is the ensemble–realization distinction of Theorem 1, with the information source made explicit. Bédard and Bergeron’s data-to-model index plays a related role: their boundary conditions locate a record within a model’s possibilities [15]. For the retained record, this corresponds to the realization coordinates and reconstruction corrections of Section 5; the omitted projection residual is separate.
Hartle–Hawking and Penrose. The Hartle–Hawking no-boundary proposal illustrates the global form above. Its law-like no-boundary condition defines a quantum state on regular histories [27]; semiclassically, that state yields a weighted family of classical histories [28]. In the present bookkeeping the condition plays the role of , replacing an independently specified initial wave function by a global prescription. If that prescription uniquely and computably selected one history, the residual realization identifier would be trivial; if it specifies only a nontrivial family, the term in Eq. (20) can remain substantial. The distinction does not depend on an interpretation of quantum probabilities.
Penrose’s Weyl-curvature hypothesis is a weaker example: it strongly restricts the early gravitational state without fixing every microscopic degree of freedom [29]. In this language it narrows the admissible history family and may lower the information required of , but need not select a unique -complexity microstate. A boundary principle can therefore replace, constrain, or only partially determine what would otherwise be free realization data.
A weaker escape: a complexity promise. Between arbitrary realization information and an explicitly supplied compact selecting rule lies an intermediate case. Suppose the observer cannot reconstruct but has a valid bound
By Eq. (15), the record then satisfies for a fixed simulator constant . The arbitrarily large residual used in Theorem 1 is therefore excluded when s is fixed, but a useful compression gap exists only when this bound is substantially below the raw record length.
If the observer is supplied an explicit threshold b with , the open-ended search of Section 3.2 is guaranteed eventually to find a program shorter than b, but the promise supplies no computable universal deadline: otherwise running the search only that long would decide whether . Thus guaranteed eventual discovery does not imply computably scheduled discovery; the formal statement and proof are given in Appendix D. A promise of existence is still weaker than being given the program itself.
An explicit witness certifies an upper bound: if a program p with is exhibited and its execution yields , then follows immediately. A method that always halts with a valid description cannot guarantee finding a shortening whenever one exists. Lower bounds and optimality require different evidence: Chaitin-type incompleteness (Appendix C.3) limits sufficiently strong lower-bound proofs in a fixed sound formal theory, without limiting explicit upper-bound witnesses.
There is another useful distinction between a set being finite and a search through it being known to finish. The promise implies that only finitely many programs shorter than s could describe , but there is no computable stage at which one knows that all halting programs of that length have appeared. By contrast, if physics supplies an explicit parametrization
with g total and the finite number of candidates J explicitly known, all candidates can be generated, simulated, and checked under a computable schedule. Physics has then replaced unrestricted program search by a decidable finite model class.
Stronger prescriptions buy stronger guarantees. A concise boundary prescription that selects the realization can make long records simple relative to the physical background; a promise guarantees eventual open-ended discovery but no computable deadline; an explicit witness settles its instance; and an explicit finite model class makes exhaustive search computably schedulable within that class. None guarantees that a supplied class contains every globally available compressor.
These restrictions may also break the arbitrary-record reductions behind the discovery and optimality barriers, so the present impossibility proof need not transfer. That is a boundary on the negative result, not a positive constructor theorem. Source accounting is the common ledger: realization information must reside in the initial state, boundary rule, realization identifier, or law; be rendered irrelevant by discarding information from the history; or come from a compact source. Moving arbitrary bits between accounts does not reduce their description cost, and reducing the burden creates compression only when the retained history is simpler than its raw code.
4. Structure Functions and the Geometry of Discovery
Source accounting asked where the case-specific information resides. The next question is how a description should divide its cost between a model and the information that identifies one record among those the model covers. Algorithmic statistics quantifies this division: its structure function balances the cost of specifying a macrostate against the information needed to locate the realized microstate within it.
4.1. Macrostates, Model Budget, and Cracks
For a finite record , read a finite set as a candidate macrostate (literally, a set of states): it declares which microrecords remain equivalent at that descriptive level. Its cost is , while bits suffice to identify the realized x once the macrostate is known. For a model-description budget , Kolmogorov’s structure function is
Thus is the residual index length of the tightest macrostate containing x that is specifiable within model budget . The trivial macrostate costs only the description of n but leaves n residual bits.
Once a model is available, extra model bits can specify progressively finer cells of : listing in canonical order and fixing k leading index bits selects a cell of size that is describable from , k, and the cell index. Hence, as long as residual information remains, ordinary refinement gives, up to the usual logarithmic precision,
A one-bit increase in model description can therefore buy roughly one bit less residual without revealing any new economy: information has merely moved from the index into the macrostate description. Candidate discoveries are drops substantially steeper than this one-for-one baseline (Figure 3).
For an interval , define the net structure gain
We call an interval a crack when is positive by an amount resolvable above the coding uncertainty: a small increase in model complexity removes more residual information than it costs. Since is discrete and Kolmogorov complexities are defined only to lower-order precision, “slope” and “discontinuity” are schematic descriptions of this finite-difference event.
Bédard and Bergeron [15] develop a more refined version of this idea. Their modified structure-function construction is designed so that ordinary slope- segments correspond to refinement within the same minimal partial model, whereas nontrivial drops mark transitions to successively richer partial models. Our uses the ordinary structure function and asks only whether an interval establishes net description gain beyond one-for-one refinement. It is therefore a weaker geometric diagnostic, used here to nominate candidate algorithmic discovery events; Section 5 asks whether such an economy can be reorganized as a reusable parametrized compression that continues to work on new data.
The corresponding constrained two-part, or MDL, function is
It does not define a second frontier. With the definitions above,
Thus is the running minimum of the same model–residual frontier after the one-for-one trade is charged explicitly. Along an ordinary refinement, is approximately constant; a crack lowers it when the new model establishes a better total description. The MDL view is therefore often more intuitive economically, but it contains no separate discovery geometry.
When a macrostate satisfies
it is an algorithmic sufficient statistic: at that precision, all compressible regularity has been assigned to the macrostate and the remaining index is irreducible [30,31].
This picture also reconnects to the discovery barrier. For a residual target r, the statement asserts the existence of a finite-set model within budget whose residual is below r. Dovetailed search finds such a model when it exists but need not halt when it does not; indeed, the unrestricted structure function is not computable [31]. The frontier is an ideal description of what can be discovered, not a computable map handed to the observer.
From the structure-function split to a coarse-graining. The structure function is an abstract retained–residual frontier. A projection gives a concrete realization of the same split. Here D denotes the finite pre-representation record at the modeling grain under discussion, and the corresponding computable record-level representation map. For a time series this may be induced by, or extend, a pointwise observation map; in the barrier theorems that observation map is already fixed and denotes its output. Fix a finite computable record domain as part of the declared frame, let , and define
The set is the fiber of : the records that share the retained value .
Choose a self-delimiting residual index identifying D inside this finite fiber. Then, at the declared finite domain/resolution and up to fixed coding overhead,
Moreover,
for the canonical index code within the finite fiber; the fixed domain is part of the conditioned frame C. Thus realizes a concrete point on, or above, the ideal structure-function frontier.
We call the retained algorithmic kernel relative to : it is the part of the record kept by the observer’s representation. The coordinate is the realization detail discarded by the lossy coarse-graining. Thus a candidate selects one particular model–residual decomposition from the broader structure-function frontier. Operationally, science keeps the kernel and throws away the residual; the lossless pair in Eq. (30) is only the AIT bookkeeping device that tells us exactly what was discarded. The kernel is what the representation keeps; it may itself remain incompressible, and modeling it is the next step.
This also refines the usual MDL decomposition. Once a reusable model M of is introduced, the data term , an actual code length, separates up to coding overhead as
where denotes meaningful coordinates of the retained kernel. The first term is not “noise”: it is retained model state or realization-specific parameter information. The second is the omitted residual. Section 5 makes this reusable interpretation explicit.
4.2. Relevance-Constrained Coarse-Graining
The main barriers fix the projection . Joint construction must also choose what to retain and what to omit. The objective supplies the reason for a relevance constraint: without one, a constant map could discard everything, while the identity map could retain everything. Neither extreme alone establishes useful compression.
For fixed , define
the macrohistory family induced by . These dynamically realizable families form a restricted subclass of the arbitrary finite-set macrostates in Eq. (23), with only the general upper bound
This bound controls the cost of naming the history family induced by a candidate ; it does not compare candidates, because changing changes both the family and the retained data.
Fix the objective and a nonempty, effectively encoded class of admissible projections under frame C. The relevance criterion states which distinctions are needed to evaluate or pursue that objective, and how their preservation is assessed. It constrains acceptable omission; the compression certificate separately charges the representation and model of what remains.
Predictive sufficiency, dynamical closure, and information about a designated target provide different relevance criteria. Predictive sufficiency identifies histories that induce the same predictive distribution for a designated future target, as in computational mechanics [32]. Exact dynamical closure requires a macro-update satisfying . Intervention compatibility additionally requires the representation and model to reproduce the effects of the specified interventions. These conditions support different scientific claims: prediction of a target, autonomous macrodynamics, and causal or interventional adequacy. Compression alone establishes none of them. For a finite target record z, retained algorithmic information is an ideal diagnostic; its evaluation is uncomputable in general, so an effective test needs a computable surrogate or restricted structure. Israeli and Goldenfeld give a constructive closure example by searching a finite class of local block maps [33]. Finiteness belongs to that example; the present formulation allows infinite classes of programs.
The constructive target is the pair : preserves the relevant distinctions and M supplies a reusable code for what it retains. The two programs may be discovered together or refined in alternation. Supplying fixes one component of this same construction task. Section 5 gives the certificate, proves the general construction barrier, and asks which effective restrictions restore a guarantee (Problem 1).
5. Algorithmic Emergence: Acquiring Reusable Regularity
5.1. Models Shared Across Observations
Figure 2 shows the setup: the observer receives a history through and acquires a pair to predict or support action. We now ask what that acquisition achieves. Let observations arrive in chunks , with supplied conditions . Here denotes a data chunk, whereas denotes the microscopic update rule. The retained sequence is . The representation may act on a whole chunk; applying the same state readout at each time is a special case. The objective determines which distinctions must survive this representation; the model describes regularities within and across the retained observations. Omitted detail belongs to , as in Section 4.
A sequence of hand images makes the distinction concrete [11]. The shared model describes a hand; changing parameters describe its pose, rotation, and shape. If the objective concerns hand movement, background detail may be omitted. A recurring gesture also makes the parameter sequence compressible. Individual images need not become shorter: the saving can come from what the sequence shares. Repeated copies of an otherwise incompressible image are another example. Storing that image once and referring back to it captures a real regularity.
Let the finite model program provide computable encoder and decoder maps, total on the declared domain of finite observation blocks. Write
Here denotes the realization coordinates or states for the block, represented by one joint code. Its length includes all framing and any corrections needed for exact reconstruction of the retained data. Approximate predictions can therefore support a lossless description of the retained observations when their correction data are included and charged. Separate parameter codes with total length are a special case; joint or sequential coding can exploit relations among them. The same interface assigns coordinates a stable role across observations, such as pose, object identity, initial state, phase, or an intervention handle. A learned neural encoder may supply those coordinates without being human-readable. Supplied conditions and deterministically recovered state can be used by both encoder and decoder; additional information must be encoded. When the coordinates and conditions are available before a future observation, the same decoder also supplies a prediction.
The complete retained description has length
The first term is the actual additional self-delimiting code needed to acquire the pair beyond the prior frame C; it is paid once. It is not the time or energy spent learning. Either component may reuse prior resources, and an inherited or supplied has zero incremental cost. The second term describes the observations through that shared model. Appendix B gives the full accounting when omitted detail is restored.
Even a large, inelegant implementation can capture regularity and repay its cost through reuse. This raises a separate question: how much of its program length is needed to reproduce the behavior that matters? Different programs and algorithms can compute the same function [34,35]. We compare implementations by a declared task that specifies the relevant input/output behavior. The objective motivates this task, but equality of objective values alone would not establish behavioral equivalence. A shortest equivalent program is the model’s functional core; for a coding task it is the minimum reusable program cost of that compression strategy.
Definition 2
(Task-relative functional core). Fix a finitely encoded task with admissible input domain and a computable interface specifying the observed output. Let denote the behavior induced by model M, computable on that domain. Write
Thefunctional core complexityof M relative to the pre-acquisition frame C is
Afunctional coreis any program attaining this minimum. If the declared behavior yields systematic coding gains, such a representative is acompressive coreof M.
For the generative interface in Eq. (35), the behavior may be ; for a predictor, it may be a next-observation map or a conditional predictive distribution. A core preserves that behavior, including its errors and memorized exceptions. It is shortest among equivalent implementations; another model may predict or compress better. The core need not be a localized subprogram of the acquired implementation.
Consider a program that predicts planetary trajectories using a Newtonian solver. A second program runs the same solver but also executes a finite Minecraft simulation before returning each prediction. The game’s state and output never affect the celestial-mechanics calculation, and its output is discarded. Both programs therefore return exactly the same predictions, although the second can contain much more code and take much longer to run. For this prediction task, they have the same functional core complexity: the additional game contributes nothing to the declared predictive behavior.
Additional code can also serve a purpose. A third implementation could store exact precomputed answers to frequent queries, returning those answers directly and using the original Newtonian solver for all other inputs. At the declared precision, its predictions remain identical, but the stored answers can reduce response time [36]. Its functional core complexity is again unchanged. These examples distinguish the minimum code needed to reproduce the predictions from the cost and speed of a particular implementation. A response deadline imposes an additional runtime requirement.
Finding a shortest equivalent program is not computable in general: for a fixed task with a single admissible input and declared behavior , the functional core complexity is , where is fixed. No total computable procedure computes this value for all w (Appendix C.2). Even guaranteed shortening is impossible in general: no total computable procedure can always return an equivalent program and shorten it whenever a shortening exists. Repeated application of such a procedure, stopping when the length no longer decreases, would yield a shortest equivalent program.
The functional core is therefore an ideal benchmark; the observer must pay for the implementation it actually acquires. Appendix B gives the condition under which repeated coding gains repay that cost.
5.2. Algorithmic Emergence: Acquisition and Reuse
Hold fixed the observer’s pre-acquisition frame , objective , and relevance requirements. They determine the admissible projection class . Changing the model or its coordinates does not change the objective. Relevance constrains acceptable omission; it does not guarantee that the retained data contain a compressible regularity.
Let be a computable baseline code length for the retained sequence, with decoding rules and supplied information fixed before candidate selection. It may be a literal code or an incumbent compressor. A sum is one allowed baseline. Use fixed encodings of the required variables so that padding cannot inflate the apparent savings, and require , where is the declared original-data baseline. These comparisons concern valid codes, not arbitrary numerical scores.
Choose fixed positive total margins and before selection. After acquisition, evaluate the pair on a new finite observation block. Write for any history or model state already available to both encoder and decoder from the construction observations. Both test codes may use this same declared information; any extra state must be paid for. Figure 4 summarizes this acquisition and reuse accounting.
Definition 3
(Compression-certified algorithmic emergence). Relative to the pre-acquisition frame C,compression-certified algorithmic emergenceoccurs when acquires a pair exposing previously unexploited regularity and meeting the following requirements.
(i) Relevance to the objective. retains the distinctions required to evaluate or pursue , under the declared accuracy and domain. The omitted detail is irrelevant, or acceptably irrelevant, under these requirements.
(ii) Reusable model and coordinates.The same finite model and coordinate interface satisfy Eq. (35) throughout the declared domain of finite observation blocks. The representation, model, and coding rules remain fixed during evaluation; their declared state and parameter values may change as observations arrive. The coordinates have declared roles that remain fixed across observations—for example, pose, object identity, state, or phase—while their values may change. Models without realization parameters are also allowed; any framing or correction data needed for reconstruction still count toward the code length.
(iii) Net compression of the retained observations.With the observations collected so far, the complete retained code beats its baseline:
(iv) Prospective reuse.On a new observation block acquired after the pair is fixed, its coordinate code satisfies
The acquisition cost paid in (iii) is not charged again. Decoding uses only the acquired pair and the declared shared information, and reconstructs the retained test block exactly.
The inequalities certify compression and reuse relative to the declared baseline. They do not by themselves establish historical novelty: a model already usable in the prior frame does not become a new acquisition merely by beating that baseline.
The certificate measures compression within the retained observations; omitting detail alone does not satisfy (iii). Savings can come from a shared object model, regular changes in parameters, or simple repetition. A stored pattern qualifies when its use provides new net savings relative to the declared baseline, with its storage cost counted. Neither margin requires every chunk to be shortened. A useful regularity can be acquired before enough observations have arrived to repay its implementation cost; the certificate then remains unmet. Passing the prospective test shows reuse on the tested observations, without guaranteeing every future continuation.
The ideal quantities involving K, the explicit certificate, and empirical estimates serve different purposes. The first describe shortest codes but are generally uncomputable. For a candidate with an established admissible coding interface, the certificate checks finite reconstruction and actual code lengths; it does not decide totality or relevance for arbitrary programs.
At the empirical level, the Coding Theorem Method estimates complexity from small-program output frequencies, and the Block Decomposition Method combines local estimates and block multiplicities [37,38]. Such scores can guide search, but passing the certificate requires a complete lossless code, including its model, arrangement information, and framing. A periodic record can have balanced symbol frequencies while admitting a short description of its repetition, so single-symbol frequencies alone can miss regularity. Perturbation-based methods provide another route to proposing generative models [39]. Any computable method that always terminates with a valid description remains subject to Proposition 1. Numerical scores alone are not descriptions.
Prospective reuse concerns coding observations obtained after selection. Advance prediction additionally requires the relevant coordinates or predictive distribution before those observations arrive.
5.3. Reusable Objects and Coarse Laws
Two examples help to make Definition 3 concrete. In both, the observer already knows the microscopic law and can simulate; the question is what a further model adds and whether its reuse repays its cost. The first model captures a recurring object, the second a law for a coarse variable. The register example in Theorem 1 illustrates a different obstruction: the record can retain so much initial-condition information that no substantial shortening exists.
A recurring object. The observer’s objective is to predict live-cell positions and retain 32-frame movies of Conway’s Game of Life, a two-dimensional binary cellular automaton [16]. On our periodic grid, all cells update simultaneously: a live cell survives with two or three live neighbors among its eight neighbors, and a dead cell becomes live with exactly three. Each microstate contains the grid G and an independent 4096-bit cyclic register s; the update applies Life to G and a one-cell cyclic shift to s. The observation map retains the grid, which is the variable required by the objective. Figure 5 defines the experiment and shows an initial five-cell configuration called a glider: its shape recurs after four Life updates, translated by one cell in each coordinate.
A model storing that cycle and translation reconstructs the past frames and predicts later ones from the same initial position, phase within the four-step cycle, and direction. Those coordinates change across runs while the model is reused. Against an incumbent that already stores the five live-cell positions and simulates Life, the model saves 65 bits per restarted run before its acquisition cost is charged. The supplied implementation passes the 128-bit construction margin after 294 runs and the same reuse margin on two subsequent runs; a single movie does not repay it. Appendix B.2 gives the complete accounting and coding conventions.
A law for coarse observations. The observer now seeks to predict which blocks of neighboring cells are fully occupied. The microscopic system is an elementary cellular automaton (ECA): a line of binary cells updated simultaneously, each from its own bit and its two nearest neighbors. Rule numbers follow Wolfram’s convention [40]. We use Israeli and Goldenfeld’s [33] example of rule 146, which outputs 1 for neighborhoods 001, 100, and 111, and 0 otherwise. Our periodic ring has 360 cells, grouped into 120 aligned triples and observed every three microscopic steps.
The readout is 1 only for the triple 111, the Boolean AND projection. Israeli and Goldenfeld showed that its future follows rule 128 exactly: a coarse cell remains occupied precisely when it and both neighboring coarse cells are occupied, so occupied intervals shrink from their ends. The initial coarse row and this law generate every later retained row, independently of the discarded microscopic detail (Figure 6b,c). In the declared code, acquisition and the first movie cost 151 bits against the microscopic simulator’s 361-bit description; a separately initialized test movie costs 121 bits with the pair held. Both savings exceed 128 bits.
The observation task changes if the observer must instead predict whether each triple contains any occupied cell, the Boolean OR readout. Equal coarse neighborhoods then have unequal successors, so none of the 256 elementary rules reproduces every observed transition. The rule minimizing one-step error produces the divergent reconstruction in Figure 6e; correcting it costs more than the microscopic baseline. This declared model search fails the certificate. Models using additional state, history, or a different coding method remain possible. Appendix B.2 supplies the search, the second coding account, and the fresh-run test. These examples make acquisition and its finite checks concrete; the theorems establish the limits of uniform discovery guarantees.
5.4. Why Discovery Can Fail
The certificate allows real discoveries, including repetition. It does not provide a method guaranteed to find them with the observations available. Consider a stream of fixed-size images that repeats a finite movie. A compact program may generate that movie and its repetitions, yet a given total method can fail to find a shortening of the movie it has received.
Proposition 2
(Discovery and regret barriers for reusable models). Fix positive total margins and . Let the modeling domain include the periodic-block family of Appendix D.6, with fixed chunk width , supplied block length m, and literal baseline . Let A be a total computable constructor that always returns a valid candidate: an admissible pair whose encoder and decoder are total on the declared domain and reconstruct its retained data.
(a) Discovery.For every sufficiently large m, there is an m-chunk observation block for which a pair satisfies Definition 3, but A returns a complete retained code of length at least and therefore fails (iii).
(b) Regret.There is a constant , independent of m, such that the best qualifying complete code on these instances has length at most . The excess length returned by A is consequently at least , and is unbounded.
Proof
(Proof idea). There are possible observation blocks. Validity makes their complete codes distinct, but only binary strings have length below . Some block is therefore left unshortened. A fixed program containing A can find the first such block by finite search, given m, and replay it. Its code cost is independent of m. A match flag encodes each predicted block; an escape flag followed by literal data keeps the code valid on other blocks. For sufficiently large m, this model passes both fixed margins. Appendix D.6 gives the complete construction and accounting. □
The short model may be extremely slow: this is a description-length result with no runtime deadline. Storing the block literally and replaying it is possible after the first copy, but does not yet save bits once storage is charged. Further matching observations can repay that cost. The model constructed in the proof already meets the construction compression margin on the first copy and meets the prospective-reuse margin on the next copy, yet the given constructor misses it. Failure concerns the observations collected so far; the same method may recognize and exploit repetition later. The witness uses a supplied representation. That is enough to defeat a general joint guarantee, since a method required to handle the full family must handle this subclass. It does not establish a separate hardness result for choosing .
Problem 1
(Joint construction in restricted families). Under which effective restrictions can one terminating algorithm find a relevant representation and a reusable compressive model whenever a qualifying pair exists? Any pair meeting Definition 3 suffices; optimality is not required.
Each finite instance w declares the construction inputs, prior frame, objective and relevance requirements, admissible representation and model classes, baselines, margins, and evaluation protocol. It specifies the effective access to those requirements; test observations arrive only after selection. Characterize families on which the same total computable constructor always returns a valid candidate, and returns a qualifying pair whenever one exists with the available construction observations and declared continuation. Candidate classes may be infinite. Where a guarantee fails, identify the obstruction; restrictions on future continuations and effective stopping rules are part of the problem.
6. Conserved Information and the Choice of Coarse-Graining
What a representation omits does not thereby leave the substrate. Under reversible microscopic dynamics, complete-state algorithmic complexity is conserved up to fixed coding constants, with the law and elapsed time supplied. Detail omitted now can reach the retained variables later. The retained variables therefore need not stay simple, and need not determine their own future. Both bear on representation choice, because a useful projection has to support modeling across time, not only at one instant.
Let be a complete finite encoded state, and assume that, at the declared resolution, its dynamics is a computable bijection. Let be a computable projection on these states and write . A self-delimiting residual index identifies within the finite set of states having the same projected value . Its length may vary between these sets. Unlike the records used in model acquisition, these variables describe one instant.
Dropping the residual makes the representation lossy; it does not by itself make the retained information easier to model.
Complete-state conservation and localization. Write the fixed reversible evolution as
for a supplied signed offset T and computable maps F and . Fix . Reversible invariance then gives
The dynamics may unfold and redistribute the complete description, but does not create new algorithmic information that grows with system size [13,41].
At the declared finite resolution, the forward map and its inverse form a computable reversible self-delimiting representation,
The allowed residual codes may therefore depend on ; we do not assume one fixed residual alphabet for every projected value. The actual coarse-graining is the subsequent lossy operation that retains and omits . With the same held fixed,
Subtracting the same identity at two times and using Eq. (42) gives the localization balance, where denotes the change between the two times,
with the common conditioning suppressed for readability.
This is the sense in which a projected macrocoordinate can become more complex while the global budget is conserved: description already present in the complete state can be relocalized into the retained kernel, the residual, and their shared algorithmic structure. Two difficulties follow for the observer, neither of them a failure of the accounting. First, the retained coordinate can gain description length as detail once held in the residual becomes visible in it. Second, and independently, the projection can discard predictive distinctions: two complete states sharing a retained value may evolve to different retained values, so no deterministic update on the retained value alone reproduces both. The observer must then enlarge the representation, condition on more of the history, or describe the retained variables probabilistically. This is failure of the exact closure condition of Section 4.2; Section 7 exhibits structures under which it holds. Neither effect amounts to information creation by the projection:
The fuller coordinate decomposition and exact localization ledger are developed in the companion paper [41]; they are not needed for the construction results here.
A minimal relocalization picture. Consider a reversible cellular automaton whose block-averaged initial field is simple while the within-block arrangement is incompressible. Most of then resides in . Under a mixing reversible rule, fine-scale differences can later affect the retained block field, so bits initially localized in the residual become necessary to specify . Equation (45) accounts for this transfer without creating information. Detail omitted at can therefore contribute to the description complexity of the later retained field; simplicity of the initial field alone gives no guarantee of an economical description across time. Theorem 1 shows how detail omitted from an initial readout can become visible across an observed history: the instantaneous readout discards all but one bit at every step, yet the history it generates reproduces the entire initial state. Each readout remains simple; the accumulated history carries the complexity. Conversely, if the seed is simple conditional on the law, then for every t; apparent macro-level richness is only the unfolding of a short program.
From states to observed histories. The balance above concerns instantaneous states; model acquisition concerns records accumulated through time. A useful projection must preserve the distinctions required by the objective across that history. Omission can reduce what must be modeled; the certificate requires additional savings within the retained information, after charging the representation and model. Section 7 examines the physical structures that make such a pair accessible: symmetry, timescale separation, and closure.
7. Macromodel Discovery in Favorable Systems
Physical structure can make useful compression possible and make its discovery tractable. Constraints on realizations and task-relevant omission reduce what must be described. Symmetries, restricted model classes, and experiments can then guide an effective search. These are separate requirements: an economical description may exist without being accessible to the observer’s available procedure. The following cases identify sources of structure relevant to the restricted joint construction asked for in Problem 1.
For the search it is useful to separate inherited microstructure from bridge structure. Symmetries, conservation laws, locality, and interaction constraints may be computably recoverable from the microlaw and can sharply shrink the admissible macro-model class without adding substantial conditional description cost. The projection, order parameter, scaling regime, timescale separation, and closure that connect levels are often the difficult part. Empirical data can then supply constitutive laws or parameters not fixed by the microscopic description.
If denotes whatever bridge information is not already available from the conditioned microscopic description and a fixed construction on a restricted class satisfies
then
Microscopic knowledge can reduce description and search cost without itself being a uniform constructor. The split is frame-relative: if some component of is derived from or learned earlier, it moves into the conditioned background C.
Symmetry and restricted coarse dynamics. Anderson’s positive lesson is that microscopic symmetry is not discarded at the higher level: it constrains admissible phases, order parameters, and collective modes [14]. His ammonia and crystal examples also show what symmetry does not do: the microscopic Hamiltonian can constrain the possible sectors without selecting the realized symmetry-broken organization or the variables in which it is simplest. Israeli and Goldenfeld give a constructive discrete analogue: within a restricted cellular-automaton search they identify block maps whose projected variables evolve by a simpler rule of their own, with no reference to the discarded detail [33,42]. This construction succeeds within the specified search class. Figure 6 shows the rule-146 example and a failed projection for the same microscopic run.
Particles to fluids. Newtonian mechanics and local conservation constrain continuum balance equations but do not by themselves close them: stresses, heat fluxes, equations of state, and transport coefficients require constitutive information. A kinetic description connects particle dynamics to hydrodynamic variables; constitutive relations then close the balance equations into fluid equations. In controlled scaling regimes the Boltzmann equation yields Euler and Navier–Stokes limits [43]; once the relevant variables and regime are identified, kinetic and statistical mechanics can also relate macroscopic coefficients to microscopic interactions. Projection methods explain the qualification: exact elimination of unresolved variables generally produces memory and fluctuating terms rather than an autonomous macrolaw [44]. Scale separation, weak memory, local equilibrium, or a restricted constitutive family are the favorable structures that make closure possible.
Microlaws constrain the joint search over . In favorable systems, exact symmetries, conservation laws, interaction topology, analytic structure, or scale hierarchy first restrict plausible projections . One then derives the evolution or statistical weight of those projected variables. If unresolved variables remain, closure has failed: the representation must be enlarged, an approximation justified, or another scaling regime chosen. Exact closure, self-averaging, or renormalization-group irrelevance can instead yield a closed equation or effective Hamiltonian M. The following examples realize different versions of this route: identifies the retained kernel, unresolved coordinates become , and the closed law supplies M.
Exact quotient: Curie–Weiss. For N spins , the mean-field Ising Hamiltonian and magnetization m are
where J is the coupling strength and h the external field; J, h, m, and the inverse temperature below are local to the Ising examples. The Hamiltonian factors exactly through the magnetization,
The all-to-all coupling and permutation symmetry therefore do more than suggest m: for energetic purposes they identify an exact quotient . Combinatorial counting gives the multiplicity of each magnetization class, and the thermodynamic limit turns energy plus multiplicity into a one-dimensional free-energy problem whose stationary condition yields the mean-field self-consistency law at inverse temperature ,
Here the microlaw removes almost the entire representation search: symmetry identifies , the microscopic arrangement within a fixed-m fiber is the residual , combinatorial counting measures that residual, and the large-N limit supplies M, the self-consistency law above [45].
Asymptotic relevance: short-range Ising and renormalization. The previous factorization fails for the nearest-neighbor Ising model. Its Hamiltonian depends on local pair structure, so microscopic configurations with the same global magnetization can have different energies. It nevertheless supports macroscopic laws. At zero field, the spin-flip symmetry singles out magnetization as the natural symmetry-breaking order parameter; translation symmetry and locality motivate block magnetizations, correlation fields, and long-wavelength modes rather than arbitrary functions of the spin configurations.
Renormalization then turns this guidance into a construction procedure: coarse-grain the microscopic field, integrate out short-scale variables, and follow the induced flow of the effective couplings. Symmetry restricts which operators may appear, locality organizes their form, and spatial dimension determines which couplings are relevant, marginal, or irrelevant. Near a fixed point, most microscopic couplings die away while a small set of relevant directions controls phase structure, scaling laws, and critical exponents [46,47]. Thus need not be an exact finite-scale quotient to support macroscopic laws: the microlaw can instead identify a representation whose discarded coordinates become asymptotically irrelevant. The renormalization-group flow identifies which operators continue to affect long-scale behavior and which become irrelevant, with the fixed-point theory supplying M.
Exact dynamical closure in a neuronal population. Montbrió, Pazó, and Roxin provide a dynamical counterpart [48]. For a globally coupled heterogeneous population of quadratic integrate-and-fire neurons, the microscopic equations first remove neuronal labels from the interaction: in the thermodynamic limit the natural object is a population voltage density rather than N independent trajectories. On the invariant Lorentzian-ansatz family, instantaneous coupling and Lorentzian heterogeneity allow the population integral to close exactly by analytic continuation. Writing for the continuum population-density state, the resulting projection
retains the population firing rate r and mean membrane potential v (in the principal-value sense), whose evolution obeys a closed pair of ordinary differential equations. Finite populations provide empirical estimates of these continuum observables. Firing rate alone is not sufficient for all collective states; failure of that smaller representation is what forces the additional macrovariable v. In the language used here, global coupling and analytic structure narrow the candidate , the closure calculation tests whether the omitted microscopic coordinates can indeed be relegated to , and the exact reduced equations are M.
These cases illustrate three distinct ways microscopic information can make macromodel discovery tractable: an exact quotient can collapse the representation space immediately; a controlled asymptotic limit can make discarded microscopic distinctions irrelevant; or an invariant/analytic closure can identify the minimal set of collective variables needed for autonomous dynamics. These are mechanisms for discovering representations and laws; satisfying Definition 3 additionally requires the declared compression margins and held-out reuse. None is a universal constructor. What they show is more specific: favorable microlaws can encode strong clues about which representation to try and can provide mathematical tests of whether the proposed representation supports a closed macrolaw.
Hybrid and macro-first construction. Hodgkin and Huxley combined membrane-level charge and circuit constraints with a mesoscopic representation and empirically inferred ionic conductances and gating kinetics [49]. Their model is neither “from microphysics” nor “from data alone.” Thermodynamics supplies the converse historical route, with stable macrolaws preceding their statistical-mechanical explanation. Reduction is therefore usually a ladder of effective levels, and science may enter that ladder from above, below, or both.
These examples suggest the positive scientific counterpart to the theorems: constrained cumulative search. Science does not enumerate all programs. It inherits representations and laws, uses experiments and interventions to discriminate candidates, restricts attention to model classes that have worked before, and retains structures that continue to compress new data under a stable parameterization. A restricted class makes the search manageable when, for example, it is finite and effectively enumerable or has a computable stopping rule.
Within such a search, the observer can certify a relative optimum. Fix , the modeling task and coding convention, a candidate class , and a computable score L. If every candidate in can be checked and a computable rule determines when the search is complete, then the observer can certify
This is optimality only within . Successful models then become part of the next observer frame C, shrinking later searches. The success of this cumulative process depends on which representations the observer inherits, encounters, and can test.
8. Discussion
Constructing a reusable model, computing answers from it, and using it to act impose different requirements. These distinctions explain how the construction barriers coexist with successful scientific modeling and how they relate to other accounts of emergence.
An acquired model’s implementation is judged against the agent’s objectives. For telehomeostatic agents, models support regulation that sustains the persistence of the agent or its kind, so predictive accuracy, response time, and resource demands matter through their effects on that objective [13]. A shortest program can be too slow for a required action, while a larger equivalent program may meet the deadline. Evolution or learning can favor the additional stored structure when its benefit outweighs its cost. The compression certificate tests coding savings after acquisition costs; adequacy for the task and effective planning are separate requirements for successful action (Section 8.3).
8.1. Construction and Use: Distinct Computational Limits
We distinguish constructing a model from using one. Construction asks whether reusable structure exists in the chosen representation and whether an agent can find it. Use asks how a declared question q can be answered from the microscopic description or from an already acquired model. These problems can fail for different reasons. In particular, description length and runtime are separate axes. The shift-register example has a history that can remain incompressible even though a bit at a specified future time is cheap to obtain. Conversely, a system started from a short generative seed may have a compact description while answering a far-future question may still require running the dynamics.
Evaluation may have a computable stopping criterion, such as a supplied finite horizon or a finite-state cycle certificate, or be open-ended. This concerns answering q, not coding the finite record. The box distinguishes construction failures from answer-side undecidability.
[Diagnostic situations in the micro-to-macro problem]
No substantial lossless compression of the retained history. Even for a finite micro-experiment, the retained history may have no substantial compression gap (Theorem 1).
Every total model builder has blind spots. Each fixed total valid description method leaves some records unshortened at every sufficiently large length, even though they have logarithmic-length descriptions (Proposition 1). The discovery and fixed-regret failures also hold for complete codes under the compression certificate, with the observations available so far (Proposition 2).
No computable macro map for a chosen variable. With open-ended computation and fixed relevant , there may be no algorithm that returns the requested answer for every microscopic specification in the family, as in the Ising-observable and phase-diagram constructions [50,51].
No computable long-run answer inside a known model. A compact update rule may be explicit and each finite step computable, while its eventual state or fixed point remains undecidable, as in a Turing machine with an unbounded tape or the renormalization-group construction [52,53]. When microscopic evaluation has a stopping rule. For eventual-time queries on a deterministic finite reachable state graph, exact cycle detection supplies a stopping rule. Instance sizes and horizons may still grow without bound across the family. A compact continuous phase space need not be finite-state, and Poincaré recurrence does not in general imply exact periodicity; geometric boundedness alone therefore supplies no such rule (Appendix C.4). When evaluation terminates, the remaining question is whether the observer can avoid replaying the microscopic trajectory.
Bedau calls a macrostate weakly emergent if it is derivable from the system’s microdynamics and external conditions, but only by simulation. His “macrostate” can also denote a property or pattern of behavior. Wolfram’s principle of computational irreducibility gives the corresponding no-shortcut thesis for rule-based systems: predicting the state at horizon t may require work equivalent to running the evolution to t [40,54]. In this terminating case the obstruction is computational cost, not the absence of an answer. Neither formulation by itself proves a complexity lower bound for every rule and every macroscopic question. Figure 7 separates possible evaluation routes from the limits on universal construction and answering guarantees.
Acquiring gives the observer an alternative route to answers within the model’s domain, through autonomous coarse dynamics, a closed relation, or an effective law for the retained variables.
Representation couples description length and runtime. Section 7 discusses Israeli and Goldenfeld’s coarse-grainings that restore predictive shortcuts for retained questions in computationally irreducible automata [33,42]; Figure 6 illustrates one exact coarse law. Section 6 gives the information accounting: with the law and elapsed time supplied, reversible dynamics conserves complete-state complexity up to coding constants while redistributing information. The projection determines which distinctions remain visible and available for prediction; computational irreducibility is therefore relative to the variables being tracked. The barriers proved here concern compression-certified construction, not the runtime of what is constructed.
When the answer computation is open-ended. An answer computation is open-ended when q depends on memory, lattice size, or elapsed time for which no finite cutoff is supplied. Such a model can encode the halting problem. When the problem specifies a total answer map and that map is noncomputable, no total algorithm realizes it on every input. This undecidability barrier, labeled the strong computational barrier in Figure 7, concerns the absence of a uniform algorithm rather than high computational cost. The open-ended resource makes the obstruction possible but does not imply it.
Physical realizations of the undecidability barrier. Open-ended computation need not require an unbounded geometric domain. Moore showed that smooth finite-dimensional dynamics, including a particle moving in a three-dimensional potential, can simulate a universal Turing machine [52,55].
More recent constructions realize Turing-complete stationary Euler and Navier–Stokes flows on compact or suitably metrized manifolds [56,57]. These fluid results concern Lagrangian reachability, not undecidability of Navier–Stokes singularity formation; Tao’s proposed time-dependent Eulerian fluid-computation route to blow-up is related but distinct [58,59]. Appendix C.4 explains how compact continua evade finite-state cycle detection and why Poincaré recurrence does not restore it. Elementary cellular automata provide a discrete universal counterpart [60].
Physical many-body systems realize the same distinction. Gu et al. encode cellular-automaton evolution in the ground states of an infinite periodic Ising lattice and obtain averaging observables that cannot be computed uniformly from the microscopic specification [50]. The infinite lattice is essential; fixed finite systems remain computable in principle. Spectral-gap undecidability likewise concerns the thermodynamic limit [61]; the obstruction persists in one-dimensional spin chains, extends to phase diagrams, and survives rotational symmetry on the square lattice [51,62,63].
Watson, Onorati, and Cubitt construct an explicit renormalization-group map whose finite iterates are computable, while determining which fixed point the flow approaches is undecidable [53]; see [64] for a recent review.
Finite-record construction persists in larger system classes. A finite-horizon observation from an infinite or computationally open-ended system remains a finite record and is still subject to the AIT construction barriers. Corollary A2 makes this inheritance precise: any effectively encoded larger class admitting a computable, finite-horizon, record-preserving embedding of the finite experiments inherits all three barriers. For the residual-information bound, the embedded law, projection, horizon, and frame must remain common across the initial-state ensemble. A constructor universal on that larger class would otherwise solve the embedded finite hard family. Gu et al. fix a macroscopic variable and ask for its value, whereas the observer here seeks a reusable representation and compressed model. A larger substrate class can present both problems.
Composition links construction to use. For fixed and q, suppose the formal problem assigns one answer to each microscopic input. If a total constructor always produced a suitable theory and a total evaluator always recovered that answer, their composition would compute the answer map. Proposition A3 states the implication: a noncomputable map forces at least one stage to fail to be total or universally correct. It does not identify the failing stage or exclude useful theories for individual systems. Appendix E gives the formal statement and its one-way scope.
The converse carries no such consequence. A computable answer need not arise from a compressed, reusable theory, and a useful theory need answer only questions in its declared domain. A constructor and evaluator may reproduce the microscopic computation. If the constructor may choose without a relevance constraint, it can evade the question by discarding the distinction that q asks it to preserve.
8.2. Emergence as Model Discovery
The AIT barriers concern the existence and discovery of short descriptions; algorithmic emergence concerns an observer’s acquisition of a reusable compressive model. Acquiring a model is a different question from whether its predictions require simulation, whether all questions about its behavior are decidable, or whether higher-level facts follow from lower-level facts. Table 1 organizes the accounts by these questions and by their criteria for model structure and causal influence. These questions can intersect: an observer can acquire a useful model whose behavior still requires simulation or admits undecidable questions. Our proposal is therefore a separate epistemic criterion, not a refinement of weak emergence. Comparisons must specify the property and levels under consideration.
Acquisition can make prediction cheaper. A macroscopic model can replace microscopic simulation with direct calculation or with a cheaper simulation of the retained variables. For the glider in Section 5.3, position, phase, and orientation determine later observations without updating every cell. Acquisition can therefore yield both description savings and faster prediction. Other acquired models may still require step-by-step evolution. Bedau’s criterion concerns which derivations of a macrostate exist, rather than which the observer has discovered [54].
An acquired model can retain computational barriers. Suppose a coarse-graining yields a compact macroscopic rule that implements a universal Turing machine when its memory is allowed to grow without bound. If the observer acquires a representation–model pair based on this rule that passes Definition 3, algorithmic emergence has occurred. Yet no algorithm decides, for every input, whether this macroscopic machine will ever halt. The undecidability concerns unbounded computation; each fixed finite-horizon record remains computable. Thus model discovery can succeed while computational barriers remain within the discovered model.
Observer-dependent emergence and acquired compression. Abrahão and Zenil formalize observation through mutual perturbation between a system and an observer equipped with a formal theory [68]. Their observer-dependent emergence concerns trajectory information still needed beyond the observation, observer history, and fixed allowances for coding, observational error, and processing. Their criterion identifies an informational deficit relative to the observer’s theory; ours identifies an acquisition that supplies reusable compression. Our specifies retained distinctions; observer–world interaction may be included in the dynamics. Theorem 1 shares the concern with information remaining after supplied knowledge, but uses different conditioning data.
An extension of the observer’s theory can remove their deficit by storing a literal answer, without yielding a shorter or reusable description. Conversely, a model can compress a selected component without eliminating the deficit for the full trajectory. Prospective reuse is therefore a different requirement, not a stronger version of their criterion. Their asymptotic notion fixes an evolving system and eventually exceeds each fixed observer’s informational resources. Proposition 1 instead fixes a total valid description method and finds simple records that it leaves unshortened at each sufficiently large length. No implication between these claims is asserted. Their reading of Bedau emphasizes informational incompressibility; here description length and the computation required to use a description remain separate.
Model structure and model acquisition. Bédard and Bergeron [15] define emergence through several minimal partial models of a record, distinguished by drops in a modified structure function. Their closing discussion relates scientific investigation to successive improvements in models, while emphasizing that optimal models cannot be computed from data in general. Their formal definition characterizes models for a record; ours concerns an observer’s transition to a usable model. Our certificate (Definition 3) tests acquisition through task relevance, actual implementation cost, a declared baseline, and later coding reuse; it does not require computing a structure function or identifying a minimal model. Section 4.1 uses the ordinary structure function and the weaker net-gain diagnostic . Earlier KT work treats emergence as a resource-limited agent’s discovery of useful coarse-grainings [20]. A drop from microscopic to macroscopic description length can reflect omission alone and does not by itself certify reusable regularity.
Compression is also the common ground with Gell-Mann and Lloyd’s effective complexity, which measures the description length of a record’s regularities [66]. Here the additional question is when an observer acquires a model that captures and reuses regularities.
Causal emergence and acquired models. When the representation and model reproduce the effects of the specified interventions (Section 4.2), the compression-and-reuse certificate can record the acquisition of a causal macroscopic model. Hoel and colleagues define causal emergence by an increase in effective information at a macro level: coarse-graining can reduce noise or degeneracy in the induced transitions enough to outweigh the smaller state space [67]. Acquiring a causal macromodel makes this comparison available; the compression certificate does not determine its outcome.
Reversible dynamics give a case in which the criteria separate. For a complete finite microscopic state space with N states, a deterministic reversible update has effective information bits under Hoel et al.’s uniform interventions. A coarse-graining with macrostates has at most bits, so it cannot increase this measure. An observer may nevertheless discover a compact, reusable model of its retained observations that passes Definition 3. The acquisition changes the observer’s descriptive resources while the microscopic dynamics remain information preserving. Algorithmic emergence thus requires neither information loss in the microscopic evolution nor an increase in effective information at the macro level.
Existence, entailment, and construction. Chalmers asks whether higher-level truths follow from the complete lower-level facts, interpreting follow as conceptual or metaphysical necessitation; strong emergence is a failure of this relation even in principle [65]. Of the two algorithmic tasks, answering a fixed macroscopic question is closer in form to Chalmers’s question than constructing a reusable model. This comparison concerns the questions asked, not an implication from undecidability to strong emergence. Necessitation is distinct from provability in a specified formal theory and from uniform effective construction of an answer or model. Under classical semantics, truth in an intended mathematical structure differs from consequence of a specified theory. In classical first-order logic, consequence across all models of the axioms and formal provability coincide by soundness and completeness. An arithmetic statement can nevertheless be true in the standard natural numbers yet unprovable from a given sound, effectively axiomatized arithmetic theory. Likewise, every finite string has a shortest program, and hence a definite , although no algorithm computes it for every string. Neither observation by itself establishes a failure of metaphysical determination.
Constructive approaches give a different meaning to truth and existence [69]. On an intuitionistic proof interpretation, the meaning of a proposition is given by what counts as a proof: an existential proof supplies a witness and a proof that it satisfies the asserted property, while a proof of an implication transforms proofs of its premise into proofs of its conclusion [70]. Where these constructions are interpreted computationally, a uniform existence proof must supply a witness-producing procedure from the specified inputs and evidence. Intuitionism does not by itself equate construction with Turing computability, and the absence of a uniform procedure does not preclude constructions in particular cases.
Knowing how to generate each finite stage of a microscopic process does not, in general, provide a uniform procedure for settling macroscopic questions about its unbounded behavior across all inputs. A classical realist can nevertheless regard each such proposition as already true or false; an intuitionistic proof interpretation need not assign it that proof-independent status. Forster, Kunze, and Lauermann’s constructive formalization treats complexity as a relation between objects and description lengths, rather than as a computable numerical function [71]; Appendix C.1 explains the description and compression witnesses.
Direct simulation constructs each finite observed record; an economical, reusable model is a further construction. Proposition 2 separates individual availability from uniform construction over the stated periodic-block family: no total computable constructor returning valid candidates guarantees finding a qualifying pair whenever one exists, even with the representation and block length supplied and no runtime deadline. Microscopic facts may fix the record without supplying a general effective route to a compact, reusable model. Algorithmic emergence concerns the observer’s acquisition of that model, giving Anderson’s distinction between reduction and construction an operational form.
Counting and diagonal arguments. The barrier theorems rest on two mechanisms. Theorem 1 is a counting argument and takes no diagonal step. Theorems 2 and 3 and Proposition 1 are diagonal arguments: they reduce to the halting theorem or use Berry-type constructions, and Proposition 2 carries their obstruction to representation–model pairs. Lawvere’s fixed-point theorem gives that diagonal family its common abstract form, shared with Cantor’s theorem and the halting, Gödel, and Tarski arguments [72,73]. In its set-theoretic form, if a family indexed by a set represents every function from that set to an output set, every self-map of the output set must have a fixed point. Boolean negation has none. Applying a represented function to its own index and negating the result therefore defeats the proposed universal representation.
Wolpert’s physical-inference limits use the same diagonal mechanism [74]. A device weakly infers a quantity when, for each question whether it has a specified value, some setup guarantees a correct answer in every possible world compatible with that setup. The device belongs to the universe it describes, so its own binary conclusion is one such quantity. Asking whether that conclusion is “no” requires the answer to equal its own negation. No setup can meet this requirement, so no device weakly infers every binary function of its universe.
The two-device result closes the same loop. Mutual weak inference would allow one device to reproduce the other’s conclusion and the other to negate the first’s. If every pair of setups can coexist (Wolpert’s setup-distinguishability), these requirements must hold together, again forcing a conclusion to equal its own negation. Thus two such devices cannot weakly infer one another. Both arguments turn on the absence of a fixed point of negation, the obstruction expressed by Lawvere’s theorem, and require neither Turing-computable devices nor infinitely many possible worlds.
The connection to algorithmic information theory also holds by reduction. Computing exact Kolmogorov complexity for a fixed optimal universal machine would allow one to decide halting. Consequently, the uncomputability of K follows from the halting theorem and hence, through this reduction, from Lawvere’s diagonal framework [75].
Berry-type arguments make the description-length accounting explicit. For a fixed sound, effectively axiomatized theory, Chaitin’s certification ceiling follows by searching its theorems: a proof that a specified string has sufficiently high complexity would enable a shorter search program to find that string, contradicting the bound [76]. In Proposition 1, the fixed total method returning valid descriptions and the length n specify the first n-bit record the method leaves unshortened. This gives the record a description of length , with a method-dependent constant. Theorems 2 and 3 use this reasoning or reductions to it; Proposition 2 transfers the obstruction to qualifying representation–model pairs within the stated periodic-block family. The constants and additive gaps require the explicit encoding estimates supplied by these proofs.
Li presented an index-based Berry construction within Lawvere’s framework [77], and Bauer obtains the Kleene–Rogers recursion theorem through a multivalued extension [78].
Counting supplies a complementary argument. With background knowledge fixed, there are binary records of length n, but fewer than programs shorter than bits, for integers . Fewer than a fraction of those records can therefore have such short descriptions. This finite fact is the whole mechanism of Theorem 1.
Neither counting nor diagonalization requires infinity. Unboundedness enters the applications separately. The halting theorem concerns arbitrary finite programs whose available memory and execution time have no fixed ceiling. Each terminating computation nevertheless uses only finitely much of both. A deterministic machine with a known finite configuration space admits a halting test by simulation until termination or repetition. Gödel’s and Tarski’s arithmetic results concern the unbounded natural numbers and families of finite formulas; no individual formula or proof is infinitely long. Chaitin’s ceiling similarly compares a finitely specified theory with claims of arbitrarily high complexity.
Physical undecidability constructions encode unbounded computation in the observed model. Gu and colleagues use an infinite periodic lattice, while spectral-gap undecidability concerns the thermodynamic limit [50,61]. The continuum examples discussed earlier show why bounded spatial extent does not imply a finite configuration space.
In our construction barriers, every observed experiment, horizon, and record is finite. The unboundedness lies in the family over which one constructor must provide its guarantee. The observer’s program is finite, while its working memory has no fixed ceiling; this does not make any observed experiment infinite. Classically, a table of shortest descriptions exists for every fixed description frame and maximum record length, but no computable procedure generates these tables uniformly as that maximum grows.
The common diagonal mechanism thus connects the limitations, while their assumptions determine which guarantee fails: universal inference, a uniform answer, a shortening, a qualifying model, or a fixed additive gap. None of these arguments alone establishes that a macrostate requires simulation or fails to follow from complete lower-level facts.
8.3. Restricted Discovery, Action, and Inherited Structure
Restricted criteria and prior knowledge can guide the search for useful models. Successful action then depends on what those models predict and on the policies the agent can find.
Computational mechanics builds macrovariables by quotienting histories with identical predictive futures [32]. Causal-emergence methods select macroscales that maximize an interventionist effective-information criterion [67], while information-theoretic methods can identify relevant variables at renormalization-group fixed points [79]. These approaches optimize a specified predictive, causal, or scale-dependent criterion within a restricted class. They show how task relevance and restricted structure can guide parts of the joint search in Problem 1; a complete acquisition must also meet the certificate’s compression and reuse requirements. Zenil and colleagues’ perturbation-based decompositions provide another method for proposing generative models [39]. Connecting perturbations of a description to physical interventions requires additional modeling assumptions or evidence.
A complementary boundary case is model universality. De les Coves and Cubitt showed that the two-dimensional Ising model with fields can emulate arbitrary finite classical spin models, and Reinhart, Engel, and De les Coves formalize such emulations as efficient, modular transformations [80,81]. This is a forward compiler once the target theory is specified; it does not infer from microscopic data which representation and macromodel are relevant. Target-specific information remains in the programmed simulator, so universal expressivity does not collapse the discovery problem.
Models, planning, and successful action. To help minimize through action, a model must predict the consequences of the actions being considered. The objective must also be attainable using the retained information and available actions, and an effective planner must find a satisfactory policy. In finite-horizon Markov decision processes, simulation lemmas bound policy-value errors using errors in the transition and reward models [82,83]. A performance guarantee then needs a margin covering both model error and planning error.
The converse asks when a successful controller can yield a reusable model of the task-relevant dynamics. Selecting actions and predicting observations are different interfaces; the same policy can succeed in systems with different dynamics. Recovering a model therefore needs further assumptions about the tasks, observations, or interventions that distinguish those systems. This remains a question for future work; successful control alone does not supply the compression certificate.
The prior frame itself may contain structure inherited through evolution, development, and scientific culture. The filter-of-time hypothesis in Pattern, Persist! proposes that surviving agent–world organizations are enriched for representations and models useful for persistence and regulation [13]. This is an evolutionary hypothesis, not a consequence of the construction barriers; it guarantees neither optimality nor transfer beyond the selecting environments.
Construction, evaluation, and successful action impose distinct requirements. These distinctions apply to scientific discovery and to cognitive model acquisition; science makes the construction problem unusually explicit.
9. Conclusion
Knowing the microscopic law and the information specifying a case can suffice to generate its observed history. It does not by itself yield a reusable higher-level model. Theorem 1 shows that even simple known dynamics and coarse-grained observation need not leave a substantial lossless shortening of the history. Every fixed total valid description method also leaves some records unshortened at each sufficiently large length, although they have logarithmic-length descriptions (Proposition 1). Open-ended search finds a shortening whenever one exists, but without a computable worst-case deadline. Allowing abstention does not restore unlimited exactness: a partial computable method correct whenever it answers can determine K on only finitely many distinct records (Theorem A2).
Algorithmic emergence is an event in the observer’s descriptive history: a relevant representation and reusable model make a previously unavailable compression usable. A pattern is what the model captures; discovery changes the available description, not the complexity of a fixed record. Compression connects these limits to discovery by making the reuse of regularity measurable. To meet the certificate, a representation must retain what the objective requires, and its model must save bits across the retained observations after acquisition costs, with further savings on new observations. Already for a stream that repeats one block, Proposition 2 shows that each total valid method can miss an already qualifying model with the observations available so far. The certificate is one-sided: passing it witnesses reusable regularity, while failing it leaves useful regularity unconfirmed. A short code for past data alone does not establish acquisition of a reusable model; the certificate therefore demands fixed coordinate roles and further savings on later observations. Repeated coding gains can repay a model’s fixed cost even when its implementation is large, so a regularity can be worth holding before that cost has been repaid.
The model’s functional core is a shortest implementation of the same declared behavior, with the same imperfections; that program need not be fastest or the best available predictor. For agents whose models serve telehomeostasis, a larger implementation can be preferable when it enables timely action.
Scientific practice makes progress by restricting the problem: physical constraints guide representation choices, experiments distinguish candidates, and successful models become resources for later searches. Under the reversible dynamics considered here, information omitted by a representation remains in the substrate and can reach the retained variables later. A useful projection must therefore preserve the distinctions required by the objective across time. The examples in Section 7 show how such restrictions can support construction.
A separate limit concerns answering a fixed macroscopic question. When a declared macroscopic answer map is noncomputable, no total computable constructor–evaluator pair can be correct on every input, even though useful models may exist for individual systems and questions.
Problem 1 asks when one terminating method can find a qualifying representation and model whenever such a pair exists in a restricted family. Further questions concern limits on predictive discovery under the same cost and algorithmic assumptions (Appendix B), and when acquiring and using a model is cheaper or supports more effective action than microscopic evaluation. These questions require accounting for description length, computation, repeated use, and the resources each acquisition adds to the observer’s frame.
Code and formalization availability
The core reductions behind the barrier results are formalized in Lean 4 in the KTAIT development, whose inventory and axiom audit are WP0195 [84]; Appendix F gives the declaration-level mapping and identifies the results proved only at paper level, including Proposition 2. Both repositories are cited in the Data Availability Statement.
Companion papers
Reversible complete-state conservation and partition-relative localization are developed in BCOM WP0218 [41]; persistence, telehomeostatic support, and agency profiles are developed in BCOM WP0216, Pattern, Persist! [13]. How resource-limited agents derive probability and Bayesian inference is WP0017 [20].
Data Availability Statement
The example code, verifiers, recorded bit counts, and reproduction instructions are openly available in the companion repository https://github.com/giulioruffini/WP0007-algorithmic-emergence, at commit b63dd83 for this version. The Lean 4 formalization supporting the machine-checked results is openly available at https://github.com/giulioruffini/KTAIT.
Appendix A. AIT Primer and Notation
This appendix states the AIT distinctions needed to read the paper and collects its notation. The formal definitions and theorem hypotheses in the body remain authoritative.
Appendix A.1. AIT Assigns Description Lengths to Individual Finite Records
Fix a universal decoding procedure U whose valid programs form a prefix-free set, so a program is self-delimiting and concatenated descriptions can be parsed without an external end marker. A program p is a lossless description of a finite binary record x when . Its cost is its bit length , and the prefix Kolmogorov complexity
is the least such cost. Changing the universal decoder changes by at most an additive constant independent of x; this is why the paper fixes one U and writes [5]. A practical compressor supplies a computable upper bound on through the length of its output. It does not in general establish that its output is shortest.
Kolmogorov complexity and Shannon entropy answer different questions. Shannon entropy is an average uncertainty or coding cost for draws from a specified probability distribution. Kolmogorov complexity is the shortest effective description length of one finite record, whether or not a probability distribution has been supplied. This paper concerns realized histories and therefore uses K as its ideal description-length benchmark; no claim identifies with the output of a practical compressor or with the thermodynamic entropy of a physical substrate.
Appendix A.2. Conditional Complexity Records What the Observer Already Knows
Conditional complexity is the length of the shortest program that outputs x when y is already available as auxiliary information. The description frame C plays this role for the observer. It contains the fixed decoder, coding conventions, admissibility rules, and descriptive resources acquired before the event under study. Holding C fixed prevents a trivial reassignment of the new model-compressor to the background: its compression certificate must charge the information added beyond resources the observer already possessed. The framework is observer-relative in this explicit sense, while the theorems compare records under the same fixed frame.
Side information also separates a common mechanism from realization-specific information. If an update rule, observation map, and time horizon are shared across a family of experiments, their codes may be conditioned on rather than charged anew for every record. The remaining program must still identify the initial state, boundary data, or other realization identifier needed to reproduce the realized history. The residual-information barrier shows that this remaining cost can be extensive even when the shared mechanism is simple.
Appendix A.3. Searching for Shorter Descriptions Has No General Stopping Rule
Candidate programs can be run in parallel. When one halts with output x, it certifies an upper bound on ; continued search may find a shorter program. No general stopping rule certifies that the shortest program has been reached or that no program below a chosen threshold will ever halt. Existence below a threshold is semidecidable, whereas nonexistence is not. This is how an optimization problem about finite records can fail to have a terminating algorithm.
Appendix A.4. The Projection Selects the Record to Be Modeled
The map turns a microscopic history into the record visible at the observer’s chosen scale. It may discard microscopic distinctions, but lossiness alone does not make a representation useful: a constant map discards every distinction and preserves no task-relevant variation. The paper therefore treats relevance as a constraint derived from the observer’s objective and measurement requirements. Once is fixed, the barrier theorems ask whether its output history is compressible and whether a general procedure can construct a short description. This is the fixed-representation subproblem. The compression-certified form of algorithmic emergence evaluates the model-compressor as a pair: determines what is retained and omitted, while the finite parametrized model M compresses the retained information and transfers to new observations. When is supplied, this is the fixed-representation subclass of the joint search described in Section 5.
Appendix A.5. Terminology Guide
Several ordinary words carry technical meanings in different parts of the paper. The following guide fixes each meaning.
- Finite instance. One specified micro-experiment and its observation record.
- Uniform procedure / open-ended family. One algorithm used across inputs whose sizes have no fixed upper bound.
- Answer computation with an effective stopping rule. Evaluation of a declared question q comes with a computable certificate of termination, such as a supplied finite horizon or exact cycle detection on a finite reachable state graph.
- Open-ended answer computation. Evaluation of q has no supplied finite stopping condition and may encode an unbounded computation. This is the setting in which answer-side undecidability may arise; it is distinct from the finite-record construction barriers.
- Description frame C. The observer together with the coding conventions and modeling resources already available before the acquisition being studied. Costs already in C are not charged again.
- Objective and task. specifies what the observer seeks to minimize. Relevance states which distinctions must be retained to evaluate or pursue it. A task specifies input/output behavior; a question q specifies an answer to be computed. Neither is identified with the objective’s numerical value.
- Representation / projection . A computable map selecting which distinctions are retained under the declared relevance requirement. “Coarse-graining” is used in this informational sense and need not mean spatial averaging.
- Retained kernel . The information kept by and modeled. The word “kernel” here means the retained task-relevant record, not a null space in linear algebra.
- Residual . The omitted detail needed only if one wants to reconstruct the original fine-grained record exactly.
- Realization identifier . Extra information selecting one realized history from a family of histories allowed by a law plus a boundary or global prescription.
- Program / algorithm / function. A program is a concrete description; an algorithm is, in the convention used here, an equivalence class under a declared notion of procedural sameness; a function is input–output behavior. Following Yanofsky, these levels are related by quotient maps rather than literal subset inclusions.
- Compressive core. A shortest program with the model’s declared task behavior, when that behavior yields systematic coding gains. Its length is .
- Model-compressor . The acquired representation–model pair. The pair is charged relative to C; determines relevant omission, while M compresses the retained kernel and is reused prospectively.
- Compression across observations. A shared model and jointly coded parameters describe the retained sequence. Individual chunks need not be shortened, and both margins are fixed total savings rather than per-chunk requirements.
Appendix A.6. Symbol Reference
Main-text symbols are listed first; symbols used only inside an appendix or a worked example are collected at the end. Tuple arguments of K use the fixed self-delimiting pairing convention specified in Section 2.
| Symbol / term | Meaning |
| Symbol / term | Meaning |
| AIT and computability | |
| U | Fixed universal prefix-free Turing machine. |
| p; | Self-delimiting program; its length in bits. For a binary string x, likewise denotes bit length. |
| Prefix Kolmogorov complexity: length of the shortest U-program that outputs x. | |
| ; | Complexity of the fixed integer encoding of n; a method-dependent description constant. Proposition 1 and Proposition 2 use such constants for their respective codes. |
| Conditional prefix complexity with y supplied on the auxiliary input tape. | |
| Relational description threshold: there exists a U-program p with and . Appendix C.1 gives the corresponding foundation-neutral reading of the barrier results. | |
| ; | Symmetric mutual algorithmic information used here: , or conditionally . |
| Complexity of a tuple under the fixed computable self-delimiting pairing convention in . | |
| ; | Finite binary strings; nonempty finite binary strings. |
| . | |
| , | Additive coding terms independent of the varying instance; constants may depend on U, , and fixed encodings. |
| total computable | Halts on every admissible input. |
| uniform | The same procedure must work over an admissible family of finite inputs with no fixed global upper size bound. |
| semidecidable | Positive instances are eventually recognized; the computation may fail to halt on negative instances. |
| dovetailing | Fair interleaving of candidate computations so every halting candidate is eventually observed. |
| compressible | In the raw-length sense used in Section 3.2, ; is a data-length benchmark, not a prefix-program length. Relative to a supplied threshold b, compressibility means . |
| microlaw | Informal term for the microscopic update rule (sometimes together with fixed background conventions). Knowing the microlaw is distinct from knowing the realized initial state . |
| Computable boundary/selection prescription used in SubSection 3.5. In the special law-selected-IC case, . Distinct from the shared description frame . | |
| Family of histories admitted jointly by the dynamics and global prescription . | |
| Family of macrohistories induced by a candidate projection over all initial states, Eq. (33). | |
| Residual realization-identifying information sufficient to identify one realized history within when the family is not a singleton. | |
| simulation / generation | Computing the finite record from the complete micro-experiment . This is distinct from finding a shorter reusable parametrized description of the macrodata. |
| macromodel construction | Discovery of a reusable parametrized model M at the observational scale. With supplied, only M remains to be found; the joint search selects . |
| algorithmic emergence | Acquisition of objective-relevant reusable regularity beyond an agent’s prior frame. Definition 3 specifies the compression-certified subclass used in the transfer results. |
| ; additive description regret | Excess output length above optimal complexity, when A returns a valid standalone description; a static gap in bits, not decision-theoretic regret. |
| Observer framework | |
| Observer-agent. Conceptually finite at each finite time; the formal barrier theorems do not depend on its architecture. | |
| Shared coding, admissibility, and pre-acquisition descriptive/modeling resources available to the observer. | |
| ; | Conceptual substrate state and finite substrate history. |
| Fixed observer/description frame, matching the notation of the persistence companion. | |
| Observer’s objective function, with its evaluation context and relevance requirements fixed before selection. It may depend on M and its parameters or evaluate the program more generally. | |
| Computable observation/projection or coarse-graining map. “Coarse” is observer-relative: the map discards distinctions irrelevant to the declared task and need not be a spatial average or a thermodynamic coarse-graining. | |
| ; | Conceptual projected observation and observation history, with . |
| Modeling Engine of observer . | |
| Full world-model retained by the observer at time t. | |
| Candidate macroscopic submodel indexed by j; this is the same role denoted in Pattern, Persist!, with j used here to avoid collision with the structure-function budget . | |
| Micro-experiments and theorem witnesses | |
| Finite micro-experiment. | |
| ; N | Finite microstate alphabet; number of microscopic degrees of freedom. |
| ; | Computable micro-update rule; encoded microstate . |
| ; r | Finite readout alphabet; for , the readout block width. Section 4.1 reuses r for a residual target, Section 7 for the population firing rate, and Corollary A1 for a compression witness. |
| H | Finite observation horizon of a micro-experiment; the macrohistory contains readouts. |
| T | Signed offset used only for reversible complete-state comparisons in Section 6, matching the companion notation. |
| Sufficiently complete finite encoded state used in Section 6 for reversible complete-state bookkeeping. | |
| Fixed computable reversible complete-state evolution over signed offset T in Section 6; for a reversible finite micro-experiment this role may be played by . | |
| ; | Complete-state projection induced on the encoded state in Section 6; retained complete-state kernel . |
| Variable-length self-delimiting index identifying inside the finite fiber of ; no equal-fiber-size assumption is made. | |
| Computable reversible self-delimiting representation . | |
| Local shorthand for the fixed conditioning tuple in the conservation/localization subsection. | |
| Binary encoding of the macrohistory generated by micro-experiment . | |
| , | Cyclic shift rule and leading-cell readout used only for the residual-information witness, Theorem 1. |
| , | Identity update and full-state readout used for the arbitrary-record embedding in Theorems 2–3. |
| Arbitrary n-bit record. In Theorem 1 it is the initial register state with ; in the computational reductions . | |
| Common conditioning information used in the conditional auxiliary result of Appendix D; is its embedded counterpart in Corollary A2. | |
| d | Complexity deficit allowed in Theorem 1; the lower bound is . |
| A, B | Total computable procedures returning standalone descriptions in the discovery/optimality results and their auxiliary variants; the input domain is specified in each statement. |
| b | Supplied program-length threshold in open-ended search and Appendix D; fixed chunk width in Proposition 2 and Appendix D.6. |
| ; | Open-ended below-threshold program search; hypothetical computable deadline excluded in Appendix D. |
| s; | Promised bound on initial-state complexity and fixed simulator overhead in SubSection 3.5. |
| g; J | Explicit total parametrization of candidate initial states in SubSection 3.5 and the (finite, known) number of candidates. Section 7 reuses J for the Ising coupling. |
| c | Fixed additive tolerance in Theorem 3. |
| ℓ | Fixed compression margin in Corollaries 1 and A1: the witness satisfies . |
| Chaitin certification constant for formal theory F and machine U. | |
| Structure, regularity, and model-compressors | |
| Finite-set model containing record x in algorithmic statistics. | |
| Budget on model complexity . | |
| Kolmogorov structure function: minimum residual index length at model-complexity budget , with value if no finite-set model is feasible. | |
| Constrained two-part/MDL length . | |
| Net structure gain over a budget interval, Eq. (25); a crack is an interval on which it is positive beyond the coding uncertainty. | |
| Shared finite model: jointly encodes retained observations through realization coordinates and regenerates them under the declared interface. | |
| ; model-compressor | Factorization separating the choice of what is modeled from the reusable model itself: defines the retained/residual split and M organizes and encodes the retained kernel. Figure 1 calls the pair a representation–model pair. In scientific applications this is a scientific model/compressor; the formal object is agent-generic. |
| ; ; | Declared task; task-relevant behavior induced by M; equivalence of implementations that realize the same behavior on the task domain. |
| Task-relative functional core complexity: length of a shortest program, relative to the declared task and pre-acquisition frame, that reproduces M’s task-relevant behavior. | |
| compressive core | A shortest representative of a model’s task-equivalence class when its behavior yields systematic coding gains. |
| D; ; | Fine-grained finite record; retained algorithmic kernel ; residual index completing the lossy projection so . |
| Fixed finite computable record domain used when interpreting the projection fiber as a finite-set model. | |
| ; q | Admissible projection class under frame C; a fixed macroscopic question in the answer-side results, distinct from the objective . |
| w; (Problem 1) | Finite code for a construction instance and the valid representation–model pair returned by one total computable constructor on that family. The code declares the construction inputs, frame, objective and relevance requirements, admissible classes, baselines, margins, and evaluation protocol. |
| Incremental self-delimiting code needed to add the pair beyond the pre-acquisition frame C. If is supplied, its incremental cost is zero and the ledger describes the fixed-representation subclass. | |
| ; | Incremental code cost of M alone relative to C; with fixed, abbreviates . These are actual code lengths, unlike the minima K and . |
| ; | Retained lossy model-code length for relative to C; corresponding lossless AIT ledger after adding the residual . Either is a compression only when shorter than its declared baseline. |
| ; | Explicit two-part length for a held model at fixed ; its unknown excess over . |
| ; | Realization coordinates or states; conditions supplied to both encoder and decoder. Information inferred from the observations is coded or recovered from already decoded data. |
| ; ; | First m observation chunks, their retained sequence, and its jointly represented realization coordinates. |
| Actual joint parameter-code length under the acquired model, including framing and corrections needed for exact retained-data reconstruction. | |
| ; | Predeclared valid code lengths for the original and retained observation sequences, under their fixed encodings and supplied conditions. Separate per-chunk sums are allowed special cases. |
| Shared history or model state available from the construction observations for prospective coding; additional state is charged. | |
| , | Fixed positive total margins for net compression and prospective reuse; independent of the number of chunks. |
| ; | Bridge information not already available from the conditioned microlaw/frame; fixed restricted-class construction map in Section 7. |
| Appendix-local and example-local symbols | |
| ; | Approximate core benchmark at tolerance under a predeclared discrepancy , Eq. (A8). |
| ; ; Q; | Declared baseline distribution and its ideal codelength; competing (semi)measure and its codelength, Proposition A1. |
| ; ; | Countable candidate codes, prefix code length of the index m, and the two-part selected codelength in Appendix B. |
| ; ; ; | Fixed acquisition cost, baseline and joint parameter-code lengths, and net gain after N observations, Eq. (A9); is an incremental saving when using sequential codes. |
| Actual joint residual-code length needed to restore the original observations, Eq. (A3). | |
| Test code fixed by the frozen pair after the construction records , Eq. (A15). | |
| Maximum halting runtime of U-programs of at most b bits, under fixed step-counting semantics. | |
| ; | Description relation and the negation of in Appendix C.1. |
| r; p (Corollary A1) | Explicit compression witness and explicit competitor that beats the constructor’s output. |
| F; | Computably axiomatized formal theory in Appendix C.3; finite-state transition map in Appendix C.4. Both are distinct from the reversible evolution . |
| ; ; e | Larger effectively encoded system class, its finite record map, and the computable record-preserving embedding in Corollary A2. |
| ; ; ; ; ; ; | Fixed family of microscopic specifications, fixed projection and question, declared answer map and answer set, and the constructor–evaluator pair in Appendix E. |
| ; ; ; | First m-chunk block left unshortened by A; complete returned code; reusable search-and-replay model; best qualifying complete length in Appendix D.6. |
| J, h, m, ; , r, v | Ising coupling, field, magnetization, and inverse temperature; neuron membrane potentials, population rate, and mean potential. Local to the examples of Section 7. |
Appendix B. From Reusable Regularity to Compression and Prediction
Repeated use can turn a model’s coding advantage into a net saving after its fixed cost is paid. This appendix gives the full code accounting, the functional-complexity benchmark, the amortization condition, and baseline-relative evidence from construction and prospective reuse.
Appendix B.1. Joint Retained Codes and the Full Record
The split used in Figure 2 separates the record from what the observer retains:
The representation keeps and omits . For a sequence, a shared model and its joint parameter code reconstruct exactly. An additional joint residual code restores the original observations. Write for its actual length, including framing, so that
A sum of separate residual-code lengths is a special case. The residual code and its interpretation must be specified by the charged representation or prior frame. Thus the usual MDL phrase “data given the model” can contain both the realization coordinates the model retains and the detail the observer omits. This is the operational counterpart of the structure-function accounting.
With common arguments suppressed, the baseline difference separates as
The first term is the saving due to the chosen representation under the declared baselines; the second is the net compression required by Definition 3(iii). Neither is an invariant measure of omitted information. Objective-based relevance justifies omission; the second comparison asks whether a model also captures regularity in what remains.
For a fixed acquired model, the sequence encoder can code parameter values jointly or sequentially. Both conventions permit dependence across observations. All descriptions remain decodable given the same declared conditions. A representation, initial state, parameter value, or observation learned from the data cannot be inserted into those conditions for free.
For a held model, the observer knows the length of its code without knowing how close that code is to optimal. Fix , supplied conditions u, and one retained kernel . Define the explicit two-part length
Here abbreviates the acquired model-code cost with fixed. The associated held-model regret is
The first quantity is computable from the held code; the second is not computable in general because its benchmark is the shortest description. A newly found shorter model witnesses suboptimality, whereas failure to find one does not certify near-optimality. Theorem 3 gives the underlying record-level obstruction, and Proposition 2(b) gives the corresponding failure for complete sequence codes among qualifying model-compressors.
Appendix B.2. Complete Coding Accounts for the Glider and Cellular Automaton
The examples in Section 5.3 compare complete codes for the same retained observations. In each comparison the pre-acquisition frame C fixes the objective, readout, dimensions, boundary conditions, horizon, and decoding conventions; it supplies no new initial state to either decoder. Both margins are fixed at 128 bits. A test block starts a separate run after the acquired model is fixed. The accompanying example code contains the codecs, verifiers, and numerical results.
Glider: finite experiment and shared information. The microstate is a periodic binary grid G together with a cyclic register s:
The grid follows Conway’s Game of Life [16]: a live cell survives with two or three live neighbors among its eight neighbors, and a dead cell becomes live with exactly three. All cells update simultaneously. The register shifts cyclically by one cell, independently of the grid. Writing F for the Life update, the retained record and horizon are
The objective requires the grid movie and none of the register. Both competing codes reconstruct that movie, so omitting the register contributes none of the compared saving.
The shared frame supplies Python execution, a generic source-module interface, Life, and the geometry. The entire acquired source, including , the glider template, its phases and directions, encoder, decoder, and fallback, contains 2366 bytes. An Elias gamma code for 2367 supplies a 23-bit length header, giving bits. This unminimized source includes Life code already available in C; charging that copy is conservative. Its byte length is an implementation cost, not an estimate of K.
A matching movie costs one mode bit, two 8-bit position coordinates, two phase bits, and two direction bits: 21 bits. Their roles remain fixed across runs. The encoder checks the observed movie; forecasting from those coordinates additionally requires them before the forecast. The code is lossless on all nonempty finite binary torus movies of side length at least eight: it searches the finitely many glider configurations and sends literal frames after an escape flag if none matches. Geometry and movie length are shared.
Glider: acquisition and reuse against a sparse incumbent. The incumbent encodes an initial grid by a mode bit, a gamma-coded live-cell count plus one, and the live-cell positions, then simulates Life. For five live cells this costs bits. A movie inconsistent with Life uses a literal-frame escape instead. The acquired model therefore saves bits per glider run before acquisition. The protocol applies both codes separately to each run, with run counts and boundaries shared. Table A2 gives a passing construction and subsequent test.
Table A2.
Glider accounting in bits against the sparse initial-grid simulator. Construction uses 294 separately initialized glider runs; testing uses two further runs after acquisition.
Table A2.
Glider accounting in bits against the sparse initial-grid simulator. Construction uses 294 separately initialized glider runs; testing uses two further runs after acquisition.
| Quantity | Construction | Test |
|---|---|---|
| Acquired representation and model | 18 951 | 0 |
| Coordinates and mode flags | 6 174 | 42 |
| Correction data | 0 | 0 |
| Complete retained code | 25 125 | 42 |
| Sparse initial-state simulation baseline | 25 284 | 172 |
| Net saving | 159 | 130 |
| Required saving | 128 | 128 |
A single movie costs bits and fails against the 86-bit incumbent. At 292 runs the net saving is 29 bits; 294 runs are needed to reach the construction margin. Two later runs save 130 bits without paying acquisition again. A continuing run with its last grid already shared is a different protocol: a simulation match flag can cost one bit, so the 21-bit coordinate code provides no positive saving.
For comparison, an incumbent without the sparse-grid code stores the 65 536 initial cell values and a mode flag, costing 65 537 bits per run. The model then saves 46 565 bits on the first movie and 65 516 on a later run; a literal 32-frame movie costs 2 097 152 bits. These are alternative declared baselines, not additional gains in Table A2. Restoring the omitted register separately costs its 4096-bit initial state, giving a complete first-record description of bits. No sampled register seed is certified incompressible. Ordinary Life need not be reversible; this example concerns relevance and compression. The verifier checks 240 configurations against independent Life evolution, including boundary crossings, escape cases, and the minimum run counts above.
Cellular automaton: objective, domain, and decoder. The microscopic states are 360-bit rows on a periodic ring, evolving synchronously by rule 146. A cell becomes 1 for the neighborhoods 001, 100, and 111 and becomes 0 for all others; each neighborhood lists the left neighbor, the cell itself, and the right neighbor. The readout applies a fixed Boolean table to each aligned triple, producing 120 bits every three microscopic steps. The initial row and forty coarse updates give 41 rows and 4 920 retained bits. The AND task requires the fully occupied triples; the OR task requires the nonempty triples. Each objective and readout is fixed before its model search.
For each task, C supplies rule 146, the readout, alignment, dimensions, timing, and a generic interpreter for elementary rules and the codes below. The incumbent sends one mode bit and the 360-bit microscopic initial row, then simulates and projects: 361 bits. Its literal escape covers records inconsistent with the microscopic law. The original microscopic movie has the same 361-bit simulation code, satisfying the baseline restriction in Definition 3. The blocky initial rows used for the figure and test are generated reproducibly by seeds 7 and 23 in the verifier; those seeds and their generating recipe are not supplied to either decoder.
The acquired module contains an eight-bit projection table and an eight-bit coarse-rule table, each preceded by a seven-bit Elias gamma header for length eight: 30 bits in total. The module stores its own copy of the task’s supplied readout; counting that redundant copy is conservative. The interpreter and coordinate interface are preinstalled, so this is a table cost, unlike the glider’s source-module cost. These costs concern different declared prior frames and are not estimates of K.
The coordinate code contains a one-bit flag and the initial 120-bit coarse row. With flag zero, the rule generates the movie exactly. With flag one, a correction list follows: for k differing bits in the autonomously generated movie, encode by an Elias gamma code and then list the k positions in increasing order, using 13 bits per position among the 4 920 entries. The decoder flips those bits. This fixes a lossless code for every binary coarse movie of the declared dimensions, including movies for which the selected rule is inaccurate.
Cellular automaton: successful construction and a failed search. For each projection the search minimizes one-step bit error over the 256 elementary rules, breaking ties by rule number. AND selects rule 128 with no errors. Exhaustive checking of all 512 nine-bit microscopic light cones verifies exact closure for this projection and timing. The model therefore needs no correction data on either the construction or fresh test run. Table A3 records the saving beyond the known microscopic simulator.
Table A3.
Cellular-automaton accounting in bits for the declared table interpreter. The model is fixed before the separately initialized test run. The OR column uses rule 182 with the specified correction list; its negative savings are failures of this code.
Table A3.
Cellular-automaton accounting in bits for the declared table interpreter. The model is fixed before the separately initialized test run. The OR column uses rule 182 with the specified correction list; its negative savings are failures of this code.
| Quantity | AND, rule 128 | OR, rule 182 |
|---|---|---|
| Acquired projection and rule tables | 30 | 30 |
| Initial coarse row and flag | 121 | 121 |
| Construction correction data | 0 | 26 307 |
| Complete construction code | 151 | 26 458 |
| Microscopic simulation baseline | 361 | 361 |
| Construction saving | 210 | |
| Test acquisition cost | 0 | 0 |
| Test initial coarse row and flag | 121 | 121 |
| Test correction data | 0 | 27 440 |
| Complete test code | 121 | 27 561 |
| Test microscopic baseline | 361 | 361 |
| Test saving | 240 | |
| Required margin on each block | 128 | 128 |
OR selects rule 182, with 1 370 one-step errors among 4 800 observed transitions. Running that rule autonomously from the initial coarse row instead produces 2 022 differing bits across the construction movie, including 54 of 120 bits in its final row. Their correction list costs bits. The held rule produces 2 109 differences on the test movie, costing correction bits. Both codes reconstruct exactly after correction, but neither beats the microscopic baseline.
For the caption’s closure count, traverse the observed transitions in time and cell order, store the first successor for each coarse neighborhood, and count subsequent disagreements: zero for AND and 1 476 for OR. This count differs from the minimum one-step error because its reference is the first successor rather than the majority successor. Any such conflict rules out an exact elementary update on all those transitions. Its magnitude depends on the counting convention. The finite search and failed certificate concern the stated rule class and code; they leave other model classes and coding methods open. The verifier checks every local light cone, all 256 rules, exact reconstruction with and without corrections, and both coding accounts.
Appendix B.3. Functional Complexity and Implementation Cost
The acquired implementation M realizes its declared task behavior. A fixed interpreter can therefore reproduce that behavior from the incremental description of M, giving
where is the actual incremental description length relative to the prior frame. The excess can include unused code or precomputed information that improves execution speed. The bound concerns description length; it gives no runtime comparison.
For a predeclared nonnegative discrepancy between behaviors (for example, the largest output disagreement over the domain), vanishing on identical behaviors, and a tolerance , let on programs that halt on every admissible input. The approximate benchmark is
with the domain, discrepancy, tolerance, and supplied information fixed before comparison. This minimizes length over admissible approximations to the given behavior. Closeness within a tolerance need not be transitive, so it does not generally define an equivalence class.
Appendix B.4. Regularity Can Precede Net Compression
Let the acquired pair have fixed incremental cost . Suppressing common conditions, write for the baseline length of the first N observations and for their joint parameter-code length under the same model. The net gain is
For a sequential implementation, define as the difference between the baseline and model’s additional code lengths on use i, including any updated state or framing. Then , up to a fixed initialization cost already included in . Separate coding of each observation is one instance of this accounting; dependence across the parameter sequence is allowed.
If , the gain eventually exceeds any fixed . For constant additional saving per use, condition (iii) holds once . This is a sufficient condition for eventual amortization, not a requirement that every use save bits. A large implementation, including one that stores a recurring image literally, can therefore become compressive through reuse.
The cost is that of the acquired implementation even if a shorter equivalent program exists. Failure to repay it on a finite batch does not show that no useful regularity has been captured. Without a usable lower bound on the accumulated savings, this observation supplies no deadline. Proposition 2 concerns failure on the available finite batch even when another pair already meets the certificate there; it does not rule out later amortization by the original method.
Appendix B.5. Compression Relative to a Declared Baseline
Let be a probability mass function on a finite or countable set of possible records z, and write
for its ideal codelength; a Kraft-valid main-text baseline is of this form up to normalization, as discussed below. Let Q be any competing probability mass function or semimeasure, , with codelength . A semimeasure is enough because Kraft’s inequality associates every prefix code with weights whose total mass is at most one.
Proposition A1
(No hypercompression relative to a baseline). Let Z be a record drawn from . For every ,
Proof.
Let . For every , . Hence
□
Thus, if the baseline generated the data, a k-bit advantage by another valid code, fixed in advance or with its selection cost charged as below, is an event of baseline probability at most . This is the no-hypercompression inequality used in MDL theory [4,85]. For orientation, a 20-bit gain has baseline probability at most about . The statement is relative to : it is evidence against that baseline, not a probability that the discovered model is metaphysically true, nor a distribution-free theorem that the future must resemble the past.
A Kraft-valid baseline code can likewise be associated with subprobability weights by Kraft’s inequality; the same algebra then bounds those weights. A literal probability statement requires a declared normalized data-generating distribution such as above. In either case, the argument does not apply to an arbitrary numerical “baseline” that violates Kraft’s inequality. This qualification matters when description length is used operationally rather than probabilistically.
Appendix B.6. Model Selection Adds a Description Cost
Suppose a search ranges over a countable family of candidate codes , and the index m is itself described by a prefix code of length with . After seeing z, the search may choose whichever candidate looks best. The resulting two-part selected codelength is
It is still a valid code, because
Consequently Proposition A1 continues to hold for data-dependent model selection only when the information needed to identify the selected model is charged. If that cost is omitted, one can manufacture apparent regularity by searching a sufficiently rich family and reporting only the winning fit.
This is the coding-theoretic reason for keeping visible in Definition 3. Information discovered in the representation is no more free than information discovered in the model M. If was already supplied or inherited in the pre-acquisition frame C, its incremental cost is zero; if it is learned during construction, its code must be paid. The retained-kernel baseline then asks a separate question: after relevance-guided omission has fixed what is to be modeled, does M compress the information that remains?
Appendix B.7. Prospective Reuse and the Prediction–Coding Link
Compression of construction records alone does not provide a measure-free bound on future success. An arbitrary continuation can always violate a pattern inferred from a finite prefix. The paper therefore freezes and its coding rules before evaluation on a new block. Its state may evolve according to those rules. The common state is recoverable by the decoder from the charged construction code and declared conditions, or is separately charged. The incumbent baseline has access to the same shared history; it need not use that information in the same way. If an incumbent already exploits a repetition, rediscovering it need not improve that baseline. The literal baseline used in Appendix D.6 ignores prior observations by design and is fixed throughout that family.
The no-hypercompression statement has an immediate conditional form. Let denote the construction records, and let the frozen pair determine a test code before a held-out observation block Z is observed. For any declared conditional baseline ,
Conditional on the construction history, a large held-out compression gain is therefore again unlikely under the baseline. Under the declared sequential evaluation and baseline, repeated prospective success is stronger evidence that the discovered regularity is reusable rather than a one-record accident.
For sequential probabilistic models under logarithmic loss, the coding–prediction connection is exact at the level of ideal codelengths. Let denote the observed prefix, and let be the probability assigned to observation given its past. The joint probability satisfies
so predictive codelength is cumulative next-observation log loss [85]. Such sequential coding records how much probability the model assigned to observations as they arrived.
Log loss is convenient, not uniquely privileged for converting prediction into compression. Consider a deterministic sequential predictor of a binary sequence , with the predictor and all of its supplied inputs known to encoder and decoder. If it makes k mistakes, the encoder can transmit the set of mistaken positions; the decoder reruns the predictor and flips its output at those positions. Given n and k, the error pattern requires
bits, plus self-delimiting framing for n and k when not supplied, and any charged model or input information. Thus ordinary binary prediction already yields a lossless code whenever its errors are sufficiently sparse. More general discrete predictors admit analogous error descriptions, with any incorrect symbol values also encoded.
Predictive success can therefore provide compressive capability even in a large model. To transfer a particular compression barrier to a predictor-discovery claim, the predictor-to-code conversion must be made explicit and the theorem’s information, cost, validity, and guarantee assumptions preserved. For log-loss predictors the conversion is immediate; for ordinary discrete prediction it can be as elementary as the mistake code above.
Appendix B.8. Relation to Solomonoff Induction
Solomonoff induction supplies an idealized predictive justification for simpler generators. Its universal predictor weights a program p by and multiplicatively dominates every lower-semicomputable source up to a constant controlled by the source’s description complexity. For a computable data-generating source, the universal predictor’s expected cumulative excess logarithmic loss relative to that source is bounded by the source’s description complexity, up to fixed coding constants [8,9,10]. This concerns effective description length, rather than raw parameter count. It gives a predictive bound within the stated source class; it does not imply that the physical world is sampled from the universal semimeasure or is itself simple.
One possible world-side motivation, explored in earlier KT work, is that random-program generation strongly favors low-complexity outputs [11]. A broader result is that many low-complexity computable input–output maps exhibit an approximately exponential bias toward simple outputs even when their inputs are sampled uniformly [12]. These considerations motivate a simplicity-biased inductive stance; they are logically separate from Proposition A1.
Compression after model costs are paid provides evidence against a declared baseline; prospective gains test reuse beyond the selecting data. These guarantees concern the model’s achieved coding performance. They supply neither a shortest equivalent implementation nor a universal method for acquiring a qualifying model.
Appendix C. Description Relations, Uncomputability, and Certification
Numerical Kolmogorov complexity packages three distinct questions: whether a short description exists, whether its least length can be computed, and whether a lower bound can be certified. This appendix separates them and then compares finite description problems with physical systems that have unbounded resources.
Appendix C.1. A Relational Reading of Kolmogorov Complexity
Standard AIT writes the shortest description length as the numerical function
That notation is convenient, but it combines two logically separate ingredients: the underlying description relation and the step of packaging its least length as a numerical value. For the first ingredient define
The complementary statement
says that no description below the threshold exists. In standard classical AIT,
Thus many statements normally written with K can instead be stated directly in terms of descriptions and thresholds.
In constructive type theory, Forster, Kunze, and Lauermann formalize Kolmogorov complexity as a functional relation rather than as a total function. They prove classical totality only under double negation and emphasize that a constructive proof of ordinary totality would, under Curry–Howard, amount to a computable function for Kolmogorov complexity. They also show that, in their setting, the unrestricted principle assigning least elements to all inhabited predicates entails the law of excluded middle, and they choose the direct short-description predicate when formalizing compressibility [71]. A complementary classical analysis appears in reverse mathematics.
For plain complexity C (not the frame C of the main text), Nies and Shafer work in with the threshold relation “” defined by existence of a description. Their base theory proves that each individual string has a least description length using the least-element principle, while it does not prove that the noncomputable map exists as a second-order function object [86]. In HOL4, Catt and Norrish take the more classical route and define Kolmogorov complexity as a function; Forster et al. note that this packaging requires classical logic together with a unique choice operator [71,87].
These examples are not criticisms of standard AIT. They show that the description relation makes fewer foundational commitments, while the familiar numerical K packages the same information in a form convenient for classical mathematics.
The three barriers in the main text can be read at this relational level. The counting proof of Theorem 1 bounds how many records can possess a short conditional description; its lower bound can be read as the absence of any conditional program shorter than b. The discovery antecedent in Theorem 2 is exactly . The optimality barrier also has a witness-based form that avoids referring to an exact optimum:
Corollary A1
(Relational optimality barrier). Let A be a total computable procedure that returns a valid description for every finite micro-experiment μ. For every there are a finite micro-experiment μ and programs such that
Thus the record has a compression witness r, and another valid description p beats the output of A by more than c bits. Both exist for each such instance; no uniform procedure is promised to find them.
Proof.
Use the arbitrary-record embedding of Eq. (7) and . Define
Because A is valid and pad is injective, distinct z require distinct returned programs. Only finitely many programs have length at most any fixed k, so the total computable function f is unbounded. Given k, search strings in length-lexicographic order for the first with . The search terminates. A fixed program with a self-delimiting encoding of k can repeat this search and output , giving a valid competitor of length .
As k grows, leaves every finite bound, since each fixed z has a fixed value and is excluded once . Hence for sufficiently large k, the explicit program that contains self-delimitingly and appends zeroes has length
call this compression witness . Choosing k large enough also gives . Then
which proves the claim with . □
In classical AIT, Corollary A1 immediately implies the numerical regret statement because for every valid description p. Conversely, if a shortest description is available as a mathematical object, Theorem 3 supplies such a competitor by choosing a shortest program. The two presentations therefore have the same classical content, while the relational form makes the observer-side witness structure explicit. The paper uses K as the conventional compact benchmark, not as something the observer is assumed to compute or possess.
The distinction also controls the possible ontological reading of the barriers. Classical AIT may treat and the truth of as determinate even when no algorithm computes the former and no fixed effective theory certifies the latter without bound. The obstacle then concerns an observer’s access to a classically given object. For a finite micro-experiment, direct simulation still generates the finite record and answers computable questions about it. Failure to find a compressed macroscopic theory therefore does not imply that the corresponding individual answer is absent or not entailed by the microscopic specification.
An unbounded problem with a noncomputable answer map presents a stronger effective obstruction: the answer to the fixed macroscopic question cannot be returned uniformly over the family. In a constructive setting, where an existential assertion carries a demand for a witness, either failure can constrain which objects or properties are available within the theory; packaging the description relation as a total numerical function itself requires additional principles. Their ontological significance therefore depends on which constructive requirements the framework attaches to existence. Neither establishes that higher-level facts fail to follow from complete lower-level facts, as Chalmers’s strong emergence requires [65,71]. Section 8.2 compares these tasks with Chalmers’s question.
Appendix C.2. Classical Numerical Complexity Is Uncomputable
The discovery and optimality reductions depend on a common classical obstruction: shortest-program length is mathematically defined but cannot be returned by one terminating algorithm on all strings. Fix a universal prefix-free machine U and write in the usual numerical notation.
Theorem A1
(Kolmogorov, Solomonoff, Chaitin). There is no total computable function that, on every input string x, outputs .
Proof.
Suppose f is total computable with for all x. Consider a program Q that, on input k, enumerates in length-lexicographic order and returns the first y with . Such a y exists, because among the strings of length n at most have descriptions shorter than n bits, so some string of each length has . The program halts, and its total description length is once k is supplied self-delimitingly. Its output therefore satisfies , while the search guarantees . For large k these are incompatible, so f does not exist. □
Theorem A2
(Exact partial computation has finite domain). If is partial computable and whenever is defined, then is finite.
Proof.
Suppose the domain were infinite. Complexity is unbounded on any infinite set, since only finitely many strings have complexity at most a given bound. Given k, dovetail f on all strings and output the first whose returned value exceeds k. This search terminates, so for a constant depending only on the fixed search and machine. Correctness gives , contradicting for sufficiently large k. □
This is the classical partial-computation obstruction stated for plain complexity in [26]; the proof above gives the prefix version used here. A partial method returning a shortest program whenever it answers would yield such an f by taking program length. An always-valid compressor can still happen to be optimal on infinitely many inputs if it also answers suboptimally elsewhere.
In the standard classical numerical presentation, is a total benchmark but no terminating algorithm returns its exact value for all strings. Operationally, progressively better descriptions can lower an upper bound, while there is no computable stage at which one can know that the current bound is final for every input. The relation is semidecidable; its complement is not, since semideciding both sides would decide the relation. This is the content used by the discovery and optimality barriers; the preceding subsection shows that it can be stated without taking a total K-function as primitive. The separate classical limit on formal certification is recorded below.
The same argument with the auxiliary tape carrying a fixed string z shows that is not computable. The conditional auxiliary result in Appendix D needs a separate argument because its conditioning string varies with the input length.
Theorem 2 uses the related fact that the set of compressible strings is undecidable. Were it decidable, then for each n one could compute the length-lexicographically first with , which exists by the counting argument of Appendix D.1. That x would then have the description “the first incompressible string of length n” of size , contradicting for large n.
The arbitrary-threshold corollary in Appendix D then follows by specialization and requires no further classical input. Undecidability of the full threshold relation is itself a consequence of the diagonal case through the computable reduction . The reverse implication is invalid for an arbitrary diagonal slice, so the diagonal case is proved directly. The decision problem “is ?” is Turing-equivalent to the halting problem [5]. Moreover, K is upper semicomputable: it can be approximated from above, with no computable bound on the rate of convergence [26]. Thus running a compressor gives an upper bound on ; there is no effective certification procedure that succeeds on every true claim that no shorter description exists.
The argument tolerates any fixed slack. For each fixed , the predicate is undecidable by the same construction: a decision procedure would let a short program output, for each n, the first string of length n that is not ℓ-compressible, describing a string of complexity at least in bits. The construction in fact excludes any decidable separator, that is, any test that accepts every string compressible by more than the margin and never accepts an incompressible one: if a decidable set P contains every x with and only x with , then the first length-n string outside P is computable from n, exists because incompressible strings lie outside P, and has while being describable in bits. This is an independent classical strengthening. Proposition 2 instead uses the direct counting construction in Appendix D.6.
Finite catalogs and the Busy Beaver stopping bound. Fix the machine U, its output convention, and a step-counting semantics. Define
with value zero when the set is empty. A known upper bound on suffices to classify halting by simulating all inputs through b bits for that many steps each. There is no total computable upper bound valid for every b, since one would decide halting for U. Andreev compares Busy Beaver functions defined through plain, prefix, and a priori complexity [25]. This bound supplies a sufficient exhaustive stopping rule, not the necessary work for deciding a particular output: a slow program may print a different string, or one with a much faster description.
More directly, no total computable function of x bounds the runtime of a fastest shortest program for x. If one did, simulating every program up to a known literal-description length bound for that many steps would identify a shortest program and compute , contradicting Theorem A1. This also excludes a bound depending only on ; it does not give a numerical runtime lower bound for any specified short record.
Suppose halting has been completely classified through b bits and all halting outputs obtained. For any string represented in that finite catalog, its shortest listed program has length exactly , because every shorter candidate is covered. An absent string has . Thus compressibility is decided for every , but exact complexity for every such string does not follow: some may require longer programs. For a particular candidate, one can sometimes exclude every shorter program as a description of x without settling all of those programs’ halting questions. Full halting classification is sufficient, not necessary.
Calude, Dinneen, and Shu report a complete halting classification through 84 program bits for their specified universal self-delimiting register machine, and 64 initial bits of its halting probability [88]. By the preceding argument, their catalog settles compressibility through 85 record bits under that output convention; it does not determine exact complexity for all those records. These cutoffs depend on the interpreter and are not a universal frontier. The obstruction is to an effective method extending such complete answers to all lengths, while individual cases and certified finite ranges can be settled.
Appendix C.3. Certification Is Theory-Relative
A logically distinct limitation concerns formal certification. Let F be a computably axiomatized formal theory, strong enough to express the relevant statements about U and sound for them. There is a constant , depending on F and U, such that F proves no statement with [5,76]. If F proved arbitrarily large lower bounds on K, a program could take a threshold m, enumerate the theorems of F until it found a proof that some string exceeded that threshold, and output that string: a description of a string of complexity greater than m in bits, a contradiction for large m.
Since certifying that a program p is shortest for x requires ruling out every shorter description, such certificates are unavailable beyond a constant fixed by F and U. The constant is theory- and machine-dependent, not a universal physical scale. This classical limitation is separate from the micro-to-macro construction barriers of the main text; it is machine-checked as chaitin_certification_ceiling.
The statement has a finite witness when true: exhibit a program shorter than b that prints x. Its negation says that every such candidate fails, and continued simulation by itself need never produce a finite certificate. In the usual classical presentation, the least threshold is packaged into the numerical quantity ; constructive formulations need not make that same move. The barrier results require only the stated limits on effective construction and certification. They do not require the additional philosophical claim that every proposition that no finite observer can establish must nevertheless correspond to an observer-independent mathematical or physical fact.
The incompleteness literature should be read with the same separation. The Gödel–Rosser theorem states, in the usual setting, that a consistent effectively axiomatized theory strong enough for elementary arithmetic has a sentence that is neither provable nor refutable in that theory [89]. This is a statement about derivability relative to a specified formal system. Calling an independent sentence “true but unprovable” adds a semantic evaluation in an intended model; that step is not supplied by syntactic independence alone. A stronger consistent effective theory may settle a particular sentence while remaining incomplete in turn. Chaitin’s certification ceiling has the same theory-relative character: changing F or the reference machine changes the constant, so the theorem rules out unbounded certification within one fixed formal system rather than every possible route to knowledge of an individual instance.
Appendix C.4. Finite-State Boundedness, Compact Continua, and Recurrence
The finite-state argument used in Section 8 depends on effective finiteness, not on geometric boundedness. Let be a computable deterministic transition on an effectively represented set with , and let membership in a target be decidable. From any initial state , simulation must either reach A or repeat a previously visited state within visits. If with , determinism gives for every ; the future is periodic from that point onward. Eventual reachability is therefore decidable by exact cycle detection. The value of N need not be supplied to the algorithm: decidable state equality and the promise of finiteness suffice for termination. A bounded-tape Turing machine is an example.
A compact positive-dimensional continuum is different. Compactness gives a finite cover at every fixed resolution , but no resolution-independent finite set of exact states. A simple example is an irrational rotation of the circle: every orbit returns arbitrarily close to its starting point but never returns exactly. More generally, Poincaré recurrence states, under the usual finite-measure and measure-preserving hypotheses, that almost every point returns infinitely often to neighborhoods of its initial state; Hamiltonian flows on finite-measure invariant regions fall under this framework because they preserve Liouville measure [90]. Recurrence is therefore much weaker than finite-state periodicity. It supplies neither exact repetition nor, in general, a computable return-time bound that could serve as a negative reachability certificate.
A compact continuum can nevertheless carry an effectively unbounded symbolic state space inside a geometrically bounded region. Moore made this explicit using generalized shifts: the two halves of a Turing tape are encoded in the digits of a point in a Cantor set, and iteration updates those digits as a Turing machine updates its tape. He showed that smooth systems with as few as three degrees of freedom, including the motion of a particle in a three-dimensional potential, can realize universal computation [52,55]. Approximate recurrence does not identify two computational states: points that agree to finite precision may differ in deeper tape digits and hence in their subsequent symbolic evolution. Continuity alone is not sufficient for universality; it merely removes the finite-state pigeonhole argument. The dynamics must organize the available fine distinctions into an effective simulation.
Compact conservative and fluid systems show that neither compactness nor conservation restores the finite-state conclusion. Cardona et al. realized a Turing-complete area-preserving return map as a stationary incompressible Euler flow on a suitably metrized [56]. Dyhr et al. construct Turing-complete stationary Navier–Stokes solutions on Hodge-admissible Riemannian three-manifolds after, in general, a non-small metric deformation; the resulting fields are harmonic and the undecidable predicate is Lagrangian reachability of a computably specified open target [57]. The result therefore does not establish Turing completeness of time-dependent Euclidean Navier–Stokes evolution, undecidability of singularity formation, or blow-up. Tao’s averaged-equation blow-up theorem and proposed fluid logic program concern that different, time-dependent Eulerian direction [58,59].
This mathematical distinction should not be confused with a claim about unlimited physical precision. Exact continuum encodings can use distinctions at progressively finer scales. If an operational physical theory instead has only finitely many distinguishable states with computable deterministic transitions, exact cycle detection restores reachability decidability. Noise or finite measurement precision may prevent an indefinitely stable realization of a continuum encoding without changing the exact mathematical theorem. Whether nature supplies the required fine-scale state capacity is therefore a separate physical question.
These continuum examples concern the answer-side undecidability barrier; the finite-record AIT construction barriers require no continuum resource and remain logically distinct, although one substrate may exhibit both. These are computability statements, not claims about whether every inaccessible value is an observer-independent property of the world.
Appendix D. Proofs and Supporting Construction Results
This appendix records the counting argument behind Theorem 1 and closes four apparent loopholes in the construction barriers: unbounded substrates, a supplied length threshold, complete knowledge of the law, and a promise that open-ended search eventually succeeds. The final subsection gives the periodic-block counting construction and the full proof of Proposition 2 for reusable model-compressors.
Appendix D.1. Counting Proof for Theorem 1
Fix a conditioning string z and write for complexity relative to U with z on the auxiliary tape. There are binary strings of length strictly less than m, so there are at most candidate programs; restricting the machine domain to a prefix-free set can only reduce this count. Since a deterministic program has one output, fewer than strings possess a description shorter than m. Hence
Setting gives (6). Because z is fixed across the ensemble, the bound is independent of the chosen shared microlaw and observation map. It fails if the conditioning information varies with the initial state: when z contains y, . The count itself needs no probability; reading it as a fraction of initial states uses the uniform counting measure, and a different prior on gives a different fraction.
Appendix D.2. Finite-Horizon Embeddings Transfer the Barriers
Let denote the class of finite micro-experiments in Definition 1. Let be an effectively encoded class of systems whose state space or physical substrate may be unbounded, but where each input specifies a finite observation horizon and hence a finite binary record .
The main witnesses are finite, but an unbounded substrate family may contain them as finite-horizon instances. Their construction barriers survive such an enlargement whenever the embedding is computable and preserves the observed records.
Corollary A2
(Inheritance under finite-horizon embeddings). Suppose there is a computable map satisfying
Then the discovery and optimality barriers hold on : no total computable procedure on can satisfy the guarantee of Theorem 2, and no such procedure can satisfy the fixed-regret guarantee of Theorem 3, even on the embedded compressible records. If, on the embedded shift-register family, the conditioning string (the encoded embedded law, projection, horizon, and frame) depends on n but not on the initial state , then the residual-information bound also transfers:
for every integer d with .
Proof.
For fixed n, Eq. (A28) gives . Since is fixed across the ensemble, the counting argument of Theorem 1 applies without change and gives Eq. (A29). For the computational statements, suppose a total computable procedure with either claimed guarantee existed on . Composing it with the computable map e would give a total computable procedure with the same validity and length guarantee on , contradicting Theorem 2 or Theorem 3. □
A standard instance adjoins an infinite inactive tape or lattice to a finite experiment, lets the original dynamics act on a marked finite region, and lets the observation map ignore the added degrees of freedom. The horizon remains finite and the record is unchanged. Thus an unbounded substrate class admitting this effective extension inherits the barriers. Set-theoretic inclusion alone is insufficient: the embedding must be computable and must preserve the observed record; the conditioning data in the residual-information ensemble must not encode the varying initial state.
Appendix D.3. A Supplied Length Threshold Does Not Restore Construction
A target code length does not provide a program meeting that target or a stopping rule for deciding that none exists. The threshold formulation therefore retains the discovery obstruction.
Corollary A3
(Arbitrary thresholds). There is no total computable procedure B that, given a finite string x and threshold , always returns a program with and is guaranteed to satisfy .
Proof.
If such a B existed, then on input x one could compute . Validity gives , while the guarantee gives the converse. Comparing the two lengths would decide whether , contradicting Appendix C. □
Thus supplying a target code length in advance does not restore a uniform guarantee. This result is an immediate threshold variant of the discovery obstruction, not a separate construction barrier.
Appendix D.4. Known Laws Do Not Make Conditional Complexity Computable
Supplying the common microlaw, observation map, horizon, and frame changes the benchmark from unconditional to conditional complexity. It does not make that benchmark computable when the same procedure must work across the growing family.
Corollary A4
(Conditional form). There is no total computable procedure that, for every nonempty string y, returns , where is the common microlaw, observation map, register length (which fixes the horizon ), and frame of Theorem 1. Conditioning on exactly known laws does not restore computability.
Proof.
Suppose were total computable. On input n, enumerate and return the first with . Such a string exists by the counting bound of Appendix D.1. Given , a program of constant size can recover n and run this search. Hence , contradicting for large n. □
The varying condition is why this statement needs its own diagonal argument rather than the standard theorem for one fixed conditioning string.
Appendix D.5. Open-Ended Discovery Has No Computable Schedule
Let denote the open-ended search that enumerates all programs shorter than b, runs them dovetailed, and returns the first that outputs x.
On a positive instance this search eventually succeeds. Turning that fact into a constructor that terminates on all inputs would require a computable deadline. No such uniform deadline exists.
Proposition A2
(No computable schedule). There is no total computable function τ such that, for every pair with , the search returns a program within computation steps.
Proof.
Suppose such a existed. Run for steps and report whether a program has been returned. If , the search returns within the allotted time; if , no program shorter than b outputs x. This would decide the threshold relation, and in particular its diagonal case , contradicting Appendix C. □
Hence a valid promise guarantees eventual discovery without supplying a computable universal deadline.
Appendix D.6. Periodic Observation Blocks: A Counting Proof
This construction proves Proposition 2 directly. It uses the same counting idea as Proposition 1, with the observation count supplied to the decoder and a reusable code that remains valid outside the constructed stream.
Domain, objective, and shared information. Fix a chunk width , frame C, and positive total margins . A block contains bits. Its stream repeats that block indefinitely; construction uses its first m chunks and testing uses the next m, obtained only after the pair has been selected. The objective is faithful reproduction of the designated chunks, measured by
with required discrepancy zero. It can be evaluated on the output of M at the queried phases. The representation is supplied: and . It retains every distinction required by this objective, and there is no omitted residual to provide a saving. This is a permitted fixed-representation subclass of the joint problem, without an auxiliary discarded bit.
The decoder is given C, b, m, the queried block length, and the starting phase. These are common across all blocks at fixed m; the data v are not supplied. A complete code includes the acquired model, the representation if not already supplied, the joint parameter code, and all framing. The baseline is literal concatenation, of length on each construction or test block. At supplied length it is a valid fixed-length code. During testing it continues to ignore past data, although those data may be available to both codes. This baseline choice is declared before selection and does not change with the returned model.
The stream also has a finite microscopic realization: store v in a cyclic register, advance by b bits per step, and read its first b bits. The update and readout conventions are fixed, with m specifying the register length. A horizon yields one block. If the constructor’s input includes a complete micro-experiment, construct this specification effectively from v; it remains input to the constructor, not free information for the returned code’s decoder.
Proof
(Full proof of Proposition 2). Fix a total computable constructor A satisfying the proposition’s validity condition. For each m and , run A on the corresponding instance and run its returned encoder on v. Let be the resulting complete code, under the fixed convention above. This is a total computable map: both A and every returned encoder terminate on their declared inputs. A single fixed decoder reconstructs v from and the common supplied conditions. Consequently, for fixed m, the map is injective. Since
there is at least one v with .
Let be the first such block in lexicographic order. A finite search computes from m: enumerate the blocks, compute each , and stop at the first whose length is at least . Only code lengths are tested; the search need not decide equivalence, totality, or optimality of arbitrary programs. Its termination follows from the stipulated totality and validity of A, and from the counting argument.
Construct one model containing A and this search. Given m and a queried phase, it computes and generates the requested finite segment of its periodic repetition. Its encoder compares the observed segment with that prediction. It emits a one-bit match code 0 if they agree; otherwise it emits 1 followed by the literal segment. The decoder regenerates the prediction in the first case and reads the supplied number of literal bits in the second. At every fixed supplied block length, the code set consisting of 0 and all strings of the appropriate literal length is prefix-free. Both maps are total for every finite segment. This is a stable phase-and-correction interface: the model carries the recurring pattern, phase is a supplied condition, and the code records whether corrections are needed. The escape case preserves validity even when the observed stream does not follow the prediction.
Let bound the actual complete construction-code length of this pair on its matching block, including the model program, its fixed framing, and the match bit. This is a constant in m; it includes the cost of describing A. The supplied identity representation has zero incremental cost. Thus, on the first block and the later block from the same periodic stream,
The model meets relevance and reuse throughout the declared domain. For every m large enough that and , it also meets both compression inequalities of Definition 3. The test block is not used to select the model: by construction, the continuing stream supplies the predicted block afterward.
The code returned by A has length at least and fails (iii), proving (a). Let be the best qualifying complete construction-code length on this instance and its declared continuation. Since , part (b) follows from the growing lower bound
□
The fixed constant uses the fact that m is supplied. If the decoder must also learn m, encode it and include that length; the corresponding comparison gains the cost of describing m. The search inside can take an enormous time, so the result supplies no practical runtime bound for a particular short block. Nor does it say that learning recurrence is impossible: after receiving more copies, a method that stores and reuses may meet the certificate. The proved failure is on the available construction block, where a compact qualifying model already exists. A universal joint constructor must handle this supplied-representation subclass, but the argument makes no separate claim about the difficulty of selecting .
Appendix E. Composition Links Construction and Answer Barriers
This appendix records the closure observation represented by the dashed implication in Figure 7. Construction and evaluation are different tasks, but a pipeline that performs both is itself an algorithm from the microscopic input to the macroscopic answer. If that answer map is noncomputable, at least one stage of every proposed universal pipeline must fail to be total or correct. This is not an additional source of undecidability. Its role is to state the one-way dependence between the answer barrier and a constructor–evaluator pipeline.
To state the argument, fix an effectively encoded family of microscopic specifications, a projection , and one macroscopic question . The formal semantics of the problem assigns each a finite answer in an effectively encoded set . Write this declared answer map as
For example, may ask whether the -projected trajectory ever enters a target set. Over an open-ended time domain, the resulting binary map can encode the halting problem. Totality of is part of the formal problem specification; the proposition does not derive it from a claim about observer-independent physical truth.
Suppose a total computable constructor turns every microscopic input into a finite model code, and a total computable evaluator extracts the answer to from that code. The proposed detour is the composition .
Proposition A3
(Composition rule for a fixed macroquestion). Let and be total computable maps. If
then is computable. Consequently, if is noncomputable, no such total and universally correct constructor–evaluator pair exists.
Proof.
Given , first compute and then . The assumed equality makes this composition an algorithm for . The final statement is the contrapositive. □
The proposition concerns a uniform pair that works for the entire family . It permits a useful model for an individual , and it says nothing about whether a model compresses, transfers to new data, or answers questions outside . It also separates the two algorithms only for interpretation: the impossibility applies to their composition, not necessarily to either stage considered alone.
The same reasoning applies to a scheme advertised as correct for varying projections and questions. Fixing any admissible produces the factorization in Proposition A3. Holding one admissible question and representation fixed gives a slice of any scheme that promises to handle them all; a single noncomputable slice therefore rules out such a scheme.
The converse is only an encoding fact. At the fixed-task level, a computable answer map always has a trivial factorization: can encode the entire microscopic input , after which decodes it and runs the algorithm for . This “model” need not compress the input, transfer to new data, or expose reusable macroscopic structure. Computability of the answer is therefore not sufficient for algorithmic emergence.
The fixed coarse-graining is an essential hypothesis. If a constructor may choose without an independent relevance constraint, it can avoid the hard question by discarding the distinction the question asks about. The obstruction applies when is supplied, when the constructor must work for every admissible , or when the relevance criterion requires the selected question to be preserved.
Appendix F. Machine-Checked Formalization
This appendix records the declaration-level correspondence between the paper and the Lean 4 development. BCOM WP0195 gives the full inventory and axiom audit.
Appendix F.1. Main Barriers and Finite-Horizon Inheritance
The declarations named below occur in the KTAIT module AlgorithmicEmergence.lean. Theorem 1 is proved outright (counting_bound, few_low_complexity_strings, most_states_incompressible). Theorem 2 corresponds to no_universal_compressor_micro; the generic string reductions are available as no_universal_compressor and no_compression_improver_of_raw. The Lean statement of the discovery reduction is formulated abstractly over any computable faithful embedding of finite strings into experiments, that is, any with ; the trivial identity/full-readout construction of Eq. (7) used in the main text, and the shift-register family of Theorem 1, are both instances of that hypothesis. The machine-checked result therefore does not depend on the particular witness chosen for exposition, and the paper’s choice makes explicit that the computational result requires only arbitrary-record representability. Corollary A2 uses the same abstract embedding: its residual-information part is counting_bound applied with the shared embedded condition, while its discovery and optimality parts are respectively no_universal_compressor_micro and no_additive_approximation_compressible. Theorem 3, in the compressible-witness form stated here, corresponds to no_additive_approximation_compressible, whose classical input is the padded-witness strengthening of non-approximability; the weaker unconditional statement is no_additive_approximation, and identification_barrier records the exact-optimality () consequence. Theorem 3 and Corollary 1 are derived here from Proposition 1; the Lean statements and their hypotheses are unaffected. Proposition 1 and Theorem A2 are proved at paper level only; no machine-checked declarations are claimed for them.
Appendix F.2. Relational, Auxiliary, and Fixed-Question Results
Corollary A1, the witness-based form of the optimality barrier, is machine-checked as relational_optimality_barrier. Its analytic inputs are named hypotheses in the same style as elsewhere: LengthUnbounded for the counting argument that a valid constructor’s returned length grows along the witness family, CompetitorBound and BoundOutrun for the competitor being describable from the search index, and WitnessCompresses for the padded compression witness. No shortest description appears in that statement or its proof. The current Lean development supplies K through its abstract AIT frame rather than constructing a global total K-function; theorem statements quantify over frames carrying that benchmark. Thus the foundational commitment is explicit in the formal interface rather than hidden in the proofs.
The threshold, conditional, and schedule results of Appendix D are machine-checked. compressor_of_threshold and threshold_undecidable_of_compressible cover the threshold specialization. conditional_complexity_uncomputable covers the varying-condition corollary. no_computable_schedule covers Proposition A2, with the step-indexed search and promise-halting exposed as named hypotheses. Proposition A3 is machine-checked as fixed_task_computable_of_factorization. The varying-projection result is query_slice_computable_of_query_complete. Its nonexistence consequence is no_query_complete_of_noncomputable_slice. The extensional converse discussed after the proposition is fixed_task_computable_iff_factorization; its identity construction uses the microscopic input as its own model code. These declarations are proved from Mathlib’s computability interface and consume no AIT hypothesis.
Appendix F.3. Reusable-Model Acquisition
The revised Proposition 2 and its periodic-block construction in Appendix D.6 are proved at paper level. They are not claimed to be covered by the existing declarations no_universal_emergence_constructor and no_bounded_regret_emergence_constructor. Those declarations express the earlier abstract transfer through program-to-model and model-to-program translations under named hypotheses; their validity is independent of this change in the manuscript.
The revised correspondence requires the joint-code validity and injectivity argument, finite search for a missed block, a fixed program implementing that search with the block length supplied, and the match/escape code’s satisfaction of the acquisition and prospective margins. The queue records these obligations. Until they are formalized and the correspondence checked, the revised proposition carries a paper-level annotation. The record-level declarations described above and the functional-core benchmarks below retain their existing scope.
Appendix F.4. Functional-Core Benchmarks
The module FunctionalCore.lean uses the namespace KTAIT.FunctionalCore. It defines the conditional exact and approximate benchmarks as functionalCoreComplexity and approximateCoreComplexity. A shortest representative is expressed by IsFunctionalCore; its existence from a supplied finite realization is existsfunctionalcore. The bound in Eq. (A7) has a typed counterpart, functionalcoreleimplementation, under explicit hypotheses that an interpreter realizes the declared behavior within a fixed description overhead. These results concern minimum description length for fixed behavior; they assert neither computable extraction nor optimal predictive quality or runtime. The probability argument in Proposition A1 is proved at paper level.
Appendix F.5. Scope of the Formalization
The conservation/coarse-graining equations in Section 6 do not introduce a separate theorem suite. Reversible complete-state invariance and the binary localization balance are direct applications of the already machine-checked localization machinery in the companion development; Eq. (46) additionally uses only the standard non-growth of Kolmogorov complexity under a fixed computable projection. We therefore state these as consequences rather than new named theorems.
The finite-catalog and runtime Busy Beaver discussion in Appendix C.2 is likewise at paper level. Semidecidability of is not formalized: it would require an operational machine semantics and a dovetailing construction, which the present development lacks. No main theorem depends on it: Theorem 2 rules out a total procedure that halts on every input and shortens whenever a shortening exists, and semidecidability is the complementary positive fact about partial search. Its negative scheduling companion is formalized: no_computable_schedule treats the search abstractly, as a step-indexed function with soundness and promise-halting as hypotheses, so the auxiliary Proposition A2 is machine-checked without the dovetailer being built.
References
- Ruffini, G. Information, Complexity, Brains and Reality (Kolmogorov Manifesto). arXiv 2007, arXiv:0704.1147. [Google Scholar] [CrossRef]
- Wallace, C.S. Statistical and Inductive Inference by Minimum Message Length. In Information Science and Statistics; Springer: New York, 2005. [Google Scholar] [CrossRef]
- Rissanen, J. Modeling by shortest data description. Automatica 1978, 14, 465–471. [Google Scholar] [CrossRef]
- Grünwald, P.D. The Minimum Description Length Principle. In Adaptive Computation and Machine Learning; MIT Press: Cambridge, MA, 2007. [Google Scholar] [CrossRef]
- Li, M.; Vitányi, P. An introduction to Kolmogorov Complexity and its applications, 3 ed.; Springer, 2008. [Google Scholar] [CrossRef]
- Schmidhuber, J. Simple Algorithmic Principles of Discovery, Subjective Beauty, Selective Attention, Curiosity & Creativity. In Proceedings of the Discovery Science; Corruble, V., Takeda, M., Suzuki, E., Eds.; Berlin, Heidelberg, 2007; pp. 26–38. [Google Scholar] [CrossRef]
- Delétang, G.; Ruoss, A.; Duquenne, P.A.; Catt, E.; Genewein, T.; Mattern, C.; Grau-Moya, J.; Wenliang, L.K.; Aitchison, M.; Orseau, L.; et al. Language Modeling Is Compression. arXiv [cs]. 2024, arXiv:2309.10668. [Google Scholar] [CrossRef]
- Solomonoff, R.J. A Formal Theory of Inductive Inference. Part I. Inf. Control 1964, 7, 1–22. [Google Scholar] [CrossRef]
- Solomonoff, R.J. Complexity-Based Induction Systems: Comparisons and Convergence Theorems. IEEE Trans. Inf. Theory 1978, 24, 422–432. [Google Scholar] [CrossRef]
- Hutter, M. Universal Artificial Intelligence: Sequential Decisions Based on Algorithmic Probability. In Texts in Theoretical Computer Science. An EATCS Series; Springer: Berlin, Heidelberg, 2005. [Google Scholar] [CrossRef]
- Ruffini, G. Models, Networks and Algorithmic Complexity. arXiv 2016, arXiv:1612.05627. [Google Scholar] [CrossRef]
- Dingle, K.; Camargo, C.Q.; Louis, A.A. Input–Output Maps Are Strongly Biased Towards Simple Outputs. Nat. Commun. 2018, 9, 761. [Google Scholar] [CrossRef]
- Ruffini, G.; Castaldo, F. Pattern, Persist! Algorithmic Persistence, Telehomeostatic Closure, and the Graded Architecture of Agency. BCOM Working Paper WP0216, local revision v17.4. DOI identifies the version family, not this local revision. [CrossRef]
- Anderson, P.W. More is different. Science 1972, 177, 393–396. [Google Scholar] [CrossRef]
- Bédard, C.A.; Bergeron, G. An Algorithmic Approach to Emergence. Entropy 2022, 24, 985. [Google Scholar] [CrossRef]
- Gardner, M. Mathematical Games: The Fantastic Combinations of John Conway’s New Solitaire Game “Life”. Sci. Am. 1970, 223, 120–123. [Google Scholar] [CrossRef]
- Reynolds, C.W. Flocks, Herds and Schools: A Distributed Behavioral Model. In Proceedings of the 14th Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH ’87), 1987; pp. 25–34. [Google Scholar] [CrossRef]
- Ronald, E.M.A.; Sipper, M.; Capcarrère, M.S. Design, Observation, Surprise! A Test of Emergence. Artif. Life 1999, 5, 225–239. [Google Scholar] [CrossRef]
- Jaynes, E.T. Probability Theory: The Logic of Science; Cambridge University Press, 2003. [Google Scholar] [CrossRef]
- Ruffini, G. Navigating Complexity: How Resource-Limited Agents Derive Probability and Generate Emergence, 2024. PsyArXiv preprint (2024). Zenodo Concept DOI BCOM WP0017, Resolv. To Curr. Depos. v0.3.1. [CrossRef]
- Ruffini, G.; Lopez-Sola, E. AIT foundations of structured experience. J. Artif. Intell. Conscious. 2022, 9, 153–191. [Google Scholar] [CrossRef]
- Ruffini, G.; Castaldo, F.; Lopez-Sola, E.; Sanchez-Todo, R.; Vohryzek, J. The Algorithmic Agent Perspective and Computational Neuropsychiatry: From Etiology to Advanced Therapy in Major Depressive Disorder. Entropy Number. 2024, 26, 953 11. [Google Scholar] [CrossRef]
- Gandy, R. Church’s Thesis and Principles for Mechanisms. In The Kleene Symposium;Studies in Logic and the Foundations of Mathematics; Barwise, J., Keisler, H.J., Kunen, K., Eds.; North-Holland: Amsterdam, 1980; Vol. 101, pp. 123–148. [Google Scholar] [CrossRef]
- Hamkins, J.D.; Lewis, A. Infinite Time Turing Machines. J. Symb. Log. 2000, 65, 567–604. [Google Scholar] [CrossRef]
- Andreev, M. Busy Beavers and Kolmogorov Complexity. arXiv 2017, arXiv:1703.05170. [Google Scholar] [CrossRef]
- Vitányi, P.M.B. How Incomputable Is Kolmogorov Complexity? Entropy 2020, 22, 408. [Google Scholar] [CrossRef]
- Hartle, J.B.; Hawking, S.W. Wave function of the Universe. Phys. Rev. D. 1983, 28, 2960–2975. [Google Scholar] [CrossRef]
- Hartle, J.B.; Hawking, S.W.; Hertog, T. The Classical Universes of the No-Boundary Quantum State. Phys. Rev. D. 2008, 77, 123537. [Google Scholar] [CrossRef]
- Penrose, R. Singularities and Time-Asymmetry. In General Relativity: An Einstein Centenary Survey; Hawking, S.W., Israel, W., Eds.; Cambridge University Press: Cambridge, 1979; pp. 581–638. [Google Scholar]
- Gács, P.; Tromp, J.T.; Vitányi, P.M.B. Algorithmic statistics. IEEE Trans. Inf. Theory 2001, 47, 2443–2463. [Google Scholar] [CrossRef]
- Vereshchagin, N.K.; Vitányi, P.M.B. Kolmogorov’s structure functions and model selection. IEEE Trans. Inf. Theory 2004, 50, 3265–3290. [Google Scholar] [CrossRef]
- Crutchfield, J.P.; Young, K. Inferring statistical complexity. Phys. Rev. Lett. 1989, 63, 105–108. [Google Scholar] [CrossRef]
- Israeli, N.; Goldenfeld, N. Coarse-Graining of Cellular Automata, Emergence, and the Predictability of Complex Systems. Phys. Rev. E 2006, 73, 026203. [Google Scholar] [CrossRef]
- Yanofsky, N.S. Galois Theory of Algorithms. In Rohit Parikh on Logic, Language and Society; Başkent, C., Moss, L.S., Ramanujam, R., Eds.; Springer, 2017; Vol. 11, Outstanding Contributions to Logic, pp. 323–347. [Google Scholar] [CrossRef]
- Blass, A.; Dershowitz, N.; Gurevich, Y. When Are Two Algorithms the Same? Bull. Symb. Log. 2009, 15, 145–168. [Google Scholar] [CrossRef]
- Michie, D. “Memo” Functions and Machine Learning. Nature 1968, 218, 19–22. [Google Scholar] [CrossRef]
- Soler-Toscano, F.; Zenil, H.; Delahaye, J.P.; Gauvrit, N. Calculating Kolmogorov Complexity from the Output Frequency Distributions of Small Turing Machines. PLoS ONE 2014, 9, e96223. [Google Scholar] [CrossRef]
- Zenil, H.; Hernández-Orozco, S.; Kiani, N.A.; Soler-Toscano, F.; Rueda-Toicen, A.; Tegnér, J. A Decomposition Method for Global Evaluation of Shannon Entropy and Local Estimations of Algorithmic Complexity. Entropy 2018, 20, 605. [Google Scholar] [CrossRef]
- Zenil, H.; Kiani, N.A.; Zea, A.A.; Tegnér, J. Causal deconvolution by algorithmic generative models. Nat. Mach. Intell. 2019, 1, 58–66. [Google Scholar] [CrossRef]
- Wolfram, S. A new kind of science; Wolfram Media, 2002. [Google Scholar]
- Ruffini, G. What Flows When Information Is Conserved? Algorithmic Localization, Materialization, and Persistence. In BCOM Working Paper WP0218; Zenodo, 2026. [Google Scholar] [CrossRef]
- Israeli, N.; Goldenfeld, N. Computational irreducibility and the predictability of complex physical systems. Phys. Rev. Lett. 2004, 92, 074105. [Google Scholar] [CrossRef]
- Golse, F. The Boltzmann Equation and Its Hydrodynamic Limits. In Handbook of Differential Equations: Evolutionary Equations; Dafermos, C.M., Feireisl, E., Eds.; Elsevier, 2005; Vol. 2, pp. 159–301. [Google Scholar] [CrossRef]
- Mori, H. Transport, Collective Motion, and Brownian Motion. Prog. Theor. Phys. 1965, 33, 423–455. [Google Scholar] [CrossRef]
- Külske, C. The Ising model: highlights and perspectives. Math. Phys. Anal. Geom. 2025, 28, 20. [Google Scholar] [CrossRef]
- Kadanoff, L.P. Scaling laws for Ising models near Tc. Phys. Phys. Fiz. 1966, 2, 263–272. [Google Scholar] [CrossRef]
- Wilson, K.G. Renormalization Group and Critical Phenomena. I. Renormalization Group and the Kadanoff Scaling Picture. Phys. Rev. B 1971, 4, 3174–3183. [Google Scholar] [CrossRef]
- Montbrió, E.; Pazó, D.; Roxin, A. Macroscopic Description for Networks of Spiking Neurons. Phys. Rev. X 2015, 5, 021028. [Google Scholar] [CrossRef]
- Hodgkin, A.L.; Huxley, A.F. A Quantitative Description of Membrane Current and Its Application to Conduction and Excitation in Nerve. J. Physiol. 1952, 117, 500–544. [Google Scholar] [CrossRef]
- Gu, M.; Weedbrook, C.; Perales, A.; Nielsen, M.A. More Really is Different. Physica D. Nonlinear Phenom. [cond-mat]. 2009, arXiv:0809.0151238, 835–839. [Google Scholar] [CrossRef]
- Bausch, J.; Cubitt, T.S.; Watson, J.D. Uncomputability of Phase Diagrams. Nat. Commun. [quant-ph]. 2021, arXiv:1910.0163112, 452. [Google Scholar] [CrossRef]
- Moore, C. Generalized shifts: unpredictability and undecidability in dynamical systems. Nonlinearity 1991, 4, 199–230. [Google Scholar] [CrossRef]
- Watson, J.D.; Onorati, E.; Cubitt, T.S. Uncomputably Complex Renormalisation Group Flows. Nat. Commun. 2022, arXiv:quant13, 7618. [Google Scholar] [CrossRef]
- Bedau, M.A. Weak Emergence. Philos. Perspect. 1997, 11, 375–399. [Google Scholar] [CrossRef]
- Moore, C. Unpredictability and Undecidability in Dynamical Systems. Phys. Rev. Lett. 1990, 64, 2354–2357. [Google Scholar] [CrossRef]
- Cardona, R.; Miranda, E.; Peralta-Salas, D.; Presas, F. Constructing Turing Complete Euler Flows in Dimension 3. Proc. Natl. Acad. Sci. USA 2021, 118, e2026818118. [Google Scholar] [CrossRef]
- Dyhr, S.; González-Prieto, Á.; Miranda, E.; Peralta-Salas, D. Turing Complete Navier–Stokes Steady States via Cosymplectic Geometry. PNAS Nexus 2026, 5, pgag131. [Google Scholar] [CrossRef]
- Tao, T. Finite Time Blowup for an Averaged Three-Dimensional Navier–Stokes Equation. J. Am. Math. Soc. 2016, 29, 601–674. [Google Scholar] [CrossRef]
- Tao, T. Searching for Singularities in the Navier–Stokes Equations. Nat. Rev. Phys. 2019, 1, 418–419. [Google Scholar] [CrossRef]
- Cook, M. Universality in Elementary Cellular Automata. Complex Syst. 2004, 15, 1–40. [Google Scholar] [CrossRef]
- Cubitt, T.; Pérez-García, D.; Wolf, M.M. Undecidability of the Spectral Gap. Nature [quant-ph]. 2015, arXiv:1502.04135528, 207–211. [Google Scholar] [CrossRef]
- Bausch, J.; Cubitt, T.S.; Lucia, A.; Pérez-García, D. Undecidability of the Spectral Gap in One Dimension. Phys. Rev. X 2020, 10, 031038. [Google Scholar] [CrossRef]
- Castilla-Castellano, L.; Lucia, A. Undecidability of the spectral gap in rotationally symmetric Hamiltonians. Phys. Rev. Res. 2026, 8, 013033. [Google Scholar] [CrossRef]
- Perales-Eceiza, Á.; Cubitt, T.; Gu, M.; Pérez-García, D.; Wolf, M.M. Undecidability in physics: A review. Phys. Rep. 2025, 1138, 1–29. [Google Scholar] [CrossRef]
- Chalmers, D.J. Strong and weak emergence. In The re-emergence of emergence: the emergentist hypothesis from science to religion; Clayton, P., Davies, P., Eds.; Oxford University Press, 2006. [Google Scholar] [CrossRef]
- Gell-Mann, M.; Lloyd, S. Effective Complexity. In Nonextensive Entropy: Interdisciplinary Applications; Gell-Mann, M., Tsallis, C., Eds.; Oxford University Press: New York, 2004; pp. 387–398. [Google Scholar] [CrossRef]
- Hoel, E.P.; Albantakis, L.; Tononi, G. Quantifying causal emergence shows that macro can beat micro. Proc. Natl. Acad. Sci. 2013, 110, 19790–19795. [Google Scholar] [CrossRef]
- Abrahão, F.S.; Zenil, H. Emergence and algorithmic information dynamics of systems and observers. Philos. Trans. R. Soc. A Math. Phys. Eng. Sci. 2022, 380, 20200429. [Google Scholar] [CrossRef]
- Hamkins, J.D. Lectures on the Philosophy of Mathematics; The MIT Press, 2021. [Google Scholar]
- Martin-Löf, P. Truth of a proposition, evidence of a judgement, validity of a proof. Synthese 1987, 73, 407–420. [Google Scholar] [CrossRef]
- Forster, Y.; Kunze, F.; Lauermann, N. Synthetic Kolmogorov Complexity in Coq. In Proceedings of the 13th International Conference on Interactive Theorem Proving (ITP 2022), Dagstuhl, Germany, 2022; Leibniz International Proceedings in Informatics (LIPIcs); Vol. 237, pp. 12:1–12:19. [Google Scholar] [CrossRef]
- Lawvere, F.W. Diagonal Arguments and Cartesian Closed Categories. In Category Theory, Homology Theory and their Applications II;Lecture Notes in Mathematics; Springer: Berlin, 1969; Vol. 92, pp. 134–145. [Google Scholar] [CrossRef]
- Yanofsky, N.S. A Universal Approach to Self-Referential Paradoxes, Incompleteness and Fixed Points. Bull. Symb. Log. 2003, 9, 362–386. [Google Scholar] [CrossRef]
- Wolpert, D.H. Physical limits of inference. Physica D. Nonlinear Phenom. 2008, 237, 1257–1281. [Google Scholar] [CrossRef]
- Chaitin, G.J.; Arslanov, A.; Calude, C. Program-Size Complexity Computes the Halting Problem. Bull. EATCS 1995, 57, 198–200. [Google Scholar]
- Chaitin, G.J. Information-Theoretic Limitations of Formal Systems. J. ACM 1974, 21, 403–424. [Google Scholar] [CrossRef]
- Li, X. Some Applications of Lawvere’s Fixpoint Theorem. Front. Philos. China 2019, 14, 490–510. [Google Scholar] [CrossRef]
- Bauer, A. On fixed-point theorems in synthetic computability. Tbil. Math. J. 2017, 10, 167–181. [Google Scholar] [CrossRef]
- Koch-Janusz, M.; Ringel, Z. Mutual information, neural networks and the renormalization group. Nat. Phys. 2018, 14, 578–582. [Google Scholar] [CrossRef]
- De las Cuevas, G.; Cubitt, T.S. Simple universal models capture all classical spin physics. Science 2016, 351, 1180–1183. [Google Scholar] [CrossRef]
- Reinhart, T.; Engel, B.; De les Coves, G. The Structure of Emulations in Classical Spin Models: Modularity and Universality. In Journal of Mathematical Physics; To appear, 2026. [Google Scholar] [CrossRef]
- Kearns, M.; Singh, S. Near-Optimal Reinforcement Learning in Polynomial Time. Mach. Learn. 2002, 49, 209–232. [Google Scholar] [CrossRef]
- Lobel, S.; Parr, R. An Optimal Tightness Bound for the Simulation Lemma, 2024. Reinf. Learn. Conf. (RLC) 2024, arXiv:2406.16249. [Google Scholar] [CrossRef]
- Ruffini, G. A Lean 4 Formalization of Kolmogorov Theory: Axioms, Theorems, and Project Status. BCOM Working Paper WP0195, 2026; concept DOI, current version v0.6.0. [CrossRef]
- Grünwald, P.D.; Roos, T. Minimum Description Length Revisited. Int. J. Math. Ind. 2019, 11, 1930001. [Google Scholar] [CrossRef]
- Nies, A.; Shafer, P. Randomness Notions and Reverse Mathematics. J. Symb. Log. 2020, 85, 271–299, [1808.02746. [Google Scholar] [CrossRef]
- Catt, E.; Norrish, M. On the Formalisation of Kolmogorov Complexity. In Proceedings of the Proceedings of the 10th ACM SIGPLAN International Conference on Certified Programs and Proofs, 2021; Association for Computing Machinery; pp. 291–299. [Google Scholar] [CrossRef]
- Calude, C.S.; Dinneen, M.J.; Shu, C.K. Computing a Glimpse of Randomness. Exp. Math. 2002, 11, 361–370. [Google Scholar] [CrossRef]
- Rosser, J.B. Extensions of Some Theorems of Gödel and Church. J. Symb. Log. 1936, 1, 87–91. [Google Scholar] [CrossRef]
- Barreira, L. Poincaré Recurrence: Old and New. In XIVth International Congress on Mathematical Physics; World Scientific, 2006; pp. 415–422. [Google Scholar] [CrossRef]
Figure 1.
Construction questions and guarantees. Once the representation is fixed, the microscopic law and case-specific information generate a finite observed history. The left branch asks whether that history has a substantial lossless compression gap; the right branch asks what a total computable constructor that always returns valid descriptions can guarantee when compression exists. The blind-spot bound applies at every sufficiently large record length for each fixed method. Proposition 2 transfers the discovery and optimality barriers to constructors that must return a representation together with a reusable model; the witness family is constructed in Appendix D.6; its guarantee concerns the available construction observations.
Figure 1.
Construction questions and guarantees. Once the representation is fixed, the microscopic law and case-specific information generate a finite observed history. The left branch asks whether that history has a substantial lossless compression gap; the right branch asks what a total computable constructor that always returns valid descriptions can guarantee when compression exists. The blind-spot bound applies at every sufficiently large record length for each fixed method. Proposition 2 transfers the discovery and optimality barriers to constructors that must return a representation together with a reusable model; the witness family is constructed in Appendix D.6; its guarantee concerns the available construction observations.

Figure 2.
Observer-relative reconstruction. Panel A illustrates the observer setup: microscopic evolution generates successive readouts through a projection (here reads one outlined cell of a drifting pattern, so the record is one bit per step); relative to its prior frame , the observer models the retained record , omits residual detail , and uses a selected representation–model pair to predict or support action in pursuit of its objective . Panel B: the observer need not receive the projection as an intrinsic feature of the substrate: it may be inherited, supplied, or learned using the observer’s prior modeling resources. Conditional on the selected projection, builds a world-model and retains candidate submodels that continue to compress, predict, or support action. The record-level barrier theorems concern this fixed-representation subproblem; the general emergence problem selects the representation and reusable model as one pair.
Figure 2.
Observer-relative reconstruction. Panel A illustrates the observer setup: microscopic evolution generates successive readouts through a projection (here reads one outlined cell of a drifting pattern, so the record is one bit per step); relative to its prior frame , the observer models the retained record , omits residual detail , and uses a selected representation–model pair to predict or support action in pursuit of its objective . Panel B: the observer need not receive the projection as an intrinsic feature of the substrate: it may be inherited, supplied, or learned using the observer’s prior modeling resources. Conditional on the selected projection, builds a world-model and retains candidate submodels that continue to compress, predict, or support action. The record-level barrier theorems concern this fixed-representation subproblem; the general emergence problem selects the representation and reusable model as one pair.

Figure 3.
Schematic structure-function frontier. Ordinary refinement trades roughly one additional model bit for one fewer residual bit (dashed slope- reference). A crack is a finite interval in which the residual falls appreciably faster; equivalently, the total two-part cost improves. The horizontal axis is description budget, not discovery time; a crack marks a candidate descriptive gain and does not establish acquisition, relevance, or reuse. Discreteness and lower-order coding terms are suppressed.
Figure 3.
Schematic structure-function frontier. Ordinary refinement trades roughly one additional model bit for one fewer residual bit (dashed slope- reference). A crack is a finite interval in which the residual falls appreciably faster; equivalently, the total two-part cost improves. The horizontal axis is description budget, not discovery time; a crack marks a candidate descriptive gain and does not establish acquisition, relevance, or reuse. Discreteness and lower-order coding terms are suppressed.

Figure 4.
Acquisition and use of a reusable model-compressor, within the observer setup of Figure 2. The prior frame C, objective , relevance requirements, and baselines precede selection. The pair’s incremental code cost is charged once; an inherited or supplied has zero incremental cost. The same model jointly encodes the retained observations through changing coordinates, while detail not required by the objective is omitted. New observation blocks test further savings with the pair fixed (Definition 3).
Figure 4.
Acquisition and use of a reusable model-compressor, within the observer setup of Figure 2. The prior frame C, objective , relevance requirements, and baselines precede selection. The pair’s incremental code cost is charged once; an inherited or supplied has zero incremental cost. The same model jointly encodes the retained observations through changing coordinates, while detail not required by the objective is omitted. New observation blocks test further savings with the pair fixed (Definition 3).

Figure 5.
The finite micro-experiment used in the worked compression example. (A) Each microstate contains the entire binary grid and register . The update applies Life to the grid and a cyclic shift to the register. The observation map retains the grid at each of 32 times; both competing codes must reconstruct this same grid movie. (B) Local windows show the first five grid states for the chosen initial configuration. Shaded cells are alive and arrows indicate one Life update. The shape at is the shape at translated one cell down and right. The intermediate shapes specify the four phases used by the glider model. Position, phase, and direction reconstruct the past frames and predict later ones under the same translation rule. The full grid has periodic boundaries.
Figure 5.
The finite micro-experiment used in the worked compression example. (A) Each microstate contains the entire binary grid and register . The update applies Life to the grid and a cyclic shift to the register. The observation map retains the grid at each of 32 times; both competing codes must reconstruct this same grid movie. (B) Local windows show the first five grid states for the chosen initial configuration. Shaded cells are alive and arrows indicate one Life update. The shape at is the shape at translated one cell down and right. The intermediate shapes specify the four phases used by the glider model. Position, phase, and direction reconstruct the past frames and predict later ones under the same translation rule. The full grid has periodic boundaries.

Figure 6.
A closed coarse law and a failed elementary-rule search for the same microscopic run. (a) Rule 146 evolves a 360-cell periodic ring from the displayed initial row. (b,d) Each aligned triple is mapped to one bit every three microscopic steps: AND is 1 only for 111; OR is 1 for any nonzero triple. (c,e) The selected coarse rule runs from the corresponding first coarse row. Rule 128 reproduces the AND record exactly. Rule 182, which minimizes one-step error among the 256 elementary rules for the OR record, diverges from it. The observed transitions have zero closure conflicts for AND and 1 476 for OR, counted by the convention in Appendix B.2. Shaded cells denote 1; time increases downward. The same projection is used on both sides of each reconstruction comparison.
Figure 6.
A closed coarse law and a failed elementary-rule search for the same microscopic run. (a) Rule 146 evolves a 360-cell periodic ring from the displayed initial row. (b,d) Each aligned triple is mapped to one bit every three microscopic steps: AND is 1 only for 111; OR is 1 for any nonzero triple. (c,e) The selected coarse rule runs from the corresponding first coarse row. Rule 128 reproduces the AND record exactly. Rule 182, which minimizes one-step error among the 256 elementary rules for the OR record, diverges from it. The observed transitions have zero closure conflicts for AND and 1 476 for OR, counted by the convention in Appendix B.2. Shaded cells denote 1; time increases downward. The same projection is used on both sides of each reconstruction comparison.

Figure 7.
Construction and answer barriers act on different coordinates. The top route covers microscopic evaluation with an effective stopping rule; the gold marker denotes the possible weak-emergence toll that no shortcut exists. A stopping rule is available when the target’s state carries bounded encoded information, so that its reachable state graph is finite, or when the question supplies a horizon. The middle route starts from a finite observer-visible record supplied by either data-generating regime. The rust barriers mark failure of a substantial lossless compression gap to exist (1) and failure of universal discovery or near-optimality guarantees (2); the green node marks successful algorithmic emergence. The gray curved arrows show that finite observations from either regime enter the same construction problem. The lower arrow is labeled by Corollary A2, which transfers the barriers through suitable record-preserving embeddings. The bottom route covers open-ended computation, whose answer map may be noncomputable, producing the undecidability barrier. Unbounded encoded information, an infinite lattice, growing memory, or a continuum state, is what makes that possible. The dashed implication is Proposition A3, a logical implication rather than an evaluation route.
Figure 7.
Construction and answer barriers act on different coordinates. The top route covers microscopic evaluation with an effective stopping rule; the gold marker denotes the possible weak-emergence toll that no shortcut exists. A stopping rule is available when the target’s state carries bounded encoded information, so that its reachable state graph is finite, or when the question supplies a horizon. The middle route starts from a finite observer-visible record supplied by either data-generating regime. The rust barriers mark failure of a substantial lossless compression gap to exist (1) and failure of universal discovery or near-optimality guarantees (2); the green node marks successful algorithmic emergence. The gray curved arrows show that finite observations from either regime enter the same construction problem. The lower arrow is labeled by Corollary A2, which transfers the barriers through suitable record-preserving embeddings. The bottom route covers open-ended computation, whose answer map may be noncomputable, producing the undecidability barrier. Unbounded encoded information, an infinite lattice, growing memory, or a continuum state, is what makes that possible. The dashed implication is Proposition A3, a logical implication rather than an evaluation route.

Table 1.
Accounts of emergence and complexity, grouped by their defining questions. These questions concern logical derivation, computation, model structure, causal influence, or an observer’s knowledge. Several can apply to different aspects of the same system.
Table 1.
Accounts of emergence and complexity, grouped by their defining questions. These questions concern logical derivation, computation, model structure, causal influence, or an observer’s knowledge. Several can apply to different aspects of the same system.
| Account | Defining question |
|---|---|
| Derivation and computation | |
| Strong emergence (Chalmers) [65] | Do higher-level facts fail to follow, even in principle, from complete lower-level facts? |
| Weak emergence (Bedau) [54] | Can a macrostate be derived from the microdynamics and the system’s external conditions only by simulation? |
| Physical undecidability [50,64] | Can one algorithm answer a specified physical question for every input in a family? |
| Model structure and prediction | |
| Minimal partial models (Bédard–Bergeron) [15] | Does the record’s modified structure function identify several minimal partial models? |
| Statistical complexity [32] | How much information is stored in the process’s optimal predictive causal states? |
| Effective complexity [66] | How long is a description of the identified regularities? |
| Causal influence | |
| Causal emergence [67] | Does a macro level have greater effective information than the micro level under specified interventions? |
| Observer knowledge and acquisition | |
| Observer-dependent emergence (Abrahão–Zenil) [68] | What trajectory information exceeds the observer’s knowledge and fixed allowances for coding, error, and processing? |
| Algorithmic emergence (this paper) | Has the observer acquired a reusable compressive model and a representation relevant to its objective? |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.