Submitted:
02 September 2026
Posted:
02 September 2026
You are already at the latest version
Abstract
Mathematical notation is standardly treated as a recording convention. This paper argues that notation is a compression technology and that the elementary arithmetic hierarchy—counting, addition, multiplication, exponentiation—is a compression hierarchy: each level compresses iterated application of the level below, and each act of compression renders cheaply expressible a class of structure that was expressible below only at prohibitive description-length cost. Primes are, in this sense, invisible from inside addition; scaling laws are invisible from inside the additive representation of magnitude. Two principles organize the argument. First, successful compression is a measurement of structure: compression is possible only where regularity exists and fails, informatively, where it does not. Second, a compressed representation converts the exploited regularity into a manipulable object governed by the arithmetic of the level below—multiplication of numbers becomes addition of exponents—so that each level is a cognitive instrument rather than a shorthand. A two-part criterion distinguishes notational compression from algorithmic improvement: the replaced representation prices all objects of its domain uniformly, and the compressed parameter acquires laws of its own. Instances of the pattern in mechanics, quantum state representation, and neural networks are outlined and reserved for separate treatment.
Keywords:
notation
; compression
; abstraction
; description length
; salience
; matrix product states
; low-rank structure
; history of arithmetic
MSC: 00A30; 01A05; 68P30; 81P68
1. Introduction
In the Roman numeral system, written long division was a specialist’s problem. The qualification is deliberate: Roman practical computation ran on the abacus, which is a positional device, and the written numerals served largely as a recording format, so the claim concerns written algorithms and not Roman arithmetical capacity. With that restriction the episode stands: the algorithm existed, the answers existed, and the mathematical facts were exactly what they are today; but the written notation made the computation expensive enough to be expert labor, and positional written notation turned the same problem into a school exercise. Nothing mathematical changed. What changed was the cost of description and manipulation—and when that cost fell, a class of questions moved from the frontier of practice into the hands of children.
This paper takes that episode as typical rather than anecdotal. The claim is that mathematical notation is a compression technology; that the elementary hierarchy of arithmetic operations is a hierarchy of compressions, each level compressing iterated application of the level below; and that each level of compression makes visible a class of structure that was formally present, but cognitively invisible, one level down. A notation is not a neutral container for content. It is an instrument, and like any instrument it determines what can be seen.
The claim must be stated carefully, because there is an easy objection with a distinguished pedigree: anything expressible in a rich notation is expressible in a poor one. Unary notation can, in principle, state every truth of number theory. This is correct and does not touch the thesis. The thesis is not about expressibility; it is about description length, and about the empirical regularity that what a notation cannot say cheaply, its users tend not to see. A terminological convention is therefore adopted at the outset and holds throughout: invisible means expressible only at a description-length cost that excludes the expression from the working practice of finite agents—never inexpressible. Invisibility is a statement about cost, and therefore a matter of degree; wherever the text asserts that a structure is invisible at a level, the assertion is that its definitional and manipulative cost at that level is prohibitive, not that it is undefinable. Everything visible at level exists at level k at exponential cost, and exponential cost, for finite agents, is a form of invisibility in exactly this sense. The parallel objection in physics is Kretschmann’s: any theory can be written in generally covariant form. Also correct; also beside the point. Covariant notation is not more expressive than coordinate notation. It is cheaper where the physics is invariant, and the cheapness is the discovery.
Section 2 and Section 3 develop the argument using nothing beyond school arithmetic; Section 3 closes with the criterion that separates the compressions at issue from ordinary improvements of algorithm. Section 4 situates the thesis in the literature on compression and on notation. Section 5 outlines—as pointers, each reserved for separate full-resolution treatment—three candidate instances of the pattern outside arithmetic. Section 6 states the restrictions the thesis requires and the empirical commitment it incurs, and Section 7 concludes. Any mathematician can verify every step of Section 2 and Section 3. That is the design: the argument is elementary; only its consequences are not.
2. The Compression Hierarchy
2.1. Level 0: Counting
The primitive representation of a natural number is a tally: one mark per unit, n marks for the number n. Description length grows linearly with the thing described—which is to say, there is no compression at all. Tallying is not nothing: it makes magnitude and order visible, and it supports the fundamental act of comparison by pairing. But it makes visible almost nothing else. Two tallies of 851 and 851 marks are equal, but establishing this is itself a laborious computation. Structure interior to a number—that 851 is , that it is one less than a multiple of four—is present in the tally in the sense that the tally determines it, and absent in the sense that no feature of the representation displays it. This is the base condition against which every subsequent level should be measured: uncompressed representations conflate everything except size. The tally is treated here as primitive, but it embeds a presupposition of its own—that the represented domain has been resolved into separable, iterable units—and that presupposition, discreteness itself as the zeroth compression, is examined in separate work.
2.2. Positional Numerals: Compression Before Operations
The decimal numeral “851” is three characters. The tally is 851 characters. Positional notation achieves description length proportional to rather than n—an exponential compression—and it achieves this by exploiting a specific regularity: every number decomposes uniquely as . Note what this means structurally: the everyday numeral system is already parasitic on exponentiation, an operation three levels up the ladder from the objects it names. The observation is anachronistic by design: historical users of place-value systems exploited a regularity they could not yet name, and it is the modern analysis of the numeral system, not its historical use, that requires the level-3 operation. We write numbers efficiently because the numeral system writes them, whether or not its users say so, as polynomial combinations of powers. A civilization’s numerals encode how far up the hierarchy its notation has climbed; and the written algorithms available to it—the division problem of the opening paragraph—are downstream of that choice.
2.3. Level 1: Addition
Addition compresses iterated succession. The expression abbreviates “the successor of the successor of the successor of 5,” collapsing a process of counting into a single binary operation. What becomes visible at this level is decomposition: a number is no longer only a magnitude but a sum, in many ways, of other numbers. Differences become objects (), comparison becomes arithmetic, and the first laws appear—commutativity and associativity, which are statements about the operation, not about any particular number. This is the first appearance of a pattern that will repeat at every level: compressing a process into an operation creates a new object (the operation itself) whose properties can be studied. Linear structure—constant absolute increments, arithmetic progressions—is the characteristic pattern of the additive level: visible here, invisible below.
2.4. Level 2: Multiplication
Multiplication compresses iterated addition: abbreviates . The compression is again from linear to logarithmic in the appropriate sense—the description of “add a to itself b times” no longer grows with b. And the structure that becomes visible is the deepest in elementary arithmetic: divisibility, factorization, and the primes.
It is worth being exact about the sense in which this structure is invisible one level down, since the invisibility at issue is the one defined in Section 1 and not a stronger one. The predicate “prime” is definable in purely additive terms: compositeness is expressibility as a sum of equal parts each exceeding one, and the sieve of Eratosthenes is an algorithm of counting and addition. A mathematics equipped only with addition could therefore study the primes. What it lacks is any notational reason to single them out and any affordable means of working with them: in the additive presentation every number is a sum of ones, the numbers 12 and 13 differ only in magnitude, and every statement about divisibility carries the definitional overhead of spelling out, additively, what division is. Make multiplication a primitive, and the situation inverts: the primes become unavoidable, because they are precisely the objects the new operation cannot decompose, and the cost of every statement about them collapses. The entire edifice of multiplicative number theory—unique factorization, congruences, the distribution of primes—is structure that exists at level 1 and becomes affordable at level 2. The instrument preceded the sighting.
Multiplication also creates the ratio, and with it proportion, rate, and area: a rectangle is a multiplication made spatial. A mathematics of similar triangles, of speeds, of scaling one quantity against another, becomes writable—and therefore thinkable—at this level.
2.5. Level 3: Exponentiation
Exponentiation compresses iterated multiplication: abbreviates . The classical illustration is the wheat-and-chessboard problem: the total is a twenty-digit number whose tally representation would exceed the number of grains it counts, whose additive description is a computation, and whose exponential description is four characters. But the compression of magnitude is the lesser gift. The greater gift is that exponential notation makes growth rate into an object.
From inside addition, a doubling sequence is just numbers getting large quickly; the additive increments are as unruly as the sequence itself, so the additive instrument finds no invariant. From inside multiplication the constant ratio becomes visible. But only exponential notation names the invariant—the exponent—and thereby makes it calculable. The laws of exponents, , say something remarkable: arithmetic reappears one level up, acting on descriptions instead of quantities. Multiplication of numbers becomes addition of exponents. This single homomorphism is the engine of the ladder, and its practical form—the logarithm—was for three centuries the most important computational technology in science. A slide rule is the observation, made physical, that level-3 structure obeys level-1 arithmetic.
The logarithm also delivers the sharpest example of notation-dependent visibility in this paper: the scaling law. A power law is, in the additive representation—a table of raw values and their increments—indistinguishable from any other monotone growth; in that representation, scaling laws are invisible in the sense of Section 1. Plot against —that is, re-describe every quantity by its exponent—and the law becomes a straight line whose slope isk. Kepler’s third law is, in these coordinates ( against ), a line of slope . It was stated within a decade of Napier’s tables, and stated in the vocabulary of proportion theory as sesquialterate—one-and-a-half-fold—which is the verbal form of a log–log slope. Whether the tables were the instrument of the discovery is a causal question of history that this paper does not decide; the representational fact suffices for the thesis: the regularity that is invisible in the table of raw values is a straight line once every quantity is re-described by its exponent.
2.6. The Engine of the Ladder
The hierarchy can now be stated as a single mechanism. At each level, a repeated process of the level below is compressed into an operation; the operation’s parameter (the count of repetitions) becomes a new mathematical object; and the previous level’s arithmetic reappears, acting on these new objects. Addition counts successions; multiplication counts additions; exponentiation counts multiplications; and the exponents themselves add. Each level is therefore not a shorthand for the level below but a new instrument trained on it: it takes what was previously an activity and makes it a quantity, at which point the quantity can enter equations, acquire laws, and reveal invariants. Table 1 summarizes the hierarchy; the structure listed as made salient at each level is, by the convention of Section 1, invisible at every level below it.
3. The Principle: Notation Is Compression, and Compression Reveals Structure
Two elementary observations turn the examples of Section 2 into a principle.
First: compression is only possible where structure exists—so achieved compression is a measurement. A description can be shorter than the thing described only by exploiting a regularity of the thing described. “” is short because the number it names is generated by a short rule; a “typical” twenty-digit number, one with no generating regularity, admits no comparably short name. This is the elementary core of the algorithmic-information viewpoint [1,2], and it has a consequence that is easy to state and easy to underestimate: when a compression scheme succeeds on an object, that success is evidence about the object, not about the scheme. The negative control is what makes the measurement honest. Incompressible content—random sequences, structureless data—defeats every scheme, and defeats them at a quantifiable cost; so a representation that shrinks has detected something real. Compression is measurement, and the compression ratio is the reading on the dial.
Second: a compressed representation converts the exploited regularity into a manipulable object—and hides what it discards. The exponent in is not merely an abbreviation; it is a number, and numbers can be added, compared, plotted, and solved for. The regularity that the notation exploits (constant ratio) becomes, in the notation, a first-class citizen (the exponent), and the structures built from it—logarithms, doubling times, complexity classes, the very distinction between polynomial and exponential—become writable and hence thinkable. Symmetrically, whatever the notation compresses away recedes from view. The exponent displays the generating rule and hides the digits; the tally displays a magnitude and hides everything. Every notation is a choice about which structures to make cheap, and the cheap structures are the ones a community of finite agents will find.
A notation is a compression scheme. Its compression ratio on a domain measures the structure of that domain. Its choice of what to compress determines which structures its users can see. A formalism operating at the wrong compression level will not merely describe its subject inefficiently—it will miss structure that a higher-level formalism makes obvious.
The principle requires a demarcation, since almost any mature formalism is cheaper than some predecessor, and the thesis would otherwise fit everything and predict nothing. Two conditions jointly distinguish the notational compression at issue here from ordinary algorithmic improvement. First, the replaced representation is level-0 relative to the domain: its description cost is uniform across the domain, so that it conflates all objects except by size—the tally prices 12 and 13 identically, as it prices every pair of equal magnitudes. Second, the compressed parameter acquires laws of its own, entering the arithmetic of the level above: exponents add, and the addition of exponents is a theorem, not a bookkeeping convenience. A representation change that shortens computation without satisfying both conditions is an algorithm; one that satisfies both is a level, and it is levels, not algorithms, that the thesis concerns. Level-0 status is moreover relative to a target structure. A dense matrix, for example, is a compressed representation with respect to linearity—it makes the linear structure of a map salient—while being a level-0 representation with respect to rank, whose description cost it distributes uniformly. “Uncompressed” is always an assertion about a domain and a structure jointly, never about a representation in isolation.
4. Related Work
The general thesis that compression is constitutive of mathematical regularity and understanding is established. Algorithmic information theory identifies regularity with compressibility [1,2], and Wolff has developed a programme in which mathematics as a whole is analyzed as information compression via the matching and unification of patterns [3]. The cognitive role of notation likewise has a literature of its own: Babbage argued that the design of signs governs the reach and speed of mathematical reasoning [4]; Peirce treated notation as an object of philosophical analysis in its own right [5]; Iverson advanced notation as a tool of thought, with executability as its test [6]; and Schlimm has recently given a systematic treatment of mathematical notations as cognitive instruments [7]. The present paper takes the general thesis as given and contributes four narrower claims: that the elementary arithmetic operations form a hierarchy in which each level compresses iterated application of the level below; that achieved compression is a measurement of the compressed object, with incompressibility as the negative control; that the compressed parameter obeys the arithmetic of the level below it, which is the mechanism by which a notation becomes an instrument rather than a shorthand; and that a two-part criterion—uniform description cost below, lawful parameter above—separates the levels at issue from ordinary improvements of algorithm.
5. Instances Beyond Arithmetic: An Outline
Three candidate instances of the pattern outside arithmetic are stated here as pointers, each with the qualification it requires; full-resolution treatment of each, at the grain the present paper gives to the primes, is reserved for separate work. Each candidate satisfies the two conditions of Section 3: a replaced representation whose description cost is uniform across its domain, and a compressed parameter with laws of its own.
Mechanics. The 3+1 notation of classical mechanics—a state on space evolving in an external parameter t—selects one temporal foliation and omits the selection from the page. What it renders derived rather than primitive is the four-dimensional kinematics: the invariance of the interval, and the unity of energy and momentum in a single four-vector, became cheap, and then standard, with the covariant notation of Minkowski [8]. The claim is restricted to kinematics. The geometric theory of gravitation required, in addition, the equivalence principle, Riemannian geometry, and a decade of conceptual work that no notation supplied; the episode supports the thesis only for what the notation in fact made inexpensive.
Quantum state representation. The state of n qubits is standardly a vector of amplitudes: a representation whose cost is uniform across Hilbert space, satisfying the first condition of Section 3 exactly—physical and Haar-random states are priced identically. The matrix-product representation writes a state as a chain of tensors of bond dimension D at cost [9,10], and the bond dimension is a lawful parameter: it bounds the entanglement entropy across every cut. The historical sequence is the one the thesis predicts, and deserves emphasis: the compressed representation was constructed first, as a working computational notation [11,12], and the characterization of the subspace on which it succeeds—the area-law corner occupied by ground states of gapped, local, one-dimensional Hamiltonians—followed fifteen years later [13]. Notation first; theorem after.
Neural networks. The dense weight matrices of a trained network are a compressed representation with respect to linearity and a level-0 representation with respect to rank. Explicit measurement read the hidden regularity out—the intrinsic dimension of objective landscapes and of fine-tuning is orders of magnitude below the notational parameter count [14,15]—and low-rank parameterizations then exploited it [16]. The negative control is in the literature: on structureless content, networks trained to memorize random labels, compression fails against a capacity bound of a few bits per parameter [17,18]. One asymmetry distinguishes this instance from the physical ones and is flagged here rather than resolved: Hamiltonian locality is antecedent to any representation of it, while the low rank of trained networks is a property of objects produced by a training process—antecedent relative to the dense notation, but possibly conferred by the dynamics that generated the object.
6. Scope and Limitations
Three restrictions delimit the thesis. (i) Ontological neutrality. The thesis concerns representation, not existence. Compression detects antecedent regularity and confers none: the primes are determined by the additive structure of prior to any multiplicative notation, exactly as area-law entanglement is determined by Hamiltonian locality prior to any tensor-network representation. Antecedence is here relative to the uncompressed representation; the further question of antecedence relative to the process that produced the object is the asymmetry flagged in Section 5, and is left open. (ii) Non-monotonicity. No claim is made that representations at higher levels dominate those at lower levels. Every compression is lossy with respect to salience: what a representation renders inexpensive, it renders salient, and what it discards, it renders derived. MPS representations are inefficient for volume-law entangled states; covariant formulations render frame-dependent quantities such as simultaneity derived rather than primitive; exponential notation suppresses digit-level information. The selection of a representation level is therefore a selection among complementary restrictions of salience, to be made relative to the structure under investigation. (iii) Cognitive, not metaphysical, scope. The thesis constrains the practice of agents operating under description-length costs; it places no restriction on expressive power, which is invariant across the hierarchy.
Within these restrictions the thesis is empirically committed, in the form the convention of Section 1 fixes. Since invisible structure remains definable, a notation lowers the cost of recognition rather than gating its possibility; the commitment is accordingly statistical: recognition of a class of structure by a research community is expected to cluster after, and not before, the availability of a notation in which the class admits short description. Of the episodes touched in this paper, the tensor-network case exhibits the ordering most cleanly—working compressed notation in 1992, characterization of the subspace in 2007; the covariance case is restricted to kinematics; the Kepler case is illustrative rather than dispositive, since the causal role of the tables in the discovery is not established; and the neural-network case supplies both the measurement and the exploitation, but on objects whose regularity may be conferred rather than found. A systematic test of the ordering claim, with cases selected in advance of inspection, is left to future work.
7. Conclusions
The hierarchy of Section 2 admits a formal continuation: tetration compresses iterated exponentiation. Tetration is not unproductive in general—hyperoperations organize parts of the analysis of algorithms and mark the boundary of the elementary functions—but it has produced no class of domain structure comparable to the primes or to growth rates, and the criterion of Section 3 supplies the reason: the tetrated parameter has acquired no arithmetic of its own. The ladder pays where the compressed parameter has laws, not merely where iteration can be named. The productive question is accordingly not the continuation of the operational ladder but the identification of regularities in contemporary mathematics that stand where the primes stood prior to multiplication: determined by existing content, displayed by no existing notation.
Section 5 outlines three candidate instances of this situation in the recent history of physics and computation; in each, a discipline operated for an extended period within a representation whose description cost was uniform across its domain, until a compressed representation was constructed and a previously latent class of structure entered the discipline’s working vocabulary. If the analysis of Section 3 is correct, the pattern is not incidental: the construction of representations is a research method admitting deliberate application, rather than a byproduct of exposition. The prospective application of that method—the design of notations selected for their compression behavior on a target domain, in advance of knowing what they will reveal—together with the full-resolution treatment of the instances outlined above, is developed in subsequent work in this programme.
Funding
This research received no external funding.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
No new data were created or analyzed in this study.
Conflicts of Interest
The author declares no conflicts of interest.
References
- Kolmogorov, A.N. Three approaches to the quantitative definition of information. Probl. Inf. Transm. 1965, 1, 1–7. [Google Scholar]
- Chaitin, G.J. On the length of programs for computing finite binary sequences. J. ACM 1966, 13, 547–569. [Google Scholar] [CrossRef]
- Wolff, J.G. Mathematics as information compression via the matching and unification of patterns. Complexity 2019, 2019, 6427493. [Google Scholar] [CrossRef]
- Babbage, C. On the influence of signs in mathematical reasoning. Trans. Camb. Phil. Soc. 1827, 2, 325–377. [Google Scholar]
- Peirce, C.S. On the algebra of logic: A contribution to the philosophy of notation. Am. J. Math. 1885, 7, 180–196. [Google Scholar] [CrossRef]
- Iverson, K.E. Notation as a tool of thought. Commun. ACM 1980, 23, 444–465. [Google Scholar] [CrossRef]
- Schlimm, D. Mathematical Notations; Cambridge Elements: Cambridge, UK; Cambridge University Press, 2025. [Google Scholar]
- Minkowski, H. Raum und Zeit. Phys. Z. 1909, 10, 104–111. [Google Scholar]
- Schollwöck, U. The density-matrix renormalization group in the age of matrix product states. Ann. Phys. 2011, 326, 96–192. [Google Scholar] [CrossRef]
- Verstraete, F.; Murg, V.; Cirac, J.I. Matrix product states, projected entangled pair states, and variational renormalization group methods for quantum spin systems. Adv. Phys. 2008, 57, 143–224. [Google Scholar] [CrossRef]
- White, S.R. Density matrix formulation for quantum renormalization groups. Phys. Rev. Lett. 1992, 69, 2863–2866. [Google Scholar] [CrossRef] [PubMed]
- Östlund, S.; Rommer, S. Thermodynamic limit of density matrix renormalization. Phys. Rev. Lett. 1995, 75, 3537–3540. [Google Scholar] [CrossRef] [PubMed]
- Hastings, M.B. An area law for one-dimensional quantum systems. J. Stat. Mech. 2007, P08024. [Google Scholar] [CrossRef]
- Li, C.; Farkhoor, H.; Liu, R.; Yosinski, J. Measuring the intrinsic dimension of objective landscapes. In Proceedings of the International Conference on Learning Representations (ICLR), Vancouver, BC, Canada, 30 April–3 May 2018. [Google Scholar]
- Aghajanyan, A.; Zettlemoyer, L.; Gupta, S. Intrinsic dimensionality explains the effectiveness of language model fine-tuning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL–IJCNLP), Online, 1–6 August 2021. [Google Scholar]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. LoRA: Low-rank adaptation of large language models. In Proceedings of the International Conference on Learning Representations (ICLR), Online, 25–29 April 2022. [Google Scholar]
- Arpit, D.; Jastrzębski, S.; Ballas, N.; Krueger, D.; Bengio, E.; Kanwal, M.S.; Maharaj, T.; Fischer, A.; Courville, A.; Bengio, Y.; et al. A closer look at memorization in deep networks. In Proceedings of the 34th International Conference on Machine Learning (ICML), Sydney, Australia, 6–11 August 2017. [Google Scholar]
- Allen-Zhu, Z.; Li, Y. Physics of language models: Part 3.3, knowledge capacity scaling laws. arXiv 2024, arXiv:2404.05405. [Google Scholar]
Table 1.
The compression hierarchy of elementary arithmetic.
| Level | Operation | Compresses | Structure Made Salient |
|---|---|---|---|
| 0 | Counting (tally) | — | Magnitude, order |
| 1 | Addition | Iterated succession | Decomposition, difference, linear growth |
| 2 | Multiplication | Iterated addition | Primes, factorization, ratio, area |
| 3 | Exponentiation | Iterated multiplication | Growth rate as object; logarithm; scaling laws; the polynomial/exponential dichotomy |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.