Submitted:
25 February 2026
Posted:
27 February 2026
Read the latest preprint version here
Abstract
We study the canonical string-based Assembly Index (ASI), defined as the minimum number of binary concatenations needed to construct a target word under full reuse. NP-completeness of ASI-DEC over general finite alphabets and an equivalence between ASI plans and straight-line programs (SLPs) under the same size convention has been established. We emphasize that all transfers between decision variants are effected by explicit polynomial-time mappings and (where needed) an explicit reparameterization of the threshold by an absolute constant or a simple affine function. The remaining technical obstacle for the binary alphabet is that a naive encoding reduction may allow an optimizer to exploit “cross-boundary” substrings created by overlaps of codewords. We give a fully self-contained binary-alphabet proof: we construct an explicit self-synchronizing (comma-free) codebook of 17 fixed-length binary codewords and prove a boundary-normalization lemma showing that optimal plans can be assumed aligned to codeword boundaries. This yields a polynomial reduction from fixed-alphabet ASI-DEC to binary ASI-DEC, proving NP-completeness over {0, 1}. Using the recalled ASI–SLP equivalence (with a short proof for completeness), we obtain NP-completeness of binary SLP-DEC. We additionally provide an explicit, fully formal translation between our binary-rule counting convention and the standard SGP size measure (sum of right-hand side lengths), showing that the NP-completeness classification transfers to common one-string SGP/SLP decision variants over {0, 1}.
Keywords:
assembly theory
; assembly index
; grammar-based compression
; computational complexity
; information theory
; complexity measures
1. Introduction, Motivation, and Related Work
Assembly Theory proposes the assembly index as a measure of compositional complexity, defined as the minimum number of binary composition (concatenation) steps needed to build an object from primitives. In [1], we formalized the canonical decision problem ASI-DEC and proved NP-completeness for general (unbounded) alphabets by an explicit correspondence between assembly plans and straight-line programs (SLPs).
In applications, the underlying alphabet is often fixed and small (e.g., bits, DNA, or small symbol sets), and reductions that rely on growing alphabets are not directly applicable. For the classical Smallest Grammar Problem (SGP) / minimum-size SLP problem, early NP-hardness proofs use an alphabet whose size grows with the instance (e.g., [3,4]). Whether such hardness persists for fixed alphabets was long open; Casel et al. proved NP-completeness for every fixed alphabet of size at least 17 [2]. For the binary alphabet, several works explicitly note that adapting known constructions is nontrivial and that the status was unclear in that line of research [2,5].
In the grammar-based compression literature, the Smallest Grammar Problem (SGP) is typically stated for a grammar generating a single string, but the size can be defined in several (polynomially related) ways: counting nonterminals, counting productions, or counting the total number of symbols on all right-hand sides. An SLP is a particularly restricted grammar where every production has the binary-concatenation form (or a terminal), and it is the standard normal form used in algorithmics on grammar-compressed strings [6]. For single-string grammars, optimizing over general grammars versus SLPs is interchangeable up to polynomial transformations by binarization of right-hand sides. This paper works with the most direct assembly-theoretic size measure: The number of binary concatenations (i.e., the number of rules of the form ). We therefore state the transfer to SLP-DEC/SGP-DEC under this convention, and we also include an explicit lemma (Section 7) showing how to convert to common SGP/SLP size conventions by an affine or polynomial-time reparameterization of the threshold. Hence the NP-completeness classification is unaffected by the choice among standard size measures.
We give a self-contained NP-completeness proof for ASI-DEC over the binary alphabet , using a synchronizing block code and a constant-overhead dictionary gadget that prevents cross-boundary reuse. We then formalize a bridge theorem showing that, under the same binary-concatenation size measure, ASI-DEC is parsimoniously equivalent to the binary SLP-DEC decision problem (and thus to the one-string SGP-DEC variants under standard size conventions). Consequently, the paper provides an NP-completeness proof for the binary SLP decision problem in the corresponding size convention, and clarifies the relation to standard grammar-size measures.
Section 2 recalls the model and decision problems. Section 3, Section 4 and Section 5 develop the coding and normalization lemmas. Section 6 proves binary NP-hardness and NP-completeness for ASI-DEC. Section 7.1 derives NP-completeness for binary SLP (and notes the implication for one-string SGP variants) and discusses implications.
2. Model, Problems, and Background
We work in the unconstrained binary-join model: At any time one may concatenate any two previously constructed strings. Reuse is free and unlimited. Terminals are available at zero cost.
Definition 1.
For a word , is the minimum number of binary concatenations needed to construct w from the terminals Σ.
Definition 2.
ASI-DEC: Given , decide whether .
Definition 3.
A Straight-Line Program (one-string grammar) SLP for a single target word w is a context-free grammar that generates exactly w and whose nonterminal productions are binary concatenations . The size is the number of such binary productions.1 Let be the minimum size of an SLP generating w.
Lemma 1.
For every word w, .
Proof.
We use the same size convention as in the main paper [1]: The size of an SLP is the number of binary concatenation productions (terminal productions are not counted).
SLP ⇒ plan. Given an SLP with m binary productions, evaluate its nonterminals in a topological order. Each production corresponds to one assembly step that concatenates the already-built strings of Y and Z to obtain the string of X. Hence .
Plan ⇒ SLP. Given an assembly plan with m concatenation steps, create one nonterminal for each intermediate object and one production for each concatenation. Choose the nonterminal of the final object as the start symbol. This yields an SLP of size m, hence .
Therefore . □
Lemma 2.
Over any alphabet,ASI-DECbelongs to NP.
Proof.
A certificate is an explicit concatenation DAG (or the list of concatenation steps with pointers to previously built objects) of size at most k. A verifier reconstructs all intermediate strings and checks that the final one equals w. Each intermediate string has length at most , hence verification runs in polynomial time. □
Remark 1.
It is known that the Smallest Grammar Problem remains NP-complete for a fixed alphabet of size at least 17 ; see, e.g., Casel et al. for fixed alphabets. This implies NP-completeness ofASI-DECover some fixed alphabet via Lemma 1.
Implications
The corollary places binary SLP (and one-string SGP variants) on the same worst-case complexity footing as the classical unbounded-alphabet setting. Unless , there is no polynomial-time exact algorithm for producing a minimum-size SLP for a given binary word. In addition, approximation-related hardness statements for grammar-based compression are sensitive to the chosen objective and reduction notion, and thus should be interpreted with care for fixed alphabets. In our setting, the bridge identifies the objective value on binary inputs under the “number of binary concatenation rules” convention; consequently, any statement formulated solely in terms of this objective carries over verbatim between ASI and SLP on .
The purpose of this paper is to remove the remaining binary-alphabet technicality by giving a stand-alone encoding-and-normalization argument that prevents “bit-level shortcuts”.
3. Why Binary Requires Extra Work: Cross-Boundary Reuse
A naive replacement of each high-level symbol by a short binary string can change the optimal assembly cost, because a binary optimizer is free to reuse any repeated bit-pattern—even one that crosses the intended symbol boundaries. This is the core obstacle for reductions to .
Suppose (for illustration only) that two distinct symbols are encoded naively as and . Then the binary string for is
so an optimizer may first build 01 once, then reuse it to build 0101, and finally build 010101. This kind of reuse does not correspond to any legal reuse of whole symbols at the high level: It arises purely from misalignment (a reuse that “slides” across the boundary). The same phenomenon appears even more starkly if one considers longer concatenations where a cross-boundary fragment happens to repeat in several places, creating spurious low-level building blocks.
To preserve NP-hardness under alphabet reduction, we need an encoding h such that no cross-boundary substring can become a reusable building block unless it corresponds to an actual concatenation of whole symbols. In our setting this is achieved by two coupled ingredients: (i) a self-synchronizing (comma-free) codebook C whose codeword boundaries are detectable in the raw bitstream, and (ii) a boundary normalization lemma showing that every optimal (or near-optimal) plan/SLP can be transformed, without increasing cost, into one whose intermediate objects are concatenations of whole codewords.
Intuitively, our markers force every true boundary to contain the anchor pattern 11110000 (end-marker followed by start-marker). Any misaligned substring that crosses such a boundary is “tied” to that unique boundary and therefore cannot reappear elsewhere, so it cannot be useful for reuse; the normalization lemma then formalizes the elimination of all such one-off fragments.
4. A Concrete Self-Synchronizing Codebook
Fix . We give an explicit injective encoding
where is a payload chosen to satisfy a simple constraint.
Definition 4.
A fixed-length codebook satisfies(SYNC)if no codeword can be formed by a nontrivial suffix-prefix overlap of two codewords: For all and all ,
In particular, since all codewords have the same fixed length L, condition extbf(SYNC) rules out any ambiguous parsing caused by shifted overlaps; hence every word in has a unique decomposition into length-L codewords.
Proposition 1.
Fix . Let be the explicit codebook from Appendix A, where each
and the payloads are chosen so that: (i) all are distinct, (ii) no contains 0000 or 1111 as a substring, (iii) each starts with 1 and ends with 0. Then:
- C satisfies(SYNC)(Definition 4);
- the boundary marker occurs in any only across genuine codeword boundaries, and at every genuine boundary it occurs exactly once.
Proof.
Item (1) follows from Lemma A10, which proves (SYNC) for the explicit codebook.
For item (2), note that by construction no payload contains 1111 or 0000 as an internal substring, and starts with 1 and ends with 0. Hence within any single codeword the substring 11110000 cannot occur (it would require 1111 followed immediately by 0000 inside the codeword). In a concatenation of codewords, the only place where 1111 is immediately followed by 0000 is at the junction between two consecutive codewords, i.e., at a genuine codeword boundary. Therefore occurs exactly once per boundary and nowhere else. □
Remark 2.
Let and be two codewords in C (so ). Consider any crossing length-L substring s obtained by taking a nonempty suffix of u and a nonempty prefix of v (so s starts strictly inside u and ends strictly inside v). If s were itself a codeword, then it would have to start with the prefix marker 0000 and end with the suffix marker 1111. But because s crosses the boundary between u and v, any occurrence of the marker 0000 at the beginning of s forces the boundary pattern
to appear internally in s (namely at the junction between the ending 1111 of u and the beginning 0000 of v). By Proposition 1, no codeword contains B internally. Hence no such crossing substring can be a codeword, which is exactly the kind of overlap excluded by(SYNC).
5. Boundary Normalization
Extend c homomorphically to words: For let
Definition 5.
A string is aligned if . An intermediate object in an assembly plan (or nonterminal in an SLP) is aligned if its value is aligned; otherwise it is misaligned.
Lemma 3.
Let . Any assembly plan that produces x (equivalently, any SLP for x) can be transformed in polynomial time into an equivalent assembly plan in which no intermediate object starts in one codeword block of x and ends in the next. Equivalently, every intermediate object that spans multiple C-blocks is aligned (i.e., lies in ), without increasing the number of concatenation steps.
Proof.
We argue in the assembly-plan model and use the plan–SLP equivalence (Lemma 1) only as a convenient notation. Let P be an assembly plan for , and let G be the corresponding SLP obtained by naming each intermediate object by a nonterminal and each concatenation step by a rule . Thus every intermediate object in P corresponds to some nonterminal X with value .
Call a nonterminal (equivalently, intermediate object) misaligned if its value is not in . Fix the boundary marker from Proposition 1. For intuition, see Appendix B for a small worked example.
We distinguish two cases for a misaligned value .
(i) Boundary-crossing case. Suppose starts in one C-block of x and ends in the next (i.e., it crosses at least one genuine codeword boundary in the unique C-factorization of x). Then contains the boundary marker as a substring (Proposition 1), and moreover this occurrence of B is tied to a unique boundary position of x. Consequently, cannot occur as a substring of x at two different positions.
In a straight-line program generating a single output string, if a nonterminal X is referenced in at least two distinct places in the derivation tree (i.e., appears at least twice on right-hand sides), then its value must occur in the output as a substring at least twice (at the corresponding two intervals). Therefore, in the boundary-crossing case, X can be referenced at most once.
(ii) Non-boundary-crossing case. If does not cross any codeword boundary of x, then it is entirely contained within a single C-block of x. Such nonterminals are harmless for our reduction because they cannot “bridge” two adjacent codeword blocks, and in particular cannot bridge the distinguished boundary between the fixed prefix D and the suffix later in Lemma 6.
Thus, every misaligned nonterminal that crosses a boundary is used at most once, and can be eliminated without increasing the number of concatenation steps by the inlining procedure below.
Whenever such a boundary-crossing misaligned nonterminal X occurs as one child in a binary rule, i.e., in a production or , replace that single occurrence by the (binary) derivation tree rooted at X (“inline” the unique use of X), and delete X from the grammar. Because X is used at most once, this does not increase the number of concatenation rules and does not change the generated string. Repeating this process eliminates all boundary-crossing misaligned nonterminals. In particular, the resulting plan is aligned in the sense that every intermediate object that spans multiple C-blocks lies in , and no intermediate object starts in one block and ends in the next. The total number of concatenation steps does not increase. □
6. Reduction to the Binary Alphabet
Let D be a fixed binary dictionary prefix containing all codewords twice:
Let be the number of C-blocks in D (here ), and define the prefix arrangement constant
This is the minimum number of binary concatenations needed to arrange m already-available codeword blocks into D; it is a constant independent of the input instance.
Given an instance of ASI-DEC over , define
In the ASI model the only terminals are the bits , so producing a codeword of length from terminals requires a constant number of concatenations.
For example, one can build in 15 concatenations by iterative doubling, and then obtain any 16-bit word by reusing suitable power-of-two blocks; in particular, each fixed codeword can be built from terminals using at most concatenations (e.g., by concatenating its 16 terminal bits in a binary tree). Therefore there is an explicit absolute bound for constructing all 17 codewords (and in fact a much smaller bound is possible with reuse, since the same intermediate blocks can be shared across codewords).
Let be the minimum number of concatenations required to build all 17 distinct codewords from (with full reuse), and let . Both and are absolute constants independent of the instance . Our use of isolates the unavoidable arrangement cost of the fixed prefix D once the block words are available; any additional cost for constructing the blocks themselves is absorbed into the constant offset and does not affect polynomial-time reducibility or NP-completeness.
The mapping is computable in time .
Lemma 4.
Let be a concatenation of m codeword blocks. Any assembly plan (or SLP) that produces D given the m codeword blocks as available reusable subobjects, using only binary concatenation has at least concatenation steps. In particular, for our fixed D with blocks we have .
Proof.
Initially one has at least m separate components (the m blocks). Each binary concatenation can reduce the number of components by at most 1. To obtain a single component equal to D therefore requires at least concatenations. □
Lemma 5.
If over , then over .
Proof.
Take an assembly plan for w of cost at most k. First build D using an optimal plan of cost T. Now simulate the plan for w on the encoded symbols by replacing each terminal with the already-built block and performing the same sequence of concatenations. This yields using at most concatenations (plus a constant to build the fixed codewords from terminals). □
Lemma 6.
If over , then over .
Proof.
Let P be an optimal assembly plan for of cost at most . By Lemma 1 view P as an SLP for , and apply Lemma 3 to obtain a normalized plan of cost at most in which no intermediate object crosses a codeword boundary between C-blocks (and, in particular, no intermediate crosses the boundary between the prefix D and the suffix ).
In , alignment implies that no intermediate object can cross the boundary between the fixed prefix D and the suffix : Otherwise it would contain the boundary marker internally, contradicting Proposition 1 and the normalization invariant. Equivalently, in the unique C-factorization of the output , every node of the concatenation DAG yields a contiguous block interval that lies either entirely within the first blocks (the prefix region) or entirely within the remaining blocks (i.e., the blocks forming ) (the suffix region).
Even if all m prefix blocks were available as atoms, arranging them into the exact prefix string D requires at least binary concatenations (Lemma 4). Hence contains at least concatenation steps whose outputs lie in the prefix region.
Delete from a set of prefix-region concatenations that witnesses this unavoidable arrangement cost. The remaining plan has cost at most
and still produces the suffix (possibly reusing any already-constructed codeword blocks as atoms). This is valid because after normalization no intermediate object used to construct a suffix-region node depends on concatenations whose outputs lie solely in the prefix region; any such dependency would require an intermediate spanning the prefix–suffix boundary, which is excluded by the normalization invariant.
Finally, project back to by replacing each block by . Alignment ensures every concatenation remains valid and produces w. Hence . □
Theorem 1.
ASI-DECover the binary alphabet is NP-complete.
Proof.
Membership in NP is Lemma 2. NP-hardness follows from Lemmas 5 and 6, which show
a polynomial-time many-one reduction from fixed-alphabet NP-complete ASI-DEC (Remark 1) to binary ASI-DEC. □
7. Binary SLP Is NP-Complete via ASI
7.1. SLP vs. SGP and Size Conventions
Lemma 7.
Let G be a context-free grammar that generates exactly one string w and whose productions have right-hand sides of length at least 1. There is a polynomial-time transformation producing an equivalent SLP that generates exactly w such that the number of binary concatenation productions in is under any standard size measure that counts the total number of symbols on right-hand sides (equivalently: Total RHS length).
Proof.
Replace every production with by introducing fresh nonterminals and a binary parse chain: , , ..., . This preserves the generated string and increases the number of binary productions by at most . Summed over all productions, the blow-up is linear in the total RHS length. □
Remark 3.
If a venue defines grammar size differently (e.g., counting nonterminals, productions, or total RHS length), Lemma 7 implies a polynomial relationship between thresholds. Thus, NP-completeness under our binary-concatenation size implies NP-completeness for the standard SGP/SLP decision variants by polynomial-time reductions.
We recall (proved in [1]) the equivalence between ASI plans and SLPs under the same size convention; for completeness we include a short proof.
Theorem 2
(Bridge theorem). For every binary word , .
Proof.
We use the same size convention as in [1]: SLP size is the number of binary concatenation rules .
SLP ⇒ ASI. Evaluate the SLP in a topological order; each rule is one assembly concatenation producing the string of X. Thus .
ASI ⇒ SLP. From an assembly plan with m concatenations, create one nonterminal per intermediate object and one production per concatenation; the final object is the start symbol. This yields an SLP of size m, so . □
Corollary 1.
Under the size convention of this paper (SLP size = number of binary concatenation rules), the decision problemSLP-DECover is NP-complete. Moreover, by standard polynomial-time reductions between common one-string grammar size measures, the corresponding one-stringSGP-DECvariants over are NP-complete as well.
Proof.
NP-membership is standard (certificate: An SLP of size at most k). NP-hardness follows by the identity reduction from binary ASI-DEC using Theorem 2:
□
7.2. From Binary-Rule Size to Standard SGP Size (Sum of RHS Lengths)
In the grammar-based compression literature, the Smallest Grammar Problem (SGP) for a single target string is often measured by the total length of all right-hand sides (RHS) of productions, sometimes with minor variations (e.g., whether to include the start symbol, terminal productions, or separators). To avoid any ambiguity when comparing to the state of the art, we make the following explicit.
Definition 6.
Let G be a context-free grammar that generates exactly one string. Define
i.e., the sum of lengths of all right-hand sides (counting terminals and nonterminals as symbols). The corresponding decision problem is: Given , is there such a grammar G with and .
Lemma 8.
Let G be an SLP in binary form (productions and ). If m is the number of binary productions and r is the number of terminal productions , then
In particular, over the binary alphabet , we may assume , hence .
Proof.
Each binary production contributes 2 to the RHS length; each terminal production contributes 1. Over , terminal productions can be shared globally (at most one for 0 and one for 1) without affecting the derived string, so we may assume . □
Lemma 9.
Let G be a grammar generating exactly one string, with . There exists, in polynomial time, an equivalent SLP generating the same string such that
Proof.
Because G generates a single string, we can (i) delete unreachable and unproductive nonterminals, and (ii) break any RHS of length into a chain of binary concatenation rules by introducing fresh nonterminals, preserving the generated string. Each created binary rule corresponds to consuming at least one symbol from some original RHS; thus the total number of created binary rules is at most . Terminal occurrences can be factored through at most two terminal rules over . The construction is standard and runs in polynomial time. □
Corollary 2.
Over the binary alphabet , NP-completeness ofSLP-DECunder binary-rule counting implies NP-completeness of the one-string SGP decision problem under the standard size measure from Definition 6.
Proof.
By Lemma 8, an SLP of m binary rules yields a grammar of RHS-size at most . Hence . Conversely, by Lemma 9, any RHS-size-K one-string grammar yields an SLP with at most K binary rules, so . Thus the problems are polynomial-time interreducible with explicit threshold mappings, and NP-completeness is preserved. □
8. Additional Remarks
Our main result concerns the binary alphabet , where rich encodings are possible. At the opposite extreme, for a unary alphabet there is exactly one word of each length, namely . In this case, the structural content of an optimal assembly plan is purely arithmetical: Every binary concatenation corresponds to adding lengths. Consequently, the minimum number of concatenations needed to build from 0 coincides with the length of the shortest addition chain for N (see, e.g., Theorem 3.1 in [7] and OEIS A003313 for values).
It is important to distinguish two input models. In the standard ASI-DEC formulation studied in this paper, the input is the word itself. For the unary alphabet this means the input size is (an explicit representation). In contrast, the classical “addition-chain decision” problem typically takes N in binary, so the input size is (a succinct representation). Thus, even if computing is nontrivial from a number-theoretic perspective, algorithms exponential in may still be polynomial in N and therefore do not imply NP-hardness for the explicit-string unary case.
For completeness we note that one can define a succinct unary variant of ASI-DEC: Given with N in binary, decide whether . Under the plan–addition-chain correspondence, this succinct unary problem is essentially equivalent to the classical shortest addition-chain decision problem. We do not study this succinct variant here; our NP-completeness result is for the explicit-string model and becomes nontrivial precisely when .
9. Conclusions
We established the computational complexity of the Assembly Index decision problem in the binary setting. Our main result (Theorem 1) shows that ASI-DEC over the alphabet is NP-complete.
The NP-hardness proof is based on an explicit fixed-length codebook satisfying (SYNC) (Definition 4 and Proposition 1) and the induced homomorphic encoding into . The key technical step is the boundary-normalization lemma (Lemma 3), which allows us to assume that any assembly plan (equivalently, any SLP under our size convention) producing a word in contains no boundary-crossing misaligned intermediate objects, i.e., no intermediate object starts in one C-block and ends in the next. This rules out cross-boundary reuse that could otherwise invalidate the reduction.
Using the fixed binary dictionary prefix
and the associated constant prefix-arrangement cost , we obtain a polynomial-time many-one reduction
formalized in Lemmas 5 and 6. In particular, the threshold mapping is explicit and involves only the additive constant .
Finally, by the bridge theorem (Theorem 2) equating assembly plans and straight-line programs over binary inputs under the same binary-concatenation size convention, the NP-completeness result transfers to the corresponding binary decision variants for SLP/SGP via the standard threshold comparison and the explicit size mappings developed in Section 7.
This paper focuses on decision problems and NP-completeness statements. Although the bridge theorem suggests related consequences for exact optimization variants under the same convention, a fully explicit optimization-level treatment would require a separate discussion of objective conventions and reduction notions, and is left for future work.
A promising direction for future research concerns approximation questions in grammar-based compression over fixed alphabets. Our Boundary-Normalization Lemma shows that, for strings in produced by our self-synchronizing encoding, boundary-crossing misaligned intermediates can be eliminated without increasing the objective under the size convention used here. This suggests that the self-synchronizing structure of the underlying codebook may serve as a useful restriction when analyzing (or designing) approximation algorithms for SLP/SGP variants over the binary alphabet via such encodings. In particular, it would be interesting to understand to what extent approximation guarantees (and known lower bounds) for grammar-based compressors are preserved under these structured binary encodings. We do not pursue approximation results in this paper.
Funding
This research received no external funding.
Acknowledgments
I thank my partners Szymon Łukaszyk and Wawrzyniec Bieniawski for inspiring the topic of this publication and their clarifications, formal corrections and improvements.
Conflicts of Interest
The author Piotr Masierak was employed by the company Łukaszyk Patent Attorneys. The author declares that the research was conducted in the absence of any commercial or financial relationship that could be construed as a potential conflict of interest.
Appendix A. Codebook Table
Binary-alphabet reductions are sometimes criticized as “non-constructive” if they only assert the existence of a synchronizing code. Here we provide an explicit 17-word codebook (Table A1) so that every step of the reduction (including checking (SYNC)) can be verified mechanically.
We use fixed-length codewords of length with a unique prefix marker and a unique suffix marker :
The chosen middle blocks satisfy three simple constraints: (i) all are distinct, (ii) contains no substring 0000 and no substring 1111, (iii) starts with 1 and ends with 0 (so that the concatenations and do not create internal occurrences of 0000 or 1111 across the marker boundaries). These conditions ensure that P occurs only as the prefix of a codeword and S occurs only as the suffix of a codeword. The absence of internal occurrences of 0000 and 1111 in Table A1 can be verified mechanically (e.g., by a short script).
Table A1.
Explicit synchronizing codebook of length 16.
| Symbol | (length 8) | (length 16) |
|---|---|---|
| 10001000 | 0000100010001111 | |
| 10001010 | 0000100010101111 | |
| 10001100 | 0000100011001111 | |
| 10001110 | 0000100011101111 | |
| 10010010 | 0000100100101111 | |
| 10010100 | 0000100101001111 | |
| 10010110 | 0000100101101111 | |
| 10011000 | 0000100110001111 | |
| 10011010 | 0000100110101111 | |
| 10011100 | 0000100111001111 | |
| 10100010 | 0000101000101111 | |
| 10100100 | 0000101001001111 | |
| 10100110 | 0000101001101111 | |
| 10101000 | 0000101010001111 | |
| 10101010 | 0000101010101111 | |
| 10101100 | 0000101011001111 | |
| 10101110 | 0000101011101111 |
Lemma A10.
The codebook in Table A1 satisfies (SYNC).
Proof.
Let C be the codebook. Consider any nontrivial suffix–prefix crossing of two codewords that has length . If were a codeword, it would have to begin with the marker and end with the marker .
Because starts strictly inside u (nontrivial suffix) or ends strictly inside v (nontrivial prefix), at least one of the following holds:
- the first position of lies strictly inside u, hence the prefix 0000 would have to occur internally in u, or
- the last position of lies strictly inside v, hence the suffix 1111 would have to occur internally in v.
However, by the construction constraints above, 0000 appears in a codeword only as its first four bits, and 1111 appears only as its last four bits. Therefore neither case is possible, so . This is exactly (SYNC). □
Appendix B. A Small Worked Example Illustrating Normalization
Let . Then is a concatenation of three 16-bit blocks. Any substring that crosses a block boundary contains the unique boundary pattern 11110000. Hence a misaligned intermediate string (one that starts inside one block and ends inside the next) cannot occur in two different places of , so it cannot be beneficially reused. Lemma 3 formalizes this intuition by eliminating all such one-off misaligned nonterminals by inlining.
[custom]
References
- Masierak, Piotr. Computational Complexity of Determining the Assembly Index. IPI Letters 2026, 4(1), 9–12. Available online: https://ipipublishing.org/index.php/ipil/article/view/315. [CrossRef]
- Casel, Katrin; Fernau, Henning; Gaspers, Serge; Gras, Benjamin; Schmid, Markus L. On the Complexity of the Smallest Grammar Problem over Fixed Alphabets. Theory of Computing Systems 2021, 65, 344–409. [Google Scholar] [CrossRef]
- Moses Charikar, Eric Lehman, Ding Liu, Rina Panigrahy, Manoj Prabhakaran, April Rasala, Amit Sahai, and Abhi Shelat. Approximating the Smallest Grammar: Kolmogorov Complexity in Natural Models. In Proceedings of the 34th Annual ACM Symposium on Theory of Computing (STOC 2002), pages 792–801. ACM, 2002. [CrossRef]
- Storer, James A.; Szymanski, Thomas G. Data Compression via Textual Substitution. Journal of the ACM 1982, 29(4), 928–951, (Extended abstract: The Macro Model for Data Compression, STOC 1978). [Google Scholar] [CrossRef]
- Hucke, Danny; Lohrey, Markus; Reh, Carl Philipp. The Smallest Grammar Problem Revisited. In String Processing and Information Retrieval (SPIRE 2016); Springer, 2016; pp. 35–49. [Google Scholar] [CrossRef]
- Lohrey, Markus. Algorithmics on SLP-compressed strings: A survey. Groups Complexity Cryptology 2012, 4(2), 241–299. [Google Scholar] [CrossRef]
- Bieniawski, Wawrzyniec; Tomski, Andrzej; ukaszyk, Szymon; Masierak, Piotr; Tworz, Szymon. Assembly Theory – Formalizing Assembly Spaces, Discovering Patterns and Bounds. Preprints.org, 2025. Preprint 202409.1581, Version 12. Posted: 29 December 2025. Available online: https://www.preprints.org/manuscript/202409.1581. [CrossRef]
| 1 | If a venue counts size differently (e.g., adds a constant per terminal declaration), all results remain valid after a polynomial-time threshold translation; see Section 7.2. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.