Preprint
Article

This version is not peer-reviewed.

Exonuclease Ribozyme Encoded in a Circular Genome: On the Initiation of the RNA World

Submitted:

22 June 2026

Posted:

23 June 2026

You are already at the latest version

Abstract
The "RNA World" hypothesis posits a reasonable scenario for the origin of life. A question central to this scenario is: which RNA species may have emerged as the first ribozyme, thereby initiating the early life stage? To date, experimental and computational studies focus on ribozymes playing a constructive role in RNA synthesis—such as those facilitating template-directed copying or nucleotide synthesis. However, the answer remains unconvincing, primarily because their functional complexity likely demands lengthy sequences, making de novo emergence improbable. Here, we propose an alternative scenario: a destructive ribozyme, which is potentially quite short, may have emerged first. This ribozyme could cleave other RNA molecules to generate building blocks for its own replication. Notably, the dilemma for a destructive ribozyme is that it may degrade molecules of its own type, thus making it probably difficult to thrive as an RNA species. Using computational modeling, we demonstrate that an exonuclease ribozyme encoded within a circular RNA genome (thus evading self-attack) can spread within an RNA pool and may subsequently give rise to those constructive ribozymes. Our findings offer a new perspective on the initiation of the RNA world, thereby enriching both theoretical and empirical frameworks for understanding the origin of life.
Keywords: 
;  ;  ;  

1. Introduction

The origin of life remains one of the most profound unsolved questions in the natural sciences. In this field, the “RNA world” is a leading hypothesis: it proposes that RNA acted as both genetic material and functional molecules in a primordial stage of life [1,2,3]. The hypothesis is grounded in three logical points: (1) “capable of undergoing Darwinian evolution” represents the critical feature of a life form; (2) heredity and functionality are two indispensable prerequisites for Darwinian evolution to occur; (3) it is improbable that two distinct classes of polymers—one carrying heredity and the other carrying functionality (with an “information linkage” between them, like DNA and proteins in modern organisms)—could have emerged simultaneously at the start. Indeed, as we know, RNA serves as the genetic material in viroids and a portion of viruses. The discovery of ribozymes promoted the formal proposal of the “RNA world” hypothesis [1,4,5]. A subsequent finding—that the catalytic core of the ribosome is RNA [6,7]—proved to be the most compelling evidence supporting the notion that “proteins were invented by RNA”. Coupled with many other lines of evidence, the “RNA world” hypothesis is now widely accepted and commonly serves as a “model scenario” for both experimental and theoretical investigations in the field [8,9,10,11].
On account of the difficulties involved in prebiotic RNA synthesis, there were proposals suggesting the potential existence of pre-RNA worlds—involving RNA-like polymers, which were subsequently taken over by RNA [2,8]. Later, owing to advances in the area of prebiotic synthesis, especially those concerning nucleotide synthesis [12,13,14,15,16], it seems much more “feasible” that the RNA world emerged de novo. Then the question becomes clearer: how could the RNA world have started from a prebiotic pool [17]?
Notably, based on the logic that “the simpler, the more likely to have emerged de novo”, it has been hypothesized that there was a “naked” stage at the beginning of the RNA world [18,19,20,21,22]. In other words, it is initially RNA molecules themselves that acted as the units of Darwinian evolution (i.e., Darwinian entities), and RNA-based “protocells” emerged later—marking the ‘first major transition’ in the evolutionary history [18,23,24]. In this context, the question of how the RNA world started is in fact: which RNA species, as a Darwinian entity, may have emerged first in a prebiotic nucleotide pool—that is, representing the most primordial ribozyme?
An RNA species capable of catalyzing template-directed RNA synthesis—and thus replication (commonly referred to as an “RNA replicase”, here REP for short)—is a long-proposed candidate for the earliest ribozyme to emerge in the RNA world [2,8,17,25,26]. For three decades, sustained experimental efforts by RNA engineering, with in vitro evolution as a core approach, have been devoted to “constructing” an RNA polymerase ribozyme analogous to the modern RNA polymerase enzymes (i.e., using mononucleotides as substrates) [27,28,29,30,31,32,33,34,35,36]. As a key advance, a polymerase ribozyme was evolved in a water-ice medium to catalyze the copying of specific RNA templates of a length comparable to its own (~200 nt) [31]. However, its structural complexity precludes self-templating, and its large size gives rise to the potentially severe problem of “strand separation” during replication. Notably, some recent efforts have shifted to developing a polymerase that uses trinucleotides as substrates (namely, a “triplet polymerase ribozyme”—essentially a template-directed ligase) [37,38,39,40], which at least partially solves the two problems (by unraveling ribozyme secondary structure and trapping dissociated RNA strands, respectively). However, like other engineered polymerase ribozymes, the triplet polymerase ribozyme also derives from the class I ligase ribozyme (which is already large) [27], thus remaining oversized (existing as a 135 + 153 nt heterodimer) [37]. A critical unsolved problem is, in fact, that: the spontaneous emergence in the initial nucleotide pool of such a large ribozyme seems implausible. In light of this, a recent study performed a de novo selection from a random sequence pool and identified a 45 nt template-directed ligase ribozyme (QT45), which can synthesize both its complementary strand using a random triplet pool and a copy of itself using defined substrates, marking a major step towards a plausible REP in the early RNA world [41].
Simultaneously, theoretical work regarding the evolutionary dynamics of REP in the RNA world has also been conducted. For instance, computer simulations demonstrated that RNA-like “replicators” could spread in a “naked” phase via replicase-like function—by favoring its own replication through limited dispersal [42,43]. More concretely, our previous simulation study suggested that the first emerging RNA replicase may have been a simple, loosely binding template-directed ligase, which would readily dissociate from the template to act as a template for its own replication [44]. This simple template-directed ligase is hypothesized to utilize both nucleotides and various oligomers as substrates aligned on the template (the REP assumed in the present modeling study is also such a simple ligase, see below). Notably, the recently reported QT45 ribozyme closely resembles this hypothetical simple ligase—that ribozyme is engineered to use trinucleotide substrates, but longer oligomers may also be ligated on the template; additionally, dinucleotides and mononucleotides can be utilized as well, albeit with lower efficiency [41].
Another candidate for the first ribozyme to emerge in the RNA world is a ribozyme catalyzing the synthesis of nucleotides (nucleotide-synthetase ribozyme, NR). In fact, the fundamental mechanism of RNA replication is merely template-directed synthesis. Therefore, the first ribozyme to emerge in the RNA world need not be an REP, especially considering the feasibility of non-enzymatic template-directed synthesis [45]—indeed, evidence is accumulating showing that such non-enzymatic synthesis can be rather efficient [46,47,48,49]. By supplying the basic building blocks for RNA synthesis, an NR could also favor its own replication and thus spread within the prebiotic pool—the scenario has been supported by our computer modeling [50] (notably, there is also modeling work demonstrating the scenario of REP-NR co-spread [51,52]). Experimental studies have confirmed the chemical feasibility of RNA-catalyzed nucleotide synthesis [53,54,55]. In fact, an important rationale for proposing NR as the earliest ribozyme also relates to the sequence length: its substrates (i.e., nucleotide precursors) are far smaller than those of REP (i.e., the template-substrate complex), potentially enabling the ribozyme itself to be shorter in sequence [50].
However, even a simple REP or NR (e.g., ~40-50 nt) might still exceed the feasible length for spontaneous emergence from a prebiotic nucleotide pool, especially given that sequence space expands exponentially with increasing sequence length. In principle, the potential lower length limitation of REP and NR may stem from their “constructive” nature—such ribozymes require a minimum length to fold into a stable, functional structure to participate in their relevant synthetic reactions. For instance, an REP should carry several key functions: general template and substrate binding, ligation of adjacent substrates on the template, and the capabilities to undergo multiple turnovers of these two processes; an NR should be capable of binding different substrates (e.g., ribose and nucleobase or other molecular moieties) simultaneously, and catalyzing their condensation. Indeed, as depicted for that tiny candidate of REP: “The low tolerance for mutations and deletions of QT45 likely reflects a high density of functional residues in the QT ribozyme core required for sustaining the multiple structural and functional requirements of a polymerase ribozyme in a small RNA motif” [41].
Actually, given the feasibility of non-enzymatic template-directed synthesis, the first emerging ribozyme may be any one that could favor its own replication [45]. For example, a “destructive” ribozyme may cleave other RNA molecules to supply building blocks for its own synthesis, thus capable of spreading in principle of Darwinian evolution. In fact, the smallest known natural ribozymes are destructive: specifically, the small self-cleaving ribozymes, which catalyze the cleavage of their own RNA strands [56,57,58]. Among these, the well-known hammerhead ribozyme can be as short as ~50 nt. Excluding its substrate segment (the target of self-cleavage), the essential catalytic segment of this ribozyme is only ~25-30 nt in length [59]. Interestingly, a recent study identified an even smaller self-cleaving ribozyme, designated the “lantern” ribozyme, which can be as short as 31 nt, with a 15-nt essential catalytic segment [60].
Remarkably, it is now clear that small self-cleaving ribozymes are common genomic elements throughout the modern living world [56,57,58]. Furthermore, in vitro selection experiments targeting non-natural self-cleaving ribozymes have shown that initially random-sequenced RNA pools can become heavily enriched with diverse self-cleaving ribozymes [61,62]. Beyond small self-cleaving ribozymes, many other well-characterized ribozymes—including group I introns, group II introns, the RNA components of spliceosomes, and RNase P—also catalyze reactions involving phosphodiester bond cleavage [63,64]. In reality, with the sole exception of peptidyl transferase (in ribosome), all other known natural ribozymes mediate RNA cleavage or related reactions [63,64]. Given the extreme conservation of the translation machinery, it is understandable that the catalytic role of RNA in ribosomal peptide synthesis cannot be substituted by proteins. By contrast, the widespread persistence of the ribozymes associated with RNA chain cleavage in modern organisms likely reflects a simple explanation: ribozymes are inherently adept at catalyzing such reactions.
One may then ask: is it plausible that a simple, destructive ribozyme emerged first in the RNA world? For instance, we may assume it is shorter than 20 nt. Such a ribozyme would be expected to cleave other RNA molecules to provide building blocks for its own replication. However, a critical question arises: a destructive ribozyme could also degrade RNA molecules of its own kind. How, then, could such a ribozyme persist and thrive as an evolving RNA species?
Relevantly, a recent simulation study conducted in our group suggests that circular genomes, which are widespread in the modern living world, may have their “root” in the RNA world [65]. Indeed, it has long been recognized that a polymer can undergo intramolecular circularization once it attains sufficient length, enabling its two ends to approach one another [66]. For RNA, such intramolecular end-to-end ligation is known to occur readily [67]. More notably, RNA products formed under some prebiotically plausible conditions have been demonstrated to be enriched in “ring-like structures” [68]. In other words, circular RNAs were presumably abundant within the primordial RNA pool and could have participated in early Darwinian evolution throughout the RNA world era [65].
In light of this, we propose that the simple, destructive ribozyme mentioned above may have been an exonuclease ribozyme (ER), which was encoded on a circular RNA—representing the first gene in the genome. In this manner, the gene could evade cleavage by the ribozyme. As for the ribozyme itself, we hypothesize that by folding into its functional conformation, it may avoid, at least to some extent, attack from other molecules of its own kind.
Therefore, the purpose of the current study is to model a scenario: could an ER encoded in a circular genome emerge from a nucleotide pool? If the answer is positive, could constructive ribozymes such as REP and NR emerge subsequently, thus potentially evolving towards a more complex RNA world? Certainly, the relevant mechanisms involved in these evolutionary processes are expected to be explored.

2. Results

2.1. About the Model

We conducted computer simulations using a Monte Carlo model similar to those employed in our previous work concerning the naked stage of the RNA world [44,50,65,69,70], particularly the most recent one that accounts for circular RNA [65]. The system consists of a two-dimensional N × N square grid (with a toroidal topology to avoid edge effects). The adoption of a two-dimensional system is partly for simplicity and partly based on prebiotic milieus such as mineral surfaces (with dispersal limitation) for the naked stage, as supported by many studies [71,72,73,74]. Molecules, including nucleotide precursors, nucleotides, and RNA, are distributed within the grid rooms. In each time step (Monte Carlo step), relevant events may occur to the molecules with defined probabilities:
  • Nucleotide precursors may be converted into nucleotides (randomly as A, G, C, or U).
  • Nucleotides may assemble into linear RNAs through random ligation.
  • A linear RNA may circularize via end-to-end ligation.
  • Both linear RNAs and circular RNAs may undergo template-directed replication using nucleotides and oligomers as substrates.
  • A circular RNA may turn into a linear one via the breaking of a certain phosphodiester bond.
  • A linear RNA may also break into smaller fragments.
  • A nucleotide may decay into a nucleotide precursor.
  • A nucleotide residue at the end of a linear RNA may also decay into a nucleotide precursor;
  • A linear RNA molecule containing a characteristic sequence (domain) is hypothesized to exert a specialized function (i.e., it may act as a ribozyme): specifically, an ER cleaves other RNA molecules at their single-stranded termini (Figure 1a); an REP catalyzes the template-directed ligation of RNA; an NR mediates the synthesis of nucleotides from their precursors.
  • Molecules may also move to an adjacent grid room.
For descriptions of associated parameters, refer to Table 1; detailed explanations are provided in the Methods section.
At the beginning of a simulation, a certain number of nucleotide precursors are introduced into the system. As time progresses, nucleotides and RNA will emerge. Conversely, the degradation of RNA and nucleotides may regenerate nucleotide precursors. Overall, the total amount of materials in the system remains constant. RNA sequences, which can replicate via template-directed synthesis, compete for these limited materials. Potentially, an RNA “species”—with a specific sequence—may spread (i.e., become thriving) in the system by virtue of its function, which favors its own replication.
Notably, the model has an “information resolution” at the nucleotide level, thus inherently suitable for studying the early Darwinian evolution, which relies essentially on the sequence-function connection of RNA. In practice, here the characteristic sequence of a functional RNA species (ribozyme) is arbitrarily presumed due to the lack of empirical examples, but this does not matter—the significant point is merely: if a characteristic sequence corresponds to a special function, can the species with this sequence prevail in “natural selection”?

2.2. The Spread of Exonuclease Ribozyme

In our model system, the exonuclease ribozyme (ER) exerts its catalytic function by cleaving the termini of other RNA molecules, thereby generating free mononucleotides (Figure 1a). These monomers may be reused for the replication of ER itself, thus promoting the spread (thriving) of ER within the system. However, an inherent dilemma exists: an ER may also cleave other ER molecules or their complementary strands, thereby hindering its spread. As noted in the Introduction, we hypothesized that if the ribozyme is encoded within a circular genome, this dilemma would be resolved, thus enabling ER to propagate as a molecular species. This expectation was supported by our simulations. In the case depicted in Figure 2a, at 1×104 Monte Carlo steps, 10 circular RNA molecules with the characteristic sequence of ER (8 nt in length; see Table 1) were inoculated into the system, alongside an equal number of control sequences. As a result, ER increased rapidly in the system (blue circles for circular ER sequences and blue line for linear ER sequences), whereas the control species did not. Notably, linear ER sequences can function as ribozymes in the system.
Then, at approximately step 5.5×106, the ER level (both linear and circular sequences) rose significantly (Figure 2a). This is because a compact circular genome is inefficient at generating a linear ribozyme, a process that requires precise cleavage or precise “transcription” (Figure 1b). Therefore, when a “non-coding sequence” was inserted into the circular genome by chance, the expanded genome outcompeted the compact ER genome and became dominant in the system (Figure S1a). By “working in a more efficient way”, the ER genome subsequently spread to a higher level. The chance event of genome enlargement may occur via end-to-end ligation of a linear ER sequence and another linear fragment (associated with the parameter PRL in the model), followed by circularization of the resulting longer linear sequence (associated with PEL).
For the case shown in Figure 2a, the expanded genome is 11 nt in length, containing a 3-nt non-coding sequence (Figure S1a). In fact, our simulations revealed that for a ribozyme with an 8-nt characteristic sequence, evolution toward an 11-nt genome represents a most prevalent scenario, although other lengths (e.g., 9 nt and 10 nt) were also observed (Figure S1b-d). This phenomenon is highly analogous to our relevant findings regarding REP and NR, in which we first proposed the plausibility of circular genomes in the RNA world [65]. To establish a more robust platform (i.e., thriving at a higher level renders populations less susceptible to random fluctuations) for subsequent analysis, we chose to initially inoculate an 11-nt expanded circular genome (rather than the 8-nt compact genome), and directly achieved stable propagation of ER at a high level (Figure 2b).
Notably, in both the case shown in Figure 2a and that in Figure 2b, we observed the numerical asymmetry between ER and its complementary sequence: linear ER is significantly greater in number than its linear complement, while circular ER is somewhat more abundant than its circular complement. At first glance, this phenomenon is somewhat unexpected because complementary sequences act as mutual replication templates. For example, in our relevant investigation regarding REP and NR, no such asymmetry exists [65] (refer to Figure S2).
First, we reasoned that the disparity in linear chains arises because ER ribozymes are partially protected from cleavage by folding (PRP=0.5; see Methods), whereas their complements are not. However, abolishing this protection midway (PRP=0 at step 2×106) had little effect (Figure S3a). We then hypothesized the key driver is the fact that “an ER molecule cleaves other molecules” per se. Indeed, conferring ER function also upon ER’s complements minimized the difference (Figure S3b). When combining the two adjustments, the difference disappeared (Figure S3c), confirming both factors contribute to the asymmetry. Notably, when linear ER-complement disparity vanishes, circular ER-complement disparity also disappears (Figure S3c). We deduced that higher copy number of circular ER complement stems from the abundant linear ER: given their similar sequence information, the circular form shares oligonucleotide substrates with the linear form; therefore, fragments detached from the linear ER template may be recruited by the circular ER template, thereby promoting the synthesis of circular ER complements. Indeed, shutting down linear RNA template activity midway (FLT=0; see Methods) immediately abolished circular ER-complement disparity (Figure S3d). Just after the shutdown (step 2×106), circular RNAs (ER and its complement) increased—because only they can act as template. But without the contribution of linear templates, linear ER reduced significantly. That is, functional ribozymes became short of, and so the circular version subsequently descended to a lower level (Figure S3d). This indicates that the origin system (before disabling the linear templates) is actually characterized by a combined linear-circular genome, consistent with a relevant conclusion of our previous work on a circular REP genome [65]: a “purer” circular genome likely emerged later in the RNA world, with distinct genome-ribozyme division of labor.
Then we turned to a key mechanism we hypothesized: Is the circular genome—thus evading attack by the ribozyme ER—really important for the spread of ER as an RNA species? Based on the case depicted in Figure 2b, at step 2×106, all circular RNA molecules were “artificially broken” into linear molecules (wherein circular ER was converted into linear ER, and the same applied to their complementary sequences); meanwhile, the system was adjusted such that RNA circularization was no longer permitted (PEL=0) and linear RNA molecules acted as templates as efficiently as circular RNAs had previously (FLT=1). As a result, the levels of linear ER and its complementary sequences, after a temporary increase (due to the “artificial breaking” of their circular counterparts), decreased sharply (Figure 2c). In other words, ER cannot propagate when the circular genome form is “banned”, even if the linear form is assumed to have no shortcoming in respect of template capability.
Similar to our previous modeling studies [44,50,70], to avoid spurious RNA degradation effects, we inoculated multiple (10) circular ER molecules here to investigate the corresponding evolutionary dynamics (Figure 2a, b). The observed spread of ER already suggests that this RNA species could have emerged de novo from the postulated nucleotide pool. In reality, the simultaneous appearance of these specific RNA molecules is highly implausible. However, a single ER molecule may have appeared repeatedly over the extended timescale of life’s origin. As illustrated in Figure 2d, we simulated one such scenario: one circular ER molecule (11 nt in length, consistent with Figure 2b) and one control molecule were inoculated into the system every 1×105 steps, and ER eventually propagated system-wide. Importantly, to rule out a “first-mover advantage”—whereby early-appearing species can exploit a more abundant supply of raw materials (i.e., nucleotide precursors)—20 circular RNA molecules with random sequences (also 11 nt) were introduced into the system prior to the periodic inoculation of ER. As evident from the snapshots (Figure 3), at step 6×105, when the critical inoculation event triggering the subsequent spread of ER occurs, the concentration of nucleotide precursors is significantly lower than that in the initial pool (step 1×104). By step 6.7×105, both circular ER (circles) and linear ER (line segments) have already proliferated; at step 8×105, they begin to spread locally. Subsequently, at step 1×106, their expansion becomes extensive (note that the model system possesses a toroidal topology). Ultimately, at step 2×106, the entire system is occupied.

2.3. About the Influence of Parameters

Firstly, to confirm that the spread of ER is attributable to the function of the ribozyme, we examined the influence of PBER (the probability of phosphodiester bond breaking catalyzed by ER; see Table 1). Indeed, as ribozyme function was gradually reduced, the spread of ER was attenuated, and when the probability was finally set to 0, spread was completely suppressed (Figure 4-PBER). Another parameter associated with the efficiency of ER is TER (the number of times an ER acts within a single time step). With the stepwise reduction of this parameter, the spread of ER was also weakened (Figure 4-TER). Next, given our hypothesis that ER first emerged in the RNA world (i.e., prior to the appearance of ribozymes such as REP), the replication of ER in the simulation relies on non-enzymatic template-directed synthesis of RNA (see Introduction). As anticipated, gradual reduction of PTL (the probability of template-directed ligation) impaired the spread of ER and eventually abolished it completely (Figure 4-PTL). As noted earlier, the folding-dependent resistance of ER ribozymes to cleavage by other ER ribozymes appears to exert only negligible effects on the system (Figure S3a). Consistent with this observation, adjustments to PRP had little impact on the spread of ER (Figure 4-PRP). This finding is somewhat unexpected and highlights that, even though “ER molecules can attack one another”, ER may still spread as an RNA species—provided that “its gene is protected within a circular genome”.
The analysis of the remaining parameters is presented in Figure S4. Notably, the default parameter values we adopted are nearly optimal for the spread of ER (as illustrated in both Figure 4 and Figure S4): altering these values—either increasing or decreasing them—could hardly improve the level of ER (see Box S1 for explanations of a few exceptions). This indicates that our parameter exploration was highly effective (see Methods for details on “parameter exploration”). In addition, the spread of ER exhibits robustness against moderate parameter variations (the specific values used in Figure 4 are given in its legend, and those in Figure S4 are listed in Table S1). This finding reinforces the credibility of the ER-related scenario proposed in our modeling study.

2.4. RNA Chain-Length and Topology Distribution: Destructive vs Constructive

As mentioned above, we noticed a quantitative distinction between ER and its complementary sequence in the system (Figure 2), which was not observed in the spreading of REP or NR (Figure S2). Then we further analyzed the RNA chain-length and topology distribution of the ER-spreading system, aiming to compare it with the systems involving these two constructive ribozymes.
Again, there is an obvious difference in chain-length distribution between the ER-spreading system and the REP-spreading system (Figure 5a). The former contains significantly more mononucleotides than the latter, which can be attributed to the function of ER—cleaving RNA molecules at their terminals. Additionally, it can be observed that from 2mers to 10mers, the ER system generally exhibits a descending trend, whereas the REP system shows an ascending trend. This difference should have arisen from the respective destructive and constructive natures of the two ribozymes. In any case, it implies that in the former system, there may be more substrates (short oligomers) available for synthesis, whereas in the latter system, there may be more parasitic RNAs (long oligomers—tending to act as templates) that interfere with ribozyme spreading.
Then, analysis of the topological distribution of 8mers to 11mers (representing the length range of a ribozyme in the model—the characteristic domain is 8-nt long and an active ribozyme should be shorter than 1.5 times of the domain; see Methods) revealed that RNA sequences lacking the ribozyme-characteristic domain (i.e., parasites; white area in the pie charts) are significantly more abundant in the REP system—with the linear fraction (empty white area) being particularly prevalent (Figure 5b). In fact, for a single-gene circular genome, even if a non-coding sequence exists, many “false breaking or synthesizing” events are still likely to occur (e.g., breaks occurring within the ribozyme-characteristic domain; refer to Figure 1b), which generate these linear parasites. Interestingly, in the ER system, the linear parasites can be largely “destructed” by the ribozyme. As a related observation, circular parasites (white area with circles) are significantly more abundant in the ER system than in the REP system—because only by adopting a circular form can parasitic RNAs evade cleavage by ER.
As another case of a constructive ribozyme, NR exhibits a spreading pattern similar to that of REP, both in terms of RNA chain-length and topology distribution (Figure S5). In addition, to provide another instance of a destructive ribozyme, we modeled a modified version of ER, which cleaves the ends of RNA molecules randomly by 1, 2 or 3 nt in length (experimental evidence for such random cleavage by ribozymes exists [75,76,77]). This ER variant, named “ER13” here, can also spread in the system (Figure S6a), and this spreading is indeed attributable to its catalytic function (Figure S6b). The spreading pattern of ER13 is highly similar to that of the original ER, except that the 2mers and 3mers are more abundant and 1mers are less so in the system (Figure S5)—a result consistent with our expectations.

2.5. The Subsequent Emergence of Constructive Ribozymes

If ER emerged as the first ribozyme in the RNA world, other ribozymes—for example, constructive ribozymes such as REP and NR, as previously proposed in the field—should have arisen subsequently. We therefore investigated whether this scenario could be recapitulated in our model system. Notably, for a multi-gene circular genome, a non-coding sequence is no longer required for efficient ribozyme production [65] (see Figure 1c for an illustration of the situation concerning a two-gene circular genome).
First, we inoculated a two-gene circular genome containing ER and REP and found that it can spread in the system (Figure 6a). When the function of REP was turned off midway, the two-gene genome faded away, leaving only the one-gene ER genome to spread (Figure 6b). This indicates that REP function indeed contributes to the propagation of the ER-REP genome. Certainly, subsequent inactivation of ER function collapsed the spread of the one-gene ER genome. Then, we inoculated a circular genome containing ER and a non-coding sequence to see whether an ER-REP genome could spread naturally via the mutation of the non-coding sequence into REP (to facilitate the mutation, the non-coding sequence was assumed to differ from the REP sequence by only one nucleotide). Interestingly, such natural emergence was observed in our simulation (Figure 6c; see Figure 7 for snapshots of the simulation case). Finally, it was revealed that the spread of a three-gene genome containing ER, REP, and NR is also plausible (Figure 6d), implying the potential for further evolution.

3. Discussion

A NASA-formulated “working definition” of life for astrobiological exploration states: “Life is a self-sustaining chemical system capable of undergoing Darwinian evolution” [78,79,80]. The definition is widely accepted because it captures the two essential aspects of life: self-sustainment and the capacity to undergo Darwinian evolution [81]. However, these two aspects are so distinct that combining them in a single sentence seems inappropriate. For instance, we can refer to a self-sustaining chemical system, but how can an individual system undergo Darwinian evolution? Darwinian evolution is meaningful only for a lineage (at least at the population level or above). Therefore, the concept of life should be split—for example, by defining “living entity” and “life form” separately, with the former related to self-sustainment and the latter to Darwinian evolution [81]. For those who prefer a concise definition, it should at least be modified to: “Life is a sort of self-sustaining chemical system capable of undergoing Darwinian evolution” [11].
In the field of the origin of life, there has long been a debate surrounding the question of “metabolism first or replication first?” [82,83,84,85,86]. From the perspective of the conceptual analysis of life mentioned above, this debate is essentially: “self-sustainment first or Darwinian evolution first?”. In fact, the “self” of a modern living system can be regarded as being “defined” by its genetic information. Even if some self-sustaining chemical systems might have existed prior to the origin of Darwinian evolution, the meaning of “self” should have shifted after Darwinian evolution emerged, becoming intrinsically linked to the newly arisen genetic information [22]. In other words, the core logic underlying the origin of life lies in the origin of Darwinian evolution, followed by the subsequent emergence of self-sustainment as we know it (for the purpose of adaptation). Therefore, in the scenario of the RNA world, the emergence of the first ribozyme (or rather, the first gene) is a key event, which initiated Darwinian evolution—and thus kick-started the origin of the entire living world. The present paper focuses on this critical issue, which has puzzled researchers for decades since the proposal of the RNA world hypothesis.
We reasoned that a destructive ribozyme, i.e., the exonuclease ribozyme (ER), which might be much smaller than those previously proposed constructive ribozymes, such as the replicase ribozyme (REP) and nucleotide-synthetase ribozyme (NR), may have emerged first. The destructive ribozyme may cleave other RNA molecules to supply building blocks for its own replication. To avoid self-attack, the ER may have been “encoded” in a circular genome. Through computer modeling, here we confirmed the plausibility of relevant evolutionary dynamics (Figure 2a, b and d; Figure 3). Wherein, the spread of ER is indeed attributed to the destructive catalytic function of this ribozyme (Figure 4-PBER and TER), and its encoding in a circular genome was shown to be crucial for this spread (Figure 2c).
Somewhat surprisingly, it was found that though the ribozyme itself, as a linear molecule, may be attacked by other ER molecules, the spread is barely affected. For example, the spread of ER is robust against the downward adjustment (even to 0) of the assumed probability that a ribozyme is protected (Figure 4-PRP and Figure S3a; note that this probability may have a little influence on the quantity of linear ER—comparing Figure S3b with Figure S3c). The reason should be that the degradation of linear ER would also generate building blocks for further synthesis of its own species.
Furthermore, it has been demonstrated that those constructive ribozymes can emerge following the spread of ER (Figure 6). Specifically, an ER-REP two-gene circular genome is capable of propagation (Figure 6a), which can be attributed to the functions of the two ribozymes (Figure 6b). The REP gene was then shown to arise naturally via mutation from a non-coding sequence within an ER genome, thereby driving the spread of the ER-REP genome (Figure 6c and Figure 7). Notably, the reaction mediated by REP (ligation) is the reverse of that mediated by ER (cleavage), making them directly related in catalytic chemistry. For instance, many small self-cleaving ribozymes catalyze RNA ligation to varying degrees [56,57,58]. In fact, the earliest engineering efforts to construct an REP were based on the group I self-splicing intron [27], which mediates cleavage and ligation reactions simultaneously. Additionally, the recent identification of that short REP candidate (45 nt in length) [41] implies that the emergence of an REP in an RNA pool dominated by an ER (e.g., 20 nt in length) may be feasible.
Remarkably, an ER-REP-NR three-gene circular genome was also found capable of spreading in the system (Figure 6d). Indeed, some previous studies by our group and others suggest that the cooperation of more than two types of ribozymes—with unlinked genes—would be difficult in the naked stage of the RNA world [51,52,87]. Interestingly, the observation here shows that at least, three types of ribozymes may co-spread in a purely spatial system (that is, without membrane boundaries of protocells), provided their genes are linked together.
As mentioned in the Introduction, ribozymes should be adept at cleaving RNA chains, and thus such functional ribozymes are still widely distributed in the protein-dominated modern living world [63,64]. However, to date, no exonuclease ribozyme (ER) has been identified. In fact, in the context here, a “rigorous” exonuclease is not required. For example, we demonstrated that a ribozyme capable of randomly cleaving an RNA substrate by 1, 2, or 3 nt at its chain end (i.e., ER13) can also propagate within the system (Figure S6). Notably, experimental work has shown that RNase P—a well-characterized nuclease with RNA as its catalytic core—may catalyze such random cleavage when acting on short RNA substrates [75,76,77]. Indeed, the catalytic RNA of RNase P is itself excessively large (300-400 nt), but interestingly, it has been revealed that small RNAs corresponding to a domain of this ribozyme (about 30 nt in length) may also perform its cleavage function, albeit with greatly reduced efficiency [77,88]. It seems not impossible to identify a small ER (e.g., no longer than 20 nt, not necessarily with very high efficiency) in the laboratory, potentially via in vitro evolution. The successful identification of such a ribozyme will provide us a glimpse of the earliest scenario of Darwinian evolution proposed here.
In the field of life’s origins, integrating experimental and theoretical approaches is significant, as neither is sufficient on its own [11]. A recent study from our group provided a paradigmatic example in which experimental clues guided corresponding theoretical modeling [89]. Here we present a reciprocal case in which theoretical modeling can inform targeted experimental investigation. Collectively, our findings regarding an ER encoded in a circular genome offer a new perspective on the critical early stage of the RNA world’s emergence, and may expand both theoretical and empirical frameworks for understanding the origin of life.

4. Methods

4.1. The Events Occurring in the Model System

Figure 8 shows a schematic of the events in each time step (i.e., Monte Carlo step), together with their associated probabilities and factors (see Table 1 for definitions). In the N × N grid system, molecules only interact with others within the same grid room. A molecule may move to an adjacent room (related probability: PMN for nucleotides and RNA, and PMNP for nucleotide precursors, respectively).
A nucleotide precursor may form a nucleotide (randomly A, G, C, or U) in a non-enzymatic way (PNF) or catalyzed by NR (PNFR). A nucleotide may also decay into its precursor (PND). Nucleotides and linear RNAs may undergo random intermolecular ligation (PRL) to form longer chains. A linear RNA may undergo intramolecular end-to-end ligation (PEL) and thus transform into a circular one. An RNA molecule may attract substrates (nucleotides or oligomers) (PAT) via base-pairing with a certain error rate (PFP). Attraction of a substrate adjacent to one already positioned on the template is easier than de novo attraction (FDA) — the so-called “primer effect”. In addition, linear RNA may act as a template less efficiently than circular RNA (FLT). Substrates aligned adjacently on the template may be ligated in a non-enzymatic way (PTL) or catalyzed by REP (PTLR), corresponding to template-directed synthesis. The substrates or the full complementary strand may separate from the template (PSP). Phosphodiester bonds within an RNA chain may break (PBB): a circular RNA may thus become linear and a linear RNA may split into fragments. A nucleotide residue at the end of an RNA chain may decay into a nucleotide precursor (PNDE). An RNA molecule may be cleaved at its chain end by ER (PBER); however, if the molecule acts as an active ribozyme, it is protected against cleavage to a certain degree due to its folding (PRP). ER, REP and NR can catalyze their respective reactions multiple times per Monte Carlo step (TER, TREP and TNR, respectively).
Notably, similar to our previous modeling work on scenarios in the RNA world, the energy issue is not considered explicitly here. For instance, nucleotides and oligonucleotides are implicitly assumed to be activated; in particular, upon their release from RNA degradation, they are assumed to be immediately reactivated so that they can be reused in subsequent RNA synthesis. In fact, such in situ activation has been shown to be feasible in laboratory experiments [90,91,92]. In reality, the energy source for primordial life may have involved chemical energy in prebiotic habitats, such as submarine hydrothermal vents [93,94,95] or terrestrial hydrothermal fields [96,97], as previously supposed. Since substrates are assumed to be always “activated”, the RNA species in the model system compete for materials but not energy. As noted earlier, the total amount of materials in the system is limited (TNPB). Certainly, in actual Darwinian evolution, competitions for materials and energy are both possible.

4.2. The Setting of Parameters

The probabilities and factors concerning the events in the system should be set in accordance with certain rules. Reactions catalyzed by ribozymes should be much more efficient than the corresponding non-enzymatic reactions, so PBER>>PBB, PTLR>>PTL and PNFR>>PNF. Template-directed ligation should be apparently more efficient than intermolecular random ligation, so PTL>>PRL. Intramolecular end-to-end ligation of a linear RNA should be easier than intermolecular random ligation, so PEL>PRL. Here, nucleotide residues within the chain are assumed to be unable to decay, as they should be protected, whereas those at the end of the chain (which are “semi-protected”) decay at a rate lower than that of free nucleotides; that is, PNDE<PND. Template-directed synthesis without the aid of a protease could not have very high fidelity, so PFP should not be set to a very small value. Other considerations may include: PBB may be higher than PRL, but lower than PNDE; PMNP should be higher than PMN; PNF should not exceed PND, and so on.
Considering the computational intensity, the total materials in the system (TNPB) are assumed to be obviously smaller in scale than the potential corresponding situation in reality; similarly, the characteristic RNA sequences for ER, REP or NR assumed here (8 nt) are likely to be much shorter than the corresponding ribozymes in reality (particularly those for REP and NR). However, such simplifications are not considered to conflict with the fundamental mechanisms that the modeling is intended to reflect.
Notably, exploring parameter values that “support hypothetical scenarios” is both sensible and valuable in theoretical modeling on the origin of life, especially given our highly limited knowledge of prebiotic environments and chemistry [98]. In fact, the evolutionary process during the origin of life is distinguished by a trend from simplicity to complexity—a highly rare phenomenon in nature. Therefore, any evolutionary scenario moving in this direction, if supported by modeling, deserves our attention: we may even infer relevant prebiotic conditions from the “favorable parameter values” underlying that scenario. This type of modeling can be referred to as “reverse modeling” [98].
In the present study, bearing the rules mentioned above in mind and with consideration of computer intensity, most parameter values were explored based on our experience from previous modeling studies on the naked RNA world [44,50,65,69,70], with particular reference to Ref. 65, which addresses circular RNAs. As mentioned in Results, the parameter exploration proved quite successful and well supports the hypothetical scenario proposed herein. The default values listed in Table 1 were adopted for the case studies illustrating our findings; these choices are certainly not strictly required—in fact, the modeling results were indicated to be robust to moderate adjustments of most parameters (see Figure 4 and Figure S4).

4.3. Some Detailed Assumptions in Consideration of Relevant Mechanisms

A linear RNA containing the characteristic sequence of a ribozyme can function as the ribozyme only if its length is less than 1.5 times that of the characteristic sequence, because an excessive number of redundant residues may severely interfere with the folding of the catalytic domain. Therefore, for an 11-nt circular single-gene genome containing a 3-nt non-coding sequence, a linear product (generated by chain breaking or “template-directed synthesis”; see Figure 1) that includes the ribozyme’s characteristic sequence should always be an active ribozyme (e.g., the cases shown in Figure 2 and Figure S2). However, for a 16-nt circular two-gene genome or a 24-nt circular three-gene genome, a linear product that includes one ribozyme’s characteristic sequence is not always an active ribozyme (e.g., the cases shown in Figure 6)—although an overly long product may become a functional ribozyme after “appropriate” chain breaking, chain-end decay, or cleavage by ER.
An RNA template may attract a substrate (nucleotide or oligomer) at any “empty” site, provided that the substrate is no longer than the empty site. The attraction aided by a preceding substrate (acting as a “primer”) on the template is easier than de novo attraction by a factor of FDA (FDA>1). A linear template has lower templating efficiency than a circular template by a factor of FLT (FLT<1). That is, for primer-aided attraction on a circular template, the corresponding probability is PAT; for de novo attraction on a circular template, the probability is PAT/FDA; for primer-aided attraction on a linear template, the probability is PAT×FLT; and for de novo attraction on a linear template, the probability is PAT×FLT/FDA.
The probability of strand separation in a duplex RNA is assumed to be PSPr, where r = n1/2 and n is the number of base pairs in the duplex. When n = 1, the probability equals PSP. As n increases, the probability decreases (since PSP, as a probability, lies between 0 and 1). In other words, strand separation becomes increasingly difficult for longer duplexes with more base pairs. The exponent 1/2 accounts for the fact that self-folding of the individual single strands may promote the dissociation of the duplex.
For the chain breaking of RNA molecules, when cleavage occurs in a single-stranded region, the rate is PBB. When cleavage occurs within a double-stranded region, the two parallel bonds may break simultaneously, with a probability of PBB3/2. The use of the exponent 3/2, rather than 2, accounts for the synergistic effect of the two cleavage events.
Nucleotide residues at the ends of a linear RNA chain (either the 3’- or 5’-end) may undergo decay. Following similar reasoning to the RNA chain-breaking described above, when the chain end is in a single-stranded state, the decay rate is PNDE; when the chain end is in a double-stranded state, the two paired nucleotide residues may decay simultaneously, with a probability of PNDE3/2.
A linear RNA may be cleaved at its 5’-end (single-stranded) by the ER with a probability of PBER; if the RNA is an active ribozyme, the probability becomes PBER×PRP, reflecting its resistance to the cleavage due to its functional conformation. Notably, in the model, the ER represents a 5’-end nuclease—in principle, however, a 3’-end nuclease is fully equivalent with respect to its capability of spreading in the system.
The probability of an RNA molecule moving is assumed to be PMN/m1/2, where m is the mass of the RNA relative to that of a nucleotide. This assumption accounts for the effect of molecular size on molecular motion. The square-root dependence is adopted here according to the Zimm model, which describes the diffusion coefficient of polymer molecules in solution [99].
(Note: The C-language source codes for the simulation program are available on GitHub, as stated in the Data Availability Statement. Besides the role of evidencing the reproducibility of the present study, the codes contain model implementation details and can facilitate readers’ in-depth understanding of the simulation mechanism.)

Supplementary Materials

The following supporting information can be downloaded at the website of this paper posted on Preprints.org. as a combined PDF file, containing Figures S1-S6, Table S1, and Box S1.

Author Contributions

Conceptualization, M.W.T.; methodology, M.W.T. and Y.C.W.; investigation, L.M.L. and J.H.Y; writing—original draft preparation, L.M.L..; writing—review and editing, M.W.T., L.M.L. and J.H.Y.; funding acquisition, M.W.T. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National Natural Science Foundation of China (No. 31571367) and Natural Science Foundation of Hubei Province (CN) (No. 2019CFB685).

Data Availability Statement

Source codes of the simulation program can be obtained from: https://github.com/mwt2001gh/er/blob/main/ER-Figure 2b.cpp (for the case in Figure 2b); https://github.com/mwt2001gh/er/blob/main/ER-Figure 6c.cpp (for the case in Figure 6c).

Conflicts of Interest

The authors declare no conflict of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of the data; in the writing of the manuscript; or in the decision to publish the results.

References

  1. Gilbert, W. The RNA world. Nature 1986, 319, 618. [Google Scholar] [CrossRef]
  2. Joyce, G.F. The antiquity of RNA-based evolution. Nature 2002, 418, 214–221. [Google Scholar] [CrossRef] [PubMed]
  3. Higgs, P.G.; Lehman, N. The RNA World: molecular cooperation at the origins of life. Nat. Rev. Genet. 2015, 16, 7–17. [Google Scholar] [CrossRef] [PubMed]
  4. Kruger, K.; Grabowski, P.E.; Zaug, A.J.; Sands, J.; Gottschling, D.E.; Cech, T.R. Self-splicing RNA: Autoexcision and autocyclization of ribosomal RNA intervening sequence of Tetrahymena. Cell 1982, 31, 147–157. [Google Scholar] [CrossRef] [PubMed]
  5. Guerrier-Takada, C.; Gardiner, K.; Marsh, T.; Pace, N.; Altman, S. The RNA moiety of ribonuclease P is the catalytic subunit of the enzyme. Cell 1983, 35, 849–857. [Google Scholar] [CrossRef] [PubMed]
  6. Nissen, P.; Hansen, J.; Ban, N.; Moore, P.B.; Steitz, T.A. The structural basis of ribosome activity in peptide bond synthesis. Science 2000, 289, 920–930. [Google Scholar] [CrossRef] [PubMed]
  7. Yusupov, M.M.; Yusupova, G.Z.; Baucom, A.; Lieberman, K.; Earnest, T.N.; Cate, J.H.D.; et al. Crystal structure of the ribosome at 5.5 angstrom resolution. Science 2001, 292, 883–896. [Google Scholar] [PubMed]
  8. Joyce, G.F.; Orgel, L.E. Chapter 2—Progress toward understanding the origin of the RNA World. In The RNA World; Gesteland, R.F., Cech, T.R., Atkins, J.F., Eds.; Cold Spring Harbor Laboratory Press: New York, NY, USA, 2006; pp. 23–56. [Google Scholar]
  9. Kun, A.; Szilágyi, A.; Könnyü, B.; Boza, G.; Zachar, I.; Szathmáry, E. The dynamics of the RNA world: insights and challenges. Ann. N. Y. Acad. Sci. 2015, 1341, 75–95. [Google Scholar] [CrossRef] [PubMed]
  10. Pressman, A.; Blanco, C.; Chen, I.A. The RNA world as a model system to study the origin of life. Curr. Biol. 2015, 25, R953–R963. [Google Scholar] [CrossRef] [PubMed]
  11. Ma, W.; Liang, Y. Chapter 12—Investigating prebiotic protocells for an understanding of the origin of life: A comprehensive perspective combining the chemical, evolutionary and historical aspects. In Probiotic Chemistry and Life’s Origin; Fiore, M., Ed.; The Royal Society of Chemistry: London, UK, 2022; pp. 347–378. [Google Scholar]
  12. Powner, M.W.; Gerland, B.; Sutherland, J.D. Synthesis of activated pyrimidine ribonucleotides in prebiotically plausible conditions. Nature 2009, 459, 239–242. [Google Scholar] [CrossRef] [PubMed]
  13. Powner, M.W.; Sutherland, J.D.; Szostak, J.W. Chemoselective multicomponent one-pot assembly of purine precursors in water. J. Am. Chem. Soc. 2010, 132, 16677–16688. [Google Scholar] [CrossRef] [PubMed]
  14. Powner, M.W.; Sutherland, J.D.; Szostak, J.W. The origin of nucleotides. Synlett 2011, 14, 1956–1964. [Google Scholar] [CrossRef]
  15. Becker, S.; Feldmann, J.; Wiedemann, S.; Okamura, H.; Schneider, C.; Iwan, K.; et al. Unified prebiotically plausible synthesis of pyrimidine and purine RNA ribonucleotides. Science 2019, 366, 76–82. [Google Scholar] [CrossRef] [PubMed]
  16. Xu, J.F.; Chmela, V.; Green, N.J.; Russell, D.A.; Janicki, M.J.; Gora, R.W.; et al. Selective prebiotic formation of RNA pyrimidine and DNA purine nucleosides. Nature 2020, 582, 60–66. [Google Scholar] [CrossRef] [PubMed]
  17. Robertson, M.P.; Joyce, G.F. The origins of the RNA world. Cold Spring Harb. Perspect. Biol. 2012, 4, a003608. [Google Scholar] [PubMed]
  18. Szathmáry, E.; Maynard-Smith, J. From replicators to reproducers: the first major transitions leading to life. J. Theor. Biol. 1997, 187, 555–571. [Google Scholar] [CrossRef] [PubMed]
  19. Szathmáry, E. The origin of replicators and reproducers. Philos. Trans. R. Soc. B 2006, 361, 1761–1776. [Google Scholar] [CrossRef]
  20. Takeuchi, N.; Hogeweg, P. Evolution of complexity in RNA-like replicator systems. Biol. Direct 2008, 3, 11. [Google Scholar] [CrossRef] [PubMed]
  21. Takeuchi, N.; Hogeweg, P. Evolutionary dynamics of RNA-like replicator systems: a bioinformatic approach to the origin of life. Phys. Life Rev. 2012, 9, 219–263. [Google Scholar] [CrossRef] [PubMed]
  22. Ma, W.T. What does “the RNA world” mean to the origin of life? Life 2017, 7, 49. [Google Scholar] [CrossRef] [PubMed]
  23. Szathmáry, E.; Maynard-Smith, J. The major evolutionary transitions. Nature 1995, 374, 227–232. [Google Scholar] [CrossRef] [PubMed]
  24. Szathmáry, E. Toward major evolutionary transitions theory 2.0. Proc. Natl. Acad. Sci. U. S. A. 2015, 112, 10104–10111. [Google Scholar] [CrossRef] [PubMed]
  25. Szostak, J.W.; Bartel, D.P.; Luisi, P.L. Synthesizing life. Nature 2001, 409, 387–390. [Google Scholar] [CrossRef] [PubMed]
  26. Joyce, G.F.; Szostak, J.W. Protocells and RNA self-replication. Cold Spring Harb. Perspect. Biol. 2018, 10, a034801. [Google Scholar] [CrossRef] [PubMed]
  27. Bartel, D.P. Chapter 5—Re-creating an RNA replicase. In The RNA World; Gesteland, R.F., Cech, T.R., Atkins, J.F., Eds.; Cold Spring Harbor Laboratory Press: New York, NY, USA, 1999; pp. 143–162. [Google Scholar]
  28. Johnston, W.K.; Unrau, P.J.; Lawrence, M.S.; Glasner, M.E.; Bartel, D.P. RNA-catalyzed RNA polymerization: accurate and general RNA-templated primer extension. Science 2001, 292, 1319–1325. [Google Scholar] [CrossRef] [PubMed]
  29. Zaher, H.S.; Unrau, P.J. Selection of an improved RNA polymerase ribozyme with superior extension and fidelity. RNA 2007, 13, 1017–1026. [Google Scholar] [CrossRef] [PubMed]
  30. Wochner, A.; Attwater, J.; Coulson, A.; Holliger, P. Ribozyme-catalyzed transcription of an active ribozyme. Science 2011, 332, 209–212. [Google Scholar] [CrossRef] [PubMed]
  31. Attwater, J.; Wochner, A.; Holliger, P. In-ice evolution of RNA polymerase ribozyme activity. Nat. Chem. 2013, 5, 1011–1018. [Google Scholar] [CrossRef] [PubMed]
  32. Horning, D.P.; Joyce, G.F. Amplification of RNA by an RNA polymerase ribozyme. Proc. Natl. Acad. Sci. U. S. A. 2016, 113, 9786–9791. [Google Scholar] [CrossRef] [PubMed]
  33. Tjhung, K.F.; Shokhirev, M.N.; Horning, D.P.; Joyce, G.F. An RNA polymerase ribozyme that synthesizes its own ancestor. Proc. Natl. Acad. Sci. U. S. A. 2020, 117, 2906–2913. [Google Scholar] [CrossRef] [PubMed]
  34. Portillo, X.; Huang, Y.T.; Breaker, R.R.; Horning, D.P.; Joyce, G.F. Witnessing the structural evolution of an RNA enzyme. eLife 2021, 10, e71557. [Google Scholar] [CrossRef] [PubMed]
  35. Cojocaru, R.; Unrau, P.J. Processive RNA polymerization and promoter recognition in an RNA World. Science 2021, 371, 1225–1232. [Google Scholar] [CrossRef] [PubMed]
  36. Papastavrou, N.; Horning, D.P.; Joyce, G.F. RNA-catalyzed evolution of catalytic RNA. Proc. Natl. Acad. Sci. U. S. A. 2024, 121, e2321592121. [Google Scholar] [CrossRef] [PubMed]
  37. Attwater, J.; Raguram, A.; Morgunov, A.S.; Gianni, E.; Holliger, P. Ribozyme-catalysed RNA synthesis using triplet building blocks. eLife 2018, 7, e35255. [Google Scholar] [CrossRef] [PubMed]
  38. Kristoffersen, E.L.; Burman, M.; Noy, A.; Holliger, P. Rolling circle RNA synthesis catalyzed by RNA. eLife 2022, 11, e75186. [Google Scholar] [CrossRef] [PubMed]
  39. McRae, E.K.S.; Wan, C.J.K.; Kristoffersen, E.L.; Hansen, K.; Gianni, E.; Gallego, I.; et al. Cryo-EM structure and functional landscape of an RNA polymerase ribozyme. Proc. Natl. Acad. Sci. U. S. A. 2024, 121, e2313332121. [Google Scholar] [CrossRef] [PubMed]
  40. Attwater, J.; Augustin, T.L.; Curran, J.F.; Kwok, S.L.Y.; Ohlendorf, L.; Gianni, E.; et al. Trinucleotide substrates under pH-freeze-thaw cycles enable open-ended exponential RNA replication by a polymerase ribozyme. Nat. Chem. 2024, 17, 1129–1137. [Google Scholar]
  41. Gianni, E.; Kwok, S.L.Y.; Wan, C.J.K.; Goeij, K.; Clifton, B.E.; Colizzi, E.S.; et al. A small polymerase ribozyme that can synthesize itself and its complementary strand. Science 2026, 391, eadt2760. [Google Scholar] [CrossRef]
  42. Szabó, P.; Scheuring, I.; Czárán, T.; Szathmáry, E. In silico simulations reveal that replicators with limited dispersal evolve towards higher efficiency and fidelity. Nature 2002, 420, 340–343. [Google Scholar] [CrossRef] [PubMed]
  43. Takeuchi, N.; Hogeweg, P. Multilevel selection in models of prebiotic evolution II: a direct comparison of compartmentalization and spatial self-organization. PLoS Comput. Biol. 2009, 5, e1000542. [Google Scholar] [CrossRef] [PubMed]
  44. Ma, W.T.; Yu, C.W.; Zhang, W.T.; Hu, J.M. A simple template-dependent ligase ribozyme as the RNA replicase emerging first in the RNA world. Astrobiology 2010, 10, 437–447. [Google Scholar] [CrossRef] [PubMed]
  45. Szostak, J.W. The eightfold path to non-enzymatic RNA replication. J. Syst. Chem. 2012, 3, 2. [Google Scholar] [CrossRef]
  46. Deck, D.; Jauker, M.; Richert, C. Efficient enzyme-free copying of all four nucleobases templated by immobilized RNA. Nat. Chem. 2011, 3, 603–608. [Google Scholar] [CrossRef] [PubMed]
  47. Kaiser, A.; Richert, C. Nucleotide-based copying of nucleic acid sequences without enzymes. J. Org. Chem. 2013, 78, 793–799. [Google Scholar] [CrossRef] [PubMed]
  48. Li, L.; Prywes, N.; Tam, C.P.; O’Flaherty, D.K.; Lelyveld, V.S.; Izgu, E.C.; et al. Enhanced nonenzymatic RNA copying with 2-aminoimidazole activated nucleotides. J. Am. Chem. Soc. 2017, 139, 1810–1813. [Google Scholar] [CrossRef] [PubMed]
  49. O’Flaherty, D.K.; Zhou, L.J.; Szostak, J.W. Nonenzymatic template-directed synthesis of mixed-sequence 3 ‘-NP-DNA up to 25 nucleotides long inside model protocells. J. Am. Chem. Soc. 2019, 141, 10481–10488. [Google Scholar] [PubMed]
  50. Ma, W.T.; Yu, C.W.; Zhang, W.T.; Hu, J.M. Nucleotide synthetase ribozymes may have emerged first in the RNA world. RNA 2007, 13, 2012–2019. [Google Scholar] [CrossRef] [PubMed]
  51. Ma, W.T.; Hu, J.M. Computer simulation on the cooperation of functional molecules during the early stages of evolution. PLoS ONE 2012, 7, e35454. [Google Scholar] [CrossRef] [PubMed]
  52. Kim, Y.E.; Higgs, P.G. Co-operation between polymerases and nucleotide synthetases in the RNA world. PLoS Comput. Biol. 2016, 12, e1005161. [Google Scholar] [CrossRef] [PubMed]
  53. Unrau, P.J.; Bartel, D.P. RNA-catalyzed nucleotide synthesis. Nature 1998, 395, 260–263. [Google Scholar] [CrossRef] [PubMed]
  54. Lau, M.W.L.; Cadieux, K.E.C.; Unrau, P.J. Isolation of fast purine nucleotide synthase ribozymes. J. Am. Chem. Soc. 2004, 126, 15686–15693. [Google Scholar] [CrossRef] [PubMed]
  55. Lau, M.W.L.; Unrau, P.J. A promiscuous ribozyme promotes nucleotide synthesis in addition to ribose chemistry. Chem. Biol. 2009, 16, 815–825. [Google Scholar] [CrossRef] [PubMed]
  56. Ferre-D’Amare, A.R.; Scott, W.G. Small self-cleaving ribozymes. Cold Spring Harb. Perspect. Biol. 2010, 2, a003574. [Google Scholar] [CrossRef] [PubMed]
  57. Weinberg, C.E.; Weinberg, Z.; Hammann, C. Novel ribozymes: discovery, catalytic mechanisms, and the quest to understand biological function. Nucleic Acids Res. 2019, 47, 9480–9494. [Google Scholar] [CrossRef] [PubMed]
  58. Peng, H.; Latifi, B.; Müller, S.; Lupták, A.; Chen, I.A. Self-cleaving ribozymes: substrate specificity and synthetic biology applications. RSC Chem. Biol. 2021, 2, 1370–1383. [Google Scholar] [CrossRef] [PubMed]
  59. Hammann, C.; Lilley, D.M.J. Folding and activity of the hammerhead ribozyme. ChemBioChem 2002, 3, 691–700. [Google Scholar] [CrossRef]
  60. Zhang, Z.; Hong, X.; Xiong, P.; Wang, J.F.; Zhou, Y.Q.; Zhan, J. Minimal twister sister-like self-cleaving ribozymes in the human genome revealed by deep mutational scanning. eLife 2024, 12, RP90254. [Google Scholar] [CrossRef] [PubMed]
  61. Williams, K.P.; Ciafre, S.; Tocchinivalentini, G.P. Selection of novel Mg2+-dependent self-cleaving ribozymes. EMBO J. 1995, 14, 4551–4557. [Google Scholar] [PubMed]
  62. Tang, J.; Breaker, R.R. Structural diversity of self-cleaving ribozymes. Proc. Natl. Acad. Sci. U. S. A. 2000, 97, 5784–5789. [Google Scholar] [CrossRef] [PubMed]
  63. Wilson, T.J.; Lilley, D.M.J. The evolution of ribozyme chemistry. Science 2009, 323, 1436–1438. [Google Scholar] [CrossRef] [PubMed]
  64. Wilson, T.J.; Liu, Y.J.; Lilley, D.M.J. Ribozymes and the mechanisms that underlie RNA catalysis. Front. Chem. Sci. Eng. 2016, 10, 178–185. [Google Scholar] [CrossRef]
  65. Luo, Y.F.; Liang, M.L.; Yu, C.W.; Ma, W.T. Circular at the very beginning: on the initial genomes in the RNA world. RNA Biol. 2024, 21, 17–31. [Google Scholar] [CrossRef] [PubMed]
  66. Winnik, M.A. End-to-end cyclization of polymer chains. Acc. Chem. Res. 1985, 18, 73–79. [Google Scholar]
  67. Mueller, S.; Appel, B. In vitro circularization of RNA. RNA Biol. 2017, 14, 1018–1027. [Google Scholar]
  68. Hassenkam, T.; Damer, B.; Mednick, G.; Deamer, D. AFM images of viroid-sized rings that self-assemble from mononucleotides through wet-dry cycling: implications for the origin of life. Life 2020, 10, 321. [Google Scholar] [CrossRef] [PubMed]
  69. Ma, W.T.; Yu, C.W.; Zhang, W.T. Monte Carlo simulation of early molecular evolution in the RNA World. Biosystems 2007, 90, 28–39. [Google Scholar] [CrossRef] [PubMed]
  70. Chen, Y.; Ma, W.T. The origin of biological homochirality along with the origin of life. PLoS Comput. Biol. 2020, 16, e1007592. [Google Scholar] [CrossRef] [PubMed]
  71. Ferris, J.P. Montmorillonite catalysis of 30-50 mer oligonucleotides: Laboratory demonstration of potential steps in the origin of the RNA world. Orig. Life Evol. Biosph. 2002, 32, 311–332. [Google Scholar] [PubMed]
  72. Ferris, J.P.; Hill, A.R.; Liu, R.; Orgel, L.E. Synthesis of long prebiotic oligomers on mineral surfaces. Nature 1996, 381, 59–61. [Google Scholar] [CrossRef] [PubMed]
  73. Ertem, G. Montmorillonite, oligonucleotides, RNA and origin of life. Orig. Life Evol. Biosph. 2004, 34, 549–570. [Google Scholar] [CrossRef] [PubMed]
  74. Franchi, M.; Gallori, E. A surface-mediated origin of the RNA world: biogenic activities of clay-adsorbed RNA molecules. Gene 2005, 346, 205–214. [Google Scholar] [PubMed]
  75. Loria, A.; Pan, T. The 3′ substrate determinants for the catalytic efficiency of the Bacillus subtilis RNase P holoenzyme suggest autolytic processing of the RNase P RNA in vivo. RNA 2000, 6, 1413–1422. [Google Scholar] [CrossRef] [PubMed]
  76. Hansen, A.; Pfeiffer, T.; Zuleeg, T.; Limmer, S.; Ciesiolka, J.; Feltens, R.; et al. Exploring the minimal substrate requirements for trans-cleavage by RNase P holoenzymes from Escherichia coli and Bacillus subtilis. Mol. Microbiol. 2001, 41, 131–143. [Google Scholar] [PubMed]
  77. Kirsebom, L.A.; Liu, F.; McClain, W.H. The discovery of a catalytic RNA within RNase P and its legacy. J. Biol. Chem. 2024, 300, 107318. [Google Scholar] [CrossRef] [PubMed]
  78. Joyce, G.F. Foreword. In Origins of Life: The Central Concepts; Deamer, D.W., Fleischaker, G.R., Eds.; Jones and Bartlett Publishers: Boston, MA, USA, 1994; pp. xi–xii. [Google Scholar]
  79. Luisi, P.L. About various definitions of life. Orig. Life Evol. Biosph. 1998, 28, 613–622. [Google Scholar] [CrossRef] [PubMed]
  80. Benner, S.A. Defining Life. Astrobiology 2010, 10, 1021–1030. [Google Scholar] [CrossRef] [PubMed]
  81. Ma, W.T. The essence of life. Biol. Direct 2016, 11, 49. [Google Scholar] [CrossRef] [PubMed]
  82. Shapiro, R. A replicator was not involved in the origin of life. IUBMB Life 2000, 49, 173–176. [Google Scholar] [CrossRef] [PubMed]
  83. Anet, F.A. The place of metabolism in the origin of life. Curr. Opin. Chem. Biol. 2004, 8, 654–659. [Google Scholar] [CrossRef] [PubMed]
  84. Pross, A. Causation and the origin of life. metabolism or replication first? Orig. Life Evol. Biosph. 2004, 34, 307–321. [Google Scholar] [CrossRef] [PubMed]
  85. Vasas, V.; Szathmáry, E.; Santos, M. Lack of evolvability in self-sustaining autocatalytic networks constraints metabolism-first scenarios for the origin of life. Proc. Natl. Acad. Sci. U. S. A. 2010, 107, 1470–1475. [Google Scholar] [PubMed]
  86. Moore, A. The mark of metabolism: another nail in the coffin of nucleic-acids-first in the origin of life? Bioessays 2014, 36, 221–222. [Google Scholar] [PubMed]
  87. Yin, S.L.; Chen, Y.; Yu, C.W.; Ma, W.T. From molecular to cellular form: modeling the first major transition during the arising of life. BMC Evol. Biol. 2019, 19, 84. [Google Scholar] [CrossRef] [PubMed]
  88. Kikovska, E.; Wu, S.Y.; Mao, G.Z.; Kirsebom, L.A. Cleavage mediated by the P15 domain of bacterial RNase P RNA. Nucleic Acids Res. 2012, 40, 2224–2233. [Google Scholar] [PubMed]
  89. Ma, W.T.; Yu, C.W. From experimental clues to theoretical modeling: evolution associated with the membrane-takeover at an early stage of life. PLoS Comput. Biol. 2025, 21, e1012763. [Google Scholar] [CrossRef] [PubMed]
  90. Jauker, M.; Griesser, H.; Richert, C. Copying of RNA sequences without pre-activation. Angew. Chem. Int. Ed. 2015, 54, 14559–14563. [Google Scholar] [CrossRef]
  91. Zhang, S.J.; Duzdevich, D.; Ding, D.; Szostak, J.W. Freeze-thaw cycles enable a prebiotically plausible and continuous pathway from nucleotide activation to nonenzymatic RNA copying. Proc. Natl. Acad. Sci. U. S. A. 2022, 119, e2116429119. [Google Scholar] [CrossRef] [PubMed]
  92. Ding, D.; Zhang, S.J.; Szostak, J.W. Enhanced nonenzymatic RNA copying with in-situ activation of short oligonucleotides. Nucleic Acids Res. 2023, 51, 6528–6539. [Google Scholar] [CrossRef] [PubMed]
  93. Russell, M.J.; Hall, A.J.; Cairns-Smith, A.G.; Braterman, P.S. Submarine hot spring and origin of life. Nature 1988, 336, 117. [Google Scholar] [CrossRef]
  94. Martin, W.; Baross, J.; Kelley, D.; Russell, M.J. Hydrothermal vents and the origin of life. Nat. Rev. Microbiol. 2008, 6, 805–814. [Google Scholar] [CrossRef] [PubMed]
  95. Colin-Garcia, M.; Heredia, A.; Cordero, G.; Camprubi, A.; Negron-Mendoza, A.; Ortega-Gutierrez, F.; et al. Hydrothermal vents and prebiotic chemistry: a review. Bol. Soc. Geol. Mex. 2016, 68, 599–620. [Google Scholar] [CrossRef]
  96. Damer, B.; Deamer, D. Coupled phases and combinatorial selection in fluctuating hydrothermal pools: a scenario to guide experimental approaches to the origin of cellular life. Life 2015, 5, 872–887. [Google Scholar] [CrossRef] [PubMed]
  97. Damer, B.; Deamer, D. The hot spring hypothesis for an origin of life. Astrobiology 2020, 20, 429–452. [Google Scholar] [PubMed]
  98. Liang, Y.Z.; Yu, C.W.; Ma, W.T. The automatic parameter-exploration with a machine-learning-like approach: powering the evolutionary modeling on the origin of life. PLoS Comput. Biol. 2021, 17, e1009761. [Google Scholar] [PubMed]
  99. Zimm, B.H. Dynamics of polymer molecules in dilute solution: viscoelasticity, flow birefringence and dielectric loss. J. Chem. Phys. 1956, 24, 269–278. [Google Scholar] [CrossRef]
Figure 1. Conceptual schemes of exonuclease ribozyme (ER) and circular genomes at the early stage of the RNA world. (a) An ER cleaves an RNA molecule at its chain end and releases a mononucleotide. (b) The upper panel corresponds to a “compact” single-gene genome. Blue lines represent the ribozyme sequence and light-blue lines represent its complementary sequence. The ribozyme must be generated through precise breaking (the arrow) of the sense chain or through “accurate RNA synthesis” on the antisense chain template—specifically, initiation at the correct position and timely dissociation of the linear sense chain prior to circularization (the triangle). The lower panel corresponds to a single-gene genome containing a non-coding sequence. Thin black lines represent the non-coding sequence, and thin grey lines represent its complement. In this case, the ribozyme can be generated by any chain-breaking event within the non-coding region (the arrow), or by dissociation of the linear sense chain that already contains the full ribozyme sequence from the template before complete circularization. (c) A two-gene genome. Red lines represent the sequence of the second ribozyme, and light-red lines represent its complement. A ribozyme may be generated by breaking the circular sense chain at the region of the other ribozyme (the arrow), or via dissociation of a linear sense chain carrying the full ribozyme sequence from the template.
Figure 1. Conceptual schemes of exonuclease ribozyme (ER) and circular genomes at the early stage of the RNA world. (a) An ER cleaves an RNA molecule at its chain end and releases a mononucleotide. (b) The upper panel corresponds to a “compact” single-gene genome. Blue lines represent the ribozyme sequence and light-blue lines represent its complementary sequence. The ribozyme must be generated through precise breaking (the arrow) of the sense chain or through “accurate RNA synthesis” on the antisense chain template—specifically, initiation at the correct position and timely dissociation of the linear sense chain prior to circularization (the triangle). The lower panel corresponds to a single-gene genome containing a non-coding sequence. Thin black lines represent the non-coding sequence, and thin grey lines represent its complement. In this case, the ribozyme can be generated by any chain-breaking event within the non-coding region (the arrow), or by dissociation of the linear sense chain that already contains the full ribozyme sequence from the template before complete circularization. (c) A two-gene genome. Red lines represent the sequence of the second ribozyme, and light-red lines represent its complement. A ribozyme may be generated by breaking the circular sense chain at the region of the other ribozyme (the arrow), or via dissociation of a linear sense chain carrying the full ribozyme sequence from the template.
Preprints 219604 g001
Figure 2. The spread of ER encoded in a circular genome. Legends: cir_er – circular RNA containing the ER sequence (i.e., the characteristic sequence of ER); cir_ercom – circular RNA containing the complement of the ER sequence; lin_er – linear RNA containing the ER sequence; lin_ercom – linear RNA containing the complement of the ER sequence; cir_ct – circular RNA containing the control sequence; cir_ctcom – circular RNA containing the complement of the control sequence. (a) At step 1×104, 10 circular RNA molecules with the ER sequence (8 nt in length; see Table 1) and an equal number of circular RNA molecules with the control sequence (8 nt; see Table 1) are inoculated into the system (at randomly chosen locations—the same applies to all inoculations described below). (b) At step 1×104, 10 circular RNA molecules with the ER sequence plus a 3-nt non-coding sequence (11 nt in total length) and an equal number of circular RNA molecules with the control sequence plus the non-coding sequence (11 nt) are inoculated into the system. (c) Based on the case depicted in (b), at step 2.5×106 (the red arrow), all circular RNA molecules are artificially cleaved into linear forms—wherein the ER sequence and its complementary sequence are retained (i.e., cleavage sites are located within the non-coding region); PEL is set to 0; FLT is set to 1. (d) At step 1 × 104, 20 circular RNA molecules (11 nt) with random sequences are inoculated into the system. One circular ER molecule (11 nt; identical to the one depicted in b) and one circular control molecule (11 nt) are inoculated into the system every 1 × 105 steps.
Figure 2. The spread of ER encoded in a circular genome. Legends: cir_er – circular RNA containing the ER sequence (i.e., the characteristic sequence of ER); cir_ercom – circular RNA containing the complement of the ER sequence; lin_er – linear RNA containing the ER sequence; lin_ercom – linear RNA containing the complement of the ER sequence; cir_ct – circular RNA containing the control sequence; cir_ctcom – circular RNA containing the complement of the control sequence. (a) At step 1×104, 10 circular RNA molecules with the ER sequence (8 nt in length; see Table 1) and an equal number of circular RNA molecules with the control sequence (8 nt; see Table 1) are inoculated into the system (at randomly chosen locations—the same applies to all inoculations described below). (b) At step 1×104, 10 circular RNA molecules with the ER sequence plus a 3-nt non-coding sequence (11 nt in total length) and an equal number of circular RNA molecules with the control sequence plus the non-coding sequence (11 nt) are inoculated into the system. (c) Based on the case depicted in (b), at step 2.5×106 (the red arrow), all circular RNA molecules are artificially cleaved into linear forms—wherein the ER sequence and its complementary sequence are retained (i.e., cleavage sites are located within the non-coding region); PEL is set to 0; FLT is set to 1. (d) At step 1 × 104, 20 circular RNA molecules (11 nt) with random sequences are inoculated into the system. One circular ER molecule (11 nt; identical to the one depicted in b) and one circular control molecule (11 nt) are inoculated into the system every 1 × 105 steps.
Preprints 219604 g002
Figure 3. Snapshots illustrating the spread of ER encoded in a circular genome. The evolutionary dynamics of this case are depicted in Figure 2d. Raw materials (nucleotide precursors) are represented by a yellow background, where the color depth indicates their quantity in the corresponding room (gray-lined grid). Circular molecules are denoted by circles (with sizes correlated), while linear molecules are denoted by line segments (with lengths correlated). Blue corresponds to ER and green corresponds to the control. The snapshot at step 10,000 represents the system before the inoculation of the 20 random-sequence RNA molecules—raw materials are abundant at this stage. Inoculation of one ER molecule every 100,000 steps is performed when the raw material availability is significantly lower—this situation can be observed in the snapshot at step 600,000, which captures the key inoculation event that subsequently drives the spread of ER (the blue arrow points to the inoculated ER molecule and the green arrow points to the inoculated control molecule). The subsequent snapshots demonstrate the gradual spread of ER throughout the system. Note that the grid has a toroidal topology.
Figure 3. Snapshots illustrating the spread of ER encoded in a circular genome. The evolutionary dynamics of this case are depicted in Figure 2d. Raw materials (nucleotide precursors) are represented by a yellow background, where the color depth indicates their quantity in the corresponding room (gray-lined grid). Circular molecules are denoted by circles (with sizes correlated), while linear molecules are denoted by line segments (with lengths correlated). Blue corresponds to ER and green corresponds to the control. The snapshot at step 10,000 represents the system before the inoculation of the 20 random-sequence RNA molecules—raw materials are abundant at this stage. Inoculation of one ER molecule every 100,000 steps is performed when the raw material availability is significantly lower—this situation can be observed in the snapshot at step 600,000, which captures the key inoculation event that subsequently drives the spread of ER (the blue arrow points to the inoculated ER molecule and the green arrow points to the inoculated control molecule). The subsequent snapshots demonstrate the gradual spread of ER throughout the system. Note that the grid has a toroidal topology.
Preprints 219604 g003
Figure 4. The influence of several key parameters on the spread of ER. The vertical axis denotes the number of circular RNA molecules containing the ER sequence, serving as an indication of ER spread level. Legends: v0 – all parameters are set to default values; v_up – the relevant parameter depicted in the top-left corner of the panel is turned up (other parameters remain default values); v_down – the relevant parameter is turned down (the legends apply to all the subfigures). Red arrows mark the critical steps at which parameter adjustments are performed. For (PBER), the default value 0.9 is turned up to 0.95, 0.98 and 0.99 at these respective points, and turned down to 0.2, 0.05 and 0 at these points. For (TER), the default value 10 is turned up to 20, 50 and 100, and turned down to 5, 2 and 1. For (PTL), the default value 0.05 is turned up to 0.1, 0.2 and 0.5, and turned down to 0.02, 0.01 and 0.005. For (PRP), the default value 0.5 is turned up to 0.8, 0.9 and 0.95, and turned down to 0.2, 0.1 and 0.
Figure 4. The influence of several key parameters on the spread of ER. The vertical axis denotes the number of circular RNA molecules containing the ER sequence, serving as an indication of ER spread level. Legends: v0 – all parameters are set to default values; v_up – the relevant parameter depicted in the top-left corner of the panel is turned up (other parameters remain default values); v_down – the relevant parameter is turned down (the legends apply to all the subfigures). Red arrows mark the critical steps at which parameter adjustments are performed. For (PBER), the default value 0.9 is turned up to 0.95, 0.98 and 0.99 at these respective points, and turned down to 0.2, 0.05 and 0 at these points. For (TER), the default value 10 is turned up to 20, 50 and 100, and turned down to 5, 2 and 1. For (PTL), the default value 0.05 is turned up to 0.1, 0.2 and 0.5, and turned down to 0.02, 0.01 and 0.005. For (PRP), the default value 0.5 is turned up to 0.8, 0.9 and 0.95, and turned down to 0.2, 0.1 and 0.
Preprints 219604 g004
Figure 5. RNA distribution of the ER-spreading and REP-spreading systems. The charts depict the states of the two systems at step 5×106, when the corresponding populations have reached a stable spread (see Figure 2b and Figure S2a for their evolutionary dynamics, respectively). (a) The chain-length distribution (the upper panel for the ER system and the lower panel for the REP system). The bars denote the percentage distribution of RNA molecules in quantity (there is no RNA longer than 11 nt). Legends: total-RNA – all RNA molecules; er – RNA molecules containing ER or its complementary sequence; rep – RNA molecules containing REP or its complementary sequence. (b) The topology distribution (the upper panel for the ER system and the lower panel for the REP system). The pies illustrate the proportion of total materials (i.e., mass calculated based on nucleotide count) within the 8–11 nt range (wherein ribozymes distribute). Legends: cir_er – circular RNA containing ER or its complementary sequence; lin_er – linear RNA containing ER or its complementary sequence; cir_rep – circular RNA containing REP or its complementary sequence; lin_rep – linear RNA containing REP or its complementary sequence; cir_others – circular RNA neither containing the ribozyme’s sequence nor its complementary sequence; lin_others – linear RNA neither containing the ribozyme’s sequence nor its complementary sequence.
Figure 5. RNA distribution of the ER-spreading and REP-spreading systems. The charts depict the states of the two systems at step 5×106, when the corresponding populations have reached a stable spread (see Figure 2b and Figure S2a for their evolutionary dynamics, respectively). (a) The chain-length distribution (the upper panel for the ER system and the lower panel for the REP system). The bars denote the percentage distribution of RNA molecules in quantity (there is no RNA longer than 11 nt). Legends: total-RNA – all RNA molecules; er – RNA molecules containing ER or its complementary sequence; rep – RNA molecules containing REP or its complementary sequence. (b) The topology distribution (the upper panel for the ER system and the lower panel for the REP system). The pies illustrate the proportion of total materials (i.e., mass calculated based on nucleotide count) within the 8–11 nt range (wherein ribozymes distribute). Legends: cir_er – circular RNA containing ER or its complementary sequence; lin_er – linear RNA containing ER or its complementary sequence; cir_rep – circular RNA containing REP or its complementary sequence; lin_rep – linear RNA containing REP or its complementary sequence; cir_others – circular RNA neither containing the ribozyme’s sequence nor its complementary sequence; lin_others – linear RNA neither containing the ribozyme’s sequence nor its complementary sequence.
Preprints 219604 g005
Figure 6. The spread of multi-gene genomes. Legends: cir_errep – circular RNA containing the ER and the REP sequence; cir_ct – circular RNA containing the control sequence; lin_er_rib – linear RNA containing the ER sequence that functions as a ribozyme (i.e., an active ribozyme is assumed to be less than 1.5 times the length of its characteristic sequence; see Methods for explanation); lin_rep_rib – linear RNA containing the REP sequence that functions as a ribozyme; cir_er – circular RNA containing the ER sequence (but no REP sequence); cir_errepnr – circular RNA containing the ER, the REP and the NR sequence; lin_nr_rib – linear RNA containing the NR sequence that functions as a ribozyme. (a) At step 1×104, 10 circular RNA molecules with the sequences of ER and REP (16 nt in length, 8 nt for one gene), together with an equal number of control RNA molecules (also 16 nt in length), are inoculated into the system. (b) Based on the case shown in (a), PTLR is set to 0 (i.e., the catalytic ability of REP is turned off) at step 2×106 (the left arrow), and PBER is set to 0 (i.e., the catalytic ability of ER is turned off) at step 6×106 (the right arrow). (c) At step 1×104, 10 circular RNA molecules with the ER sequence and a non-coding sequence (16 nt in length, 8 nt for each), together with an equal number of control RNA molecules (also 16 nt in length), are inoculated into the system. The non-coding sequence is assumed to be “GAGUCACU”, which differs by only one residue from the REP sequence (“GAGUCUCU”, see Table 1 for reference). (d) At step 1×104, 10 circular RNA molecules with the sequences of ER, REP and NR (24 nt in length, 8 nt for each gene), together with an equal number of control RNA molecules (also 24 nt in length), are inoculated into the system.
Figure 6. The spread of multi-gene genomes. Legends: cir_errep – circular RNA containing the ER and the REP sequence; cir_ct – circular RNA containing the control sequence; lin_er_rib – linear RNA containing the ER sequence that functions as a ribozyme (i.e., an active ribozyme is assumed to be less than 1.5 times the length of its characteristic sequence; see Methods for explanation); lin_rep_rib – linear RNA containing the REP sequence that functions as a ribozyme; cir_er – circular RNA containing the ER sequence (but no REP sequence); cir_errepnr – circular RNA containing the ER, the REP and the NR sequence; lin_nr_rib – linear RNA containing the NR sequence that functions as a ribozyme. (a) At step 1×104, 10 circular RNA molecules with the sequences of ER and REP (16 nt in length, 8 nt for one gene), together with an equal number of control RNA molecules (also 16 nt in length), are inoculated into the system. (b) Based on the case shown in (a), PTLR is set to 0 (i.e., the catalytic ability of REP is turned off) at step 2×106 (the left arrow), and PBER is set to 0 (i.e., the catalytic ability of ER is turned off) at step 6×106 (the right arrow). (c) At step 1×104, 10 circular RNA molecules with the ER sequence and a non-coding sequence (16 nt in length, 8 nt for each), together with an equal number of control RNA molecules (also 16 nt in length), are inoculated into the system. The non-coding sequence is assumed to be “GAGUCACU”, which differs by only one residue from the REP sequence (“GAGUCUCU”, see Table 1 for reference). (d) At step 1×104, 10 circular RNA molecules with the sequences of ER, REP and NR (24 nt in length, 8 nt for each gene), together with an equal number of control RNA molecules (also 24 nt in length), are inoculated into the system.
Preprints 219604 g006
Figure 7. Snapshots illustrating the emergence of an ER-REP genome naturally from an ER genome-spreading system. The evolutionary dynamics of this case are depicted in Figure 6c. Raw materials (nucleotide precursors) are represented by a yellow background, where the color depth indicates their quantity in the corresponding room (gray-lined grid). Circular molecules are denoted by circles (with sizes correlated), while linear molecules are denoted by line segments (with lengths correlated). Blue circles denote circular RNA molecules containing the ER sequence (but no REP sequence). Green circles denote circular RNA molecules containing the control sequence. Blue line segments denote linear RNA molecules containing the ER sequence that function as ribozymes. Magenta circles denote circular RNA molecules containing both the ER and REP sequences. Red line segments denote linear RNA molecules containing the REP sequence that function as ribozymes. The snapshot at step 10,000 shows the inoculation of 10 circular ER molecules (containing a non-coding sequence) and 10 control RNA molecules. The snapshot at step 100,000 shows the spread of the circular ER genome. The snapshot at step 1,600,000 shows the appearance of circular RNA molecules containing both the ER and REP sequences (the arrow). The subsequent snapshots demonstrate the gradual spread of the circular ER-REP genome throughout the system and the simultaneous fading away of the single-gene genome containing ER.
Figure 7. Snapshots illustrating the emergence of an ER-REP genome naturally from an ER genome-spreading system. The evolutionary dynamics of this case are depicted in Figure 6c. Raw materials (nucleotide precursors) are represented by a yellow background, where the color depth indicates their quantity in the corresponding room (gray-lined grid). Circular molecules are denoted by circles (with sizes correlated), while linear molecules are denoted by line segments (with lengths correlated). Blue circles denote circular RNA molecules containing the ER sequence (but no REP sequence). Green circles denote circular RNA molecules containing the control sequence. Blue line segments denote linear RNA molecules containing the ER sequence that function as ribozymes. Magenta circles denote circular RNA molecules containing both the ER and REP sequences. Red line segments denote linear RNA molecules containing the REP sequence that function as ribozymes. The snapshot at step 10,000 shows the inoculation of 10 circular ER molecules (containing a non-coding sequence) and 10 control RNA molecules. The snapshot at step 100,000 shows the spread of the circular ER genome. The snapshot at step 1,600,000 shows the appearance of circular RNA molecules containing both the ER and REP sequences (the arrow). The subsequent snapshots demonstrate the gradual spread of the circular ER-REP genome throughout the system and the simultaneous fading away of the single-gene genome containing ER.
Preprints 219604 g007
Figure 8. The events occurring in the model and their associated parameters. Legends: Np – nucleotide precursor; Nt – nucleotide; ER – exonuclease ribozyme; REP – RNA replicase; NR – nucleotide-synthetase ribozyme. Solid arrows denote chemical reactions and dashed arrows represent other events. The events related to template-directed synthesis are depicted here with respect to a template segment (within purple square brackets), which may belong to either a circular RNA or a linear RNA. Note that the region shown here represents one grid room in the N × N grid of the model system. See the text for detailed explanations.
Figure 8. The events occurring in the model and their associated parameters. Legends: Np – nucleotide precursor; Nt – nucleotide; ER – exonuclease ribozyme; REP – RNA replicase; NR – nucleotide-synthetase ribozyme. Solid arrows denote chemical reactions and dashed arrows represent other events. The events related to template-directed synthesis are depicted here with respect to a template segment (within purple square brackets), which may belong to either a circular RNA or a linear RNA. Note that the region shown here represents one grid room in the N × N grid of the model system. See the text for detailed explanations.
Preprints 219604 g008
Table 1. Parameters in the model.
Table 1. Parameters in the model.
Probabilities Descriptions Default Values
PAT An RNA template attracting a substrate (nucleotide or oligomer) 0.5
PBB A phosphodiester bond breaking within an RNA chain 1×10-6
PBER A phosphodiester bond breaking catalyzed by ER 0.9
PEL The end-to-end ligation of an RNA chain (circularization) 5×10-6
PFP The false base-pairing when an RNA template attracts a substrate 0.001
PMN The movement of nucleotides 0.005
PMNP The movement of nucleotide precursors 0.01
PND A nucleotide decaying into its precursor 0.05
PNDE A nucleotide residue decaying at RNA’s chain end 0.001
PNF A nucleotide forming from its precursor (non-enzymatic) 0.05
PNFR A nucleotide forming from its precursor catalyzed by NR 0.9
PRL The random ligation of nucleotides and RNAs 1×10-7
PRP A ribozyme being protected (from the cleaving of ER) due to its folding 0.5
PSP The separation of a base pair 0.5
PTL The template-directed ligation (non-enzymatic) 0.05
PTLR The template-directed ligation catalyzed by REP 0.9
Others Descriptions Default Values
N The system is defined as an N × N grid 30
TNPB Total nucleotide precursors introduced in the beginning 1×105
TER The times for an ER to function in a time step 10
TREP The times for an REP to function in a time step 10
TNR The times for an NR to function in a time step 10
FDA The factor concerning the de novo attraction of a substrate 5
FLT The factor for a linear RNA acting as a template 0.5
CSER The characteristic sequence of ER ACGAACUG
CSREP The characteristic sequence of REP GAGUCUCU
CSNR The characteristic sequence of NR CUGCAUCA
CSCT The characteristic sequence of the control RNA GUGAUCCA
Notes: The probabilities are listed in alphabetical order. The simulation cases in this study adopted the default parameter values unless otherwise explicitly stated. Detailed explanations of the parameters and guidelines for their configuration are provided in the Methods section.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings