Preprint
Article

This version is not peer-reviewed.

Flipons, Repeat RNAs, and IDRs Enable the Rapid 2-D Search of Large Genomes Through the Holographic Hashing of Condensate Surfaces

Submitted:

25 August 2026

Posted:

27 August 2026

You are already at the latest version

Abstract
Questions arise about how well early bacterial models of gene regulation generalize to eukaryotes and whether metazoan repetitive sequences are merely genomic spandrels. The repeats are recognized by sequence-specific RNA-binding proteins that play multiple roles in RNA biology. A subset, called flipons, adopts non-B-DNA conformations under physiological conditions without any sequence change or backbone cleavage. Other low-complexity sequences encode intrinsically disordered regions (IDRs) of proteins and promote the formation of membrane-less condensates with specific cellular functions. Here, I propose that metazoan repeats enable a fast readout of genetic information by facilitating massively parallel searches of large genomes, far exceeding that possible with diffusion alone. The locally transcribed repeats and flipon subsets seed condensates that display information about local DNA microstates on their surfaces, creating a 2-dimensional holographic hash of their contents. The hash enables transcription factors to rapidly discover critical cognate binding sites, facilitating rapid cellular responses to perturbation. The information displayed on condensate surfaces is constantly updated by both outward- and inward-directed processes. The rewrite further optimizes responses and increases the phenotypic variability on which natural selection acts. Applying the Principle of Least Action to the Lagrangian demonstrates that this strategy minimizes search times across arbitrary volumes.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  

Introduction

A fundamental question is how cells retrieve information from the genome not only in real time but also in constant time, to ensure rapid responses to perturbations. Do they use strategies equivalent to those used to optimize information retrieval from large computational databases? These constant-time strategies are based on search width or depth and differ widely in their implementation. They rely on keys that assign a unique identifier to each location that is then associated with a particular feature of the data. Various strategies are used to generate unique keys for retrieving data. The Bloom filter, for example, ignores spaces that lack a particular key [1]. The elastic search strategy generates keys that optimize the size of data clusters, reducing the search time [2]. Other searches are hierarchical and use non-overlapping key matches to progress through the tree, reducing the search space at each step [3]. Here I discuss how cells can implement constant-time search strategies enabled by surface hashing and the layering of biological condensates. I will trace the development of these ideas from a historical perspective, as our understanding has changed dramatically as we progressed from the first studies of prokaryotes to the remarkable discovery that the large genomes of eukaryotes were richly riddled with self-replicating repeat elements. Contrary to expectations, the repetitive elements and the epigenetic modifications they engender actually speed genomic search. The condensates they enable create surface hashes, providing a holographic representation of their contents. The hash enables a 2-dimensional search of the 3D cellular space, facilitating rapid responses to perturbations.

Early Days

The early perceptions of genome structure were shaped by the organisms most amenable to laboratory study. Those species that could be cheaply grown, proliferated rapidly, and were easy to mutagenize were the focus of the early studies. Bacteria and their viruses yielded valuable insights [4]. The genomes were small, enabling the mapping of DNA to proteins to pathways. Studies on the regulation of genes induced by lactose, through the lac repressor binding to its cognate operator DNA sequence, set the stage. From this work, the operon concept emerged, with regulatory elements controlling gene expression in ways responsive to environmental changes [5]. Around 92% of bacterial transcription factors (TFs) engage binding sites (TFBS) of length 8-20 bp, with less than 1% with BS of 4-7 bp [6], with a mean of 15.9 (standard deviation (SD) =6.0) [7]. These findings led to the notion of one TFBS for one TF in one genome [8], a concept that was later generalized to genetic studies of multicellular organisms.

Bacterial and Metazoan Transcription Factors

Still, even with the small bacterial genomes studied, there were surprises. For example, the search by the lac repressor for its cognate binding site was two orders of magnitude faster than a simple 3-dimensional diffusion rate would predict [9], and the residence time at the lac operator was far shorter than expected based on the measured in vitro off-rate. Facilitated diffusion models were proposed to account for this discrepancy, invoking either sliding of a TF along the DNA helix or hopping from the current low-affinity binding site to another nearby DNA surface until the cognate binding site was found [10,11]. This regulatory scheme relies on the elaboration of TF that are sequence -specific. It is also DNA-centric, as the major grooves of B-DNA can accommodate a protein α-helix capable of contacting base-specific residues, and the minor groove is shallow enough to reveal base-dependent variations in its contours. Double-stranded A-RNA cannot be sensed in either of these ways.
Applying these bacterial concepts to larger genomes of multicellular organisms was not straightforward. In contrast to bacteria, TFBS tend to be shorter, with a mean length of 9.9 bps (SD = 4.2), despite a genome that is 2-3 orders of magnitude larger [7]. Even allowing for the masking of many potential TFBS by heterochromatin, regulatory sites are still greatly outnumbered by the spurious TFBS in open chromatin. Further, a single TF is expected to bind to ~25% of regulatory regions by chance alone. Although the interactions are specific, they are non-functional [8]. A model requiring a specific combination of TFs to regulate gene expression was then invoked. The right combination of 6-8 nucleotide-binding sites produces a unique 30-base array for each promoter, or for a set of promoters that were coordinately expressed. For such a complex to bind specifically to a single locus in the human genome requires a cluster of between 10-15 TFs bound in a 500-1Kb window. This model requires high expression of each TF to saturate the multitude of other available nuclear targets that are non-functional [8].

Repeat Sequences and Selection

A further challenge to generalizing the concepts developed using bacterial genetics was the composition of metazoan genomes. Although much larger than those of bacteria, whole genome sequencing (WGS) did not reveal the expected expansion of protein-coding genes. Other explanations were needed to explain the additional phenotypic complexity. The discovery of RNA splicing and base-specific editing suggested a way to increase the protein repertoire [12,13]. Further, one splice could be conditioned on the splicing of another gene, as exemplified by the splicing cascade that determines sex in Drosophila species [14]. The splices made could also vary with epigenetic marks on histones within open chromatin regions that affect splice-site choice, or on nascent RNAs that alter their folding or processing [15]. Post-transcription, noncoding RNAs, such as miRNAs, could also vary the available proteome by controlling the time of a protein’s translation or by the assembly of protein complexes on an RNA undergoing translation [16,17].
The use of noncoding RNA to control phenotype was proposed earlier, based on in vitro studies of the reannealing time of denatured DNA in several eukaryotic species. This work demonstrated that a large percentage of multicellular genomes were derived from repeat sequences of various types [18,19]. Later, WGS of the human genome revealed that at least 50% of the genome is derived from endogenous retroelements (EREs), which can copy their RNA back into the genome at other locations. The exact percentage depends on how ERE length and sequence identity are defined [20]. EREs require a reverse transcriptase (RT) to spread. The enzyme is encoded by long interspersed nuclear elements (LINEs), long terminal repeat (LTR) elements, and endogenous retrovirus (ERV) EREs, but not by the human genome [21]. A different class of EREs hijacks the RT to transpose; short interspersed nuclear elements (SINEs) are an example. SINES contain no protein-coding sequence. Instead, they abstract the control sequences necessary to drive their forward and reverse transcription [22]. In this process, they also incorporate adjacent sequences that contain TFBS motifs [23]. Given the frequency of SINEs, with over 1 million copies in the human genome, they initially appeared to have little informational value. They were termed either junk DNA, which accumulated over time, or selfish DNA, which existed only to replicate itself [24,25,26]. Such elements were selected against in bacterial genomes where rapid proliferation furthers survival, with advantageous mutations rapidly increasing in population frequency.

Repeats and the Proteome

Surprisingly, the evidence for positive selection of repeat elements in metazoans has accumulated [27,28,29,30]. Some repeats were required to protect chromosomal ends and scaffold attachment sites, ensuring proper chromosomal segregation during division [31,32]. Others produced proteins essential to the evolution of placentas [33]. A further set provided promoters and enhancers to enable tissue-specific gene expression [34]. Many were expressed during early development and helped scaffold the nucleus and position nucleosomes, thereby separating active chromosomal regions from heterochromatin. This arrangement left segments of DNA that remained open and available to bind sequence-specific transcription factors. Their transcripts can fold into stable tetraloops and pseudoknots, enabling sequence- and structure- specific recognition by proteins, just like the RNAs from protein-coding genes [35].
Computational analyses of the genome revealed that some repeat sequences were distributed non-randomly through protein-coding regions, with CG nucleotides enriched in promoters and 5ˊ-untranslated regions (UTRs) [36] to TA-rich sequences in 3ˊ-UTRs, and polypyrimidine tracts near 3ˊ splice sites [37,38]. Many of the transcripts from these elements are bound by a diverse set of sequence-specific proteins called heteronuclear ribonucleoproteins (hnRNPs) [39] that are found in nuclear bodies or granules that represent regions of liquid-liquid phase separation (often referred to as condensates) (reviewed in [40]). These proteins are abundant in the cell, with many at concentrations up to 100 million copies per cell, and a subset is expressed in a cell-specific manner [39,41]. Other work reported the genome-wide role of noncoding RNA clusters in organizing chromatin territories [42,43]. The condensates they form with hnRNP and other RNA-binding proteins (RBPs) represent one way to index a large genome. Their surface features capture information about their content, enabling a TF to search for and locate a TFBS faster than by diffusion.

Repeats and Flipons

The idea that repeats also seed condensates by forming alternative nucleic acid structures is only recent [44,45]. The insights are not from the recent discovery of these structures. Historically, as the Watson-Crick model of B-DNA was gaining widespread acceptance, a variety of other DNA conformations were experimentally revealed, including three-stranded triplexes and four-stranded quadruplexes (Figure 1A) [46,47]. Then R-Loops were reported, in which an RNA displaces a DNA strand with the same sequence from a B-DNA helix [48]. Later, left-handed Z-DNA turned up in the first synthetic DNA crystal ever solved (Figure 1B) [49] and then a four-stranded structure in which cytosine base pairs intercalate, called an i-motif, was described [50]. Each of these structures has a canonical repeat motif that favors its formation. Z-DNA is associated with alternating pyrimidine/purine repeats, with guanine-based repeats most favorable. GQ are formed from G-tetrads that stack in a number of different ways to form a quadruplex [51,52,53], while triplexes are formed by the docking of a third strand to a homopurine repeat, and i-motifs require a long cytosine strand. Like the sequence repeats recognized by hnRNP, the distribution of these motifs was not random. For example, potential Z-DNA and GQ motifs were found in promoters and 5ˊ-UTRs [54,55,56], with around 300,000 of each mapped throughout the genome [57,58]. At the time of discovery, none of these alternative conformations had an obvious biological role, and many were formed at the bench only under non-physiological conditions. Rather, these repeats were associated with genomic instability. G-quadruplexes (GQ) and triplexes (TRX) initially attracted attention for their proposed roles in Mendelian and repeat-expansion diseases, particularly those affecting the neuromuscular system [59,60,61,62]. They were of interest in their own right, but lacked any experimental evidence of a cellular function.

Biological Functions of Flipons

Over time, conditions were identified that favor the transition of these repeat elements, called flipons, from B-DNA to an alternative structure. For Z-DNA, unwinding of the B-DNA helix produced by negative supercoiling of plasmids produced sufficient strain to force the flip [65,66]. The energy required to initiate Z-DNA formation and to propagate the flip to adjacent segments was then determined using closed circular plasmids with different flipon inserts [67]. This work revealed that the energetic cost of forming Z-DNA depended on base composition and motif length. Base modifications, such as 5-methyl cytosine and 8-oxo-guanosine, were also found to promote Z-DNA formation [68,69].
Over the past 5 years, numerous papers have been published linking Z-DNA and Z-RNA to various infectious, inflammatory, neurological, and malignant disease outcomes. These events involve a highly specific Z-DNA and Z-RNA (collectively called ZNA) binding domain, called Zα, present in only two proteins in the human genome, ADAR and ZBP1 [70,71,72,73]. Their role, with ADAR negatively regulating interferon responses and ZBP1inducing various inflammatory and cell death pathways, is now well validated and the subject of numerous recent papers [74,75,76,77,78,79,80]. Evidence is also accumulating that other proteins can bind ZNA, TRX, and GQ, playing many different roles in the cell. These flipon conformations are often bound by proteins known to bind B-DNA in a sequence-specific manner, usually through the same binding domain [81,82,83].
Protein recognition of alternative flipon structures is widespread in eukaryotes. One example is provided by the two crystal structures of the yeast RAP1 bound either to its cognate binding site in B-DNA or to a GQ [84,85]. The interactions involve different faces of the same helix. Similarly, the ability of the large zinc-finger transcription family to dock to alternative flipon conformations is now recognized [83]. The binding affinity (KD) for Z-DNA and G-quadruplexes can be nanomolar, but the on- and off-rates can vary, allowing both stable docking modes when the off-rate is low, and rapid scanning when the off-rate is fast. For example, both the Zα domain that binds ZNA and the helicase domain of SMARCA4 that binds a wide range of GQs are extremely slow [70,86].

Flipons and Genetic Encoding

Flipons within the repeat genome represent an alternative way of encoding genetic information, affecting our understanding of how cells work and how organisms evolve. As B-DNA, flipons have little informational value due to the high frequency of repeat motifs. However, this is not true for all flipons. A small flipon subset marks active regions of the genome where sufficient free energy is available to power the transition. They are positioned in promoters and other regulatory segments. Depending on their structure, flipons can interact with different sets of proteins, changing the readout of genetic information and altering cellular programming. Flipons can also act mechanically, operating as actuators. In this role, they act as topological switches that modulate transcription, affecting both the elongation and reinitiation phases of the cycle [87]. In such situations, Z-DNA is generated by the negative supercoiling generated by an RNA polymerase as it plows through the helix [88]. Recent work suggests that the general transcription factor IIE subunit 1 (encoded by GTF2E1) can bind to Z-DNA and reinitiate the reassembly of pre-initiation complexes [89]. ZNAs also modulate innate immunity and cell death pathways, with the single Zα domain on ADAR1 negatively regulating interferon responses and the two Zα domains on ZBP1 initiating cell death pathways. GQs also facilitate a number of outcomes. They can localize TFs to promoters and splice sites, and also regulate translation by engaging initiation factors [51,59,90,91,92,93]. Triplex formation can occur locally in DNA direct repeats that fold back to form the third strand, or with non-coding transcripts, either in cis or trans, and with small RNAs [94,95,96,97,98,99,100]

Flipon-Enabled Search of Repeat-Rich Genomes

Targeting of the non-B-DNA flipon conformations reduces the search space for TF to find a cognate binding site. They help maintain an open chromatin region because histones optimally wrap around right-handed DNA conformers and poorly around alternative DNA structures. GQs in particular provide a docking site from which TF can initiate scanning of the local region to find a cognate TFBS, while TRX provides a different way to nucleate condensates associated with DNA damage. The transcripts from active genomic regions help scaffold condensates by engaging sequence-specific hnRNP. These condensates capture the composition gradient between GC-rich promoters, ERE-rich enhancers, polypyrimidine-rich spliceosome sites, and AT-rich transcript termination sites. The merging of condensates formed at these different sites modulates RNA production, processing, and quality control. These interactions are transient, forming and dissolving within seconds to minutes rather than hours or days [101,102,103]. Measurements performed using a variety of experimental approaches across a range of organisms confirm their dynamic nature and their regulation by suppressive factors [104]. Flipons thereby provide another way of indexing the genome by seeding condensate formation.

Condensate-Enabled Search of Repeat-Rich Genomes

Overall, these diverse findings help resolve the dilemma posed by genomes that grow larger through the accrual of noncoding and repetitive sequences. How do TFs with short recognition motifs rapidly localize to functional sites? Indexing the genome by condensates is one solution to this problem. The flipon state allows the identification of active sites in the repeat genome, setting the stage for compaction of inactive regions into heterochromatin and their exclusion from the search space. The sites where alternative nucleic acid structures form help nucleate other condensate forms, which are assembled using region- and sequence-specific RBP scaffolds. The condensates are self-assembling, automatically creating a distinctive surface. Protein inclusion is based on domain-specific docking of functional complexes. These can be assembled cotranslationally, with proteins made at different times and locations in the cell [17]. Other interactions depend on IDRs, in which charge properties rather than the exact sequence determine whether a protein is included or excluded. The interactions are transient with fast on- and off-rates. Each type of interaction can vary through context-specific protein modifications such as phosphorylation, ubiquitylation, cysteine, or acylation. The adducts produced reflect cell state. They allow condensates to grow, shrink, or maintain equilibrium by reengaging components as fast as they dissociate. At equilibrium, where the chemical potentials are equal, the concentration of a component may be much lower than in the condensate, maintained by interactions with other constituents [105].

Condensate Surfaces and Optimized Search Strategies

Due to the process of self-assembly, the surfaces of the condensates index their contents and provide a simple search mechanism akin to the lateral diffusion models proposed for the LacI repressor in bacteria [11,106]. A simple intuition is shown in Figure 2A, where a cubic lattice is formed from well-ordered colored spheres. This close packing enables a massively parallelized Cartesian-grid search of all sphere surfaces to locate a subset of properties associated with a particular sphere color. Searching for a feature on the surface of a sphere is faster than searching the entire volume of the cube. The speed is increased by a factor of 2r/π (the volume of a cube with sides of length 2r divided by the surface area of a sphere of radius r). In this example, color is used as a key. Increasing the number of keys on the surface generates a hash that uniquely identifies each sphere, with the embedded information quantified by the Shannon entropy. Thereby, aggregates present low-complexity surfaces composed of single proteins. The efficiency of search can be increased by increasing the surface area of a sphere relative to other spheres, with the key density δ given by:
δ =   i n p i log p i 4 π r 2
where pi represents the probability of finding the ith key on the surface that contains n keys. The formula can be easily adjusted for irregular surfaces. The parameter δ is tunable by modifying the surface informational molecules and by varying the condensate size, thereby altering the hash density and the percent occupancy of the lattice. Surfaces with only a single copy of a key require a search of the entire surface and will be confounded by the noise arising from the protein’s association with unrelated condensates. When the search time is limited, the use of such low-density keys will be disfavored, setting a signal-to-noise threshold.
The number of spheres and their size define the unoccupied (free) space in the cube. In Figure 2A, ~48% (1 - π/6) of the cube's volume is outside the spheres, allowing free diffusion of small molecules that can fit into the gaps, such as ions or nucleotides. These agents can act on much larger proteins that are fixed in place to modify their state and, subsequently, the surfaces and cellular scaffolds they contact. These outward-directed changes dynamically update the surface hash, providing new search keys for this subset of spheres. The search can be further enhanced by spatially addressing each sphere to a particular cube location, much as zip codes direct RNA delivery [107] and signal peptides target proteins [108]. The key search can then be restricted to these locations.

An example of Search Optimization Based on the Principle of Least Action

The search strategy can be optimized using the principle of least action first proposed by Maupertuis in 1744. The principle describes how a physical system naturally chooses a path between two points that minimizes a specific quantity. The optimal path for a particular physical system can be determined using a widely used approach described by Lagrange in 1788. Figure 3A shows a Lagrangian (boxed in blue) defined by two parameters, α and β, for calculating the minimum time required to search an arbitrary 3D volume using surface-bounded operations. For each combination of α and β, the number of spheres and their radius are varied to capture different outcomes. The plots in Figure 3B-D show how the cost surface behaves for a surface constraint value K = 500, with the red line indicating the feasible path and a red arrow pointing to the minimum reachable search time. These examples reveal that the optimization landscape is highly sensitive to hardware constraints, with the minimum search time varying by approximately 20-fold (1.59 s versus 31.71s) depending on the specific combination of α and β.
In this system, α represents the time required to read a unit of surface data. If the time spent processing the surfaces, α, is high, the system reduces dwell time by shrinking the sphere size and increasing the number of readers. β corresponds to the time between surfaces, representing the non-productive time spent traversing empty space. The system optimizes by using fewer large spheres, increasing efficiency by the search of larger continuous surfaces. Counterintuitively, large volumes with small C and small volumes with large C have the same K and the same α and β optimizations. Here, lower C corresponds to a more efficient packing scheme. Lowering the volume does not increase search efficiency, because many redundant checks are performed on each surface.

Condensate Surfaces and Adaptations

The use of surfaces for the 2D hashing of their contents allows cells to optimize their response times by varying both the number and size of condensates in different compartments, such as the cytoplasm and nucleus. The different spaces and chemistry enable separate optimizations for each. Other variations can also affect operational efficiency. The packing density of the condensates can affect how quickly a particular condensate turns over. The available space, as measured by the excluded volume, plays a critical role. In contrast to globular states, which have a radius of gyration (rg) that scales as L1/3 (where L is their length), fully unfolded polymers have a rg that scales as L3/5, reflecting their increased entropy [109]. Figure 2B illustrates the reduced free volume when spheres are replaced with their unfolded constituents. The physical barriers created by unfolded proteins negatively impact search efficiency. However, disordered proteins will diffuse more easily through crowded cellular spaces than rigid globular spheres [110]. Unfolding will also expose previously buried charged residues, altering the protein interactions with condensate surfaces [111]. Both the folding kinetics and the diffusion place limits on the rate of dynamic rehashing.
Studies reveal that in cells, the excluded volume is much larger than for close-packed spheres, with estimates centered within the range of approximately 20-30% [112]. The lower concentration of proteins in solution than in the condensates allows diffusion to proceed without aggregation [105]. As with many other processes, the concentration is likely close to criticality, with small changes in the cell state triggering condensate formation, regulating their charge and incorporation into condensates [111]. With larger structures, the cytoskeleton provides guidance and assists delivery. Here, an ice-breaker model may operate, in which components at the bow divert, dissolve, or derivatize condensates and other obstacles ahead of the complex. In this example, changes to the condensate and its surface hash are inward-directed.
Events that happen below the surface can also rewrite the condensate surface. The rehash dynamically captures and provides updates of internal events. For example, a condensate with acetylated surface proteins conveys different information than one with methylated surface proteins. Condensate scaffolded by the same proteins may produce different outcomes due to distinct modifications: one may promote gene expression, while another suppresses neighboring loci. The updates may help offset changes in cell state. Subsequently, a change in hashing may reverse the current pattern, with previously repressed genes now expressed, and vice versa. The effects of knocking out a core component may then vary with context, either enhancing or suppressing a particular output, thereby producing contradictory experimental results.

Condensate Layering to Enhance Functionality

In biological systems, the simple search design implemented using condensates is further modified. The layering of condensates creates another level of regulation (Figure 1C). The layers can incorporate filters to control entry into an inner space of the condensate and control output. Nucleoli are one example, enabling the stepwise maturation of ribosomal RNAs before assembly into ribosomes, with the inner layer established by nucleolin interactions with G-quadruplexes (Figure 1A) [63,113]. Another example is paraspeckles, where the layers are based on the block copolymer structure of nuclear paraspeckle assembly transcript 1 (NEAT1) [114]. This arrangement allows the sequential, ordered assembly of protein components as the NEAT1 gene is transcribed [115]. Variations in the RNA components used to assemble components also enable optimized search over time
Once formed, access to the various spaces in a condensate depends on the mesh size and the component size. In biological systems, exclusion of inert particles larger than 4 nm is reported [116], the size of a globular 30-kd protein. However, partitioning of much larger proteins into condensates, such as antibodies and ribosomes, is observed when charge interactions overcome the entropic barriers involved [117]. Polymeric modifications based on ubiquitin or ubiquitin-like proteins, such as SUMO [118], PARylation, and lipidation, which modify these interactions, can tune exit and entry [119]. Protein modifications that affect charge, such as phosphorylation or acylation, can affect the location of a protein within a condensate. The layering may restrict physical access to internal spaces, further filtering what reaches the core. The layers also filter what emerges from those spaces. Such filters play an essential role in preventing the release of defectively processed RNAs. In the simplest of such schemes, epigenetic marks are added cotranscriptionally to introns. If these marks are not removed by splicing, the RNA is degraded rather than released into an outer layer. The persistence of such marks then triggers transcript triage. Layering has been modeled using DNA structures called nanostars that vary in the number of arms and in the sequence specificity of their single-stranded ends. Pairing of these ends with other nanostars creates three-dimensional forms that vary with the number and geometry of each nanostar. Nanostars can be assembled into layers, with the outer layer controlling the assembly's stability and its propensity to merge with other complexes [120]. Each layer can be designed to optimize particular chemistries.

Inward-Directed Rehashing

In a cell, the hash on the outer layer of a condensate can be modified from within. As a result, the hashing dynamically changes, altering where components are delivered long after they were first made. The new outer hash can modulate merges with other condensates. For example, merges between stress granules and processing bodies (P-bodies) enable the transfer of materials between them and the efficient disposal of unneeded cytoplasmic RNAs. In other cases, the merging of promoter and spliceosome condensates may depend on rehashing and occur only after each has assembled correctly. Otherwise, the components of each condensate disperse, producing premature termination of transcription. A related approach to regulating these processes involves modifying cellular proteins that can initially bridge distinct types of condensates. The cargo adaptor proteins involved in condensate protein-quality control provide an example [121]
Rapid hash updates also occur within the nucleus. This process potentially enables the rapid search of large genomes by TFs to find their critical TFBS. In the simplest model, condensate seeding relies initially on random transcription initiation. Only at some loci is the negative supercoiling sufficient to transition flipons to an alternative state and maintain a region of open chromatin structure. The exact structure formed depends on the sequence. Similarly, the RNA transcripts produced will localize hnRNP to the repeat motifs they recognize. Together, the flipons and hnRNP will generate a surface hash of the underlying repeat structure. The flipons will recruit structure-specific proteins, updating the hash of that region. Similarly, the hnRNP will recruit a different set of cellular proteins through domain- or IDR-mediated interaction, leading to a further surface rehash. In contrast, a subset of repeats will engage small, sequence-specific RNA, such as piRNAs, that mask these elements by promoting their packaging into heterochromatin. This form of negative hashing masks EREs and other repeat elements, thereby suppressing their retrotransposition to other genomic loci. The multi-layered formation and re-hashing of condensates reduce the search space for TFs to find their cognate binding sites. The design enables a parallelized Cartesian grid search to facilitate fast responses.

Condensate Contents Are Also Dynamically Updated

The dynamic re-hashing of condensates is promoted by the dissolution of the transcriptional condensates. This process allows the components to be updated as the RNA mesh is dismantled, thereby reflecting the cell's current state. The adoption of secondary structures, such as GQ, TRX, and Z-RNAs within a condensate may normally be restricted by helicases and the tethering of protein components to actin filaments and microtubules. The interactions of condensates with the cytoskeleton likely result in tensegrity arising from the actin-generated strain and the compressive resistance of tubulin structures (Figure 1D). These forces tension the mesh, preventing pore collapse and the formation of dsRNA from hybridization of repeat elements [122,123]. Upon dissolution and release of tension, the now-relaxed single-stranded RNA can form GQ and Z-RNA. These structures may facilitate the retention of TFs that recognize these alternative conformations in addition to their cognate binding sites. Indeed, one surprise from the ENCODE project was the identification of hot loci that engage a large number of TF, but lack their TFBSs [124].
It also allows competition for the available binding sites. Over time, the formation of higher-affinity complexes will result from an increase in the local concentration of tight binders as the cycle repeats. Indeed, it is possible that flipons can help assemble dimeric TFs that also undergo a cycle of association and dissociation. In an example of ZNAs, the rapid dissociation of helix-loop-helix and leucine zippers enables partner exchange. Transient docking with ZNAs may bring new partners into close proximity, leading to the assembly of novel combinations. Dimers that bind to and are stabilized by a nearby TFBS will increase their local concentration. In these scenarios, condensates index target sites to promote localization of high-affinity TFs. This strategy does not require that each TF have a unique motif in the genome[125].

Conclusions and Future Studies

Scientific advances lead to new concepts beyond what was imagined or even conceived as possible by earlier generations. After describing differences in genomic regulation between prokaryotes and metazoans, we review recent data on the roles repetitive elements play in metazoan genomes to lay the foundation for the rest of the paper. The emergence of distinct forms of genetic encoding by flipons provides one example of how our understanding has changed, and how genome readout is regulated. Meanwhile, new tools have sparked renewed interest in the functions of the condensates observed many decades ago [126]. When combined, the new and old findings suggest that condensates empower the rapid searches of large genomes, enabling fast responses. Their surface features provide a hash of their contents, enabling rapid localization of effectors to their targets. Here we apply the Principle of Least Action to the Lagrangian to show how this strategy minimizes total execution by optimizing search parameters.
The programmability of these surfaces by both inner- and outward-directed processes allows further enhancements over time. It also generates phenotypic diversity on which natural selection can act. The condensate density is a critical factor in this system's adaptability and may lead to dysfunction if the excluded volume becomes too high. The hash features that are unlocked by system keys also require regulation. If levels are too high, they are prone to aggregation; if too low, the search takes longer to locate the feature on a surface. The use of condensates by metazoans overcomes the three-dimensional diffusion limits in a different way from bacteria, thereby explaining the importance of the repeat genome in these distinct evolutionary strategies. Their holographic representation of cellular contents enables rapid two-dimensional search strategies and rapid responses. Future studies aimed at understanding and modulating the surface hash of condensates will lead to therapeutic strategies to control condensate formation, size, and fate. Progress in this direction has already been made [120,127,128].

Funding

This work was not externally funded.

Conflicts of Interest

The author is the founder of InsideOutBio and declares no competing interests.

References

  1. Bloom, B. H. 1970 Space/time trade-offs in hash coding with allowable errors. Commun. ACM 13, 422–426. [CrossRef]
  2. Farach-Colton, M.; Krapivin, A.; Kuszmaul, W. Optimal Bounds for Open Addressing Without Reordering. 2024 IEEE 65th Annu. Symp. Found. Comput. Sci. (FOCS) 2024, 594–605. [Google Scholar] [CrossRef]
  3. Fredman, M. L.; Komlós, J.; Szemerédi, E. 1984 Storing a Sparse Table with 0(1) Worst Case Access Time. J. ACM 31, 538–544. [CrossRef]
  4. Luria, S. E.; Delbrück, M. 1943 Mutations of Bacteria from Virus Sensitivity to Virus Resistance. Genetics 28, 491–511. [CrossRef] [PubMed]
  5. Jacob, F.; Monod, J. 1961 Genetic regulatory mechanisms in the synthesis of proteins. J. Mol. Biol. 3, 318–356. [CrossRef] [PubMed]
  6. Borges Farias, A.; Sganzerla Martinez, G.; Galan-Vasquez, E.; Nicolas, M. F.; Perez-Rueda, E. Predicting bacterial transcription factor binding sites through machine learning and structural characterization based on DNA duplex stability. Brief. Bioinform. 2024, 25. [Google Scholar] [CrossRef] [PubMed]
  7. Stewart, A. J.; Hannenhalli, S.; Plotkin, J. B. 2012 Why transcription factor binding sites are ten nucleotides long. Genetics 192, 973–985. [CrossRef] [PubMed]
  8. Wunderlich, Z.; Mirny, L. A. 2009 Different gene regulation strategies revealed by analysis of binding motifs. Trends Genet. 25, 434–440. [CrossRef] [PubMed]
  9. Riggs, A. D.; Bourgeois, S.; Cohn, M. The lac repressor-operator interaction. 3. Kinetic studies. J. Mol. Biol. 1970, 53, 401–417. [Google Scholar] [CrossRef] [PubMed]
  10. Winter, R. B.; Berg, O. G.; von Hippel, P. H. Diffusion-driven mechanisms of protein translocation on nucleic acids. 3. The Escherichia coli lac repressor--operator interaction: kinetic measurements and conclusions. Biochemistry 1981, 20, 6961–6977. [Google Scholar] [CrossRef] [PubMed]
  11. Hammar, P.; Leroy, P.; Mahmutovic, A.; Marklund, E. G.; Berg, O. G.; Elf, J. 2012 The lac repressor displays facilitated diffusion in living cells. Science 336, 1595–1598. [CrossRef] [PubMed]
  12. Berget, S. M.; Moore, C.; Sharp, P. A. 1977 Spliced segments at the 5' terminus of adenovirus 2 late mRNA. Proc. Natl. Acad. Sci. U S A 74, 3171–3175. [CrossRef] [PubMed]
  13. Chow, L. T.; Gelinas, R. E.; Broker, T. R.; Roberts, R. J. An amazing sequence arrangement at the 5' ends of adenovirus 2 messenger RNA. Cell 1977, 12, 1–8. [Google Scholar] [CrossRef] [PubMed]
  14. Baker, B. S.; Nagoshi, R. N.; Burtis, K. C. Molecular genetic aspects of sex determination in Drosophila. Bioessays 1987, 6, 66–70. [Google Scholar] [CrossRef] [PubMed]
  15. Luco, R. F.; Allo, M.; Schor, I. E.; Kornblihtt, A. R.; Misteli, T. 2011 Epigenetics in alternative pre-mRNA splicing. Cell 144, 16–26. [CrossRef] [PubMed]
  16. Horvitz, H. R. 2003 Worms, Life, and Death DOI:Nobel Lecture. ChemBioChem 4, 697–711. [CrossRef] [PubMed]
  17. Mayr, C. 2017 Regulation by 3'-Untranslated Regions. Annu Rev. Genet. 51, 171–194. [CrossRef] [PubMed]
  18. Britten, R. J.; Davidson, E. H. 1969 Gene Regulation for Higher Cells: A Theory. Science 165, 349–357. [CrossRef] [PubMed]
  19. Schmid, C. W.; Deininger, P. L. 1975 Sequence organization of the human genome. Cell 6, 345–358. [CrossRef] [PubMed]
  20. Hoyt, S. J.; Storer, J. M.; Hartley, G. A.; Grady, P. G. S.; Gershman, A.; de Lima, L. G.; Limouse, C.; Halabian, R.; Wojenski, L.; Rodriguez, M.; et al. 2022 From telomere to telomere: The transcriptional and epigenetic state of human repeat elements. Science 376, eabk3112. [CrossRef] [PubMed]
  21. Xiong, Y.; Eickbush, T. H. 1990 Origin and evolution of retroelements based upon their reverse transcriptase sequences. EMBO J. 9, 3353–3362. [CrossRef] [PubMed]
  22. Batzer, M. A.; Deininger, P. L. 2002 Alu repeats and human genomic diversity. Nat. Rev. Genet. 3, 370–379. [CrossRef] [PubMed]
  23. Hermant, C.; Torres-Padilla, M. E. 2021 TFs for TEs: the transcription factor repertoire of mammalian transposable elements. Genes Dev. 35, 22–39. [CrossRef] [PubMed]
  24. Ohno, S. Evolution by Gene Duplication, 1 ed.; Springer-Verlag Berlin: Heidelberg, 1970. [Google Scholar]
  25. Brenner, S. 1998 Refuge of spandrels. Curr. Biol. 8. [CrossRef] [PubMed]
  26. Dawkins, R. 2006 The selfish gene, 30th anniversary ed.; Oxford University Press: Oxford; New York.
  27. Kamal, M.; Xie, X.; Lander, E. S. 2006 A large family of ancient repeat elements in the human genome is under strong selection. Proc. Natl. Acad. Sci. U S A 103, 2740–2745. [CrossRef] [PubMed]
  28. Gemayel, R.; Vinces, M. D.; Legendre, M.; Verstrepen, K. J. 2010 Variable tandem repeats accelerate evolution of coding and regulatory sequences. Annu Rev. Genet. 44, 445–477. [CrossRef] [PubMed]
  29. Kapusta, A.; Kronenberg, Z.; Lynch, V. J.; Zhuo, X.; Ramsay, L.; Bourque, G.; Yandell, M.; Feschotte, C. 2013 Transposable elements are major contributors to the origin, diversification, and regulation of vertebrate long noncoding RNAs. PLoS Genet. 9, e1003470. [CrossRef] [PubMed]
  30. Cagliani, R.; Forni, D.; Mozzi, A.; Sarama, R.; Pozzoli, U.; Fumagalli, M.; Sironi, M. 2026 Positive Selection Targeted Primate Genes that Encode Transposable Element Repressors. Genome Biol. Evol. 18. [CrossRef] [PubMed]
  31. Blackburn, E. H.; Gall, J. G. A tandemly repeated sequence at the termini of the extrachromosomal ribosomal RNA genes in Tetrahymena. J. Mol. Biol. 1978, 120, 33–53. [Google Scholar] [CrossRef] [PubMed]
  32. Blackburn, E. H.; Greider, C. W.; Szostak, J. W. 2006 Telomeres and telomerase: the path from maize, Tetrahymena and yeast to human cancer and aging. Nat. Med. 12, 1133–1138. [CrossRef] [PubMed]
  33. Kokosar, J.; Kordis, D. 2013 Genesis and regulatory wiring of retroelement-derived domesticated genes: a phylogenomic perspective. Mol. Biol. Evol. 30, 1015–1031. [CrossRef] [PubMed]
  34. Peaston, A. E.; Evsikov, A. V.; Graber, J. H.; de Vries, W. N.; Holbrook, A. E.; Solter, D.; Knowles, B. B. 2004 Retrotransposons regulate host genes in mouse oocytes and preimplantation embryos. Dev. Cell 7, 597–606. [CrossRef] [PubMed]
  35. Shen, L. X.; Tinoco, I., Jr. The structure of an RNA pseudoknot that causes efficient frameshifting in mouse mammary tumor virus. J. Mol. Biol. 1995, 247, 963–978. [Google Scholar] [CrossRef] [PubMed]
  36. Antequera, F.; Bird, A. 2018 CpG Islands: A Historical Perspective. Methods Mol. Biol. 1766, 3–13. [Google Scholar] [CrossRef] [PubMed]
  37. Ruskin, B.; Green, M. R. 1985 Role of the 3' splice site consensus sequence in mammalian pre-mRNA splicing. Nature 317, 732–734. [CrossRef] [PubMed]
  38. Frendewey, D.; Keller, W. Stepwise assembly of a pre-mRNA splicing complex requires U-snRNPs and specific intron sequences. Cell 1985, 42, 355–367. [Google Scholar] [CrossRef] [PubMed]
  39. Dreyfuss, G.; Kim, V. N.; Kataoka, N. 2002 Messenger-RNA-binding proteins and the messages they carry. Nat. Rev. Mol. Cell Biol. 3, 195–205. [CrossRef] [PubMed]
  40. Belmont, A. S. 2022 Nuclear Compartments: An Incomplete Primer to Nuclear Compartments, Bodies, and Genome Organization Relative to Nuclear Architecture. Cold Spring Harb. Perspect. Biol. 14. [CrossRef] [PubMed]
  41. Kamma, H.; Portman, D. S.; Dreyfuss, G. 1995 Cell type-specific expression of hnRNP proteins. Exp. Cell Res. 221, 187–196. [CrossRef] [PubMed]
  42. Quinodoz, S. A.; Jachowicz, J. W.; Bhat, P.; Ollikainen, N.; Banerjee, A. K.; Goronzy, I. N.; Blanco, M. R.; Chovanec, P.; Chow, A.; Markaki, Y.; et al. 2021 RNA promotes the formation of spatial compartments in the nucleus. Cell 184, 5775–5790 e5730. [CrossRef] [PubMed]
  43. Asamitsu, S.; Iwasaki, Y. W. 2025 Biomolecular liquid‒liquid phase separation associated with repetitive genomic elements. Polym. Journal. 57, 785–797. [CrossRef]
  44. Herbert, A. 2019 A Genetic Instruction Code Based on DNA Conformation. Trends Genet. 35, 887–890. [CrossRef] [PubMed]
  45. Zhang, Y.; Yang, M.; Duncan, S.; Yang, X.; Abdelhamid, M. A. S.; Huang, L.; Zhang, H.; Benfey, P. N.; Waller, Z. A. E.; Ding, Y. 2019 G-quadruplex structures trigger RNA phase separation. Nucleic Acids Res. 47, 11746–11754. [CrossRef] [PubMed]
  46. Felsenfeld, G.; Davies, D. R.; Rich, A. 1957 Formation of a Three-Stranded Polynucleotide Molecule. J. Am. Chem. Soc. 79, 2023–2024. [CrossRef]
  47. Gellert, M.; Lipsett, M. N.; Davies, D. R. 1962 Helix Formation by Guanylic Acid. Proc. Natl. Acad. Sci. 48, 2013–2018. [CrossRef] [PubMed]
  48. Thomas, M.; White, R. L.; Davis, R. W. 1976 Hybridization of RNA to double-stranded DNA: formation of R-loops. Proc. Natl. Acad. Sci. U S A 73, 2294–2298. [CrossRef] [PubMed]
  49. Wang, A. H.; Quigley, G. J.; Kolpak, F. J.; Crawford, J. L.; van Boom, J. H.; van der Marel, G.; Rich, A. 1979 Molecular structure of a left-handed double helical DNA fragment at atomic resolution. Nature 282, 680–686. [CrossRef] [PubMed]
  50. Gehring, K.; Leroy, J. L.; Gueron, M. A tetrameric DNA structure with protonated cytosine.cytosine base pairs. Nature 1993, 363, 561–565. [Google Scholar] [CrossRef] [PubMed]
  51. Lightfoot, H. L.; Hagen, T.; Tatum, N. J.; Hall, J. 2019 The diverse structural landscape of quadruplexes. FEBS Lett. 593, 2083–2102. [CrossRef] [PubMed]
  52. Sundaresan, S.; Uttamrao, P. P.; Kovuri, P.; Rathinavelan, T. 2024 Entangled World of DNA Quadruplex Folds. ACS Omega 9, 38696–38709. [CrossRef] [PubMed]
  53. Granzhan, A.; Mouawad, L. 2026 Topological rules and anomalies in intramolecular G-quadruplex folding: a comprehensive study. Nucleic Acids Res. 54. [CrossRef] [PubMed]
  54. Schroth, G. P.; Chou, P. J.; Ho, P. S. 1992 Mapping Z-DNA in the human genome. Computer-aided mapping reveals a nonrandom distribution of potential Z-DNA-forming sequences in human genes. J. Biol. Chem. 267, 11846–11855. [CrossRef]
  55. Huppert, J. L.; Balasubramanian, S. 2005 Prevalence of quadruplexes in the human genome. Nucleic Acids Res. 33, 2908–2916. [CrossRef] [PubMed]
  56. Cer, R. Z.; Bruce, K. H.; Mudunuri, U. S.; Yi, M.; Volfovsky, N.; Luke, B. T.; Bacolla, A.; Collins, J. R.; Stephens, R. M. 2011 Non-B DB: a database of predicted non-B DNA-forming motifs in mammalian genomes. Nucleic Acids Res. 39, D383–391. [CrossRef] [PubMed]
  57. Umerenkov, D.; Herbert, A.; Konovalov, D.; Danilova, A.; Beknazarov, N.; Kokh, V.; Fedorov, A.; Poptsova, M. 2023 Z-flipon variants reveal the many roles of Z-DNA and Z-RNA in health and disease. Life Sci. Alliance 6. [CrossRef] [PubMed]
  58. Konovalov, D.; Umerenkov, D.; Herbert, A.; Poptsova, M. 2025 GQ-DNABERT reveals GQ proximal enhancer-promoter interactions associated with tissue-specific transcription. Nucleic Acids Res. 53. [CrossRef] [PubMed]
  59. Maizels, N. 2015 G4-associated human diseases. EMBO Rep. 16, 910–922. [CrossRef] [PubMed]
  60. Sauer, M.; Paeschke, K. 2017 G-quadruplex unwinding helicases and their function in vivo. Biochem Soc. Trans. 45, 1173–1182. [CrossRef] [PubMed]
  61. Khristich, A. N.; Mirkin, S. M. On the wrong DNA track: Molecular mechanisms of repeat-mediated genome instability. J. Biol. Chem. 2020, 295, 4134–4170. [Google Scholar] [CrossRef] [PubMed]
  62. Wang, E.; Thombre, R.; Shah, Y.; Latanich, R.; Wang, J. 2021 G-Quadruplexes as pathogenic drivers in neurodegenerative disorders. Nucleic Acids Res. 49, 4816–4830. [CrossRef] [PubMed]
  63. Chen, L.; Dickerhoff, J.; Zheng, K. W.; Erramilli, S.; Feng, H.; Wu, G.; Onel, B.; Chen, Y.; Wang, K. B.; Carver, M.; et al. 2025 Structural basis for nucleolin recognition of MYC promoter G-quadruplex. Science 388, eadr1752. [CrossRef] [PubMed]
  64. Abramson, J.; Adler, J.; Dunger, J.; Evans, R.; Green, T.; Pritzel, A.; Ronneberger, O.; Willmore, L.; Ballard, A. J.; Bambrick, J.; et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 2024, 630, 493–500. [Google Scholar] [CrossRef] [PubMed]
  65. Singleton, C. K.; Klysik, J.; Stirdivant, S. M.; Wells, R. D. 1982 Left-handed Z-DNA is induced by supercoiling in physiological ionic conditions. Nature 299, 312–316. [CrossRef] [PubMed]
  66. Peck, L. J.; Wang, J. C. 1983 Energetics of B-to-Z transition in DNA. Proc. Natl. Acad. Sci. U S A 80, 6206–6210. [CrossRef] [PubMed]
  67. Ho, P. S.; Ellison, M. J.; Quigley, G. J.; Rich, A. A computer aided thermodynamic approach for predicting the formation of Z-DNA in naturally occurring sequences. Embo J. 1986, 5, 2737–2744. [Google Scholar] [CrossRef] [PubMed]
  68. Behe, M.; Felsenfeld, G. 1981 Effects of methylation on a synthetic polynucleotide: the B--Z transition in poly(dG-m5dC).poly(dG-m5dC). Proc. Natl. Acad. Sci. U S A 78, 1619–1623. [CrossRef] [PubMed]
  69. Wang, J.; Wang, S.; Zhong, C.; Tian, T.; Zhou, X. 2015 Novel insights into a major DNA oxidative lesion: its effects on Z-DNA stabilization. Org. Biomol. Chem. 13, 8996–8999. [CrossRef] [PubMed]
  70. Herbert, A.; Alfken, J.; Kim, Y. G.; Mian, I. S.; Nishikura, K.; Rich, A. 1997 A Z-DNA binding domain present in the human editing enzyme, double-stranded RNA adenosine deaminase. Proc. Natl. Acad. Sci. U S A 94, 8421–8426. [CrossRef] [PubMed]
  71. Schade, M.; Turner, C. J.; Kuhne, R.; Schmieder, P.; Lowenhaupt, K.; Herbert, A.; Rich, A.; Oschkinat, H. 1999 The solution structure of the Zα domain of the human RNA editing enzyme ADAR1 reveals a prepositioned binding surface for Z-DNA. Proc. Natl. Acad. Sci. U S A 96, 12465–12470. [CrossRef] [PubMed]
  72. Schwartz, T.; Rould, M. A.; Lowenhaupt, K.; Herbert, A.; Rich, A. 1999 Crystal structure of the Zα domain of the human editing enzyme ADAR1 bound to left-handed Z-DNA. Science 284, 1841–1845. [CrossRef] [PubMed]
  73. Schwartz, T.; Behlke, J.; Lowenhaupt, K.; Heinemann, U.; Rich, A. 2001 Structure of the DLM-1-Z-DNA complex reveals a conserved family of Z-DNA-binding proteins. Nat. Struct. Biol. 8, 761–765. [CrossRef] [PubMed]
  74. Herbert, A. 2020 Mendelian disease caused by variants affecting recognition of Z-DNA and Z-RNA by the Za domain of the double-stranded RNA editing enzyme ADAR. Eur. J. Hum. Genet. 28, 114–117. [CrossRef] [PubMed]
  75. Zhang, T.; Yin, C.; Boyd, D. F.; Quarato, G.; Ingram, J. P.; Shubina, M.; Ragan, K. B.; Ishizuka, T.; Crawford, J. C.; Tummers, B.; et al. 2020 Influenza Virus Z-RNAs Induce ZBP1-Mediated Necroptosis. Cell 180, 1115–1129. [CrossRef] [PubMed]
  76. de Reuver, R.; Dierick, E.; Wiernicki, B.; Staes, K.; Seys, L.; De Meester, E.; Muyldermans, T.; Botzki, A.; Lambrecht, B. N.; Van Nieuwerburgh, F.; et al. 2021 ADAR1 interaction with Z-RNA promotes editing of endogenous double-stranded RNA and prevents MDA5-dependent immune activation. Cell Rep. 36, 109500. [CrossRef] [PubMed]
  77. Jiao, H.; Wachsmuth, L.; Wolf, S.; Lohmann, J.; Nagata, M.; Kaya, G. G.; Oikonomou, N.; Kondylis, V.; Rogg, M.; Diebold, M.; et al. 2022 ADAR1 averts fatal type I interferon induction by ZBP1. Nature 607, 776–783. [CrossRef] [PubMed]
  78. Zhang, T.; Yin, C.; Fedorov, A.; Qiao, L.; Bao, H.; Beknazarov, N.; Wang, S.; Gautam, A.; Williams, R. M.; Crawford, J. C.; et al. 2022 ADAR1 masks the cancer immunotherapeutic promise of ZBP1-driven necroptosis. Nature 606, 594–602. [CrossRef] [PubMed]
  79. de Reuver, R.; Verdonck, S.; Dierick, E.; Nemegeer, J.; Hessmann, E.; Ahmad, S.; Jans, M.; Blancke, G.; Van Nieuwerburgh, F.; Botzki, A.; et al. 2022 ADAR1 prevents autoinflammation by suppressing spontaneous ZBP1 activation. Nature 607, 784–789. [CrossRef] [PubMed]
  80. Hubbard, N. W.; Ames, J. M.; Maurano, M.; Chu, L. H.; Somfleth, K. Y.; Gokhale, N. S.; Werner, M.; Snyder, J. M.; Lichauco, K.; Savan, R.; et al. 2022 ADAR1 mutation causes ZBP1-dependent immunopathology. Nature 607, 769–775. [CrossRef] [PubMed]
  81. Herbert, A. 2023 Z-DNA and Z-RNA: Methods-Past and Future. Methods Mol. Biol. 2651, 295–329. [CrossRef] [PubMed]
  82. Diallo, M. A.; Pirotte, S.; Hu, Y.; Morvan, L.; Rakus, K.; Suarez, N. M.; PoTsang, L.; Saneyoshi, H.; Xu, Y.; Davison, A. J.; et al. 2023 A fish herpesvirus highlights functional diversities among Zalpha domains related to phase separation induction and A-to-Z conversion. Nucleic Acids Res. 51, 806–830. [CrossRef] [PubMed]
  83. Herbert, A. The evolutionary entanglement of flipons with zinc fingers and retroelements has engendered a large family of Z-DNA and G-quadruplex binding proteins. Open Biol. 2025, 15, 250171. [Google Scholar] [CrossRef] [PubMed]
  84. Konig, P.; Giraldo, R.; Chapman, L.; Rhodes, D. The crystal structure of the DNA-binding domain of yeast RAP1 in complex with telomeric DNA. Cell 1996, 85, 125–136. [Google Scholar] [CrossRef] [PubMed]
  85. Traczyk, A.; Liew, C. W.; Gill, D. J.; Rhodes, D. 2020 Structural basis of G-quadruplex DNA recognition by the yeast telomeric protein Rap1. Nucleic Acids Res. 48, 4562–4571. [CrossRef] [PubMed]
  86. Madden, S. K.; Tannahill, D.; Balasubramanian, S. 2026 Dissecting the Binding Interactions of the Chromatin Remodeler SMARCA4 with G-Quadruplex DNA. Biochemistry 65, 670–677. [CrossRef] [PubMed]
  87. Herbert, A. 2023 Flipons and small RNAs accentuate the asymmetries of pervasive transcription by the reset and sequence-specific microcoding of promoter conformation. J. Biol. Chem. 299, 105140. [CrossRef] [PubMed]
  88. Liu, L. F.; Wang, J. C. 1987 Supercoiling of the DNA template during transcription. Proc. Natl. Acad. Sci. U S A 84, 7024–7027. [CrossRef] [PubMed]
  89. Herbert, A. The ancient Z-DNA and Z-RNA specific Zalpha fold has evolved modern roles in immunity and transcription through the natural selection of flipons. R Soc. Open Sci. 2024, 11, 240080. [Google Scholar] [CrossRef] [PubMed]
  90. Fleming, A. M.; Zhou, J.; Wallace, S. S.; Burrows, C. J. 2015 A Role for the Fifth G-Track in G-Quadruplex Forming Oncogene Promoter Sequences during Oxidative Stress: Do These "Spare Tires" Have an Evolved Function? ACS Cent. Sci. 1, 226–233. [CrossRef] [PubMed]
  91. Fay, M. M.; Lyons, S. M.; Ivanov, P. 2017 RNA G-Quadruplexes in Biology: Principles and Molecular Mechanisms. J. Mol. Biol. 429, 2127–2147. [CrossRef] [PubMed]
  92. Spiegel, J.; Adhikari, S.; Balasubramanian, S. 2020 The Structure and Function of DNA G-Quadruplexes. Trends Chem. 2, 123–136. [CrossRef] [PubMed]
  93. Herbert, A. A Compendium of G-Flipon Biological Functions That Have Experimental Validation. Int. J. Mol. Sci. 2024, 25. [Google Scholar] [CrossRef] [PubMed]
  94. Frank-Kamenetskii, M. D.; Mirkin, S. M. 1995 Triplex DNA structures. Annu Rev. Biochem. 64, 65–95. [CrossRef] [PubMed]
  95. Toscano-Garibay, J. D.; Aquino-Jarquin, G. 2014 Transcriptional regulation mechanism mediated by miRNA-DNA*DNA triplex structure stabilized by Argonaute. Biochim Biophys. Acta 1839, 1079–1083. [Google Scholar] [CrossRef] [PubMed]
  96. Li, Y.; Syed, J.; Sugiyama, H. 2016 RNA-DNA Triplex Formation by Long Noncoding RNAs. Cell Chem. Biol. 23, 1325–1333. [CrossRef] [PubMed]
  97. Senturk Cetin, N.; Kuo, C. C.; Ribarska, T.; Li, R.; Costa, I. G.; Grummt, I. 2019 Isolation and genome-wide characterization of cellular DNA:RNA triplex structures. Nucleic Acids Res. 47, 2306–2321. [CrossRef] [PubMed]
  98. Warwick, T.; Brandes, R. P.; Leisegang, M. S. 2023 Computational Methods to Study DNA:DNA:RNA Triplex Formation by lncRNAs. In Noncoding RNA. [CrossRef] [PubMed]
  99. Hisey, J. A.; Masnovo, C.; Mirkin, S. M. 2024 Triplex H-DNA structure: the long and winding road from the discovery to its role in human disease. NAR Mol. Med. 1, ugae024. [CrossRef] [PubMed]
  100. Herbert, A. 2026 The Chromaverse Is Colored by Triplexes Formed Through the Interactions of Noncoding RNAs with HNPRNPU, TP53, AGO, REL Proteins, Intrinsically-Disordered Regions, and Flipons. Int. J. Mol. Sci. 27. [CrossRef] [PubMed]
  101. Erickson, B.; Sheridan, R. M.; Cortazar, M.; Bentley, D. L. 2018 Dynamic turnover of paused Pol II complexes at human promoters. Genes Dev. 32, 1215–1225. [CrossRef] [PubMed]
  102. Hasegawa, Y.; Struhl, K. 2019 Promoter-specific dynamics of TATA-binding protein association with the human genome. Genome Res. 29, 1939–1950. [CrossRef] [PubMed]
  103. Pimmett, V. L.; Dejean, M.; Fernandez, C.; Trullo, A.; Bertrand, E.; Radulescu, O.; Lagha, M. 2021 Quantitative imaging of transcription in living Drosophila embryos reveals the impact of core promoter motifs on promoter state dynamics. Nat. Commun. 12, 4504. [CrossRef] [PubMed]
  104. Szczurek, A. T.; Dimitrova, E.; Kelley, J. R.; Blackledge, N. P.; Klose, R. J. The Polycomb system sustains promoters in a deep OFF state by limiting pre-initiation complex formation to counteract transcription. Nat. Cell Biol. 2024, 26, 1700–1711. [Google Scholar] [CrossRef] [PubMed]
  105. Hyman, A. A.; Weber, C. A.; Julicher, F. 2014 Liquid-liquid phase separation in biology. Annu Rev. Cell Dev. Biol. 30, 39–58. [CrossRef] [PubMed]
  106. von Hippel, P. H.; Berg, O. G. 1989 Facilitated target location in biological systems. J. Biol. Chem. 264, 675–678. [CrossRef]
  107. Singer, R. H. RNA zipcodes for cytoplasmic addresses. Curr. Biol. 1993, 3, 719–721. [Google Scholar] [CrossRef] [PubMed]
  108. Blobel, G.; Sabatini, D. D. 1971 Ribosome-Membrane Interaction in Eukaryotic Cells. In Biomembranes; DOI, Ed.; pp. 193–195.
  109. Grosberg, A. I. U.; Khokhlov, A. R. Statistical physics of macromolecules; AIP Press: New York, 1994. [Google Scholar]
  110. Speer, S. L.; Stewart, C. J.; Sapir, L.; Harries, D.; Pielak, G. J. 2022 Macromolecular Crowding Is More than Hard-Core Repulsions. Annu Rev. Biophys. 51, 267–300. [CrossRef] [PubMed]
  111. Muñoz, M. A. 2018 Colloquium: Criticality and dynamical scaling in living systems. Rev. Mod. Phys. 90. [CrossRef]
  112. Ellis, R. J. 2001 Macromolecular crowding: obvious but underappreciated. Trends Biochem Sci. 26, 597–604. [CrossRef] [PubMed]
  113. Lafontaine, D. L. J.; Riback, J. A.; Bascetin, R.; Brangwynne, C. P. 2021 The nucleolus as a multiphase liquid condensate. Nat. Rev. Mol. Cell Biol. 22, 165–182. [CrossRef] [PubMed]
  114. Yamazaki, T.; Yamamoto, T.; Yoshino, H.; Souquere, S.; Nakagawa, S.; Pierron, G.; Hirose, T. 2021 Paraspeckles are constructed as block copolymer micelles. EMBO J. 40, e107270. [CrossRef] [PubMed]
  115. Nakagawa, S.; Hirose, T. 2012 Paraspeckle nuclear bodies--useful uselessness? Cell Mol. Life Sci. 69, 3027–3036. [CrossRef] [PubMed]
  116. Grigorev, V.; Wingreen, N. S.; Zhang, Y. 2025 Conformational Entropy of Intrinsically Disordered Proteins Bars Intruders from Biomolecular Condensates. PRX Life 3. [CrossRef] [PubMed]
  117. Kelley, F. M.; Ani, A.; Pinlac, E. G.; Linders, B.; Favetta, B.; Barai, M.; Ma, Y.; Singh, A.; Dignon, G. L.; Gu, Y.; et al. 2025 Controlled and orthogonal partitioning of large particles into biomolecular condensates. Nat. Commun. 16, 3521. [CrossRef] [PubMed]
  118. Jansen, N. S.; Vertegaal, A. C. O. 2021 A Chain of Events: Regulating Target Proteins by SUMO Polymers. Trends Biochem Sci. 46, 113–123. [CrossRef] [PubMed]
  119. Ditlev, J. A.; Case, L. B.; Rosen, M. K. 2018 Who's In and Who's Out-Compositional Control of Biomolecular Condensates. J. Mol. Biol. 430, 4666–4684. [CrossRef] [PubMed]
  120. Skipper, K.; Wickham, S. F. J. 2026 DNA nanostars that self-assemble into core-shell condensate microdroplets. Nanoscale Horiz. 11, 1320–1331. [CrossRef] [PubMed]
  121. Rajendran, A.; Castaneda, C. A. 2025 Protein quality control machinery: regulators of condensate architecture and functionality. Trends Biochem Sci. 50, 106–120. [CrossRef] [PubMed]
  122. Ingber, D. E.; Wang, N.; Stamenovic, D. Tensegrity, cellular biophysics, and the mechanics of living systems. Rep. Prog. Phys. 2014, 77, 046603. [Google Scholar] [CrossRef] [PubMed]
  123. Shin, Y.; Chang, Y. C.; Lee, D. S. W.; Berry, J.; Sanders, D. W.; Ronceray, P.; Wingreen, N. S.; Haataja, M.; Brangwynne, C. P. 2018 Liquid Nuclear Condensates Mechanically Sense and Restructure the Genome. Cell 175, 1481–1491 e1413. [CrossRef] [PubMed]
  124. Ramaker, R. C.; Hardigan, A. A.; Goh, S. T.; Partridge, E. C.; Wold, B.; Cooper, S. J.; Myers, R. M. 2020 Dissecting the regulatory activity and sequence content of loci with exceptional numbers of transcription factor associations. Genome Res. 30, 939–950. [CrossRef] [PubMed]
  125. Herbert, A. 2025 Control of Gene Expression by Proteins That Bind Many Alternative Nucleic Acid Structures Through the Same Domain. Int. J. Mol. Sci. 27. [CrossRef] [PubMed]
  126. Banani, S. F.; Lee, H. O.; Hyman, A. A.; Rosen, M. K. 2017 Biomolecular condensates: organizers of cellular biochemistry. Nat. Rev. Mol. Cell Biol. 18, 285–298. [CrossRef] [PubMed]
  127. Dai, Y.; You, L.; Chilkoti, A. 2023 Engineering synthetic biomolecular condensates. In Nat Rev Bioeng.; pp. 1–15. [CrossRef] [PubMed]
  128. Li, S.; Kim, Y.; Wang, K.; Payson, E. J.; Tang, A. A.; Villalba Nieto, M.; Osmanovic, D.; Yang, M.; Dilao, D.; Bermudez, A.; et al. 2026 Programmable artificial RNA condensates in mammalian cells. Nat. Nanotechnol. [CrossRef] [PubMed]
Figure 1. Condensates from the inside out. Flipons can nucleate condensates in active genomic regions. A. Nucleolin bound to a parallel-stranded G-quadruplex (based on protein database structure 9CB5 [63]). B. The Zα domain engages left-handed Z-DNA that forms with a B-DNA segment, as modeled with AlphaFold 3 [64]. C. The condensate that forms the nucleolus has layers. The unique components are labeled. D. Stress granules are modulated by attachment to microtubules and actin-myosin filaments of the cell’s cytoskeleton. Images drawn with nanobanana2.
Figure 1. Condensates from the inside out. Flipons can nucleate condensates in active genomic regions. A. Nucleolin bound to a parallel-stranded G-quadruplex (based on protein database structure 9CB5 [63]). B. The Zα domain engages left-handed Z-DNA that forms with a B-DNA segment, as modeled with AlphaFold 3 [64]. C. The condensate that forms the nucleolus has layers. The unique components are labeled. D. Stress granules are modulated by attachment to microtubules and actin-myosin filaments of the cell’s cytoskeleton. Images drawn with nanobanana2.
Preprints 230163 g001
Figure 2. Search based on the Surface Hash of Condensates. A. The optimal packing of equal-sized spheres within a cubic lattice occupies over 50% of the volume. Surface features, in this case colors, enable rapid search to locate a subset of colored spheres that contain the required information. B. The uncondensed elements of the pink spheres occupy more volume and act as a barrier that impedes diffusion of other molecules through that space. C. A cellular implementation of sphere packing produced discrete clusters that optimize information storage and retrieval. The condensates that form spherical structures are shown with porous surfaces that allow free diffusion of the small molecules required by enzymes that dynamically rehash their surfaces by modifying their contents in an outward-directed process. The reduced number of spheres increases the free space for folding and unfolding of the protein's intrinsically disordered regions, which are shown as wavy lines. The decrease in sphere number also increases the diffusion volume. The surface hash on each condensate allows rapid searches by excluding certain genomic regions, analogous to a Bloom filter, while targeting transcription factors and complexes to others. The hash can be modified by inward-directed processes triggered by environmental events that perturb the cell. Images drawn with nanobanana2.
Figure 2. Search based on the Surface Hash of Condensates. A. The optimal packing of equal-sized spheres within a cubic lattice occupies over 50% of the volume. Surface features, in this case colors, enable rapid search to locate a subset of colored spheres that contain the required information. B. The uncondensed elements of the pink spheres occupy more volume and act as a barrier that impedes diffusion of other molecules through that space. C. A cellular implementation of sphere packing produced discrete clusters that optimize information storage and retrieval. The condensates that form spherical structures are shown with porous surfaces that allow free diffusion of the small molecules required by enzymes that dynamically rehash their surfaces by modifying their contents in an outward-directed process. The reduced number of spheres increases the free space for folding and unfolding of the protein's intrinsically disordered regions, which are shown as wavy lines. The decrease in sphere number also increases the diffusion volume. The surface hash on each condensate allows rapid searches by excluding certain genomic regions, analogous to a Bloom filter, while targeting transcription factors and complexes to others. The hash can be modified by inward-directed processes triggered by environmental events that perturb the cell. Images drawn with nanobanana2.
Preprints 230163 g002
Figure 3. Varying the number of search surfaces dramatically alters the reaction time. A. Lagrangian and parameters used to apply the principle of least action to determine the search time when the surface search (α) and 3D diffusion (β) parameters were varied to determine how search time varied with sphere surface area and number when spheres were optimally packed in an arbitrary volume. B, C, and D illustrate how search time can be minimized by leveraging 2D surface-area operations. The feasible path for the surface constraint K = 500 is mapped onto the cost surface for different combinations of α and β, with a red arrow indicating the optimal search time.
Figure 3. Varying the number of search surfaces dramatically alters the reaction time. A. Lagrangian and parameters used to apply the principle of least action to determine the search time when the surface search (α) and 3D diffusion (β) parameters were varied to determine how search time varied with sphere surface area and number when spheres were optimally packed in an arbitrary volume. B, C, and D illustrate how search time can be minimized by leveraging 2D surface-area operations. The feasible path for the surface constraint K = 500 is mapped onto the cost surface for different combinations of α and β, with a red arrow indicating the optimal search time.
Preprints 230163 g003
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.