Submitted:
16 September 2026
Posted:
17 September 2026
You are already at the latest version
Abstract
The structural complexity of eukaryotic genomes remains a major open challenge in molecular biology. Genome architecture operates across multiple scales—from chromatin loops that mediate regulatory interactions, to isochores, chromatin domains and spatial territories. Genes within these regions are differentially distributed, replicated and expressed, underscoring the importance of long-range correlations for coordinated cellular functions. Here, we describe a previously unrecognized and evolutionarily conserved structural feature, the DNA strand parity, of eukaryotic genome. This principle governs the distribution of virtually all major genomic components. Across eleven evolutionary distant animal and plant species genes, exons, introns and all major transposable element (TE) classes occur in nearly equal number on the plus and minus DNA strands and display highly similar length distributions. The parity also extends to the start positions of TE fragments and to local sequence composition, as illustrated for LINE elements. This property provides a unifying explanation for longstanding observations of sequence-level symmetry, including the symmetry described by Chargaff’s second rule for single DNA strands. We further show that gene and TE lengths follow power-law-like distributions identical on either strand, with correlated scaling exponents across species, suggesting conserved fractal-like constraints during genome evolution. Together, these results establish strand parity as a universal principle of eukaryotic genome organization and support the view that it is an emergent property of the global fractal architecture of the genome.
Keywords:
genome organization
; strand parity of genes and transposons
; fractal-like scaling behavior
1. Introduction
Genome organization in eukaryotes is governed by a hierarchical architecture that integrates chromatin interactions, sequence composition, and spatial nuclear compartmentalization. Investigations into chromatin–chromatin interactions within interphase nuclei have revealed the presence of A/B compartments—large, long-range interacting chromosomal regions shaped by the fractal globule model, which facilitates coordinated gene regulation. These compartments correlate with key genomic features, including gene density, chromatin activity, DNA accessibility, and histone modifications [1]. At higher resolution, topologically associating domains (TADs) have been defined as self-interacting chromosomal units, now recognized as fundamental building blocks of genome organization [2,3,4]. At even finer scales, dynamic chromatin loops mediate transient interactions between regulatory elements and gene promoters, with their length and stability modulated by cellular context. During mitosis, A/B compartments and TADs are no longer detectable; nonetheless, structural genome organization persists, as evidenced by chromosome banding patterns. R/G bands reflect distinct chromatin states and are associated with differential gene content (e.g., constitutive vs tissue-specific), replication timing, and the spatial distribution of LINE and SINE transposons. These bands also align with GC-rich and GC-poor isochores—compositionally homogeneous DNA regions considered a fundamental layer of genomic structure [5]. Together, these observations underscore the high degree of order and non-randomness that governs genome architecture [6]. Despite substantial progress in elucidating the relationship between genome structure and function, many aspects remain unresolved. One such foundational yet enigmatic feature is Chargaff’s second rule (Chargaff’s PR2), also known as the Symmetry Principle [7,8]. It posits that for any small oligonucleotide, the frequency approximates that of its reverse complement within the same DNA strand [9]. This principle holds across double-stranded eukaryotic, bacterial, viral, and archaeal genomes. Although its mechanistic basis is still under investigation, it may reflect a deeper, evolutionarily conserved parity rule. In this study, we describe a previously unrecognized property of genome architecture: strand parity in the number and size of genomic elements, observed across a broad phylogenetic spectrum. Specifically, we demonstrate that in all analyzed genomes (eleven species across distant evolutionary lineages), the frequency and length distributions of genes are strikingly similar between the plus and minus strands. This symmetry extends to other genomic features, including exons, introns, and all classes of transposable elements (TEs). Moreover, the distributions of genes and LINEs length on both strands conform to a shared power-law pattern across species. These findings suggest a solid relationship between genomic architecture and evolutionary dynamics, revealing a novel layer of genome organization rooted in strand-level symmetry and fractal-like scaling behavior.
2. Materials and Methods
We analyzed the genomes of eleven animal and plant species representing distantly related evolutionary taxa: Homo sapiens (Human), Mus musculus (Mouse), Bos taurus (Cow), Monodelphis domestica (Opossum), Gallus gallus (Chicken), Xenopus tropicalis (Xenopus), Danio rerio (Zebrafish), Drosophila melanogaster (Drosophila), Caenorhabditis elegans (C. elegans), Arabidopsis thaliana (Arabidopsis), and Zea mays (Maize). Scientific names were used at first mention and throughout all tables. In the main text, capitalized common names were used for readability, except for Arabidopsis, C. elegans, Drosophila, and Xenopus, in accordance with standard genetic nomenclature. Species names and assembly sources were listed in Supplementary Table S1.
Gene datasets were obtained from the NCBI Genome Annotation database [10], selecting the reference assembly for each species. An exception was C. elegans, whose genes—absent from NCBI—were retrieved from the UCSC Genome Browser [11]. Annotated genes included RefSeq-curated protein-coding (NM) and long non-coding RNA (NR) transcripts located on complete genomic molecules; predicted genes and genes on unplaced scaffolds were excluded. For C. elegans, only NM genes were included due to the lack of distinction among NR-type genes. Similarly, M. domestica and X. tropicalis contained only NM genes in the NCBI database. For each gene, only the representative transcript was analyzed, and chromosome coordinates were used to calculate gene sequence length. UCSC data were also used to define chromosome coordinates for G/R bands, exons, and introns. Coordinates for A/B compartments were derived from [12], and gene data for the hg19 assembly were retrieved from UCSC.
Datasets for TEs were obtained from the UCSC Genome Browser. For LINEs, SINEs, LTRs, and DNA transposons, all elements annotated in the RepeatMasker track were included without filtering; both complete and fragmented LINE elements have been included. DNA sequence length was calculated using start and end coordinates in the genomic sequence. TE start positions were defined using the internal coordinates within the repeat sequence.
Given the large number of observations per dataset (in some cases hundreds of thousands), analyses were performed using both individual sequence values and binned representations of the data. To characterize the global structure of sequence-length distributions, a rank–size representation was adopted. For each dataset, sequences were ordered by increasing length, and an integer rank was assigned to each observation. Rank–size plots were constructed by reporting sequence length (b) as a function of sequence rank. This representation allowed direct visual comparison of distributional patterns across strands and species without assuming a specific underlying probability distribution.
To further investigate the distributional properties of sequence lengths, empirical frequency distributions were computed by grouping sequence lengths into classes. For each dataset, sequences were assigned to bins defined over the observed length range; identical bins were used for gene lengths among vertebrates (15 kb), among non-vertebrates (3 kb), and for LINE lengths among all species (200b), enabling general comparison. The corresponding frequencies were used for statistical analyses. Given the substantial differences in absolute gene lengths across species, the frequency classes were ordered by increasing sequence length, and their ordinal position was used as a normalized index suitable for comparative purposes.
Log–log plots were constructed by reporting class frequencies as a function of class rank, defined as the ordinal index of each size class after ordering by increasing sequence length. This approach enabled comparison across datasets with heterogeneous length scales by focusing on relative ordering rather than absolute size. In the log–log representation, a subset of the distribution exhibited an approximately linear behavior over a large range of class ranks. This region was interpreted as a power-law-like scaling regime, indicating scale-invariant behavior over that interval. Linear regression analysis was performed within the identified scaling region to estimate the slope and the coefficient of determination (), parameters used as descriptive measures of the extent and robustness of the scaling behavior.
The scaling exponents and line slopes obtained from single-strand values were used to compare the strands and thereby define their parity. The scaling exponents and slopes obtained from values aggregated across both strands were used to compare species and thus enable evolutionary interspecific comparisons.
Scaling extents were expressed in two complementary ways. The first, the rank-based extent, represented the number of decades covered on the rank axis; because this measure depends on the number of bins and may underestimate the true physical scale of the distribution, we also reported the size-based extent, in which the lower and upper bounds represented the size interval over which the power-law-like behavior held. This representation provided a descriptive characterization of the physical range over which scale-invariant behavior was maintained.
To investigate global patterns in the TE distribution of genomic start positions, rank–start-position plots were constructed, providing an intuitive, assumption-free visualization of their complex patterns across strands and species. Empirical frequency distributions were also computed, and their histograms were used for visual and statistical comparison.
Statistical analyses of gene and TE lengths, including chi-square tests and linear regression, were performed using STATGRAPHICS Centurion XVIII and Microsoft Excel. We defined statistical parity between datasets as cases where p-values exceeded , a threshold commonly accepted when comparing very large samples. This analytical framework enabled the identification of conserved distributional patterns across genomic datasets and provided a flexible approach for comparative analyses of large-scale genomic features.
3. Results
3.1. Gene Length Strand Parity Across Phylogenetically Distant Species
Strand data were analyzed for eleven animal and plant species, spanning distant evolutionary lineages: Arabidopsis (Arabidopsis thaliana), C.elegans (Caenorhabditis elegans), Chicken (Gallus gallus), Cow (Bos taurus), Drosophila (Drosophila melanogaster), Human (Homo sapiens), Maize (Zea mays), Mouse (Mus musculus), Opossum (Monodelphis domestica), Xenopus (Xenopus tropicalis) and Zebrafish (Danio rerio) (Supplementary Table S1). Gene numbers on the plus and minus strands are approximately equal, and in some cases impressively almost identical as in Mouse (12,373 vs. 12,359),Opossum (10,867 vs 10,804), Drosophila (7,963 vs 7,968) (Supplementary Table S2). Gene length distributions substantially overlap between strands, with reverse-plotted curves revealing striking visual symmetry for all species (Figure 1A, Supplementary Figure S1). In the human genome, the strand parity is evident not only at genome-wide but also within defined subregions including entire chromosomes, chromosomal subdomains such as G and R bands, and A/B epigenomic compartments. Furthermore, genes located on the opposite strands of chromosome 1, for example, exhibit comparable numbers of exons and introns with fully symmetrical length distributions (Supplementary Figure S2). Importantly, strand parity does not imply perfect one-to-one correspondence between strands; rather, it emerges as a robust statistical property, with only minor deviations (typically ). Such pervasive strand parity encompassing structural genomic elements across evolutionary distant taxa has not previously been documented.
3.2. Strand Parity Extends to Transposable Elements
Transposable elements comprise roughly half of many eukaryotic genomes and are therefore major determinants of genomic structure [13]. LINEs, very ancient transposon ubiquitous across eukaryotes, exhibit, as observed for genes, full strand parity in number and length distribution encompassing full-length and fragmented elements across all species examined (Figure 1C, Supplementary Figure S3). Nearly identical counts are observed in Opossum (1,268,157 vs 1,268,230), Xenopus (140,402 vs 139,512), an impressive symmetry given the very large number of elements involved (Supplementary Table S2). Parity extends to the start positions, defined by UCSC annotations as the consensus sequence nucleotide at which each LINE element begins which, to our knowledge, are analyzed here for the first time. In humans, these start positions display a complex multimodal distribution that was nearly identical on the plus and minus strands, revealing a deep symmetry conserved in all species (Figure 1E-F and Figure 2). Their frequency-class analysis reveals prominent peaks in mammals (3,200–3,400 and 5,800–6,400 nt) and partially conserved positions in the other species, occurring on both strands with nearly perfect reciprocity. These peaks correspond to amplified fragments with shared start positions, likely reflecting the conservation of regions containing functional domains under positive selection. LINE elements have indeed been shown to harbor transcription factors and architectural protein binding sites, form secondary structures, and participate in long-range regulatory interactions [14,15]. Strand parity extends beyond LINEs to all other TE classes, including LTR, DNA, and SINE elements; again, a nearly identical numbers and length distributions occurrr between strands across all species examined, with only minor taxon-specific deviations (Figure 3, Supplementary Table S3). Start position distributions also retain strand reciprocity in every TE class, indicating that the observed symmetry is not restricted to specific TE types but represented a global principle of genome organization.
In the human genome, the parity rule can be demonstrated even for all major LINE, LTR, DNA, and SINE transposon families, which show full strand symmetry in number, size, and start positions (Supplementary Figure S4). Furthermore, for the LINE-L1 family on human chromosome 22, the analysis of ATGC content reveals a clear identity between the two strands for fragments belonging to the same start position class. This pattern indicates, albeit indirectly, a substantial DNA sequence-level equivalence between corresponding fragments on opposite strands (Supplementary Figure S5).
3.3. Gene and Transposon Lengths Show Power-Law-Like Distributions
Further key findings are that gene and LINE lengths follow almost identical power-law-like scaling on both strands. For genes by using identical homogeneous length classes (15kb bins in vertebrates, 3kb bins in non-vertebrates) the log-log representation of their frequencies F as a function of class ranks r exhibited an approximately linear behavior over a wide range of class ranks, as determined by linear regression analysis (Supplementary Table S4A). In humans, the gene-length distributions on either strand follow the relation over 16 size classes, covering of genes (Figure 1B). Scaling exponents were nearly identical (), as were the intercepts. Power-law-like scaling persists for genes within G/R bands, A/B compartments, and for exons and introns (Supplementary Figure S2). Comparable scaling exponents are consistently observed for the two strands across all other species, with high goodness of fit () (Supplementary Figure S1, Supplementary Table S4).
LINE elements (classes of 200b bins) exhibit a similar pattern of power-law-like length scaling in log-log plots, showing matching exponents between strands and a high number of overlapping classes in all species (Figure 1D, Supplementary Figure S3, Supplementary Table S4B). In mammals, a second scaling region is observed on both strands, particularly pronounced in mouse. The primary scaling region substantially captures the size variation of the most LINE fragments, whereas the secondary region, ubiquitous in mammals and present also in Xenopus, corresponds to the distinct size variation of full-length LINEs. The strong correspondence of power-law exponents between the two strands provides formal evidence for the widespread strand parity in gene and LINE lengths.
3.4. Evolutionary Implications of Power-Law-Like Scaling for Genome Structure
The power-law-like scaling analysis reveals additional general features. When comparing species, aggregating lengths from both strands to increase statistical power, the scaling exponents follow a clear evolutionary gradient (Figure 4).
For genes, mammals show higher-magnitude scaling exponents from -1.6362 to -1.8791, consistent with broader size ranges; non-mammalian vertebrates exhibit reduced values from -1.9797 to -2.1041, whereas non-vertebrates display steeper slopes, from -2.0769 to -5.6131, and narrower size distributions (Supplementary Table S5A). Over evolutionary time, genomes underwent increases in gene number, size and complexity, driven by duplication and recombination of existing genes or gene fragments. In this context, the scaling exponents identified here could be interpreted as expansion factors that shaped the evolutionary trajectories of gene size within each taxon. For LINEs, variation in the scaling exponent of the primary scaling region shows an opposite trend relative to genes, separating vertebrates, ranging from -2.2130 to -2.7611, from the other species, ranging from -1.4141 to -1.7138 (Supplementary Table S5B). From an evolutionary standpoint, these values likely represent the degree of LINE fragmentation, which was particularly pronounced in vertebrates. The limited variation observed in mammals reflects the partial structural modification of full-length elements, suggesting the involvement of distinct mechanisms.
Scaling exponents and extents are reported to reflect dense genome packing and self-similar, fractal-like spatial organization. Here, we quantified the scaling extent as a descriptive measure of the genomic intervals over which power-law behavior was observed for gene and LINE lengths. Because rank-based extents underestimated the underlying physical size range, we also provide size-based extents, defined by the lower and upper bounds of the distributions, for descriptive purposes. Gene lengths exhibit scaling over more than three orders of magnitude, whereas LINE elements span two orders of magnitude, consistent with their more restricted size range. These broad intervals indicate robust power-law behavior, as scaling over two to three decades, beyond statistical artefacts, is widely regarded as evidence of genuine self-similarity in natural systems [16,17]. Together, these results support a fractal-like organization of genomes.
4. Discussion
4.1. A Unified Framework for Genome Strand Parity
In this study, we identify strand parity as a universal principle of genomic architecture, encompassing all major genomic elements, both in terms of number and size, in phylogenetically distant eukaryotic from both animal and plant lineages. Its breadth implies deep evolutionary conservation over billions of years, sustained by genomic processes linked to strong selective pressure. A related and longstanding genomic feature is Chargaff’s PR2 which states that within any single strand of a double-stranded DNA molecule, the number of adenine and guanine is approximately matched by that of thymine and cytosine [7]. Later generalized as sequence or inverse symmetry, Chargaff’s PR2 has been documented in double-stranded prokaryotic and eukaryotic genomes, and extended to short oligonucleotides, which occur at frequencies closely matching those of their reverse complements. Despite extensive investigation, the mechanistic basis of Chargaff’s PR2 has remained unresolved. Our findings suggest that these symmetries may reflect deeper, genome-wide constraints rather than an independent rule. Because the reverse complement of any oligonucleotide corresponds to the same oligonucleotide on the opposite strand, Chargaff’s PR2 and symmetry would naturally arise if oligonucleotides were constrained to occur in equal numbers on both strands of the DNA duplex. Consistent with this idea, we find that LINE-L1 fragments, with defined start positions and full parity in both number and length, also display ATGC composition between strands, indicative of equivalent nucleotide sequence. This global parity in LINE-L1 fragment therefore provides a plausible structural basis for the parity described by the Chargaff’s PR2 and by oligonucleotide reverse complement equivalence. In this view, the strand-parity rule should share conceptual ground with Chargaff’s PR2 and related symmetry principles, for years extensively documented [16,18,19,20].
4.2. Functional and Evolutionary Inferences
Power-law behavior has been documented in numerous biological context [21,22]. Genomic studies, although never focused on single DNA strands, have reported power-law distributions in exon spacers, transposons, conserved noncoding elements, -untranslated regions, protein lengths, single-cell transcriptomes and related structures [23,24], proposing several models for their origin. In our analysis, the rank-frequency power-law relationship for gene and LINE length, while primarily descriptive, provides a robust basis for cross-strand and cross-species comparison. The resulting scaling exponent and extent values serve as proxies for the degree and strength of scale invariance, pointing to an underlying self-similar, fractal-like organization of the genome. Multiple lines of evidence support this interpretation. Both chromatin topology and sequence organization have been shown to exhibit hierarchical, scale-invariant patterns consistent with fractal-like genomic architecture, in agreement with the trends observed in our rank–frequency analysis. Empirical power-law relationships between genomic distance and contact probability in eukaryotic nuclei support the fractal globule model of chromatin folding, which describes a knot-free, dynamically accessible chromatin state that facilitates long-range regulatory interactions [1]. Power-law distributions of homologous repetitive elements, including LINEs, have likewise been proposed to stabilize this fractal conformation [25]. Fractal properties have also been identified directly in genomic DNA sequences, spanning repetitive elements, CpG islands, coding and non-coding regions, whole chromosomes, and complete genomes [18,21,26,27,28,29,30,31,32,33]. In this framework, the gene and TE length variation observed here may have been shaped by topological constraints imposed by fractal genome organization. Evolutionary increases in gene number, size, and structural complexity, arising from chromosomal duplication and rearrangement, would have been guided by such constraints to preserve large-scale chromosome organization, while additional evolutionary processes such as TE transposition or fragmentation, would have followed the same pattern, possibly playing a role in the conformation [15,34,35,36]. One possible interpretation is that this symmetry reflects underlying dynamical constraints shaping genome architecture. While general fractal architecture provides a chromatin configuration that supports long-range regulatory interactions, specific topological features may diverge among species in relation to gene compositional complexity. In this sense the gradient in the scaling exponent may represent the intrinsic capacity of genomes to generate structural complexity and so serve as a proxy for genome evolvability, by offering insights into the mechanisms that drive evolutionary innovation [15,35]. Based on our comprehensive findings, we propose that strand parity is an emergent property of the universal fractal architecture of genomes, arising from topological constraints imposed by its functional organization. Such constraints would necessarily apply to all genomic elements irrespective of strand orientation, leading to parity as a natural consequence.
4.3. Evolutionary Implications of Strand Parity
The evolutionary basis of strand parity remains to be fully elucidated, yet several mechanistic considerations suggest that this property may confer structural and functional advantages to eukaryotic genomes. First, a balanced distribution of genes and transposable elements between the two strands may stabilize chromosomal topology by minimizing asymmetric torsional stress during replication, transcription and chromatin folding. Such symmetry could support the maintenance of higher-order genome architecture, particularly within fractal-like chromatin domains where spatial packing constraints are strong. Second, strand parity may enhance replicative robustness. Comparable densities of functional and repetitive elements on both strands reduce strand-specific replication delays, fork stalling and asymmetric mutational loads, thereby contributing to long-term genome stability. Third, a symmetric arrangement of transcription units may mitigate persistent conflicts between replication and transcription, a well-known source of genomic instability. Finally, strand parity may be maintained naturally from the cumulative effect of stochastic events such as inversions, segmental duplications and recombinations, which redistribute genomic elements across strands and preserve genomes in the equilibrium state. Experimental evidence indicates that disruptions of higher-order genome architecture, including its fractal-like organization, are a recurrent feature of human disease, particularly cancer. Multiple tumor types display loss of TAD integrity, aberrant compartment switching, and large-scale rewiring of chromatin interactions, all of which reflect a breakdown of the self-similar structural principles that normally govern genome folding. Such alterations are associated with transcriptional deregulation, replication stress, and increased genomic instability. In this context, the widespread strand parity described here may represent one of the structural constraints whose perturbation contributes to pathological chromatin states [37,38,39,40,41,42]. The conservation of strand parity across evolution suggests that it supports a stable and energetically favorable genomic configuration; its disruption, as observed in several malignancies, may therefore be part of the broader collapse of fractal genome organization that accompanies disease progression. Together, these considerations suggest that strand parity may represent an evolutionarily advantageous configuration that optimizes genome stability, regulatory efficiency and structural coherence across eukaryotic lineages.
Supplementary Materials
The following supporting information can be downloaded at the website of this paper posted on Preprints.org.
Author Contributions
Conceptualization: I.S. and G.M. Methodology and Genomic Analysis: I.S. Data collection and formal analysis: I.S. and G.M. Data interpretation: I.S., G.M. and A. M. Project supervision and funding acquisition: A.M. Writing—Original Draft Preparation: I.S. Writing—review and editing: I.S., G.M., and A.M. All authors have read and agreed to the published version of the manuscript.
Funding
This work was supported by a grant from Associazione Italiana per la Ricerca sul Cancro (IG23284) to A.M.
Informed Consent Statement
Not applicable
Data Availability Statement
The data supporting the findings of this study are available in Article and its Supplementary Information, and from the corresponding authors on reasonable request.
Conflicts of Interest
The authors declare no conflicts of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.
Abbreviations
The following abbreviations are used in this manuscript:
| Chargaff’s PR2 | Chargaff’s second rule |
| TADs | topologically associating domains |
| TE | transposable element |
References
- Lieberman-Aiden, E.; van Berkum, N.L.; Williams, L.; Imakaev, M.; Ragoczy, T.; Telling, A.; Amit, I.; Lajoie, B.R.; Sabo, P.J.; Dorschner, M.O.; et al. Comprehensive mapping of long-range interactions reveals folding principles of the human genome. Science 2009, 326, 289–293. [Google Scholar] [CrossRef] [PubMed]
- Dixon, J.R.; Selvaraj, S.; Yue, F.; Kim, A.; Li, Y.; Shen, Y.; Hu, M.; Liu, J.S.; Ren, B. Topological domains in mammalian genomes identified by analysis of chromatin interactions. Nature 2012, 485, 376–380. [Google Scholar] [CrossRef] [PubMed]
- Dixon, J.R.; Gorkin, D.U.; Ren, B. Chromatin Domains: The Unit of Chromosome Organization. Mol. Cell 2016, 62, 668–680. [Google Scholar] [CrossRef] [PubMed]
- Dekker, J.; Heard, E. Structural and functional diversity of Topologically Associating Domains. FEBS Lett. 2015, 589, 2877–2884. [Google Scholar] [CrossRef] [PubMed]
- Bernardi, G. Chromosome Architecture and Genome Organization. PLoS ONE 2015, 10, e0143739. [Google Scholar] [CrossRef] [PubMed]
- Misteli, T. The Self-Organizing Genome: Principles of Genome Architecture and Function. Cell 2020, 183, 28–45. [Google Scholar] [CrossRef] [PubMed]
- Rudner, R.; Karkas, J.D.; Chargaff, E. Separation of B. subtilis DNA into complementary strands. 3. Direct analysis. Proc. Natl. Acad. Sci. 1968, 60, 921–922. [Google Scholar] [CrossRef] [PubMed]
- Rudner, R.; Karkas, J.D.; Chargaff, E. Separation of B. subtilis DNA into complementary strands, I. Biological properties. Proc. Natl. Acad. Sci. 1968, 60, 630–635. [Google Scholar] [CrossRef] [PubMed]
- Prabhu, V.V. Symmetry observations in long nucleotide sequences. Nucleic Acids Res. 1993, 21, 2797–2800. [Google Scholar] [CrossRef] [PubMed]
- Bethesda (MD) National Library of Medicine (US). National Center for Biotechnology Information (NCBI). 1988. Available online: https://www.ncbi.nlm.nih.gov/ (accessed on 10 September 2026).
- Casper, J.; Speir, M.L.; Raney, B.J.; et al. The UCSC Genome Browser database: 2026 update. Nucleic Acids Res. 2026, 54, D1331–D1335. [Google Scholar] [CrossRef] [PubMed]
- Rao, S.S.P.; Huntley, M.H.; Durand, N.C.; Stamenova, E.K.; Bochkov, I.D.; Robinson, J.T.; Sanborn, A.L.; Machol, I.; Omer, A.D.; Lander, E.S.; et al. A 3D map of the human genome at kilobase resolution reveals principles of chromatin looping. Cell 2014, 159, 1665–1680. [Google Scholar] [CrossRef] [PubMed]
- Smit, A.F.A.; Hubley, R.; Green, P. RepeatMasker Genomic Datasets. 2013–2026. Available online: https://www.repeatmasker.org/genomicDatasets/RMGenomicDatasets.html.
- Choudhary, M.N.K.; Quaid, K.; Xing, X.; Schmidt, H.; Wang, T. Widespread contribution of transposable elements to the rewiring of mammalian 3D genomes. Nat. Commun. 2023, 14, 634. [Google Scholar] [CrossRef] [PubMed]
- Nishihara, H. Transposable elements as genetic accelerators of evolution: contribution to genome size, gene regulatory network rewiring and morphological innovation. Genes Genet. Syst. 2020, 94, 269–281. [Google Scholar] [CrossRef] [PubMed]
- Albrecht-Buehler, G. Asymptotically increasing compliance of genomes with Chargaff’s second parity rules through inversions and inverted transpositions. Proc. Natl. Acad. Sci. 2006, 103, 17828–17833. [Google Scholar] [CrossRef] [PubMed]
- West, G.B.; Brown, J.H. The origin of allometric scaling laws in biology from genomes to ecosystems: towards a quantitative unifying theory of biological structure and organization. J. Exp. Biol. 2005, 208, 1575–1592. [Google Scholar] [CrossRef] [PubMed]
- Almirantis, Y.; Provata, A.; Li, W. Noether’s Theorem as a Metaphor for Chargaff’s 2nd Parity Rule in Genomics. J. Mol. Evol. 2022, 90, 231–238. [Google Scholar] [CrossRef] [PubMed]
- Cristadoro, G.; Degli Esposti, M.; Altmann, E.G. The common origin of symmetry and structure in genetic sequences. Sci. Rep. 2018, 8, 15817. [Google Scholar] [CrossRef] [PubMed]
- Afreixo, V.; Bastos, C.A.; Garcia, S.P.; Rodrigues, J.M.; Pinho, A.J.; Ferreira, P.J. The breakdown of the word symmetry in the human genome. J. Theor. Biol. 2013, 335, 153–159. [Google Scholar] [CrossRef] [PubMed]
- Moreno, P.A.; Velez, P.E.; Martinez, E.; Garreta, L.E.; Diaz, N.; Amador, S.; Tischer, I.; Gutierrez, J.M.; Naik, A.K.; Tobar, F.; et al. The human genome: a multifractal analysis. BMC Genom. 2011, 12, 506. [Google Scholar] [CrossRef] [PubMed]
- Koonin, E.V.; Wolf, Y.I.; Karev, G.P. Molecular Biology Intelligence Unit; Springer, 2006; p. 257. [Google Scholar]
- Sellis, D.; Provata, A.; Almirantis, Y. Alu and LINE1 distributions in the human chromosomes: evidence of global genomic organization expressed in the form of power laws. Mol. Biol. Evol. 2007, 24, 2385–2399. [Google Scholar] [CrossRef] [PubMed]
- van Nimwegen, E. Scaling laws in the functional content of genomes. Trends Genet. 2003, 19, 479–484. [Google Scholar] [CrossRef] [PubMed]
- Polychronopoulos, D.; Tsiagkas, G.; Athanasopoulou, L.; Sellis, D.; Almirantis, Y. Chaos, Information Processing and Paradoxical Games. In Chaos, Information Processing and Paradoxical Games; World Scientific, 2014; pp. 221–252. [Google Scholar]
- Albrecht-Buehler, G. Fractal genome sequences. Gene 2012, 498, 20–27. [Google Scholar] [CrossRef] [PubMed]
- Alvarez-Ballesteros, Y.A.; Quiroz-Juarez, M.A.; Del-Rio-Correa, J.L.; Escobar-Ruiz, A.M. Exploring the multifractal behavior of the human genome T2T-CHM13v2.0: Graphical representations and cytogenetics. Comput. Biol. Med. 2025, 197, 110977. [Google Scholar] [CrossRef] [PubMed]
- Cattani, C.; Pierro, G. On the fractal geometry of DNA by the binary image analysis. Bull. Math. Biol. 2013, 75, 1544–1570. [Google Scholar] [CrossRef] [PubMed]
- Correia, J.P.; Silva, R.; Anselmo, D.H.A.L.; Vasconcelos, M.S.; da Silva, L.R. Multifractal Properties of Human Chromosome Sequences. Fractal Fract. 2024, 8, 312. [Google Scholar] [CrossRef]
- Garte, S. Fractal properties of the human genome. J. Theor. Biol. 2004, 230, 251–260. [Google Scholar] [CrossRef] [PubMed]
- Yu, Z.G.; Anh, V.; Lau, K.S. Measure representation and multifractal analysis of complete genomes. Phys. Rev. E Stat. Nonlin. Soft Matter Phys. 2001, 64, 031903. [Google Scholar] [CrossRef] [PubMed]
- Zhou, L.Q.; Yu, Z.G.; Deng, J.Q.; Anh, V.; Long, S.C. A fractal method to distinguish coding and non-coding sequences in a complete genome based on a number sequence representation. J. Theor. Biol. 2005, 232, 559–567. [Google Scholar] [CrossRef] [PubMed]
- Almirantis, Y.; Provata, A. An evolutionary model for the origin of non-randomness, long-range order and fractality in the genome. Bioessays 2001, 23, 647–656. [Google Scholar] [CrossRef]
- Lowe, C.B.; Bejerano, G.; Haussler, D. Thousands of human mobile element fragments undergo strong purifying selection near developmental genes. Proc. Natl. Acad. Sci. U. S. A. 2007, 104, 8005–8010. [Google Scholar] [CrossRef] [PubMed]
- Vaishnav, E.D.; de Boer, C.G.; Molinet, J.; Yassour, M.; Fan, L.; Adiconis, X.; Thompson, D.; Levin, J.Z.; Cubillos, F.A.; Regev, A. The evolution, evolvability and engineering of gene regulatory DNA. Nature 2022, 603, 455–463. [Google Scholar] [CrossRef] [PubMed]
- Grechishnikova, D.; Poptsova, M. Conserved 3’ UTR stem-loop structure in L1 and Alu transposons in human genome: possible role in retrotransposition. BMC Genom. 2016, 17, 992. [Google Scholar] [CrossRef] [PubMed]
- Spielmann, M.; Lupiáñez, D.G.; Mundlos, S. Structural variation in the 3D genome. Nat. Rev. Genet. 2018, 19, 453–467. [Google Scholar] [CrossRef] [PubMed]
- Rao, S.S.P.; Huang, S.C.; Glenn St Hilaire, B.; et al. Cohesin loss eliminates all loop domains. Cell 2017, 171, 305–320. [Google Scholar] [CrossRef] [PubMed]
- Metze, K.; Adam, R.; Florindo, J.B. The fractal dimension of chromatin – a potential molecular marker for carcinogenesis, tumor progression and prognosis. Expert Rev. Mol. Diagn. 2019, 19, 299–312. [Google Scholar] [CrossRef] [PubMed]
- Bedin, V.; Adam, R.L.; de Sá, B.C.; Landman, G.; Metze, K. Fractal dimension of chromatin is an independent prognostic factor for survival in melanoma. BMC Cancer 2010, 10, 260. [Google Scholar] [CrossRef] [PubMed]
- Hnisz, D.; Weintraub, A.S.; Day, D.S.; et al. Activation of proto-oncogenes by disruption of chromosome neighborhoods. Science 2016, 351, 1454–1458. [Google Scholar] [CrossRef] [PubMed]
- Amodeo, M.E.; Eyler, C.E.; Johnstone, S.E. Rewiring cancer: 3D genome determinants of cancer hallmarks. Curr. Opin. Genet. Dev. 2025, 91, 102307. [Google Scholar] [CrossRef] [PubMed]
Figure 1.
DNA strand parity for gene and LINE sequences in the human genome. A,C Rank– size plots of ordered gene and LINE sequence lengths in the Plus and Minus strands, with Minus-strand values plotted as negative Y-coordinates to visualize separately the two strands otherwise superposed. B,D Power-law-like scaling regimes of gene and LINE sequence lengths, inferred from the linearity in log–log space and quantified by linear regression line equations are reported to allow direct comparison between strands. E Rank–start-position plots of LINE start coordinate in both strands, with Plus-strand values plotted in descending order to visualize the two strands separately. F Frequency distributions of LINE start positions using 200b- bins, with Minus-strand values plotted as negative Y-coordinates for optimal visualization.
Figure 1.
DNA strand parity for gene and LINE sequences in the human genome. A,C Rank– size plots of ordered gene and LINE sequence lengths in the Plus and Minus strands, with Minus-strand values plotted as negative Y-coordinates to visualize separately the two strands otherwise superposed. B,D Power-law-like scaling regimes of gene and LINE sequence lengths, inferred from the linearity in log–log space and quantified by linear regression line equations are reported to allow direct comparison between strands. E Rank–start-position plots of LINE start coordinate in both strands, with Plus-strand values plotted in descending order to visualize the two strands separately. F Frequency distributions of LINE start positions using 200b- bins, with Minus-strand values plotted as negative Y-coordinates for optimal visualization.

Figure 2.
DNA strand parity for LINE sequence start positions across all species. A column Rank– start-position plots of LINE start coordinate in the Plus and Minus strands, with Plus-strand values plotted in descending order to visualize the two strands separately. B column Frequency distributions of LINE start positions using 200b-bins, with Minus-strand values plotted as negative Y-coordinates for optimal visualization. a: A and B data referring to Mouse, Cow, Opossum, Chicken, Xenopus; b A and B data referring to Zebrafish, Drosophila, C. elegans, Maize, Arabidopsis.
Figure 2.
DNA strand parity for LINE sequence start positions across all species. A column Rank– start-position plots of LINE start coordinate in the Plus and Minus strands, with Plus-strand values plotted in descending order to visualize the two strands separately. B column Frequency distributions of LINE start positions using 200b-bins, with Minus-strand values plotted as negative Y-coordinates for optimal visualization. a: A and B data referring to Mouse, Cow, Opossum, Chicken, Xenopus; b A and B data referring to Zebrafish, Drosophila, C. elegans, Maize, Arabidopsis.

Figure 3.
DNA strand parity for LTR, DNA, and SINE transposon sequences in the human genome. A-C-E Rank–size and B-D-F Rank–start-position plots of ordered lengths of LTR, DNA, and SINE transposons in the Plus and Minus strands, with Plus-strand values shown in descending order to visualize the two strands separately.
Figure 3.
DNA strand parity for LTR, DNA, and SINE transposon sequences in the human genome. A-C-E Rank–size and B-D-F Rank–start-position plots of ordered lengths of LTR, DNA, and SINE transposons in the Plus and Minus strands, with Plus-strand values shown in descending order to visualize the two strands separately.

Figure 4.
Gene and LINE size variation across genomes: comparison of power-law-like slopes and scaling exponents among species. Power-law-like slopes refer to values computed from linear regression on data aggregated across the two strands. A Gene size power-law-like slopes in vertebrate species, calculated using identical bin classes (15kb-bins). B LINE size power-law-like slopes in all species (200b-bins), shown for both the principal and secondary scaling regions. C Gene size scaling exponents and their standard errors, displaying an increasing trend across vertebrates, partially associated with genome complexity. C LINE size scaling exponents and their standard errors, showing a decreasing trend across species, with a clear distinction between vertebrates and non-vertebrates.
Figure 4.
Gene and LINE size variation across genomes: comparison of power-law-like slopes and scaling exponents among species. Power-law-like slopes refer to values computed from linear regression on data aggregated across the two strands. A Gene size power-law-like slopes in vertebrate species, calculated using identical bin classes (15kb-bins). B LINE size power-law-like slopes in all species (200b-bins), shown for both the principal and secondary scaling regions. C Gene size scaling exponents and their standard errors, displaying an increasing trend across vertebrates, partially associated with genome complexity. C LINE size scaling exponents and their standard errors, showing a decreasing trend across species, with a clear distinction between vertebrates and non-vertebrates.

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.