Submitted:
15 July 2026
Posted:
15 July 2026
You are already at the latest version
Abstract
Transposable elements (TEs) constitute nearly half of the human genome and shape chromatin organization, gene regulation and genome evolution. We have large gaps in understanding their influence to physiology and pathology for now. In humans, the most active elements—LINE-1 (L1), Alu, and SVA retain some copies with the ability to evade epigenetic repression and mobilize via target-primed reverse transcription (TPRT), whereas copies become inactive through fragmentation, mutation, or nesting, a process where a TE segments integrates into another TE segment. TE activity contributes to genomic instability and has been implicated in aging, cancer, neurological disorders, chromatin organization, and epigenetic regulation. Studying TEs is challenging due to their repetitive and polymorphic nature. Recent advances in sequencing technologies, including short- and long-read platforms, combined with specialized bioinformatic pipelines, now allow more comprehensive characterization of TE insertions, deletions, expression, and epigenetic status. Computational approaches vary in sensitivity, specificity, and resource requirements, and their performance is influenced by sequencing modality, coverage, and the reference genome used. Assembly-based and read-based methods, as well as tools integrating methylation or single-cell data, provide complementary insights into TE biology. Here we review the biology of active human TE. Survey state of the art short and long-read pipelines for TE analysis. And highlight their applications in studies of aging cancer and other complex diseases. We also provide practical guidance for selecting appropriate sequencing strategies and tools for TE-focused projects, and discuss emerging approaches and open questions in the field.
Keywords:
transposable elements
; LINE-1
; Alu
; SVA
; long-read sequencing
; Oxford Nanopore Technologies
; Pacific Biosciences
; bioinformatics
; epigenetics
; structural variation
Key messages:
- Transposable elements (TEs) are not merely “junk DNA”; some remain active and mobilize through target-primed reverse transcription (TPRT), while others contribute to genome regulation and evolution.
- Long-read sequencing technologies, including Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PB), provide improved resolution for repetitive genomic regions and enable characterization of complex TE insertions and structural rearrangements.
- Modern TE analysis pipelines integrate read-based, assembly-based, methylation-aware, and single-cell approaches to investigate TE insertions, expression, epigenetic regulation, and disease associations.
- Selection of an appropriate TE analysis workflow depends on sequencing modality, computational resources, biological question, and the balance between sensitivity and specificity.
1. Introduction
Repetitive sequences constitute a substantial fraction of eukaryotic genomes [1], with transposable elements (TEs) representing their primary contributors [2,3]. In humans, approximately 45% of the genome is derived from TE sequences, and some estimates suggest that up to two-thirds of the genome may originate from highly fragmented ancestral TE insertions whose origins are no longer readily identifiable [4,5].
Recent advances in high-throughput sequencing technologies, particularly long-read sequencing platforms, together with increasingly sophisticated bioinformatic pipelines, have enabled more comprehensive characterization of the TE landscape [6]. These developments provide important insights into the evolutionary history, regulation, and functional impact of TE in humans and other organisms.
Multiple TE classes have been identified across all genomes analysed to date and can be classified according to their propagation mechanisms and sequence homology [7]. A comprehensive overview of TE classification is beyond the scope of this review, which focuses primarily on active human TE and their analysis using long-read sequencing approaches. Such information can be found in the reviews of human TE [8,9] For researchers entering the field, inconsistent TE nomenclature and overlapping classification systems may present a substantial challenge. Community resources such as TE Hub [10] provide useful summaries of current classification frameworks and TE-related reviews. Table 1 summarizes the major retrotransposon families that remain active in the human genome.
Table 1.
Summary of active mobile elements in the human genome [8]. Values are approximate and may vary between reference genomes, populations, and individuals.
Table 1.
Summary of active mobile elements in the human genome [8]. Values are approximate and may vary between reference genomes, populations, and individuals.
| Name | L1 | Alu | SVA |
|---|---|---|---|
| Length (bp) | 6000 | 285 | 300–3000 |
| Copies in genome | 500,000 | 1,100,000 | 3,000–7,500 |
| Active elements | 40–50 | ~850 | 20–50 |
| Autonomy | Autonomous | Dependent on L1 | Dependent on L1 |
| Transcribed | Yes | Yes | Yes |
| Translated | ORF0, ORF1, ORF2 | No | Non-AUG translation |
| Required for transposition | ORF1p, ORF2p | ORF2p | ORF2p |
Table 2.
TE content of the HG38 and T2T reference genomes [11]. Almost half of the genome is measurably taken up by TE. Some estimates hypothesize that up to two thirds of the genome used to be TE, but became so fragmented that they no longer appear like TE [5]. These differences highlight the importance of reference selection in experiments as tiny differences can skew results when evaluated over a whole genome.
Table 2.
TE content of the HG38 and T2T reference genomes [11]. Almost half of the genome is measurably taken up by TE. Some estimates hypothesize that up to two thirds of the genome used to be TE, but became so fragmented that they no longer appear like TE [5]. These differences highlight the importance of reference selection in experiments as tiny differences can skew results when evaluated over a whole genome.
| TE | HG38% | T2T% |
| L1 | 17.36 | 16.77 |
| Other LINE | 4.07 | 3.90 |
| Alu | 10.43 | 10.09 |
| Other SINE | 2.80 | 2.68 |
| SVA | 0.15 | 0.15 |
| LTR | 9.16 | 8.84 |
| DNA transposon | 3.71 | 3.58 |
Recent telomere-to-telomere sequencing efforts and large-scale analyses of repetitive DNA suggest that most major human TE families have now been identified [11,12]. Accurate TE detectio remains challenging, especially in short-read dataset, because the high copy number and sequence similarity of TE complicate unambiguous mapping to their true genomic loci [4]. This review focuses on recent computational approaches and bioinformatic tools developed for TE analysis in human genomic datasets, with particular emphasis on long-read sequencing applications.
1.1. LINE-1
LINE-1 (L1) elements constitute approximately 17% of the human genome [3], and more than 30% of the genome reflects L1-driven activity through other retrotransposons that exploit the L1-encoded retrotransposition machinery [13,14,15]. L1 is the only currently active autonomous TE family in humans [13,14,15], with each genome containing roughly 500,000 L1 copies in various states of truncation and fragmentation. L1 mobilizes through target-primed reverse transcription (TPRT) Figure 1.
All active L1 copies belong to the L1Pa1 lineage—also referred to as L1HS (human-specific) or L1TA (transcriptionally active). A small subset of these elements, often termed “hot L1s” (30–40 copies per genome),these L1 elements accounts for approximately 80% of retrotransposition events [16]. Loss of L1 activity arises through multiple mechanisms, including 5’ truncation, point mutations [17], nesting insertions [18], premature termination of TPRT, and epigenetic repression such as hypermethylation. TPRT events frequently generate truncated L1 copies incapable of retrotransposition. On average, two human genomes differ by around 285 L1 insertions [19].
An intact L1Pa1 element is approximately 6 kb in length, comprising a 5’ untranslated region (UTR) containing an internal RNA polymerase II promoter, followed by three open reading frames (ORFs). ORF0 (213 bp) encodes a regulatory enhancer in the antisense direction [20], ORF1 (1122 bp) encodes a 40-kDa RNA-binding protein, and ORF2 (3852 bp) encodes a 150-kDa protein with endonuclease and reverse transcriptase activity. ORF1p and ORF2p are essential for L1 retrotransposition, whereas ORF0p enhances retrotransposition frequency. Non-autonomous elements such as Alu and SVA rely solely on L1-encoded ORF2p for mobilization. L1 elements terminate in a 3’ UTR enriched in adenine and thymine bases.
Current estimates place the L1 germline retrotransposition rate at approximately one new insertion per 20–200 births, while short-read whole-genome sequencing in three-generation pedigrees suggests a rate of one insertion per roughly 63 births [21]. Insertions often occur at motifs that deviate from the canonical TTAAAA L1 endonuclease target site by up to two mismatches.
Because L1 is subject to strong epigenetic repression, most retrotransposition events occur in the germline when the genome is hypomethylated. Insertions often undergo 5’ truncation through nesting, in which a new L1 copy integrates into a pre-existing element. When L1 activity bypasses host repression, somatic retrotransposition occurs [17,22,23]. Somatic insertions typically feature a truncated 5’ end, an intact 3’ poly(A) tail, and flanking target site duplications (TSDs) of approximately 15 bp generated by L1 endonuclease during integration. Many bioinformatic pipelines identify somatic insertions by their absence in matched normal tissue DNA [24], although adjacent tissues may already exhibit early molecular alterations prior to overt tumorigenesis, potentially confounding detection [25].
Figure 1.
The schematic oveview of L1 mediated TPRT. L1 RNA is transcribed, exported and traslated to produce ORF1p and ORF2p which form a ribonucleoprotein complex. The endonuclease activity or ORF2p nics a genomic target site and the resulting 3’-OH group primes reverse transcription of L1 RNA leading to integration of a new copy and in some cases, mobilization of non-autonomous elements such as Alu, SVA or processed pseudogenes [26]. Figure created with BioRender.com.
Figure 1.
The schematic oveview of L1 mediated TPRT. L1 RNA is transcribed, exported and traslated to produce ORF1p and ORF2p which form a ribonucleoprotein complex. The endonuclease activity or ORF2p nics a genomic target site and the resulting 3’-OH group primes reverse transcription of L1 RNA leading to integration of a new copy and in some cases, mobilization of non-autonomous elements such as Alu, SVA or processed pseudogenes [26]. Figure created with BioRender.com.

1.2. Alu
Alu elements are non-autonomous, primate-specific retrotransposons [27] and constitute approximately 11% of the human genome. With over one million copies present in various states of fragmentation, Alu represents the most numerous TE family by copy number. Intact Alu elements are roughly 300 bp in length, are transcribed by RNA polymerase III, and rely on L1-encoded ORF2p for reverse transcription. Although ORF2p binds L1 RNA with higher affinity than Alu RNA in the presence of ORF1p, Alu elements can undergo target-primed reverse transcription using ORF2p alone, whereas L1 elements require ORF1p. Alu likely originated from non-coding 7SL RNA, distinguishing it from most other SINEs, which generally derive from tRNA ancestors. Alu elements emerged approximately 65 million years ago, and current retrotransposition estimates range from one new insertion per approximately 20 births (phylogenetic inference) to one per approximately 40 births (short-read whole-genome sequencing) [3,21,28].
Alu retrotransposition has significantly shaped human evolution; at least 5% of alternatively spliced internal exons in the human genome originate from Alu sequences [29]. Similar to L1, Alu elements comprise hierarchical subfamilies in which older lineages have become immobile due to truncation, mutation, or host repression, allowing younger subfamilies to dominate current activity.
1.3. SVA
SVA (SINE–VNTR–Alu) elements are composite, non-autonomous retrotransposons [30] and represent the youngest active TE family in humans. Emerging approximately 6 million years ago, SVAs consist of a (CCCTCT)n hexameric tandem repeat, a reversed Alu-like region, and a GC-rich variable number tandem repeat (VNTR) domain [30]. SVA elements range from roughly 300 bp to several kilobases in length, and many copies remain retrotranspositionally competent. Their mobilization requires an intact L1 ORF2p and RNA polymerase II for transcription. SVAs occupy only approximately 0.2% of the human genome, with roughly 7,500 annotated copies. However, due to their variable-length structure, the number of full-length, potentially active SVAs remains difficult to determine accurately. Unlike other active non-autonomous TEs in humans, SVA transcription is driven by RNA polymerase II rather than RNA polymerase III [22].
2. TE Contributions to Human Health and Disease
The first TE identified in humans was an L1 insertion causing insertional mutagenesis in exon 14 of the F8 gene, resulting in hemophilia [31]. The earliest disease-causing Alu insertion was later detected in an intron of the NF1 gene, where it disrupted normal gene function and caused neurofibromatosis [32]. TE insertions within exons or promoter regions often have deleterious effects on gene expression. As Figure 1 shows TPRT promotes double stranded breaks and somatic insertions. These mechanisms imply that uncontrolled TE activity contributes to genomic instability.
Some studies revealed increased L1 ORF1 activation correlates to several types of cancer [33,34] and some assume ORF2 also correlates as it codes protein with double standed DNA clevage function but L1 orf2 is hard to measure and only recently became possible to reliably measure them [35].
In vivo studies in mice have shown that L1-mediated activation of the cGAS–STING (cyclic GMP-AMP synthase–stimulator of interferon genes) pathway can accelerate aging in cardiovascular tissue [36] There is a possibility that L1 activation can also accelarate ageing in humans.
Despite these detrimental effects, TEs also provide beneficial regulatory functions [37], shaping chromatin structure, gene expression, and genome evolution.
Higher-order chromatin organization is critical for long-range promoter–enhancer interactions and complex gene regulatory networks [38]. CTCF (CCCTC-binding factor), a conserved architectural protein, binds directly to specific repetitive elements [39], and L1 elements contribute to the formation of topologically associated domains (TADs), suggesting an ancient and conserved relationship between TEs and three-dimensional genome organization [40].
TE activity also contributes to somatic genome mosaicism in human brain tissue. For example, the hippocampus contains an average of 13.7 L1 insertions per individual [41], and it has been hypothesized that this activity underlies aspects of neural plasticity required for synapse formation. In mice embrios in vivo experiments have shown that L1 are involved in neural progenitor cell differentiation [42]. L1 elements likely also take part in human brain development. A study has shown increased L1 silencing results in reduced cerebral organoids in brain development [43].
TEs are additionally likely involved in tissue regeneration and homeostasis: dental pulp stem cell–derived osteoblasts show lower TE methylation levels [44], and in axolotls, L1 reactivation occurs at the onset of limb regeneration [45], suggesting that loosening TE repression is part of the regenerative programs.
2.1. Regulation of TE Activity
As the previous segment explained carefully controlled and regulated TE can have many beneficial effects. This could be one of the reasons they occupy such a large percentage of the human genome. On the other hand uncontrolled TEs pose a threat to genomic stability due to their potential for unchecked proliferation, illegitimate recombination, and the generation of double-stranded DNA breaks during TPRT [17,34,46,47,48,49,50]. Consequently, host organisms have evolved multiple mechanisms to suppress TE activity. Additionally, immobile TEs can be domesticated, serving as scaffolds for chromatin organization and stability.
Genome-wide screens in the K562 cancer cell line identified numerous factors modulating L1 activity [51]. Hypermethylation of CpG islands within TE promoters is a major repression mechanism observed in differentiated tissues [46,52]. Tumor suppressors, such as P53, directly inhibit L1 transcription [50], while DNA repair factors (e.g., RAD51C, BRCA1/2, RAD54L) suppress TPRT [53]. The Human Silencing Hub (HUSH) complex, comprising TASOR, MPP8, and Periphilin, mediates L1 silencing in somatic cells via methylation-dependent chromatin condensation [54]. Genome-wide CRISPR screens in K562 cells revealed that METTL3/14 can both activate L1 transcription and repress SVA transcription, whereas CTBP1 and DBR1 act as L1 repressors, suggesting potential competition between L1 and SVA at the transcriptional level [55].
The Krüppel-associated box zinc finger protein (KRAB-ZFP) family provides another major host defense mechanism. Approximately two-thirds of KRAB-ZFPs recognize multiple sequences across diverse repetitive element families (LTRs, LINEs, SINEs, SVAs, simple repeats) [56]. Upon binding, KRAB-ZFPs recruit KAP1 (TRIM28) to induce sequence-specific post-transcriptional silencing through methylation, maintaining genomic integrity. In germline cells, PIWI proteins and PIWI-interacting RNAs (piRNAs) guide cleavage of complementary TE transcripts, preventing transcription and mobilization [49]. Whether additional innate immune mechanisms exist in somatic tissues remains an open question.
These intricate layers of TE regulation underscore the need for robust computational tools capable of accurately detecting, annotating, and interpreting TE insertions across diverse human tissues and diseases.
3. Sequencing Techniques for TE Analysis
As discussed in the previous section, TEs occupy an important role in the genome and possess a complex regulatory system and could be used as prognostic or predictive biomarkers upon further understanding. This section goes over sequencing techniques and bioinformatic tools and pipelines developed in order to uncover transposable element insertions, deletions, locations, disease connections, regulations, chromatin structure and many other goals in mind to help in the planning and execution of future research.
Since the completion of the Human Genome Project [57] sequencing of the DNA have become significantly faster and cheaper and more accessible. This rapid technological leap also allowed more special experimental methods to acquire more information not deducible from Sanger sequencing. The platforms most commonly used in TE-focused sequencing studies are Illumina (IL), Oxford Nanopore Technologies (ONT), and Pacific Biosciences (PB): Illumina(IL). As many laboratories are bottlenecked by the type of sequencing platform available this section presents the most used platforms.
Currently the most widespread Next generation technology is based on sequencing by synthesis with fluorescently labelled nucleotides on the Illumina platform [58]. It has well described difficulties at properly mapping long repetitive sequences including TE [59]. The difficulties arise from the fact that Illumina reads are often shorter than 150 bp thus large repetitive sequences cannot be mapped unambiguously. This has led to a plethora of specialized techniques, tools. An independent comprehensive benchmark for TE detection in Illumina short read data have been performed [60].
For choosing sequencing protocols, reagents and informatics pipelines combined with Illumina, SequenceEng [61] could be a useful starting point. It is an interactive database of Illumina sequencing based pipelines that is very helpful in selecting an analysis method for specific task from a long list of possibilities, complete with methods and references. In the effort of making the bioinformatics of Illumina require less IT knowledge many widely used tools are bundled together in the Galaxy online platform [62]. These make the barrier to entry in TE research the lowest compared to other sequencing platforms.
Third-generation sequencing (TGS) technologies have substantially improved TE analysis by enabling long-read sequencing, real-time signal acquisition, and detection of selected epigenetic modifications. The repetitive and fragmented nature of TE creates substantial challenges for alignment and assembly algorithms [63]. Because many TE copies share near-identical sequences, short reads frequently cannot be mapped unambiguously to a single genomic locus [64]. Long-read sequencing has been particularly transformative for the analysis of active human TE families such as L1, Alu, and SVA. These elements frequently generate polymorphic insertions, truncated copies, inversions, and nested retrotransposition [18] events that are difficult to reconstruct using short-read sequencing alone. Long reads can span entire insertion loci together with their flanking genomic regions, enabling improved breakpoint resolution, more accurate genotyping, and characterization of insertion hallmarks including poly(A) tails, target-site duplications and complex structural rearrangements.
ONT platforms generate long reads typically on the range of 10 kb or more [65], with ultra-long protocols producing reads exceeding hundreds of kilobases and, in some cases, reaching megabase scale. These long reads enable spanning of entire TE insertions and other repetitive regions, reducing mapping ambiguity and improving structural variant detection. ONT sequencing operates by measuring changes in ionic current as native DNA or RNA molecules pass through nanopores, and the absence of PCR amplification also reduces amplification bias. This allows preservation and direct detection of base modifications such as DNA methylation [66], making ONT particularly valuable for studying both TE insertions and their epigenetic regulation. Drawbacks include a higher raw error rate compared with short-read platforms [67], which can be mitigated by increased coverage, consensus polishing, or hybrid assembly strategies combining ONT and Illumina data.
The Pacific Biosciences (PacBio, PB) platform employs single-molecule real-time (SMRT) sequencing to generate long reads, typically 10–25 kb, with highly accurate circular consensus reads (HiFi) reaching 99.9% accuracy [68]. PB is particularly advantageous for TE analysis because the combination of read length and accuracy allows confident detection of full-length TE insertions, structural variants, and complex rearrangements. SMRT sequencing can also indirectly detect epigenetic modifications, such as DNA methylation via polymerase kinetics [69], enabling studies on TE regulation. Limitations include higher per-sample cost and lower throughput compared with Illumina, which may be a bottleneck for large population studies.
Although ONT and PB both overcome many limitations of short-read sequencing for TE analysis, their strengths differ substantially. ONT provides substantially longer reads and direct detection of base modifications, making it particularly suitable for methylation-aware TE studies, resolving large repetitive regions, and detecting complex insertions spanning multiple kilobases [70]. In contrast, PacBio HiFi sequencing offers superior per-base accuracy, which improves breakpoint resolution and characterization of highly similar TE subfamilies. Consequently, the optimal platform depends on the primary biological question, sequencing budget, and required balance between read length, throughput, and nucleotide-level accuracy.
Despite their advantages, long-read technologies retain several limitations relevant to TE analysis. Ultra-long sequencing protocols require high molecular weight DNA and careful sample preparation, which may not be feasible for archived or degraded samples [71]. In addition, long-read datasets typically require substantially greater storage capacity and computational resources during alignment, assembly, and polishing. For ONT data, sequencing chemistry and basecalling models may also influence methylation detection and insertion accuracy, complicating reproducibility between studies.
4. Bioinformatic Pipelines for Transposable Element Analysis
As described in the previous section, TEs play important roles in genome regulation, disease, and evolution. This section reviews sequencing techniques and computational pipelines developed to detect TE insertions and deletions, assess their genomic locations, quantify expression, and explore epigenetic regulation, providing guidance for experimental design and bioinformatic analysis.
Key research and diagnostic questions include: Where are the TEs located? Does an insertion disrupt a gene? Is it positioned near regulatory elements such as exons or promoters, and could it alter gene expression? Addressing these questions requires consideration of the fundamental properties of TEs: they are repetitive, often highly fragmented, and newly inserted elements are frequently structurally incomplete or aberrant [7]. While the main premise of these tools are highly similar. Some of the differences in what type of question they answer, and what key questions they answer, where is it placed in a pipeline highlighted at Table 3. It is also worth noting that while some of these tools don’t directly answer these, it may be possible to use specialized R packages or other downstream analysis tools, measurements to answer questions like: Are these insertions near specific genes? How many of these L1 insertion are full length or have intact orf? How methylated are the insertions? Figure 2 show the framework of TE analysis pipelines and the many possible changes making comparison between studies hard.
As a vast amount of TE in the human genome is fractured, permanently immobilized, or otherwise inactive. The goal is to distinguish potentially active or polymorphic TE insertions, deletions from the large background of ancient, shared, and inactive TE-derived sequences present in the human genome. A common strategy in TE analysis is to identify sequence variants relative to a reference genome and subsequently annotate variants overlapping known TE sequences. By retaining only insertions absent from the reference genome and classified as TE-derived, the candidate search space can be substantially reduced. Although most TE detection pipelines follow this general principle, they differ considerably in how candidate insertions are identified, filtered, reconstructed, and validated.
4.1. Read-Based TE Insertion Detection
Read-based approaches operate directly on aligned sequencing reads and infer TE insertions from discordant alignments, clipped reads, split reads, or characteristic hallmarks of retrotransposition such as poly(A) tails and target-site duplications. Because these methods avoid computationally intensive genome assembly, they are typically faster and require fewer computational resources. They are particularly suitable for studies involving large cohorts or heterogeneous tumor samples where detecting low-frequency insertions may take priority over reconstructing full insertion architecture. However, increased sensitivity often comes at the expense of specificity, and read-based approaches may be more susceptible to false-positive predictions in highly repetitive genomic regions. Tools such as PALMER [72], xTea [73], sTELLeR [74], and TradetION [75] primarily follow this strategy.
4.2. Assembly-Assisted TE Reconstruction
Assembly-based approaches attempt to locally or globally reconstruct genomic sequence prior to TE annotation. By rebuilding insertion loci directly from sequencing reads, these methods can achieve improved breakpoint resolution and more accurate reconstruction of complex insertions, truncations, inversions, and nested [18] retrotransposition events. Such approaches are advantageous when validating potentially deleterious TE insertions or studying complex structural rearrangements. However, assembly quality is strongly dependent on sequencing depth and read length, and assembly-based workflows are generally computationally demanding due to the additional assembly and polishing steps. TELR [76] and TrEMOLO [77] represent examples of assembly-assisted TE detection pipelines. Notably, TrEMOLO distinguishes between “insider” insertions incorporated into the genome assembly and “outsider” insertions supported only by aligned reads, highlighting the limitations of assembly completeness in repetitive regions, and the difference between different approaches.
Table 3.
Conceptual classification of long-read transposable element detection tools according to workflow architecture, preprocessing requirements, computational complexity, and organism specificity.
Table 3.
Conceptual classification of long-read transposable element detection tools according to workflow architecture, preprocessing requirements, computational complexity, and organism specificity.
| Tool | ONT | PB | Detection strategy | Pipeline scope | Required preprocessing | Complexity | Organism scope | Key strategies / primary application |
|---|---|---|---|---|---|---|---|---|
| Alignment and read-based approaches | ||||||||
| PALMER | ✓ | ✓ | Read-based | Detection only | Alignment | Medium | Human-focused | Originally developped for L1HS focused detection. Later used as benchmark-oriented detection of human TE families (L1, Alu, SVA, HERV-K) using long-read sequencing. Detects hallmark features of retrotransposition. |
| xTea | ✓ | ✓ | Read-based | Detection + genotyping | Alignment | Medium | Adaptable | Multi-platform TE insertion detection supporting short-read and long-read sequencing with machine-learning-based genotyping. |
| sTELLeR | ✓ | ✓ | DBSCAN clustering | VCF annotation | Alignment + SV calling | Low | Adaptable | Lightweight DBSCAN-based TE insertion detection with low computational requirements and fast runtimes. |
| TradetION | ✓ | × | Read-based | Somatic TE workflow | Alignment + SV calling | Medium | Human-focused | Detection of somatic and germline TE insertions in ONT tumour-normal paired samples with TPRT hallmark detection. |
| MEIGA-PAV | ✓ | ✓ | Hybrid SV analysis | Complex rearrangement analysis | Variant calling | High | Human-focused | Detection of complex TE-associated rearrangements including inversions and potentially active L1 elements with intact ORFs. |
| Assembly-based approaches | ||||||||
| TELR | ✓ | ✓ | Assembly-based | Detection + local assembly | Alignment | High | Adaptable | High-precision reconstruction and annotation of non-reference TE insertions using local assembly and polishing. |
| TrEMOLO | ✓ | ✓ | Assembly-based | Integrated TE analysis | Assembly + alignment | High | Adaptable | Assembly-supported TE characterization.Distinguishes “insider” and “outsider” TE insertions using assembled genomes and aligned reads with graphical summaries. |
| Hybrid and end-to-end workflows | ||||||||
| GraffiTE | ✓ | ✓ | Hybrid | Full pipeline | Alignment + optional assembly | Medium–High | Adaptable | Flexible TE-associated structural variant detection and genotyping pipeline supporting batch analysis and multiple input types. |
| Retroinspector | ✓ | × | Read-based | Full workflow + visualization | Raw reads or alignment | High | Human-focused | Integrated workflow combining alignment, SV calling, TE annotation and built-in visualization/report generation. |
| TLDR | ✓ | × | Read-based | Methylation-aware detection | Alignment | Medium–High | Adaptable | Simultaneous TE insertion and methylation analysis using ONT long-read sequencing. |
4.3. Integrated End-to-End Workflows
A third category includes integrated, hybrid workflows that combine multiple analytical stages including alignment, structural variant calling, TE annotation, and visualization into unified pipelines. These integrated pipelines emphasize workflow standardization and simplified execution, although this may reduce flexibility for highly customized analyses. For example, Retroinspector performs alignment, structural variant detection, annotation, and graphical summarization within a single Snakemake workflow, whereas GraffiTE [78] combines structural variant discovery with TE-focused genotyping and annotation in a Nextflow pipeline. Such integrated pipelines can reduce technical barriers for non-specialized laboratories but may provide less flexibility for custom tailored analysis strategies.
Most current TE detection methods remain fundamentally library-based. Since the development of RepeatMasker [79], curated TE consensus databases such as Repbase [12] and Dfam [80] have formed the basis of TE annotation pipelines. These approaches rely on sequence similarity and are highly effective for identifying previously characterized TE families. However, their performance depends heavily on the completeness and quality of the underlying TE libraries. This limitation becomes particularly important in non-human organisms where TE catalogues remain incomplete. Although machine learning approaches are increasingly being explored for TE classification and insertion detection, their adoption remains substantially more limited than library-based approaches. sTELLeR [74] is one example for complementing DBSCAN machine-learning approaches with sequence similarity.
Despite the advantages of long-read sequencing, short-read sequencing data remain substantially more widely available in many research settings. Consequently, TE detection from short-read datasets remains highly relevant, particularly for retrospective analyses of existing clinical cohorts [64]. Large-scale benchmarking studies have demonstrated that specialized short-read pipelines can achieve robust TE detection performance in specific settings. For example, MELT showed strong performance in exome sequencing datasets [60], whereas xTea demonstrated improved detection accuracy in whole-genome sequencing data [73]. These comparisons highlight that optimal pipeline selection is highly dependent on experimental design, sequencing modality, biological question being addressed and the value of independent benchmarking.
Long-read sequencing has also enabled expansion of TE analysis beyond insertion detection alone. ONT-based methylation-aware sequencing allows simultaneous characterization of TE insertions and their epigenetic state. Tools such as TLDR [81] integrate insertion detection with methylation profiling, enabling investigation of whether newly inserted elements remain epigenetically repressed or transcriptionally active. For those who have chosen a different pipeline and methylation information is available, using Modkit, Methylartist [82] or other methylation analysis in downstream analysis could complement insertion data with epigenetic information.
4.4. Single-Cell TE Analysis
Similarly, single-cell sequencing approaches are beginning to address cell-type-specific TE activity and somatic mosaicism. In heterogeneous tissues such as tumors or neuronal populations, bulk sequencing may obscure low-frequency TE insertions restricted to specific cell populations. CELLO-seq [83] extends long-read sequencing to single-cell TE expression analysis, while MATES [84] applies machine learning approaches to TE quantification in single-cell datasets. These methods remain computationally demanding but represent an important future direction for understanding TE dynamics at cellular resolution.
Computational requirements vary substantially between TE analysis strategies and are influenced by factors including alignment complexity, genome assembly, variant calling, methylation analysis, and machine learning integration. Read-based heuristic filtering approaches generally require fewer computational resources than assembly-based reconstruction methods [76,77]. Likewise, workflows incorporating methylation-aware basecalling, genome polishing, or single-cell analyses introduce additional computational overhead. Importantly, the computational complexity classifications provided in Table 3 represent qualitative methodological estimates rather than direct benchmarking measurements, as runtime and memory usage remain highly dependent on sequencing depth, read length, sample number, and available hardware infrastructure.
All discussed tools are open source and typically installed from repositories through command line environments managers such as Conda, Mamba, or through the use of containers. Their use requires computational expertise, and analyses can be computationally demanding, particularly for long-read WGS data. A detailed summary of tool availability and pipeline scope, organism scope is provided in Table 3, with download links listed in Section 6.
5. Practical Considerations and Guidelines for TE-Focused Studies
5.1. Selection of Appropriate TE Analysis Pipelines
This section presents a couple of possible scenarios, research questions and discusses what would be the optimal TE pipeline for that particular situation. Table 4 Lists the different Research objectives, their key consideration and some recommended tools for that task. These recommendations are intended as general guidance. Multiple tools may be suitable for a given study depending on sequencing technology, sample quality, computational resources and desired balance between sensitivity and specificity.
Table 4.
Practical guidance table on which TE analysis tools are useful for each mentioned Research objectives mentioned in this article. It is very likely that several other questions can be answered by these tools.
Table 4.
Practical guidance table on which TE analysis tools are useful for each mentioned Research objectives mentioned in this article. It is very likely that several other questions can be answered by these tools.
| Research objective | Key requirement | Suitable approaches |
|---|---|---|
| Population cohorts | Scalability, multi platform support | xTea |
| Clinical validation | High confidence | TELR, TrEMOLO |
| Tumour-normal-pairing | Somatic comparison | TradetION |
| Using existing illumina data | Short-read support | MELT, xTea |
| Methylation analysis | Native ONT reads | TLDR |
| Single-cell analysis | Single-cell level resolution | CELLO-seq, MATES |
| Structural characterization | Complex insertions | MEIGA-PAV |
| Exploratory analysis | High sensitivity | Retroinspector, GraffiTE |
5.1.1. Retrospective Analysis of Different Cohorts
Research questions: Which non-reference TE insertions are associated with a disease across hundreds or thousands of genomes?
Methodological considerations:
Large cohorts such as the cancer genome atlas [85] often contain sequencing data generated using different platforms and library preparations. Therefore, compatibility across sequencing technologies, scalability, and standardized output become more important than reconstructing every insertion in maximal detail.
Recommended analytical strategy:
When both long read and short read samples are available tools supporting multiple sequencing technologies, such as xTea, are well suited for these analyses because they can process both Illumina and long-read datasets using a unified analytical framework. When only short read data is available MELT demonstrated strong performance in independent benchmarking of exome sequencing datasets [60].
5.1.2. Clinical Validation
Research question. Does a patient carry a pathogenic TE insertion affecting a disease-associated gene?
Methodological considerations. In clinical investigations, confidence in individual insertion calls is generally more important than processing speed. Assembly-assisted reconstruction can improve breakpoint resolution and facilitate downstream validation.
Recommended analytical strategy. Assembly-based tools such as TELR or TrEMOLO may facilitate downstream orthogonal validation owing to improved breakpoint reconstruction at the cost of higher computational cost.
5.1.3. Tumour-Normal Comparison
Research question. Which TE insertions are somatically acquired during tumor development?
Methodological considerations. The analytical workflow should distinguish germline from somatic insertions while accounting for tumour heterogeneity.
Recommended analytical strategy. TradetION was specifically developed for paired tumour-normal analyses and therefore directly addresses this experimental design.
5.1.4. Epigenetic Regulations
Research question: Are the TE insertions epigenetically silenced?
Methodological considerations: DNA methylation information must be available in addition to the sequence.
Recommended analytical strategy: TLDR was designed with this research question and directly addresses this question. Furthermore combining other tools with downstream methylation analysis with Modkit or Methylartist could also answer this.
5.1.5. Single Cell Biology
Research question: Which cell populations express transposable elements?
Difficulty: Bulk sequencing averages signal across cells. Cell level expression data is lost.
Recommended analytical strategy: Single cell approaches preserve cell identity. Single cell barcoding combined expression analysis followed up CELLO-seq or MATES are appropriate pipelines developed to solve this question.
5.1.6. Structural Characterization
Research question: Are the TE insertions truncated, inverted or associated with 3’ transduction events? Do they have intact open reading frames? Are they potentially active or aberrant?
Methodological consideration: Beyond localization, structural characterization is required.
Recommended analytical strategy MEIGA-PAV was designed to address this research question.
5.1.7. Low Tumour Purity Analysis
Research question: Where are TE insertions in a liquid biopsy sample or other low tumour purity sample?
Methodological considerations: Setting a high read filter or only accepting a TE that appears in a de novo assembly means it is unlikely to detect a TE insertion if it only appears in a tiny fraction of total reads. Assembly needs sufficient supporting reads.
Recommended analytical strategy An adjustable read based tool like Retroinspector or GraffiTE could detect an insertion that would not appear in a tool not sensitive enough like one based on de novo assembly.
No single workflow is optimal for all research objectives. Pipeline selection should primarily be driven by the biological question, sequencing modality, desired sensitivity, available computational resources, and downstream analyses rather than by overall tool popularity.
Finally it is important to remind the readers that the accuracy of TE detection depends strongly on sequencing quality. Read length, sequencing depth, N50, alignment quality, and basecalling accuracy influence both sensitivity and precision. Consequently, appropriate quality control should precede any biological interpretation.
5.2. Current Limitations of TE Analysis
Despite major advances in sequencing technologies and computational methods, several important limitations continue to constrain transposable element (TE) analysis. These limitations affect detection accuracy, reproducibility, computational requirements, and biological interpretation.
In human genomes, curated databases such as Dfam 3.9 and Repbase are generally assumed cover the vast majority of currently recognized human TE families in the era of telomere-to-telomere sequencing. However, Repbase is now under commercial licence, which may remain a practical barrier for some laboratories, institutions, especially in underdeveloped regions. In non-human organisms, incomplete TE annotation and limited availability of curated TE libraries remain substantial challenges and may reduce detection sensitivity [86,87] Consequently, researchers must carefully select the TE consensus libraries used during analysis, particularly when using tools restricted to predefined TE families. TE insertions absent from the selected database cannot be detected. Conversely, tools hardwired to broad databases, such as Retroinspector with the Dfam 3.9 human library, may recover large numbers of inactive or highly fragmented TE copies that are not biologically relevant for a given study, increasing downstream filtering requirements and computational burden.
Another important limitation is that many TE analysis pipelines were developed for highly specific research questions. Default parameters are therefore not universally applicable, as hyperparameters such as minimum read support, alignment thresholds, or structural variant filtering criteria can substantially influence sensitivity and precision depending on sequencing depth, tumor purity, and experimental design [73,74,88]. For example, highly heterogeneous tumour samples or samples with low tumour purity may require relaxed filtering thresholds to recover low-frequency somatic insertions, whereas clinical diagnostic workflows may prioritize specificity and reproducibility over maximal sensitivity. Sensitivity estimates reported by different tools are often difficult to compare directly because benchmarking datasets, sequencing depth, read lengths, and validation criteria vary substantially between studies.
A major unresolved issue in the field is the lack of comprehensive independent benchmarking studies across recently developed tools. Existing evaluations are often performed by the developers themselves and commonly compare new methods only against older approaches rather than against the most recent alternatives [73,74,78]. A recent preprint has begun benchmarking several long-read TE detection tools, including PALMER, TLDR, sTELLeR, and xTea [88]; however, the rapidly expanding number of pipelines highlights the need for larger standardized benchmarking efforts. Such studies would improve reproducibility, clarify optimal use cases for each tool, and facilitate more objective selection of analysis workflows.
Computational reproducibility also remains challenging. Although most TE analysis tools are publicly available and open source, practical implementation often requires substantial bioinformatic expertise. Installation may be complicated by dependency conflicts, outdated packages, variable documentation quality, or incompatibilities between software versions. In addition, many pipelines rely on whole-genome sequencing data and therefore require substantial storage capacity, memory allocation, and wall-clock runtime. These technical barriers may limit adoption in smaller laboratories lacking dedicated computational infrastructure.
Reference genome selection introduces an additional source of bias. As illustrated in Table 2, substantial differences exist between commonly used human reference genomes such as GRCh38 and T2T assemblies. TE insertions identified relative to one reference genome may represent common polymorphisms absent from another assembly rather than novel insertion events. This issue is particularly relevant in underrepresented populations, where population-specific germline variants may be incorrectly interpreted as rare or disease-associated insertions if they are not represented in the reference genome.
6. Future Directions and Open Questions
Transposable elements have profoundly shaped the human genome, participating in an ongoing evolutionary dynamic as hosts evolve mechanisms to suppress their activity. Although most TE subfamilies have lost their ability to mobilize and now persist as genomic fossils, their continued presence underscores their biological significance.
Understanding the contributions of currently active and replicating TE requires a detailed investigation on their role in disease. Alu and L1 ORF1p [27,47] transcript levels correlate with disease severity and may serve as biomarkers for poor prognosis. These observations suggest that TE activation may contribute to disease progression, although causality remains incompletely established.
The broader adoption of long-read sequencing, along with improved bioinformatic pipelines and experimental approaches specifically designed for WGS and TE detection, is expected to improve our understanding of the mechanistic links between TEs and human diseases in the near future. Integrating long-read sequencing, functional studies, and standardized analysis pipelines will be essential to fully understand TE biology and their roles in human health and disease.
At the time of writing, many of the listed TE analysis tools reviewed here have not been independently benchmarked. Such benchmarking could establish community standards for TE insertion detection, facilitate reproducibility, and provide developers with critical feedback for improving and harmonizing analysis pipelines.
Future methodological developments are likely to include pangenome references [89], graph genomes, standardized benchmark datasets, improved TE annotations, and increasingly integrated long-read sequencing workflows. Direct RNA sequencing [90], Advances in machine learning approaches, T2T references [11], single cell multiomics. These improvements could slowly evolve the TE research field from answering simple questions like: Where are TE insertions located? towards answering complex questions regarding TE: Are TE involved in ageing, if yes how? What mechanisms explain the association between L1 expression and cancer? Can TE activity be therapeutically targeted?
Future pipelines will likely integrate insertion detection, methylation profiling, transcriptomic analysis, and structural characterization into unified workflows rather than treating these analyses independently. Prospective clinical studies will be required to determine whether TE insertions, expression profiles, or methylation signatures can serve as reliable diagnostic, prognostic, or predictive biomarkers.
Author Contributions
D.V and B.M conceived the review idea and wrote the manuscript N.S and A.K contributed their previous expertise in TE methylation and its analysis I.T reviewed the article.
Funding
This study was funded by the Hungarian Scientific Research Funds NRDI-FK0201NEPE/TKPNKTA-47 and RRF-2.3.1-21-2022-00003.
Data Availability Statement
The source codes for all bioinformatic tools discussed in this review are publicly available at the following repositories:
- Retroinspector: https://github.com/javiercguard/retroinspector
- TradetION: https://github.com/panummi/TraDetIONS
- GraffiTE: https://github.com/cgroza/GraffiTE
- MEIGA-PAV: https://github.com/MEIGA-tk/MEIGA-PAV
- CELLO-seq: https://github.com/MarioniLab/CELLOseq
Conflicts of Interest
The authors declare no conflicts of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| TE | Transposable Element |
| L1 | long interspersed nuclear element-1 |
| SVA | SINE-VNTR-Alu |
| TPRT | target-primed reverse transcription |
| ONT | Oxford Nanopore Technologies |
| PB | Pacific Biosciences |
| WGS | Whole Genome Sequencing |
References
- Liehr, T. Repetitive elements in humans. Int. J. Mol. Sci. 2021, 22, 2072. [Google Scholar] [CrossRef] [PubMed]
- Mangiavacchi, A.; Liu, P.; Della Valle, F.; Orlando, V. New insights into the functional role of retrotransposon dynamics in mammalian somatic cells. Cell. Mol. Life Sci. 2021, 78, 5245–5256. [Google Scholar] [CrossRef] [PubMed]
- Hancks, D.C.; Kazazian, H.H. Active human retrotransposons: Variation and disease. Curr. Opin. Genet. Dev. 2012, 22, 191–203. [Google Scholar] [CrossRef] [PubMed]
- Lee, E.; Iskow, R.; Yang, L.; Gokcumen, O.; Haseley, P.; Luquette, L.J.; Lohr, J.G.; Harris, C.C.; Ding, L.; Wilson, R.K.; et al. Landscape of somatic retrotransposition in human cancers. Science 2012, 337, 967–971. [Google Scholar] [CrossRef] [PubMed]
- de Koning, A.P.; Gu, W.; Castoe, T.A.; Batzer, M.A.; Pollock, D.D. Repetitive elements may comprise over two-thirds of the human genome. PLoS Genet. 2011, 7. [Google Scholar] [CrossRef] [PubMed]
- Szakállas, N.; Kalmár, A.; Rada, K.R.; Kucarov, M.; Linkner, T.R.; Barták, B.K.; Takács, I.; Molnár, B. Methodological comparison of short-read and long-read sequencing methods on colorectal cancer samples. Int. J. Mol. Sci. 2025, 26, 9254. [Google Scholar] [CrossRef] [PubMed]
- Snowbarger, J.; Koganti, P.; Spruck, C. Evolution of repetitive elements, their roles in homeostasis and human disease, and potential therapeutic applications. Biomolecules 2024, 14, 1250. [Google Scholar] [CrossRef] [PubMed]
- Shan, M.; Fatima, A.; Amjad, Z.; Khan, Z. Transposable elements and epigenetic regulation. Plant Transposable Elem. 2023, 147–178. [Google Scholar] [CrossRef]
- Wells, J.N.; Feschotte, C. A field guide to eukaryotic transposable elements. Annu. Rev. Genet. 2020, 54, 539–561. [Google Scholar] [CrossRef] [PubMed]
- Elliott, T.A.; Heitkam, T.; Hubley, R.; Quesneville, H.; Suh, A.; Wheeler, T.J. TE Hub: A Community-Oriented Space for Sharing and Connecting Tools, Data, Resources, and Methods for Transposable Element Annotation. Mobile DNA 12, 16. [CrossRef] [PubMed]
- Hoyt, S.J.; Storer, J.M.; Hartley, G.A.; Grady, P.G.; Gershman, A.; de Lima, L.G.; Limouse, C.; Halabian, R.; Wojenski, L.; Rodriguez, M.; et al. From telomere to telomere: The transcriptional and epigenetic state of human repeat elements. Science 2022, 376. [Google Scholar] [CrossRef] [PubMed]
- Kojima, K.K. Human transposable elements in Repbase: Genomic footprints from fish to humans. Mob. DNA 2018, 9. [Google Scholar] [CrossRef] [PubMed]
- Jurka, J. Subfamily Structure and evolution of the human L1 family of repetive sequences. J. Mol. Evol. 1989, 29, 496–503. [Google Scholar] [CrossRef] [PubMed]
- Beck, C.R.; Collier, P.; Macfarlane, C.; Malig, M.; Kidd, J.M.; Eichler, E.E.; Badge, R.M.; Moran, J.V. Line-1 retrotransposition activity in human genomes. Cell 2010, 141, 1159–1170. [Google Scholar] [CrossRef] [PubMed]
- Wang, P.J. Tracking Line1 retrotransposition in the Germline. Proc. Natl. Acad. Sci. 2017, 114, 7194–7196. [Google Scholar] [CrossRef] [PubMed]
- Brouha, B.; Schustak, J.; Badge, R.M.; Lutz-Prigge, S.; Farley, A.H.; Moran, J.V.; Kazazian, H.H. Hot l1s account for the bulk of retrotransposition in the human population. Proc. Natl. Acad. Sci. 2003, 100, 5280–5285. [Google Scholar] [CrossRef] [PubMed]
- Lee, E.; Iskow, R.; Yang, L.; Gokcumen, O.; Haseley, P.; Luquette, L.J.; Lohr, J.G.; Harris, C.C.; Ding, L.; Wilson, R.K.; et al. Landscape of somatic retrotransposition in human cancers. Science 2012, 337, 967–971. [Google Scholar] [CrossRef] [PubMed]
- SanMiguel, P.; Tikhonov, A.; Jin, Y.K.; Motchoulskaia, N.; Zakharov, D.; Melake-Berhan, A.; Springer, P.S.; Edwards, K.J.; Lee, M.; Avramova, Z.; et al. Nested retrotransposons in the intergenic regions of the maize genome. Science 1996, 274, 765–768. [Google Scholar] [CrossRef] [PubMed]
- Lanciano, S.; Philippe, C.; Sarkar, A.; Pratella, D.; Domrane, C.; Doucet, A.J.; van Essen, D.; Saccani, S.; Ferry, L.; Defossez, P.A.; et al. Locus-level L1 DNA methylation profiling reveals the epigenetic and transcriptional interplay between L1s and their integration sites. Cell Genom. 2024, 4, 100498. [Google Scholar] [CrossRef] [PubMed]
- Denli, A.M.; Narvaiza, I.; Kerman, B.E.; Pena, M.; Benner, C.; Marchetto, M.C.; Diedrich, J.K.; Aslanian, A.; Ma, J.; Moresco, J.J.; et al. Primate-Specific ORF0 Contributes to Retrotransposon-Mediated Diversity. Cell 2015, 163, 583–593. [Google Scholar] [CrossRef] [PubMed]
- Feusier, J.; Watkins, W.S.; Thomas, J.; Farrell, A.; Witherspoon, D.J.; Baird, L.; Ha, H.; Xing, J.; Jorde, L.B. Pedigree-based estimation of human mobile element retrotransposition rates. Genome Res. 2019, 29, 1567–1577. [Google Scholar] [CrossRef] [PubMed]
- Helman, E.; Lawrence, M.S.; Stewart, C.; Sougnez, C.; Getz, G.; Meyerson, M. Somatic retrotransposition in human cancer revealed by whole-genome and exome sequencing. Genome Res. 2014, 24, 1053–1063. [Google Scholar] [CrossRef] [PubMed]
- Belancio, V.P.; Deininger, P.L.; Roy-Engel, A.M. Line dancing in the human genome: Transposable elements and disease. Genome Med. 2009, 1, 97. [Google Scholar] [CrossRef] [PubMed]
- Goerner-Potvin, P.; Bourque, G. Computational tools to unmask transposable elements. Nat. Rev. Genet. 2018, 19, 688–704. [Google Scholar] [CrossRef] [PubMed]
- Braakhuis, B.J.; Leemans, C.R.; Brakenhoff, R.H. Using tissue adjacent to carcinoma as a normal control: An obvious but questionable practice. J. Pathol. 2004, 203, 620–621. [Google Scholar] [CrossRef] [PubMed]
- Esnault, C.; Maestre, J.; Heidmann, T. Human LINE Retrotransposons Generate Processed Pseudogenes. Nat. Genet. 24, 363–367. [CrossRef] [PubMed]
- Ade, C.; Roy-Engel, A.M.; Deininger, P.L. Alu elements: An intrinsic source of human genome instability. Curr. Opin. Virol. 2013, 3, 639–645. [Google Scholar] [CrossRef] [PubMed]
- Bennett, E.A.; Keller, H.; Mills, R.E.; Schmidt, S.; Moran, J.V.; Weichenrieder, O.; Devine, S.E. Active alu retrotransposons in the human genome. Genome Res. 2008, 18, 1875–1883. [Google Scholar] [CrossRef] [PubMed]
- Sorek, R.; Lev-Maor, G.; Reznik, M.; Dagan, T.; Belinky, F.; Graur, D.; Ast, G. Minimal conditions for exonization of intronic sequences. Mol. Cell 2004, 14, 221–231. [Google Scholar] [CrossRef] [PubMed]
- Chu, C.; Lin, E.W.; Tran, A.; Jin, H.; Ho, N.I.; Veit, A.; Cortes-Ciriano, I.; Burns, K.H.; Ting, D.T.; Park, P.J. The landscape of human SVA retrotransposons. Nucleic Acids Res. 2023, 51, 11453–11465. [Google Scholar] [CrossRef] [PubMed]
- Kazazian, H.H.; Wong, C.; Youssoufian, H.; Scott, A.F.; Phillips, D.G.; Antonarakis, S.E. Haemophilia a resulting from de novo insertion of L1 sequences represents a novel mechanism for mutation in man. Nature 1988, 332, 164–166. [Google Scholar] [CrossRef] [PubMed]
- Wallace, M.R.; Andersen, L.B.; Saulino, A.M.; Gregory, P.E.; Glover, T.W.; Collins, F.S. A de novo alu insertion results in neurofibromatosis type 1. Nature 1991, 353, 864–866. [Google Scholar] [CrossRef] [PubMed]
- Pradhan, B.; Cajuso, T.; Katainen, R.; Sulo, P.; Tanskanen, T.; Kilpivaara, O.; Pitkänen, E.; Aaltonen, L.A.; Kauppi, L.; Palin, K. Detection of subclonal L1 transductions in colorectal cancer by long-distance inverse-PCR and nanopore sequencing. Sci. Rep. 2017, 7. [Google Scholar] [CrossRef] [PubMed]
- Scott, E.; Devine, S. The role of somatic L1 retrotransposition in human cancers. Viruses 2017, 9, 131. [Google Scholar] [CrossRef] [PubMed]
- Nielsen, M.I.; Wolters, J.C.; Bringas, O.G.R.; Jiang, H.; Di Stefano, L.H.; Oghbaie, M.; Hozeifi, S.; Nitert, M.J.; van Pijkeren, A.; Smit, M.; et al. Targeted Detection of Endogenous LINE-1 Proteins and ORF2p Interactions. Mob. DNA 2025, 16, 3. [Google Scholar] [CrossRef] [PubMed]
- Yang, C.; Du, H.; Liu, S.; Xu, P.; Wang, Y.; Zhou, Y.; Yuan, H.; Li, Y.; Shen, J.; Yuan, X.; et al. Targeting age-related line-1 activation alleviates cardiac aging. Nat. Aging 2026, 6, 414–429. [Google Scholar] [CrossRef] [PubMed]
- Ilık, E.A.; Yang, X.; Zhang, Z.Z.; Aktaş, T. Transcriptional and post-transcriptional regulation of transposable elements and their roles in development and disease. Nat. Rev. Mol. Cell Biol. 2025. [Google Scholar] [CrossRef] [PubMed]
- Bonev, B.; Cavalli, G. Organization and function of the 3D genome. Nat. Rev. Genet. 2016, 17, 661–678. [Google Scholar] [CrossRef] [PubMed]
- Schmidt, D.; Schwalie, P.; Wilson, M.; Ballester, B.; Gonçalves, A.; Kutter, C.; Brown, G.; Marshall, A.; Flicek, P.; Odom, D. Waves of retrotransposon expansion remodel genome organization and CTCF binding in multiple mammalian lineages. Cell 2012, 148, 832. [Google Scholar] [CrossRef]
- Hong, Y.; Bie, L.; Zhang, T.; Yan, X.; Jin, G.; Chen, Z.; Wang, Y.; Li, X.; Pei, G.; Zhang, Y.; et al. SAFB Restricts Contact Domain Boundaries Associated with L1 Chimeric Transcription. Mol. Cell 2024, 84, 1637–1650.e10. [Google Scholar] [CrossRef] [PubMed]
- Upton, K.; Gerhardt, D.; Jesuadian, J.; Richardson, S.; Sánchez-Luque, F.; Bodea, G.; Ewing, A.; Salvador-Palomeque, C.; vanderKnaap, M.; Brennan, P.; et al. Ubiquitous L1 mosaicism in hippocampal neurons. Cell 2015, 161, 228–239. [Google Scholar] [CrossRef] [PubMed]
- Mangoni, D.; Simi, A.; Lau, P.; Armaos, A.; Ansaloni, F.; Codino, A.; Damiani, D.; Floreani, L.; Di Carlo, V.; Vozzi, D.; et al. LINE-1 Regulates Cortical Development by Acting as Long Non-Coding RNAs. Nat. Commun. 2023, 14, 4974. [Google Scholar] [CrossRef] [PubMed]
- Garza, R.; Atacho, D.A.M.; Adami, A.; Gerdes, P.; Vinod, M.; Hsieh, P.; Karlsson, O.; Horvath, V.; Johansson, P.A.; Pandiloski, N.; et al. LINE-1 Retrotransposons Drive Human Neuronal Transcriptome Complexity and Functional Diversification. Sci. Adv. 9, eadh9543. [CrossRef] [PubMed]
- Prucksakorn, T.; Mutirangura, A.; Pavasant, P.; Subbalekha, K. Altered methylation levels in line-1 in dental pulp stem cell–derived osteoblasts. Int. Dent. J. 2025, 75, 1269–1276. [Google Scholar] [CrossRef] [PubMed]
- Zhu, W.; Kuo, D.; Nathanson, J.; Satoh, A.; Pao, G.M.; Yeo, G.W.; Bryant, S.V.; Voss, S.R.; Gardiner, D.M.; Hunter, T. Retrotransposon long interspersed nucleotide element-1 (line-1) is activated during salamander limb regeneration. Dev. Growth Differ. 2012, 54, 673–685. [Google Scholar] [CrossRef] [PubMed]
- Chuong, E.B.; Elde, N.C.; Feschotte, C. Regulatory activities of transposable elements: from conflicts to benefits. Nat. Rev. Genet. 2017, 18, 71–86. [Google Scholar] [CrossRef] [PubMed]
- Ardeljan, D.; Taylor, M.S.; Ting, D.T.; Burns, K.H. The Human Long Interspersed Element-1 Retrotransposon: An Emerging Biomarker of Neoplasia. Clin. Chem. 2017, 63, 816–822. [Google Scholar] [CrossRef] [PubMed]
- Burns, K.H. Transposable elements in cancer. Nat. Rev. Cancer 2017, 17, 415–424. [Google Scholar] [CrossRef] [PubMed]
- Ozata, D.M.; Gainetdinov, I.; Zoch, A.; O’Carroll, D.; Zamore, P.D. Piwi-interacting RNAS: Small RNAS with big functions. Nat. Rev. Genet. 2018, 20, 89–108. [Google Scholar] [CrossRef] [PubMed]
- Tiwari, B.; Jones, A.E.; Caillet, C.J.; Das, S.; Royer, S.K.; Abrams, J.M. P53 directly represses human LINE1 transposons. Genes Dev. 2020, 34, 1439–1451. [Google Scholar] [CrossRef] [PubMed]
- Liu, N.; Lee, C.H.; Swigut, T.; Grow, E.; Gu, B.; Bassik, M.C.; Wysocka, J. Selective silencing of euchromatic l1s revealed by genome-wide screens for L1 regulators. Nature 2017, 553, 228–232. [Google Scholar] [CrossRef] [PubMed]
- Weisenberger, D.J. Analysis of repetitive element DNA methylation by methylight. Nucleic Acids Res. 2005, 33, 6823–6836. [Google Scholar] [CrossRef] [PubMed]
- Bona, N.; Crossan, G.P. Fanconi anemia DNA crosslink repair factors protect against line-1 retrotransposition during Mouse Development. Nat. Struct. Mol. Biol. 2023, 30, 1434–1445. [Google Scholar] [CrossRef] [PubMed]
- Tunbak, H.; Enriquez-Gasca, R.; Tie, C.H.; Gould, P.A.; Mlcochova, P.; Gupta, R.K.; Fernandes, L.; Holt, J.; van der Veen, A.G.; Giampazolias, E.; et al. The Hush Complex is a gatekeeper of type I interferon through epigenetic regulation of line-1s. Nat. Commun. 2020, 11. [Google Scholar] [CrossRef] [PubMed]
- Li, X.; Bie, L.; Wang, Y.; Hong, Y.; Zhou, Z.; Fan, Y.; Yan, X.; Tao, Y.; Huang, C.; Zhang, Y.; et al. Line-1 transcription activates long-range gene expression. Nat. Genet. 2024, 56, 1494–1502. [Google Scholar] [CrossRef] [PubMed]
- Yang, P.; Wang, Y.; Macfarlan, T.S. The role of Krab-ZFPs in transposable element repression and mammalian evolution. Trends Genet. 2017, 33, 871–881. [Google Scholar] [CrossRef] [PubMed]
- Powledge, T.M. Human Genome Project Completed. Genome Biol. 2003, 4, spotlight–20030415–01. [Google Scholar] [CrossRef]
- Uhlen, M.; Quake, S.R. Sequential Sequencing by Synthesis and the Next-Generation Sequencing Revolution. Trends Biotechnol. 2023, 41, 1565–1572. [Google Scholar] [CrossRef] [PubMed]
- Witherspoon, D.J.; Zhang, Y.; Xing, J.; Watkins, W.S.; Ha, H.; Batzer, M.A.; Jorde, L.B. Mobile Element Scanning (ME-Scan) Identifies Thousands of Novel Alu Insertions in Diverse Human Populations. Genome Res. 2013, 23, 1170–1181. [Google Scholar] [CrossRef] [PubMed]
- Wijngaard, R.; Demidov, G.; O’Gorman, L.; Corominas-Galbany, J.; Yaldiz, B.; Steyaert, W.; De Boer, E.; Vissers, L.E.L.M.; Kamsteeg, E.J.; Pfundt, R.; et al. Mobile Element Insertions in Rare Diseases: A Comparative Benchmark and Reanalysis of 60,000 Exome Samples. Eur. J. Hum. Genet. 2024, 32, 200–208. [Google Scholar] [CrossRef] [PubMed]
- Zhang, Y.; Manjunath, M.; Kim, Y.; Heintz, J.; Song, J.S. SequencEnG: An Interactive Knowledge Base of Sequencing Techniques. Bioinformatics 2019, 35, 1438–1440. [Google Scholar] [CrossRef] [PubMed]
- Jalili, V.; Afgan, E.; Gu, Q.; Clements, D.; Blankenberg, D.; Goecks, J.; Taylor, J.; Nekrutenko, A. The Galaxy Platform for Accessible, Reproducible and Collaborative Biomedical Analyses: 2020 Update. Nucleic Acids Res. 2020, 48, 8205–8207. [Google Scholar] [CrossRef] [PubMed]
- Snowbarger, J.; Koganti, P.; Spruck, C. Evolution of repetitive elements, their roles in homeostasis and human disease, and potential therapeutic applications. Biomolecules 2024, 14, 1250. [Google Scholar] [CrossRef] [PubMed]
- Keane, T.M.; Wong, K.; Adams, D.J. RetroSeq: Transposable Element Discovery from next-Generation Sequencing Data. Bioinformatics 2013, 29, 389–390. [Google Scholar] [CrossRef] [PubMed]
- Delahaye, C.; Nicolas, J. Sequencing DNA with Nanopores: Troubles and Biases. PLoS ONE 2021, 16, e0257521. [Google Scholar] [CrossRef] [PubMed]
- Rand, A.C.; Jain, M.; Eizenga, J.M.; Musselman-Brown, A.; Olsen, H.E.; Akeson, M.; Paten, B. Mapping DNA Methylation with High-Throughput Nanopore Sequencing. Nat. Methods 2017, 14, 411–413. [Google Scholar] [CrossRef] [PubMed]
- Rang, F.J.; Kloosterman, W.P.; De Ridder, J. From Squiggle to Basepair: Computational Approaches for Improving Nanopore Sequencing Read Accuracy. Genome Biol. 2018, 19, 90. [Google Scholar] [CrossRef] [PubMed]
- Veselovsky, V.; Romanov, M.; Zoruk, P.; Larin, A.; Babenko, V.; Morozov, M.; Strokach, A.; Zakharevich, N.; Khamidova, S.; Danilova, A.; et al. Comparative Evaluation of Sequencing Platforms: Pacific Biosciences, Oxford Nanopore Technologies, and Illumina for 16S rRNA-based Soil Microbiome Profiling. Front. Microbiol. 2025, 16, 1633360. [Google Scholar] [CrossRef] [PubMed]
- Yang, Y.; Scott, S.A. DNA Methylation Profiling Using Long-Read Single Molecule Real-Time Bisulfite Sequencing (SMRT-BS). In Functional Genomics; Kaufmann, M., Klinger, C., Savelsbergh, A., Eds.; Springer New York: New York, NY, 2017; Vol. 1654, pp. 125–134. [Google Scholar] [CrossRef] [PubMed]
- Logsdon, G.A.; Vollger, M.R.; Eichler, E.E. Long-Read Human Genome Sequencing and Its Applications. Nat. Rev. Genet. 21, 597–614. [CrossRef] [PubMed]
- Amarasinghe, S.L.; Su, S.; Dong, X.; Zappia, L.; Ritchie, M.E.; Gouil, Q. Opportunities and Challenges in Long-Read Sequencing Data Analysis. Genome Biol. 21, 30. [CrossRef] [PubMed]
- Zhou, W.; Emery, S.B.; Flasch, D.A.; Wang, Y.; Kwan, K.Y.; Kidd, J.M.; Moran, J.V.; Mills, R.E. Identification and characterization of occult human-specific line-1 insertions using long-read sequencing technology. Nucleic Acids Res. 2019, 48, 1146–1163. [Google Scholar] [CrossRef] [PubMed]
- Chu, C.; Borges-Monroy, R.; Viswanadham, V.V.; Lee, S.; Li, H.; Lee, E.A.; Park, P.J. Comprehensive identification of transposable element insertions using multiple sequencing technologies. Nat. Commun. 2021, 12. [Google Scholar] [CrossRef] [PubMed]
- Bilgrav Saether, K.; Eisfeldt, J. Detecting transposable elements in long-read genomes using Steller. Bioinformatics 2024, 40. [Google Scholar] [CrossRef] [PubMed]
- Nummi, P.; Cajuso, T.; Norri, T.; Taira, A.; Kuisma, H.; Välimäki, N.; Lepistö, A.; Renkonen-Sinisalo, L.; Koskensalo, S.; Seppälä, T.T.; et al. Structural features of somatic and germline retrotransposition events in humans. Mob. DNA 2025, 16. [Google Scholar] [CrossRef] [PubMed]
- Han, S.; Dias, G.B.; Basting, P.J.; Viswanatha, R.; Perrimon, N.; Bergman, C. Local Assembly of long reads enables phylogenomics of transposable elements in a polyploid cell line. Nucleic Acids Res. 2022, 50. [Google Scholar] [CrossRef] [PubMed]
- Mohamed, M.; Sabot, F.; Varoqui, M.; Mugat, B.; Audouin, K.; Pélisson, A.; Fiston-Lavier, A.S.; Chambeyron, S. Tremolo: Accurate transposable element allele frequency estimation using long-read sequencing data combining assembly and mapping-based approaches. Genome Biol. 2023, 24. [Google Scholar] [CrossRef] [PubMed]
- Groza, C.; Chen, X.; Wheeler, T.J.; Bourque, G.; Goubert, C. A unified framework to analyze transposable element insertion polymorphisms using graph genomes. Nat. Commun. 2024, 15. [Google Scholar] [CrossRef] [PubMed]
- Hubley, S. Repeatmasker Home Page, 2013.
- Storer, J.; Hubley, R.; Rosen, J.; Wheeler, T.J.; Smit, A.F. The DFAM community resource of transposable element families, sequence models, and genome annotations. Mob. DNA 2021, 12. [Google Scholar] [CrossRef] [PubMed]
- Ewing, A.D.; Smits, N.; Sanchez-Luque, F.J.; Faivre, J.; Brennan, P.M.; Richardson, S.R.; Cheetham, S.W.; Faulkner, G.J. Nanopore sequencing enables comprehensive transposable element epigenomic profiling. Mol. Cell 2020, 80. [Google Scholar] [CrossRef] [PubMed]
- Cheetham, S.W.; Kindlova, M.; Ewing, A.D. Methylartist: Tools for Visualizing Modified Bases from Nanopore Sequence Data. Bioinformatics 2022, 38, 3109–3112. [Google Scholar] [CrossRef] [PubMed]
- Berrens, R.V.; Yang, A.; Laumer, C.E.; Lun, A.T.L.; Bieberich, F.; Law, C.T.; Lan, G.; Imaz, M.; Bowness, J.S.; Brockdorff, N.; et al. Locus-Specific Expression of Transposable Elements in Single Cells with CELLO-seq. Nat. Biotechnol. 2022, 40, 546–554. [Google Scholar] [CrossRef] [PubMed]
- Wang, R.; Zheng, Y.; Zhang, Z.; Song, K.; Wu, E.; Zhu, X.; Wu, T.P.; Ding, J. MATES: A Deep Learning-Based Model for Locus-Specific Quantification of Transposable Elements in Single Cell. Nat. Commun. 2024, 15, 8798. [Google Scholar] [CrossRef] [PubMed]
- The Cancer Genome Atlas Research Network; Weinstein, J.N.; Collisson, E.A.; Mills, G.B.; Shaw, K.R.M.; Ozenberger, B.A.; Ellrott, K.; Shmulevich, I.; Sander, C.; Stuart, J.M. The Cancer Genome Atlas Pan-Cancer Analysis Project. Nat. Genet. 2013, 45, 1113–1120. [Google Scholar] [CrossRef] [PubMed]
- Ou, S.; Su, W.; Liao, Y.; Chougule, K.; Agda, J.R.A.; Hellinga, A.J.; Lugo, C.S.B.; Elliott, T.A.; Ware, D.; Peterson, T.; et al. Benchmarking Transposable Element Annotation Methods for Creation of a Streamlined, Comprehensive Pipeline. Genome Biol. 2019, 20, 275. [Google Scholar] [CrossRef] [PubMed]
- RepeatModeler2 for Automated Genomic Discovery of Transposable Element Families | PNAS. Available online: https://www.pnas.org/doi/abs/10.1073/pnas.1921046117.
- Seymen, N.; Santos, R.; Lakshmanan, R.; Topp, S.; Al-Chalabi, A.; Al Khleifat, A.; Breen, G.; Dobson, R.J.; Quinn, J.P.; Karimi, M.M.; et al. Systematic Comparative Benchmarking of Computational Methods for the Detection of Transposable Elements in Long-Read Sequencing Data. 2025. [Google Scholar] [CrossRef]
- Wang, T.; Antonacci-Fulton, L.; Howe, K.; Lawson, H.A.; Lucas, J.K.; Phillippy, A.M.; Popejoy, A.B.; Asri, M.; Carson, C.; Chaisson, M.J.P.; et al. The Human Pangenome Project: A Global Resource to Map Genomic Diversity. Nature 2022, 604, 437–446. [Google Scholar] [CrossRef] [PubMed]
- Garalde, D.R.; Snell, E.A.; Jachimowicz, D.; Sipos, B.; Lloyd, J.H.; Bruce, M.; Pantic, N.; Admassu, T.; James, P.; Warland, A.; et al. Highly Parallel Direct RNA Sequencing on an Array of Nanopores. Nat. Methods 2018, 15, 201–206. [Google Scholar] [CrossRef] [PubMed]
Figure 2.
Conceptual framework for transposable element analysis using modern sequencing technologies. Showing the major analytical decision points in TE analysis pipelines.
Figure 2.
Conceptual framework for transposable element analysis using modern sequencing technologies. Showing the major analytical decision points in TE analysis pipelines.

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.