Preprint
Review

This version is not peer-reviewed.

Comprehensive Survey and Guidance for Building High Quality Genome Assembly

Submitted:

02 July 2026

Posted:

03 July 2026

You are already at the latest version

Abstract
Genome sequence information is the primordial need for studying species genetics, evolutionary history, disease mechanisms, risk prediction, adaptation and many more. Eventually large-scale initiatives are underway to sequence unknown genomes from various species including humans with various phenotypic states to gain insights into gene function and genetic diversity. However, assembling large genomes remains a computational challenge due to several factors including sequence complexity, continuous growth in sequencing throughput, lack of suitable benchmarking for the selection of optimal combination of tools etc. Additionally, its quality evaluation is another crucial step to perform but varying methods often lead to arbitrary comparisons. Moreover, high sequencing error rates necessitate the error correction and consensus sequence generation steps in genome assembly. To rectify these sequencing errors several polishing tools are already in use but their in-depth survey is still lacking. Appropriate selection of tools can enhance consensus quality and can generate precise assembly. Hence, a comprehensive survey is always beneficial to build a pipeline incorporating multiple evaluation indicators, including contiguity, accuracy, completeness, and contamination along with a proper guidance to select optimal tool and follow the right steps to achieve more accurate and complete genome assembly.
Keywords: 
;  ;  ;  ;  ;  

I. Introduction

Genome sequence information aids in studying species genetics and evolutionary history, enabling genetic improvement, promoting personalized treatment, and guiding biopharmaceuticals. Therefore, the high-quality assembly of a genome sequence is the critical foundation for understanding the complex biology of an organism, the genetic variation within a species, and the pathology of a tumor. It helps to predict phenotypic risk in diseases like type I diabetes, reveals cis and trans interactions between chromosomes, and provides treatment options for lung cancer patients (Glusman, Cox & Roach, 2014; Niederst et al., 2015; Mantere et al., 2019), and many more. Assembling genomes remains a challenging problem in computational genomics and is becoming more and more essential due to exponential growth in DNA sequencing throughput. Numerous large-scale projects currently focus on sequencing the genomes of multiple species, including humans, animals, plants, and microorganisms (Vega, 2019). The conclusion of the Human Genome Project provides valuable information regarding gene functionality and genetic diversity (Lander et al., 2001; Venter et al., 2001). However, the existing reference genome is not representative of the entire human population as highlighted by studies on human reference pangenome graphs and the Human Pangenome Reference Consortium (Lander et al., 2001; Venter et al., 2001; Wang et al., 2023). Therefore, efforts are made to achieve high-quality, continuous assembly with scaffolds that approximate chromosomal proportions. The availability and abundance of high-quality genomes provide plenty of new opportunities in genetic research. Genomic information and technology hold the potential to revolutionize healthcare by enhancing patient outcomes, ensuring improved quality and safety, and enabling cost efficiencies. An individual’s genetic makeup and variations are crucial in determining disease risk across various life stages, including prenatal, newborn, childhood, and adulthood (McCormick & Calzone, 2016; Riess et al., 2024). Many individuals with rare diseases do not have any molecular diagnosis, and the causative variants and associated genes for over 50% of these conditions remain unidentified (Ferreira, 2019). Genetic information can be leveraged for screening, offering a more precise understanding of health issues, optimizing medication choices including targeted therapies based on the genomic underpinnings of diseases and aiding in effective symptom management (McCormick & Calzone, 2016). However, there is a lack of patient genome data in publicly available databases.
DNA sequencing technology has revolutionized genomics since its inception in 1975, the dideoxy chain-termination method or first-generation sequencing by Sanger (Sanger, 1975). Gradually, it is followed by second-generation sequencing based on the sequencing-by-synthesis method, which offers high throughput and accuracy, thereby significantly contributing to the success of the Human Genome Project. (Margulies et al., 2005; Schuster, 2008). However, it has a short read length of 150-400 bp, which makes complex genome sequencing challenging due to high duplication and multiple heterozygosity. The emerging long-read sequencing technology (third generation) overcomes the limitations of short-read sequencing. It has significantly changed the landscape of whole-genome sequencing and empowered to decode genetic information (Goodwin, McPherson & McCombie, 2016). Advancements in long-read sequencing technology have enhanced genome assembly quality, resulting in more comprehensive chromosomal-level data. This improvement facilitates various applications, including genome annotation, mutation detection, evolutionary studies, gene function exploration, and comparative genomics. PacBio and Oxford nanopore technology (ONT) are widely used platforms for long-read sequencing where PacBio generates continuous long reads (CLR) or circular consensus sequencing reads (CCS referred to as HiFi). ONT generates ultra-long reads, making the genome completion process easier. The average lengths of PacBio and ONT read are 10-20 and 100kb, respectively. Both platforms have error rates ranging from 10 to 15%, with CLR errors being randomly distributed and ONT reads having limitations like GC bias (Marks et al., 2019). The preparation of paired-end and mate-pair libraries, along with an increase in sequencing coverage depth, contributes to enhancing the accuracy and completeness of genomic data to a certain degree. Few studies suggested that short reads could be required to refine long read assembly even more (Jain et al., 2018; Wang et al., 2023).
The increasing availability of sequence data necessitates the development of more efficient analysis algorithms to obtain fragments that span a longer range across chromosomes during genome assembly. There are two main types of genome assembly: reference-based assembly suitable for resequencing and de novo assembly more applicable for new species. De novo assembly reduces errors in the reference genome, and chromosome rearrangement, and promotes the development of Pan genomes (Gao et al., 2018; Padovani et al., 2019). De novo assembly uses numerous algorithms like the de-Bruijn graph (DBG), overlap-layout-consensus (OLC) string-graph (SG), and many other hybrid approaches (Dida & Yi, 2021). Methodologically, DBG chops reads into short k-mers with overlapping edges, resulting in edges or node paths to create a graph. Whereas, OLC generates a consensus sequence using an overlap reads layout. SG is a simplified OLC, where sequence reads and non-transitive edges result in suffix-to-prefix overlaps (Li et al., 2012; Liao et al., 2019). These methods aim to build an assembly graph using overlapping reads, simplifying it by removing tips and bubbles followed by traversal to obtain the contigs (Compeau, Pevzner & Tesler, 2011). As mentioned earlier, assembling complex genomes is challenging due to high duplication and multiple heterozygosity, which cannot be sequenced in a single pass using existing tools. Hybrid genome assembly techniques are employed to address intricate DNA segments. This method integrates short, precise second-generation sequencing data with longer, albeit less accurate, third-generation sequencing data to resolve complex DNA segments effectively. Researchers employ various hybrid assemblers like MaSuRCA (https://github.com/alekseyzimin/masurca), Unicycler (https://github.com/rrwick/Unicycler), HASLR (https://github.com/vpc-ccg/haslr), and WENGEN (https://github.com/adigenova/wengan) to enhance assembly quality (Wick et al., 2017; Zimin et al., 2017; Haghshenas et al., 2020; Genova et al., 2021).
After de novo assembly, its quality assessment is crucial for users to achieve optimal results and for developers to enhance assembly algorithms. However, varying methods of quality evaluation can lead to arbitrary comparisons. A pipeline incorporating multiple evaluation indicators is needed for comprehensive appraisal. Moreover, several quality polishing tools are also available but require accurate sequence information for this task. Most of the polishers use short reads due to their higher accuracy in correcting errors in genome assembly. Some polishers utilize the raw signal from the sequencing platform to improve the genome assembly quality. Selecting the right next-generation sequencing (NGS) technologies and implementing efficient bioinformatics workflows are essential for achieving high-quality genome assemblies at the chromosome level. A notable gap exists in comprehensive information on selecting the right de novo assembly tools and NGS platforms for human genome assembly.
Previous reviews have primarily focused on the assembly of short-sequence reads. However, recently published reviews have contributed valuable insights into the assembly of long sequencing reads (Flicek & Birney, 2009; Horner et al., 2009; Alkan, Sajjadian & Eichler, 2010; Jackman & Birol, 2010; Miller, Koren & Sutton, 2010; Poszhkiewicz & Studholme, 2010; Li et al., 2011; Narzisi & Mishra, 2011; He et al., 2013; Sohan & Nam, 2018 & Liao et al., 2019). Additionally, few reviews summarized step-by-step methods that can be utilized with minimal time or resources (Del angel et al., 2018; Jung et al., 2020). A recently published review presents a comprehensive overview of existing methods, technologies, and challenges associated with de novo genome assembly, which involves reconstructing a genome from sequencing data without relying on a reference genome (Zhang et al., 2022). These reviews have offered significant insights; however, specific areas have been overlooked or insufficiently highlighted in these analyses. These reviews did not include comparative benchmarking or in-depth analyses of hybrid assemblies that integrate short- and long-read data. Moreover, the previous reviews have not adequately discussed the processes of post-assembly refinement and error correction that enhance assembly precision. Furthermore, the significance of quality control (QC) throughout the assembly process has not been thoroughly addressed. Typically, previous reviews concentrate on fundamental methods and fail to investigate emerging technologies or the advanced computational challenges linked to large or complex genomes. Consequently, the absence of comprehensive examinations of specific computational techniques and detailed comparisons of algorithms can restrict their relevance for a variety of genomic projects. This review explores the influence of sequencing methods, assembly tools, polishing tools, and sequencing depth on genome construction. It helps to select the appropriate combination of tools for precise genome assembly, especially for large genomes. Genome sequencing, assembly, and annotation projects vary significantly because of the unique traits of each subject’s genome. There are four critical elements to consider when starting a new genome project: the size of the genome, the levels of ploidy and heterozygosity, the GC content, and the complexity of the genome. These elements will directly impact the quality and expenses of genome sequencing, assembly, and annotation. Therefore, it is essential to understand these properties before commencing the genome assembly (Del Angel et al., 2018; Jung et al., 2019).

II. Genome Complexity

More sophisticated biological organizations are the result of evolution and the evolved organisms are more complex than the primitive ones. The C-value represents the amount of DNA present in the haploid genome of an organism. Generally, it increases with the complexity of organisms ranging from prokaryotes to invertebrates, vertebrates, and plants (Choi, Kwon & Kim, 2020). However, it has been observed that some lower organisms have significantly higher C-values than the more evolved ones. Additionally, the C-values among related species vary by an order of magnitude. This lack of correlation between biological complexity and C-values is termed the C-value paradox. The classic C-value paradox is explained by various factors such as introns in genes, regulatory elements, pseudogenes, multiple copies of genes, intergenic sequences, and repetitive DNA. Eukaryotic organisms have larger genomes (2.3 Mbp to 150 Gbp) than prokaryotic (140 kbp to 15 Mbp) ones due to the presence of an intron and repetitive region (Corradi et al., 2010; Han et al., 2013; Lakhotia, 2023). Similar to the concept of the C-value, it was hypothesized that the protein-coding genes (G-value) would be greater in biologically more complex groups. However, classical cytogenetic studies revealed that protein-coding genes in organisms vary within a narrow range, resulting in a “G-value paradox” (Lakhotia, 2023). The G-value paradox is often attributed to various genomic features, including the ability of one gene to produce multiple mature mRNAs through alternative splicing, highly regulated gene expression at both transcriptional and translational levels, and the numerous functions of proteins during development (Choi, Kwon & Kim, 2020). Rapid genomics advancements have provided answers to the C-value and G-value paradoxes.
Since the human genome was completed fifteen years ago, genome sequence data has increased exponentially. As a result, the global dataset now contains hundreds of additional species across eukaryotic phylogeny. DNA Data Bank of Japan (DDBJ) (https://www.ddbj.nig.ac.jp/), GenBank (https://www.ncbi.nlm.nih.gov/genbank/) and the European Nucleotide Archive (ENA) (https://www.ebi.ac.uk/ena) collectively provide comprehensive repositories for nucleotide sequence data from a wide range of organisms. In addition to these primary biological databases, several genome browsers are available publicly for studying the genome (Supplementary Table S1). Genome sequence information has enabled broad-scale genome content examination and addressed long-standing questions in genome biology. Genome assembly projects vary by species due to differences in genome properties. Therefore, standardized reporting of genome properties could improve the characterization of hundreds of existing sequences with a standardized bioinformatics pipeline and curated database that facilitates large-scale comparative analyses. However, expanding the dataset to encompass detailed information about larger genomes will be more challenging than survey sequencing (Elliott & Gregory, 2015). The genome assembly of an organism initially requires the knowledge of genome properties including estimated genome size, repeats, heterozygosity, ploidy level, GC content, transposable elements, and, structural and functional annotation. The genome’s size is a key factor in influencing the volume of data that needs to be arranged and analyzed. Genome assembly requires a certain number of sequences, resulting in more data requirements for large genome assemblies. Thus, determining the genome size before proceeding with sequencing is vital. Usually, flow cytometry and k-mer frequency distribution methods are used for reliable genome size estimation. Apart from this, various public databases are also available to understand the genome size and complexity of fungi (http://www.zbi.ee/fungal-genomesize), animals (http://www.genomesize.com/), and plants (http://data.kew.org/cvalues). If the target species’ information is unavailable from a public database, selecting a closely related species is a practical option (Del Angel et al., 2018; Swathi et al., 2018; Jung et al., 2020).
Moreover, large genomes are highly repetitive which causes difficulties in the sequencing and assembly generation process. Around 50% of the human genome contains non-random repeat elements, such as LINEs, SINEs, LTRs, and STRs. These repetitive elements can cause misarrangements or gaps in assembly. However, technological advancements like long-read sequencing may help to overcome these challenges (Huddleston et al., 2014; McCoy et al., 2014). Larger genomes like lungfish, aquatic salamanders, pine, and lilies may require sorting chromosomes for sequencing or using haploid DNA for new assembly methods (Lucas et al., 2014; Neale et al., 2014). A genome with higher ploidy, heterozygosity rates, and repeating counterparts requires more computing resources than assembling a small genome. Regarding ploidy, haploid genomes are easy to assemble due to a single contiguous sequence without heterozygosity. However, diploids and polyploids are problematic in assembly due to heterozygosity between genome copies in a single individual. Most of the assemblers are based on the haploid genome. These assembly algorithms aim to combine allelic differences into a single consensus sequence, resulting in a final haploid assembly. For highly heterozygous genomes, sequence reads from homologous alleles can be too different to be assembled. These alleles will be assembled separately (Del Angel et al., 2018; Jung et al., 2020). The highly heterozygous genome also leads to more fragmented assemblies or creates doubt about the contigs’ homology. The level of heterozygosity largely depends upon the size of the population. Large population sizes often result in high heterozygosity levels. Therefore, inbreeding or creating doubled haploid individuals has been attempted to reduce heterozygosity (Berthelot et al., 2014; Jung et al., 2019; Linsmith et al., 2019; Zhang et al., 2020). However, these methods are not feasible for most species in natural environments. Finally, the GC content of the genome can also cause issues in assembly. Illumina sequencing can struggle with regions with low or high GC content, leading to low or no coverage in that region. Due to inadequate read coverage in GC-poor or GC-rich regions assembly gets fragmented and lowers the completeness of assembly (Chen et al., 2013). However, this can be addressed by increasing coverage or using a non-biased sequencing technology like PacBio or Nanopore (Del Angel et al., 2018; Jung et al., 2019).

III. NGS Raw Data from the Sequencing Platform

In 1977, Sanger’s ‘chain-termination’ or dideoxy technique revolutionized DNA sequencing technology. It enables the sequencing of clonal DNA populations. First-generation sequencing devices produce reads that are shorter than one kilobase. Researchers use ‘shotgun sequencing’ techniques to analyze large DNA fragments. In shotgun sequencing overlapping DNA segments were cloned, sequenced independently, and then assembled into a single, lengthy continuous sequence (Anderson, 1981). Large-scale dideoxy sequencing initiatives led to the emergence of another approach that paved the way for the first wave of DNA sequencers. Second-generation sequencing uses a massively parallel sequencing approach. This parallelization led to an orders-of-magnitude increase in the yield of sequencing operations. Several NGS platforms with different strategies have been developed for preparing sequence libraries and detecting signals, like Roches 454, and Illumina’s GA. The performance comparison shows variations in output capacity for different platforms (Supplementary Table S2) (Liao et al., 2019). Advances in NGS technology have increased genomic information but it also contains more errors than traditional methods. It may introduce bias and mismatches, making it less convenient for data analysis. Next-generation platforms like Illumina, SOLiD, PGM, and 454 systems require local clonal amplification to increase signal-to-noise ratios. Illumina is the leading NGS platform, offering high throughput, base quality, and low per-base cost. It uses bridge amplification for colony generation and sequencing. It generates up to 180 million reads per HiSeq2000 lane. The HiSeqX10 and NovaSeq6000 short-read sequencers are the latest platforms offered by Illumina (Logsdon et al., 2020). In contrast to the Illumina sequencer, the Roche 454 system provides a longer read length along with an accuracy rate of 99.9%. It is ideal for applications requiring long read length, such as de novo assembly. On the other hand, the SOLiD system has low error and 99.94% accuracy. However, it is less suited for de novo assembly due to short reads and lengthy run durations (Huang et al., 2017). Another rival sequencer, the DNBSEQ-T7 (formerly MGISEQ-T7) was created by Complete Genomics and MGI Technology (Logsdon et al., 2020). DNBSEQ-T7 is a novel sequencing platform that builds upon BGISEQ-500 by utilizing combinatorial probe anchor synthesis and DNA nanoball technology to produce short reads on a massive scale (Drmanac et al., 2010). Using the BGISEQ-500 sequencer, the first dataset of a complete human genome was developed, drawing on the well-known cell line HG001 (NA12878). It has been reported that the data of BGISEQ-500 PE100 and HiSeq2500 PE150 had a similar single nucleotide polymorphism (SNP) detection accuracy (Huang et al., 2017). Subsequently, the comparison of BGISEQ-500 and HiSeqX10 sequencing platforms showed the high concordance between the two platforms for germline genotypes and single nucleotide variations (SNVs) detected from SNP arrays, as well as indels (Patch et al., 2018). In addition, various comparative studies on variant calling pipelines across sequencing platforms have been reported (Chen et al., 2019; Kim et al., 2021). A recent study estimates the quality of 7 short-read-based whole-genome sequencers, 2 MGI platforms (BGISEQ-500 and DNBSEQ-T7), and 5 Illumina platforms (HiSeq2000, HiSeq2500, HiSeq4000, HiSeqX10, and NovaSeq6000) (Kim et al., 2021). They showed that MGI platforms had a higher concordance rate for SNP genotyping, while Illumina platforms had comparable levels of quality, sequencing uniformity, and coverage. Short reads may not accurately align with the reference genome, leading to the discarding of reads and gaps in the sequencing data. This poses challenges in sequencing regions with pseudogenes or repeated genes, especially when the read length exceeds that of the repetitive region (Goodwin, McPherson & McCombie, 2016).
Long-read sequencing or third-generation sequencing technology can overcome the challenges faced by NGS. It can resolve difficult-to-map genes, and identify co-inherited alleles, haplotype information, and phase de novo mutations. It generates long reads for genome-finishing applications. Third-generation sequencing technologies allow for the direct sequencing of individual DNA molecules, a significant advancement over earlier sequencing methods. It does not require pre-amplification steps, thereby reducing the errors of PCR amplification. Two popular long-read sequencing platforms that measure disturbances in electric current to sequence DNA are PacBio and ONT. It offers read length advantages, with average read lengths over 10 kb and maximum read lengths over 60 kb. However, limited yield, high error rate, and cost per base hinder large-scale sequencing projects. In comparison with Illumina, PacBio CLR data has low single-pass accuracy. However, these errors can be corrected using polishing. The latest innovation from PacBio is the creation of high-fidelity (HiFi) sequencing reads. This is the first kind of data that is both very accurate (>99%) and lengthy (>10 kb). PacBio HiFi long-reads with excellent base quality, are the greatest option for haploid contig assembly but with low throughput and high cost (Wang et al., 2023). Conversely, ONT has a special pore chemistry and can outperform PacBio by an order of magnitude by producing continuous sequences ranging from hundreds to thousands of kilobases in length. However, ONT long reads constitute a small fraction of the entire read length distribution (Logsdon et al., 2020). There are three standard ONT platforms: MinION, GridION x5, and PromethION. The flow cell capacities of each platform vary, enabling different data generation capabilities. Studies show that PacBio and ONT sequencing methods can assemble the human genome in less than 100 minutes (Chin & Khalak, 2019; Shafin et al., 2019). This is significantly faster than the 100 CPU hours required for aligning 30X short-read Illumina data. A comparison of the assembly time for long-read sequencing data revealed that the assembler operated four times slower with ONT data than with PacBio data. This discrepancy is likely attributed to the higher incidence of systematic errors in the raw sequencing data from ONT, which leads to diminished accuracy compared to PacBio data (Jain et al., 2018). The study conducted by Wang et al. (2023) showed that PacBio-HiFi long-reads produce assemblies with low base errors, while Oxford Nanopore long-reads produce the most continuous contigs that may need further polishing for quality improvement.

IV. Major Problems with the Sequencing Read from Different Platforms

As discussed in the previous section, advances in DNA sequencing have enabled the production and curation of large genomic data sets in several species. The features of reads generated from the sequencing platform significantly affect the downstream analysis of data. Therefore, maintaining consistent read quality is crucial to minimize the variations in analysis output. The primary challenges in the current sequence assembly include read length, sequencing errors, repetitive regions, and sequencing bias. The short read platform has necessitated considerable computational resources to synthesize short reads that cover the entire genome, posing a substantial challenge in NGS. The assembly of next-generation short reads encounters difficulties when dealing with repetitive fragments that exceed the length of the reads. Third-generation sequencing (TGS) technologies can produce large reads. So it can overcome the problem of computing time for short reads generated through second-generation sequencing (SGS). It can also solve the larger scale repetitive regions but has limited efficiency due to the size of repetitive regions. Another way to solve the repeat problem is using large insert-size paired-end (PE) reads (Liao et al., 2019). Several assemblers, like SOAPdenovo, ABySS, and Velvet, use PE sequencing information to reduce gaps in the genome assembly (Cahill et al., 2010; Li et al., 2010). Furthermore, compressed data structure-based analysis showed the potential for de novo assembly of large genomes (Simpson & Durbin, 2012). Regardless of the advantages, TGS platforms continue to cause worry because of their higher sequencing error rate (5–20% for TGS and 1% for SGS). These errors can harm the precision of assembly (Goodwin, McPherson & McCombie, 2016; Giordano et al., 2017). The error correction step is crucial for ONT read assembly. However, the upcoming improvements in read quality may produce ONT HiFi reads, eliminating the need for a specialized error correction step. NGS platforms generate several errors like substitutions, insertions, and deletions, leading to mis-assemblies and complex DBG (Liao et al., 2019). Sequencing error profiles may include base call error enrichment, compositional bias towards high-GC sequences, and inaccurate determination of simple sequence repeats. Illumina sequencing can struggle with regions having low or high GC content. This can be compensated with increased coverage or by using sequencing technology (such as PacBio or Nanopore) that doesn’t show such bias (Del Angel et al., 2018). Homopolymer errors are another common issue in long-read sequencing technologies. It can cause reading frame shifts in protein-coding regions, incorrect genome annotation, and more downstream analysis (Jain et al., 2015; Weirather et al., 2017; Wenger et al., 2019). Moreover, sequencing reads face sequence bias problems, resulting in uneven read depth distribution across the genome. This creates difficulties in assembling low-depth reads into contigs or aligning them to reference sequences even if mistakenly assembled. Existing de novo assembly algorithms face a significant challenge due to uneven coverage. Various de Bruijn-based assemblers like Velvet, SOAPdenovo, and ABySS use average coverage cutoffs to prune out low-coverage regions to reduce complexity but it affects the effective length and genome fraction of the final assembly (Zerbino et al., 2008; Simpson et al., 2009). Reports suggest that combining an assembler with assembly based on read classification results in the generation of more precise and larger contigs and scaffolds than when the assembler is used independently (Liao et al., 2020).

V. NGS Data Analysis

NGS data analysis for genome assembly plays a vital role in genomics, focusing on reconstructing an organism’s entire genome from fragmented DNA sequences produced by high-throughput sequencing technologies. The fundamental steps involved in genome assembly are illustrated in Figure 1. The process initiates with QC of the raw reads, utilizing tools such as FastQC to evaluate read quality and eliminate low-quality sequences or adapter contamination. Subsequently, the cleaned reads are assembled into longer sequences known as contigs through alignment and overlap methods, employing tools like SPAdes, Canu, or Velvet, tailored to the specific sequencing platform (e.g., Illumina, PacBio, or Oxford Nanopore). Once the contigs are formed, they undergo further processing into scaffolds representing larger DNA segments. This is typically achieved by leveraging paired-end or mate-pair read information to connect contigs based on known distances. The assembly is then refined using tools such as Pilon or Arrow to rectify errors and enhance accuracy. To evaluate the assembly’s quality, metrics like continuity and completeness are computed. The final assembled genome can then be annotated to identify genes, regulatory elements, and other functional characteristics. This genome assembly process is crucial for various applications, including species discovery, evolutionary research, and disease investigation. A schematic (Figure 1) and detailed discussion on the genome assembly process have been summarized below.
(1) 
Pre-Processing of Raw reads
The sequencing data may contain errors, biases, and uncertainties that cause computational and statistical challenges during downstream analysis (Aird et al., 2011; Nakamura et al., 2011; Allhoff et al., 2013). Therefore, a QC process such as filtering low-quality reads is needed to obtain a reliable conclusion in downstream analysis. Pre-processing raw reads involves the elimination of known contamination, ambiguous bases, adapter sequences, and low-quality bases. Generally, there are two methods for handling low-quality nucleotides in raw reads. The first approach, which is applied by many pre-processing programs such as Quake (http://www.cbcb.umd.edu/software/quake) is to superimpose the reads. Then low-frequency patterns are modified based on the most common sequences in read overlap (Kelley, Schatz & Salzberg, 2010). Usually, this approach works at the k-mer level. Based on this approach, some de novo assemblers such as ALLPATH-LG, SOAPdenovo2, and ABySS have in-built error correction steps (Simpson et al., 2009; Gnerre et al., 2011; Luo et al., 2012).
The second strategy for dealing with low-quality nucleotides focuses on surgically eliminating the low-quality regions—a process known as read “trimming”—rather than altering the initial read dataset. FASTQ data QC and pre-processing can be resolved using various tools such as Quake, PIQA, Trimmomatic, Trimgalore (https://github.com/FelixKrueger/TrimGalore), Seqtk (https://github.com/lh3/seqtk), CutAdapt, NGS QC Toolkit (https://github.com/mjain-lab/NGSQCToolkit) and FastQC (Andrews, 2010; Kelley, Schatz & Salzberg, 2010; Martin, 2011; Liu et al., 2012; Andrews, 2014; Bolger, Lohse & Usadel, 2014). FastQC (https://sourceforge.net/projects/fastqc.mirror/) is a Java-based tool with per-base and per-read quality profiling features (Andrews, 2010). In contrast, Cutadapt (https://github.com/marcelm/cutadapt/) is a widely used tool for trimming adapters and includes read-filtering functionalities (Martin, 2011). Trimmomatic (http://www.usadellab.org/cms/index.php?page=trimmomatic) is another popular trimming adapter tool that performs quality trimming using algorithms like sliding window cutting (Bolger, Lohse & Usadel, 2014). Recently, SOAPnuke (https://github.com/BGI-flexlab/SOAPnuke) has been released for adapter trimming and read filtering using MapReduce on Hadoop systems (Chen et al., 2018). In earlier practices, the pre-processing and QC for FASTQ data relied on multiple software applications. Due to the need for repeated data readings and loading, these approaches proved inefficient and unproductive regarding input/output. Fastp (https://github.com/OpenGene/fastp), on the other hand, is an exceptionally rapid tool designed for base correction, read filtering, and quality assessment of FASTQ data. It has been reported that Fastp trimming reduces the execution time and computing resource requirements while improving the analysis’s excellence and consistency (Chen et al., 2018). Recently, a web tool (e-QA NGS) has been reported to be capable of performing various tasks such as quality assessment, trimming, and quality filtering through an intuitive and automated workflow (Gaia et al., 2023).
In the majority of existing NGS research, trimming has been widely used, particularly before transcriptome assembly, metagenome reconstruction, RNA-Seq, epigenetic studies, comparative genomics, genome assembly, and metagenome reconstruction (Karlsson et al., 2012; Schott et al., 2012; Tomato Genome Consortium, 2012; Giorgi, Del-Fabbro & Licausi, 2013; Liu et al., 2013). Despite its widespread use, neither a thorough evaluation of the impact of trimming on typical NGS studies nor a systematic comparison of the current tools have been created. Numerous techniques have been separately detailed in the literature, but their applicability has only been demonstrated in certain genome assembly scenarios (Cox, Peterson & Biggs, 2010; Schmieder & Edwards, 2011; Smeds & Kunstner, 2011). However, a study explained the basic techniques that drive the workings of the nine most widely used trimming applications (Fabbro et al., 2013). They showed that the behavior of trimmers varies and is mostly influenced by the settings employed. If the quality threshold is too high, it may significantly reduce the quantity of the surviving dataset. Conversely, a low-quality threshold leads to retaining an excessive number of noisy or random readings, which lowers the quality of the dataset and needlessly increases computing demands. Therefore, the researcher must decide the optimal trade-off between read loss and dataset quality.
(2) 
Generation of genome assembly from NGS data
The downstream analysis of high throughput sequencing reads with different error profiles is challenging, especially for genome assembly of large genomes. In bioinformatics, genome assembly refers to aligning and combining fragments from a longer DNA sequence to reconstruct the original sequence. (Sohn & Nam, 2018). Sequencing data can be assembled using two primary methods: reference-guided genome assembly and de novo genome assembly. The former method is easier with reference genome or proteome information than the earlier one. De novo assembly is ideal for identifying new species or species genetic diversity, reducing errors in reference genomes, and promoting Pan genomics development by reducing chromosome rearrangement deviations (Zhang et al., 2022). Here a range of tools available for genome sequence assembly specifically highlighting approaches for de novo assembly has been explored.
(a) Reference-based assembly
Over the decade, several BioGenome projects have been initiated to create a reference genome for all known eukaryotic species (Hotaling, Kelley & Frandsen, 2021; Formenti et al., 2022). The Genome Reference Consortium (GRC) is dedicated to enhancing the human reference genome assembly by correcting errors and adding sequences to meet basic and clinical research needs. In 2017, the GRC published the GRCh38 assembly of the human reference genome. The latest patch was published in 2022, GRCh38.p14 (“GRCh38.p14 - hg38 - Genome - Assembly - NCBI”. www.ncbi.nlm.nih.gov. Retrieved 2022-08-19 and “GRCh38.p14- hg38 - Genome - Assembly - NCBI - Statistics Report”. www.ncbi.nlm.nih.gov. Retrieved 2022-08-18). It has 349 gaps, a significant improvement from the first version with 150,000 gaps. Later, the assembly focuses on telomeres, centromeres, and repetitive sequences to remove gaps and errors. In 2022, the Telomere-to-Telomere Consortium published the first completely assembled reference genome without gaps (“T2T-CHM13v2.0 - Genome - Assembly - NCBI”. www.ncbi.nlm.nih.gov. Retrieved 2022-08-16 and “Telomere-to-Telomere”. NHGRI. Retrieved 2022-08-16). The next assembly release for the human genome is currently “indefinitely postponed” (Schneider et al., 2017; Nurk et al., 2022). Reference-guided assembly techniques can be used for the genome assembly process when more and more genomes are sequenced. The first step involves the aligning of PE and mate-pair reads to an available reference genome of similar species to create contiguous DNA regions (contigs) of consensus sequences. Generally, read alignment was accomplished using the Burrows-Wheeler Aligner (BWA) (https://github.com/lh3/bwa) method. Its revision algorithm BWA-MEM, is used for aligning particularly for reads longer than 70 bps. However, BWA-backtrack may perform better for shorter reads (Patch et al., 2018). On the other hand, BWA-SW, BLAT (https://genome.ucsc.edu/cgi-bin/hgBlat), SSAHA2, and Ira were used for long-read alignment. Long-read aligners must be flexible about alignment gaps due to the prevalence of indels in long reads, which may be the primary cause of sequencing errors in long sequencing technologies (Li & Durbin, 2010; Ren & Chaisson, 2021). Gaps between contigs can be filled through scaffolding, PCR and sequencing, or BAC cloning. It is not always possible to fill these gaps. In that instance, a reference assembly is produced with numerous scaffolds. After the alignment of reads against a reference genome, genotype calling pipelines are crucial for large-scale whole genome sequencing (WGS) efforts. Various variations calling systems discover a multitude of minor genomic variants, including insertions and deletions (INDELs) and SNPs. GATK (https://github.com/broadinstitute/gatk/releases) is widely recognized as the predominant variant calling pipeline. However, alternatives such as DeepVariant (https://github.com/google/deepvariant) and Illumina’s DRAGEN (https://github.com/seqfu/docker-dragen) have demonstrated superior in F1 statistics, precision, and recall. In 2019, a study by Chen et al. evaluated the performance, consistency, and operational efficiency of 27 combinations of sequencing platforms and variant calling pipelines. They tested three variant calling pipelines—Genome Analysis Tool Kit HaplotypeCaller (https://github.com/oicr-gsi/haplotypeCaller), Strelka2 (https://github.com/Illumina/strelka), and Samtools-Varscan2 (https://github.com/samtools/samtools and https://github.com/Jeltje/varscan2) for the data set generated by different platforms including BGISEQ500, MGISEQ2000, HiSeq4000, NovaSeq, and HiSeq Xten. Out of three variant calling pipelines, Strelka2 demonstrated exceptional performance in detecting accuracy and processing efficiency on both BGI and Illumina platforms for the GIAB datasets. Similarly, Betschart et al. (2022) compared six WGS data pre-processing pipelines, including two mapping and alignment approaches (GATK utilizing BWA-MEM2 2.2.1, DRAGEN 3.8.4) and three variant calling pipelines (GATK 4.2.4.1, DRAGEN 3.8.4 and DeepVariant 1.1.0). They showed that DRAGEN outperformed GATK in variant calling. However, DRAGEN and DeepVariant performed similarly with slight advantages for Indels and SNVs, respectively, over GATK. Variant calls in one platform differ from another possibly due to read length differences, error rate, and the same analysis methodology. Short reads have higher alignment bias and error, especially in AT-rich regions with repetitive, non-coding DNA (Patch et al., 2018). Reference-based methods like GATK can infer human variations from short-read sequences but only cover about 90% of the reference human genome assembly. These approaches are accurate for single-nucleotide variants and short insertions and deletions, but struggle for de novo genome assembly, structural variations, and phasing relationships without transmission information or haplotype panels (Shafin et al., 2020). Long reads can also be used for genotypic analysis to cover the whole genome. However, it has been found that nanopore genotyping accuracy presently falls behind short-read sequencing devices. This is due to its inability to distinguish between heterozygous and homozygous alleles, a limitation arising from the error rate and coverage depth of nanopore sequencing data (Jain et al., 2018). Moreover, reference-based assembly faces limitations due to limited representation of existing human reference genome, lack of large structural variants, and poor representation of complex regions (Wang et al., 2023). Therefore, de novo genome assembly is required to create high-quality reference genomes using a combination of precise short reads with long noisy reads.
(b) De novo assembly
De novo genome assembly involves reconstructing an organism’s genome using smaller sequenced fragments, all without a reference genome. Many de novo sequence assembly methods have been developed in recent years to manage and combine the vast quantity of small sequence reads to create larger segments known as contigs (Baker, 2012). However, selecting a suitable assembler remains a difficult task. Different sequencing platforms produce reads of varying lengths and accuracy. Therefore, the assembly process requires distinct algorithms from both short and long-read technologies. Various types of assemblers used for short and long-read assembly based on these different algorithms are given in Supplementary Tables S3 and S4. All the available de novo assemblers often use two different kinds of algorithms: greedy and graph method algorithms. Greedy algorithm-based assemblers locate local optima in alignments whereas graph algorithm assembler aims to reach the global optima. Greedy algorithms such as OLC, begin by merging the short sequence reads with the best overlap to generate contigs (Miller, Koren & Sutton, 2010; Khan et al., 2018). The OLC method detects overlaps between reads and creates a directed graph based on all vs. all pairwise comparisons. The graph makes each sequence a node, with an edge formed between two nodes whose sequences overlap (Liao et al., 2019; Dida & Yi, 2021). The algorithm identifies the Hamiltonian traversal path of the graph, combining all overlapping sequences in the nodes into the genome sequence. Some assemblers that use the OLC approach include Arachne (ftp://ftp.broadinstitute.org/pub/crd/nightly/arachne/), Ray (https://github.com/sebhtml/ray), Canu (https://github.com/marbl/canu), Minimus (http://amos.sourceforge.net/docs/pipeline/minimus.html), CABOG (Celera assembler) (https://www.cbcb.umd.edu/software/celera-assembler), Newbler (https://github.com/etheleon/newbler), Flye (https://github.com/mikolmogorov/Flye), Shasta (https://github.com/chanzuckerberg/shasta) and MIRA (https://github.com/bachev/mira). The OLC technique performs better with longer sequence reads (Khan et al., 2018). Another slightly modified algorithm is DBG, which replaces reads with shorter sequences represented as k-mers of a fixed length (k) (Li et al., 2012). Then contigs are formed by merging the adjacent k-mers that overlap by k-1. Dividing sequences into smaller sizes can enhance the crisis of different initial read lengths (Khan et al., 2018). DBG requires precise reads and initially eliminates part of the reads’ capacity to resolve repetitions longer than k bases. This method, like OLC, removes the necessity of storing pairwise overlaps and features a graph structure that resembles the genome’s repeat structure. The DBG-based approach has proven effective for assembling short reads under 100 bp, and has also been successfully applied to longer reads. It is also used for complex genome assembly (Zhang et al., 2011). Some preferred de novo methods for next-generation sequencing assembly are ALLPATHS (https://bioinformaticshome.com/db/tool/Allpaths-LG), Velvet (https://github.com/dzerbino/velvet), SPAdes (https://github.com/ablab/spades), SOAPdenovo (https://github.com/aquaskyline/SOAPdenovo2) and ABySS (https://www.bcgsc.ca/resources/software/abyss). The Velvet assembler is a standard tool for assembling small to medium-sized genomes using short reads (Zerbino & Birney, 2008). Other short read assemblers such as ABySS, SOAPdenovo and ALLPATHS can be used for large genome assembly (Simpson et al., 2009). Long sequences and high error rate single-molecule sequencing reads are better suited for overlap-based techniques than DBG-based ones (Dida & Yi, 2021). The string graph algorithm is a modification of the OLC algorithm that performs a global overlap graph by eliminating unnecessary sequences (Li et al., 2012). In theory, the construction of the string graph assembly is comparable to that of a DBG. Its benefit is that it takes the entire length of a read sequence rather than breaking them down into k-mers (Dida & Yi, 2021). It is primarily used for small genome assembly (Zhang et al., 2011).
Various studies were carried out that use different de novo algorithms for human genome assembly (Supplementary Table S5). The objective of genome assemblers is to reconstruct the entire genome using the fewest consecutive fragments while achieving the highest base accuracy and minimizing computational resources. The assembly algorithm was initially developed for short reads due to their high accuracy. Short-read DBG assemblers achieve goals of high accuracy and low computational resources, while long-read assemblers excel at reconstructing the complete genome in the fewest consecutive pieces. In 2011, Gnerre et al. demonstrated the ability of DBG (ALLPATHS-LG) assembler to perform genome assembly of human and mouse genome datasets generated on the Illumina platform. The resulting draft assemblies showed high accuracy (≥99.95%), good long-range connectivity, and genome coverage. In addition to DBG, research has been carried out on the string and OLC strategies. A study reveals that string-based and OLC assemblers are suitable for very short and longer reads respectively, in small genomes with millions of short reads with limited computational power (Zhang et al., 2011). They also showed that the time and memory cost of a string-based assembler (SGA) is directly proportional to the size of the dataset. The influence is further affected by the complexity inherent in the dataset. Simpson & Durbin (2012) showed that SGA uses the least memory compared to ABySS, Velvet, and SOAPdenovo assemblers. However, DBG assemblers (ABySS, Velvet, and SOAPdenovo) were more computationally efficient than SGA as the latter required more CPU hours. This increased demand is largely attributed to the time taken by SGA to construct the FM index. Overall comparison studies showed that SGA is less efficient than DBG such as ALLPATHS-LG and SOAPdenovo (Li et al., 2010; Gnerre et al., 2011). However, using short reads for de novo assembly faces challenges like less genome coverage and more computational time. Long-read sequencing makes de novo human genome assembly more tractable. De novo assembly of the human genome using nanopore sequencing has been reported (Jain et al., 2018). They use short read correction to improve the accuracy of assembly. Although this initial effort required 53 ONT MinION flow cells and 150,000 CPU hours for assembly. To facilitate fast human genome assembly, the use of nanopore in combination with proximity-ligation (HiC) sequencing has been reported (Shafin et al., 2020). They showed that the Shasta toolkit and nanopore sequencing enable efficient de novo assembly of 11 human genomes in 9 days. They required 3 ONT PromethION flow cells per sample to achieve 63X coverage. For the successful completion of a haploid human genome assembly, Shasta required 6 hours on a single commercial compute node. The emergence of innovative algorithms involves utilizing PacBio sequencing for the de novo assembly of genomes. It has been reported that the assembly produced using PacBio Hifi reads outperforms the Illumina-Nanopore hybrid and Nanopore assemblies. De novo genome assemblers using HiFi reads need less genome coverage than those of the former approaches (Gavrielatos et al., 2021). Using lengthy, high-fidelity sequencing reads, Hifiasm is a present-day de novo assembler that accurately depicts haplotype information in a phased assembly graph (Cheng et al., 2021). It has been shown that HiFiasm outperformed other assemblers (A5-Miseq, canu, falcon, flye, hinge, SGA, and Spades) in the human genome dataset in terms of genome fraction, total aligned length, and total contigs length. However, it does encounter challenges related to mismatches and mis-assemblies (Dida & Yi, 2021).
Modern long-read assemblers have adopted the OLC paradigm and new algorithms have accelerated all-versus-all read comparison. However, assembling uncorrected long reads gives more work to consensus polishers which improves the base accuracy of assembled contig sequences. Long-read assemblers typically perform a single round of long-read polishing followed by several rounds using third-party tools. Hybrid assemblers combine short and long reads for precise genome assembly, addressing the challenges faced by short and long read assemblers. WENGAN is a hybrid assembly algorithm that offers high quality at low computational cost. It enables the de novo assembly of human genomes using sequencing data from ONT PromethION, PacBio Sequel, Illumina, and MGI technology (Genova et al., 2021). Long-read sequencing assembly in human genomes often requires large amounts of DNA from homogeneous cell lines without considering cell heterogeneity, which could significantly impact haplotype assembly results. The systematic analysis of de novo human genome assembly using single cells sequenced on HiFi or ONT platforms has been reported (Porubsky et al., 2021; Xie et al., 2022). Recently, comprehensive studies conducted on benchmarking the multi-platform sequencing technologies (PacBio Hifi, CLR, and ONT) for human genome assembly showed that PacBio HiFi long-reads can generate the best-quality contig assembly with Hifiasm, HiCanu, or Flye (Gonzalez-Garica et al., 2023; Wang et al., 2023). Using PacBio HiFi long reads is the best way to achieve an assembly with minimal foundational errors. In contrast, Wang et al. (2023) were able to create the greatest number of continuous contigs using Oxford Nanopore long-reads. Still, the task of creating high-quality human diploid assemblies remains a challenging one. A study reveals that graph-based assembly approaches using PacBio HiFi long-reads and parent-child data can produce high-quality diploid human genomes (Jarvis et al., 2022). However, the reference-free method has been reported for diploid de novo genome assembly. This method combines the continuous long-read or high-fidelity sequencing data with the chromosome-wide phasing and scaffolding powers of single-cell strand sequencing to generate a fully phased de novo genome assembly (Porubsky et al., 2021). Recently, a study benchmarks four diploid assemblers that can construct phased assemblies without parental data (Wang et al., 2023). All are compatible with hybrid approaches, with hifiasm being the best pipeline for assembling phased human genomes with PacBio HiFi and Hi-C data.
Today, the increasing volume of sequencing data for biological research makes the availability of numerous genome assemblers. These assemblers generated valuable assemblies that included a substantial representation of their genes and overall genome structure. However, the significant degree of variation across the entries indicates that there is likely still much space for advancement in the field of genome assembly. The methods that are effective for assembling the genome of one species might not be effective for another. These create difficulties in choosing the best assembly of reads from high-throughput sequencing systems. Therefore, it is essential to confirm the accuracy of the information included in the assembled sequence before moving further with downstream data processing.
(3) 
Evaluation of genome assembly quality
Genome quality evaluation is a complex and challenging task due to the uncertainty of the real genome sequence. Using multiple strategies to assess assembly quality is a common solution, but selective use can be arbitrary and inconvenient for fair comparison. Genomic projects often use different metrics to assess the quality of genome assemblies, but these methods can vary across projects and lead to different outcomes. This results in a lack of confidence in the quality of published genomes. Moreover, the consensus on appropriate metrics for evaluating assembly quality remains unresolved. To address this, a pipeline incorporating multiple evaluation indicators is needed for a more comprehensive and efficient appraisal. Different assembly evaluation tools focus on four key metrics: contiguity, accuracy or correctness, completeness, and contamination. Some tools evaluate these metrics by comparing the results of the assembled genome to a related reference genome. While others use heuristics to ensure consistency between assembly results and sequencing reads (El-Metwally, Hamouda & Tarek, 2021). Length metrics, specifically N50/NG50 and L50/LG50 values, are a widely accepted measure of assembly contiguity. The sequencing length of the shortest contig at 50% of the total assembly length is known as the N50. NG50 is similar to N50 except it is related to genome size rather than total contig length. NA50 and NGA50 are respectively, similar to N50 and NG50, but contigs are replaced by aligned blocks (Bradnam et al., 2013; Gurevich et al., 2013). The N50 value is not robust enough to measure assembly contiguity. Therefore, its effectiveness is constantly under scrutiny. It has been noted that misassembly can lead to higher N50 and biased evaluations of assembly quality (Phillippy, Schatz & Pop, 2008). Furthermore, incorporating improperly joined sequences in N50s can enhance the accuracy of genome consensus sequences. Thus, a more informed decision should consider real contiguity and errors in each assembly. Therefore, NG(X) plots offer a more comprehensive view across all thresholds (1–100%) (Bradnam et al., 2013). Moreover, Gap number and contig number are also widely used to indicate the continuity of assembly (Liu et al., 2021; Ma et al., 2021b; Qin et al., 2021b). It has been suggested to adjust the ratio of contig counting to chromosome pair number to get better assembly evaluation (Wang & Wang, 2023). The second parameter for evaluating the quality of assembly is correctness or accuracy. The precision of the assembly is evaluated by identifying the quantity and categories of mis-assemblies, such as mismatches, indels, and mis-joins, which are considered the least favourable types. Mis-joins include inversions, relocations, and translocations where distant loci are improperly joined in assembly (Zhang et al., 2023). Base-level accuracy is evaluated by mapping NGS reads onto the assembly and calling for homozygous variations. But it may suffer from non-specific alignment and sequencing imbalance (Liu et al., 2020; Li et al., 2021; Oppenheimer et al., 2021; Qin et al., 2021a). These mapping-induced errors can be avoided by comparing the k-mer spectra. However, the heterozygous and repetitive regions present in the genome can affect the accuracy of genome assembly (Mapleson et al., 2017; Rhie et al., 2020). Structural-level accuracy delves deeper into the precision of assembly processes. It focuses on bigger genomic configurations rather than just individual nucleotides. To evaluate structural correctness, reference-based methods (such as QUAST) can be utilized as benchmarks to find structural variants in reference genome assemblies. Most of the approaches to evaluate structure-level accuracy are based on either whole-genome sequencing reads, or manual checks assisted by reference genome, Hi-C, or Bionano data (Hunt et al., 2013; Putnam et al., 2016; Mikheenko et al., 2018; Chen et al., 2021; Khelik et al., 2020; Qin et al., 2021b). Inspector is a reference-free long-read de novo assembly evaluator that accurately reports structural and small-scale assembly errors (Chen et al., 2021). It evaluates assemblies solely based on third-generation sequencing reads, which provide the most precise representations of target genomes. However, the current assembly releases lack scores that accurately reflect the importance of correctness in QC. The third important parameter to evaluate the quality of assembly is completeness. It can be assessed by various means like comparing the assembled genome length to the estimated genome size, comparing the k-mer spectrum to high-accuracy sequencing reads, or mapping WGS reads to the assembly (Shen et al., 2020; Rhie et al., 2020; Xu et al., 2022). The assessment of genome completeness is conducted by examining universally distributed genes across species, referred to as orthologs. Benchmarking Universal Single-Copy Orthologs (BUSCO) (https://gitlab.com/ezlab/busco/-/releases#5.8.1) and CEGMA (https://github.com/KorfLab/CEGMA_v2) are evaluation strategies for conserving gene sets in assembly. These tools assess gene status and infer the effect of assembly (Parra, Bradnam & Korf, 2007; Simao et al., 2015; Seppey, Manni & Zdobnoy, 2019). BUSCO offers a comprehensive overview of complete single-copy, duplicated, fragmented, and missing genes. The current metrics for completeness evaluation primarily focus on gene space but often overlook challenging areas like tandem repeats, including centromeres, telomeres, ribosomal loci, and organelle genomes. LAI is a tool used to assess the completeness of repetitive genomic regions by estimating intact LTR retro-elements (Ou, Chen & Jiang, 2018). Overall, the genome assembly quality assessment is based on 3C principles: continuity, correctness, and completeness. However, the best assembly for a genome is not well defined. It varies from project to project and depends upon whether to maximize contig and scaffold length or minimize misassemblies. Currently, the most commonly used measures are N50 and BUSCO/CEGMA which address only two of the 3C. Over 90 assemblies published in the past two years were analyzed, with 98% using BUSCO for assembly completeness and 91% using contig N50 for assembly continuity. However, only 22% and 40% used sequencing read for base-level accuracy and genome completeness, making assembly quality evaluation incomplete (Zhang et al., 2023).
Historically, extensive research has been conducted on creating methods for comparing various assemblers. Assembly competitions such as GAGE (http://gage.cbcb.umd.edu/) and Assemblathon (https://assemblathon.org/) assess various assemblers using common datasets (Earl et al., 2011; Salzberg et al., 2012; Bradnam et al., 2013). GAGE-based evaluations are limited to dataset assemblies with a known reference genome. It makes them unsuitable for evaluating assemblies of previously un-sequenced genomes. QUAST (https://github.com/ablab/quast) is a new tool for assessing assembly quality, offering a range of metrics for various users (Gurevich et al., 2013). It can evaluate assemblies with or without a reference genome, allowing researchers to assess new species assemblies without a finished reference genome. Initially, quality evaluation described completeness qualitatively, and mis-assemblies are often overlooked. Subsequently, tools like Recognition of Errors in Assemblies using Paired Reads (REAPR) (http://www.sanger.ac.uk/science/tools/reapr#download-and-installation) are developed specifically for large genomes and NGS data. It focuses on the precise scoring of each base and automatically detects misassemblies without requiring a reference sequence (Hunt et al., 2013). These tools offer corrected assembly statistics and enable the quantitative comparison of multiple assemblies. LASER (https://github.com/lucian-ilie/LASER) is an innovative program for assessing large genome assemblies, building on the foundation established by the QUAST program. (Khiste & Ilie, 2015). It offers significant performance improvements, being 5.6 times faster and requiring half the memory required by the QUAST program for human genome assemblies. QUAST provides various metrics, including N50, the number of erroneous contigs, and the coverage of genes. However, it is not well visualized or aligned with the reference genome. Icarus integrates with QUAST, providing interactive visualizations of assembly alignments (Mikheenko et al., 2016). It can visualize genes, operons, and read coverage distribution along the genome using annotation and coverage distribution tracks and enabling a better understanding of assembly elements and algorithm behavior. However, Icarus is suitable for studies involving reference genomes and non-model organisms, where a related genome is available. Moreover, QUAST-LG, an extension of QUAST has been developed to evaluate large genome assemblies (Mikheenko et al., 2018). Despite the availability of various assembly evaluation tools for calculating individual metrics, comprehensive evaluations of multiple assembly features are surprisingly lacking. There’s no systematic way to determine the best assembly for a specific genome and dataset. Instead, several genome assemblies are generated, and the best one is based on statistics, homology analysis, and maps. This makes evaluating genome assembly challenging due to the need for multiple software packages and parameter debugging. The emergence of assembly reconciliation algorithms to improve the quality of genome assemblies brings us one step closer to a completed genome. It can merge multiple draft assemblies to enhance the contiguity and prevent errors rather than relying on random guessing among drafts (Alhakami, Mirebrahim & Lonardi, 2017). Recently, some tools have been reported that evaluate the genome quality comprehensively. The Genome Assembly Evaluation Pipeline (GAEP) (https://github.com/zy-optimistic/GAEP), WebQUAST, GenomeQC (https://github.com/HuffordLab/GenomeQC) and dnAQET (https://www.fda.gov/media/133407/download?attachment) are tools that help researchers to assess the quality and accuracy of genome assemblies comprehensively (Yavas, Hong & Xiao, 2019; Manchanda et al., 2020; Mikheenko et al., 2023; Wang et al., 2023; Zhang et al., 2023). These tools reduce the cumbersome task of installing multiple software and debugging parameters.
Overall, the best assembler and parameter settings for a specific genome and dataset cannot be determined systematically. Because the assembly parameters such as contiguity and correctness of an assembly significantly differ among different assemblers and genomes (Salzberg et al., 2012). The correlation analysis indicated that most metrics are uncorrelated, suggesting that each metric holds its distinct importance. However, a study suggests a collection of metrics with statistical indices for a thorough assessment of the assembly, offering a final assembly score as a standard for producing high-quality genome assemblies (Wang et al., 2023).
(4) 
Tools for Improving the quality of genome assembly
Long-read sequencing technologies have significantly improved genome assembly work by providing long-range evidence to resolve challenging repetitive regions in complex genomes. However, assemblies from nanopore data have high error rates, making it difficult to create accurate final sequences. Therefore, it is essential to implement error correction to prevent substantial effects on downstream analysis. Assembly pipelines often use resource-intensive error correction and consensus-generation steps to manage high error rates in these technologies (Chin et al., 2013; Loman, Quick & Simpson, 2015). Fast assemblers like Miniasm (https://github.com/lh3/miniasm/releases/tag/v0.3), Ra (https://github.com/lbcb-sci/ra/releases/tag/0.2.1) and wtdbg2 (https://github.com/ruanjue/wtdbg2/releases/tag/v2.5) are available that speed up structurally correct contig assembly by reducing read error correction time, but still have more base-level errors due to polishing for error correction (Kundu, R., Casey, J. & Sung, 2019). Therefore, polishing tools are essential for precise long-read assemblies, especially those produced by fast, error-correction-free assemblers. To improve the quality of the assembly, accurate reads from the same DNA source are recommended such as short Illumina reads or expensive PacBio HiFi reads. Various tools have been reported to improve the quality of assembly (Supplementary Table S6). These tools can be categorized as ‘Sequencer-bound’ and ‘General’. Sequencer-bound polishers require raw signal-level information from a specific sequencer, allowing them to polish reads from that sequencer such as NanoPolish from ONT, Quiver and Arrow from PacBio. On the other hand, General polishers are tools designed to enhance assembly quality by utilizing the raw reads produced by any sequencer. The general polisher includes Pilon, Racon, wtpoa-cns, and Apollo polishers (Walker et al., 2014; Vaser et al., 2017). These polishers use either short read or long read for error correction. However, some publishers have been reported to be able to utilize both types of reads.
Pilon (https://github.com/broadinstitute/pilon) and ntEdit, (https://github.com/bcgsc/ntEdit) polishers are designed to utilize short reads for improving short read assembly. However, they are not well-suited for large genomes due to their limited resources (Walker et al., 2014). POLCA and NextPolish are top-rated methods for polishing long-read assemblies using Illumina short reads. Both outperformed Pilon in correcting sequence errors using human and plant genomes (Hu et al., 2020; Zimin & Salzberg, 2020). NextPolish (https://github.com/Nextomics/NextPolish) doesn’t carry out local reassembly, it solely concentrates on fixing small-scale defects, such as small Indels and SNVs. Medaka (https://github.com/nanoporetech/medaka) uses long reads for polishing. It uses neural networks to generate consensus sequences. It outperforms sequence-graph and signal-based methods (NanoPolish) and offers better results (Medaka, 2022). Although long reads can resolve the sequence, the low confidence region was not included in the error counting in nanopore assemblies. This highlights the importance of closely examining and devising strategies to address or identify these problematic areas (Luan et al., 2024). Racon (https://github.com/isovic/racon) and wtpoa-cns (https://github.com/ruanjue/wtdbg2/releases/tag/v2.5) are hybrid polishing tools specifically designed to efficiently handle short reads and long, noisy sequences (Ruan & Li, 2020). Racon due to its ultra-fast speed and resource-efficient scaling on large genomes. It can use either long or short reads in one run. Generally, long-read polishing followed by short-read polishing is recommended for better accuracy. Moreover, Apollo can use both types of reads in a single run, but it is slow. Racon is currently the most widely used polisher due to its speed and accuracy, and its results confirm that it produces more accurate results than other polishers (Vaser et al., 2017). Polishers like wtpoa-cns and Apollo require alignment information of reads on the computationally expensive draft assembly (Warren et al., 2019). HyPo (https://github.com/kensung-lab/hypo) is a hybrid polisher that uses short and long reads to polish long-read genome assemblies, using unique genomic k-mers to selectively polish segments (Kundu et al., 2019). It produces more precise assemblies in one-third of the time and utilizes only half the memory compared to Racon. The hybrid approach can be applied in two distinct manners: one involves utilizing short reads to refine the consensus obtained from long reads, while the other incorporates short reads during the initial read correction phase of the assembly process (Luan et al., 2024). The FM-index Long Read Corrector (FMLRC) (https://github.com/holtjma/fmlrc) represents a hybrid approach to error correction. It dynamically reassembles error-prone long sequences using the Full-text Minute-space index of a Burrows-Wheeler transform built from accurate short reads and proving consistently precise and efficient polishing (Wang et al., 2018; Zhang, Jain & Aluru, 2020; Mak et al., 2023). NeuralPolish (https://github.com/huangnengCSU/NeuralPolish) is a novel polishing method that uses base-called reads to correct assembly errors by aligning them to the draft assembly. This alignment is input into a neural network for prediction by constructing an alignment matrix. Recent research reveals that NeuralPolish is more effective than other polishing methods in achieving higher accuracy in assembly and fewer errors, significantly improving the assembly precision across multiple assemblers (Huang et al., 2021).
Short-read polishing tools that utilize alignments may not be able to correct errors in repeats. Polypolish (https://github.com/rrwick/Polypolish) is a new short-read polishing tool that addresses the issue of inaccuracies in repeats by using a different type of short-read alignment as input. Instead of confining each read to a specific location, Polypolish aligns each read across all feasible positions. This method ensures that errors found in repetitive areas are managed through these alignments, allowing it to rectify mistakes that other tools might miss (Wick & Holt, 2022). Recent genome polishing tools, ntLink, and JASPER, improve draft genome assemblies using long-read sequencing data (Coombe et al., 2023; Guo, Salzberg & Zimin, 2023). ntLink (https://github.com/bcgsc/ntLink) uses minimizer-based mappings to order and orient input sequences into scaffolds, and recent improvements include overlap detection, gap-filling, and incode scaffolding iterations. It uses Bloom filters to compute and store k-mer quality information. Whereas, JASPER (https://github.com/alguoo314/jasper) is a Jellyfish-based genome polishing tool that uses a database of k-mer counts from Jellyfish polishing reads to correct assembled contigs, instead of aligning reads to the assembly. Compared to other k-mer-based and alignment-based polishing approaches, this technique is quicker and more precise, highlighting its capability and effectiveness in error detection and correction within consensus data. Various recent studies evaluate the potential of different polishers to improve assembly quality (Shafin et al., 2020; Wang et al., 2023; Wong et al., 2023; Luan et al., 2024). The recent study showed that the short-read-only contig polishing strategy outperforms long-read-only, indicating the lower error rates in Illumina short-reads. However, the study does not provide any evidence to support the claim that hybrid polishing outperforms the short-reads-only approach (Wang et al., 2023). Furthermore, it has been reported that increasing polishing iterations improves contigs’ accuracy but also increases computational resource demand. The initial round demonstrated considerable enhancement, while subsequent rounds resulted in only slight gains in sequence accuracy. Therefore, changing a polisher is better for significant assembly accuracy. However, the sequence in which the polishing tools were used is also important. It has been reported that mistakes were introduced when less accurate tools were used after more precise ones (Luan et al., 2024). In a recent study, DEGAP (Dynamic Elongation of a Genome Assembly Path) represents a novel approach to gap-filling software, designed to resolve gap regions by taking advantage of the dual properties of HiFi reads, which ensure both accuracy and extended length. This software identifies discrepancies in reads and provides two modes, ‘GapFiller’ and ‘CtgLinker,’ to eliminate or reduce gaps in genomic data. DEGAP has proven effective in analyzing intricate genomic regions across multiple projects and holds the potential for the widespread creation of gap-free genomes (Huang et al., 2024).

VI. Conclusions

  • Sequencing technologies have evolved significantly staring from Sanger’s chain-termination method to the latest TGS platforms, improving read lengths and genome finishing. Despite their benefits, new technologies still face challenges related to error rates, sequencing biases, yield, cost-effectiveness, and repetitive regions. Long-read sequencing technologies hold the potential to further genomic discoveries by providing deeper insights into genome structure, variation, and function. As these technologies advance, they will enhance our understanding of genomics, leading to more precise and comprehensive analyses across various species.
  • Genome assembly from NGS data involves complex processes and algorithms tailored to specific sequencing technologies and genome characteristics. The choice of assembly methods and tools is crucial for achieving accuracy, efficiency, and successful genome assembly outcomes.
  • A comprehensive approach, using a range of tools and metrics, is necessary to assess contiguity, accuracy, completeness, and contamination in genome assemblies. Existing tools like QUAST and BUSCO provide valuable insights, but a standardized evaluation framework is needed to ensure the reproducibility and reliability of genome assemblies.
  • The combination of long-read technologies and error correction methods continues to improve assembly quality, though establishing universal standards for genome assembly quality remains a challenge.

References

  1. Aird, D.; Ross, M. G.; Chen, W. S.; Danielsson, M.; Fennell, T.; Russ, C.; Jaffe, D.B.; Nusbaum, C.; Gnirke, A. Analyzing and minimizing PCR amplification bias in Illumina sequencing libraries. Genome biology 2011, 12, 1–14. [Google Scholar] [CrossRef]
  2. Alhakami, H.; Mirebrahim, H.; Lonardi, S. A comparative evaluation of genome assembly reconciliation tools. Genome biology 2017, 18, 1–14. [Google Scholar] [CrossRef]
  3. Alkan, C.; Sajjadian, S.; Eichler, E. E. Limitations of next-generation genome sequence assembly. Nature methods 2011, 8(1), 61–65. [Google Scholar] [PubMed]
  4. Allhoff, M.; Schonhuth, A.; Martin, M.; Costa, I. G.; Rahmann, S.; Marschall, T. Discovering motifs that induce sequencing errors. BMC bioinformatics 2013, 14, 1–10. [Google Scholar] [CrossRef]
  5. Anderson, S. Shotgun DNA sequencing using cloned DNase I-generated fragments. Nucleic acids research 1981, 9(13), 3015–3027. [Google Scholar] [CrossRef] [PubMed]
  6. Andrews, S. (2010). FastQC: a quality control tool for high throughput sequence data http://www. bioinformatics. babraham. ac. uk/projects/fastqc. Babraham Bioinformatics.
  7. Andrews, S. (2014). FastQC a quality-control tool for high-throughput sequence data http://www. Bioinformaticsbabraham. ac. uk/projects/fastqc.
  8. Baker, M. De novo genome assembly: what every biologist should know. Nature methods 2012, 9(4), 333–337. [Google Scholar] [CrossRef]
  9. Berthelot, C.; Brunet, F.; Chalopin, D.; Juanchich, A.; Bernard, M.; Noel, B.; Bento, P.; Da Silva, C.; Labadie, K.; Alberti, A.; Aury, J.M.; Louis, A.; Dehais, P.; Bardou, p.; Montfort, J.; Klopp, C. The rainbow trout genome provides novel insights into evolution after whole-genome duplication in vertebrates. Nature communications 2014, 5(1), 1–10. [Google Scholar] [CrossRef]
  10. Betschart, R. O.; Thiery, A.; Aguilera-Garcia, D.; Zoche, M.; Moch, H.; Twerenbold, R.; Zeller, T.; Blankenberg, S.; Ziegler, A. Comparison of calling pipelines for whole genome sequencing: an empirical study demonstrating the importance of mapping and alignment. Scientific Reports 2022, 12(1), 21502. [Google Scholar] [CrossRef] [PubMed]
  11. Bolger, A. M.; Lohse, M.; Usadel, B. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics 2014, 30(15), 2114–2120. [Google Scholar] [CrossRef] [PubMed]
  12. Bradnam, K. R.; Fass, J. N.; Alexandrov, A.; Baranay, P.; Bechner, M.; Birol, I.; Boisvert, S.; Chapman, J.A.; Chapuis, G.; Chikhi, R.; Chitsaz, H.; Chou, W. C.; Corbeil, J.; Fabbro, C. D.; Docking, T. R. Assemblathon 2: evaluating de novo methods of genome assembly in three vertebrate species. Gigascience 2013, 2(1), 2047–217X. [Google Scholar] [CrossRef]
  13. Cahill, M. J.; Koser, C. U.; Ross, N. E.; Archer, J. A. Read length and repeat resolution: exploring prokaryote genomes using next-generation sequencing technologies. PloS one 2010, 5(7), e11518. [Google Scholar] [CrossRef] [PubMed]
  14. Chen, J.; Li, X.; Zhong, H.; Meng, Y.; Du, H. Systematic comparison of germline variant calling pipelines cross multiple next-generation sequencers. Scientific reports 2019, 9(1), 9345. [Google Scholar] [CrossRef] [PubMed]
  15. Chen, S.; Zhou, Y.; Chen, Y.; Gu, J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 2018, 34(17), i884–i890. [Google Scholar] [CrossRef] [PubMed]
  16. Chen, Y. C.; Liu, T.; Yu, C. H.; Chiang, T. Y.; Hwang, C. C. Effects of GC bias in next-generation-sequencing data on de novo genome assembly. PloS one 2013, 8(4), e62856. [Google Scholar] [PubMed]
  17. Chen, Y.; Chen, Y.; Shi, C.; Huang, Z.; Zhang, Y.; Li, S.; Li, Y.; Ye, J.; Yu, C.; Li, Z.; Zhang, X.; Wang, J.; Yang, H.; Fang, L.; Chen, Q. SOAPnuke: a MapReduce acceleration-supported software for integrated quality control and preprocessing of high-throughput sequencing data. Gigascience 2018, 7(1), gix120. [Google Scholar] [PubMed]
  18. Chen, Y.; Zhang, Y.; Wang, A. Y.; Gao, M.; Chong, Z. Accurate long-read de novo assembly evaluation with Inspector. Genome Biology 2021, 22, 1–21. [Google Scholar] [CrossRef]
  19. Cheng, H.; Concepcion, G. T.; Feng, X.; Zhang, H.; Li, H. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nature methods 2021, 18(2), 170–175. [Google Scholar] [CrossRef] [PubMed]
  20. Chin, C. S.; Khalak, A. Human genome assembly in 100 minutes. BioRxiv 705616; 2019.
  21. Chin, C. S.; Alexander, D. H.; Marks, P.; Klammer, A. A.; Drake, J.; Heiner, C.; Clum, A.; Copeland, A.; Huddleston, J.; Eichler, E. E.; Turner, S. W.; Korlach, J. Nonhybrid, finished microbial genome assemblies from long-read SMRT sequencing data. Nature methods 2013, 10(6), 563–569. [Google Scholar] [CrossRef] [PubMed]
  22. Cho, Y. S.; Kim, H.; Kim, H. M.; Jho, S.; Jun, J.; Lee, Y. J.; Chae, K. S.; Kim, C. G.; Kim, S.; Eriksson, A.; Edwards, J. S.; Lee, S.
  23. Choi, I. Y.; Kwon, E. C.; Kim, N. S. The C-and G-value paradox with polyploidy, repeatomes, introns, phenomes and cell economy. Genes & genomics 2020, 42, 699–714. [Google Scholar]
  24. Compeau, P. E.; Pevzner, P. A.; Tesler, G. How to apply de Bruijn graphs to genome assembly. Nature biotechnology 2011, 29(11), 987–991. [Google Scholar] [CrossRef] [PubMed]
  25. Coombe, L.; Warren, R. L.; Wong, J.; Nikolic, V.; Birol, I. ntLink: a toolkit for de novo genome assembly scaffolding and mapping using long reads. Current Protocols 2023, 3(4), e733. [Google Scholar] [CrossRef] [PubMed]
  26. Corradi, N.; Pombert, J. F.; Farinelli, L.; Didier, E. S.; Keeling, P. J. The complete sequence of the smallest known nuclear genome from the microsporidian Encephalitozoon intestinalis. Nature communications 2010, 1(1), 77. [Google Scholar] [CrossRef] [PubMed]
  27. Cox, M. P.; Peterson, D. A.; Biggs, P. J. SolexaQA: At-a-glance quality assessment of Illumina second-generation sequencing data. BMC bioinformatics 2010, 11, 1–6. [Google Scholar]
  28. Del Angel, V. D.; Hjerde, E.; Sterck, L.; Capella-Gutierrez, S.; Notredame, C.; Pettersson, O. V.; Amselem, J.; Bouri, L.; Bocs, S.; Klopp, C.; Gibrat, J. F.; Vlasova, A.; Leskosek, B. L.; Soler, L.; Binzer-Panchal, M.; Lantz, H. Ten steps to get started in Genome Assembly and Annotation; F1000Research 7, 2018. [Google Scholar]
  29. Dida, F.; Yi, G. Empirical evaluation of methods for de novo genome assembly. PeerJ Computer Science 2021, 7, e636. [Google Scholar] [PubMed]
  30. Drmanac, R.; Sparks, A. B.; Callow, M. J.; Halpern, A. L.; Burns, N. L.; Kermani, B. G.; Carnevali, P.; Nazarenko, I.; Nilsen, G. B.; Yeung, G.; Dahl, F.; Fernandez, A.; Staker, B.; Pant, K. P.; Baccash, J. Human genome sequencing using unchained base reads on self-assembling DNA nanoarrays. Science 2010, 327(5961), 78–81. [Google Scholar] [CrossRef] [PubMed]
  31. Earl, D.; Bradnam, K.; St John, J.; Darling, A.; Lin, D.; Fass, J.; Yu, H. O.; Buffalo, V.; Zerbino, D. R.; Diekhans, M.; Nguyen, N.; Ariyaratne, P. N.; Sung, W. K.; Ning, Z.; Haimel, M. Assemblathon 1: a competitive assessment of de novo short read assembly methods. Genome Research 2011, 21(12), 2224–2241. [Google Scholar] [CrossRef] [PubMed]
  32. Elliott, T. A.; Gregory, T. R. What’s in a genome? The C-value enigma and the evolution of eukaryotic genome content. Philosophical Transactions of the Royal Society B: Biological Sciences 2015, 370(1678), 20140331. [Google Scholar]
  33. El-Metwally, S.; Hamouda, E.; Tarek, M. A roadmap to sequence assembly evaluation tools. Current Bioinformatics 2021, 16(5), 644–661. [Google Scholar] [CrossRef]
  34. Fabbro, C.; Scalabrin, S.; Morgante, M.; Giorgi, F. M. An extensive evaluation of read trimming effects on Illumina NGS data analysis. PloS one 2013, 8(12), e85024. [Google Scholar] [CrossRef] [PubMed]
  35. Ferreira, C. R. The burden of rare diseases. American journal of medical genetics Part A 2019, 179(6), 885–892. [Google Scholar] [CrossRef] [PubMed]
  36. Firtina, C.; Kim, J. S.; Alser, M.; Senol Cali, D.; Cicek, A. E.; Alkan, C.; Mutlu, O. Apollo: a sequencing-technology-independent, scalable and accurate assembly polishing algorithm. Bioinformatics 2020, 36(12), 3669–3679. [Google Scholar] [PubMed]
  37. Flicek, P.; Birney, E. Sense from sequence reads: methods for alignment and assembly. Nature methods 2009, 6 (Suppl 11), S6–S12. [Google Scholar] [CrossRef] [PubMed]
  38. Formenti, G.; Theissinger, K.; Fernandes, C.; Bista, I.; Bombarely, A.; Bleidorn, C.; Ciofi, C.; Crottini, A.; Godoy, J. A.; Hoglund, J.; Malukiewicz, J.; Mouton, A.; Oomen, R. A.; Paez, S.; Palsboll, P. J. The era of reference genomes in conservation genomics. Trends in ecology & evolution 2022, 37(3), 197–202. [Google Scholar] [CrossRef]
  39. Gaia, A. S. C.; de Oliveira, M. S.; da Silva Moia, G.; dos Santos, V. C.; Alves, J. T. C.; de Sá, P. H. C. G.; de Oliveira Veras, A. A. e-QA NGS: a user-friendly tool to preprocessing data from next generation sequencing. Peer Review 2023, 5(3), 91–105. [Google Scholar]
  40. Gao, S. H.; Yu, H. Y.; Wu, S. Y.; Wang, S.; Geng, J. N.; Luo, Y. F.; Hu, S. N. Advances of sequencing and assembling technologies for complex genomes. Yi Chuan Hereditas 2018, 40(11), 944–963. [Google Scholar] [PubMed]
  41. Gavrielatos, M.; Kyriakidis, K.; Spandidos, D. A.; Michalopoulos, I. Benchmarking of next and third generation sequencing technologies and their associated algorithms for de novo genome assembly. Molecular Medicine Reports 2021, 23(4), 1–1. [Google Scholar] [CrossRef]
  42. Genome Reference Consortium. (2022-05-09). “GenomeRef: GRCh38.p14 is now released!”. GRC Blog (GenomeRef). Retrieved 2022-08-19.
  43. Genova, A. D.; Buena-Atienza, E.; Ossowski, S.; Sagot, M. F. Efficient hybrid de novo assembly of human genomes with WENGAN. Nature Biotechnology 2021, 39(4), 422–430. [Google Scholar] [PubMed]
  44. Giordano, F.; Aigrain, L.; Quail, M. A.; Coupland, P.; Bonfield, J. K.; Davies, R. M.; Tischler, G.; Jackson, D. K.; Keane, T. M.; Li, J.; Yue, J. X. De novo yeast genome assemblies from MinION, PacBio and MiSeq platforms. Scientific reports 2017, 7(1), 3935. [Google Scholar] [CrossRef] [PubMed]
  45. Giorgi, F. M.; Del-Fabbro, C.; Licausi, F. Comparative study of RNA-seq-and microarray-derived coexpression networks in Arabidopsis thaliana. Bioinformatics 2013, 29(6), 717–724. [Google Scholar] [CrossRef] [PubMed]
  46. Glusman, G.; Cox, H. C.; Roach, J. C. Whole-genome haplotyping approaches and genomic medicine. Genome medicine 2014, 6, 1–16. [Google Scholar] [CrossRef]
  47. Gnerre, S.; Maccallum, I.; Przybylski, D.; Ribeiro, F. J.; Burton, J. N.; Walker, B. J.; Sharpe, T.; Hall, G.; Shea, T. P.; Sykes, S.; Berlin, A. M.; Aird, D.; Costello, M.; Daza, R.; Williams, L. High-quality draft assemblies of mammalian genomes from massively parallel sequence data. Proceedings of the National Academy of Sciences of the United States of America 2011, 108(4), 1513–1518. [Google Scholar] [PubMed]
  48. Gonzalez-Garcia, L.; Guevara-Barrientos, D.; Lozano-Arce, D.; Gil, J.; Díaz-Riaño, J.; Duarte, E.; Andrade, G.; Bojaca, J. C.; Hoyos-Sanchez, M. C.; Chavarro, C.; Guayazan, N.; Chica, L. A.; Acosta, M. C. B.; Bautista, E.; Trujillo, M.; Duitama, J. New algorithms for accurate and efficient de novo genome assembly from long DNA sequencing reads. Life Science Alliance 2023, 6(5). [Google Scholar] [CrossRef] [PubMed]
  49. Goodwin, S.; McPherson, J. D.; McCombie, W. R. Coming of age: ten years of next-generation sequencing technologies. Nature reviews genetics 2016, 17(6), 333–351. [Google Scholar] [CrossRef] [PubMed]
  50. Guo, A.; Salzberg, S. L.; Zimin, A. V. JASPER: A fast genome polishing tool that improves accuracy of genome assemblies. PLoS computational biology 2023, 19(3), e1011032. [Google Scholar] [CrossRef] [PubMed]
  51. Gurevich, A.; Saveliev, V.; Vyahhi, N.; Tesler, G. QUAST: quality assessment tool for genome assemblies. Bioinformatics 2013, 29(8), 1072–1075. [Google Scholar] [CrossRef] [PubMed]
  52. Haghshenas, E.; Asghari, H.; Stoye, J.; Chauve, C.; Hach, F. HASLR: fast hybrid assembly of long reads. Iscience 2020, 23(8). [Google Scholar] [CrossRef] [PubMed]
  53. Han, K.; Li, Z. F.; Peng, R.; Zhu, L. P.; Zhou, T.; Wang, L. G.; Li, S. G.; Zhang, X. B.; Hu, W.; Wu, Z. H.; Qin, N.; Li, Y. Z. Extraordinary expansion of a Sorangium cellulosum genome from an alkaline milieu. Scientific reports 2013, 3(1), 2101. [Google Scholar] [CrossRef] [PubMed]
  54. He, Y.; Zhang, Z.; Peng, X.; Wu, F.; Wang, J. De novo assembly methods for next generation sequencing data. Tsinghua Science and Technology 2013, 18(5), 500–514. [Google Scholar] [CrossRef]
  55. Heydari, M.; Miclotte, G.; Demeester, P.; Van de Peer, Y.; Fostier, J. Evaluation of the impact of Illumina error correction tools on de novo genome assembly. BMC bioinformatics 2017, 18, 1–13. [Google Scholar] [CrossRef]
  56. Horner, D. S.; Pavesi, G.; Castrignano, T.; De Meo, P. D. O.; Liuni, S.; Sammeth, M.; Picardi, E.; Pesole, G. Bioinformatics approaches for genomics and post genomics applications of next-generation sequencing. Briefings in bioinformatics 2010, 11(2), 181–197. [Google Scholar] [PubMed]
  57. Hotaling, S.; Kelley, J. L.; Frandsen, P. B. Toward a genome sequence for every animal: Where are we now? Proceedings of the National Academy of Sciences 2021, 118(52), e2109019118. [Google Scholar] [CrossRef]
  58. Hu, J.; Fan, J.; Sun, Z.; Liu, S. NextPolish: a fast and efficient genome polishing tool for long-read assembly. Bioinformatics 2020, 36(7), 2253–2255. [Google Scholar] [PubMed]
  59. Huang, J.; Liang, X.; Xuan, Y.; Geng, C.; Li, Y.; Lu, H.; Qu, S.; Mei, X.; Chen, H.; Yu, T.; Sun, N.; Rao, J.; Wang, J.; Zhang, W.; Chen, Y. A reference human genome dataset of the BGISEQ-500 sequencer. GigaScience 2017, 6(5), 1–9. [Google Scholar] [CrossRef] [PubMed]
  60. Huang, N.; Nie, F.; Ni, P.; Luo, F.; Gao, X.; Wang, J. NeuralPolish: a novel Nanopore polishing method based on alignment matrix construction and orthogonal Bi-GRU Networks. Bioinformatics 2021, 37(19), 3120–3127. [Google Scholar] [PubMed]
  61. Huang, Y.; Wang, Z.; Schmidt, M. A.; Su, H.; Xiong, L.; Zhang, J. DEGAP: Dynamic elongation of a genome assembly path. Briefings in Bioinformatics 2024, 25(3), bbae194. [Google Scholar] [CrossRef] [PubMed]
  62. Huddleston, J.; Ranade, S.; Malig, M.; Antonacci, F.; Chaisson, M.; Hon, L.; Sudmant, P. H.; Graves, T. A.; Alkan, C.; Dennis, M. Y.; Wilson, R. K.; Turner, S. W.; Korlach, J.; Eichler, E. E. Reconstructing complex regions of genomes using long-read sequencing technology. Genome Research 2014, 24(4), 688–696. [Google Scholar] [CrossRef] [PubMed]
  63. Hunt, M.; Kikuchi, T.; Sanders, M.; Newbold, C.; Berriman, M.; Otto, T. D. REAPR: a universal tool for genome assembly evaluation. Genome biology 2013, 14, 1–10. [Google Scholar] [CrossRef]
  64. Jackman, S. D.; Birol, İ. Assembling genomes using short-read sequencing technology. Genome Biology 2010, 11, 1–4. [Google Scholar] [CrossRef]
  65. Jain, M.; Fiddes, I. T.; Miga, K. H.; Olsen, H. E.; Paten, B.; Akeson, M. Improved data analysis for the MinION nanopore sequencer. Nature methods 2015, 12(4), 351–356. [Google Scholar] [CrossRef] [PubMed]
  66. Jain, M.; Koren, S.; Miga, K. H.; Quick, J.; Rand, A. C.; Sasani, T. A.; Tyson, J. R.; Beggs, A. D.; Dilthey, A. T.; Fiddes, I. T.; Malla, S.; Marriott, H.; Nieto, T.; O’Grady, J.; Olsen, H. E. Nanopore sequencing and assembly of a human genome with ultra-long reads. Nature biotechnology 2018, 36(4), 338–345. [Google Scholar] [CrossRef] [PubMed]
  67. Jarvis, E. D.; Formenti, G.; Rhie, A.; Guarracino, A.; Yang, C.; Wood, J.; Tracey, A.; Thibaud-Nissen, F.; Vollger, M. R.; Porubsky, D.; Cheng, H.; Asri, M.; Logsdon, G. A.; Carnevali, P.; Chaisson, M. J. P. Semi-automated assembly of high-quality diploid human reference genomes. Nature 2022, 611(7936), 519–531. [Google Scholar] [PubMed]
  68. Jung, H.; Ventura, T.; Chung, J. S.; Kim, W. J.; Nam, B. H.; Kong, H. J.; Kim, Y. O.; Jeon, M. S.; Eyun, S. I. Twelve quick steps for genome assembly and annotation in the classroom. PLoS computational biology 2020, 16(11), e1008325. [Google Scholar] [CrossRef] [PubMed]
  69. Jung, H.; Winefield, C.; Bombarely, A.; Prentis, P.; Waterhouse, P. Tools and strategies for long-read sequencing and de novo assembly of plant genomes. Trends in plant science 2019, 24(8), 700–724. [Google Scholar] [CrossRef] [PubMed]
  70. Karlsson, F. H.; Fåk, F.; Nookaew, I.; Tremaroli, V.; Fagerberg, B.; Petranovic, D.; Backhed, F.; Nielsen, J. Symptomatic atherosclerosis is associated with an altered gut metagenome. Nature communications 2012, 3(1), 1245. [Google Scholar] [CrossRef] [PubMed]
  71. Kelley, D. R.; Schatz, M. C.; Salzberg, S. L. Quake: quality-aware detection and correction of sequencing errors. Genome biology 2010, 11, 1–13. [Google Scholar] [CrossRef]
  72. Khan, A. R.; Pervez, M. T.; Babar, M. E.; Naveed, N.; Shoaib, M. A comprehensive study of de novo genome assemblers: current challenges and future prospective. Evolutionary Bioinformatics 2018, 14, 1176934318758650. [Google Scholar] [CrossRef] [PubMed]
  73. Khelik, K.; Sandve, G. K.; Nederbragt, A. J.; Rognes, T. NucBreak: location of structural errors in a genome assembly by using paired-end Illumina reads. BMC bioinformatics 2020, 21, 1–11. [Google Scholar]
  74. Khiste, N.; Ilie, L. Laser: Large genome assembly evaluator. BMC research notes 2015, 8, 1–3. [Google Scholar] [CrossRef]
  75. Kim, B. C.; Manica, A.; Oh, T. K. An ethnically relevant consensus Korean reference genome is a step towards personal reference genomes. Nature communications 2016, 7(1), 1–13. [Google Scholar] [CrossRef]
  76. Kim, H.M.; Jeon, S.; Chung, O.; Jun, J.H.; Kim, H.S.; Blazyte, A.; Lee, H.Y.; Yu, Y.; Cho, Y.S.; Bolser, D.M.; Bhak, J. Comparative analysis of 7 short-read sequencing platforms using the Korean Reference Genome: MGI and Illumina sequencing benchmark for whole-genome sequencing. GigaScience 2021, 10(3), 1–9. [Google Scholar]
  77. Kundu, R.; Casey, J.; Sung, W. K. HyPo: super fast & accurate polisher for long read genome assemblies. Biorxiv 2019, 20, 2019–12. [Google Scholar]
  78. Lakhotia, S. C. C-value paradox: Genesis in misconception that natural selection follows anthropocentric parameters of ‘economy’and ‘optimum’. BBA advances 2023, 4, 100107. [Google Scholar] [PubMed]
  79. Lander, E.S.; Linton, L.M.; Birren, B. Initial sequencing and analysis of the human genome. Nature 2001, 409, 860–921. [Google Scholar] [CrossRef] [PubMed]
  80. Li, H.; Durbin, R. Fast and accurate long-read alignment with Burrows–Wheeler transform. Bioinformatics 2010, 26(5), 589–595. [Google Scholar] [PubMed]
  81. Li, R.; Di, L.; Li, J.; Fan, W.; Liu, Y.; Guo, W.; Liu, W.; Liu, L.; Li, Q.; Chen, L.; Chen, Y. A body map of somatic mutagenesis in morphologically normal human tissues. Nature 2021, 597(7876), 398–403. [Google Scholar] [CrossRef] [PubMed]
  82. Li, R.; Zhu, H.; Ruan, J.; Qian, W.; Fang, X.; Shi, Z.; Li, Y.; Li, S.; Shan, G.; Kristiansen, K.; Li, S.; Yang, H.; Wang, J.; Wang, J. De novo assembly of human genomes with massively parallel short read sequencing. Genome Research 2010, 20(2), 265–272. [Google Scholar] [PubMed]
  83. Li, Z.; Chen, Y.; Mu, D.; Yuan, J.; Shi, Y.; Zhang, H.; Gan, J.; Li, N.; Hu, X.; Liu, B.; Yang, B.; Fan, W. Comparison of the two major classes of assembly algorithms: overlap-layout-consensus and de-bruijn-graph. Briefings in functional genomics 2012, 11(1), 25–37. [Google Scholar] [PubMed]
  84. Liao, X.; Li, M.; Zou, Y.; Wu, F. X.; Wang, J. Current challenges and solutions of de novo assembly. Quantitative Biology 2019, 7(2), 90–109. [Google Scholar] [CrossRef]
  85. Liao, X.; Li, M.; Zou, Y.; Wu, F.; Pan, Y.; Luo, F.; Wang, J. Improving de novo Assembly Based on Read Classification. IEEE/ACM Transactions on Computational Biology and Bioinformatics 2020, 17(1), 177–188. [Google Scholar] [PubMed]
  86. Linsmith, G.; Rombauts, S.; Montanari, S.; Deng, C. H.; Celton, J. M.; Guérif, P.; Liu, C.; Lohaus, R.; Zurn, J. D.; Cestaro, A.; Bassil, N. V.; Bakker, L. V.; Schijlen, E.; Gardiner, S. E.; Lespinasse, Y. Pseudo-chromosome–length genome assembly of a double haploid “Bartlett” pear (Pyrus communis L.). Gigascience 2019, 8(12), giz138. [Google Scholar] [CrossRef] [PubMed]
  87. Liu, H.; Wang, X.; Wang, G.; Cui, P.; Wu, S.; Ai, C.; Hu, N.; Li, A.; He, B.; Shao, X.; Wu, Z. The nearly complete genome of Ginkgo biloba illuminates gymnosperm evolution. Nature Plants 2021, 7(6), 748–756. [Google Scholar] [CrossRef] [PubMed]
  88. Liu, J.; Seetharam, A.S.; Chougule, K.; Ou, S.; Swentowsky, K.W.; Gent, J.I.; Llaca, V.; Woodhouse, M.R.; Manchanda, N.; Presting, G.G.; Kudrna, D.A. Gapless assembly of maize chromosomes using long-read technologies. Genome Biology 2020, 21, 1–17. [Google Scholar] [CrossRef]
  89. Liu, L.; Li, Y.; Li, S.; Hu, N.; He, Y.; Pong, R.; Lin, D.; Lu, L.; Law, M. Comparison of next-generation sequencing systems. BioMed research international 2012, 2012(1), 251364. [Google Scholar] [CrossRef]
  90. Liu, S.; Li, W.; Wu, Y.; Chen, C.; Lei, J. De novo transcriptome assembly in chili pepper (Capsicum frutescens) to identify genes involved in the biosynthesis of capsaicinoids. PloS one 2013, 8(1), e48156. [Google Scholar] [CrossRef] [PubMed]
  91. Logsdon, G. A.; Vollger, M. R.; Eichler, E. E. Long-read human genome sequencing and its applications. Nature Reviews Genetics 2020, 21(10), 597–614. [Google Scholar] [CrossRef] [PubMed]
  92. Loman, N. J.; Quick, J.; Simpson, J. T. A complete bacterial genome assembled de novo using only nanopore sequencing data. Nature methods 2015, 12(8), 733–735. [Google Scholar] [CrossRef] [PubMed]
  93. Luan, T.; Commichaux, S.; Hoffmann, M.; Jayeola, V.; Jang, J. H.; Pop, M.; Rand, H.; Luo, Y. Benchmarking short and long read polishing tools for nanopore assemblies: achieving near-perfect genomes for outbreak isolates. BMC genomics 2024, 25(1), 679. [Google Scholar] [CrossRef] [PubMed]
  94. Lucas, S. J.; Akpınar, B. A.; Simkova, H.; Kubaláková, M.; Dolezel, J.; Budak, H. Next-generation sequencing of flow-sorted wheat chromosome 5D reveals lineage-specific translocations and widespread gene duplications. BMC genomics 2014, 15, 1–18. [Google Scholar]
  95. Luo, R.; Liu, B.; Xie, Y.; Li, Z.; Huang, W.; Yuan, J.; He, G.; Chen, Y.; Pan, Q.; Liu, Y.; Tang, J.; Wu, G.; Zhang, H.; Shi, Y.; Liu, Y. SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler. Gigascience 2012, 1(1), 2047–217X. [Google Scholar]
  96. Ma, Z.; Zhang, Y.; Wu, L.; Zhang, G.; Sun, Z.; Li, Z.; Jiang, Y.; Ke, H.; Chen, B.; Liu, Z.; Gu, Q. High-quality genome assembly and resequencing of modern cotton cultivars provide resources for crop improvement. Nature Genetics 2021b, 53(9), 1385–1391. [Google Scholar]
  97. Mak, Q. C.; Wick, R. R.; Holt, J. M.; Wang, J. R. Polishing de novo nanopore assemblies of bacteria and eukaryotes with FMLRC2. Molecular Biology and Evolution 2023, 40(3), msad048. [Google Scholar] [CrossRef] [PubMed]
  98. Manchanda, N.; Portwood, J. L.; Woodhouse, M. R.; Seetharam, A. S.; Lawrence-Dill, C. J.; Andorf, C. M.; Hufford, M. B. GenomeQC: a quality assessment tool for genome assemblies and gene structure annotations. BMC genomics 2020, 21, 1–9. [Google Scholar] [CrossRef]
  99. Mantere, T.; Kersten, S.; Hoischen, A. Long-read sequencing emerging in medical genetics. Frontiers in genetics 2019, 10, 432668. [Google Scholar] [CrossRef] [PubMed]
  100. Mapleson, D.; Garcia Accinelli, G.; Kettleborough, G.; Wright, J.; Clavijo, B. J. KAT: a K-mer analysis toolkit to quality control NGS datasets and genome assemblies. Bioinformatics 2017, 33(4), 574–576. [Google Scholar] [PubMed]
  101. Margulies, M.; Egholm, M.; Altman, W. E.; Attiya, S.; Bader, J. S.; Bemben, L. A.; Berka, J.; Braverman, M. S.; Chen, Y. J.; Chen, Z.; Dewell, S. B.; Du, L.; Fierro, J. M.; Gomes, X. V.; Godwin, B. C. Genome sequencing in microfabricated high-density picolitre reactors. Nature 2005, 437(7057), 376–380. [Google Scholar] [CrossRef] [PubMed]
  102. Marks, P.; Garcia, S.; Barrio, A. M.; Belhocine, K.; Bernate, J.; Bharadwaj, R.; Bjornson, K.; Catalanotti, C.; Delaney, J.; Fehr, A.; Fiddes, I. T.; Galvin, B.; Heaton, H.; Herschleb, J.; Hindson, C. Resolving the full spectrum of human genome variation using Linked-Reads. Genome Research 2019, 29(4), 635–645. [Google Scholar] [CrossRef] [PubMed]
  103. Martin, M. Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet. journal 2011, 17(1), 10–12. [Google Scholar] [CrossRef]
  104. McCormick, K. A.; Calzone, K. A. The impact of genomics on health outcomes, quality, and safety. Nursing management 2016, 47(4), 23–26. [Google Scholar] [CrossRef] [PubMed]
  105. McCoy, R. C.; Taylor, R. W.; Blauwkamp, T. A.; Kelley, J. L.; Kertesz, M.; Pushkarev, D.; Petrov, D. A.; Fiston-Lavier, A. S. Illumina TruSeq synthetic long-reads empower de novo assembly and resolve complex, highly-repetitive transposable elements. PloS one 2014, 9(9), e106689. [Google Scholar] [PubMed]
  106. Medaka, O. N. T. (2018). sequence correction provided by ONT Research. Github. Available from: https://github. com/nanoporetech/medaka.
  107. Mikheenko, A.; Prjibelski, A.; Saveliev, V.; Antipov, D.; Gurevich, A. Versatile genome assembly evaluation with QUAST-LG. Bioinformatics 2018, 34(13), i142–i150. [Google Scholar] [CrossRef] [PubMed]
  108. Mikheenko, A.; Saveliev, V.; Hirsch, P.; Gurevich, A. WebQUAST: online evaluation of genome assemblies. Nucleic Acids Research 2023, 51(W1), 601–606. [Google Scholar] [CrossRef]
  109. Mikheenko, A.; Valin, G.; Prjibelski, A.; Saveliev, V.; Gurevich, A. Icarus: visualizer for de novo assembly evaluation. Bioinformatics 2016, 32(21), 3321–3323. [Google Scholar] [CrossRef] [PubMed]
  110. Miller, J. R.; Koren, S.; Sutton, G. Assembly algorithms for next-generation sequencing data. Genomics 2010, 95(6), 315–327. [Google Scholar] [CrossRef] [PubMed]
  111. Nakamura, K.; Oshima, T.; Morimoto, T.; Ikeda, S.; Yoshikawa, H.; Shiwa, Y.; Ishikawa, S.; Linak, M. C.; Hirai, A.; Takahashi, H.; Altaf-Ul-Amin, Md.; Ogasawara, N.; Kanaya, S. Sequence-specific error profile of Illumina sequencers. Nucleic acids research 2011, 39(13), e90–e90. [Google Scholar] [CrossRef] [PubMed]
  112. Narzisi, G.; Mishra, B. Comparing de novo genome assembly: the long and short of it. PloS one 2011, 6(4), e19175. [Google Scholar] [CrossRef] [PubMed]
  113. Neale, D. B.; Wegrzyn, J. L.; Stevens, K. A.; Zimin, A. V.; Puiu, D.; Crepeau, M. W.; Cardeno, C.; Koriabine, M.; Holtz-Morris, A. E.; Liechty, J. D.; Martínez-García, P. J.; Vasquez-Gross, H. A.; Lin, B. Y.; Zieve, J. J.; Dougherty, W. M. Decoding the massive genome of loblolly pine using haploid DNA and novel assembly strategies. Genome biology 2014, 15, 1–13. [Google Scholar] [CrossRef]
  114. Niederst, M. J.; Hu, H.; Mulvey, H. E.; Lockerman, E. L.; Garcia, A. R.; Piotrowska, Z.; Sequist, L. V.; Engelman, J. A. The allelic context of the C797S mutation acquired upon treatment with third-generation EGFR inhibitors impacts sensitivity to subsequent treatment strategies. Clinical Cancer Research 2015, 21(17), 3924–3933. [Google Scholar] [PubMed]
  115. Nurk, S.; Koren, S.; Rhie, A.; Rautiainen, M.; Bzikadze, A. V.; Mikheenko, A.; Vollger, M. R.; Altemose, N.; Uralsky, L.; Gershman, A.; Aganezov, S.; Hoyt, S. J.; Diekhans, M.; Logsdon, G. A.; Alonge, M. The complete sequence of a human genome. Science 2022, 376(6588), 44–53. [Google Scholar] [CrossRef] [PubMed]
  116. Oppenheimer, J.; Rosen, B.D.; Heaton, M.P.; Vander Ley, B.L.; Shafer, W.R.; Schuetze, F.T.; Stroud, B.; Kuehn, L.A.; McClure, J.C.; Barfield, J.P.; Blackburn, H.D. A reference genome assembly of American Bison, Bison bison bison. Journal of Heredity 2021, 112(2), 174–183. [Google Scholar] [CrossRef] [PubMed]
  117. Ou, S.; Chen, J.; Jiang, N. Assessing genome assembly quality using the LTR Assembly Index (LAI). Nucleic acids research 2018, 46(21), e126–e126. [Google Scholar] [CrossRef] [PubMed]
  118. Padovani de Souza, K.; Setubal, J. C.; Ponce de Leon F de Carvalho, A. C.; Oliveira, G.; Chateau, A.; Alves, R. Machine learning meets genome assembly. Briefings in bioinformatics 2019, 20(6), 2116–2129. [Google Scholar] [PubMed]
  119. Parra, G.; Bradnam, K.; Korf, I. CEGMA: a pipeline to accurately annotate core genes in eukaryotic genomes. Bioinformatics 2007, 23(9), 1061–1067. [Google Scholar] [CrossRef] [PubMed]
  120. Paszkiewicz, K.; Studholme, D. J. De novo assembly of short sequence reads. Briefings in bioinformatics 2010, 11(5), 457–472. [Google Scholar] [CrossRef] [PubMed]
  121. Patch, A. M.; Nones, K.; Kazakoff, S. H.; Newell, F.; Wood, S.; Leonard, C.; Holmes, O.; Xu, Q.; Addala, V.; Creaney, J.; Robinson, B. W.; Fu, S.; Geng, C.; Li, T.; Zhang, W. Germline and somatic variant identification using BGISEQ-500 and HiSeq X Ten whole genome sequencing. PloS one 2018, 13(1), e0190264. [Google Scholar] [CrossRef] [PubMed]
  122. Phillippy, A. M.; Schatz, M. C.; Pop, M. Genome assembly forensics: finding the elusive mis-assembly. Genome biology 2008, 9, 1–13. [Google Scholar] [CrossRef]
  123. Porubsky, D.; Ebert, P.; Audano, P. A.; Vollger, M. R.; Harvey, W. T.; Marijon, P.; Ebler, J.; Munson, K. M.; Sorensen, M.; Sulovari, A.; Haukness, M.; Ghareghani, M.; Human Genome Structural Variation Consortium; Lansdorp, P. M.; Paten, B. Fully phased human genome assembly without parental data using single-cell strand sequencing and long reads. Nature biotechnology 2021, 39(3), 302–308. [Google Scholar] [PubMed]
  124. Putnam, N. H.; O’Connell, B. L.; Stites, J. C.; Rice, B. J.; Blanchette, M.; Calef, R.; Troll, C. J.; Fields, A.; Hartley, P. D.; Sugnet, C. W.; Haussler, D.; Rokhsar, D. S.; Green, R. E. Chromosome-scale shotgun assembly using an in vitro method for long-range linkage. Genome research 2016, 26(3), 342–350. [Google Scholar] [PubMed]
  125. Qin, L.; Hu, Y.; Wang, J.; Wang, X.; Zhao, R.; Shan, H.; Li, K.; Xu, P.; Wu, H.; Yan, X.; Liu, L. Insights into angiosperm evolution, floral development and chemical biosynthesis from the Aristolochia fimbriata genome. Nature Plants 2021a, 7(9), 1239–1253. [Google Scholar]
  126. Qin, P.; Lu, H.; Du, H.; Wang, H.; Chen, W.; Chen, Z.; He, Q.; Ou, S.; Zhang, H.; Li, X.; Li, X. Pan-genome analysis of 33 genetically diverse rice accessions reveals hidden genomic variations. Cell 2021b, 184(13), 3542–3558. [Google Scholar]
  127. Ren, J.; Chaisson, M. J. lra: A long read aligner for sequences and contigs. PLOS Computational Biology 2021, 17(6), e1009078. [Google Scholar] [CrossRef] [PubMed]
  128. Rhie, A.; Walenz, B. P.; Koren, S.; Phillippy, A. M. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. Genome biology 2020, 21, 1–27. [Google Scholar] [CrossRef]
  129. Riess, O.; Sturm, M.; Menden, B.; Liebmann, A.; Demidov, G.; Witt, D.; Casadei, N.; Admard, J.; Schutz, L.; Ossowski, S.; Taylor, S.; Schaffer, S.; Schroeder, C.; Dufke, A.; Haack, T. Genomes in clinical care. Npj Genomic Medicine 2024, 9(1), 20. [Google Scholar] [CrossRef] [PubMed]
  130. Ruan, J.; Li, H. Fast and accurate long-read assembly with wtdbg2. Nature methods 2020, 17(2), 155–158. [Google Scholar] [PubMed]
  131. Salzberg, S. L.; Phillippy, A. M.; Zimin, A.; Puiu, D.; Magoc, T.; Koren, S.; Treangen, T. J.; Schatz, M. C.; Delcher, A. L.; Roberts, M.; Marçais, G.; Pop, M.; Yorke, J. A. GAGE: A critical evaluation of genome assemblies and assembly algorithms. Genome Research 2012, 22(3), 557–567. [Google Scholar] [PubMed]
  132. Sanger, F.; Coulson, A. R. A rapid method for determining sequences in DNA by primed synthesis with DNA polymerase. Journal of molecular biology 1975, 94(3), 441–448. [Google Scholar] [CrossRef] [PubMed]
  133. Schmieder, R.; Edwards, R. Fast identification and removal of sequence contamination from genomic and metagenomic datasets. PloS one 2011, 6(3), e17288. [Google Scholar] [CrossRef] [PubMed]
  134. Schneider, V. A.; Graves-Lindsay, T.; Howe, K.; Bouk, N.; Chen, H. C.; Kitts, P. A.; Murphy, T. D.; Pruitt, K. D.; Thibaud-Nissen, F.; Albracht, D.; Fulton, R. S.; Kremitzki, M.; Magrini, V.; Markovic, C.; McGrath, S. Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly. Genome research 2017, 27(5), 849–864. [Google Scholar] [CrossRef] [PubMed]
  135. Schott, T.; Kondadi, P. K.; Hänninen, M. L.; Rossi, M. Microevolution of a zoonotic Helicobacter population colonizing the stomach of a human host before and after failed treatment. Genome biology and evolution 2012, 4(12), 1310–1315. [Google Scholar] [CrossRef] [PubMed]
  136. Schuster, S. C. Next-generation sequencing transforms today’s biology. Nature methods 2008, 5(1), 16–18. [Google Scholar] [PubMed]
  137. Seppey, M.; Manni, M.; Zdobnov, E. M. BUSCO: Assessing Genome Assembly and Annotation Completeness. Methods in molecular biology 2019, 1962, 227–245. [Google Scholar] [CrossRef] [PubMed]
  138. Shafin, K.; Pesout, T.; Lorig-Roach, R.; Haukness, M.; Olsen, H. E.; Bosworth, C.; Armstrong, J.; Tigyi, K.; Maurer, N.; Koren, S.; Sedlazeck, F. J.; Marschall, T.; Mayes, S.; Costa, V.; Zook, J. M. Efficient de novo assembly of eleven human genomes using PromethION sequencing and a novel nanopore toolkit. BioRxiv 2019, 715722. [Google Scholar]
  139. Shafin, K.; Pesout, T.; Lorig-Roach, R.; Haukness, M.; Olsen, H.E.; Bosworth, C.; Armstrong, J.; Tigyi, K.; Maurer, N.; Koren, S.; Sedlazeck, F.J. Nanopore sequencing and the Shasta toolkit enable efficient de novo assembly of eleven human genomes. Nature biotechnology 2020, 38(9), 1044–1053. [Google Scholar] [CrossRef] [PubMed]
  140. Shen, C.; Du, H.; Chen, Z.; Lu, H.; Zhu, F.; Chen, H.; Meng, X.; Liu, Q.; Liu, P.; Zheng, L.; Li, X.; Dong, J.; Liang, C.; Wang, T. The chromosome-level genome sequence of the autotetraploid alfalfa and resequencing of core germplasms provide genomic resources for alfalfa research. Molecular Plant 2020, 13(9), 1250–1261. [Google Scholar] [CrossRef] [PubMed]
  141. Simao, F. A.; Waterhouse, R. M.; Ioannidis, P.; Kriventseva, E. V.; Zdobnov, E. M. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 2015, 31(19), 3210–3212. [Google Scholar] [CrossRef] [PubMed]
  142. Simpson, J. T.; Durbin, R. Efficient de novo assembly of large genomes using compressed data structures. Genome Research 2012, 22(3), 549–556. [Google Scholar] [PubMed]
  143. Simpson, J. T.; Wong, K.; Jackman, S. D.; Schein, J. E.; Jones, S. J.; Birol, I. ABySS: a parallel assembler for short read sequence data. Genome Research 2009, 19(6), 1117–1123. [Google Scholar] [CrossRef] [PubMed]
  144. Smeds, L.; Künstner, A. ConDeTri-a content dependent read trimmer for Illumina data. PloS one 2011, 6(10), e26314. [Google Scholar] [PubMed]
  145. Sohn, J. I.; Nam, J. W. The present and future of de novo whole-genome assembly. Briefings in bioinformatics 2018, 19(1), 23–40. [Google Scholar] [PubMed]
  146. Swathi, A.; Shekhar, M. S.; Katneni, V. K.; Vijayan, K. K. Genome size estimation of brackishwater fishes and penaeid shrimps by flow cytometry. Molecular biology reports 2018, 45, 951–960. [Google Scholar] [CrossRef] [PubMed]
  147. Tomato Genome Consortium, X. The tomato genome sequence provides insights into fleshy fruit evolution. Nature 2012, 485(7400), 635. [Google Scholar] [CrossRef]
  148. Toolkit for processing sequences in FASTA/Q formats. GitHub. Available online: https://github.com/lh3/seqtk (accessed on 5 January 2018).
  149. Vaser, R.; Sovic, I.; Nagarajan, N.; Sikic, M. Fast and accurate de novo genome assembly from long uncorrected reads. Genome research 2017, 27(5), 737–746. [Google Scholar] [CrossRef] [PubMed]
  150. Vega, L. Fundamentals of genetics; Waltham Abbey, Essex; Scientific e-Resources, 2019. [Google Scholar]
  151. Venter, J. C.; Adams, M. D.; Myers, E. W.; Li, P. W.; Mural, R. J.; Sutton, G. G.; Smith, H. O.; Yandell, M.; Evans, C. A.; Holt, R. A.; Gocayne, J. D.; Amanatides, P.; Ballew, R. M.; Huson, D. H.; Wortman, J. R. The sequence of the human genome. Science 2001, 291(5507), 1304–1351. [Google Scholar] [CrossRef] [PubMed]
  152. Walker, B. J.; Abeel, T.; Shea, T.; Priest, M.; Abouelliel, A.; Sakthikumar, S.; Cuomo, C. A.; Zeng, Q.; Wortman, J.; Young, S. K.; Earl, A. M. Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement. PloS one 2014, 9(11), e112963. [Google Scholar] [CrossRef] [PubMed]
  153. Wang, J. R.; Holt, J.; McMillan, L.; Jones, C. D. FMLRC: Hybrid long read error correction using an FM-index. BMC bioinformatics 2018, 19(1), 50. [Google Scholar] [CrossRef] [PubMed]
  154. Wang, J.; Veldsman, W. P.; Fang, X.; Huang, Y.; Xie, X.; Lyu, A.; Zhang, L. Benchmarking multi-platform sequencing technologies for human genome assembly. Briefings in bioinformatics 2023, 24(5), bbad300. [Google Scholar] [CrossRef] [PubMed]
  155. Wang, P.; Wang, F. A proposed metric set for evaluation of genome assembly quality. Trends in Genetics 2023, 39(3), 175–186. [Google Scholar] [PubMed]
  156. Wang, T.; Antonacci-Fulton, L.; Howe, K.; Lawson, H. A.; Lucas, J. K.; Phillippy, A. M.; Popejoy, A. B.; Asri, M.; Carson, C.; Chaisson, M. J. P.; Chang, X.; Cook-Deegan, R.; Felsenfeld, A. L.; Fulton, R. S.; Garrison, E. P. The Human Pangenome Project: a global resource to map genomic diversity. Nature 2022, 604(7906), 437–446. [Google Scholar] [CrossRef] [PubMed]
  157. Warren, R. L.; Coombe, L.; Mohamadi, H.; Zhang, J.; Jaquish, B.; Isabel, N.; Jones, S. J. M.; Bousquet, J.; Bohlmann, J.; Birol, I. ntEdit: scalable genome sequence polishing. Bioinformatics 2019, 35(21), 4430–4432. [Google Scholar] [CrossRef] [PubMed]
  158. Weirather, J. L.; de Cesare, M.; Wang, Y.; Piazza, P.; Sebastiano, V.; Wang, X. J.; Buck, D.; Au, K. F. Comprehensive comparison of Pacific Biosciences and Oxford Nanopore Technologies and their applications to transcriptome analysis; F1000Research 6, 2017. [Google Scholar]
  159. Wenger, A. M.; Peluso, P.; Rowell, W. J.; Chang, P. C.; Hall, R. J.; Concepcion, G. T.; Ebler, J.; Fungtammasan, A.; Kolesnikov, A.; Olson, N. D.; Topfer, A.; Alonge, M.; Mahmoud, M.; Qian, Y.; Chin, C. S. Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome. Nature biotechnology 2019, 37(10), 1155–1162. [Google Scholar] [CrossRef] [PubMed]
  160. Wick, R. R.; Holt, K. E. Polypolish: short-read polishing of long-read bacterial genome assemblies. PLoS computational biology 2022, 18(1), e1009802. [Google Scholar] [PubMed]
  161. Wick, R. R.; Judd, L. M.; Gorrie, C. L.; Holt, K. E. Unicycler: resolving bacterial genome assemblies from short and long sequencing reads. PLoS computational biology 2017, 13(6), e1005595. [Google Scholar] [CrossRef] [PubMed]
  162. Wong, J.; Coombe, L.; Nikolić, V.; Zhang, E.; Nip, K.M.; Sidhu, P.; Warren, R.L.; Birol, I. Linear time complexity de novo long read genome assembly with GoldRush. Nature Communications 2023, 14(1), 2906. [Google Scholar] [CrossRef] [PubMed]
  163. Xie, H.; Li, W.; Hu, Y.; Yang, C.; Lu, J.; Guo, Y.; Wen, L.; Tang, F. De novo assembly of human genome at single-cell levels. Nucleic Acids Research 2022, 50(13), 7479–7492. [Google Scholar] [PubMed]
  164. Xu, T.; Li, Y.; Zheng, W.; Sun, Y. A chromosome-level genome assembly of the blackspotted croaker (Protonibea diacanthus). Aquaculture and Fisheries 2022, 7(6), 616–622. [Google Scholar] [CrossRef]
  165. Yavas, G.; Hong, H.; Xiao, W. dnAQET: a framework to compute a consolidated metric for benchmarking quality of de novo assemblies. BMC genomics 2019, 20, 1–16. [Google Scholar] [CrossRef]
  166. Zerbino, D. R.; Birney, E. Velvet: algorithms for de novo short read assembly using de Bruijn graphs. Genome Research 2008, 18(5), 821–829. [Google Scholar] [CrossRef] [PubMed]
  167. Zhang, H.; Jain, C.; Aluru, S. A comprehensive evaluation of long read error correction methods. BMC genomics 2020, 21(6), 889. [Google Scholar] [CrossRef] [PubMed]
  168. Zhang, L.; Li, S.; Luo, J.; Du, P.; Wu, L.; Li, Y.; Zhu, X.; Wang, L.; Zhang, S.; Cui, J. Chromosome-level genome assembly of the predator Propylea japonica to understand its tolerance to insecticides and high temperatures. Molecular ecology resources 2020, 20(1), 292–307. [Google Scholar] [PubMed]
  169. Zhang, T.; Zhou, J.; Gao, W.; Jia, Y.; Wei, Y.; Wang, G. Complex genome assembly based on long-read sequencing. Briefings in bioinformatics 2022, 23(5), bbac305. [Google Scholar] [CrossRef] [PubMed]
  170. Zhang, W.; Chen, J.; Yang, Y.; Tang, Y.; Shang, J.; Shen, B. A practical comparison of de novo genome assembly software tools for next-generation sequencing technologies. PloS one 2011, 6(3), e17915. [Google Scholar] [CrossRef] [PubMed]
  171. Zhang, Z.; Yang, C.; Veldsman, W. P.; Fang, X.; Zhang, L. Benchmarking genome assembly methods on metagenomic sequencing data. Briefings in Bioinformatics 2023, 24(2), bbad087. [Google Scholar] [CrossRef] [PubMed]
  172. Zimin, A. V.; Salzberg, S. L. The genome polishing tool POLCA makes fast and accurate corrections in genome assemblies. PLoS computational biology 2020, 16(6), e1007981. [Google Scholar] [CrossRef] [PubMed]
  173. Zimin, A. V.; Puiu, D.; Luo, M. C.; Zhu, T.; Koren, S.; Marçais, G.; Yorke, J.; Salzberg, S. L. Hybrid assembly of the large and highly repetitive genome of Aegilops tauschii, a progenitor of bread wheat, with the MaSuRCA mega-reads algorithm. Genome research 2017, 27(5), 787–792. [Google Scholar] [PubMed]
Figure 1. Complete pipeline for genome assembly building and post analysis along with the tool names in red (https://www.biorender.com/).
Figure 1. Complete pipeline for genome assembly building and post analysis along with the tool names in red (https://www.biorender.com/).
Preprints 221362 g001
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings