Submitted:
02 July 2026
Posted:
03 July 2026
You are already at the latest version
Abstract
Keywords:
I. Introduction
II. Genome Complexity
III. NGS Raw Data from the Sequencing Platform
IV. Major Problems with the Sequencing Read from Different Platforms
V. NGS Data Analysis
- (1)
- Pre-Processing of Raw reads
- (2)
- Generation of genome assembly from NGS data
- (3)
- Evaluation of genome assembly quality
- (4)
- Tools for Improving the quality of genome assembly
VI. Conclusions
- Sequencing technologies have evolved significantly staring from Sanger’s chain-termination method to the latest TGS platforms, improving read lengths and genome finishing. Despite their benefits, new technologies still face challenges related to error rates, sequencing biases, yield, cost-effectiveness, and repetitive regions. Long-read sequencing technologies hold the potential to further genomic discoveries by providing deeper insights into genome structure, variation, and function. As these technologies advance, they will enhance our understanding of genomics, leading to more precise and comprehensive analyses across various species.
- Genome assembly from NGS data involves complex processes and algorithms tailored to specific sequencing technologies and genome characteristics. The choice of assembly methods and tools is crucial for achieving accuracy, efficiency, and successful genome assembly outcomes.
- A comprehensive approach, using a range of tools and metrics, is necessary to assess contiguity, accuracy, completeness, and contamination in genome assemblies. Existing tools like QUAST and BUSCO provide valuable insights, but a standardized evaluation framework is needed to ensure the reproducibility and reliability of genome assemblies.
- The combination of long-read technologies and error correction methods continues to improve assembly quality, though establishing universal standards for genome assembly quality remains a challenge.
References
- Aird, D.; Ross, M. G.; Chen, W. S.; Danielsson, M.; Fennell, T.; Russ, C.; Jaffe, D.B.; Nusbaum, C.; Gnirke, A. Analyzing and minimizing PCR amplification bias in Illumina sequencing libraries. Genome biology 2011, 12, 1–14. [Google Scholar] [CrossRef]
- Alhakami, H.; Mirebrahim, H.; Lonardi, S. A comparative evaluation of genome assembly reconciliation tools. Genome biology 2017, 18, 1–14. [Google Scholar] [CrossRef]
- Alkan, C.; Sajjadian, S.; Eichler, E. E. Limitations of next-generation genome sequence assembly. Nature methods 2011, 8(1), 61–65. [Google Scholar] [PubMed]
- Allhoff, M.; Schonhuth, A.; Martin, M.; Costa, I. G.; Rahmann, S.; Marschall, T. Discovering motifs that induce sequencing errors. BMC bioinformatics 2013, 14, 1–10. [Google Scholar] [CrossRef]
- Anderson, S. Shotgun DNA sequencing using cloned DNase I-generated fragments. Nucleic acids research 1981, 9(13), 3015–3027. [Google Scholar] [CrossRef] [PubMed]
- Andrews, S. (2010). FastQC: a quality control tool for high throughput sequence data http://www. bioinformatics. babraham. ac. uk/projects/fastqc. Babraham Bioinformatics.
- Andrews, S. (2014). FastQC a quality-control tool for high-throughput sequence data http://www. Bioinformaticsbabraham. ac. uk/projects/fastqc.
- Baker, M. De novo genome assembly: what every biologist should know. Nature methods 2012, 9(4), 333–337. [Google Scholar] [CrossRef]
- Berthelot, C.; Brunet, F.; Chalopin, D.; Juanchich, A.; Bernard, M.; Noel, B.; Bento, P.; Da Silva, C.; Labadie, K.; Alberti, A.; Aury, J.M.; Louis, A.; Dehais, P.; Bardou, p.; Montfort, J.; Klopp, C. The rainbow trout genome provides novel insights into evolution after whole-genome duplication in vertebrates. Nature communications 2014, 5(1), 1–10. [Google Scholar] [CrossRef]
- Betschart, R. O.; Thiery, A.; Aguilera-Garcia, D.; Zoche, M.; Moch, H.; Twerenbold, R.; Zeller, T.; Blankenberg, S.; Ziegler, A. Comparison of calling pipelines for whole genome sequencing: an empirical study demonstrating the importance of mapping and alignment. Scientific Reports 2022, 12(1), 21502. [Google Scholar] [CrossRef] [PubMed]
- Bolger, A. M.; Lohse, M.; Usadel, B. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics 2014, 30(15), 2114–2120. [Google Scholar] [CrossRef] [PubMed]
- Bradnam, K. R.; Fass, J. N.; Alexandrov, A.; Baranay, P.; Bechner, M.; Birol, I.; Boisvert, S.; Chapman, J.A.; Chapuis, G.; Chikhi, R.; Chitsaz, H.; Chou, W. C.; Corbeil, J.; Fabbro, C. D.; Docking, T. R. Assemblathon 2: evaluating de novo methods of genome assembly in three vertebrate species. Gigascience 2013, 2(1), 2047–217X. [Google Scholar] [CrossRef]
- Cahill, M. J.; Koser, C. U.; Ross, N. E.; Archer, J. A. Read length and repeat resolution: exploring prokaryote genomes using next-generation sequencing technologies. PloS one 2010, 5(7), e11518. [Google Scholar] [CrossRef] [PubMed]
- Chen, J.; Li, X.; Zhong, H.; Meng, Y.; Du, H. Systematic comparison of germline variant calling pipelines cross multiple next-generation sequencers. Scientific reports 2019, 9(1), 9345. [Google Scholar] [CrossRef] [PubMed]
- Chen, S.; Zhou, Y.; Chen, Y.; Gu, J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 2018, 34(17), i884–i890. [Google Scholar] [CrossRef] [PubMed]
- Chen, Y. C.; Liu, T.; Yu, C. H.; Chiang, T. Y.; Hwang, C. C. Effects of GC bias in next-generation-sequencing data on de novo genome assembly. PloS one 2013, 8(4), e62856. [Google Scholar] [PubMed]
- Chen, Y.; Chen, Y.; Shi, C.; Huang, Z.; Zhang, Y.; Li, S.; Li, Y.; Ye, J.; Yu, C.; Li, Z.; Zhang, X.; Wang, J.; Yang, H.; Fang, L.; Chen, Q. SOAPnuke: a MapReduce acceleration-supported software for integrated quality control and preprocessing of high-throughput sequencing data. Gigascience 2018, 7(1), gix120. [Google Scholar] [PubMed]
- Chen, Y.; Zhang, Y.; Wang, A. Y.; Gao, M.; Chong, Z. Accurate long-read de novo assembly evaluation with Inspector. Genome Biology 2021, 22, 1–21. [Google Scholar] [CrossRef]
- Cheng, H.; Concepcion, G. T.; Feng, X.; Zhang, H.; Li, H. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nature methods 2021, 18(2), 170–175. [Google Scholar] [CrossRef] [PubMed]
- Chin, C. S.; Khalak, A. Human genome assembly in 100 minutes. BioRxiv 705616; 2019.
- Chin, C. S.; Alexander, D. H.; Marks, P.; Klammer, A. A.; Drake, J.; Heiner, C.; Clum, A.; Copeland, A.; Huddleston, J.; Eichler, E. E.; Turner, S. W.; Korlach, J. Nonhybrid, finished microbial genome assemblies from long-read SMRT sequencing data. Nature methods 2013, 10(6), 563–569. [Google Scholar] [CrossRef] [PubMed]
- Cho, Y. S.; Kim, H.; Kim, H. M.; Jho, S.; Jun, J.; Lee, Y. J.; Chae, K. S.; Kim, C. G.; Kim, S.; Eriksson, A.; Edwards, J. S.; Lee, S.
- Choi, I. Y.; Kwon, E. C.; Kim, N. S. The C-and G-value paradox with polyploidy, repeatomes, introns, phenomes and cell economy. Genes & genomics 2020, 42, 699–714. [Google Scholar]
- Compeau, P. E.; Pevzner, P. A.; Tesler, G. How to apply de Bruijn graphs to genome assembly. Nature biotechnology 2011, 29(11), 987–991. [Google Scholar] [CrossRef] [PubMed]
- Coombe, L.; Warren, R. L.; Wong, J.; Nikolic, V.; Birol, I. ntLink: a toolkit for de novo genome assembly scaffolding and mapping using long reads. Current Protocols 2023, 3(4), e733. [Google Scholar] [CrossRef] [PubMed]
- Corradi, N.; Pombert, J. F.; Farinelli, L.; Didier, E. S.; Keeling, P. J. The complete sequence of the smallest known nuclear genome from the microsporidian Encephalitozoon intestinalis. Nature communications 2010, 1(1), 77. [Google Scholar] [CrossRef] [PubMed]
- Cox, M. P.; Peterson, D. A.; Biggs, P. J. SolexaQA: At-a-glance quality assessment of Illumina second-generation sequencing data. BMC bioinformatics 2010, 11, 1–6. [Google Scholar]
- Del Angel, V. D.; Hjerde, E.; Sterck, L.; Capella-Gutierrez, S.; Notredame, C.; Pettersson, O. V.; Amselem, J.; Bouri, L.; Bocs, S.; Klopp, C.; Gibrat, J. F.; Vlasova, A.; Leskosek, B. L.; Soler, L.; Binzer-Panchal, M.; Lantz, H. Ten steps to get started in Genome Assembly and Annotation; F1000Research 7, 2018. [Google Scholar]
- Dida, F.; Yi, G. Empirical evaluation of methods for de novo genome assembly. PeerJ Computer Science 2021, 7, e636. [Google Scholar] [PubMed]
- Drmanac, R.; Sparks, A. B.; Callow, M. J.; Halpern, A. L.; Burns, N. L.; Kermani, B. G.; Carnevali, P.; Nazarenko, I.; Nilsen, G. B.; Yeung, G.; Dahl, F.; Fernandez, A.; Staker, B.; Pant, K. P.; Baccash, J. Human genome sequencing using unchained base reads on self-assembling DNA nanoarrays. Science 2010, 327(5961), 78–81. [Google Scholar] [CrossRef] [PubMed]
- Earl, D.; Bradnam, K.; St John, J.; Darling, A.; Lin, D.; Fass, J.; Yu, H. O.; Buffalo, V.; Zerbino, D. R.; Diekhans, M.; Nguyen, N.; Ariyaratne, P. N.; Sung, W. K.; Ning, Z.; Haimel, M. Assemblathon 1: a competitive assessment of de novo short read assembly methods. Genome Research 2011, 21(12), 2224–2241. [Google Scholar] [CrossRef] [PubMed]
- Elliott, T. A.; Gregory, T. R. What’s in a genome? The C-value enigma and the evolution of eukaryotic genome content. Philosophical Transactions of the Royal Society B: Biological Sciences 2015, 370(1678), 20140331. [Google Scholar]
- El-Metwally, S.; Hamouda, E.; Tarek, M. A roadmap to sequence assembly evaluation tools. Current Bioinformatics 2021, 16(5), 644–661. [Google Scholar] [CrossRef]
- Fabbro, C.; Scalabrin, S.; Morgante, M.; Giorgi, F. M. An extensive evaluation of read trimming effects on Illumina NGS data analysis. PloS one 2013, 8(12), e85024. [Google Scholar] [CrossRef] [PubMed]
- Ferreira, C. R. The burden of rare diseases. American journal of medical genetics Part A 2019, 179(6), 885–892. [Google Scholar] [CrossRef] [PubMed]
- Firtina, C.; Kim, J. S.; Alser, M.; Senol Cali, D.; Cicek, A. E.; Alkan, C.; Mutlu, O. Apollo: a sequencing-technology-independent, scalable and accurate assembly polishing algorithm. Bioinformatics 2020, 36(12), 3669–3679. [Google Scholar] [PubMed]
- Flicek, P.; Birney, E. Sense from sequence reads: methods for alignment and assembly. Nature methods 2009, 6 (Suppl 11), S6–S12. [Google Scholar] [CrossRef] [PubMed]
- Formenti, G.; Theissinger, K.; Fernandes, C.; Bista, I.; Bombarely, A.; Bleidorn, C.; Ciofi, C.; Crottini, A.; Godoy, J. A.; Hoglund, J.; Malukiewicz, J.; Mouton, A.; Oomen, R. A.; Paez, S.; Palsboll, P. J. The era of reference genomes in conservation genomics. Trends in ecology & evolution 2022, 37(3), 197–202. [Google Scholar] [CrossRef]
- Gaia, A. S. C.; de Oliveira, M. S.; da Silva Moia, G.; dos Santos, V. C.; Alves, J. T. C.; de Sá, P. H. C. G.; de Oliveira Veras, A. A. e-QA NGS: a user-friendly tool to preprocessing data from next generation sequencing. Peer Review 2023, 5(3), 91–105. [Google Scholar]
- Gao, S. H.; Yu, H. Y.; Wu, S. Y.; Wang, S.; Geng, J. N.; Luo, Y. F.; Hu, S. N. Advances of sequencing and assembling technologies for complex genomes. Yi Chuan Hereditas 2018, 40(11), 944–963. [Google Scholar] [PubMed]
- Gavrielatos, M.; Kyriakidis, K.; Spandidos, D. A.; Michalopoulos, I. Benchmarking of next and third generation sequencing technologies and their associated algorithms for de novo genome assembly. Molecular Medicine Reports 2021, 23(4), 1–1. [Google Scholar] [CrossRef]
- Genome Reference Consortium. (2022-05-09). “GenomeRef: GRCh38.p14 is now released!”. GRC Blog (GenomeRef). Retrieved 2022-08-19.
- Genova, A. D.; Buena-Atienza, E.; Ossowski, S.; Sagot, M. F. Efficient hybrid de novo assembly of human genomes with WENGAN. Nature Biotechnology 2021, 39(4), 422–430. [Google Scholar] [PubMed]
- Giordano, F.; Aigrain, L.; Quail, M. A.; Coupland, P.; Bonfield, J. K.; Davies, R. M.; Tischler, G.; Jackson, D. K.; Keane, T. M.; Li, J.; Yue, J. X. De novo yeast genome assemblies from MinION, PacBio and MiSeq platforms. Scientific reports 2017, 7(1), 3935. [Google Scholar] [CrossRef] [PubMed]
- Giorgi, F. M.; Del-Fabbro, C.; Licausi, F. Comparative study of RNA-seq-and microarray-derived coexpression networks in Arabidopsis thaliana. Bioinformatics 2013, 29(6), 717–724. [Google Scholar] [CrossRef] [PubMed]
- Glusman, G.; Cox, H. C.; Roach, J. C. Whole-genome haplotyping approaches and genomic medicine. Genome medicine 2014, 6, 1–16. [Google Scholar] [CrossRef]
- Gnerre, S.; Maccallum, I.; Przybylski, D.; Ribeiro, F. J.; Burton, J. N.; Walker, B. J.; Sharpe, T.; Hall, G.; Shea, T. P.; Sykes, S.; Berlin, A. M.; Aird, D.; Costello, M.; Daza, R.; Williams, L. High-quality draft assemblies of mammalian genomes from massively parallel sequence data. Proceedings of the National Academy of Sciences of the United States of America 2011, 108(4), 1513–1518. [Google Scholar] [PubMed]
- Gonzalez-Garcia, L.; Guevara-Barrientos, D.; Lozano-Arce, D.; Gil, J.; Díaz-Riaño, J.; Duarte, E.; Andrade, G.; Bojaca, J. C.; Hoyos-Sanchez, M. C.; Chavarro, C.; Guayazan, N.; Chica, L. A.; Acosta, M. C. B.; Bautista, E.; Trujillo, M.; Duitama, J. New algorithms for accurate and efficient de novo genome assembly from long DNA sequencing reads. Life Science Alliance 2023, 6(5). [Google Scholar] [CrossRef] [PubMed]
- Goodwin, S.; McPherson, J. D.; McCombie, W. R. Coming of age: ten years of next-generation sequencing technologies. Nature reviews genetics 2016, 17(6), 333–351. [Google Scholar] [CrossRef] [PubMed]
- Guo, A.; Salzberg, S. L.; Zimin, A. V. JASPER: A fast genome polishing tool that improves accuracy of genome assemblies. PLoS computational biology 2023, 19(3), e1011032. [Google Scholar] [CrossRef] [PubMed]
- Gurevich, A.; Saveliev, V.; Vyahhi, N.; Tesler, G. QUAST: quality assessment tool for genome assemblies. Bioinformatics 2013, 29(8), 1072–1075. [Google Scholar] [CrossRef] [PubMed]
- Haghshenas, E.; Asghari, H.; Stoye, J.; Chauve, C.; Hach, F. HASLR: fast hybrid assembly of long reads. Iscience 2020, 23(8). [Google Scholar] [CrossRef] [PubMed]
- Han, K.; Li, Z. F.; Peng, R.; Zhu, L. P.; Zhou, T.; Wang, L. G.; Li, S. G.; Zhang, X. B.; Hu, W.; Wu, Z. H.; Qin, N.; Li, Y. Z. Extraordinary expansion of a Sorangium cellulosum genome from an alkaline milieu. Scientific reports 2013, 3(1), 2101. [Google Scholar] [CrossRef] [PubMed]
- He, Y.; Zhang, Z.; Peng, X.; Wu, F.; Wang, J. De novo assembly methods for next generation sequencing data. Tsinghua Science and Technology 2013, 18(5), 500–514. [Google Scholar] [CrossRef]
- Heydari, M.; Miclotte, G.; Demeester, P.; Van de Peer, Y.; Fostier, J. Evaluation of the impact of Illumina error correction tools on de novo genome assembly. BMC bioinformatics 2017, 18, 1–13. [Google Scholar] [CrossRef]
- Horner, D. S.; Pavesi, G.; Castrignano, T.; De Meo, P. D. O.; Liuni, S.; Sammeth, M.; Picardi, E.; Pesole, G. Bioinformatics approaches for genomics and post genomics applications of next-generation sequencing. Briefings in bioinformatics 2010, 11(2), 181–197. [Google Scholar] [PubMed]
- Hotaling, S.; Kelley, J. L.; Frandsen, P. B. Toward a genome sequence for every animal: Where are we now? Proceedings of the National Academy of Sciences 2021, 118(52), e2109019118. [Google Scholar] [CrossRef]
- Hu, J.; Fan, J.; Sun, Z.; Liu, S. NextPolish: a fast and efficient genome polishing tool for long-read assembly. Bioinformatics 2020, 36(7), 2253–2255. [Google Scholar] [PubMed]
- Huang, J.; Liang, X.; Xuan, Y.; Geng, C.; Li, Y.; Lu, H.; Qu, S.; Mei, X.; Chen, H.; Yu, T.; Sun, N.; Rao, J.; Wang, J.; Zhang, W.; Chen, Y. A reference human genome dataset of the BGISEQ-500 sequencer. GigaScience 2017, 6(5), 1–9. [Google Scholar] [CrossRef] [PubMed]
- Huang, N.; Nie, F.; Ni, P.; Luo, F.; Gao, X.; Wang, J. NeuralPolish: a novel Nanopore polishing method based on alignment matrix construction and orthogonal Bi-GRU Networks. Bioinformatics 2021, 37(19), 3120–3127. [Google Scholar] [PubMed]
- Huang, Y.; Wang, Z.; Schmidt, M. A.; Su, H.; Xiong, L.; Zhang, J. DEGAP: Dynamic elongation of a genome assembly path. Briefings in Bioinformatics 2024, 25(3), bbae194. [Google Scholar] [CrossRef] [PubMed]
- Huddleston, J.; Ranade, S.; Malig, M.; Antonacci, F.; Chaisson, M.; Hon, L.; Sudmant, P. H.; Graves, T. A.; Alkan, C.; Dennis, M. Y.; Wilson, R. K.; Turner, S. W.; Korlach, J.; Eichler, E. E. Reconstructing complex regions of genomes using long-read sequencing technology. Genome Research 2014, 24(4), 688–696. [Google Scholar] [CrossRef] [PubMed]
- Hunt, M.; Kikuchi, T.; Sanders, M.; Newbold, C.; Berriman, M.; Otto, T. D. REAPR: a universal tool for genome assembly evaluation. Genome biology 2013, 14, 1–10. [Google Scholar] [CrossRef]
- Jackman, S. D.; Birol, İ. Assembling genomes using short-read sequencing technology. Genome Biology 2010, 11, 1–4. [Google Scholar] [CrossRef]
- Jain, M.; Fiddes, I. T.; Miga, K. H.; Olsen, H. E.; Paten, B.; Akeson, M. Improved data analysis for the MinION nanopore sequencer. Nature methods 2015, 12(4), 351–356. [Google Scholar] [CrossRef] [PubMed]
- Jain, M.; Koren, S.; Miga, K. H.; Quick, J.; Rand, A. C.; Sasani, T. A.; Tyson, J. R.; Beggs, A. D.; Dilthey, A. T.; Fiddes, I. T.; Malla, S.; Marriott, H.; Nieto, T.; O’Grady, J.; Olsen, H. E. Nanopore sequencing and assembly of a human genome with ultra-long reads. Nature biotechnology 2018, 36(4), 338–345. [Google Scholar] [CrossRef] [PubMed]
- Jarvis, E. D.; Formenti, G.; Rhie, A.; Guarracino, A.; Yang, C.; Wood, J.; Tracey, A.; Thibaud-Nissen, F.; Vollger, M. R.; Porubsky, D.; Cheng, H.; Asri, M.; Logsdon, G. A.; Carnevali, P.; Chaisson, M. J. P. Semi-automated assembly of high-quality diploid human reference genomes. Nature 2022, 611(7936), 519–531. [Google Scholar] [PubMed]
- Jung, H.; Ventura, T.; Chung, J. S.; Kim, W. J.; Nam, B. H.; Kong, H. J.; Kim, Y. O.; Jeon, M. S.; Eyun, S. I. Twelve quick steps for genome assembly and annotation in the classroom. PLoS computational biology 2020, 16(11), e1008325. [Google Scholar] [CrossRef] [PubMed]
- Jung, H.; Winefield, C.; Bombarely, A.; Prentis, P.; Waterhouse, P. Tools and strategies for long-read sequencing and de novo assembly of plant genomes. Trends in plant science 2019, 24(8), 700–724. [Google Scholar] [CrossRef] [PubMed]
- Karlsson, F. H.; Fåk, F.; Nookaew, I.; Tremaroli, V.; Fagerberg, B.; Petranovic, D.; Backhed, F.; Nielsen, J. Symptomatic atherosclerosis is associated with an altered gut metagenome. Nature communications 2012, 3(1), 1245. [Google Scholar] [CrossRef] [PubMed]
- Kelley, D. R.; Schatz, M. C.; Salzberg, S. L. Quake: quality-aware detection and correction of sequencing errors. Genome biology 2010, 11, 1–13. [Google Scholar] [CrossRef]
- Khan, A. R.; Pervez, M. T.; Babar, M. E.; Naveed, N.; Shoaib, M. A comprehensive study of de novo genome assemblers: current challenges and future prospective. Evolutionary Bioinformatics 2018, 14, 1176934318758650. [Google Scholar] [CrossRef] [PubMed]
- Khelik, K.; Sandve, G. K.; Nederbragt, A. J.; Rognes, T. NucBreak: location of structural errors in a genome assembly by using paired-end Illumina reads. BMC bioinformatics 2020, 21, 1–11. [Google Scholar]
- Khiste, N.; Ilie, L. Laser: Large genome assembly evaluator. BMC research notes 2015, 8, 1–3. [Google Scholar] [CrossRef]
- Kim, B. C.; Manica, A.; Oh, T. K. An ethnically relevant consensus Korean reference genome is a step towards personal reference genomes. Nature communications 2016, 7(1), 1–13. [Google Scholar] [CrossRef]
- Kim, H.M.; Jeon, S.; Chung, O.; Jun, J.H.; Kim, H.S.; Blazyte, A.; Lee, H.Y.; Yu, Y.; Cho, Y.S.; Bolser, D.M.; Bhak, J. Comparative analysis of 7 short-read sequencing platforms using the Korean Reference Genome: MGI and Illumina sequencing benchmark for whole-genome sequencing. GigaScience 2021, 10(3), 1–9. [Google Scholar]
- Kundu, R.; Casey, J.; Sung, W. K. HyPo: super fast & accurate polisher for long read genome assemblies. Biorxiv 2019, 20, 2019–12. [Google Scholar]
- Lakhotia, S. C. C-value paradox: Genesis in misconception that natural selection follows anthropocentric parameters of ‘economy’and ‘optimum’. BBA advances 2023, 4, 100107. [Google Scholar] [PubMed]
- Lander, E.S.; Linton, L.M.; Birren, B. Initial sequencing and analysis of the human genome. Nature 2001, 409, 860–921. [Google Scholar] [CrossRef] [PubMed]
- Li, H.; Durbin, R. Fast and accurate long-read alignment with Burrows–Wheeler transform. Bioinformatics 2010, 26(5), 589–595. [Google Scholar] [PubMed]
- Li, R.; Di, L.; Li, J.; Fan, W.; Liu, Y.; Guo, W.; Liu, W.; Liu, L.; Li, Q.; Chen, L.; Chen, Y. A body map of somatic mutagenesis in morphologically normal human tissues. Nature 2021, 597(7876), 398–403. [Google Scholar] [CrossRef] [PubMed]
- Li, R.; Zhu, H.; Ruan, J.; Qian, W.; Fang, X.; Shi, Z.; Li, Y.; Li, S.; Shan, G.; Kristiansen, K.; Li, S.; Yang, H.; Wang, J.; Wang, J. De novo assembly of human genomes with massively parallel short read sequencing. Genome Research 2010, 20(2), 265–272. [Google Scholar] [PubMed]
- Li, Z.; Chen, Y.; Mu, D.; Yuan, J.; Shi, Y.; Zhang, H.; Gan, J.; Li, N.; Hu, X.; Liu, B.; Yang, B.; Fan, W. Comparison of the two major classes of assembly algorithms: overlap-layout-consensus and de-bruijn-graph. Briefings in functional genomics 2012, 11(1), 25–37. [Google Scholar] [PubMed]
- Liao, X.; Li, M.; Zou, Y.; Wu, F. X.; Wang, J. Current challenges and solutions of de novo assembly. Quantitative Biology 2019, 7(2), 90–109. [Google Scholar] [CrossRef]
- Liao, X.; Li, M.; Zou, Y.; Wu, F.; Pan, Y.; Luo, F.; Wang, J. Improving de novo Assembly Based on Read Classification. IEEE/ACM Transactions on Computational Biology and Bioinformatics 2020, 17(1), 177–188. [Google Scholar] [PubMed]
- Linsmith, G.; Rombauts, S.; Montanari, S.; Deng, C. H.; Celton, J. M.; Guérif, P.; Liu, C.; Lohaus, R.; Zurn, J. D.; Cestaro, A.; Bassil, N. V.; Bakker, L. V.; Schijlen, E.; Gardiner, S. E.; Lespinasse, Y. Pseudo-chromosome–length genome assembly of a double haploid “Bartlett” pear (Pyrus communis L.). Gigascience 2019, 8(12), giz138. [Google Scholar] [CrossRef] [PubMed]
- Liu, H.; Wang, X.; Wang, G.; Cui, P.; Wu, S.; Ai, C.; Hu, N.; Li, A.; He, B.; Shao, X.; Wu, Z. The nearly complete genome of Ginkgo biloba illuminates gymnosperm evolution. Nature Plants 2021, 7(6), 748–756. [Google Scholar] [CrossRef] [PubMed]
- Liu, J.; Seetharam, A.S.; Chougule, K.; Ou, S.; Swentowsky, K.W.; Gent, J.I.; Llaca, V.; Woodhouse, M.R.; Manchanda, N.; Presting, G.G.; Kudrna, D.A. Gapless assembly of maize chromosomes using long-read technologies. Genome Biology 2020, 21, 1–17. [Google Scholar] [CrossRef]
- Liu, L.; Li, Y.; Li, S.; Hu, N.; He, Y.; Pong, R.; Lin, D.; Lu, L.; Law, M. Comparison of next-generation sequencing systems. BioMed research international 2012, 2012(1), 251364. [Google Scholar] [CrossRef]
- Liu, S.; Li, W.; Wu, Y.; Chen, C.; Lei, J. De novo transcriptome assembly in chili pepper (Capsicum frutescens) to identify genes involved in the biosynthesis of capsaicinoids. PloS one 2013, 8(1), e48156. [Google Scholar] [CrossRef] [PubMed]
- Logsdon, G. A.; Vollger, M. R.; Eichler, E. E. Long-read human genome sequencing and its applications. Nature Reviews Genetics 2020, 21(10), 597–614. [Google Scholar] [CrossRef] [PubMed]
- Loman, N. J.; Quick, J.; Simpson, J. T. A complete bacterial genome assembled de novo using only nanopore sequencing data. Nature methods 2015, 12(8), 733–735. [Google Scholar] [CrossRef] [PubMed]
- Luan, T.; Commichaux, S.; Hoffmann, M.; Jayeola, V.; Jang, J. H.; Pop, M.; Rand, H.; Luo, Y. Benchmarking short and long read polishing tools for nanopore assemblies: achieving near-perfect genomes for outbreak isolates. BMC genomics 2024, 25(1), 679. [Google Scholar] [CrossRef] [PubMed]
- Lucas, S. J.; Akpınar, B. A.; Simkova, H.; Kubaláková, M.; Dolezel, J.; Budak, H. Next-generation sequencing of flow-sorted wheat chromosome 5D reveals lineage-specific translocations and widespread gene duplications. BMC genomics 2014, 15, 1–18. [Google Scholar]
- Luo, R.; Liu, B.; Xie, Y.; Li, Z.; Huang, W.; Yuan, J.; He, G.; Chen, Y.; Pan, Q.; Liu, Y.; Tang, J.; Wu, G.; Zhang, H.; Shi, Y.; Liu, Y. SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler. Gigascience 2012, 1(1), 2047–217X. [Google Scholar]
- Ma, Z.; Zhang, Y.; Wu, L.; Zhang, G.; Sun, Z.; Li, Z.; Jiang, Y.; Ke, H.; Chen, B.; Liu, Z.; Gu, Q. High-quality genome assembly and resequencing of modern cotton cultivars provide resources for crop improvement. Nature Genetics 2021b, 53(9), 1385–1391. [Google Scholar]
- Mak, Q. C.; Wick, R. R.; Holt, J. M.; Wang, J. R. Polishing de novo nanopore assemblies of bacteria and eukaryotes with FMLRC2. Molecular Biology and Evolution 2023, 40(3), msad048. [Google Scholar] [CrossRef] [PubMed]
- Manchanda, N.; Portwood, J. L.; Woodhouse, M. R.; Seetharam, A. S.; Lawrence-Dill, C. J.; Andorf, C. M.; Hufford, M. B. GenomeQC: a quality assessment tool for genome assemblies and gene structure annotations. BMC genomics 2020, 21, 1–9. [Google Scholar] [CrossRef]
- Mantere, T.; Kersten, S.; Hoischen, A. Long-read sequencing emerging in medical genetics. Frontiers in genetics 2019, 10, 432668. [Google Scholar] [CrossRef] [PubMed]
- Mapleson, D.; Garcia Accinelli, G.; Kettleborough, G.; Wright, J.; Clavijo, B. J. KAT: a K-mer analysis toolkit to quality control NGS datasets and genome assemblies. Bioinformatics 2017, 33(4), 574–576. [Google Scholar] [PubMed]
- Margulies, M.; Egholm, M.; Altman, W. E.; Attiya, S.; Bader, J. S.; Bemben, L. A.; Berka, J.; Braverman, M. S.; Chen, Y. J.; Chen, Z.; Dewell, S. B.; Du, L.; Fierro, J. M.; Gomes, X. V.; Godwin, B. C. Genome sequencing in microfabricated high-density picolitre reactors. Nature 2005, 437(7057), 376–380. [Google Scholar] [CrossRef] [PubMed]
- Marks, P.; Garcia, S.; Barrio, A. M.; Belhocine, K.; Bernate, J.; Bharadwaj, R.; Bjornson, K.; Catalanotti, C.; Delaney, J.; Fehr, A.; Fiddes, I. T.; Galvin, B.; Heaton, H.; Herschleb, J.; Hindson, C. Resolving the full spectrum of human genome variation using Linked-Reads. Genome Research 2019, 29(4), 635–645. [Google Scholar] [CrossRef] [PubMed]
- Martin, M. Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet. journal 2011, 17(1), 10–12. [Google Scholar] [CrossRef]
- McCormick, K. A.; Calzone, K. A. The impact of genomics on health outcomes, quality, and safety. Nursing management 2016, 47(4), 23–26. [Google Scholar] [CrossRef] [PubMed]
- McCoy, R. C.; Taylor, R. W.; Blauwkamp, T. A.; Kelley, J. L.; Kertesz, M.; Pushkarev, D.; Petrov, D. A.; Fiston-Lavier, A. S. Illumina TruSeq synthetic long-reads empower de novo assembly and resolve complex, highly-repetitive transposable elements. PloS one 2014, 9(9), e106689. [Google Scholar] [PubMed]
- Medaka, O. N. T. (2018). sequence correction provided by ONT Research. Github. Available from: https://github. com/nanoporetech/medaka.
- Mikheenko, A.; Prjibelski, A.; Saveliev, V.; Antipov, D.; Gurevich, A. Versatile genome assembly evaluation with QUAST-LG. Bioinformatics 2018, 34(13), i142–i150. [Google Scholar] [CrossRef] [PubMed]
- Mikheenko, A.; Saveliev, V.; Hirsch, P.; Gurevich, A. WebQUAST: online evaluation of genome assemblies. Nucleic Acids Research 2023, 51(W1), 601–606. [Google Scholar] [CrossRef]
- Mikheenko, A.; Valin, G.; Prjibelski, A.; Saveliev, V.; Gurevich, A. Icarus: visualizer for de novo assembly evaluation. Bioinformatics 2016, 32(21), 3321–3323. [Google Scholar] [CrossRef] [PubMed]
- Miller, J. R.; Koren, S.; Sutton, G. Assembly algorithms for next-generation sequencing data. Genomics 2010, 95(6), 315–327. [Google Scholar] [CrossRef] [PubMed]
- Nakamura, K.; Oshima, T.; Morimoto, T.; Ikeda, S.; Yoshikawa, H.; Shiwa, Y.; Ishikawa, S.; Linak, M. C.; Hirai, A.; Takahashi, H.; Altaf-Ul-Amin, Md.; Ogasawara, N.; Kanaya, S. Sequence-specific error profile of Illumina sequencers. Nucleic acids research 2011, 39(13), e90–e90. [Google Scholar] [CrossRef] [PubMed]
- Narzisi, G.; Mishra, B. Comparing de novo genome assembly: the long and short of it. PloS one 2011, 6(4), e19175. [Google Scholar] [CrossRef] [PubMed]
- Neale, D. B.; Wegrzyn, J. L.; Stevens, K. A.; Zimin, A. V.; Puiu, D.; Crepeau, M. W.; Cardeno, C.; Koriabine, M.; Holtz-Morris, A. E.; Liechty, J. D.; Martínez-García, P. J.; Vasquez-Gross, H. A.; Lin, B. Y.; Zieve, J. J.; Dougherty, W. M. Decoding the massive genome of loblolly pine using haploid DNA and novel assembly strategies. Genome biology 2014, 15, 1–13. [Google Scholar] [CrossRef]
- Niederst, M. J.; Hu, H.; Mulvey, H. E.; Lockerman, E. L.; Garcia, A. R.; Piotrowska, Z.; Sequist, L. V.; Engelman, J. A. The allelic context of the C797S mutation acquired upon treatment with third-generation EGFR inhibitors impacts sensitivity to subsequent treatment strategies. Clinical Cancer Research 2015, 21(17), 3924–3933. [Google Scholar] [PubMed]
- Nurk, S.; Koren, S.; Rhie, A.; Rautiainen, M.; Bzikadze, A. V.; Mikheenko, A.; Vollger, M. R.; Altemose, N.; Uralsky, L.; Gershman, A.; Aganezov, S.; Hoyt, S. J.; Diekhans, M.; Logsdon, G. A.; Alonge, M. The complete sequence of a human genome. Science 2022, 376(6588), 44–53. [Google Scholar] [CrossRef] [PubMed]
- Oppenheimer, J.; Rosen, B.D.; Heaton, M.P.; Vander Ley, B.L.; Shafer, W.R.; Schuetze, F.T.; Stroud, B.; Kuehn, L.A.; McClure, J.C.; Barfield, J.P.; Blackburn, H.D. A reference genome assembly of American Bison, Bison bison bison. Journal of Heredity 2021, 112(2), 174–183. [Google Scholar] [CrossRef] [PubMed]
- Ou, S.; Chen, J.; Jiang, N. Assessing genome assembly quality using the LTR Assembly Index (LAI). Nucleic acids research 2018, 46(21), e126–e126. [Google Scholar] [CrossRef] [PubMed]
- Padovani de Souza, K.; Setubal, J. C.; Ponce de Leon F de Carvalho, A. C.; Oliveira, G.; Chateau, A.; Alves, R. Machine learning meets genome assembly. Briefings in bioinformatics 2019, 20(6), 2116–2129. [Google Scholar] [PubMed]
- Parra, G.; Bradnam, K.; Korf, I. CEGMA: a pipeline to accurately annotate core genes in eukaryotic genomes. Bioinformatics 2007, 23(9), 1061–1067. [Google Scholar] [CrossRef] [PubMed]
- Paszkiewicz, K.; Studholme, D. J. De novo assembly of short sequence reads. Briefings in bioinformatics 2010, 11(5), 457–472. [Google Scholar] [CrossRef] [PubMed]
- Patch, A. M.; Nones, K.; Kazakoff, S. H.; Newell, F.; Wood, S.; Leonard, C.; Holmes, O.; Xu, Q.; Addala, V.; Creaney, J.; Robinson, B. W.; Fu, S.; Geng, C.; Li, T.; Zhang, W. Germline and somatic variant identification using BGISEQ-500 and HiSeq X Ten whole genome sequencing. PloS one 2018, 13(1), e0190264. [Google Scholar] [CrossRef] [PubMed]
- Phillippy, A. M.; Schatz, M. C.; Pop, M. Genome assembly forensics: finding the elusive mis-assembly. Genome biology 2008, 9, 1–13. [Google Scholar] [CrossRef]
- Porubsky, D.; Ebert, P.; Audano, P. A.; Vollger, M. R.; Harvey, W. T.; Marijon, P.; Ebler, J.; Munson, K. M.; Sorensen, M.; Sulovari, A.; Haukness, M.; Ghareghani, M.; Human Genome Structural Variation Consortium; Lansdorp, P. M.; Paten, B. Fully phased human genome assembly without parental data using single-cell strand sequencing and long reads. Nature biotechnology 2021, 39(3), 302–308. [Google Scholar] [PubMed]
- Putnam, N. H.; O’Connell, B. L.; Stites, J. C.; Rice, B. J.; Blanchette, M.; Calef, R.; Troll, C. J.; Fields, A.; Hartley, P. D.; Sugnet, C. W.; Haussler, D.; Rokhsar, D. S.; Green, R. E. Chromosome-scale shotgun assembly using an in vitro method for long-range linkage. Genome research 2016, 26(3), 342–350. [Google Scholar] [PubMed]
- Qin, L.; Hu, Y.; Wang, J.; Wang, X.; Zhao, R.; Shan, H.; Li, K.; Xu, P.; Wu, H.; Yan, X.; Liu, L. Insights into angiosperm evolution, floral development and chemical biosynthesis from the Aristolochia fimbriata genome. Nature Plants 2021a, 7(9), 1239–1253. [Google Scholar]
- Qin, P.; Lu, H.; Du, H.; Wang, H.; Chen, W.; Chen, Z.; He, Q.; Ou, S.; Zhang, H.; Li, X.; Li, X. Pan-genome analysis of 33 genetically diverse rice accessions reveals hidden genomic variations. Cell 2021b, 184(13), 3542–3558. [Google Scholar]
- Ren, J.; Chaisson, M. J. lra: A long read aligner for sequences and contigs. PLOS Computational Biology 2021, 17(6), e1009078. [Google Scholar] [CrossRef] [PubMed]
- Rhie, A.; Walenz, B. P.; Koren, S.; Phillippy, A. M. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. Genome biology 2020, 21, 1–27. [Google Scholar] [CrossRef]
- Riess, O.; Sturm, M.; Menden, B.; Liebmann, A.; Demidov, G.; Witt, D.; Casadei, N.; Admard, J.; Schutz, L.; Ossowski, S.; Taylor, S.; Schaffer, S.; Schroeder, C.; Dufke, A.; Haack, T. Genomes in clinical care. Npj Genomic Medicine 2024, 9(1), 20. [Google Scholar] [CrossRef] [PubMed]
- Ruan, J.; Li, H. Fast and accurate long-read assembly with wtdbg2. Nature methods 2020, 17(2), 155–158. [Google Scholar] [PubMed]
- Salzberg, S. L.; Phillippy, A. M.; Zimin, A.; Puiu, D.; Magoc, T.; Koren, S.; Treangen, T. J.; Schatz, M. C.; Delcher, A. L.; Roberts, M.; Marçais, G.; Pop, M.; Yorke, J. A. GAGE: A critical evaluation of genome assemblies and assembly algorithms. Genome Research 2012, 22(3), 557–567. [Google Scholar] [PubMed]
- Sanger, F.; Coulson, A. R. A rapid method for determining sequences in DNA by primed synthesis with DNA polymerase. Journal of molecular biology 1975, 94(3), 441–448. [Google Scholar] [CrossRef] [PubMed]
- Schmieder, R.; Edwards, R. Fast identification and removal of sequence contamination from genomic and metagenomic datasets. PloS one 2011, 6(3), e17288. [Google Scholar] [CrossRef] [PubMed]
- Schneider, V. A.; Graves-Lindsay, T.; Howe, K.; Bouk, N.; Chen, H. C.; Kitts, P. A.; Murphy, T. D.; Pruitt, K. D.; Thibaud-Nissen, F.; Albracht, D.; Fulton, R. S.; Kremitzki, M.; Magrini, V.; Markovic, C.; McGrath, S. Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly. Genome research 2017, 27(5), 849–864. [Google Scholar] [CrossRef] [PubMed]
- Schott, T.; Kondadi, P. K.; Hänninen, M. L.; Rossi, M. Microevolution of a zoonotic Helicobacter population colonizing the stomach of a human host before and after failed treatment. Genome biology and evolution 2012, 4(12), 1310–1315. [Google Scholar] [CrossRef] [PubMed]
- Schuster, S. C. Next-generation sequencing transforms today’s biology. Nature methods 2008, 5(1), 16–18. [Google Scholar] [PubMed]
- Seppey, M.; Manni, M.; Zdobnov, E. M. BUSCO: Assessing Genome Assembly and Annotation Completeness. Methods in molecular biology 2019, 1962, 227–245. [Google Scholar] [CrossRef] [PubMed]
- Shafin, K.; Pesout, T.; Lorig-Roach, R.; Haukness, M.; Olsen, H. E.; Bosworth, C.; Armstrong, J.; Tigyi, K.; Maurer, N.; Koren, S.; Sedlazeck, F. J.; Marschall, T.; Mayes, S.; Costa, V.; Zook, J. M. Efficient de novo assembly of eleven human genomes using PromethION sequencing and a novel nanopore toolkit. BioRxiv 2019, 715722. [Google Scholar]
- Shafin, K.; Pesout, T.; Lorig-Roach, R.; Haukness, M.; Olsen, H.E.; Bosworth, C.; Armstrong, J.; Tigyi, K.; Maurer, N.; Koren, S.; Sedlazeck, F.J. Nanopore sequencing and the Shasta toolkit enable efficient de novo assembly of eleven human genomes. Nature biotechnology 2020, 38(9), 1044–1053. [Google Scholar] [CrossRef] [PubMed]
- Shen, C.; Du, H.; Chen, Z.; Lu, H.; Zhu, F.; Chen, H.; Meng, X.; Liu, Q.; Liu, P.; Zheng, L.; Li, X.; Dong, J.; Liang, C.; Wang, T. The chromosome-level genome sequence of the autotetraploid alfalfa and resequencing of core germplasms provide genomic resources for alfalfa research. Molecular Plant 2020, 13(9), 1250–1261. [Google Scholar] [CrossRef] [PubMed]
- Simao, F. A.; Waterhouse, R. M.; Ioannidis, P.; Kriventseva, E. V.; Zdobnov, E. M. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 2015, 31(19), 3210–3212. [Google Scholar] [CrossRef] [PubMed]
- Simpson, J. T.; Durbin, R. Efficient de novo assembly of large genomes using compressed data structures. Genome Research 2012, 22(3), 549–556. [Google Scholar] [PubMed]
- Simpson, J. T.; Wong, K.; Jackman, S. D.; Schein, J. E.; Jones, S. J.; Birol, I. ABySS: a parallel assembler for short read sequence data. Genome Research 2009, 19(6), 1117–1123. [Google Scholar] [CrossRef] [PubMed]
- Smeds, L.; Künstner, A. ConDeTri-a content dependent read trimmer for Illumina data. PloS one 2011, 6(10), e26314. [Google Scholar] [PubMed]
- Sohn, J. I.; Nam, J. W. The present and future of de novo whole-genome assembly. Briefings in bioinformatics 2018, 19(1), 23–40. [Google Scholar] [PubMed]
- Swathi, A.; Shekhar, M. S.; Katneni, V. K.; Vijayan, K. K. Genome size estimation of brackishwater fishes and penaeid shrimps by flow cytometry. Molecular biology reports 2018, 45, 951–960. [Google Scholar] [CrossRef] [PubMed]
- Tomato Genome Consortium, X. The tomato genome sequence provides insights into fleshy fruit evolution. Nature 2012, 485(7400), 635. [Google Scholar] [CrossRef]
- Toolkit for processing sequences in FASTA/Q formats. GitHub. Available online: https://github.com/lh3/seqtk (accessed on 5 January 2018).
- Vaser, R.; Sovic, I.; Nagarajan, N.; Sikic, M. Fast and accurate de novo genome assembly from long uncorrected reads. Genome research 2017, 27(5), 737–746. [Google Scholar] [CrossRef] [PubMed]
- Vega, L. Fundamentals of genetics; Waltham Abbey, Essex; Scientific e-Resources, 2019. [Google Scholar]
- Venter, J. C.; Adams, M. D.; Myers, E. W.; Li, P. W.; Mural, R. J.; Sutton, G. G.; Smith, H. O.; Yandell, M.; Evans, C. A.; Holt, R. A.; Gocayne, J. D.; Amanatides, P.; Ballew, R. M.; Huson, D. H.; Wortman, J. R. The sequence of the human genome. Science 2001, 291(5507), 1304–1351. [Google Scholar] [CrossRef] [PubMed]
- Walker, B. J.; Abeel, T.; Shea, T.; Priest, M.; Abouelliel, A.; Sakthikumar, S.; Cuomo, C. A.; Zeng, Q.; Wortman, J.; Young, S. K.; Earl, A. M. Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement. PloS one 2014, 9(11), e112963. [Google Scholar] [CrossRef] [PubMed]
- Wang, J. R.; Holt, J.; McMillan, L.; Jones, C. D. FMLRC: Hybrid long read error correction using an FM-index. BMC bioinformatics 2018, 19(1), 50. [Google Scholar] [CrossRef] [PubMed]
- Wang, J.; Veldsman, W. P.; Fang, X.; Huang, Y.; Xie, X.; Lyu, A.; Zhang, L. Benchmarking multi-platform sequencing technologies for human genome assembly. Briefings in bioinformatics 2023, 24(5), bbad300. [Google Scholar] [CrossRef] [PubMed]
- Wang, P.; Wang, F. A proposed metric set for evaluation of genome assembly quality. Trends in Genetics 2023, 39(3), 175–186. [Google Scholar] [PubMed]
- Wang, T.; Antonacci-Fulton, L.; Howe, K.; Lawson, H. A.; Lucas, J. K.; Phillippy, A. M.; Popejoy, A. B.; Asri, M.; Carson, C.; Chaisson, M. J. P.; Chang, X.; Cook-Deegan, R.; Felsenfeld, A. L.; Fulton, R. S.; Garrison, E. P. The Human Pangenome Project: a global resource to map genomic diversity. Nature 2022, 604(7906), 437–446. [Google Scholar] [CrossRef] [PubMed]
- Warren, R. L.; Coombe, L.; Mohamadi, H.; Zhang, J.; Jaquish, B.; Isabel, N.; Jones, S. J. M.; Bousquet, J.; Bohlmann, J.; Birol, I. ntEdit: scalable genome sequence polishing. Bioinformatics 2019, 35(21), 4430–4432. [Google Scholar] [CrossRef] [PubMed]
- Weirather, J. L.; de Cesare, M.; Wang, Y.; Piazza, P.; Sebastiano, V.; Wang, X. J.; Buck, D.; Au, K. F. Comprehensive comparison of Pacific Biosciences and Oxford Nanopore Technologies and their applications to transcriptome analysis; F1000Research 6, 2017. [Google Scholar]
- Wenger, A. M.; Peluso, P.; Rowell, W. J.; Chang, P. C.; Hall, R. J.; Concepcion, G. T.; Ebler, J.; Fungtammasan, A.; Kolesnikov, A.; Olson, N. D.; Topfer, A.; Alonge, M.; Mahmoud, M.; Qian, Y.; Chin, C. S. Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome. Nature biotechnology 2019, 37(10), 1155–1162. [Google Scholar] [CrossRef] [PubMed]
- Wick, R. R.; Holt, K. E. Polypolish: short-read polishing of long-read bacterial genome assemblies. PLoS computational biology 2022, 18(1), e1009802. [Google Scholar] [PubMed]
- Wick, R. R.; Judd, L. M.; Gorrie, C. L.; Holt, K. E. Unicycler: resolving bacterial genome assemblies from short and long sequencing reads. PLoS computational biology 2017, 13(6), e1005595. [Google Scholar] [CrossRef] [PubMed]
- Wong, J.; Coombe, L.; Nikolić, V.; Zhang, E.; Nip, K.M.; Sidhu, P.; Warren, R.L.; Birol, I. Linear time complexity de novo long read genome assembly with GoldRush. Nature Communications 2023, 14(1), 2906. [Google Scholar] [CrossRef] [PubMed]
- Xie, H.; Li, W.; Hu, Y.; Yang, C.; Lu, J.; Guo, Y.; Wen, L.; Tang, F. De novo assembly of human genome at single-cell levels. Nucleic Acids Research 2022, 50(13), 7479–7492. [Google Scholar] [PubMed]
- Xu, T.; Li, Y.; Zheng, W.; Sun, Y. A chromosome-level genome assembly of the blackspotted croaker (Protonibea diacanthus). Aquaculture and Fisheries 2022, 7(6), 616–622. [Google Scholar] [CrossRef]
- Yavas, G.; Hong, H.; Xiao, W. dnAQET: a framework to compute a consolidated metric for benchmarking quality of de novo assemblies. BMC genomics 2019, 20, 1–16. [Google Scholar] [CrossRef]
- Zerbino, D. R.; Birney, E. Velvet: algorithms for de novo short read assembly using de Bruijn graphs. Genome Research 2008, 18(5), 821–829. [Google Scholar] [CrossRef] [PubMed]
- Zhang, H.; Jain, C.; Aluru, S. A comprehensive evaluation of long read error correction methods. BMC genomics 2020, 21(6), 889. [Google Scholar] [CrossRef] [PubMed]
- Zhang, L.; Li, S.; Luo, J.; Du, P.; Wu, L.; Li, Y.; Zhu, X.; Wang, L.; Zhang, S.; Cui, J. Chromosome-level genome assembly of the predator Propylea japonica to understand its tolerance to insecticides and high temperatures. Molecular ecology resources 2020, 20(1), 292–307. [Google Scholar] [PubMed]
- Zhang, T.; Zhou, J.; Gao, W.; Jia, Y.; Wei, Y.; Wang, G. Complex genome assembly based on long-read sequencing. Briefings in bioinformatics 2022, 23(5), bbac305. [Google Scholar] [CrossRef] [PubMed]
- Zhang, W.; Chen, J.; Yang, Y.; Tang, Y.; Shang, J.; Shen, B. A practical comparison of de novo genome assembly software tools for next-generation sequencing technologies. PloS one 2011, 6(3), e17915. [Google Scholar] [CrossRef] [PubMed]
- Zhang, Z.; Yang, C.; Veldsman, W. P.; Fang, X.; Zhang, L. Benchmarking genome assembly methods on metagenomic sequencing data. Briefings in Bioinformatics 2023, 24(2), bbad087. [Google Scholar] [CrossRef] [PubMed]
- Zimin, A. V.; Salzberg, S. L. The genome polishing tool POLCA makes fast and accurate corrections in genome assemblies. PLoS computational biology 2020, 16(6), e1007981. [Google Scholar] [CrossRef] [PubMed]
- Zimin, A. V.; Puiu, D.; Luo, M. C.; Zhu, T.; Koren, S.; Marçais, G.; Yorke, J.; Salzberg, S. L. Hybrid assembly of the large and highly repetitive genome of Aegilops tauschii, a progenitor of bread wheat, with the MaSuRCA mega-reads algorithm. Genome research 2017, 27(5), 787–792. [Google Scholar] [PubMed]

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).