Preloader

Multiplex de Bruijn graphs enable genome assembly from long, high-fidelity reads

  • Nurk, S. et al. The complete sequence of a human genome. bioRxiv https://doi.org/10.1101/2021.05.26.445798 (2021).

  • Miller, D. E. et al. Targeted long-read sequencing identifies missing disease-causing variation. Am. J. Hum. Genet. 108, 1436–1449 (2021).

    CAS 
    Article 

    Google Scholar 

  • Wenger, A. M. et al. Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome. Nat. Biotechnol. 37, 1155–1162 (2019).

    CAS 
    Article 

    Google Scholar 

  • Compeau, P. E., Pevzner, P. A. & Tesler, G. How to apply de Bruijn graphs to genome assembly. Nat. Biotechnol. 29, 987–991 (2011).

    CAS 
    Article 

    Google Scholar 

  • Bankevich, A. et al. SPAdes: a new genome assembly algorithm and its applications to single cell sequencing. J. Comput. Biol. 19, 455–477 (2012).

    CAS 
    Article 

    Google Scholar 

  • Zerbino, D. R. & Birney, E. Velvet: algorithms for de novo short read assembly using de Bruijn graphs. Genome Res. 18, 821–829 (2008).

    CAS 
    Article 

    Google Scholar 

  • Nurk, S. et al. HiCanu: accurate assembly of segmental duplications, satellites, and allelic variants from high-fidelity long reads. Genome Res. 30, 1291–1305 (2020).

    CAS 
    Article 

    Google Scholar 

  • Cheng, H. et al. Haplotype-resolved de novo assembly with phased assembly graphs. Nat. Methods 18, 170–175 (2021).

    CAS 
    Article 

    Google Scholar 

  • Myers, E. W. The fragment assembly string graph. Bioinformatics 21, ii79–ii85 (2005).

    CAS 
    PubMed 

    Google Scholar 

  • Li, D., Liu, C. M., Luo, R., Sadakane, K. & Lam, T. W. MEGAHIT: an ultra-fast single-node solution for large and complex metagenomics assembly via succinct de Bruijn graph. Bioinformatics 31, 1674–1676 (2015).

    CAS 
    Article 

    Google Scholar 

  • Pevzner, P., Tang, H. & Tesler, G. De novo repeat classification and fragment assembly. Genome Res. 14, 1786–1796 (2004).

    CAS 
    Article 

    Google Scholar 

  • Ye, C., Ma, Z. S., Cannon, C. H., Pop, M. & Yu, D. W. Exploiting sparseness in de novo genome assembly. BMC Bioinformatics 13, S1 (2012).

    Article 

    Google Scholar 

  • Kolmogorov, M. et al. metaFlye: scalable long-read metagenome assembly using repeat graphs. Nat. Methods 17, 1103–1110 (2020).

    CAS 
    Article 

    Google Scholar 

  • Rautiainen, M. & Marschall, T. MBG: minimizer-based sparse de Bruijn graph construction. Bioinformatics 37, 2476–2478 (2021).

    CAS 
    Article 

    Google Scholar 

  • Bloom, B. H. Space/time tradeoffs in hash coding with allowable errors. Commun. ACM 13, 422–426 (1970).

    Article 

    Google Scholar 

  • Karp, R. M. & Rabin, M. O. Efficient randomized pattern-matching algorithms. IBM J. Res. Dev. 31, 249–260 (1987).

    Article 

    Google Scholar 

  • Kolmogorov, M., Yuan, J., Lin, Y. & Pevzner, P. A. Assembly of long error-prone reads using repeat graphs. Nat. Biotechnol. 37, 540 (2019).

    CAS 
    Article 

    Google Scholar 

  • Mc Cartney, A. M. et al. Chasing perfection: validation and polishing strategies for telomere-to-telomere genome assemblies. Preprint at https://www.biorxiv.org/content/10.1101/2021.07.02.450803v1 (2021).

  • Roberts, M., Hayes, W., Hunt, B. R., Mount, S. M. & Yorke, J. A. Reducing storage requirements for biological sequence comparison. Bioinformatics 20, 3363–3369 (2004).

    CAS 
    Article 

    Google Scholar 

  • Chikhi, R. & Rizk, G. Space-efficient and exact de Bruijn graph representation based on a Bloom filter. Algorithms Mol. Biol. 8, 22 (2013).

    Article 

    Google Scholar 

  • Minkin, I., Pham, S. & Medvedev, P. TwoPaCo: an efficient algorithm to build the compacted de Bruijn graph from many complete genomes. Bioinformatics 33, 4024–4032 (2017).

    CAS 
    PubMed 

    Google Scholar 

  • Bzikadze, A. V. & Pevzner, P. A. Automated assembly of centromeres from ultra-long error-prone reads. Nat. Biotechnol. 38, 1309–1316 (2020).

    CAS 
    Article 

    Google Scholar 

  • Peng, Y., Leung, H. C. M., Yiu, S. M. & Chin, F. Y. L. IDBA—a practical iterative de Bruijn graph de novo assembler. in Research in Computational Molecular Biology. RECOMB 2010. Lecture Notes in Computer Science Vol. 6044, 426–440 (2010).

  • Mikheenko, A., Prjibelski, A., Saveliev, V., Antipov, D. & Gurevich, A. Versatile genome assembly evaluation with QUAST-LG. Bioinformatics 34, i142–i150 (2018).

    CAS 
    Article 

    Google Scholar 

  • Garg, S. et al. Chromosome-scale, haplotype-resolved assembly of human genomes. Nat. Biotechnol. 39, 309–312 (2021).

    CAS 
    Article 

    Google Scholar 

  • Idury, R. M. & Waterman, M. S. A new algorithm for DNA sequence assembly. J. Comput. Biol. 2, 291–306 (1995).

    CAS 
    Article 

    Google Scholar 

  • Pevzner, P. A., Tang, H. & Waterman, M. S. An Eulerian path approach to DNA fragment assembly. Proc. Natl Acad. Sci. USA 98, 9748–9753 (2001).

    CAS 
    Article 

    Google Scholar 

  • Chin, C. S. et al. Nonhybrid, finished microbial genome assemblies from long-read SMRT sequencing data. Nat. Methods 10, 563–569 (2013).

    CAS 
    Article 

    Google Scholar 

  • Chin, C. et al. Phased diploid genome assembly with single-molecule real-time sequencing. Nat. Methods 13, 1050–1054 (2016).

    CAS 
    Article 

    Google Scholar 

  • Koren, S. et al. Hybrid error correction and de novo assembly of single-molecule sequencing reads. Nat. Biotechnol. 30, 693–700 (2012).

    CAS 
    Article 

    Google Scholar 

  • Roberts, R. J., Carneiro, M. O. & Schatz, M. C. The advantages of SMRT sequencing. Genome Biol. 14, 405 (2013).

    Article 

    Google Scholar 

  • Ruan, J. & Li, H. Fast and accurate long-read assembly with wtdbg2. Nat. Methods 17, 155–158 (2020).

    CAS 
    Article 

    Google Scholar 

  • Mikheenko, A., Valin, G., Prjibelski, A., Saveliev, V. & Gurevich, A. Icarus: visualizer for de novo assembly evaluation. Bioinformatics 32, 3321–3323 (2016).

    CAS 
    Article 

    Google Scholar 

  • Hon, T. et al. Highly accurate long-read HiFi sequencing data for five complex genomes. Sci. Data 7, 399 (2020).

    CAS 
    Article 

    Google Scholar 

  • Tvedte, E. S. et al. Comparison of long-read sequencing technologies in interrogating bacteria and fly genomes. G3 (Bethesda) 11, jkab083 (2021).

    Article 

    Google Scholar 

  • Source link