Sparse Binary Relation Representations for Genome Graph Annotation

  • Mikhail Karasikov
  • Harun Mustafa
  • Amir Joudaki
  • Sara Javadzadeh-No
  • Gunnar RätschEmail author
  • André KahlesEmail author
Conference paper
Part of the Lecture Notes in Computer Science book series (LNCS, volume 11467)


High-throughput DNA sequencing data is accumulating in public repositories, and efficient approaches for storing and indexing such data are in high demand. In recent research, several graph data structures have been proposed to represent large sets of sequencing data and to allow for efficient querying of sequences. In particular, the concept of labeled de Bruijn graphs has been explored by several groups. While there has been good progress towards representing the sequence graph in small space, methods for storing a set of labels on top of such graphs are still not sufficiently explored. It is also currently not clear how characteristics of the input data, such as the sparsity and correlations of labels, can help to inform the choice of method to compress the graph labeling. In this work, we present a new compression approach, Multi-BRWT, which is adaptive to different kinds of input data. We show an up to 29% improvement in compression performance over the basic BRWT method, and up to a 68% improvement over the current state-of-the-art for de Bruijn graph label compression. To put our results into perspective, we present a systematic analysis of five different state-of-the-art annotation compression schemes, evaluate key metrics on both artificial and real-world data and discuss how different data characteristics influence the compression performance. We show that the improvements of our new method can be robustly reproduced for different representative real-world datasets.


Sparse binary matrices Binary relations Genome graph annotation Compression 



We would like to thank the members of the Biomedical Informatics group for fruitful discussions and critical questions, and Torsten Hoefler and Mario Stanke for constructive feedback on the graph setup. Harun Mustafa and Mikhail Karasikov are funded by the Swiss National Science Foundation grant #407540_167331 “Scalable Genome Graph Data Structures for Metagenomics and Genome Annotation” as part of Swiss National Research Programme (NRP) 75 “Big Data”.


  1. 1.
    UK10K Project (2015). Accessed 2 Nov 2018.
  2. 2.
    Agarwala, R., et al.: Database resources of the national center for biotechnology information. Nucleic Acids Res. (2017).
  3. 3.
    Alipanahi, B., Muggli, M.D., Jundi, M., Noyes, N., Boucher, C.: Resistome SNP Calling via Read Colored de Bruijn Graphs. bioRxiv (2018).
  4. 4.
    Almodaresi, F., Pandey, P., Patro, R.: Rainbowfish: a succinct colored de Bruijn graph representation. In: Schwartz, R., Reinert, K. (eds.) 17th International Workshop on Algorithms in Bioinformatics (WABI 2017). Leibniz International Proceedings in Informatics (LIPIcs), vol. 88, pp. 18:1–18:15. Schloss Dagstuhl–Leibniz-Zentrum fuer Informatik, Dagstuhl (2017).
  5. 5.
    Álvarez-García, S., Brisaboa, N.: Compressed k2-triples for full-in-memory RDF engines. arXiv preprint arXiv:1105.4004 (2011)
  6. 6.
    Baraniuk, R., Davenport, M., DeVore, R., Wakin, M.: A simple proof of the restricted isometry property for random matrices. Constr. Approximation (2008).
  7. 7.
    Barbay, J., Claude, F., Navarro, G.: Compact binary relation representations with rich functionality. Inf. Comput. (2013).
  8. 8.
    Brisaboa, N.R., Ladra, S., Navarro, G.: k2-trees for compact web graph representation. In: Karlgren, J., Tarhio, J., Hyyrö, H. (eds.) SPIRE 2009. LNCS, vol. 5721, pp. 18–30. Springer, Heidelberg (2009). Scholar
  9. 9.
    Church, D.M., et al.: Modernizing reference genome assemblies. PLoS Biol. 9(7), e1001,091 (2011).
  10. 10.
    Ghosh, M., Gupta, I., Gupta, S., Kumar, N.: Fast compaction algorithms for NoSQL databases. In: Proceedings of International Conference on Distributed Computing Systems (2015).
  11. 11.
    Gog, S., Beller, T., Moffat, A., Petri, M.: From theory to practice: plug and play with succinct data structures. In: Gudmundsson, J., Katajainen, J. (eds.) SEA 2014. LNCS, vol. 8504, pp. 326–337. Springer, Cham (2014). Scholar
  12. 12.
    Iqbal, Z., Caccamo, M., Turner, I., Flicek, P., McVean, G.: De novo assembly and genotyping of variants using colored de Bruijn graphs. Nat. Genet. (2012).
  13. 13.
    Kokot, M., Długosz, M., Deorowicz, S.: KMC 3: counting and manipulating k-mer statistics. Bioinformatics 33(17), 2759–2761 (2017).
  14. 14.
    Muggli, M.D., et al.: Succinct colored de Bruijn graphs. Bioinformatics (2017).
  15. 15.
    Mustafa, H., et al.: Dynamic compression schemes for graph coloring. Bioinformatics (2018).
  16. 16.
    Navarro, G., Providel, E.: Fast, small, simple rank/select on bitmaps. In: Klasing, R. (ed.) SEA 2012. LNCS, vol. 7276, pp. 295–306. Springer, Heidelberg (2012). Scholar
  17. 17.
    Novak, A., Paten, B.: A graph extension of the positional Burrows Wheeler transform and its applications. Algorithms Mol. Biol. (2016).
  18. 18.
    Okanohara, D., Sadakane, K.: Practical entropy-compressed rank/select dictionary. In: Proceedings of the Meeting on Algorithm Engineering & Expermiments, pp. 60–70. Society for Industrial and Applied Mathematics (2007).
  19. 19.
    O’Leary, N.A., et al.: Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation. Nucleic Acids Res. (2016).
  20. 20.
    Pandey, P., Almodaresi, F., Bender, M.A., Ferdman, M., Johnson, R., Patro, R.: Mantis: a fast, small, and exact large-scale sequence-search index. Cell Syst. (2018).
  21. 21.
    Raman, R., Raman, V., Satti, S.R.: Succinct indexable dictionaries with applications to encoding k-ary trees, prefix sums and multisets. ACM Trans. Algorithms 3(4) (2007).
  22. 22.
    Solomon, B., Kingsford, C.: Improved search of large transcriptomic sequencing databases using split sequence bloom trees. J. Comput. Biol. 25(7), 755–765 (2018).
  23. 23.
    Stephens, Z.D., et al.: Big data: astronomical or genomical? PLoS Biol. (2015).
  24. 24.
    Weigel, D., Mott, R.: The 1001 genomes project for Arabidopsis thaliana. Genome Biol. 10(5), 107 (2009).
  25. 25.
    Zhang, G.: Bird sequencing project takes off. Nature 522(7554), 34–34 (2015).,

Copyright information

© Springer Nature Switzerland AG 2019

Authors and Affiliations

  1. 1.Department of Computer ScienceETH ZurichZurichSwitzerland
  2. 2.University Hospital Zurich, Biomedical Informatics ResearchZurichSwitzerland
  3. 3.SIB Swiss Institute of BioinformaticsLausanneSwitzerland

Personalised recommendations