Submitted:
22 September 2023
Posted:
27 September 2023
You are already at the latest version
Abstract
Keywords:
1. Introduction
- CSR++, a new graph data structure that supports fast in-place updates without sacrificing read-only performance or memory consumption.
- Our thorough evaluation that shows that CSR++ achieves the best of both read-only and update-friendly worlds.
- An in-depth analysis of the design space, regarding memory allocation, segment size, and synchronization mechanisms, to further improve the read and update performance of CSR++.
2. Background & Related Work
2.1. Graph Representations
2.1.1. Adjacency Matrices and Lists
2.1.2. Compressed Sparse Row (CSR)
2.2. Graph Mutations
2.2.1. In-Place Updates
2.2.2. Batching
2.2.3. Multi-Versioning and Deltas
3. CSR++: Design and Implementation
3.1. Graph Topology and Properties
3.1.1. Segments
3.1.2. Vertices
- length (4 bytes): The vertex degree. A length of -1 indicates a deleted vertex.
- neighbors (8 bytes): A pointer to the set of neighbors. As a space optimization, if length = 1 , this field directly contains the neighbor’s vertex ID.
- edge_properties (8 bytes): A pointer to the set of edge properties. As a space optimization, this field can be disabled in case the graph does not define edge properties.
3.1.3. Edges
- deleted_flag (2 bytes): For logical deletion of edges.
- vertex_id (2 bytes): The index of the neighbor in the segment; using 16 bits allows for segments with a capacity NUM_V_SEG of up to 65,536 entries.
- segment_id (4 bytes): The segment ID where the neighbor is stored.
3.1.4. Properties
3.1.5. Additional Structures
3.1.6. Synchronization
3.2. Update Protocols
3.2.1. Vertex and Edge Insertion
- Group the edges by their source vertices and convert both source and destination user keys to internal keys. The new vertices are inserted in CSR++ and each acquires a new internal ID. We keep this step sequential in CSR++, as it is very lightweight (see Section 4.6).
- Sort the new edges (parallel for each source vertex), and then insert them into the direct and reverse maps (also parallel for each source vertex).
- Sort the final edge arrays using a technique that merges two sorted arrays (i.e., the old edges and the new ones) and reallocate edge properties (parallel for each modified segment) according to the new order of edges.
3.2.2. Vertex and Edge Deletion
3.3. Algorithms on Top of CSR++
4. Evaluation
- How does CSR++ perform on read-only and on update workloads?
- How much memory does CSR++ consume on these workloads?
4.1. Sensitivity Analysis: Segment Size
4.2. Sensitivity Analysis: Improving Update Performance with HTM
4.3. Sensitivity Analysis: Memory Allocators
4.4. Read-Only: Algorithms
4.5. Read-Only: Sequential and Random Scans
4.6. Updates: Vertex Insertions
4.7. Updates: Batch Edge Insertions
4.8. Updates: Edge Insertions with Properties
4.9. Updates: Memory Consumption
4.10. Updates: Edge Deletions
4.11. Analytics after Graph Updates
4.12. Memory Footprint
5. Concluding Remarks
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Dhulipala, L.; Blelloch, G.; Shun, J. Julienne: a Framework For Parallel Graph Algorithms Using Work-efficient Bucketing. SPAA, 2017.
- Haubenschild, M.; Then, M.; Hong, S.; Chafi, H. ASGraph: a Mutable Multi-versioned Graph Container With High Analytical Performance. GRADES, 2016.
- Macko, P.; Marathe, V.J.; Margo, D.W.; Seltzer, M.I. LLAMA: efficient Graph Analytics Using Large Multiversioned Arrays. ICDE, 2015.
- Sevenich, M.; Hong, S.; van Rest, O.; Wu, Z.; Banerjee, J.; Chafi, H. Using Domain-specific Languages For Analytic Graph Databases. PVLDB 2016, 9, 1257–1268. [Google Scholar] [CrossRef]
- Shun, J.; Blelloch, G.E. Ligra: a Lightweight Graph Processing Framework For Shared Memory. PPoPP, 2013.
- Zhang, K.; Chen, R.; Chen, H. NUMA-Aware Graph-Structured Analytics. PPoPP, 2015.
- Page, L.; Brin, S.; Motwani, R.; Winograd, T. The PagerRank Citation Ranking: Bringing Order To The Web. Technical report, Stanford InfoLab, 1999.
- Dias, V.; Teixeira, C.H.C.; Guedes, D.; Meira, W.; Parthasarathy, S. Fractal: a General-Purpose Graph Pattern Mining System. SIGMOD, 2019.
- Kankanamge, C.; Sahu, S.; Mhedbhi, A.; Chen, J.; Salihoglu, S. Graphflow: an Active Graph Database. SIGMOD, 2017.
- Mawhirter, D.; Wu, B. AutoMine: harmonizing High-level Abstraction And High Performance For Graph Mining. SOSP, 2019.
- Neo4j. http://www.neo4j.org.
- Raman, R.; van Rest, O.; Hong, S.; Wu, Z.; Chafi, H.; Banerjee, J. PGX.ISO: parallel And Efficient In-memory Engine For Subgraph Isomorphism. GRADES, 2014.
- Sakr, S.; Elnikety, S.; He, Y. G-SPARQL: a Hybrid Engine For Querying Large Attributed Graphs. ACM CIKM, 2012.
- van Rest, O.; Hong, S.; Kim, J.; Meng, X.; Chafi, H. PGQL: a Property Graph Query Language. GRADES, 2016.
- PGQL: Property Graph Query Language. http://pgql-lang.org/.
- Staudt, C.L.; Sazonovs, A.; Meyerhenke, H. NetworKit: a Tool Suite For Large-Scale Complex Network Analysis. Network Science 2016, 4, 508–530. [Google Scholar]
- Wheatman, B.; Xu, H. Packed Compressed Sparse Row: a Dynamic Graph Representation. HPEC, 2018.
- Cheng, R.; Chen, E.; Hong, J.; Kyrola, A.; Miao, Y.; Weng, X.; Wu, M.; Yang, F.; Zhou, L.; Zhao, F. Kineograph: taking The Pulse Of A Fast-changing And Connected World. EuroSys, 2012.
- Madduri, K.; Bader, D.A. Compact Graph Representations And Parallel Connectivity Algorithms For Massive Dynamic Network Analysis. IPDPS, 2009.
- Kyrola, A.; Blelloch, G.; Guestrin, C. GraphChi: large-Scale Graph Computation on Just a PC. OSDI, 2012.
- Kumar, P.; Huang, H.H. GraphOne: a Data Store for Real-Time Analytics on Evolving Graphs. ACM Trans. Storage 2020, 15. [Google Scholar] [CrossRef]
- Ediger, D.; McColl, R.; Riedy, J.; Bader, D.A. STINGER: high performance data structure for streaming graphs. HPEC, 2012. [CrossRef]
- De Leo, D.; Boncz, P. Teseo and the Analysis of Structural Dynamic Graphs. Proc. VLDB Endow. 2021, 14, 1053–1066. [Google Scholar] [CrossRef]
- Bender, M.A.; Demaine, E.D.; Farach-Colton, M. Cache-Oblivious B-Trees. SIAM Journal on Computing 2005, 35, 341–358. [Google Scholar] [CrossRef]
- Firmli, S.; Trigonakis, V.; Lozi, J.P.; Psaroudakis, I.; Weld, A.; Chiadmi, D.; Hong, S.; Chafi, H. CSR++: a Fast, Scalable, Update-Friendly Graph Data Structure. OPODIS, 2021.
- GFE driver code. https://github.com/cwida/gfe_driver.
- Gartner top 10 data and analytics trends for 2019. https://www.gartner.com/smarterwithgartner/gartner-top-10-data-analytics-trends/.
- Sun, W.; Fokoue, A.; Srinivas, K.; Kementsietsidis, A.; Hu, G.; Xie, G.T. SQLGraph: an Efficient Relational-Based Property Graph Store. SIGMOD, 2015.
- Hong, S.; Chafi, H.; Sedlar, E.; Olukotun, K. Green-Marl: a DSL For Easy And Efficient Graph Analysis. ASPLOS, 2012.
- SPARQL query language For RDF. http://www.w3.org/TR/rdf-sparql-query/.
- Tinkerpop, Gremlin. https://github.com/tinkerpop/gremlin/wiki.
- Bratsas, C.; Chondrokostas, E.; Koupidis, K.; Antoniou, I. The Use of National Strategic Reference Framework Data in Knowledge Graphs and Data Mining to Identify Red Flags. Data 2021, 6. [Google Scholar] [CrossRef]
- Berners-Lee, T.; Hendler, J.; Lassila, O. ; others. The Semantic Web. Scientific American 2001, 284. [Google Scholar]
- Zeng, K.; Yang, J.; Wang, H.; Shao, B.; Wang, Z. A Distributed Graph Engine For Web Scale RDF Data. PVLDB 2013, 6. [Google Scholar]
- Property Graph model. https://github.com/tinkerpop/blueprints/wiki/Property-Graph-Model.
- Oracle Parallel Graph AnalytiX (PGX). https://www.oracle.com/middleware/technologies/parallel-graph-analytix.html.
- Hong, S.; Depner, S.; Manhardt, T.; van der Lugt, J.; Verstraaten, M.; Chafi, H. PGX.D: a Fast Distributed Graph Processing Engine. SC, 2015.
- Firmli, S.; Chiadmi, D. A Review Of Engines For Graph Storage And Mutations. EMENA-ISTL, 2020.
- Trigonakis, V.; Lozi, J.P.; Faltín, T.; Roth, N.P.; Psaroudakis, I.; Delamare, A.; Haprian, V.; Iorgulescu, C.; Koupy, P.; Lee, J.; Hong, S.; Chafi, H. aDFS: An Almost Depth-First-Search Distributed Graph-Querying System. 2021 USENIX Annual Technical Conference (USENIX ATC 21). USENIX Association, 2021, pp. 209–224.
- Roth, N.P.; Trigonakis, V.; Hong, S.; Chafi, H.; Potter, A.; Motik, B.; Horrocks, I. PGX.D/Async: a Scalable Distributed Graph Pattern Matching Engine. GRADES, 2017.
- Mariappan, M.; Vora, K. GraphBolt: dependency-Driven Synchronous Processing of Streaming Graphs. EuroSys, 2019.
- Ediger, D.; Riedy, J.; Bader, D.A.; Meyerhenke, H. Tracking structure of streaming social networks. IPDPSW, 2011.
- Boost adjacency list documentation. https://www.boost.org/doc/libs/1_67_0//libs/graph/doc/adjacency_list.html.
- Besta, M.; Fischer, M.; Kalavri, V.; Kapralov, M.; Hoefler, T. Practice of Streaming and Dynamic Graphs: concepts, Models, Systems, and Parallelism. CoRR 2019, abs/1912.12740, [1912.12740].
- Kallimanis, N.D.; Kanellou, E. Wait-free concurrent graph objects with dynamic traversals. OPODIS, 2016.
- Busato, F.; Green, O.; Bombieri, N.; Bader, D.A. Hornet: an Efficient Data Structure for Dynamic Sparse Graphs and Matrices on GPUs. HPEC, 2018. [CrossRef]
- Feng, G.; Meng, X.; Ammar, K. DISTINGER: a distributed graph data structure for massive dynamic graph processing. IEEE Big Data, 2015. [CrossRef]
- Green, O.; Bader, D.A. cuSTINGER: supporting dynamic graph algorithms for GPUs. HPEC, 2016. [CrossRef]
- Wheatman, B.; Xu, H. A Parallel Packed Memory Array to Store Dynamic Graphs. In 2021 Proceedings of the Symposium on Algorithm Engineering and Experiments (ALENEX); pp. 31–45. [CrossRef]
- Besta, M.; Hoefler, T. Accelerating Irregular Computations with Hardware Transactional Memory and Active Messages. CoRR 2020, abs/2010.09135, [2010.09135].
- Herlihy, M.; Moss, J.E.B. Transactional Memory: architectural Support for Lock-Free Data Structures. 1993, ISCA ’93. [CrossRef]
- Paradies, M.; Lehner, W.; Bornhövd, C. GRAPHITE: an Extensible Graph Traversal Framework For Relational Database Management Systems. SSDBM, 2015.
- Green-Marl code. https://github.com/stanford-ppl/Green-Marl.
- Falsafi, B.; Guerraoui, R.; Picorel, J.; Trigonakis, V. Unlocking Energy. USENIX ATC, 2016.
- OpenMP. https://www.openmp.org.
- Iosup, A.; Hegeman, T.; Ngai, W.L.; Heldens, S.; Prat-Pérez, A.; Manhardto, T.; Chafio, H.; Capotă, M.; Sundaram, N.; Anderson, M.; Tănase, I.G.; Xia, Y.; Nai, L.; Boncz, P. LDBC Graphalytics: a Benchmark for Large-Scale Graph Analysis on Parallel and Distributed Platforms. Proc. VLDB Endow. 2016, 9, 1317–1328. [Google Scholar] [CrossRef]
- LLAMA code. https://github.com/goatdb/llama.
- David, T.; Guerraoui, R.; Trigonakis, V. Everything You Always Wanted to Know about Synchronization but Were Afraid to Ask. Proceedings of the Twenty-Fourth ACM Symposium on Operating Systems Principles; Association for Computing Machinery: New York, NY, USA, 2013; SOSP ’13, p. 33–48. [CrossRef]
- Evans, J. A Scalable Concurrent malloc(3) Implementation for FreeBSD. BSDCan, 2006.
- Hunter, A.H.; Kennelly, C.; Gove, D.; Ranganathan, P.; Turner, P.J.; Moseley, T.J. Beyond malloc efficiency to fleet efficiency: a hugepage-aware memory allocator. OSDI, 2021.
| 1 | This article extends the OPODIS ’20 Conference publication by Firmli et al. [25] with (i) a more in-depth analysis of CSR++ in terms of design and performance, (ii) a sensitivity analysis of different design parameters, namely the segment sizes in CSR++, the use of different memory allocators, and synchronization with Intel’s Hardware Transactional Memory (HTM), and (iii) an extended performance evaluation, which compares CSR++’s performance to three more graph data structures, namely GraphOne [21], Teseo [23], and STINGER [22], and includes a new set of experiments from an external graph update benchmark framework, the GFE driver [26], as well as new data sets. |
| 2 | CSR++ can support different growing factors than to enable tuning edge insertion and memory consumption performance. |















| Data Set | #Vertices | #Edges | Source |
|---|---|---|---|
| 41 million | 1.4 billion | real-world graph | |
| LiveJournal | 4.8 million | 68 milion | real-world graph |
| Graph500-22 | 2.3 million | 64 million | synthetic |
| Uniform-24 | 8 million | 260 million | synthetic |
| Name | Type | Configuration |
|---|---|---|
| CSR++ | Segmentation based | Pre-allocated extra space for new edges. Deletion support enabled only on deletion workloads, in order to have fair comparison to LLAMA that does not support deletions by the default. |
| BGL [43] | Adjacency list | Bidirectional with default parameters. |
| CSR [53] | CSR | Implementation in the Green-Marl library [53]. |
| LLAMA [57] | CSR with delta logs | Read- and space-optimized with explicit linking. The fastest overall variant of LLAMA. Deletion support enabled only on deletion workloads. |
| STINGER [22] | Blocked Adjacency List | Linked list of blocks storing up to 14 edges. |
| GraphOne [21] | Multi-level Adjacency List and Circular Edge Log | Ignored archiving phase. |
| Teseo [23] | Transactional Fat Tree based on Packed Memory Arrays | Asynchronous rebalances delayed to 200ms and 1MB maximum leaf capacity. |
| Algorithm | Description |
|---|---|
| PageRank | Computes ranking scores for vertices based on their incoming edges. |
| Weakly Connected Components (WCC) | Computes affinity of vertices within a network. |
| Breadth-First Search (BFS) | Traverses the graph starting from a root vertex, visits neighbors, and stores the distance of vertices from the root vertex, as well as parents. |
| Weighted PageRank | Computes ranking scores like the original PageRank, but with weights and allows a weight associated with every edge. It requires accesses to edge properties. |
| Segment Size | 8 | 32 | 128 | 512 | 1024 | 2048 | 4096 | 16,384 | 32,768 |
| Memory overhead in bytes | 869,616 | 217,368 | 54,288 | 13,536 | 6768 | 3384 | 1656 | 360 | 144 |
| #Vertices | 10 K | 100 K | 1 M | 10 M |
|---|---|---|---|---|
| Time (ms)—0 vertex properties | 1.6 | 11 | 120 | 1188 |
| × (ms)—50 vertex properties | 10 | 32 | 181 | 1259 |
| Graph Structure | LiveJournal | Twitter-12 | Twitter-20 | Twitter-100 | |
|---|---|---|---|---|---|
| CSR | 0.53 | 11.09 | 11.09 | 11.09 | 11.09 |
| CSR++ read-only | 0.57 | 11.54 | - | - | - |
| CSR++ | 0.82 | 16.55 | 16.55 | 16.55 | 16.55 |
| LLAMA | 0.58 | 11.56 | 21.66 | 27.03 | 78.00 |
| LLAMA implicit linking | 0.58 | 11.56 | 19.02 | 23.99 | 73.64 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).