Relational-Core Graph Analytics Querying graphs at SQL scale, and why the node/edge model is a performance tax, not a truer picture of connected data

📅 2026-09-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了使用列式关系引擎和图查询语言处理大规模图分析问题,证明其性能优于原生图引擎,并指出节点/边模型在关系表中是冗余的。
📝 Abstract
A durable assumption holds that graph analytics requires a purpose-built graph engine, and that relational systems are ill-suited to connected data. We argue the opposite for the workloads enterprises actually run. A columnar relational engine fronted by a graph query language matches or exceeds native graph engines on analytical graph queries, and - decisively - scales past the point where in-memory graph engines fail. We further argue that the node/edge property graph is not a more faithful model of connected data but a re-encoding of relationships that already exist explicitly in relational tables; reconstructing them at query time is pure overhead. We present ClickGraph and its Databricks-dialect sibling DeltaGraph, systems that translate Cypher directly onto the native relational schema - the tables, columns, and foreign keys as they already exist - and execute in place on ClickHouse, Databricks, or in-process on lakehouse files, with no import and no separate cluster. Because the output is ordinary SQL, an underperforming query is an open optimization surface: it can be rewritten, and the engine itself extended. We support the argument with a peer system's own published benchmark, in which a columnar engine outruns Neo4j by two-to-four orders of magnitude, and with reproducible measurements across the LDBC Social Network Benchmark suite.
Problem

Research questions and friction points this paper is trying to address.

graph analytics
relational systems
node/edge model
performance tax
connected data
Innovation

Methods, ideas, or system contributions that make the work stand out.

columnar relational engine
graph query language
analytical graph queries
Cypher
relational schema