๐ค AI Summary
This work addresses the prevalence of structural and semantic errors in Cypher queries generated by large language models, which can trigger database anomalies. The authors propose a defensive pre-execution validation framework that interposes a four-stage verification pipeline between query generation and execution on production databases. This pipeline integrates multi-stage syntactic and constraint checks, mirrored graph execution, cost-aware execution plan gating, and a language model feedback loop. The framework achieves zero false positives in intercepting structural errorsโthe first such result to dateโand clearly delineates the boundary between structural and semantic validation. Evaluated on CypherBench, the system introduces only 5.6 ms of latency, attains an average correction success rate of 89% across five language models, and detects all three categories of structural errors with 100% recall and zero false positives on template-based datasets.
๐ Abstract
Language models acting as agents over knowledge graphs generate Cypher queries that fail structurally (crashing at the database) or semantically (executing but returning wrong results). We place a pre-execution gate between query generation and a production Neo4j database. The gate validates structure through a four-backend chain culminating in execution against a mirror graph at 5.6 ms median latency. Structurally broken queries are routed to a corrector that iterates structured error feedback through a language model. On seven CypherBench schemas (2348 questions, ACL 2025) the pipeline maintains generation accuracy on every model tested, confirming it operates as a safe defensive layer. The corrector achieves 81% to 95% success across five models (mean 89%). On a template-generated corpus across nine schemas the gate catches 100% of parse errors, 100% of constraint violations, and 100% of schema-reference errors in path queries with labelled endpoints, at zero false positives across 1135 queries. Property sibling-swaps where the substituted name is valid on the target label score 0%, marking the formal boundary where structural validation ends and semantic validation must begin. A planner-based cost gate flags catastrophic plan structures before execution.