🤖 AI Summary
This study addresses the limitation of existing Text-to-Cypher benchmarks, which support only single-turn queries and fail to reflect multi-turn conversational scenarios, by constructing the first conversational Text-to-Cypher evaluation benchmark. Methodologically, it designs a dual-evaluation protocol comprising a guided Oracle and a fully autonomous Agent, spanning seven knowledge graphs and thirteen dialogue phenomena. Furthermore, it introduces the novel concept of "autonomy divergence," revealing that error management constitutes a core competency independent of generation capability. Experimental results demonstrate that even the best-performing model achieves merely 64.7% execution accuracy, with session-level correctness falling below 5%. These findings highlight critical bottlenecks in autonomous decision-making and fault tolerance during multi-turn interactions, thereby establishing new research challenges for the field.
📝 Abstract
Graph databases are increasingly queried through natural language, yet every existing benchmark evaluates isolated single-turn queries rather than the multi-turn sessions through which analysts actually work. We introduce CypherTurn, the first benchmark for conversational Text-to-Cypher evaluation, comprising 721 sessions and 5,927 turns across 7 knowledge graphs and 13 conversational phenomena. We evaluate 15 models under a guided oracle protocol and a fully autonomous agentic protocol, yielding four findings. First, the best model reaches only 64.7% execution accuracy, and session-level correctness remains below 5%. Second, despite strong overall rank correlation, frontier models exhibit a consequential reordering of the top of the leaderboard under autonomous operation, a phenomenon we term the Autonomy Divergence, which reveals error-management as a partially independent capability from raw generation skill. Third, scaling action budgets from x3 to x10 fails to close the autonomy gap, as the strongest frontier models self-limit to approximately two actions per turn regardless of available budget. Fourth, single-turn Cypher fine-tuning degrades multi-turn instruction following, while architecture-appropriate specialization outperforms several frontier models. These results establish CypherTurn as an open challenge for conversational graph database reasoning. Code and data are available at https://github.com/BarryQ/CypherTurn.