DatalogBench: Evaluating Large Language Models on Text-to-Datalog Synthesis

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic evaluation for Datalog programs generated by large language models (LLMs) by constructing a benchmark comprising 136 tasks and proposing reliable metrics based on execution verification and mutation analysis. Through experiments involving six prompting strategies applied to six LLMs and coding agents, results indicate that direct prompting achieves a maximum match rate of 68.4%, whereas coding agents attain 83.8% while effectively eliminating most compilation errors. The research precisely identifies semantic and compilation errors, revealing that recursive reasoning remains a core challenge for current models. Overall, this work establishes a new paradigm for evaluating the logical programming capabilities of LLMs.
📝 Abstract
Datalog underpins reasoning tasks such as program analysis, but its programs are hard to write. Existing synthesizers automate this task but require users to state their intent as input-output examples. Large language models (LLMs) suggest a more natural route, text-to-Datalog synthesis from a natural-language question, yet how well they do so has not been systematically evaluated. We present DatalogBench, a benchmark of 136 text-to-Datalog synthesis tasks curated from existing Datalog-based artifacts. Synthesized programs are graded by execution on held-out inputs against an oracle validated by mutation analysis. Across six LLMs and four prompting configurations, exact match peaks at 68.4%, and relation descriptions or an input-output example have only modest, model-dependent effects. Under direct prompting, most failures occur at compile time, typically because a model invents auxiliary predicates that it never declares or types consistently. Two coding agents reach up to 83.8% and eliminate nearly all such failures, leaving mostly semantic errors concentrated in recursive tasks. DatalogBench thus identifies recursive reasoning and decomposition as open challenges for current LLMs and agents, and offers a reliable, execution-grounded measure of both.
Problem

Research questions and friction points this paper is trying to address.

Text-to-Datalog synthesis
Large Language Models
Benchmark evaluation
Datalog
Recursive reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Text-to-Datalog Synthesis
Benchmark Evaluation
Large Language Models
Coding Agents
Recursive Reasoning
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yuan Li
The State Key Laboratory of Blockchain and Data Security, Zhejiang University
H
Hanyun Jiang
The State Key Laboratory of Blockchain and Data Security, Zhejiang University
G
Guowei Tian
The State Key Laboratory of Blockchain and Data Security, Zhejiang University
Chengpeng Wang
Chengpeng Wang
Purdue University
AI for CodeProgram AnalysisSoftware Engineering
Peisen Yao
Peisen Yao
Zhejiang University
Programming LanguagesSoftware EngineeringLogic and VerificationSecurity