datalog programming

Designs and implements declarative analyses and encodings as Datalog programs: write relations and rules (logical clauses), program in Datalog, and formalize or extend Datalog semantics to capture needed constructs. Build hybrid workflows that combine Datalog with SMT filtering and neuro‑symbolic or probabilistic predicates to filter infeasible inferences, emit satisfying models, and perform constraint‑driven or probabilistic rule-based reasoning.

datalogprogramming

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of concisely expressing Datalog-style logical rules and queries within Lean, a highly expressive yet complex proof assistant based on the Calculus of Inductive Constructions (CIC). To this end, the authors propose a shallowly embedded domain-specific language (DSL) that enables, for the first time, a seamless integration of Datalog into Lean. The DSL supports declarative definitions of facts and rules, backward-chaining queries, and—crucially—the automatic translation of Datalog queries into theorems accompanied by proof scripts, thereby establishing bidirectional interoperability with Lean’s native reasoning framework. The effectiveness of the approach is demonstrated through three representative case studies, which collectively illustrate its expressiveness in rule formulation, readability of queries, and capability to support formal verification.

bidirectional interoperabilityDatalogdomain-specific language

Datalog with First-Class Facts

Nov 01, 2024
TG
Thomas Gilray
🏛️ Washington State University | Syracuse University | University of Illinois at Chicago

Datalog natively supports only flat atomic facts, making it inefficient for modeling and reasoning over recursive hierarchical structures (e.g., ASTs, derivation trees). Existing extensions—such as Datalog± and Soufflé—are either hampered by high-order quantification (compromising implementation simplicity and decidability) or constrained by algebraic data types lacking native indexing and rule-triggering mechanisms. This paper introduces *first-order facts*: a novel paradigm elevating structured facts to first-class entities, uniquely identified by Skolem terms. This preserves Datalog’s decidability while enabling native structural modeling, efficient indexing, and rule evaluation. Technically, the approach integrates Skolemization-based representation, MPI-optimized parallel communication, and custom indexing strategies. Experiments on diverse benchmarks demonstrate order-of-magnitude throughput improvements over state-of-the-art systems—including Nemo, Vlog, RDFox, and Soufflé—and scalable performance up to thousands of threads.

Datalog struggles with tree-structured data handlingDL$^{exists!}$ aims to improve efficiency in parallel reasoningExisting extensions complicate reasoning over recursive data types

Existing Datalog engines typically employ a uniform physical relation representation, which struggles to simultaneously optimize performance across diverse mixed operations—such as insertion, lookup, and containment checks—under varying workloads. This work presents the first systematic analysis of the relationship between seven-dimensional workload characteristics and physical representations, introducing a decision tree–based adaptive selection mechanism that dynamically matches recursive Datalog programs with their optimal physical representation. Experimental results demonstrate that this approach significantly improves evaluation efficiency and clearly identifies the key workload dimensions governing representation choice, thereby establishing a new paradigm for performance optimization in Datalog engines.

Datalogphysical representationsrecursive queries

FlowLog: Efficient and Extensible Datalog via Incrementality

Nov 02, 2025
HZ
Hangdong Zhao
🏛️ Microsoft Gray Systems Lab | University of Wisconsin, Madison | Meta Platforms Inc.

Existing Datalog systems struggle to balance efficiency and generality: specialized engines (e.g., Soufflé) achieve high performance but lack flexibility, while database-embedded approaches (e.g., RecStep) offer modularity yet hinder integration of Datalog-specific optimizations. This paper introduces FlowLog, a novel Datalog engine built upon incremental computation. Its key contributions are: (1) an explicit per-rule relational intermediate representation that decouples recursive control from logical execution; (2) preservation of Datalog-aware optimizations at the logical layer while reusing database execution primitives; and (3) novel techniques—structured optimization, lateral information propagation, and recursion-aware Boolean specialization. Implemented atop Differential Dataflow, FlowLog unifies batch and incremental evaluation via semi-naïve evaluation, logical fusion, subplan reuse, and algebraic specialization. Experiments demonstrate that FlowLog significantly outperforms state-of-the-art Datalog engines and modern databases across diverse recursive workloads, achieving high performance, strong scalability, and architectural simplicity.

Addressing high volatility in recursive workloads with robustness-first approachBridging efficiency and extensibility trade-offs in Datalog systemsSeparating recursive control from logical plans for optimization flexibility

Boolean Matrix Logic Programming

Aug 19, 2024
LA
L. Ai
🏛️ Imperial College London

Traditional logic programming relies on CPU-based symbolic computation, limiting scalability for large-scale Datalog inference. Method: This paper introduces Boolean Matrix Logic Programming (BMLP), the first approach to model bottom-up Datalog evaluation as Boolean matrix operations, supporting arbitrary binary recursive rules—including both linear and nonlinear variants. Its core innovation is a composable Boolean matrix inference framework: it defines modular logical operators and establishes a rigorous, semantics-preserving mapping from Datalog programs to sparse Boolean matrix transformations. Optimization of Boolean matrix multiplication and sparse storage further enhances efficiency. Results: On datasets with millions of facts, BMLP achieves 30× speedup over general-purpose systems (e.g., LogicBlox) and 9× over specialized systems (e.g., RDFox), significantly improving both scalability and execution efficiency of logic programming.

Accelerating datalog query evaluation using GPU matrix operationsEnabling efficient bottom-up inference for linear recursive programsScaling logic programming performance on large knowledge graphs

Latest Papers

What's happening recently
View more

Static filtering techniques have long been overlooked in Datalog and lack support for Answer Set Programming (ASP). This work presents the first unified generalization of static filtering to both Datalog and ASP, introducing a more expressive form of filtering predicates together with a corresponding theoretical framework. The approach enables efficient reasoning through logic program rewriting, static analysis, and controllable approximation strategies. While preserving full expressiveness, the method substantially enhances the performance of rule-based systems, achieving order-of-magnitude speedups on representative benchmarks and real-world datasets.

answer set programmingDatalogoptimization

Existing Datalog engines struggle to simultaneously achieve efficiency, scalability, and extensible semantics in static analysis, while also lacking robust support for rule debugging and incremental updates. This work proposes a novel approach that compiles Soufflé-style Datalog programs into executable Differential Dataflow programs, yielding a high-performance, memory-efficient static analysis framework capable of millisecond-scale incremental recomputation. The framework natively supports non-standard semantics—such as k-core analysis—and integrates in-browser performance profiling and rule-tuning capabilities. Evaluated on 24 real-world static analysis benchmarks, the system outperforms state-of-the-art engines in both runtime performance and scalability.

Datalogefficiencyextensibility

This work addresses a critical limitation in probabilistic Datalog-based program analysis: while individual alarms are often sound, they may be mutually exclusive, leading developers to investigate combinations that cannot co-occur in practice. To tackle this issue, the paper introduces PPProbe, the first conflict extractor specifically designed for this setting. PPProbe formalizes inconsistencies as minimal unsatisfiable subsets (MUSes) and innovatively integrates structural information from Datalog derivation graphs with probabilistic weights. It employs a structure-aware, bottom-up search strategy enhanced by UNSAT-core-based pruning to efficiently enumerate conflicts. Experimental evaluation on 70 benchmarks demonstrates that PPProbe achieves 2.5× to 24× higher throughput than the state-of-the-art MUS enumerator and, on average, filters out 47.7% of mutually exclusive alarms.

alarm rankingconflict extractionminimal unsatisfiable subsets

This work proposes NeuroLog, the first end-to-end vulnerability discovery framework that operates without requiring a build environment. Addressing the limitations of traditional static analysis—which depends on complete build setups—and large language models (LLMs)—which struggle with precise interprocedural data-flow tracking—NeuroLog leverages LLMs to extract type-aware data-flow facts function by function. These facts are then composed into cross-function paths using Datalog rules. Infeasible paths are pruned via an SMT solver, which also generates a SAT model used by the LLM to synthesize verifiable crash-inducing inputs. Evaluated on multiple open-source projects, NeuroLog reproduces eight known CVEs, including CVE-2023-38545 (CVSS 9.8), and discovers five new memory-safety vulnerabilities in libarchive—four previously unreported—all confirmed by AddressSanitizer. Each analysis completes in 37 seconds at a cost of approximately $0.005.

C/C++ source codecross-function dataflowlarge language models

Hot Scholars

CL

Carsten Lutz

Professor of Computer Science, University of Leipzig
Knowledge RepresentationArtificial IntelligenceLogic in Computer ScienceTheoretical Computer Science
JD

Jens Dietrich

Victoria University of Wellington
Software SecurityProgram AnalysisSoftware Engineering
YS

Yihao Sun

Syracuse University
Computer Science/Programming Language/HPC
KM

Kristopher Micinski

Syracuse University
Programming LanguagesStatic AnalysisAutomated ReasoningReverse Engineering