Score
Designs and implements formal encodings that translate Boolean and conditional constraints, rules, modalities, parameters, permutations and state descriptions into logical or model-native representations (e.g., propositional, modal, or solver input formats) so that downstream inference or search procedures can operate on them. This work produces domain and solution encodings, semantics‑preserving mappings, and interpretable rule/parameter encodings that retain information needed for correct constraint satisfaction, reconstruction of solutions, and reasoning about conditional or modal behaviors.
This paper addresses the challenges of embedding Constraint Handling Rules (CHR) in host languages and the disconnect between theoretical foundations and practical implementations. We propose the first higher-order embedding framework for CHR grounded in category theory. Methodologically, we introduce the novel application of initial algebra semantics combined with monad transformer techniques to formally model both CHR syntax and operational semantics, thereby unifying abstract theoretical definitions with executable semantics. Our main contributions are: (1) a general embedding framework that ensures both formal rigor and implementability; (2) a complete correctness proof establishing semantic soundness and completeness; (3) the first abstract execution algorithm for CHR, enabling modular and language-agnostic reasoning about rule application; and (4) a dual-language prototype implementation in Haskell and Python, empirically validating the framework’s effectiveness, expressiveness, and practical utility across diverse programming environments.
This paper addresses the lack of systematic support for theory morphisms and logical relations in the $λΠ$-calculus modulo rewriting framework. Methodologically, it introduces a unified extension mechanism that formally integrates both concepts for the first time within this framework and designs a pattern-based invariant verification procedure, reducing the proof of translation invariants to finite, decidable propositional checks. The main contributions are: (1) a structurally clear, machine-verifiable formalization of inductive translations—e.g., type erasure; (2) the first fully verified type-erasure instance in $λΠ$-calculus modulo rewriting; and (3) a reusable methodology for rigorously verifying the correctness of translations between formal systems.
To address the low accuracy of large language models (LLMs) on symbolic reasoning tasks—such as logic puzzles—this paper proposes the LLM-Constraint Hybrid framework. First, an LLM (Llama 3.1 70B) automatically formalizes natural-language puzzle descriptions into Logic.py, a logic-oriented domain-specific language; subsequently, a constraint solver performs exact symbolic execution. This two-stage paradigm enables the first end-to-end, LLM-driven logical modeling and symbolic solving pipeline, overcoming the accuracy limitations of pure LLM-based direct reasoning. Evaluated on the ZebraLogicBench benchmark, our method achieves 90.2% accuracy—surpassing the strongest baseline by 65.1 percentage points and establishing a new state-of-the-art. The results empirically validate the effectiveness and scalability of tightly integrating LLMs with symbolic reasoning tools.
Neural network interpretability lacks formal foundations and verifiability. Method: This paper introduces abstract interpretation—a rigorous program analysis framework—into Transformer interpretability research for the first time, establishing the first axiomatic framework for defining and verifying compositional, approximate semantic characterizations of model computations. It integrates circuit-level neuron analysis, attention pattern tracing, and logical trajectory reconstruction to reverse-engineer mechanistic behavior, and formally proves that the resulting explanations satisfy all axioms. Contribution/Results: The approach successfully reconstructs the complete stepwise 2-SAT solving algorithm implemented by a Transformer—including syntactic parsing and variable enumeration/evaluation—demonstrating the first verified white-box reconstruction from black-box behavior. This work provides a verifiable, reproducible theoretical foundation and technical methodology for mechanistic interpretability.
This work addresses the challenge of efficiently integrating Constraint Handling Rules (CHR) into the purely relational logic programming language miniKanren to support sophisticated constraint solving. We present the first seamless embedding of CHR into miniKanren, enabling tight coordination between constraint propagation and relational search while preserving the logical completeness of both mechanisms. The resulting system, chrKanren, significantly extends miniKanren’s expressiveness and demonstrates practical utility through applications in user-defined data structure semantics unification and the synthesis of MYTH-style relational interpreters, thereby validating the effectiveness and applicability of our approach.
This work addresses the longstanding absence of efficient decision procedures based on restricted nondeterministic matrix (RNmatrix) semantics, which has hindered automated theorem proving for paraconsistent, intuitionistic, and modal logics. The authors encode RNmatrix semantics—along with its row-elimination criteria—as SMT problems, leveraging off-the-shelf SMT solvers to decide formula validity and construct countermodels, thereby yielding a general-purpose automated prover. Their approach achieves the first complete implementation of a decision procedure for the entire $C_n$ hierarchy of paraconsistent logics, outperforming the state-of-the-art tools in this domain while attaining competitive performance in intuitionistic and modal logics. This breakthrough effectively overcomes the historical bottleneck of lacking scalable and practical reasoning mechanisms grounded in RNmatrix semantics.
This work addresses the inefficiency of traditional finite-domain propagation methods that handle difference constraints $x - y \leq d$ individually. It presents the first global propagator for difference constraints equipped with an explanation mechanism, unifying all such constraints into a single model and enforcing bounds consistency via shortest-path algorithms. The propagator is seamlessly integrated into a lazy clause generation (LCG) solving framework, overcoming the limitations of constraint-by-constraint propagation. For the first time in constraint programming, this approach enables synergistic global reasoning over difference constraints and conflict explanation. Experimental results demonstrate that the proposed method significantly outperforms standard propagation strategies in terms of solving efficiency.
This study investigates whether large language models (LLMs) perform faithful logical reasoning in legal tasks or merely rely on heuristic approximations. By systematically evaluating three paradigms—pure LLM classification, LLM-based formalization followed by natural language inference, and rigorous formal reasoning using the Z3 solver—on a manually relabeled subset of ContractNLI, the work reveals a systematic gap between legal pragmatic interpretation and formal entailment. It identifies three characteristic failure modes in LLM-driven formalization, including “scope sanitization.” Although formalization improves task accuracy, the experiments demonstrate that all models exhibit pervasive logical unfaithfulness, underscoring that high predictive accuracy does not guarantee logical correctness.
This work addresses the challenge that traditional knowledge systems, due to their separation of storage and computation, struggle to perform automated, structured domain-constrained reasoning. To overcome this limitation, the paper proposes a Representation-Computation Unified (RCU) paradigm, which embeds domain semantics directly into data by treating domains as structured fields within predicates (e.g., is_a(Apple, Company, @Business)). This enables domain-scoped inference without external rules. The core contributions include a novel representation method embedding domains into predicates, three key inference mechanisms—closure over domain fibers, typed inheritance, and write-time cycle detection—and their formal theoretical foundation. A symbolic reasoning engine implemented in 2,400 lines of Python and Prolog, based on a quadruple model, supports multi-constraint queries and arc-consistency solving. Empirical validation on ICD-11 multiple inheritance resolution and CBT clinical temporal reasoning demonstrates effectiveness, with complexity O(m(N/K)²), highlighting the critical impact of domain lattice sparsity on performance.