Score
Designs and implements systems that combine neural models and symbolic reasoning to translate natural-language or structured inputs into formal, solver-ready specifications such as typed logical rules, finite-domain constraints, or program-like specifications. Builds pipelines that synthesize or refine these symbolic artifacts, integrate them with solvers and symbolic reasoners, and emit provenance metadata for traceability and debugging.
Neural-symbolic approaches face practical limitations due to weak semantic generalization, the difficulty of predefining complex rules, and growing skepticism about their competitiveness in the era of large language models. This work presents the first task-oriented systematic survey of neural-symbolic AI, focusing on integrating symbolic systems to enhance the interpretability and reasoning capabilities of black-box models. By analyzing task-specific hybrid architectures in domains such as natural language processing and computer vision, the study highlights the real-world utility of neural-symbolic methods. It further provides reproducible code and in-depth annotations, offering researchers a practical design guide for developing interpretable AI systems tailored to concrete tasks, thereby fostering the continued evolution of this paradigm amid the rise of large models.
This work addresses the challenge of automatically generating ACSL formal specifications for C programs. We propose a neuro-symbolic collaborative approach that integrates the DeepSeek-R1 large language model with the Frama-C toolchain—specifically its EVA abstract interpreter and PathCrawler path coverage analyzer—to systematically uncover how code defects influence specification generation tendencies. We design a user-controllable, intent- or implementation-oriented prompting mechanism that explicitly distinguishes high-level behavioral specifications from low-level implementation-specific ones. Furthermore, we introduce a multi-stage symbolic-augmented reasoning paradigm, dynamically injecting symbolic analysis feedback into the LLM’s generation process. Experimental evaluation demonstrates substantial improvements in specification accuracy and semantic soundness: critical specification error rates decrease by 37%, and the method enables on-demand generation of high-quality, mechanically verifiable ACSL annotations.
This work addresses the challenge of format and semantic errors produced by small language models when translating natural language into first-order logic (FOL), which undermines the reliability of symbolic reasoning. To mitigate this, the authors propose a staged incremental reasoning framework: first, a large language model synthesizes training data to supervise fine-tuning of the small model; then, the translation process is decoupled into predicate generation and FOL formulation stages. An external verification module is introduced to detect and correct predicate arity errors, thereby enhancing translation accuracy. Evaluated on four logical reasoning benchmarks, the approach significantly reduces error rates, improves predicate coverage, and boosts overall reasoning performance, bringing small models closer to reliable, verifiable symbolic reasoning systems.
This work addresses critical challenges in safety-critical rule-based systems—namely poor scalability, fragility, and goal mis-specification—which often lead to reward hacking and failures in formal verification. To overcome these limitations, the authors propose a neuro-symbolic causal framework that integrates first-order logic abductive trees, structural causal models, and deep reinforcement learning within a MAPE-K control loop. A novel meta-layer architecture enables the automatic synthesis and formal verification of rules from natural language objectives. This meta-layer comprises a goal/rule synthesizer and a rule verification engine, which iteratively generate necessary and sufficient causal rule sets grounded in legal and safety principles provided by human experts. Evaluated in an autonomous driving scenario, the approach successfully derives a minimal yet complete rule set, formally encoded as logical constraints, demonstrating its modularity, traceability, and practical applicability.
Existing neural-symbolic NLP approaches rely on static solver integration, limiting adaptability to diverse formal reasoning paradigms. Method: We propose the first adaptive multi-paradigm neural-symbolic reasoning framework that leverages large language models (LLMs) to infer implicit reasoning paradigms—e.g., first-order logic, constraint solving, or inductive reasoning—from natural language questions, and dynamically orchestrates specialized symbolic solvers via an automated formalization interface and multi-logic solver scheduler. Contribution/Results: Experiments demonstrate significant improvements over strong baselines: +27% accuracy over GPT-4o and +6% over DeepSeek-V3.1 on multi-paradigm reasoning tasks. Under zero-shot and chain-of-thought settings, our method boosts GPT-4o’s performance by up to 10%. To our knowledge, this is the first framework enabling end-to-end, natural-language-driven, adaptive formal reasoning.
Reactive synthesis faces dual challenges of high algorithmic complexity and the difficulty of writing formal specifications. This work proposes a neurosymbolic approach that, for the first time, incorporates natural language specifications into reactive synthesis by leveraging a large reasoning model to generate Verilog circuits and integrating a model checker to provide symbolic feedback for iterative refinement. The method establishes an end-to-end natural synthesis pipeline that outperforms existing specialized tools on benchmarks from the annual synthesis competition. Notably, it achieves performance comparable to hand-crafted formal specifications when using natural language inputs and scales to the synthesis of undecidable parameterized systems.
This paper describes our system submitted to SemEval-2026 Task 11: Disentangling Content and Formal Reasoning in Large Language Models. We present an efficient modular neuro-symbolic approach, combining a symbolic prover with small reasoning LLMs (4B parameters). The system consists of an LLM-based parser that translates natural language syllogisms to a first-order logic (FOL) representation, an automated theorem prover, and two optional modules: machine translation for multilingual inputs and a symbolic retrieval component for the identification of relevant premises. The system achieves competitive accuracy and relatively low content effect on most subtasks. Our ablations show that this approach outperforms LLM-based zero-shot baselines in this parameter size range, but also reveal limited multilingual capabilities of small LLMs. Finally, we include a discussion of the task's main ranking metric and analyze its limitations.
This work addresses the limited reliability of large language models in tasks requiring explicit symbolic structures, multi-step reasoning, and uncertainty representation. The authors propose a neuro-symbolic framework that compiles natural language reasoning problems into executable Narsese programs, leveraging the OpenNARS runtime to ensure semantic alignment through program execution. Key contributions include the creation of NARS-Reasoning-v0.1—the first benchmark integrating natural language, first-order logic, and executable Narsese—and the introduction of a language structure-aware (LSP) training paradigm. The approach combines deterministic FOL-to-Narsese compilation, LoRA fine-tuning of Phi-2, and three-label (True/False/Uncertain) supervised learning. Experimental results demonstrate that the benchmark effectively supports supervised fine-tuning and enables interpretable evaluation grounded in program execution.
Existing approaches that manually translate domain theories into neural network architectures suffer from limited generality, verifiability, and scalability, often failing to guarantee strict alignment between models and prior knowledge. This work proposes a theory compiler that, for the first time, enables automatic mapping from formalized, typed domain theories to provably correct neural architectures. The method integrates formal methods, type systems, neural architecture design, and statistical learning theory through a universal theory language, a compositional compilation algorithm, and formal verification criteria, with large language models assisting in end-to-end compilation. The resulting architectures are theoretically guaranteed to achieve generalization performance comparable to or better than handcrafted designs while substantially reducing reliance on large training datasets.