Score
Designs, implements, and evaluates systems that automatically derive logical conclusions from formal premises, such as theorem provers, decision procedures, model checkers, SMT/SAT solvers, and proof search or proof-assistant automation; analyzes their inference algorithms, encodings, soundness, completeness, and performance trade-offs to produce correct and efficient automated deduction.
This work addresses the critical challenge of reliably integrating automated reasoning tools—such as theorem provers, SAT/SMT solvers, and termination analyzers—with proof assistants to build highly trustworthy systems. It presents a systematic survey and comparative analysis of two principal technical approaches: certification and formal verification. The study examines core methodologies including logical encoding, result replay and checking, and integration mechanisms within proof assistants. By elucidating the respective strengths and limitations of these methods and illustrating them through multiple successful case studies, the paper offers clear methodological guidance for constructing high-assurance automated reasoning systems, thereby substantially enhancing the verifiability and trustworthiness of their outputs.
Theorem provers exhibit unstable performance, poor scalability, and low reproducibility in behavioral verification—particularly for real-time and safety-critical software. Method: We systematically reproduce and extend existing benchmarks to construct an empirical evaluation framework featuring diverse behavioral models and logical specifications, enabling rigorous assessment of robustness, scalability, and reproducibility across mainstream theorem provers. Contribution/Results: Our study is the first to empirically establish strong correlations between irregular solver performance and structural problem characteristics—specifically temporal constraint density and state-space distribution. Leveraging these insights, we propose adaptive heuristic strategies and a self-optimizing solver architecture. The approach delivers measurable stability improvements for just-in-time verification in CI/CD pipelines and AI-augmented IDEs, significantly enhancing the practicality and trustworthiness of automated logical verification in high-assurance software development.
Despite its efficacy in isolated projects, deductive verification has yet to achieve broad industrial adoption. To identify root barriers and key enablers, this paper conducts semi-structured interviews with 30 practitioners, followed by thematic analysis. We systematically uncover fundamental obstacles—including high proof maintenance overhead, limited automation, poor tool usability, and lack of workflow integration—as well as critical enabling factors. Diverging from prior work, we empirically establish *usability* and *workflow adaptability* as core dimensions governing adoption. Based on these findings, we propose three actionable improvement principles: (1) enhancing automation support for proof construction and evolution; (2) reducing proof maintenance burden through modularization and abstraction; and (3) deepening integration with IDEs and CI/CD pipelines. Our empirically grounded insights provide concrete, evidence-based guidance for tool developers, practitioners, and researchers—bridging the gap between academic verification techniques and engineering practice.
This work proposes a novel paradigm termed “agent-based proof automation” to address the high cost of manually crafting lengthy formal proof scripts. In this approach, human experts supply key mathematical insights, while large language model (LLM) agents autonomously generate and iteratively refine proof scripts within the Lean 4 environment. Relying solely on off-the-shelf LLMs and lightweight verification tools, the method demonstrates— for the first time—the capacity for efficient, large-scale collaborative formal verification by LLM agents. Evaluated on the 14,000-line System Capless type safety proof, the system successfully completed 189 out of 217 tasks (87% success rate), with only 16% of the tasks requiring human intervention.
This paper addresses the lack of a unified metatheoretic characterization for program logics handling multi-branching effects—such as nondeterminism and probabilism. We propose a novel program logic framework centered on algebraic choice structures. Methodologically, we are the first to embed algebraic effects modeling directly into the core of Hoare logic, integrating modal semantics with a relatively complete proof system that supports general loops and uniform reasoning across effect types (e.g., nondeterministic and probabilistic). Our main contributions are: (1) the first relatively complete proof system for Hoare logic strictly extending it to cover multiple branching effects; (2) a unified metatheoretic account of multi-result programs; and (3) formal support for cross-model reuse of proof fragments—enabling verification transfer between distinct semantic models (e.g., relational, probabilistic, or game-based interpretations).
This study presents the first systematic evaluation of state-of-the-art large language models’ ability to generate formally verifiable proofs in graduate-level mathematics, including analysis, algebra, probability, and logic. The authors construct a private benchmark that pairs natural language mathematical problems with their formalized statements in Lean 4 and employ an agent-based framework to automatically assess whether model-generated proofs are accepted by the Lean 4 proof checker. Experimental results show that the best base model achieves an accuracy of 33.5%, while other models exhibit substantially lower performance. The work further provides a detailed empirical analysis of tool usage, failure modes, reasoning costs, and latency, offering valuable insights into the current capabilities and limitations of large language models in formal mathematical reasoning.
This work addresses the lack of reproducibility, auditability, and robustness to execution failures in current evaluations of logical reasoning agents. To overcome these limitations, the authors propose an “agentified evaluation” paradigm that models the evaluation process itself as agent behavior, introducing a standardized framework capable of task dispatching, execution budget control, output parsing, and structured failure logging. Built upon a unified agent interface, the framework integrates Z3Py-based automated formalization, SMT solving, and failure classification to enable automated, structured, and auditable assessment. Evaluated on a cleaned FOLIO validation set, the automated formalization agent achieves an accuracy of 86.70%, substantially outperforming a chain-of-thought baseline at 73.89%.
Existing automated proof synthesis methods struggle with complex theorems in interactive theorem provers and rely heavily on expert knowledge. This work presents the first systematic analysis of failed proof attempts, uncovering critical correlations between human expert proof patterns and successful proofs. Building on these insights, we propose Pattern-Guided Tactic Search (PGTS), a novel approach that integrates deep learning–driven proof synthesis, empirical analysis of proof scripts, and heuristic tactic search guided by expert-derived patterns. Experimental results demonstrate that PGTS improves upon existing tools by proving 8.05% more theorems on standard benchmarks on average and achieves a 20% higher success rate on previously unproven theorems, while also generating more concise proof scripts.
While current large language models can automatically fill proof holes (i.e., eliminate 'sorries') in interactive theorem proving, their generated formalizations often fail expert review due to ill-conceived definitions, insufficiently general theorems, or suboptimal API design. This work presents a semi-autonomous formalization of Grothendieck’s vanishing theorem as a case study and introduces expert review as a central criterion for evaluating the quality of automated formalizations. By integrating large language model assistance, interactive proving, and an iterative refactoring-compression pipeline, the study systematically assesses the high-level design usability of automatically generated content. The findings reveal that measuring success solely by 'sorry' closure is markedly inadequate; expert-driven refactoring substantially improves formalization quality, underscoring the critical role of expert acceptability in evaluating automated formalization efforts.
This work addresses the high degree of manual effort and tediousness inherent in existing automated reasoning algorithms—such as those for hyper-exponential quantifier elimination—for complexity analysis. The paper proposes a higher-order abstract interpretation framework grounded in operator semantics, which automatically abstracts symbolic programs into numerical recurrence relations. By integrating termination analysis, fixed-point theory, and SMT solving techniques, the method enables fully automated derivation and verification of asymptotic upper bounds on computational complexity. This approach substantially reduces human intervention while significantly enhancing the automation, efficiency, and scalability of complexity analysis for intricate algorithms.