Score
Designs, builds, and analyzes formal specifications, mathematical models, and verification artifacts that establish correctness, safety, liveness, or other precise properties of systems, using model checking, theorem proving, type systems, and formal specification languages to generate and discharge proof obligations and counterexamples. Develops and integrates formal methods, tools, and workflows into development processes and conducts research on new formalisms, automation techniques, and toolchains to make formal reasoning scalable and usable in practice.
This work proposes a human-AI collaborative paradigm for formal software specification that mitigates the traditional barriers to industrial adoption—namely, the notational complexity and high expertise threshold—while preserving the benefits of early error detection and explicit invariants. The approach employs an intermediate language blending natural language with lightweight LaTeX mathematical notation, enabling AI-assisted review, refinement, and code generation. Crucially, it distinguishes between components requiring rigorous formalization and those amenable to flexible treatment. By deeply integrating AI into the specification authoring and verification workflow, this method achieves “correct-by-construction” development in a case study on organizational knowledge growth simulation, significantly reducing costs while ensuring early validation and design correctness.
This study addresses the high cost, error-proneness, and poor maintainability of manually writing Linear Temporal Logic (LTL) specifications—a core bottleneck in formal verification. We systematically survey and evaluate automated LTL specification mining methods. First, we propose a unified, multi-paradigm classification framework—covering constraint solving, neural networks, enumeration-based search, formal language inference, and grammar-guided learning—marking the first such comprehensive taxonomy. Second, we introduce a standardized evaluation methodology assessing scalability, interpretability, and noise robustness across approaches. Our analysis clarifies the applicability boundaries and inherent limitations of each paradigm and yields a practical, industry-oriented selection guide. The work significantly advances the automation level and reliability of LTL specification acquisition, bridging the gap between theoretical mining techniques and real-world verification practice.
To address the time-consuming and error-prone nature of manually constructing logical specifications for complex systems, this paper proposes an automated approach for inferring formal logical specifications from process event logs. Methodologically, it integrates workflow mining, pattern-driven logical translation, SMT solving (via Z3), and automated theorem proving (via Vampire) to achieve end-to-end generation of verifiable logical specifications from process models. Key contributions include: (i) the first unified empirical evaluation of specification quality on diverse, real-world event logs; (ii) a systematic analysis of how noise impacts specification structure and testability; and (iii) formal guarantees of satisfiability, internal consistency, and requirement conformance for generated specifications. Experimental results demonstrate high verification success rates and engineering practicality—even on noisy, real-world logs.
Natural language requirements are ill-suited for direct use in formal verification. Method: This work proposes an automated property generation framework integrating large language models (LLMs) with formal verification tools. It introduces an assertion generation mechanism extending beyond Linear Temporal Logic (LTL) to support numerical constraints and compositional system behavior modeling; integrates Claude 3.5 Sonnet with the ESBMC bounded model checker; and employs human-in-the-loop supervision to calibrate output quality. Contribution/Results: We first identify and characterize systematic impacts of LLM-induced model connection errors and numerical approximations on verification outcomes—reducing false positives and uncovering previously overlooked falsifiable scenarios. Evaluated on nine cyber-physical systems from Lockheed Martin, our approach achieves 46.5% verification accuracy—on par with NASA’s CoCoSim—while substantially lowering the barrier to formal verification and enhancing defect detection capability.
Automated verification in separation logic (SL) has long relied on ad hoc heuristics, lacking a systematic metatheory and suffering from poor scalability. Method: This paper establishes the first general SL metatheory grounded in category theory and algebraic structures—specifically functors, homomorphisms, and modules over rings—systematically integrating abstract algebra into SL automation. The framework supports compositional model instantiation and modular predicate synthesis for any data structure admitting an algebraic characterization. All results are formally verified in Isabelle/HOL, and an automatic algebraic instantiation algorithm is developed. Contribution/Results: Experiments demonstrate fully automated algebraic modeling of complex imperative program semantics—including lists, trees, and graphs—and yield inference engines whose performance matches state-of-the-art hand-crafted systems. This approach decisively overcomes the scalability limitations inherent in heuristic-based methods.
This work presents the first systematic investigation into the capability of large language models (LLMs) to generate program specifications involving higher-order logical constructs, which are essential for expressing complex verification properties yet remain beyond the reach of existing LLMs that predominantly handle basic syntactic forms. The authors design four syntactic configurations spanning different levels of abstraction and establish a comprehensive evaluation framework to assess a range of representative LLMs on standard verification benchmarks. Experimental results demonstrate that LLMs can effectively produce valid higher-order logical expressions; moreover, integrating logical constructs with base syntax significantly enhances verification efficacy and robustness without substantially increasing verification overhead. The study also reveals distinct advantages of two refinement paradigms in specification generation.
This study addresses the limitations of existing SysML verification approaches, which are often tool-dependent and restricted to performance properties, lacking support for automated validation of behavioral and interface requirements. To overcome these shortcomings, this work proposes a tool-agnostic, automated verification workflow driven by SysML test cases, integrating UML Testing Profile and behavioral diagram constructs to enable unified validation of multidimensional attributes—including behavior, timing, and state responses. The methodology was developed through a mixed-methods research strategy combining literature review and stakeholder interviews, and its efficacy was empirically validated across two independent SysML toolchains. The approach not only transcends the constraints of conventional parametric methods but also enables automatic traceability of verification results back to the original model elements.
This work addresses the challenge of providing verifiable safety assurance for robotic systems in safety-critical domains, where traditional assurance cases rely on manually generated evidence that is costly, error-prone, and difficult to maintain. The paper proposes a model-based automated approach that deeply integrates formal verification into the assurance workflow. It employs RoboChart—a domain-specific modeling language with formal semantics—to capture system designs, and introduces a template-driven mechanism to automatically translate natural-language requirements into formal assertions. These assertions are then discharged through a combination of model checking and theorem proving tools, yielding formally verified evidence that can be seamlessly integrated into assurance cases. Case studies demonstrate that the proposed method significantly enhances the reliability, maintainability, and degree of automation in safety argumentation.
This work addresses the challenge of formally verifying mature, safety-critical industrial C++ codebases by strategically integrating theorem proving (PVS) and model checking (SeaHorn), augmented with large language models to assist in specification construction. The approach is applied to the core order book algorithm of Stellar’s SDEX blockchain module. The verification effort successfully establishes critical correctness properties—including state consistency and unreachability of erroneous states—uncovers discrepancies between documentation and implementation, and produces reusable formal artifacts. These assets enable continuous validation of invariants during future code evolution, thereby enhancing long-term reliability and maintainability of the system.
Traditional formal verification lacks mechanisms for knowledge accumulation and cross-system reuse, making it difficult to transfer specifications, contracts, and proofs. This work proposes a novel paradigm that integrates artificial intelligence with formal methods, pioneering the combination of large language models and graph-based representations to enable semantic guidance across heterogeneous notations and abstraction levels. By leveraging automated contract synthesis, semantic artifact reuse, and compositional refinement theory, the authors construct a hybrid reasoning framework that ensures formal reliability while supporting continuous synthesis and migration of verification artifacts. This approach lays the foundation for a cumulative and evolvable verification ecosystem, paving the way toward scalable, knowledge-driven next-generation verification systems.