Score
Designs and constructs formal specifications of programming-language semantics — operational, denotational, or hybrid — by defining core/welterweight calculi, typing and reduction rules, and formal mappings between semantic levels. Uses these formal models to reason about and analyze language behaviors and properties such as structural subtyping interactions, generics, type unions/sets, and type soundness.
This work addresses the scalability and practicality challenges in modeling operational semantics for programming languages. Methodologically, it introduces a mathematically lightweight yet semantically precise and extensible operational semantics framework, formalizing program computation steps to uniformly support semantic equivalence, reduction semantics, static analysis, compiler correctness proofs, and program property verification. Its key contributions are: (i) systematic modeling of multi-paradigm language features using minimal, accessible mathematical machinery—balancing theoretical rigor with engineering utility; and (ii) significantly enhanced portability and reusability of semantic models, demonstrated through successful formal verification of multiple production compilers and static analyzers. The framework provides a unified, scalable semantic foundation for programming language design, specification standardization, and trustworthy software construction.
Gradually typed languages require modeling recursion, errors, memory allocation, and dynamic type tags simultaneously, while verifying complex metatheoretic properties—including type equality reasoning and graduality—yet existing approaches are repetitive, ad hoc, and lack reusability. Method: We introduce guarded domain theory to gradual type semantics for the first time, unifying these key features; by combining guarded recursion with denotational semantics, we support step-indexed logical relations while preserving modularity and reusability. Contribution/Results: We formally construct a complete denotational model of a simple gradually typed λ-calculus in Guarded Cubical Agda. This yields the first mechanized proofs of βη-equivalence and graduality theorems. Our framework provides a trustworthy, general, and extensible semantic foundation for gradual type systems, enabling rigorous formal verification of their metatheory.
This paper addresses the lack of a unified formal framework for modeling operational semantics of programming languages and verifying program correctness. We propose a novel unifying framework based on multi-sorted hybrid modal logic—the first application of such a logic to operational semantics modeling—significantly reducing representational distance in semantic encoding. Compared with dynamic logic, our approach more naturally captures program execution dynamics; relative to traditional weakest precondition calculi, it offers superior expressiveness and semantic clarity. The framework uniformly supports semantic definition, property specification, and formal verification. Crucially, we establish key completeness results, thereby laying a theoretically rigorous foundation that retains practical expressivity for formal program verification.
Formal program specifications are notoriously difficult, error-prone, and inefficient to write manually. To address this, we propose a two-stage LLM-driven approach: dialogue-guided specification synthesis followed by mutation-based verification. First, multi-turn dialogues model complex semantic requirements; second, four mutation operators—insertion, replacement, deletion, and reordering—enable verifiability-driven selection, eliminating reliance on rigid templates or syntactic grammars. Our method integrates code understanding, prompt engineering, and heuristic verifiability assessment. Evaluated on SV-COMP and a custom Java benchmark comprising 385 programs, it generates 279 verifiable specifications. These achieve significantly higher completeness and accuracy than pure-LLM baselines and classical tools (e.g., Houdini, Daikon). To our knowledge, this is the first approach to achieve both high coverage and formal verifiability in fully automated specification generation.
This paper addresses untyped lambda calculus with global read-write state, developing a unified semantic framework for effectful functional computation. Methodologically, it pioneers the integration of intersection type systems with monadic algebraic effects semantics, concurrently defining operational semantics, denotational semantics, and a type system—while proving type preservation under reduction and expansion of state-term configurations. The contributions are threefold: (1) establishing completeness of type safety and convergence characterization; (2) employing intersection types to precisely capture termination behavior of stateful computations; and (3) providing a theoretically rigorous foundation—combining semantic precision and type-based guarantees—for functional languages with global state.
This work addresses the absence of a unified categorical bialgebraic denotational semantics for higher-order languages that simultaneously guarantees congruence of bisimilarity and coherence of denotational equivalence. Building upon the higher-order abstract GSOS framework, it realizes— for the first time—the bialgebraic semantic vision proposed by Turi and Plotkin by employing locally final coalgebras as the semantic domain for behavioral endofunctors. The resulting syntax-agnostic, compositional denotational semantics is parametric in the construction of the semantic domain, thereby uniformly accommodating both typed and untyped higher-order languages, as well as those featuring probabilistic or nondeterministic effects. This approach not only subsumes existing models such as step-indexed semantics but also ensures the compositionality of bisimilarity and semantic consistency across language variants.
This work presents the first systematic investigation into the capability of large language models (LLMs) to generate program specifications involving higher-order logical constructs, which are essential for expressing complex verification properties yet remain beyond the reach of existing LLMs that predominantly handle basic syntactic forms. The authors design four syntactic configurations spanning different levels of abstraction and establish a comprehensive evaluation framework to assess a range of representative LLMs on standard verification benchmarks. Experimental results demonstrate that LLMs can effectively produce valid higher-order logical expressions; moreover, integrating logical constructs with base syntax significantly enhances verification efficacy and robustness without substantially increasing verification overhead. The study also reveals distinct advantages of two refinement paradigms in specification generation.
This work addresses the limitations of relying on abstract syntax trees in program correctness verification by proposing an intrinsically defined interpreter based on Hoare logic derivations. By introducing an entry-indexing technique, the approach supports total correctness reasoning, well-founded functions, dynamic-frame-based local reasoning, and behavioral subtyping, and for the first time formally integrates all these features into a unified Hoare logic system. Implemented in Rocq, the mechanized interpreter constitutes the first fully formalized dynamic-frame Hoare logic, successfully verifying the correctness of complex programs with dynamic semantics. This provides a more intrinsic and extensible semantic foundation for program verification.
This work proposes a human-AI collaborative paradigm for formal software specification that mitigates the traditional barriers to industrial adoption—namely, the notational complexity and high expertise threshold—while preserving the benefits of early error detection and explicit invariants. The approach employs an intermediate language blending natural language with lightweight LaTeX mathematical notation, enabling AI-assisted review, refinement, and code generation. Crucially, it distinguishes between components requiring rigorous formalization and those amenable to flexible treatment. By deeply integrating AI into the specification authoring and verification workflow, this method achieves “correct-by-construction” development in a case study on organizational knowledge growth simulation, significantly reducing costs while ensuring early validation and design correctness.