Score
Designs and implements formal language and program representations such as abstract syntaxes, parsers, abstract syntax trees, control‑flow and other graph abstractions, and interface/API specifications, together with formal semantics and declarative specifications that support analysis. Develops and applies abstraction methods and refinement strategies (e.g., abstract interpretation, model/policy/graph abstraction), and builds techniques to compress, segment, cluster, embed, and compare high‑dimensional execution traces or state‑action sequences and to enable flow‑sensitive analyses.
Model checking temporal properties of safety-critical embedded C programs faces significant challenges due to the difficulty of constructing accurate, tractable abstractions manually. Method: This paper proposes an automated abstraction modeling approach grounded in verified component contracts. It directly translates state-transition contracts into language-agnostic flow graphs (FGs), integrating static analysis and abstract interpretation to achieve lightweight, high-fidelity abstraction. Contribution/Results: The work establishes, for the first time, a formal semantic mapping from contracts to flow graphs, enabling compositional model checking and drastically reducing manual abstraction effort. Experiments on real-world safety-critical C code demonstrate fully automated construction of high-precision abstract models. The approach improves feasibility, efficiency, and scalability of timing property verification and has been integrated into an end-to-end prototype toolchain.
Modular control-flow handling in abstract interpretation and supporting multiple analysis strategies—such as path- vs. flow-sensitivity, forward vs. backward directionality, and upper vs. lower approximations—traditionally relies on complex monad transformers, leading to implementation brittleness and poor composability. Method: This paper introduces the *cumulative abstract semantics* framework, the first to incorporate *scoped effects* into abstract interpretation. It decouples syntactic structure from semantic behavior via two classes of effect handlers: *syntax-resolving* and *domain-semantics-introducing*. A single syntax-driven interpreter suffices to generate diverse dynamic evaluators and static analyzers. Contribution/Results: The framework eliminates heavyweight data structures, preserving expressiveness while drastically reducing implementation complexity for multi-strategy analyses. It enhances maintainability, composability, and modularity—providing a concise, unified, and extensible theoretical and practical foundation for modular program analysis.
This work addresses three fundamental challenges in model checking—insufficient precision, state-space explosion, and spurious counterexamples—by systematically reducing model checking to program verification. Methodologically, we design MOKA, a domain-specific language that encodes ACTL and universal μ-calculus formulas as programs; within an abstract interpretation framework, we construct a Kleene algebraic semantic model and introduce locally complete abstractions coupled with counterexample-guided dynamic domain refinement, synergistically combining under-approximation and abstraction for controllable precision enhancement. Our contributions are threefold: (1) the first rigorous reduction of model checking to program verification under abstract interpretation; (2) support for non-partitioning abstractions, significantly reducing false positives; and (3) theoretical guarantees for complete detection of violating initial states. The resulting analyzer is general-purpose and precision-tunable.
Software architecture suffers from ambiguous abstraction concepts and inadequate tool support. Method: This work systematically reconstructs the seminal 1995 architectural model and proposes, for the first time, a practice-grounded conceptual framework for architectural abstraction—elevating component composition relationships to system-level abstractions that are formally modelable and verifiable. It integrates architectural description language (ADL) design, abstract modeling, prototype tool development, and diachronic historical analysis. Contribution/Results: The study establishes software architecture as an independent concern with rigorous theoretical foundations. Its outcomes catalyzed a surge in ADL research, laid the groundwork for model-based systems engineering (MBSE), and continue to inform the design of cloud-native, microservice, and AI-driven architectures. The framework significantly enhances the expressiveness, formal verifiability, and engineering applicability of architectural abstractions.
Manual tuning of abstraction strategies in static program analysis is labor-intensive and struggles to balance precision and efficiency. Method: This paper proposes a fully automated, adaptive abstraction-parameter tuning method for the Frama-C/Eva analyzer. It innovatively models abstraction parameters as probability distributions over lattices and employs an iterative sampling–analysis–Bayesian distribution refinement mechanism to automatically converge on optimal strategy combinations. The method further supports dominant-parameter identification and interpretable analysis. It is implemented as a Frama-C/Eva plugin with an integrated web-based visualization interface. Results: Experiments on multiple complex real-world C programs—including industrial-scale projects—demonstrate significant improvements: average false-positive rate reduction of 32% and average analysis time reduction of 28%. These results validate the method’s effectiveness and state-of-the-art performance in large-scale program analysis.
This work addresses the challenge of identifying structural and semantic similarities across imperative programs written in different languages by proposing a unified graph representation that integrates abstract syntax trees with neural semantic embeddings. The approach transforms annotated programs into typed, attributed graphs and leverages CodeBERT and SentenceTransformer to generate rich semantic embeddings. By constructing consistent graph representations across multilingual verification datasets—including C/ACSL, Java/JML, and Dafny—it achieves, for the first time, joint modeling of syntactic structure and formal semantics. This unified framework offers a viable pathway for cross-language reuse of verification artifacts and demonstrates strong generality and effectiveness across diverse programming languages and specification frameworks.
This work addresses the prevailing lack of systematic understanding of foundational formal theories in current AI compiler design, which hinders rigorous evaluation of the completeness and desirability of intermediate representations and compilation abstractions. For the first time, it systematically establishes precise correspondences between core mechanisms of MLIR—such as term rewriting systems, refinement calculi, and abstract interpretation—and classical formal theories. By grounding compiler abstractions in formal semantics, the paper clarifies the theoretical underpinnings of these constructs, articulates a precise notion of “design completeness,” and provides assessable criteria and guiding principles to navigate trade-offs between engineering pragmatism and theoretical ideals.
This work addresses the lack of client-agnostic precision metrics for abstract domains in static analysis. It proposes the MCAI framework, which introduces model counting into abstract interpretation for the first time, enabling quantitative precision evaluation independent of specific analysis tasks. By encoding concrete semantics and abstract values as logical formulas, MCAI systematically assesses widely used abstract domains such as Interval, Octagon, and KnownBit. The evaluation reveals that Interval often matches the precision of Octagon—suggesting that many octagonal constraints are redundant—and that bit-level KnownBit significantly outperforms word-level abstractions. These findings provide both empirical evidence and theoretical grounding to inform the selection of abstract domains in practice.
This work presents the first systematic investigation into the capability of large language models (LLMs) to generate program specifications involving higher-order logical constructs, which are essential for expressing complex verification properties yet remain beyond the reach of existing LLMs that predominantly handle basic syntactic forms. The authors design four syntactic configurations spanning different levels of abstraction and establish a comprehensive evaluation framework to assess a range of representative LLMs on standard verification benchmarks. Experimental results demonstrate that LLMs can effectively produce valid higher-order logical expressions; moreover, integrating logical constructs with base syntax significantly enhances verification efficacy and robustness without substantially increasing verification overhead. The study also reveals distinct advantages of two refinement paradigms in specification generation.
Static program analysis faces significant challenges in uniformly modeling stack/heap memory behaviors and value semantics across multiple programming languages, which hinders precise detection of memory safety issues such as buffer overflows and null pointer dereferences. To address this limitation, this work proposes a generic memory analysis framework grounded in abstract interpretation. The framework introduces a novel, parameterizable partitioned state abstraction mechanism that decouples value analysis from memory structure analysis, enabling flexible and modular composition of stack and heap modeling through customizable abstract domains. Formally rigorous and language-agnostic by design, the framework has been implemented to provide unified support for C/C++, Java, and Python, substantially enhancing static detection capabilities for a broad range of memory-related errors.