Score
Designs and implements faithful ports of algorithms and software components between programming languages, including the associated cross-language tests and benchmarks; validates outputs against reference implementations and detects computational, heuristic, or idiomatic discrepancies to ensure equivalent behavior.
This work proposes a novel paradigm that bridges the long-standing divide between testing and formal verification in traditional software validation, enabling them to synergistically enhance both efficiency and quality. Grounded in Design by Contract, the approach leverages the counterexample generation capability of SMT solvers to transform formal verification tools into an integrated engine for automated testing and repair. Within a unified framework, the method simultaneously achieves three key objectives: automatic generation of test cases for faulty programs, construction of regression test suites with full coverage for correct programs, and correctness-guaranteed program repair. This represents the first integration of verification, testing, and repair into a single cohesive methodology.
This work evaluates large language models’ (LLMs) capability in code functional equivalence detection—a critical task for assessing their performance on semantics-preserving code transformations such as rewriting and translation. To this end, we introduce CETBench, the first controllable and scalable benchmark for this task. CETBench employs a program-transformation-based paradigm: leveraging static analysis and predefined semantics-preserving (e.g., variable renaming, control-flow restructuring) or semantics-breaking transformations, it systematically generates high-quality equivalent and non-equivalent code pairs. Experiments reveal that state-of-the-art LLMs exhibit high sensitivity to subtle semantics-preserving modifications, exposing fundamental limitations in deep semantic code understanding; their accuracy drops significantly on transformed samples. We further propose a lightweight supervised fine-tuning approach that substantially improves discrimination accuracy across diverse models, demonstrating both generalizability and practical utility.
This work addresses the significant disparity in verifiability among semantically equivalent yet structurally diverse programs, a key bottleneck in generating high-assurance software. The authors propose Diversify2Verify, a novel approach that leverages large language models to synthesize diverse recursive and imperative implementations of the same task, integrates the Why3 platform for automatic contract inference and formal verification, and introduces a verifier-guided annotation repair mechanism to enhance verifiability. This study is the first to systematically expose the verifiability gap across equivalent program variants and establishes a new paradigm wherein implementation diversity drives improved verification success. Evaluated on a benchmark of 73 tasks, the method yields 154 verifiable programs after two rounds of repair, with at least one successfully verified variant for 67.1% of the tasks—substantially outperforming baseline approaches.
Source-to-source transformations—such as loop unrolling, tiling, and fusion—in compiler optimization lack reliable correctness verification, posing significant risks in high-level synthesis (HLS) and production compilers. Method: This paper introduces the first equivalence checking framework for compiler optimizations based on e-graphs and equality saturation. It deeply integrates MLIR into the e-graph infrastructure and devises a hybrid dynamic/static rewriting strategy to achieve comprehensive functional equivalence verification. Contribution/Results: The framework successfully verifies over 100 optimization combinations on PolyBenchC; verifies >100K lines of MLIR code within 40 minutes, exhibiting linear and predictable runtime scaling; and discovers and localizes two classes of semantics-breaking bugs in mlir-opt for the first time. By bridging a critical gap in automated equivalence verification for HLS and compiler optimizations, this work establishes a new paradigm for trustworthy compiler optimization.
Automated verification of interactive console I/O programs in Haskell education remains challenging due to the dynamic, history-dependent nature of student implementations. Method: We propose a lightweight, formal behavioral specification language that uniquely integrates global state and execution history, expressed via regex-like syntax; its trace-based semantics enable probabilistic testing and scalable verification through *sampleable validity*. Contribution/Results: Our system automatically validates student submissions against behavioral specifications and supports pedagogical closed-loop applications—including real-time feedback generation, example solution synthesis, and exercise randomization. Empirical evaluation demonstrates substantial improvements in test coverage and pedagogical adaptability while preserving formal rigor. To our knowledge, this is the first framework for verifying interactive behaviors in functional programming education that simultaneously achieves theoretical soundness and practical deployability.
Existing compiler testing techniques are often ill-suited for transpilers, as they typically lack multiple equivalent implementations and may produce non-executable output code. This work introduces metamorphic testing to transpiler validation by proposing the notion of “mutation consistency”: it defines metamorphic relations at the source-code level to verify whether structurally consistent and expected changes manifest in the generated code when the input DSL program undergoes semantics-preserving mutations. This approach enables defect detection without requiring execution of the generated code. We develop a mutation-based modeling method for metamorphic relations, a source-level structural consistency analysis mechanism, and implement an automated tool, MCP-Tester. Evaluated on real-world technology migration cases, our method effectively uncovers transpiler bugs that elude conventional fuzzing approaches.
Traditional program equivalence checking offers only binary judgments, failing to characterize the scope and conditions under which patches affect program behavior. This work proposes a quantitative partial equivalence analysis method that integrates symbolic execution with a numerical-domain-optimized range-search heuristic to precisely identify regions in the input space where original and patched programs exhibit consistent or divergent behaviors, and to quantify the degree of their differences. By elevating patch impact analysis from qualitative to quantitative, the approach provides reliable lower-bound estimates for equivalence. Experimental evaluation on 90 CVE patches and the Juliet test suite demonstrates its effectiveness, and within EqBench, it successfully uncovered five C program pairs erroneously labeled as equivalent, accurately pinpointing the conditions causing behavioral divergence.
This work addresses the limited capability of large language models (LLMs) in automatically translating programs into formal specifications suitable for model checking, a key bottleneck in their application to software verification. To systematically evaluate and advance this capability, the authors introduce Model-Bench, the first benchmark specifically designed for program-to-formal-model translation. Constructed from HumanEval, MBPP, and LiveCodeBench, Model-Bench comprises 400 Python programs paired with reference formal specifications. Leveraging a pipeline that integrates LLMs, program modeling, and model-checking techniques, the study assesses the ability of current models to generate verifiable specifications. Experimental results uncover critical limitations in existing LLMs for this task and provide clear directions for future improvements in bridging the gap between natural-language-driven code generation and formal verification.
This work proposes a hybrid concrete-symbolic interpretation method to efficiently verify semantic equivalence between original and optimized programs in MLIR, ensuring the correctness of optimization transformations. The approach supports diverse syntactic, scheduling, and memory representations and theoretically achieves linear-time complexity for equivalence checking. Building upon this method, the authors develop a formal verifier for a subset of MLIR and successfully apply it to the AMD MLIR-AIR and MLIR-AIE toolchains as well as the standard mlir-opt infrastructure. Evaluation across hundreds of benchmark variants demonstrates the verifier’s effectiveness in validating optimization pipelines, significantly enhancing the reliability of compiler optimizations within the MLIR ecosystem.