Score
Designs and implements deterministic compilers that translate algorithmic descriptions, computation graphs, or problem instances into solver-executable code or other deterministic computation artifacts (solver-code compilation, deterministic computation). Builds verifiers and verifiable back-translation tools that map solver outputs back to the original representation and perform deterministic checks or evaluators to prove the correctness of the compiled code and its results (deterministic verification, verifiable back-translation, verifiable reward/evaluation).
This work addresses the long-standing absence of explicit deterministic polynomial-time solvers for concrete NP-complete problems by presenting the first complete deterministic Turing machine implementation for specific problems such as SAT and Subset-Sum. Building upon an NP-verifier simulation framework, the approach extends verification mechanisms to deterministic FNP solving—without increasing the polynomial time complexity—through techniques including dynamic computation graphs, feasible graph construction, and verification-path traversal. A fully functional simulator is implemented in Python, and experimental results demonstrate that the system strictly adheres to polynomial time bounds while effectively generating valid witnesses for satisfiable instances. The source code is publicly released to ensure transparency and reproducibility of the results.
This study addresses the lack of quantitative comparison between certified compilation and fully formal verification in terms of development cost, performance impact, and practical feasibility. Leveraging the first verified compiler developed by a coding agent under human supervision, the work systematically evaluates both approaches in implementing common optimizations—such as unreachable code elimination, dead assignment elimination, and constant propagation/folding—focusing on development overhead, algorithmic choices, proof complexity, and runtime efficiency. The results show that formal verification incurs an order-of-magnitude higher development cost, leading the agent to favor less efficient algorithms and narrower optimization scopes; certificate checking becomes a performance bottleneck. Nevertheless, with modern coding agents, both methods remain practically feasible for the selected optimizations, offering the first quantitative insights into the key trade-offs imposed by verification, including increased development burden, performance compromises, and heightened supervision requirements.
This work addresses the challenge that large language model (LLM) agents often produce redundant, exploratory, and non-deterministic execution trajectories that are difficult to reuse. To overcome this, the authors propose a skill-guided framework that extracts reusable structures from noisy trajectories and compiles them into near-deterministic workflows. The core innovations include a dependency inference mechanism based on evidence tuples—establishing strong dependencies only when parameters are uniquely traceable and flagging ambiguous relations as suspect—along with fine-grained binding-type categorization. The method integrates trajectory clustering, dependency rule mining, deterministic replay, and leave-one-out validation into a unified pipeline. Experiments demonstrate high precision (0.928) and recall (0.943) in dependency identification on the T1 dataset; for Venmo tasks, API calls are reduced from 34 to 11 while passing 15 of 21 test cases, and the system correctly rejects ill-posed or irreversible intents in Spotify and Todoist scenarios.
This work addresses the challenges in Constraint Horn Clause (CHC) solving—specifically, the difficulty of modeling bit-vector and low-level semantics, and the limited expressiveness of existing CHC solvers. We propose a compositional CHC solving framework that leverages mature software verifiers (e.g., SeaHorn, Ultimate) as backend engines. The framework features a unified CHC intermediate representation, semantics-aware preprocessing supporting bit-vectors and nonlinear arithmetic, and a multi-strategy orchestration mechanism to synergize complementary strengths across tools. Its key contribution is the first systematic adaptation of industrial-grade software verification technology to CHC solving, overcoming inherent limitations of conventional SMT-based and abstract-interpretation approaches in modeling low-level program semantics. Experimental evaluation on bit-vector CHC benchmarks demonstrates significant improvements in both solving rate and efficiency, empirically validating the feasibility and effectiveness of using general-purpose software verifiers as CHC backends.
Formal verification of compilers incurs high maintenance costs, especially when modifications necessitate extensive re-verification. Method: This paper introduces the first trusted rewriting engine framework for Coq, modeling compilers as collections of algebraic rewrite rules—each independently verifiable. It employs theorem-driven modeling, metaprogramming-based automated synthesis, and proof reuse to enable rule-level formal verification and automatic composition. Contribution/Results: The framework decouples rule verification from compiler construction, significantly reducing verification and maintenance overhead. Evaluated in the Fiat Cryptography toolchain, the generated command-line compiler achieves approximately 1000× speedup over prior verified counterparts. Moreover, its proofs are more concise and exhibit substantially higher reusability across compiler transformations.
This work systematically evaluates the potential of large language models (LLMs) for automatic code optimization in high-performance computing (HPC), where traditional approaches often struggle to balance performance and correctness. The study introduces a novel methodology that leverages multi-level abstractions and goal-oriented prompting to guide LLMs in directly generating optimized C code. Evaluated on the PolyBench benchmark suite, this approach is compared against conventional auto-tuning frameworks that rely on schedule representations. Experimental results demonstrate that LLM-generated C code achieves superior performance and effectiveness, highlighting the critical influence of compiler optimization abstractions on LLM guidance. These findings establish a promising new direction toward verifiable, LLM-driven code optimization for HPC applications.
Traditional program analysis relies on control flow graphs and separate forward or backward data-flow analyses, resulting in complex structures that are difficult to formally verify. This work proposes a forward-directed formal method based on prophecy and history variables, tightly integrating program analysis into operational semantics and establishing correctness and optimality of transformations via subset-inclusion constraints. We present the first machine-verified framework supporting prophecy and history variables, eliminating the need for explicit control flow graphs, abstraction/concretization functions, and Galois connections. By extending the operational semantics of a domain-specific language and employing forward simulation, we implement formally verified program transformations within the Nexis compiler. This approach yields the first machine-checked proofs of both correctness and optimality for dead code elimination and lazy code motion optimizations.
This work addresses the challenge of reconciling program determinism with pipeline optimization in asynchronous dataflow compilation by introducing Wavelet, a framework that achieves the first end-to-end formal verification for such compilers. Wavelet integrates a capability-based type system augmented with memory fences and the Lean theorem prover to provide modular, semantics-preserving proofs for core compiler transformations, guaranteeing that the generated dataflow graphs satisfy both forward simulation and determinism. Experimental results demonstrate that code produced by Wavelet matches the size of that generated by the unverified RipTide compiler while offering rigorous correctness guarantees.
This work addresses the challenge of efficiently eliminating redundant logic in AI-generated code, particularly under constrained verification budgets. The authors formulate redundancy removal as a candidate statement ranking and scheduling task, proposing to control deletion order rather than relying on model confidence scores. They introduce a two-stage hybrid scheduling mechanism that combines a static shortest-first rule with a neural ranking model, prioritizing high-potential deletion candidates within limited verification resources while ensuring behavioral equivalence through execution-based validation. Experimental results on the MBPP benchmark demonstrate a 9.5% improvement in verified deletion coverage and six additional successfully solved tasks. Notably, even without in-domain validation, the static prefix strategy alone guarantees non-degrading performance.