π€ AI Summary
This study addresses the challenge of automatically formalizing C code, where reliance on a single success-rate metric conflates translation rejections with silent errorsβthe latter being more hazardous yet difficult to detect. To overcome this, we construct a deterministic exporter based on Code Property Graphs (CPGs) that maps C source code to Lean 4 semantics. We further introduce a methodological protocol featuring a three-tier correctness discipline and the separation of unparseable constructs, integrated with static analysis and a trust ledger mechanism to precisely identify and flag silent errors. Evaluated on the SQLite source code, our approach achieves a 60.7% hole-free translation rate, reveals a 2.4Γ hidden performance gap obscured by conventional metrics, and effectively uncovers potential silent error defects.
π Abstract
Verifying large C codebases requires translating them into formally checkable semantics, but most autoformalization work reports a single aggregate success rate that conflates two different failure modes: code the translator declined to handle and code the translator translated incorrectly. We present a deterministic code-property-graph (CPG)-based exporter that translates C source into a small Lean~4 core semantics under a three-tier correctness discipline: hole-free, call-closed, and dynamic-hole-risk, that keeps these modes distinct. Applied to a re-export of the SQLite source tree (8{,}602 functions, over 1.5~million AST nodes), the exporter translates 5{,}222 functions (60.7\%) hole-free, of which only 2{,}134 (24.8\%) are call-closed, a $2.4\times$ gap a single rate would hide. Both figures are reported in a full per-construct trust ledger rather than a single score. Our main contribution is methodological: a protocol for separating constructs that are fundamentally unresolvable by whole-program static analysis (such as public API boundaries) from constructs that only look that way. For example, we retracted our own ``impossible'' classification of function-pointer/vtable dispatch after tracing a concrete counterexample in the target codebase. Separately, a search for ``more holes closed'' surfaced a latent silent-wrong-answer bug: a translation that succeeded with an incorrect result rather than declining, a failure mode we argue is more dangerous than any hole, and one a hole-count-only evaluation would never surface.