Functionally Equivalent or Not? Graph-Grounded Differential Surrogate Execution for Code Equivalence

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of insufficient test coverage, misleading textual similarity, and non-executable environments in code equivalence determination by proposing the FEAgent framework. This approach integrates typed program graph alignment with differential agent execution, employing branch-aware input generation and double-blind LLM prediction of observable behaviors to produce an evidence ledger with explicit uncertainty, thereby enabling auditable equivalence judgments that effectively bridge the gap between testing and formal verification. Evaluated on EquiBench and SWE-bench, the framework successfully identifies 216 mislabeled benchmark pairs and 94 defective patches, revealing behavioral divergences overlooked by existing unit tests.
📝 Abstract
Determining whether two programs are functionally equivalent is central to code modernization, patch validation, refactoring, and code-generation evaluation. Yet the usual signals are incomplete: tests cover only finite inputs, textual similarity confuses implementation with behavior, and unconstrained LLM judgments are difficult to audit. Direct execution is often impossible when a program depends on an obsolete, licensed, unavailable, or unsafe environment. We introduce FEAgent, a selective equivalence assessor agent that combines typed program-graph evidence with differential surrogate execution. FEAgent first aligns public interfaces and behaviorally relevant graph anchors, then issues bounded queries over call-flow, control-flow, data-flow, type, import, and effect relations. Next, a branch-aware generator agent proposes discriminating inputs, and two blinded LLM surrogates independently predict source and target observables. Every claim and predicted divergence is recorded in an evidence ledger. A deterministic reconciler then returns EQUIVALENT, INEQUIVALENT, or UNCLEAR rather than forcing a verdict when paths are uncovered or evidence conflicts. We evaluate FEAgent on function-level equivalence and repository-level bug patches, where the existing oracle is a benchmark label or a passing test suite. Every disagreement with that oracle is adjudicated by direct execution, revealing errors in benchmark labels and behavioral divergences missed by unit-test-only scoring. On EquiBench, execution confirms FEAgent's disagreements with published labels on 216 of 1,200 evaluated pairs (18.0%); on SWE-bench Verified, 94 of 331 test-passing agent patches (28.4%) diverge from the reference patch. FEAgent thus serves as an audit layer between testing and formal verification, keeping its evidence reviewable and its uncertainty explicit without claiming a proof of equivalence.
Problem

Research questions and friction points this paper is trying to address.

code equivalence
functional equivalence
program analysis
patch validation
surrogate execution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Code Equivalence
Differential Surrogate Execution
Program Graph
Evidence Ledger
FEAgent
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Amit Kachroo
AWS AI Labs, Santa Clara, California, USA
Like Hui
Like Hui
AWS AI Labs, Santa Clara, California, USA
Haitao Mao
Haitao Mao
Unknown affiliation
Y
Yuhao Zhang
AWS AI Labs, New York, USA
N
Nguyen Vo
AWS AI Labs, Santa Clara, California, USA