🤖 AI Summary
This study investigates whether code agents can repair failing software while preserving underlying logic and rebuild industrial engines from open specifications, addressing the measurement error challenges inherent in executable verification. To this end, this work proposes ReviveBench, a benchmark that establishes an evaluation framework incorporating hidden validators, native environment calibration, identifier obfuscation, and consistency analysis. Furthermore, contamination control experiments and auditing mechanisms are introduced to expose measurement errors, alongside three checking strategies designed to calibrate evaluation standards. Empirical results demonstrate that the strongest models achieve full success on revival tasks, while several models meet reconstruction metrics. Additionally, the proposed framework identifies and rectifies 28 defects within existing validators, thereby enhancing the reliability of automated code generation evaluation.
📝 Abstract
Can coding agents restore software that no longer runs while preserving its underlying methods, and reconstruct industrial software engines from open specifications? Here we introduce ReviveBench, a benchmark with two task families evaluated by hidden verifiers calibrated against native execution environments, established engineering tools, or purpose-built reference implementations. The revival family comprises ten tasks involving dependency incompatibilities, deleted core modules, legacy builds, and a GPU-based foundation model. Every starting workspace fails verification, and the strongest evaluated model passes all ten tasks in at least one run each. In contamination-control experiments, identifier obfuscation reduces line similarity to the original implementations from 0.51--0.96 to 0.03--0.44 without reducing the observed pass rate of any evaluated model. On repositories created after the stated knowledge cutoffs, the strongest model passes eight of nine runs. The reconstruction family comprises thirteen tasks spanning numerical, geometric, hardware, and transactional systems (e.g. CAD and CRM). Two models meet the benchmark's pass criteria on all thirteen, although our audit shows that the CFD task cannot establish numerical-solver capability. Benchmark construction and auditing uncover 28 verifier defects, including 24 false negatives and two false positives. These findings show that executable verification can itself introduce substantial measurement error. We present three practical checks: test whether prescribed methods can reach the grading thresholds, investigate agreement among independently generated candidates, and recompute diagnostics from submitted artifacts. ReviveBench thus provides both an evaluation of software revival and engine reconstruction, and cases in validating the verifiers used to measure coding agents for software design.