Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the recurring errors and low efficiency of large language models in upgrading scientific codebases by proposing a certificate-driven evolutionary search framework. The method introduces an executable certificate mechanism that leverages failed candidates to generate hard constraints, preventing error recurrence through rejection, backtracking, and targeted repair logic. It further integrates formal verification, numerical comparison, physics-based validation, and GPU safety testing to ensure porting correctness. Experimental results demonstrate that the proposed framework achieves 13.78× to 23.54× function-level speedups when porting Geant4 code from CPU to GPU, improving the correctness pass rate from 55% to 90%, with certain cases even surpassing expert-written implementations.
📝 Abstract
The upgrade and rewriting of large scientific codebases has traditionally been a major challenge. While evolutionary search with large language models (LLMs) can port and accelerate legacy code, repair feedback in prompts alone does not prevent subsequent candidates from repeating the same errors. We introduce Certificate-Driven Evolutionary Search (CDES), which extends evolutionary search with enforceable restrictions derived from failed candidates, recorded as certificates of assumptions, checker evidence, and justified restrictions. Its control logic enforces these restrictions through rejection, backtracking, and targeted repair while preserving compatible edits. We apply CDES to CPU-to-GPU translation of two particle-simulation functions from the Geant4 toolkit, evaluated with a harness that goes beyond unit tests to combine formal checks, numerical comparisons, physics checks, and GPU safety tests. Generated implementations achieve 13.78x and 23.54x function-level speedups over CPU code, including data conversion and transfers; for one function, GPU throughput exceeds an expert implementation by 14.9%, reaching 16.1% when complementary components are combined. In an ablation over execution settings, certificate feedback increases the fraction of candidates passing required correctness checks from 55% to 90%.
Problem

Research questions and friction points this paper is trying to address.

scientific code porting
legacy code optimization
evolutionary search
large language models
error repetition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Certificate-Driven Evolutionary Search
LLM-based Code Porting
Self-Improving Agentic Harness
GPU Optimization
Scientific Software