🤖 AI Summary
This study investigates whether general-purpose coding agents can undertake rigorous scientific performance engineering. Leveraging Codex and Claude Code, the research introduces an “executable scientific contract” mechanism to guide and validate hypothesis-driven optimization experiments for fixed-radius nearest neighbor search algorithms, autonomously refactoring a PyTorch implementation into a dependency-free, high-performance C++/CUDA library. The results demonstrate that goal-directed agents can effectively assume the role of experimental performance engineers. The generated standalone library precisely reproduces the original results while achieving a 1.6× speedup over the baseline GPU-accelerated PyTorch implementation through its synchronous NumPy interface, with consistent performance maintained across diverse hardware architectures.
📝 Abstract
Coding agents can pursue persistent objectives across many tool-use turns, but evidence that general-purpose agents can conduct rigorous scientific performance engineering remains limited. We present a repository-scale case study in which off-the-shelf Codex and Claude Code agents optimize fixed-radius nearest-neighbor (FRNN) search for particle tracking. Starting from a PyTorch-dependent CUDA implementation, the agents follow an executable goal that specifies exact-correctness tests, profiling requirements, and acceptance criteria without prescribing code transformations. In the primary sequential trajectory, they autonomously remove the PyTorch dependency and conduct hypothesis-driven optimization experiments. The resulting standalone C++/CUDA library exactly reproduces the targeted reference result. Its synchronous NumPy interface achieved 1.6-fold speedup over the original GPU-resident PyTorch interface, despite including host transfers. Similar speedups were observed across different GPU architectures and software stacks. An independent optimization rerun followed a different sequence of hypotheses and reached even better performance on the target workload. These results show that goal-persistent coding agents can act as experimental performance engineers, and that executable scientific contracts are needed both to guide and to validate their optimization.