🤖 AI Summary
This work addresses the low success rate in automatic migration from PyTorch to JAX, primarily caused by API discrepancies and dynamic execution behaviors. The authors propose an autonomous agent system that integrates in-context learning (ICL) with an execution oracle. By executing the original PyTorch module to capture ground-truth tensor states as immutable references, the system automatically generates test cases to iteratively guide a large language model in refining the translated JAX code. This approach presents the first closed-loop integration of an execution oracle with ICL-driven self-debugging, achieving high numerical equivalence and reliability across frameworks while maintaining low computational overhead. Evaluated on prominent models including SAM, T5, and CodeWhisperer, the method attains a 91% numerical equivalence rate—substantially outperforming both baseline approaches (9%) and instruction-based self-debugging strategies (27%).
📝 Abstract
Translating deep learning models from PyTorch's flexible, object-oriented design to JAX's functional, stateless setup is usually a manual and error-prone task. Automated migration is challenging because Large Language Models (LLMs) struggle with strict and dynamic API alignment and are prone to mistakes for exacting operations. We propose a fully autonomous system that combines In-Context Learning (ICL) with oracle-driven self-debugging. First, we curated an ICL context that serves as a strict reference for idiomatic JAX styling and test case generation. Second, instead of depending on the LLM to deduce mathematical outputs, we run the source PyTorch modules to get their actual dynamic tensor states. This creates an unchangeable execution oracle. We then use an autonomous agentic loop to synthesize tests based on the oracle data. The test cases are executed repeatedly, and the traceback is sent back to the LLM for self-correction. Ablations show that combining ICL references with oracle grounding and self-debugging greatly outperforms pure instructional and basic agentic baselines. This improvement does not add an excessive computational overhead. Our lightweight pipeline achieves 91% numerical equivalence (compared to baseline: 9%, instruction + self-debugging: 27%) on neural modules, providing a highly reliable, scalable blueprint for cross-framework migration. This has been validated across several state-of-the-art models including SAM (segment anything), T5, Code Whisper amongst others showing high numerical equivalency. Code: https://github.com/AI-Hypercomputer/accelerator-agents/tree/main/MaxCode