🤖 AI Summary
This study addresses the challenge that long-horizon software engineering agents struggle to distinguish effective actions from redundant exploration when relying solely on terminal supervision. To overcome this, we propose an asynchronous milestone credit assignment framework that leverages workflow runtime signals to generate fine-grained process supervision without auxiliary reward models or external evaluators. Core innovations include a novel differential attribution mechanism based on navigation and verification potentials, alongside an asynchronous shadow probing technique that conceals verification latency. These components are integrated with reinforcement learning and sandbox parallel replay to achieve efficient credit assignment. Experiments demonstrate that our approach significantly enhances agent performance on representative long-horizon tasks, validating the effectiveness of runtime signals as a source of process supervision.
📝 Abstract
Long-horizon software engineering (SWE) agents trained with reinforcement learning with verifiable rewards (RLVR) typically receive only terminal outcome supervision, making it difficult to distinguish productive actions from redundant exploration or functional regressions. We propose SWE-MILE, an asynchronous potential-induced milestone credit assignment framework that derives fine-grained process supervision from workflow runtime, without auxiliary reward models or external evaluators. SWE-MILE quantifies task-relevant file exposure and test-state alignment as navigation and verification potentials, respectively. Differences in these potentials attribute milestone progress and regressions to individual actions, while discounted backward credit propagates supervision to preceding steps. To efficiently acquire intermediate verification states, SWE-MILE further introduces asynchronous shadow probing, which replays repository-changing actions in an isolated sandbox and runs verification in parallel with the agent's primary interaction, largely hiding verification latency. The resulting process credit augments terminal outcome advantages and provides informative learning signals. Experiments on two representative long-horizon SWE tasks demonstrate substantial improvements in agent performance, highlighting workflow runtime signals as a practical source of process supervision for long-horizon SWE agents.