🤖 AI Summary
This work addresses the challenge of dynamically selecting between low-cost recovery and high-cost upgrade actions for coding agents after execution failures, under a limited budget. The authors propose a recovery routing mechanism with heterogeneous action spaces, which trains a supervised routing policy using execution trajectories and incorporates a Conformal Risk Control (CRC) layer to enable dynamic cost adjustment without retraining, thereby guaranteeing bounded marginal expected cost. Experiments across five coding benchmarks demonstrate the complementary benefits of low-cost recovery and high-cost upgrading: compared to fixed strategies, pure prompt routing, and cascaded baselines, the CRC-calibrated policy achieves significantly higher solve rates than always-upgrading approaches while incurring only 35% of the average recovery cost in the GPT-5.4-nano/GPT-5.4 setting.
📝 Abstract
Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer. Existing cost-aware systems typically treat such failures as cascade decisions: try a cheap model first, then escalate hard cases to a stronger and more expensive model. In coding, however, execution feedback can also make further cheap-model recovery worthwhile, raising a budgeted deployment question: when should an agent spend more cheap compute, and when should it escalate? We formulate this post-failure decision as recovery routing over heterogeneous actions and train a supervised router from execution rollouts. To make the same router usable under changing budgets, we add a Conformal Risk Control (CRC) layer that selects a deployment-time cost penalty without retraining and provides marginal expected-cost control under exchangeability. Across held-out failures from five coding benchmarks, cheap recovery and escalation exhibit complementary success patterns. The calibrated frontier improves over fixed actions, prompt-only routers, and a binary cascade baseline; in the main GPT-5.4-nano/GPT-5.4 setting, one CRC-calibrated frontier point exceeds always-escalate solve rate while using 35% of its mean recovery cost. Code is available at https://github.com/Qijia-He/agent-budget-control.