🤖 AI Summary
This study addresses the high token overhead incurred by frequent environment interactions in coding agents, as well as the inherent difficulty of balancing efficiency with task success rates. To this end, we propose HERO, a framework that models multi-turn interactions through hierarchical reinforcement learning. By revealing the intrinsic relationship between efficiency variance and entropy, HERO introduces an entropy-aware policy optimization mechanism that prioritizes task resolution while promoting efficient reasoning at both the trajectory and episode levels. This design enables an end-to-end training paradigm for token efficiency without sacrificing critical information. Extensive evaluations on the SWE-bench benchmark demonstrate that HERO significantly outperforms existing methods, achieving a superior trade-off between high resolution rates and low token consumption.
📝 Abstract
Recently, coding agents have emerged as a dominant paradigm for real-world software engineering (SWE) scenarios, which solve complex tasks through multi-turn interactions with development environments. However, frequent interactions with environments would inevitably introduce substantial token overhead, leading to high usage costs and latency. Although recent studies have explored reducing token usage by context manipulation and interaction limits at inference time, these approaches focus on improving token efficiency while overlooking the risk of discarding task-relevant information, thus struggling to balance the trade-off between resolution rate and token efficiency. In this paper, we study a more general paradigm without suffering from the limitation, i.e., training token-efficient coding agents with promising resolution performance, which is a highly-practical yet less-explored problem. To this end, we reveal two core observations in SWE scenarios: i) Efficiency Variation: successful resolution could be achieved with fewer tokens; ii) Entropy Correlation: unproductive behaviors are associated with turn-level entropy. Motivated by observations, we propose a novel reinforcement learning framework, dubbed HERO. Specifically, HERO prioritizes task resolution over token efficiency during policy optimization and encourages efficient reasoning patterns at both trajectory and turn levels. Extensive experiments on SWE-bench Verified and SWE-bench Multilingual demonstrate that HERO achieves a favorable trade-off between resolution rate and token efficiency compared with state-of-the-art coding agents and reinforcement learning methods.