🤖 AI Summary
This study addresses the limitations of existing agent evaluations in distinguishing genuine capability abstraction from code leakage, alongside state management and cross-domain abstraction bottlenecks in long-horizon software engineering. We propose a benchmark grounded in a "no-solution overlap" axiom, positioning static skill libraries as execution guides rather than shortcuts. The evaluation integrates LLM-simulated interactions, 48-hour multi-trajectory execution, and human expert verification. Our findings demonstrate that while skill abstraction cannot substitute for fundamental reasoning, it effectively mitigates context bloat and reduces overall coding time by over 55%. This work thereby advances a paradigm shift in intelligent agents from superficial pattern matching toward deep capability transfer.
📝 Abstract
While large language models have achieved remarkable success in isolated code generation, authentic software engineering requires sustained reasoning, complex state management, and continuous cross-domain abstraction. However, current evaluations of skill evolution in autonomous agents suffer from a critical identifiability problem: they structurally confound genuine capability abstraction with rote solution leakage (i.e., copying highly similar code from historical training data). To resolve this, we introduce EngramBench, a rigorous, capability-grounded benchmark governed by the strict axiom of capability overlap without solution overlap. Comprising 30 diverse learning tasks and 13 unseen transfer tasks, EngramBench challenges agents to navigate interactive, multi-hour development cycles driven by LLM-simulated users. Our extensive evaluation across 48 multi-hour execution trajectories -- corroborated by human-expert validation -- reveals a profound insight into procedural memory. We demonstrate that static skill banks do not magically bypass the "last mile" of exact code implementation, which remains bottlenecked by the base model's inherent reasoning limits. However, they serve as an indispensable execution compass. By navigating agents away from catastrophic, token-heavy trial-and-error, genuine capability abstraction slashes redundant context bloat and reduces overall coding time by over 55%. Ultimately, EngramBench shifts the evaluation paradigm from trivial pattern matching to the verifiable measurement of deep, cross-domain capability transfer.