Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the decision-making and learning efficiency bottlenecks faced by language agents in environments requiring long sequences of low-level actions. We propose a code abstraction-based approach to constructing reusable skill libraries, which achieves hierarchical action control by integrating semantic skills with primitive actions to mitigate abstraction leakage. Using NetHack as a benchmark, we develop the CodeHack skill library and validate it through zero-shot prompting, supervised fine-tuning, and reinforcement learning. Experimental results demonstrate that our method nearly triples game progress while reducing inference costs by 86%. Furthermore, reinforcement learning training yields a 7.2-fold improvement in dungeon exploration depth. These findings indicate that the proposed framework effectively balances productivity with flexibility for language-driven agents in complex environments.
📝 Abstract
Language agents struggle to act and learn in environments that require long sequences of low-level actions. Code-based abstractions can make these agents more productive by letting them invoke reusable skills instead of repeatedly selecting individual actions. The code handles recurring local decisions, while the language model decides which skills to use and how to combine them. Yet abstractions are leaky, and situations beyond a skill's capabilities may require a return to primitive actions. Motivated by this tradeoff between productivity and flexibility, we systematically study how code-based action abstraction affects the performance, inference cost, and learning of language agents. We study this in NetHack, a challenging, long-horizon game environment, using CodeHack, our library of code-based skills with natural-language descriptions. We use this library to compare agents restricted to primitives with those using semantic skills alone or in combination with primitives. We evaluate these agents in three settings: zero-shot prompting, supervised fine-tuning, and reinforcement learning. Across a broad zero-shot evaluation on NetHack, we find that compared with primitives, skills nearly triple game progression, while reducing inference cost per episode by 86%. Combining skills with primitives retains much of this benefit while preserving a path back down to low-level actions. Finally, in RL, we find that skill-based agents learn significantly faster than agents acting on primitives, achieving a 7.2x larger average gain in dungeon level over the same training budget. These results show that a supplied skill library can improve performance, efficiency, and learning, while retaining primitives provides flexibility when the library is insufficient. We release CodeHack together with training and evaluation code.
Problem

Research questions and friction points this paper is trying to address.

Language Agents
Code-based Abstraction
Action Abstraction
Long-horizon Tasks
NetHack
Innovation

Methods, ideas, or system contributions that make the work stand out.

Code-based Abstraction
Language Agents
Skill Library
Reinforcement Learning
NetHack
🔎 Similar Papers
No similar papers found.