🤖 AI Summary
This study addresses the challenge faced by LLM agents in adaptively selecting the optimal execution granularity when atomic tools and composite skills coexist. To this end, we propose CIPO, a novel framework that formulates skill invocation as an adaptive granularity decision problem for the first time. Specifically, CIPO constructs an executable skill library through budget-constrained trajectory mining and employs counterfactual imagination to generate both atomic and skill-level branches from identical states. The divergence between these branches serves as a supplementary reward signal to optimize the granularity decision policy via reinforcement learning. Extensive experiments across multiple benchmarks demonstrate that CIPO significantly improves task success rates and decision efficiency, enabling precise, state-dependent granularity switching.
📝 Abstract
Large language model (LLM) agents solve complex tasks through multi-step interactions with external tools. These interactions often contain recurring local tool sequences. Treating such sequences as composite"Skills"can shorten tool-use trajectories and reduce repeated low-level decisions. However, when atomic tools and composite skills coexist, skill use becomes a policy problem: the agent must decide whether the current state requires atomic fine control or skill-level abstraction. In this paper, we argue that effective skill use should be studied as adaptive tool granularity selection. The most direct training signal for this problem is to compare the consequences of atomic and skill choices available from the same state. Based on this view, we propose CIPO, a Counterfactual Imagination Policy Optimization framework for adaptive tool granularity. CIPO constructs executable skills through budget-constrained mining of successful tool-use trajectories and trains granularity decisions with counterfactual branch rollouts. For each base rollout, CIPO branches at the first eligible granularity decision and replaces the chosen action with a feasible atomic or skill alternative. The paired outcome difference serves as a supplementary reward for policy optimization. Experiments across multiple benchmarks and model backbones show that CIPO improves task success and decision efficiency over baselines. Further analyses show that CIPO learns effective skill use by improving the choice between atomic tools and composite skills based on the current state, without simply increasing skill frequency.