🤖 AI Summary
This study addresses the limited reasoning capabilities of small-model agents in long-horizon tasks and the difficulty of balancing cloud invocation costs against success rates. To this end, we propose Sibyl, a framework that enables efficient collaboration between large and small language models through a three-stage training paradigm. First, self-evolving reinforcement learning optimizes on-demand consultation decisions. Second, deterministic disagreement state mining identifies critical moments for seeking assistance. Finally, consultation-aware reinforcement learning internalizes large-model guidance into the small model. Experimental results demonstrate that a model with merely 0.6B parameters outperforms baselines by 95.2% and 80.4% in success rate on ALFWorld and WebShop, respectively, while requiring only 0.8 to 3.9 cloud invocations on average.
📝 Abstract
Small language models (SLMs) offer a promising foundation for on-device agents through low-latency, resource-efficient inference, yet limited reasoning and planning capabilities constrain their performance on long-horizon tasks requiring multi-step interaction with the environment. Step-level collaboration between SLMs and larger cloud-hosted models can bridge this gap, but identifying states that warrant cloud assistance remains challenging: the contribution of each cloud call is entangled with subsequent actions and can be assessed only from the final task outcome. Compounding this challenge, the SLM must balance two competing objectives: maximizing task success and minimizing cloud calls. To address this, we propose Sibyl, an algorithm that trains SLM agents to selectively consult cloud models at the step level and internalize their guidance for subsequent decisions, achieving strong task performance with minimal cloud reliance. Sibyl follows a three-stage training pipeline that (1) builds a robust base policy through consultation-free self-evolving reinforcement learning (RL); (2) cold-starts consultation behavior via decisive-disagreement state mining; and (3) jointly optimizes consultation decisions and guidance internalization through consultation-aware RL. Experiments on ALFWorld and WebShop demonstrate that Sibyl, using only a 0.6B-parameter model, outperforms state-of-the-art baselines, including agent training and routing methods, by 95.2% and 80.4% in success rate while averaging only 0.8 and 3.9 cloud calls per trajectory, respectively.