Sibyl: An Efficient Small-large Model Collaboration Framework for Long-horizon Tasks

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited reasoning capabilities of small-model agents in long-horizon tasks and the difficulty of balancing cloud invocation costs against success rates. To this end, we propose Sibyl, a framework that enables efficient collaboration between large and small language models through a three-stage training paradigm. First, self-evolving reinforcement learning optimizes on-demand consultation decisions. Second, deterministic disagreement state mining identifies critical moments for seeking assistance. Finally, consultation-aware reinforcement learning internalizes large-model guidance into the small model. Experimental results demonstrate that a model with merely 0.6B parameters outperforms baselines by 95.2% and 80.4% in success rate on ALFWorld and WebShop, respectively, while requiring only 0.8 to 3.9 cloud invocations on average.
📝 Abstract
Small language models (SLMs) offer a promising foundation for on-device agents through low-latency, resource-efficient inference, yet limited reasoning and planning capabilities constrain their performance on long-horizon tasks requiring multi-step interaction with the environment. Step-level collaboration between SLMs and larger cloud-hosted models can bridge this gap, but identifying states that warrant cloud assistance remains challenging: the contribution of each cloud call is entangled with subsequent actions and can be assessed only from the final task outcome. Compounding this challenge, the SLM must balance two competing objectives: maximizing task success and minimizing cloud calls. To address this, we propose Sibyl, an algorithm that trains SLM agents to selectively consult cloud models at the step level and internalize their guidance for subsequent decisions, achieving strong task performance with minimal cloud reliance. Sibyl follows a three-stage training pipeline that (1) builds a robust base policy through consultation-free self-evolving reinforcement learning (RL); (2) cold-starts consultation behavior via decisive-disagreement state mining; and (3) jointly optimizes consultation decisions and guidance internalization through consultation-aware RL. Experiments on ALFWorld and WebShop demonstrate that Sibyl, using only a 0.6B-parameter model, outperforms state-of-the-art baselines, including agent training and routing methods, by 95.2% and 80.4% in success rate while averaging only 0.8 and 3.9 cloud calls per trajectory, respectively.
Problem

Research questions and friction points this paper is trying to address.

Small Language Models
Long-horizon Tasks
Model Collaboration
Cloud Consultation
Multi-step Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Small-Large Model Collaboration
Reinforcement Learning
Long-horizon Tasks
Selective Consultation
Guidance Internalization
🔎 Similar Papers
Z
Zhewei Fang
Hangzhou Dianzi University
Yuxin Zhang
Yuxin Zhang
Fudan University
Distributed Machine LearningEdge AISatellite Internet
Z
Zhenwei Shao
Hangzhou Dianzi University
M
Mengze Li
The Hong Kong University of Science and Technology
Z
Zheng Lin
University of Luxembourg
L
Long Chen
Simon Fraser University
Z
Zhou Yu
Hangzhou Dianzi University
Zhe Chen
Zhe Chen
Fudan University
Satellite InternetEdge AIWireless Sensing
Z
Zhiwen Chen
Alibaba Group
Zhaode Wang
Zhaode Wang
Alibaba
C
Chengfei Lv
Alibaba Group