🤖 AI Summary
This study addresses the low collaboration efficiency and redundant context transmission between small language models (SLMs) and large language models (LLMs) under black-box API access. We reformulate this collaboration as an information acquisition process subject to budget constraints and propose a three-stage Reinforcement Learning with Verifiable Rewards (RLVR) framework. In this framework, the SLM leads the reasoning process while selectively issuing targeted queries to the LLM, jointly optimizing invocation decisions, query construction, and information integration strategies. Experimental results demonstrate that our approach significantly improves the performance-cost trade-off, surpassing the oracle routing upper bound in certain scenarios while exhibiting strong cross-model transferability.
📝 Abstract
Collaboration between a small language model (SLM) and a large language model (LLM) offers an opportunity to combine the efficiency of smaller models with the strong reasoning capabilities of larger ones. Existing approaches primarily frame such collaboration as a computation allocation problem, determining which model should handle each portion of the reasoning process. In black-box API-based settings, however, this paradigm can be inefficient due to coarse-grained delegation or repeated transmission of context across model switches. In this work, we instead formulate SLM-LLM collaboration as an information acquisition problem, under an API budget constraint. The SLM remains the primary reasoner and selectively queries a black-box LLM advisor only when needed, issuing targeted queries rather than delegating the reasoning process itself. To realize this strategy, we develop a three-stage RLVR framework that learns whether to call the advisor, how to formulate useful queries, and how to integrate the collaboration into the reasoning process by jointly refining advisor invocation and information use. Across mathematical reasoning and coding tasks, our approach improves the performance--cost tradeoff over existing collaboration baselines and, in some settings, matches or exceeds oracle problem-level routing. Finally, we show that our strategy can transfer to other advisor model families, without further training.