RSI-Router: Evolving Subtask-Level LLM Routing and Skills for Cost-Efficient Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of balancing task performance and inference cost for LLM-based agents in long-horizon tasks by proposing a recursive self-improvement framework that enables efficient collaboration between large and small models. The proposed method comprises four phases: subtask mining, routing evolution, skill evolution, and Pareto selection. Through multi-objective optimization, it establishes a closed loop integrating subtask-level model routing with automated skill generation. Experimental results across five benchmarks demonstrate that the proposed approach surpasses existing baselines while requiring approximately 50% of the inference cost. Furthermore, it significantly achieves a superior Pareto frontier, yielding simultaneous improvements in both cost-efficiency and task performance.
📝 Abstract
Practical deployment of large language model (LLM) agents requires strong task performance at affordable inference cost. For long-horizon agentic tasks, this performance-cost trade-off can be improved through within-task large-small model collaboration, as smaller models can handle some stages even when they cannot solve the full task. In this paper, we introduce RSI-router, a routing framework that constructs subtask-level model assignments and model-specific skills through recursive self-improvement over accumulated experience. Each iteration consists of four stages: Subtask Mining derives subtask definitions and identification rules from training trajectories; Routing Strategy Evolution proposes and evaluates diverse model assignments; Model-Specific Skill Evolution compares routed and large-model-only trajectories to diagnose failures and develop reusable execution skills; and Pareto-Optimal Router Selection updates the Pareto population using historical and newly generated routers while retaining dominated routers as experience for subsequent evolution. Routing between DeepSeek-V4.1-Flash and Qwen3.5-9B, RSI-router consistently surpasses the DeepSeek-only baseline at roughly half the inference cost (48.3%) across five agentic benchmarks. In particular, on ALFWorld, ScienceWorld, and WebShop, it cuts inference cost by 74.7-82.2% while simultaneously improving performance; on Terminal-Bench 2.0, it achieves a 16.7% relative performance gain at 18.0% lower cost. Moreover, RSI-router establishes a stronger performance--cost Pareto frontier than 9 routing methods.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
inference cost
model routing
subtask-level collaboration
performance-cost trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Routing
Recursive Self-Improvement
Subtask Mining
Pareto Optimization
Cost-Efficient Agents