Measured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM Serving

📅 2026-09-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过设计基于语言模型的路由器,从异构LLM池中选择最合适的模型来处理查询,以减少能源消耗同时保持任务性能。
📝 Abstract
Large language models (LLMs) and agentic AI systems are creating rapidly growing inference energy demands as model sizes grow and reasoning trajectories extend. While in practice, many queries do not require the capabilities of the largest available model, and routinely directing such queries to a high-capability model can introduce unnecessary, considerable computation and energy consumption. In this paper, we investigate whether adaptive routing across a heterogeneous pool of LLMs can reduce this energy burden without substantially compromising task performance. We design a language-model-based router that reads in each query and selects an answer model from a fixed candidate pool. The candidate models are first profiled through an offline tournament that records their correctness, latency, power, and GPU energy for each query. Using these measurements, the router is trained through supervised fine-tuning followed by group relative policy optimization (GRPO) with the tailored paradigms. Results demonstrate that learned routing can selectively allocate expensive model capacity based on query context and improve the accuracy-energy tradeoff in multi-LLM serving. Across seven benchmark tasks, we also observe a sharp accuracy-energy phase transition among routers, providing practical insights into improving energy efficiency while maintaining LLM performance.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Energy Efficiency
Inference Energy Demands
Adaptive Routing
Model Capacity
Innovation

Methods, ideas, or system contributions that make the work stand out.

adaptive routing
language-model-based router
energy efficiency
supervised fine-tuning
group relative policy optimization (GRPO)
🔎 Similar Papers
No similar papers found.