🤖 AI Summary
This study addresses the limitations of local optimization in existing inference systems by introducing an agent self-evolution mechanism for the end-to-end optimization of complex reasoning engines. Specifically, we construct a micro-SGLang engine in which autonomous agents iteratively evolve both code and architecture, inheriting historical experience to achieve continuous improvement without human intervention. Experimental results demonstrate that this approach yields a 3.27× throughput improvement for the Qwen3-0.6B model, surpassing mainstream state-of-the-art engines such as vLLM while strictly preserving numerical precision and task correctness. This work establishes a novel paradigm for the automated global acceleration of inference systems.
📝 Abstract
Inference systems determine how fast and how cheaply language models can be served, so making them faster has direct practical value. However, prior work focuses mostly on optimizing certain parts such as kernels or memory within the large system. In this work, we take a holistic approach and apply agentic self-evolution to optimize the whole system end-to-end. Our SEIS (Self-Evolving Inference Systems) autonomously optimizes the entire mini-sglang engine without human intervention through iterative sessions with inherited experiences and code changes. Serving Qwen3-0.6B on H100, the resulting engine reaches 3.27X the throughput of the original mini-sglang implementation and beats SOTA engines like vLLM, TensorRT-LLM, and SGLang in the single-request workload. The correctness of the optimized inference engine by SEIS is tested in terms of numerical difference and downstream accuracy on math and long-context retrieval tasks. The code and session histories show that the speedup comes from redesigning the whole engine and that building on earlier sessions beats independent attempts. These results suggest that agentic self-evolution can optimize a complex system end-to-end. The evaluation also has to evolve with the engine, and letting agents evolve it is a natural next step.