Serve Now or Improve Later? Scheduling Self-Evolution in Online Agent Systems

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the GPU resource contention between online agent self-evolution and real-time serving, which causes delayed returns and high recovery costs. We propose LearnSched, a scheduler that incorporates capability reuse windows and computation regression times into a unified investment model to dynamically determine evolution actions. Methodologically, we construct a state-aware scheduling framework integrating candidate progress, checkpoint overhead, and recovery paths for global optimization, alongside a one-step counterfactual reasoning algorithm to select optimal execution strategies. Experimental results demonstrate that our approach significantly improves net system value under cold-recovery scenarios, validating that preserving recoverable states effectively shortens capacity regression cycles and reduces scheduling complexity.
📝 Abstract
Online agents can improve future service by constructing reusable tools, guidance, or model states, but this work competes with current requests for the same GPUs. Exploiting idle compute for self-evolution faces a fundamental systems constraint: benefits arrive only after an artifact is published and used, while pausing evolution leaves service capacity waiting for memory release and runtime recovery. An investment worth completing may therefore be worth postponing. We present LearnSched, a state-aware scheduler that brings the reuse window of a capability and the timely return of compute into a common investment model. LearnSched incorporates candidate progress, checkpoint overhead, and the current recovery path into action values. Under the same information and capacity constraints, it uses one-step counterfactual rollout to choose progress, checkpointing, or waiting relative to a fully costed window policy. We characterize handoff costs through independent A100 component measurements and evaluate action selection in 1920 paired finite-model scenarios. When cold recovery is expensive, waiting for longer execution windows can improve net value by avoiding handoffs; in evaluated warm-recovery scenarios, the strong window policy already achieves the same value. Retaining recoverable state can shorten capacity return and reduce the need for complex evolution scheduling.
Problem

Research questions and friction points this paper is trying to address.

online agent systems
self-evolution scheduling
GPU resource contention
checkpoint overhead
capacity recovery
Innovation

Methods, ideas, or system contributions that make the work stand out.

state-aware scheduler
self-evolution
counterfactual rollout
checkpoint overhead
recovery cost
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yangbo Wei
Shanghai Jiao Tong University
J
Junhong Qian
Shanghai Jiao Tong University
Zhen Huang
Zhen Huang
National University of Defense Technology
distributed storageNLPmachine learning
Z
Zhenyu Su
Shanghai Jiao Tong University
Q
Qifan Wang
Shanghai Jiao Tong University
S
Shaoqiang Lu
Shanghai Jiao Tong University
Rumin Zhang
Rumin Zhang
Ningbo Institute of Digital Twin, Eastern Institute of Technology
Chen Wu
Chen Wu
Institute of Computing Technology, Chinese Academy of Sciences
Information RetrievalNatural Language ProcessingAdversarial Attack
L
Lei He
Eastern Institute of Technology