🤖 AI Summary
This study addresses the GPU resource contention between online agent self-evolution and real-time serving, which causes delayed returns and high recovery costs. We propose LearnSched, a scheduler that incorporates capability reuse windows and computation regression times into a unified investment model to dynamically determine evolution actions. Methodologically, we construct a state-aware scheduling framework integrating candidate progress, checkpoint overhead, and recovery paths for global optimization, alongside a one-step counterfactual reasoning algorithm to select optimal execution strategies. Experimental results demonstrate that our approach significantly improves net system value under cold-recovery scenarios, validating that preserving recoverable states effectively shortens capacity regression cycles and reduces scheduling complexity.
📝 Abstract
Online agents can improve future service by constructing reusable tools, guidance, or model states, but this work competes with current requests for the same GPUs. Exploiting idle compute for self-evolution faces a fundamental systems constraint: benefits arrive only after an artifact is published and used, while pausing evolution leaves service capacity waiting for memory release and runtime recovery. An investment worth completing may therefore be worth postponing. We present LearnSched, a state-aware scheduler that brings the reuse window of a capability and the timely return of compute into a common investment model. LearnSched incorporates candidate progress, checkpoint overhead, and the current recovery path into action values. Under the same information and capacity constraints, it uses one-step counterfactual rollout to choose progress, checkpointing, or waiting relative to a fully costed window policy. We characterize handoff costs through independent A100 component measurements and evaluate action selection in 1920 paired finite-model scenarios. When cold recovery is expensive, waiting for longer execution windows can improve net value by avoiding handoffs; in evaluated warm-recovery scenarios, the strong window policy already achieves the same value. Retaining recoverable state can shorten capacity return and reduce the need for complex evolution scheduling.