ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of evaluation benchmarks for self-improvement in dynamic environments among large language model (LLM) agents by proposing EESD, a formal definition supporting hidden strategy evolution, alongside streaming environment datasets and evaluation benchmarks. Employing continual learning frameworks such as RAG, Mem0, and SkillOpt, the research systematically evaluates the adaptive capabilities of six LLM categories across diverse domains, including retail. This work bridges existing gaps in benchmark scale and knowledge diversity, revealing a significant disparity between task execution proficiency and experiential learning. Furthermore, it identifies high adaptation costs and insufficient exploration as core bottlenecks constraining autonomous agent evolution.
📝 Abstract
Large language model agents are increasingly deployed to perform complex tasks in real-world environments. However, the knowledge required for correct behavior in these environments is often implicit, undisclosed, and subject to change over time. Recent continual-learning harnesses seek to address this challenge by enabling agents to improve from serving experience. Yet the effectiveness and limitations of these methods are not yet well characterized. Existing benchmarks provide only partial coverage: some explicitly provide the target knowledge, others assume a static environment, and those that support continual adaptation remain limited in scale and knowledge diversity. To enable systematic evaluation, we formalize an evolving-environment streaming dataset (EESD), in which agents must infer, apply, and revise latent environment knowledge from interaction and outcome feedback as hidden policies evolve, and introduce ServeLearnBench, spanning retail support, banking, and sales-pitch generation with 53 environment windows and 7,718 tasks. We evaluate five learning harnesses (RAG, Mem0, SkillOpt, Continual Harness, and Prime) across six models (GPT-5.6 Terra, Opus 5, Kimi K3, GLM-5.3, DeepSeek V4.1 Flash, and GLM-5.3 Flash), covering 28 model-harness pairs and 252 learning runs. Our evaluation reveals three main findings: a substantial gap remains between task capability and learning from experience; continual adaptation is costly and can degrade already-correct behavior; and insufficient exploration emerges as a key bottleneck to effective adaptation. Overall, ServeLearnBench provides a controlled testbed for diagnosing these limitations and tracking progress toward agents that continually and reliably improve through serving experience.
Problem

Research questions and friction points this paper is trying to address.

LLM agents
continual learning
serving experience
evolving environment
benchmark evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

ServeLearnBench
Evolving-Environment Streaming Dataset
Continual Learning
LLM Agents
Self-Improvement
🔎 Similar Papers
No similar papers found.