ReLiveGym: Evaluating Long-Lived Agents over Weeks of Replayed Reality

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of evaluation frameworks for long-horizon agents facing persistent monitoring, precise decision-making, and adaptation challenges in dynamic environments. We construct a diagnostic environment based on real-world data stream replay and propose a temporal replay mechanism that establishes action timing as a critical design axis. Furthermore, we introduce a hindsight feedback learning framework to systematically evaluate large language model agents on sparse action-taking and continuous learning over multi-week spans. Our findings reveal that optimal architectures dynamically evolve across tasks and models, demonstrating that model selection, timing mechanisms, and feedback utilization exert decisive influences on long-term agent performance.
📝 Abstract
As large language model (LLM) agents become widely adopted, they are increasingly deployed for tasks that require persistent monitoring or recurring actions (e.g., market analysis). These agents are expected to operate unattended for days or weeks, act at the right timing, and adapt to the dynamic environment over time. These challenges are not fully captured in the existing long-horizon agent work, as they often consider a static environment that is not temporally changing. We introduce ReLiveGym, a diagnostic evaluation environment of long-lived tasks in which agents act sparsely over simulated weeks of chronologically replayed real-world news, market, and social-media streams. The tasks span diverse levels of time sensitivity, reasoning intensity, and recurrence. Across eight base language models, we investigate how model choice and harness design affect agent performance on such long-lived tasks. Our results show that how agents determine when to act arises as an important harness-design axis for long-lived tasks; and that the optimal design varies across tasks and sometimes model choices as well. We also evaluate how continuous learning from hindsight feedback affects performance and addresses failure modes observed in these long-lived tasks. These findings indicate model choice, action timing mechanism, and use of feedback as important considerations in the design of long-lived agents. Code: https://github.com/SaharaLabsAI/ReLiveGym
Problem

Research questions and friction points this paper is trying to address.

long-lived agents
dynamic environments
action timing
long-horizon tasks
LLM agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Long-lived Agents
Dynamic Environment Evaluation
Action Timing Mechanism
Continuous Learning
Hindsight Feedback
🔎 Similar Papers
No similar papers found.