Breaking the Environment Wall: Evolving LLM Agent Environments for Recursive Self-Improvement

πŸ“… 2026-09-24
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the significant performance degradation of LLM-based agents in environments characterized by fragmented information, noise interference, and dynamic changes. To this end, it proposes Env-Rethink, a system that introduces a novel environment-adaptive reconstruction mechanism. By constructing context maps and event logs, combined with offline trajectory learning, the method identifies environmental noise and dynamically adjusts state-evidence relationships. It further generates highly challenging virtual event histories to drive recursive agent self-improvement. Built upon a 27B post-trained model, Env-Rethink is validated across 30 tasks involving nine models. Experimental results demonstrate that the system improves scoring pass rates by over 15.1%, significantly enhancing both the robustness and generalization capabilities of LLM agents operating in complex, noisy environments.
πŸ“ Abstract
Many real-world tasks (e.g., office workflows, scientific experimentation) require LLM agents to interact repeatedly with their environments for context-dependent operations. However, such environments are often not agent-ready. First, information is often scattered and fragmented across the environment. Second, relevant evidence in the environment is often mixed with misleading information and conflicting versions. Third, environments evolve over time, introducing new noise and more challenging tasks. These challenges can substantially degrade performance for state-of-the-art AI agents (e.g., from 83.9% to 57.6%). To address these challenges, we propose Env-Rethink (a system with 27B post-trained model) that supports three main capabilities: (1) It adaptively builds Collection Maps (for organizing related files) and Event Logs (for contextualizing cross-data relationships) to supplement necessary context; (2) It further leverages the post-trained model (through offline trajectory learning) to identify underlying noise issues in the environment; (3) It ultimately evolves environments through virtual event histories that alter environmental states and evidence relationships, producing more tricky ones for further agent improvement. Experiments show that Env-Rethink can effectively improve downstream task performance (with over 15.1% rubric pass rate improvement across nine models on 30 tasks).
Problem

Research questions and friction points this paper is trying to address.

LLM agents
environment readiness
information fragmentation
noise and conflicts
dynamic environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Agent
Environment Evolution
Recursive Self-Improvement
Offline Trajectory Learning
Context Organization
Y
Yukai Wu
Shanghai Jiao Tong University; Theseus Labs; Tencent Hunyuan
Y
Yuanjing Yang
Shanghai Jiao Tong University; Theseus Labs
L
Le Zhou
Shanghai Jiao Tong University; Theseus Labs
S
Shaokun Han
Shanghai Jiao Tong University; Theseus Labs
Haoyu Wang
Haoyu Wang
University of Pennsylvania, Shanghai Jiao Tong University
Natural Language ProcessingComputer VisionKnowledge Graph
Z
Zirui Tang
Shanghai Jiao Tong University; Theseus Labs
W
Weihuang Zheng
Tencent Hunyuan
M
Maxm Pan
Tencent Hunyuan
Xuanhe Zhou
Xuanhe Zhou
Assistant Professor, Shanghai Jiao Tong University
Data ManagementArtificial Intelligence
Fan Wu
Fan Wu
Professor, Department of Computer Science and Engineering, Shanghai Jiao Tong University
Wireless NetworkingMobile ComputingAlgorithmic Game Theory and Its Applications