TTSE: A Two-Track Online Self-Evolution Framework

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出TTSE框架,通过FACT和TIP双轨机制解决大语言模型在持续交互环境中的自我进化问题,实验证明其优于单轨方法。
📝 Abstract
As Large Language Model (LLM) agents are applied in continuously interactive environments, driving the evolution of their own capabilities becomes a core problem for achieving long-term autonomy. Currently, environmental knowledge is typically treated as an external fixed input rather than as part of the agent's ongoing evolution. Reinforcement learning methods usually optimize policies through environmental interaction but tend to adapt only to fixed task distributions or single environments. This paper proposes TTSE (Two-Track Self-Evolution), a dual-track online self-evolution framework that separates evolving knowledge into FACT (environmental facts, whose reliability is continuously verified through interaction evidence) and TIP (task-conditioned implementation procedures). From a decision-theoretic perspective, we decompose the agent's excess risk into environment-representation regret and conditional-execution regret, characterize the conditions under which environment-conditioned policies strictly outperform condition-agnostic policies, and bound the downstream risk in terms of FACT identification error and cross-condition mismatch cost. In practice, TTSE's ablation experiments on GDPevo validate the advantage of dual-track evolution. On the classic agent task benchmarks ALFWorld and ScienceWorld, TTSE further demonstrates superior task adaptation. Moreover, TTSE is broadly compatible with existing skill self-evolution methods; combined with the Bayesian-Agent algorithm, a single-track ablation validates the dual-track advantage, substantially improving the aggregate score across the five major domains of SOPBench over three independent repetitions. Finally, on the real end-to-end task benchmark PinchBench, TTSE is integrated into a general agent framework via retrieval-based injection and stably outperforms the baseline across three independent runs.
Problem

Research questions and friction points this paper is trying to address.

Large Language Model
Self-Evolution
Continuously Interactive Environments
Reinforcement Learning
Task Adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Two-Track Self-Evolution
FACT and TIP Separation
Environment-Conditioned Policies
Excess Risk Decomposition
Task Adaptation
🔎 Similar Papers
No similar papers found.
R
Ruimin Pei
Poisson Lab, Huawei
Y
Yongkang Wu
Poisson Lab, Huawei
S
Shangyi Zheng
Poisson Lab, Huawei
Y
Yaqing Zhang
Poisson Lab, Huawei
D
Deyang Li
Poisson Lab, Huawei
J
Jianjun Tao
Poisson Lab, Huawei
Xinyu Zhang
Xinyu Zhang
Huawei Poisson Lab
Information RetrievalGenerative Search
X
Xiang Zhang
Poisson Lab, Huawei