FlowState: Execution State as Memory for Long-Horizon LLM Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the dilemma faced by LLM agents in long-horizon tasks, where retaining complete interaction histories incurs prohibitive costs while compressing them leads to critical information loss. To resolve this, we propose an execution-state-as-memory framework featuring a semantically typed state node architecture that decouples persistent storage from on-demand access. Through incremental state update (ISU) and progressive state access (PSA) mechanisms, the framework enables dynamic reuse of historical information and re-evaluation of prior decisions. Experimental results demonstrate that, when instantiated with the DeepSeek-V4-Flash model, our approach improves success rates by 4.55% and 13.95% on the MemoryArena and τ³-Bench benchmarks, respectively, while reducing token consumption by over 40%. These findings indicate that the proposed method effectively unifies performance and efficiency for long-horizon agent tasks.
📝 Abstract
Long-horizon tasks require LLM agents to continually draw on information from earlier interactions. However, retaining the full history increases context costs, while compressing it risks losing details needed later, and the relevance of historical information often becomes apparent as the task progresses. To address these challenges, we propose FlowState, which treats execution state as memory that can be retained and revisited across requests, unifying current decision-making with the reuse of historical information. FlowState preserves semantically typed state nodes, their relations, and references to raw tool observations, separating persistent retention from on-demand access. Within a single execution loop, Incremental State Update (ISU) maintains the current state based on new inputs and feedback, while Progressive State Access (PSA) progressively reveals historical states and supporting evidence as needed during reasoning. Together, these mechanisms enable agents to reassess prior decisions in light of new information and guide subsequent actions. Compared with a full-context baseline using the same DeepSeek-V4-Flash model, FlowState improves the average success rate on MemoryArena and the average pass rate on $τ^3$-Bench by 4.55 and 13.95 percentage points, respectively, while reducing total token consumption by 43.2% and 40.6%. These results demonstrate the performance and efficiency advantages of FlowState on long-horizon tasks.
Problem

Research questions and friction points this paper is trying to address.

Long-horizon tasks
LLM agents
Memory management
Context cost
Information retention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Execution State Memory
Incremental State Update
Progressive State Access
Long-Horizon LLM Agents
Structured State Representation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Minghao Li
Minghao Li
Beihang University
Natural Language Processing
B
Bangyan Li
Ant International, Ant Group
Z
Zifan Wang
Ant International, Ant Group
Y
Yulong Li
Ant International, Ant Group
H
Hu Xu
Ant International, Ant Group
Gan Zhang
Gan Zhang
Guangzhou Institute of Geochemistry, CAS
J
Jingtong Wu
Ant International, Ant Group
Wenqiang Xu
Wenqiang Xu
Shanghai Jiao Tong University
Computer visionRobotics