🤖 AI Summary
This study addresses the memory bottleneck caused by context accumulation in long-horizon LLM agent tasks, as well as the information loss and decision inconsistencies introduced by existing compression methods. We propose ARSM, a lightweight, training-free framework that introduces an autoregressive self-compressive generation space. By employing a Hierarchical Action-Reaction (HAR) micro-chain trajectory abstraction mechanism to reorganize interaction histories and leveraging an atomic operation-based dynamic state machine to regulate hierarchical memory, ARSM achieves in-situ reasoning compression. This enables model outputs to simultaneously execute external actions and update internal states without requiring auxiliary models or specialized optimization. Evaluations on benchmarks such as WebShop demonstrate that our approach significantly reduces token consumption while preserving task performance, offering an efficient and cost-effective solution for scalable autonomous agents.
📝 Abstract
While Large Language Model (LLM)-based agents demonstrate strong capabilities in long-horizon tasks by interleaving reasoning with external environment interactions, the continuous accumulation of context rapidly creates a critical memory bottleneck. Existing memory compression methods rely on task-specific optimization or external auxiliary models, introducing significant computational overhead. Furthermore, the resulting compressed representations tend to lose structured relationships, leading to information dilution, attention collapse, and degraded decision consistency. To address these limitations, we propose Auto-Regressive State Machine (ARSM), a lightweight training-free framework that enables in-situ reasoning compression through structured state evolution. ARSM introduces two key components: (i) a trajectory abstraction mechanism that reorganizes interaction histories into compact Hypothesis-Action-Result (HAR) micro-chains; (ii) a dynamic state machine that regulates hierarchical memory through atomic operations and a compression-control parameter. These components are unified within an auto-regressive, self-compressive generation space, where each model output jointly performs external action execution and internal state updates. We evaluate ARSM on Webshop, Multi-Objective Multi-Hop QA, and SWE-Bench Lite datasets. Experimental results show that ARSM maintains the task performance while simultaneously reducing token consumption, offering a practical, cost-effective route toward scalable autonomous agents for long-horizon tasks.