Score
Designs, implements, and evaluates mechanisms and policies for maintaining and updating bounded context windows in stateful or sequence-processing systems — including sliding-window maintenance, truncation and expansion strategies, alignment of overlapping contexts for synthesis, and caching and retrieval of past trajectories to preserve continuity and efficiency.
Large language models (LLMs) excel at complex contextual understanding but exhibit pronounced capability asymmetry—struggling to stably generate long, equally sophisticated texts. Method: We systematically establish a unified “context engineering” framework, proposing a four-dimensional taxonomy encompassing retrieval, generation, processing, and management. Based on a systematic review and architectural analysis of 1,300+ papers, we construct the first comprehensive context engineering technology map; identify the intrinsic mechanisms underlying the understanding–generation capability mismatch; and delineate architectural integration pathways for four key application paradigms: retrieval-augmented generation, memory modeling, tool integration, and multi-agent coordination. Contribution/Results: The work delivers a standardized conceptual framework, a strategic technology roadmap, and identified critical breakthrough directions—providing both theoretical foundations and practical guidance for developing advanced context-aware AI systems.
To address semantic drift, reasoning degradation, and context explosion in long-horizon software engineering agents operating over ultra-large-scale codebases—stemming from uncontrolled context growth—this paper introduces the “Context-as-Tool” (CAT) paradigm, which explicitly models context management as callable, learnable tools. Methodologically: (1) we construct a structured workspace that decouples high-fidelity short-term interactions from compressed long-term memory; (2) we design CAT-GENERATOR, a trajectory-level supervised framework enabling milestone-driven proactive compression; and (3) we develop SWE-Compressor, a context-aware compression model. Evaluated on SWE-Bench-Verified, our approach achieves a 57.6% task success rate—significantly outperforming ReAct baselines and static compression methods—while ensuring robustness and scalability of long-range reasoning under fixed context budgets.
This work addresses the challenge that large language models face in maintaining coherent state memory across long-horizon, multi-session tasks due to constraints imposed by fixed context windows and KV cache limitations. The authors propose a model-agnostic cognitive state plane that, for the first time, formalizes cognitive states as dynamically evolving structures integrating episodic, semantic, and procedural memory. Grounded in formal models from cognitive psychology, the framework incorporates information-theoretically constrained selective encoding, goal-conditioned retrieval, reconstructive synthesis, and adaptive forgetting mechanisms. It further integrates KV-aware algorithms and write-path poisoning defenses to support enterprise-grade deployment. Evaluations across six domain-specific benchmarks demonstrate significant performance improvements on long-horizon intelligent tasks without requiring context window expansion or model retraining.
This work addresses the inefficiency of traditional token-level context modeling in distinguishing between recollective, summarizing, and local information. The authors propose a novelty-driven memory mechanism that dynamically partitions context into three components: a content-addressable novelty cache for retrievable details, a recurrent state for compressed summaries, and a sliding window for recent local context. This architecture uniquely scales memory capacity with the amount of distinct information rather than raw token count, yielding an auditable and interpretable working memory structure. Integrating a Dirichlet-process-inspired novelty-gated attention mechanism, the system achieves full-attention performance in character-level control tasks with roughly half the attention cost and outperforms both full-attention and fixed-budget baselines on a thousand-event healthcare claims prediction task, while enabling human inspection of stored memory contents.
This work addresses decision failures in long-horizon embodied agents caused by task-state inconsistencies—such as phase locking and actuator-context mismatches—by introducing an auditable task-state alignment framework. The approach models task phases as explicit contracts and generates runtime evidence bundles through hierarchical task representations. Alignment of the task frontier is achieved via scoped, local update operations: continue, refine, shift, elevate, and repair. This is the first framework to integrate contract-based phase modeling with evidence-driven mechanisms, ensuring task consistency under multi-actuator coordination. Experimental results demonstrate that the method effectively diagnoses and mitigates state misalignments, substantially reducing unnecessary replanning while enhancing system robustness and interpretability.
This work addresses the degradation of recent interaction influence in long-horizon agents due to recursive context compression, which often leads to action blocking, redundant exploration, and cross-run instability. To mitigate these issues, the authors propose TRACE, a framework that employs a verifier-guided closed-loop contrastive mechanism to refine natural language compression prompts through boundary-localized evaluation—without updating the underlying model. By integrating paired closed-loop continuations with summary preference learning, TRACE substantially enhances the reliability of context compression. Evaluated on the AppWorld benchmark, the method consistently outperforms existing baselines in task performance, multi-run stability, and context-to-execution efficiency.
In long-horizon tasks, LLM agents frequently trigger KV cache invalidation due to context engineering techniques such as offloading and compression, leading to substantial increases in first-token latency. This work reveals, for the first time, that context transitions exhibit segment-wise decomposability. Building on this insight, the authors propose a proactive programming model that asynchronously pre-executes transition operations and introduce an interference-aware scheduler. Without modifying existing agent logic, this approach proactively generates target KV caches, enabling seamless, non-blocking context replacement. Integrated into mainstream agent frameworks and LLM serving systems, the method reduces first-token latency by up to 11.9× and entirely eliminates the performance overhead associated with context transitions.
Shared state profoundly influences the performance and fault tolerance of stream processing, service-oriented, and continual learning systems, yet existing approaches often treat access control, hardware-aware execution, memory management, and long-term evolution in isolation. This work reframes state management as a runtime control problem and introduces a contract-driven blueprint centered on state objects, control planes, coupling paths, evaluation boundaries, and pending contracts. Building upon this foundation, we develop a unified analytical framework encompassing state-access scheduling, state-aware execution, and state evolution reuse. Through systematic scheduling, runtime control, and cross-layer coupling analysis, our approach identifies critical anti-patterns and advances a perturbation-aware evaluation paradigm, thereby establishing both theoretical foundations and practical design guidelines for state control in distributed systems.