Score
Designs and builds models, input encoders, prompting strategies, exemplars, and policies that explicitly incorporate external, spatial, temporal, or task-specific context into predictions, planning, or generation (including language-conditioned and motion planning pipelines), and implements the input‑transformation and agent loop code that feeds that context to models. Develops evaluation and ablation protocols for in‑context, few‑/one‑shot and long‑context behavior — measuring how context changes outputs and adapting policies or planners accordingly, and comparing performance across model variants.
Large language models (LLMs) excel at complex contextual understanding but exhibit pronounced capability asymmetry—struggling to stably generate long, equally sophisticated texts. Method: We systematically establish a unified “context engineering” framework, proposing a four-dimensional taxonomy encompassing retrieval, generation, processing, and management. Based on a systematic review and architectural analysis of 1,300+ papers, we construct the first comprehensive context engineering technology map; identify the intrinsic mechanisms underlying the understanding–generation capability mismatch; and delineate architectural integration pathways for four key application paradigms: retrieval-augmented generation, memory modeling, tool integration, and multi-agent coordination. Contribution/Results: The work delivers a standardized conceptual framework, a strategic technology roadmap, and identified critical breakthrough directions—providing both theoretical foundations and practical guidance for developing advanced context-aware AI systems.
This study investigates whether persistent context files—such as AGENTS.md—enhance the task correctness of AI coding agents in real-world codebases. Through controlled ablation experiments on 17 authentic repository tasks using Claude Code and Codex, the work employs gold-standard testing, failure mode categorization, and equivalence testing to rigorously evaluate context injection strategies in a multi-agent, realistic setting for the first time. Results indicate that context files do not significantly improve correctness for either agent (with an upper bound of ≤15 percentage points) and fail to convert near-correct outputs into passing solutions. The primary cause of failure stems from insufficient implementation capability rather than missing knowledge. Furthermore, task difficulty exhibits agent-specific characteristics (Spearman ρ = 0.75), clarifying the source of contradictory findings in prior literature.
This work proposes a training-free multi-agent optimization framework designed to simultaneously satisfy correctness and performance requirements for system code under the constraint of accessing large language models solely through API calls. The approach decouples the optimization context into three orthogonal dimensions—semantic summarization, directional guidance, and experience sampling—and, for the first time, establishes a functional isomorphism in textual latent space that emulates reinforcement learning components: state representation, policy gradients, and experience replay. Efficient directed evolution is achieved through multi-agent collaboration, code-language abstraction, trajectory-guided directional distillation, and priority-based exemplar retrieval. Evaluated on the ADRS benchmark, the method outperforms the current state-of-the-art by 33.3% in performance while reducing token consumption by 29.0%.
This work addresses the limited interpretability and generalizability of policies in grid-world tasks. We propose Iterative Procedural Planning (IPP), a framework that leverages large language models (LLMs)—including GPT-4o-mini and five others—to generate executable, verifiable code-based policies that directly map environment states to action sequences, bypassing traditional search or reinforcement learning. IPP introduces a novel task-feedback-driven code refinement mechanism, integrating pseudocode guidance, curriculum-style prompting, and multi-stage prompting, while enabling cross-task policy reuse. On the GRASP benchmark, IPP achieves new state-of-the-art performance; compared to direct code generation, 5 out of 6 LLMs show 10%–10× improvements in success rate; relative to direct solving, it boosts success rates by 63% on MiniGrid and 116% on GRASP. Policy reuse reduces amortized computational cost by 400×.
Current AI coding agents often produce code requiring extensive debugging due to insufficient contextual understanding, thereby diminishing development efficiency. This work proposes a three-phase preparation methodology inspired by the culinary concept of “mise en place”—termed MEP—comprising context anchoring, collaborative specification formulation, and task decomposition, which enhances agent coding performance through structured contextualization. The study introduces “contextual fluency” as a novel developer competency, integrating backward design principles and theories of tacit knowledge externalization. It employs structured documentation, human–AI collaborative dialogues, and dependency-aware task logging to operationalize this approach. In a hackathon setting, just two hours of preparatory work enabled multiple AI agents to concurrently and effectively construct a complete educational platform, substantially reducing the overall development cycle.
To address semantic drift, reasoning degradation, and context explosion in long-horizon software engineering agents operating over ultra-large-scale codebases—stemming from uncontrolled context growth—this paper introduces the “Context-as-Tool” (CAT) paradigm, which explicitly models context management as callable, learnable tools. Methodologically: (1) we construct a structured workspace that decouples high-fidelity short-term interactions from compressed long-term memory; (2) we design CAT-GENERATOR, a trajectory-level supervised framework enabling milestone-driven proactive compression; and (3) we develop SWE-Compressor, a context-aware compression model. Evaluated on SWE-Bench-Verified, our approach achieves a 57.6% task success rate—significantly outperforming ReAct baselines and static compression methods—while ensuring robustness and scalability of long-range reasoning under fixed context budgets.
This work addresses the limitation of large language model–driven software engineering agents, which struggle to maintain the long-term context required for complex, multi-step tasks due to finite context windows. The study presents the first systematic investigation into implicit context compression methods for such tasks, proposing an In-Context Autoencoder that compresses contextual information into continuous embeddings to circumvent length constraints. While the approach demonstrates strong performance on single-step code comprehension and commonsense reasoning benchmarks, it exhibits significant degradation in multi-step agent-based coding tasks. These findings reveal a fundamental limitation of current implicit compression techniques when applied to complex, temporally extended software engineering scenarios, offering critical insights for future research directions in context management for intelligent software agents.
This work addresses the absence of benchmarks evaluating language model agents’ ability to consistently adhere to complex, constraint-laden instructions—such as corporate policy manuals—over long contexts and multi-turn tool interactions. The authors introduce the first benchmark for this challenge, comprising 65 tasks across five professional domains, which requires agents to operate within a simulated office environment (e.g., email, calendar, chat) guided by dynamic policy manuals ranging from 20 to 124 pages. Leveraging expert-authored, non-redundant manuals and 824 deterministic scoring rules, the benchmark enables fully automated, stringent evaluation where all criteria must be satisfied. Experiments reveal that even the best-performing configuration among 30 state-of-the-art models passes only 36.2% of tasks, with most scoring below 25%, exposing systemic deficiencies in policy compliance and behavioral consistency.