Score
Designs, builds, and evaluates systems, protocols, and algorithms that create, represent, manage, compress, propagate, retrieve, and window the contextual state provided to or maintained by a model — covering context modeling, context compression, context propagation, context retrieval, and long-context handling. This work includes engineering context management systems and model context protocols, and devising strategies for context/state management, in‑context learning, contextual personalization, and effective use of limited context windows.
Large language models (LLMs) excel at complex contextual understanding but exhibit pronounced capability asymmetry—struggling to stably generate long, equally sophisticated texts. Method: We systematically establish a unified “context engineering” framework, proposing a four-dimensional taxonomy encompassing retrieval, generation, processing, and management. Based on a systematic review and architectural analysis of 1,300+ papers, we construct the first comprehensive context engineering technology map; identify the intrinsic mechanisms underlying the understanding–generation capability mismatch; and delineate architectural integration pathways for four key application paradigms: retrieval-augmented generation, memory modeling, tool integration, and multi-agent coordination. Contribution/Results: The work delivers a standardized conceptual framework, a strategic technology roadmap, and identified critical breakthrough directions—providing both theoretical foundations and practical guidance for developing advanced context-aware AI systems.
This work addresses the lack of rigorous mathematical definitions for context and its perception mechanisms in existing research, which hinders effective characterization of system-environment interactions. To bridge this gap, the paper presents a formal context modeling framework that deeply integrates context awareness with computational theory. Built upon extended Turing machines and interactive computation models, and further enriched by distributed computing principles, the proposed framework establishes a solid mathematical foundation for context operations. It inherently supports interactivity and networked capabilities while exhibiting strong robustness and scalability. The efficacy of the framework is empirically validated through cloud-native context-aware applications.
Multi-agent systems face fundamental challenges including disorganized context management, low collaboration efficiency, and poor scalability. To address these, this paper proposes the Model Context Protocol (MCP), a novel unified theoretical framework and scalable coordination paradigm. MCP introduces protocol-driven context modeling, hierarchical coordination scheduling, domain-adaptive knowledge injection, and a multi-granularity evaluation benchmark—enabling dynamic, cross-agent context awareness and semantic alignment while overcoming the limitations of static role assignment. Experimental results demonstrate that, in enterprise knowledge management and collaborative scientific research scenarios, MCP improves task completion efficiency by 42%, reduces communication overhead by 37%, and ensures stable coordination among up to one hundred agents. MCP establishes a standardized, reusable infrastructure for context-aware coordination in multi-agent systems.
This paper addresses security and privacy risks arising from the Model Context Protocol (MCP) in AI model–external tool interoperability. We first systematically define MCP’s full lifecycle—encompassing creation, execution, and update phases—and develop a stage-specific threat taxonomy with corresponding mitigation strategies. Integrating protocol design principles, security threat modeling, privacy risk analysis, and industry ecosystem surveys, we propose an MCP security governance guideline, a compatibility mapping across major platforms, and a sustainable development roadmap. Our core contribution is the establishment of the first comprehensive MCP lifecycle security model, unifying technical implementation, platform integration, and ecosystem evolution into a coherent research paradigm. This work provides both theoretical foundations and practical benchmarks for trustworthy AI interoperability. (136 words)
This study addresses the frequent inefficiencies in human-AI collaboration caused by incomplete contextual information, which often leads to excessive iteration and suboptimal output quality. To mitigate this, the authors propose a structured context construction framework that integrates a five-role context package—comprising authority, exemplars, constraints, evaluation criteria, and metadata—within a four-stage workflow encompassing review, design, construction, and audit. Notably, this work pioneers the incorporation of information theory and reliability engineering principles into context quality assessment, yielding a reusable and auditable collaboration framework. Empirical results from 200 interaction trials demonstrate that the approach reduces the average number of iterations from 3.8 to 2.0, increases first-pass success rates from 32% to 55%, and achieves a final task success rate of 91.5%.
This work addresses the limitations of existing context compression methods in long-horizon agent tasks, which suffer from information loss and rigid triggering mechanisms that hinder adaptive reasoning. To overcome these challenges, the authors propose the Agentic Context Management (ACM) framework, which introduces a novel mechanism inspired by human short-term and long-term memory interactions. ACM equips agents with dedicated context-editing tools to achieve lossless compression, autonomously decide when to compress, and leverage an external memory system for on-demand retrieval. By incorporating a post-training pipeline to generate high-quality demonstration data, ACM significantly enhances performance in search and programming tasks, reduces peak token consumption, enables extended exploration horizons, and yields more consistent solutions.
This work addresses the current lack of standardized methods for precisely characterizing the structure and dynamic evolution of input contexts in large language model (LLM) agent systems. To bridge this gap, the paper introduces ACDL—a formal, human-readable, and implementation-agnostic context description language that enables clear modeling of role-based message sequences, dynamic content, time-indexed references, and conditional or iterative constructs, complemented by visual representations. ACDL fills a critical semantic communication gap in LLM agent design and has been successfully employed to document multiple existing systems and their variants. An accompanying open-source toolchain and illustrative examples are provided to facilitate adoption within the research community and support its integration into academic discourse and system documentation.
This study addresses the tendency of large language models (LLMs) to overlook reference data embedded in server instructions within Model Context Protocol (MCP) environments, instead inefficiently invoking search tools and wasting computational resources. Through 54,000 controlled trials, the authors systematically evaluate the behavior of 24 prominent LLMs on legal information retrieval tasks under MCP, employing tool ablation, a 2³ factorial design, and cross-model-family analysis. The work reveals— for the first time—that this behavior stems from preference rather than capability deficits. Notably, removing the search tool yields over 98% accuracy in 23 out of 24 models, and combining three targeted prompting interventions restores accuracy above 86% for 20 out of 24 models even when the tool is present. The findings advocate for MCP hosts to explicitly prioritize server-provided instructions.
Large language models are constrained by fixed context windows and often require context compression to align with agent states, yet this process lacks theoretical foundations. This work formally defines the context compression problem for the first time and introduces two game-theoretic models—context selection and context generation—establishing their equivalence to one-way communication complexity. Theoretical analysis reveals the existence of query sets for which generative strategies strictly require less budget than selective ones. Empirical evaluation further demonstrates the suboptimality of commercial APIs (e.g., Anthropic) on set membership query tasks. Our study provides a theoretical benchmark and an evaluation framework for context compression methods.