Score
Designs, builds, and evaluates pipelines that select, retrieve, filter, and fuse contextual information presented to a model or service while encoding and enforcing explicit boundaries on what context is accessible. Work includes defining allowlisted model-accessible fields, implementing context acquisition/incorporation and retrieval-and-fusion mechanisms, specifying implementation and hardware capability limits, and preventing raw-data leakage through sanitization and access-control policies.
Large language models (LLMs) excel at complex contextual understanding but exhibit pronounced capability asymmetry—struggling to stably generate long, equally sophisticated texts. Method: We systematically establish a unified “context engineering” framework, proposing a four-dimensional taxonomy encompassing retrieval, generation, processing, and management. Based on a systematic review and architectural analysis of 1,300+ papers, we construct the first comprehensive context engineering technology map; identify the intrinsic mechanisms underlying the understanding–generation capability mismatch; and delineate architectural integration pathways for four key application paradigms: retrieval-augmented generation, memory modeling, tool integration, and multi-agent coordination. Contribution/Results: The work delivers a standardized conceptual framework, a strategic technology roadmap, and identified critical breakthrough directions—providing both theoretical foundations and practical guidance for developing advanced context-aware AI systems.
This paper addresses security and privacy risks arising from the Model Context Protocol (MCP) in AI model–external tool interoperability. We first systematically define MCP’s full lifecycle—encompassing creation, execution, and update phases—and develop a stage-specific threat taxonomy with corresponding mitigation strategies. Integrating protocol design principles, security threat modeling, privacy risk analysis, and industry ecosystem surveys, we propose an MCP security governance guideline, a compatibility mapping across major platforms, and a sustainable development roadmap. Our core contribution is the establishment of the first comprehensive MCP lifecycle security model, unifying technical implementation, platform integration, and ecosystem evolution into a coherent research paradigm. This work provides both theoretical foundations and practical benchmarks for trustworthy AI interoperability. (136 words)
This work addresses the limited introspectability, visualizability, and interoperability with external tools in existing information retrieval (IR) pipelines, which hinder their interpretability and integration efficiency. To overcome these limitations, the paper introduces novel operations within the PyTerrier framework that enable structured introspection, interactive visualization, and interoperability via the Model Context Protocol (MCP). These capabilities facilitate transparent inspection and dynamic exploration of IR workflows, significantly enhancing pipeline transparency, debuggability, and cross-tool integration. The proposed approach provides researchers, students, and AI agents with more effective means to understand, analyze, and utilize IR systems.
This work addresses the security risks posed by Model Context Protocol (MCP) servers, which often expose high-risk capabilities such as file system access, network requests, and command execution that can be exploited if not properly audited. To mitigate this, we present mcp-sec-audit, the first security auditing framework specifically designed for the MCP protocol. Our approach combines static pattern matching with dynamic sandboxed fuzz testing powered by Docker and eBPF to automatically identify and assess these hazardous capabilities. The framework supports extensible rule configuration and fully automated detection, and has been validated on Python-based MCP server implementations. It accurately generates actionable hardening recommendations, thereby significantly enhancing the overall security posture of the MCP ecosystem.
This study addresses the lack of systematic guidance on contextualization strategies for large language model (LLM) agents operating in structured data environments, particularly concerning effectiveness and efficiency across multi-file, large-scale schemas. Using SQL generation as a proxy task, the work presents the first systematic evaluation of eleven models across four context formats—YAML, Markdown, JSON, and TOON—at schema scales ranging from 10 to 10,000 tables. The findings reveal that model capability tiers critically determine optimal context architecture: tailored strategies significantly improve performance, with state-of-the-art models gaining 2.7% accuracy under native file-based contexts, while open-source models average a 7.7% decline. Moreover, native file-based agents scale efficiently to ten-thousand-table schemas while maintaining high navigation accuracy.
This study addresses the privacy risks posed by enterprise large language model (LLM) agents, which, while enhancing operational efficiency, are prone to leaking sensitive information through internal contextual cues. For the first time, the paper introduces the Contextual Integrity (CI) theory to evaluate privacy in enterprise LLM agents and constructs CI-Work, a benchmark simulating five canonical enterprise information-flow scenarios. Using dense retrieval and multi-directional workflow modeling, the authors systematically assess mainstream LLMs, revealing a counterintuitive trade-off between task utility and privacy preservation. Results show that privacy violation rates range from 15.8% to 50.9%, with information leakage reaching up to 26.7%. Notably, merely scaling model size or increasing reasoning depth fails to mitigate these issues, underscoring the need for context-centric privacy-preserving architectures.
论文提出SCOUT系统,通过选择性上下文优化解决大型语言模型在使用外部工具时面临的上下文饱和和工具发现难题。
This work addresses the limitations of existing privacy and AI compliance assessment methods, which often assume complete contextual information despite real-world scenarios frequently involving ambiguity or missing context. To bridge this gap, the paper introduces ContextLens, a novel framework that—without requiring model training—integrates rule-based reasoning with large language models through semi-formalized inference to guide structured responses to legal compliance queries. ContextLens explicitly models compliance risks under incomplete context and identifies unknown or ambiguous elements. Evaluated on benchmarks aligned with the GDPR and the EU AI Act, the approach significantly outperforms current methods, not only improving judgment accuracy but also effectively surfacing contextual uncertainties inherent in compliance assessments.
This work addresses the lack of a trust mechanism for third-party tool servers in the Model Context Protocol (MCP) and its susceptibility to unauthorized invocations. We propose the first security extension that requires no modifications to the existing protocol or APIs. Our approach introduces offline-signed admission assertions, server-level tool allowlists, and configurable enforcement policies ranging from warnings to outright rejection. Security and consistency are ensured through URI-distributed signed assertions, pinned trust root verification, tamper-resistant audit logs, and machine-verifiable test vectors. The solution has been integrated into the enclawed-oss and enclaved distributions and validated through formal security analysis and LLM adversarial evaluation. The resulting specification conforms to RFC 2119 and is ready for direct adoption as an MCP appendix.
本文系统分析了AI代理程序中上下文组装设计的安全风险,揭示了两种新型攻击方式:消息角色上下文特权提升和跨范围上下文特权提升,并对12个实际应用进行了安全评估。
AEGIS通过利用大型语言模型的推理能力,定义细粒度的安全措施,防止MCP工具跨域资源滥用。