Score
Designs and evaluates plans and workflows that identify, acquire, and incorporate missing contextual information needed for downstream tasks. This includes diagnosing what context is absent, prioritizing and selecting context sources, sequencing and scheduling actions to obtain or request context, and specifying fallback or adaptive steps to support reliable generation when full context is unavailable.
This study addresses the frequent inefficiencies in human-AI collaboration caused by incomplete contextual information, which often leads to excessive iteration and suboptimal output quality. To mitigate this, the authors propose a structured context construction framework that integrates a five-role context package—comprising authority, exemplars, constraints, evaluation criteria, and metadata—within a four-stage workflow encompassing review, design, construction, and audit. Notably, this work pioneers the incorporation of information theory and reliability engineering principles into context quality assessment, yielding a reusable and auditable collaboration framework. Empirical results from 200 interaction trials demonstrate that the approach reduces the average number of iterations from 3.8 to 2.0, increases first-pass success rates from 32% to 55%, and achieves a final task success rate of 91.5%.
This study addresses the high context management costs and frequent omission of critical information encountered by LLM agents during long-horizon tasks. To mitigate these issues, we propose a context optimization method grounded in execution dependency graphs and tool-flow analysis. Specifically, this approach introduces observation reuse signals to integrate execution dependencies into the context rendering mechanism, combining persistent dependency graphs with semantic retrieval for precise information filtering. Experimental results demonstrate that, under a fixed budget of 6K tokens, the proposed method achieves performance comparable to full-history baselines while reducing inference costs by 10.2%–32.2%. These findings indicate that our approach effectively balances information completeness with computational efficiency, offering a practical solution for resource-constrained agentic workflows.
Current AI evaluation methods often operate in abstraction from real-world deployment contexts, failing to assess an AI system’s capacity to sustainably generate value within specific organizations. This work proposes a novel “contextual specification” framework that leverages qualitative modeling and collaborative stakeholder analysis to transform ambiguous, context-dependent elements into clearly defined, nameable constructs. By explicitly delineating the attributes, behaviors, and outcomes that warrant evaluation, the framework establishes a set of observable and measurable context-sensitive metrics. This approach provides organizations with an actionable evaluation roadmap, effectively bridging the gap between technical performance and business value, thereby substantially enhancing the relevance and efficacy of AI deployment decisions.
Large language models (LLMs) face two key challenges in code completion: limited context window capacity and susceptibility to noisy, irrelevant context. This paper investigates how context granularity—file-level versus block-level—and retrieval ranking strategies affect generation quality. We propose a static-analysis-driven, block-level context retrieval method that enables fine-grained, semantically relevant context extraction, followed by optimized context composition and ordering. Experiments on Python code completion show that our approach improves completion accuracy by 6% over the best-performing file-level retrieval baseline and by 16% over a no-context baseline. Our core contributions are: (1) empirical validation that block-level context significantly enhances effectiveness under strict context-length constraints; and (2) the first lightweight, deployable retrieval framework that jointly integrates static program analysis with context sequence control. This work establishes a practical, production-ready paradigm for context optimization in industrial code completion systems.
Business processes frequently fail or underperform at runtime due to dynamic changes in contextual variables—such as workflow data and environmental states—yet existing BPM systems lack generic, non-domain-specific mechanisms for context adaptation. To address this, we propose a versatile context engine based on Complex Event Processing (CEP), embedded within BPM systems to support both initialization-time configuration and runtime reconfiguration of process instances. The engine integrates CEP, business rule engines, and context-aware architecture to enable, for the first time, context-driven dynamic compensation and intervention at decision points. Experimental evaluation demonstrates that our approach significantly enhances process responsiveness to runtime context changes, overcoming the limitations of static variable instantiation. It bridges a critical gap in BPM research by establishing a foundational framework for runtime context adaptability—marking the first general-purpose, context-adaptive mechanism for BPM systems.
This work addresses the frequent failure of AI agents in production environments due to errors or omissions in contextual sources such as system prompts, knowledge bases, or tool descriptions—a problem exacerbated by the reliance on manual log inspection for maintenance, which does not scale. To overcome this, the authors propose an automated context engineering framework that operates without explicit user feedback by mining implicit dissatisfaction signals (e.g., corrections, rephrasings, or task abandonment) from historical interaction trajectories. The framework integrates multi-component causal attribution with an exploratory validation strategy to automatically diagnose and repair contextual defects. Key contributions include the first verifiable simulation benchmark for context debugging, a taxonomy of six failure types, and a causal attribution and active verification mechanism applicable across heterogeneous context sources. Experiments demonstrate 72.7% root-cause attribution accuracy and 82% end-to-end repair effectiveness over 60 dissatisfaction trajectories, confirming the approach’s capability for efficient self-repair of context-layer faults.
This work addresses the challenges of context overflow, outdated state tracking, and escalating reasoning costs in enterprise workflows caused by verbose tool responses from large language model (LLM) agents. Focusing on a Microsoft Dynamics 365 expense reimbursement scenario, the authors propose a context management approach that integrates recent tool interaction pruning with automated summarization. Leveraging GPT-5 and Claude Sonnet 4.5 models via the Model Context Protocol, the method enables efficient context compression and multidimensional evaluation. Experimental results demonstrate that the approach achieves a 91.6% task completion rate and 99.64% monetary coverage using only 550k tokens and 5.79 hours of compute, significantly outperforming a full-history baseline while substantially reducing resource consumption without compromising performance.
This work addresses the unreliable behaviors of current AI agents, which often stem from low-quality runtime contexts, yet systematic metrics for context engineering remain lacking. The paper proposes ProofAgent-Harness, an evaluation framework that models context quality as an auditable layer independent of behavioral performance. By incorporating a multi-reviewer consensus mechanism and context isolation design, it establishes a seven-dimensional assessment schema encompassing role clarity, safeguard coverage, instruction consistency, and others. Experimental results demonstrate that individual context quality dimensions significantly predict corresponding behavioral outcomes—for instance, foundational adequacy correlates with hallucination resistance, and tool-mode quality predicts tool-use accuracy—thereby validating context quality as a leading indicator of agent reliability.
This study addresses the lack of a unified theoretical foundation in traditional requirements engineering (RE) quality assessment, which often fails to integrate artifact- and process-oriented perspectives and overlooks information transmission efficiency. To bridge this gap, the paper proposes a holistic theoretical framework that models RE as a flow of information particles among stakeholders, developers, testers, and artifacts, with information flow as its core construct. Building on this model, the authors develop a simulation system to capture dynamic interactions and information exchanges across roles. The simulation reveals how high-quality requirements specifications can be inadvertently bypassed in agile environments and yields actionable insights for improving RE processes. This work establishes a theoretical basis for optimizing information flow, enhancing RE effectiveness, and understanding the underlying causes of success or failure in requirements engineering practices.
为解决长时任务中上下文管理问题,通过扩展工具集并引入细粒度强化学习方法,提出ContextPilot框架以实现高效主动的上下文管理。