Score
Designs and implements mechanisms that traverse recorded task dependency trees or graphs top-down to detect mismatches or corrupted task outputs and recover by selectively recomputing only the affected tasks while reusing correct child results. Builds schedulers, replay controllers, and propagation logic (e.g., via futures/promises) that perform targeted replay/cross-validation-based detection and minimize recomputation overhead across dependent tasks and cluster resources.
This work addresses the lack of a unified framework in current LLM agent workflows, which hinders method comparison and reproducibility. To resolve this, we propose the Agent Computation Graph (ACG) framework, which models workflows as computation graphs and adopts “structure determines timing” as a core principle. The framework explicitly distinguishes between reusable templates, runtime instance graphs, and execution traces, enabling a systematic categorization of static and dynamic optimization approaches. Through a comprehensive literature review and conceptual modeling, we develop a multidimensional evaluation framework that integrates structural properties, establishes precise terminology, and defines standardized evaluation criteria. This foundation supports a reproducible and highly comparable research paradigm for optimizing LLM agent workflows.
Repairing execution traces of robots and cyber-physical systems under Temporal Behavior Tree (TBT) specifications remains challenging due to the prohibitive computational cost and poor scalability of existing Mixed-Integer Linear Programming (MILP)-based approaches. Method: This paper proposes two efficient trace repair strategies—incremental repair and landmark-guided repair—leveraging the robust semantics of TBTs. Both methods approximate MILP-based repair using linear programming and piecewise iterative optimization, thereby avoiding combinatorial explosion. Contribution/Results: The proposed methods significantly improve scalability and efficiency: on trace instances with over 100,000 time steps, repair completes within ≤10 minutes, whereas MILP fails due to memory overflow. Furthermore, the framework enables fault attribution analysis and high-quality training sample generation, establishing a novel paradigm for trustworthy operation and maintenance of TBT-driven systems.
This work addresses the limitations of existing AI agent workflows, which rely on implicit dialogue states and struggle to ensure stability of intermediate artifacts, isolate irrelevant updates, and propagate changes precisely. To overcome these challenges, the paper proposes modeling AI-native workflows as directed acyclic graphs (DAGs) and introduces the concept of execution lineage. By leveraging explicit dependency tracking, identity-based identification of intermediate artifacts, and an identity-aware replay mechanism, this approach achieves deterministic computation graphs in AI agents for the first time. The method guarantees precise change propagation, zero contamination across unrelated branches, and preserves both upstream stability and cross-artifact consistency. Evaluated on a policy memo updating task, DAG-based replay attains 100% fidelity in final outputs, substantially outperforming iterative baseline approaches.
This work addresses the lack of auditable construction records in existing code generation models, which hinders error tracing and localized repair. The authors propose a contract-annotated task graph approach that simultaneously outputs code and a responsibility-labeled construction trace, binding each task’s implementation, provenance, verification evidence, and intervention history. Upon verification failure, a conservative locator maps evidence to specific graph nodes or dependency branches, enabling bounded repair only in affected regions while freezing and reusing the rest. This is the first method to treat responsibility-preserving task graphs as a unified output structure for both code generation and repair, establishing full traceability from decisions to code. Experiments show pass@1 rates of 82.5–83.0% on APPS and 75.0–82.0% on ClassEval; auditing reveals a task-to-code trace coverage of 0.9725, with 26 of 60 failure cases correctly localized and 17 successfully repaired to pass validation.
This work addresses the heightened risk of silent data corruption (SDC) in large-scale supercomputing clusters, where existing replication-based fault-tolerance mechanisms struggle to accommodate asynchronous many-task (AMT) runtimes that support dynamic task generation and work stealing. The authors propose a lightweight SDC detection and recovery mechanism tailored for nested fork-join programs. By recording the task dependency tree, comparing results, and performing a top-down identification of corrupted tasks, the approach selectively re-executes only the affected tasks while reusing correct results from their subtasks. This method achieves precise, localized recovery under dynamic scheduling for the first time, substantially reducing fault-tolerance overhead. Experimental results demonstrate negligible detection and recovery costs, correctness guarantees, and extensibility to future-based task models.
This work addresses the persistent degradation of agent reasoning and tool use caused by erroneous memories—such as contamination, staleness, or misattribution—which existing approaches struggle to correct without discarding valid knowledge. The paper formalizes, for the first time, the post-failure memory recovery problem and introduces a dependency-guided rollback repair mechanism. By constructing a typed memory-action dependency graph, the method tracks downstream effects at runtime, selectively deactivates unreliable memories, and replays only those computations relevant to the final answer. Evaluated on a controlled benchmark of 150 cases, the approach achieves an 85.3% recovery rate—surpassing the best baseline (77.3%)—while fully eliminating error sources and preserving all benign memories. In 50 stress-test scenarios, it attains a 68.0% recovery rate and significantly outperforms baselines, achieving the highest statement invalidation F1 score of 0.669.
This work addresses the inefficiency of general-purpose language agents in self-repair, which often stems from a lack of fine-grained failure diagnosis, leading to blind context expansion and conflation of distinct error types. To overcome this, the authors propose DARC, a novel framework that prioritizes diagnosis before repair: it first analyzes failure patterns across a task family using a development set, selects appropriate repair interventions, and employs a validator to freeze the optimal success-cost strategy, thereby enforcing a causal “diagnose-then-repair” workflow. By designing recovery-oriented interfaces that integrate failure mode analysis, pruning of a shared repair library, and strategy freezing, DARC significantly improves task success rates while reducing interaction steps or retrieval overhead across diverse environments—including ALFWorld, AppWorld, and XBRL Finance—outperforming both standard foundation agents and existing general-purpose repair methods.