Score
Designs and constructs directed acyclic task graphs that decompose high-level tasks into atomic subtasks and explicitly encode subtask dependencies and ordering constraints. Builds planning and analysis methods to sequence and schedule those subtasks across stages, trace graph evolution during decomposition and execution, and reason about correctness and resource dependencies.
This work addresses the lack of a unified framework in current LLM agent workflows, which hinders method comparison and reproducibility. To resolve this, we propose the Agent Computation Graph (ACG) framework, which models workflows as computation graphs and adopts “structure determines timing” as a core principle. The framework explicitly distinguishes between reusable templates, runtime instance graphs, and execution traces, enabling a systematic categorization of static and dynamic optimization approaches. Through a comprehensive literature review and conceptual modeling, we develop a multidimensional evaluation framework that integrates structural properties, establishes precise terminology, and defines standardized evaluation criteria. This foundation supports a reproducible and highly comparable research paradigm for optimizing LLM agent workflows.
Existing large language model agents tackling multi-step tasks often rely on costly recomputation or task-specific fine-tuning, resulting in poor generalization and limited reusability of intermediate results. This work proposes the Atomic Task Graph (ATG) framework, which— for the first time—explicitly models task decomposition and execution dependencies using a unified directed acyclic graph. During planning, ATG recursively decomposes high-level tasks; during execution, it enables parallel scheduling and local backtracking for error recovery. Notably, ATG operates effectively across diverse tasks without any training, achieving significant performance gains over strong baselines on three interactive benchmarks using only lightweight 7B–8B parameter models, while simultaneously improving both task success rates and execution efficiency.
ROS 2’s publish-subscribe model lacks native support for enforcing priority and data-dependency constraints in directed acyclic graph (DAG)–structured tasks, resulting in out-of-order callback execution, inconsistent multi-input matching policies, and DAG semantics sustained solely through ad hoc programming conventions—rendering systems prone to instability and crashes. To address this, we propose the Function-as-Subtask (FasS) API: a declarative interface that explicitly models data flow via function parameters and return values, thereby enforcing DAG structure at the API level and eliminating reliance on developer discipline. We implement a native DAG-aware scheduler in Rust and design a system integration layer compatible with Linux’s sched_ext subsystem. Experimental evaluation demonstrates that FasS guarantees semantic fidelity while delivering a production-ready, real-time–capable DAG scheduling infrastructure.
This work addresses the limitation of existing research agents that oversimplify complex scientific projects into single tasks, resulting in ambiguous task boundaries, disorganized execution, and missing deliverables—challenges that hinder long-horizon, multi-objective, and dependency-sensitive research planning. To overcome this, the authors propose a graph-guided, project-level planning approach that explicitly decomposes a research project into executable task compositions with clearly attributed contributions and explicit dependencies, leveraging an innovative atomic representation and a directed provenance graph. A lightweight Bernoulli block model optimizes task selection, generating standardized task contracts that specify objectives, dependencies, and constraints, enabling seamless decoupled integration with arbitrary executors. Evaluated on ten scientific benchmarks, the method achieves an average quality score of 7.15, significantly outperforming baselines (4.58 and 5.31), and when integrated with AutoResearchClaw, boosts downstream task accuracy from 0.536 to 0.759.
In multicore systems, cache coherence and task execution are deeply intertwined, yet existing task-graph modeling approaches either rely on predefined structures or target specific schedulers, commonly neglecting coherence interactions—leading to a mismatch between design assumptions and runtime behavior. This work introduces CoTAM, the first framework to explicitly model how cache coherence affects task dependencies. CoTAM decouples coherence effects via runtime behavioral analysis and employs a data-driven learning mechanism to dynamically infer weighted task dependencies, thereby generating coherence-aware, general-purpose task graphs. Experimental results demonstrate that CoTAM significantly outperforms implicit modeling methods, improving task-graph accuracy and adaptability under dynamic workloads, and effectively bridging the semantic gap between system-level design abstractions and actual runtime execution.
Long-horizon collaborative tasks for dual robotic arms face challenges including complex spatiotemporal dependencies among subtasks, difficulty in dynamic action allocation, and limited expressiveness of linear programming formulations. This paper proposes the first LLM-driven DAG-structured task decomposition framework, which automatically parses high-level instructions into directed acyclic graphs (DAGs) encoding dependency constraints, and integrates environment perception to enable real-time, dynamic action allocation and parallel adaptive execution across both arms. The method breaks away from predefined operational paradigms, supporting end-to-end, interpretable, and generalizable collaborative planning. Evaluated on the Dual-Arm Kitchen benchmark, it achieves a 52.8% efficiency gain over single-arm systems, improves success rate by 48% and reduces LLM query count by 84.1% compared to conventional dual-arm planners, significantly enhancing robustness and scalability in complex scenarios.
Current prompt graphs lack a clear definition and standardized terminology, resulting in conceptual ambiguity in practice. This work addresses this gap by proposing a formal definition of prompt graph engineering through conceptual analysis, gray literature review, and systematic categorization. It identifies prompt graphs as first-class, executable, and improvable engineering artifacts and establishes four necessary constitutive conditions along with inclusion and exclusion criteria for operational validation. The proposed definition demonstrates consistent applicability across six major frameworks—including LangGraph and DSPy—thereby offering the field its first operational framework and shared vocabulary. Building on this foundation, the paper outlines a future research agenda structured around four key design tensions inherent to prompt graph development.
This work proposes the first arbitrarily scalable and automatically verifiable task-graph benchmark designed to evaluate language agents’ ability to retain, update, combine, and discard contextual information during complex reasoning. The benchmark constructs task graphs from natural language questions paired with executable Python solvers, modeling tasks through typed intermediate states such as scalars and lists. It enables flexible control over task length, dependency structure, distractors, and value types. Experimental results reveal that while Qwen3.5-27B excels on isolated tasks, its accuracy drops by up to 33.3% on complex tasks involving branching dependencies, effectively exposing a critical bottleneck in current agents’ context management capabilities.
This work addresses the challenge that agents struggle to efficiently reuse successful experiences when repeatedly performing similar tasks, often resulting in redundant reasoning and excessive interaction rounds. To overcome this limitation, the paper introduces a novel framework that formalizes procedural skills as parameterized finite state machine (PFSM) subgraphs and automatically extracts, verifies, and reuses structured skills through distillation and compilation of successful execution trajectories. Evaluated on the ALFWorld and WebArena benchmarks, the proposed method significantly improves task success rates while reducing the number of required interactions, demonstrating effectiveness across language models of varying scales.