Score
Designs and implements methods that decompose complex, long-horizon goals into ordered or DAG-structured subtasks and modular agent pipelines, including algorithms to learn or train decomposition policies and represent dependencies among subtasks. Builds decomposition-guided retrieval and skill-selection mechanisms that identify ready (frontier) subtasks, chain subtask outputs into final synthesis, retain and pass critical observations, and handle sequential learning, noise filtering, and subtask-aware execution.
In the era of large language models, agent workflows face critical challenges in scalability, controllability, and security. To address these, this paper presents a systematic literature review and proposes, for the first time, a dual-dimensional taxonomy—spanning functional capabilities (task planning, multi-agent collaboration, tool integration) and architectural characteristics (role definition, orchestration process, specification languages). Through comparative analysis of over twenty representative academic and industrial systems, we identify recurring design patterns and persistent technical bottlenecks. We further introduce security-enhanced orchestration optimization strategies and pinpoint core gaps, including the lack of standardization and insufficient multimodal integration. This work establishes a foundational theoretical framework and practical guidelines for the design, evaluation, and evolution of agent workflows, advancing the field toward structured, trustworthy, and multimodal-cooperative paradigms.
This work addresses the lack of a unified framework in current LLM agent workflows, which hinders method comparison and reproducibility. To resolve this, we propose the Agent Computation Graph (ACG) framework, which models workflows as computation graphs and adopts “structure determines timing” as a core principle. The framework explicitly distinguishes between reusable templates, runtime instance graphs, and execution traces, enabling a systematic categorization of static and dynamic optimization approaches. Through a comprehensive literature review and conceptual modeling, we develop a multidimensional evaluation framework that integrates structural properties, establishes precise terminology, and defines standardized evaluation criteria. This foundation supports a reproducible and highly comparable research paradigm for optimizing LLM agent workflows.
This paper addresses three key challenges in cooperative multi-agent reinforcement learning: (1) heavy reliance on human priors for task decomposition, (2) low sample efficiency, and (3) opaque credit assignment. We propose the first end-to-end, model-agnostic framework for learning task decomposition. Our method automatically discovers an optimal symbolic task decomposition—formalized as a Reward Machine—directly from environment interactions, dynamically partitioning the global task into assignable subtasks while jointly optimizing agent policies. Key contributions include: (1) fully automated discovery of task structure without manual specification of decomposition hierarchies; (2) integration of task-conditioned neural architectures with the formal semantics of Reward Machines, thereby ensuring both policy generalizability and interpretable, semantically grounded credit assignment; and (3) substantial improvements in sample efficiency and convergence stability across multiple benchmarks featuring strongly coupled agent dynamics.
To address coarse-grained task decomposition and rigid collaboration mechanisms in single-agent systems for complex tasks, this paper proposes a modular multi-agent architecture grounded in large language models (LLMs). Our approach enables fine-grained, semantically driven hierarchical task parsing, constraint-aware subtask partitioning, and environment-feedback-guided dynamic scheduling. Key innovations include a global consistency preservation mechanism, lightweight communication routing, and a collaboration balancing algorithm—collectively reducing communication overhead while enhancing policy adaptability. Extensive experiments demonstrate that our method significantly outperforms existing baselines across multiple metrics: task success rate, decomposition efficiency, subtask coverage, and collaboration balance. Moreover, the system exhibits markedly improved overall robustness and execution stability.
Existing large language model agents struggle to effectively decompose complex tasks, retrieve appropriate skills, and generate executable multi-skill composition plans. This work proposes SkillWeaver, a framework comprising a three-stage pipeline—task decomposition, skill retrieval, and dependency-aware DAG planning—and introduces the first iterative Skill-Aware Decomposition (SAD) mechanism, which leverages a retrieval feedback loop to enhance alignment between subtasks and the skill repository. The study also constructs CompSkillBench, the first benchmark dedicated to compositional skills. Experimental results demonstrate that a single SAD iteration improves decomposition accuracy from 51.0% to 67.7% while reducing context consumption by over 99%, and achieves a 35.6% relative planning gain on unseen skill categories.
This work addresses the challenge that existing large language models struggle to decouple planning from execution in deep research tasks, leading to ambiguous credit assignment and optimization difficulties. To resolve this, the authors propose DecomposeR, a novel framework that explicitly models research plans as typed directed acyclic graphs (DAGs) and employs a two-stage reinforcement learning approach to separately optimize a planner and an answerer. In the first stage, the planner is trained to generate query decompositions and corresponding DAG structures; in the second, the answerer executes retrieval and synthesis following the graph. The method introduces fine-grained rewards targeting planner tokens and graph components, enabling effective decoupled optimization. Evaluated on standard long-form question answering benchmarks, DecomposeR outperforms strong open-source baselines by 5.1–8.0 points, significantly improving both planning quality and final answer accuracy.
This work addresses the challenges faced by long-horizon agents in complex tasks, where entangled global context leads to high cognitive load, rapid error propagation, and costly recovery. To mitigate these issues, the paper proposes a Task-Decoupled Planning (TDP) framework that introduces, for the first time, a training-agnostic task decoupling mechanism. Specifically, a supervisor decomposes the task into a directed acyclic graph of subgoals, enabling a planner and executor to perform localized reasoning and replanning within bounded scopes, thereby isolating subtasks. This approach significantly enhances robustness and efficiency in long-horizon settings, outperforming strong baselines on TravelPlanner, ScienceWorld, and HotpotQA while reducing token consumption by up to 82%.
This work addresses a key limitation in traditional decomposition-based program synthesis: reliance on ground-truth subgoals that ignore the solver’s actual capabilities, often yielding logically correct but practically unsolvable subtasks. To overcome this, the authors propose Solver-Aware Decomposition (SAD), a framework that retains supervision from ground-truth subgoals while incorporating feedback signals from a frozen synthesizer to refine the decomposer. This approach reveals, for the first time, that decomposition quality depends more on the solver than on the task itself, demonstrating that ground-truth subgoals are not universally optimal. By combining supervised learning with reinforcement learning—using synthesizer loss as a reward signal—the method guides decomposition toward subgoals better aligned with the solver’s strengths. Experiments in two programming domains show substantial gains in synthesis success rate and end-to-end accuracy, even solving tasks previously intractable to existing methods.
This work addresses the limited planning generalization of large language model (LLM) agents in unseen scenarios by proposing a dynamic policy learning framework that integrates generalized planning with hierarchical task decomposition. The approach automatically extracts and reuses parameterized policy components from successful executions to construct a composable policy library. Central to the method are hierarchical component learning (HCL-GP), semantic-driven policy retrieval, and a dynamic reuse mechanism that enables cross-task knowledge transfer. Evaluated on the AppWorld benchmark, the proposed method achieves task success rates of 98.2% on standard tasks and 97.8% on challenging ones—representing a 15.8 percentage point improvement over static composition. Notably, it elevates the success rate of open-source LLM agents from near zero to 62.5%, substantially enhancing their task generalization capabilities.
Existing large language model agent systems struggle to meet the demands of production environments—such as simplicity, controllability, and predictable inference costs—due to their high complexity, unbounded reasoning expenses, and unpredictable behavior. To address these limitations, this work proposes a practical, utility-driven agent design framework that employs “pseudo-tools” to enforce modularity, replaces dynamic planning with fixed workflows, and integrates a dedicated learning algorithm to jointly optimize component performance. The approach innovatively applies multi-objective optimization to balance inference cost and response quality, while supporting result fusion across multiple systems. Experimental results demonstrate that the proposed method significantly reduces inference costs and improves accuracy across diverse tasks, outperforming handcrafted dynamic planning baselines.
This work addresses the scarcity of high-quality, verifiable, and diverse task data that hinders large-scale training of terminal-based intelligent agents. Existing synthetic approaches often suffer from a disconnect between task generation and execution and rely heavily on pre-existing repositories, limiting diversity and scalability. To overcome these limitations, the authors propose modeling the task synthesis process itself as a Terminal-Bench–formatted terminal task, enabling closed-loop iterative generation, execution, and validation within real containerized environments. Their method enhances diversity and realism through multi-stage task specification, decoupling of task dimensions, and augmentation with external materials, while employing an LLM-as-Judge mechanism for quality filtering. Using only 3,221 synthesized trajectories for fine-tuning, Qwen3-14B and Qwen3-32B achieve Avg Pass@1 scores of 22.5% and 31.8%, respectively, on Terminal-Bench 2.0—significantly outperforming concurrent methods with substantially less training data.
This work addresses the limitations of existing reinforcement learning–based task decomposition approaches, which often suffer from reward hacking—such as repetitive subtask generation—due to the direct use of retrieval metrics as rewards, leading to poor out-of-domain generalization in tool-augmented settings. To mitigate this, the authors propose a preference-guided counterfactual task decomposition framework that introduces counterfactual causal reasoning into task decomposition for the first time. By employing counterfactual rewards to quantify the causal contribution of each decomposition step to retrieval performance, the method severs spurious correlations between superficial features and retrieval metrics. Additionally, a structured preference reward provides fine-grained supervision over the logical coherence and atomicity of decomposed steps. Experiments on the newly introduced mobile multi-turn interaction benchmark, MTDTool, demonstrate that the proposed approach substantially alleviates repetitive decomposition and consistently outperforms state-of-the-art methods in retrieval effectiveness, decomposition quality, and out-of-domain generalization.