Score
Designs and implements methods that decompose complex, long-horizon goals into ordered or DAG-structured subtasks and modular agent pipelines, including algorithms to learn or train decomposition policies and represent dependencies among subtasks. Builds decomposition-guided retrieval and skill-selection mechanisms that identify ready (frontier) subtasks, chain subtask outputs into final synthesis, retain and pass critical observations, and handle sequential learning, noise filtering, and subtask-aware execution.
In the era of large language models, agent workflows face critical challenges in scalability, controllability, and security. To address these, this paper presents a systematic literature review and proposes, for the first time, a dual-dimensional taxonomy—spanning functional capabilities (task planning, multi-agent collaboration, tool integration) and architectural characteristics (role definition, orchestration process, specification languages). Through comparative analysis of over twenty representative academic and industrial systems, we identify recurring design patterns and persistent technical bottlenecks. We further introduce security-enhanced orchestration optimization strategies and pinpoint core gaps, including the lack of standardization and insufficient multimodal integration. This work establishes a foundational theoretical framework and practical guidelines for the design, evaluation, and evolution of agent workflows, advancing the field toward structured, trustworthy, and multimodal-cooperative paradigms.
This study addresses the challenges of cumulative error propagation in planning and the difficulty of retrieving from large-scale tool libraries when agents execute long-horizon tasks. To this end, we propose a tool-aware recursive decomposition method. The core innovation lies in constructing a hierarchical structure of tool capabilities to recursively decompose complex tasks into subtask trees. This architecture enables dynamic alignment between subtasks and available tools alongside on-demand retrieval, thereby optimizing context utilization efficiency and effectively mitigating error propagation. Experimental results demonstrate that the proposed approach improves the end-to-end success rate by up to 40 percentage points over baseline methods on complex real-world tasks.
This paper addresses three key challenges in cooperative multi-agent reinforcement learning: (1) heavy reliance on human priors for task decomposition, (2) low sample efficiency, and (3) opaque credit assignment. We propose the first end-to-end, model-agnostic framework for learning task decomposition. Our method automatically discovers an optimal symbolic task decomposition—formalized as a Reward Machine—directly from environment interactions, dynamically partitioning the global task into assignable subtasks while jointly optimizing agent policies. Key contributions include: (1) fully automated discovery of task structure without manual specification of decomposition hierarchies; (2) integration of task-conditioned neural architectures with the formal semantics of Reward Machines, thereby ensuring both policy generalizability and interpretable, semantically grounded credit assignment; and (3) substantial improvements in sample efficiency and convergence stability across multiple benchmarks featuring strongly coupled agent dynamics.
To address coarse-grained task decomposition and rigid collaboration mechanisms in single-agent systems for complex tasks, this paper proposes a modular multi-agent architecture grounded in large language models (LLMs). Our approach enables fine-grained, semantically driven hierarchical task parsing, constraint-aware subtask partitioning, and environment-feedback-guided dynamic scheduling. Key innovations include a global consistency preservation mechanism, lightweight communication routing, and a collaboration balancing algorithm—collectively reducing communication overhead while enhancing policy adaptability. Extensive experiments demonstrate that our method significantly outperforms existing baselines across multiple metrics: task success rate, decomposition efficiency, subtask coverage, and collaboration balance. Moreover, the system exhibits markedly improved overall robustness and execution stability.
Existing large language model agents struggle to effectively decompose complex tasks, retrieve appropriate skills, and generate executable multi-skill composition plans. This work proposes SkillWeaver, a framework comprising a three-stage pipeline—task decomposition, skill retrieval, and dependency-aware DAG planning—and introduces the first iterative Skill-Aware Decomposition (SAD) mechanism, which leverages a retrieval feedback loop to enhance alignment between subtasks and the skill repository. The study also constructs CompSkillBench, the first benchmark dedicated to compositional skills. Experimental results demonstrate that a single SAD iteration improves decomposition accuracy from 51.0% to 67.7% while reducing context consumption by over 99%, and achieves a 35.6% relative planning gain on unseen skill categories.
This work addresses the challenges faced by long-horizon agents in complex tasks, where entangled global context leads to high cognitive load, rapid error propagation, and costly recovery. To mitigate these issues, the paper proposes a Task-Decoupled Planning (TDP) framework that introduces, for the first time, a training-agnostic task decoupling mechanism. Specifically, a supervisor decomposes the task into a directed acyclic graph of subgoals, enabling a planner and executor to perform localized reasoning and replanning within bounded scopes, thereby isolating subtasks. This approach significantly enhances robustness and efficiency in long-horizon settings, outperforming strong baselines on TravelPlanner, ScienceWorld, and HotpotQA while reducing token consumption by up to 82%.
本文针对多技能任务中环境反馈稀疏延迟的问题,提出RLDS方法,通过分解子任务奖励来优化策略更新。
This work addresses the challenge that existing large language models struggle to decouple planning from execution in deep research tasks, leading to ambiguous credit assignment and optimization difficulties. To resolve this, the authors propose DecomposeR, a novel framework that explicitly models research plans as typed directed acyclic graphs (DAGs) and employs a two-stage reinforcement learning approach to separately optimize a planner and an answerer. In the first stage, the planner is trained to generate query decompositions and corresponding DAG structures; in the second, the answerer executes retrieval and synthesis following the graph. The method introduces fine-grained rewards targeting planner tokens and graph components, enabling effective decoupled optimization. Evaluated on standard long-form question answering benchmarks, DecomposeR outperforms strong open-source baselines by 5.1–8.0 points, significantly improving both planning quality and final answer accuracy.
This work addresses a key limitation in traditional decomposition-based program synthesis: reliance on ground-truth subgoals that ignore the solver’s actual capabilities, often yielding logically correct but practically unsolvable subtasks. To overcome this, the authors propose Solver-Aware Decomposition (SAD), a framework that retains supervision from ground-truth subgoals while incorporating feedback signals from a frozen synthesizer to refine the decomposer. This approach reveals, for the first time, that decomposition quality depends more on the solver than on the task itself, demonstrating that ground-truth subgoals are not universally optimal. By combining supervised learning with reinforcement learning—using synthesizer loss as a reward signal—the method guides decomposition toward subgoals better aligned with the solver’s strengths. Experiments in two programming domains show substantial gains in synthesis success rate and end-to-end accuracy, even solving tasks previously intractable to existing methods.
Existing large language model agent systems struggle to meet the demands of production environments—such as simplicity, controllability, and predictable inference costs—due to their high complexity, unbounded reasoning expenses, and unpredictable behavior. To address these limitations, this work proposes a practical, utility-driven agent design framework that employs “pseudo-tools” to enforce modularity, replaces dynamic planning with fixed workflows, and integrates a dedicated learning algorithm to jointly optimize component performance. The approach innovatively applies multi-objective optimization to balance inference cost and response quality, while supporting result fusion across multiple systems. Experimental results demonstrate that the proposed method significantly reduces inference costs and improves accuracy across diverse tasks, outperforming handcrafted dynamic planning baselines.
This work addresses the scarcity of high-quality, verifiable, and diverse task data that hinders large-scale training of terminal-based intelligent agents. Existing synthetic approaches often suffer from a disconnect between task generation and execution and rely heavily on pre-existing repositories, limiting diversity and scalability. To overcome these limitations, the authors propose modeling the task synthesis process itself as a Terminal-Bench–formatted terminal task, enabling closed-loop iterative generation, execution, and validation within real containerized environments. Their method enhances diversity and realism through multi-stage task specification, decoupling of task dimensions, and augmentation with external materials, while employing an LLM-as-Judge mechanism for quality filtering. Using only 3,221 synthesized trajectories for fine-tuning, Qwen3-14B and Qwen3-32B achieve Avg Pass@1 scores of 22.5% and 31.8%, respectively, on Terminal-Bench 2.0—significantly outperforming concurrent methods with substantially less training data.