Score
Design, build, or analyze systems that decompose complex, nested, or multi-step user instructions into an ordered sequence of actionable subgoals or steps and produce explicit stepwise plans or annotations indicating the currently active step. Ensure the decomposition maintains plan coherence across steps and long horizons, respects nested constraints and conditions, and supports mechanisms (e.g., attention alignment or subgoal labels) that connect model internals to the active subgoal for downstream planners or executors.
Large language models (LLMs) exhibit insufficient reasoning capabilities and low success rates in long-horizon, embodied task planning for real-world robots. Method: We propose a closed-loop hierarchical subgoal planning framework that constructs a cross-level subgoal tree: a base LLM performs coarse-grained task decomposition, while an environment-state-driven leaf-node termination model dynamically assesses subgoal completion and triggers the next-level planning, thereby closing the perception–planning–execution loop. Contribution/Results: Our key innovation lies in decoupling task decomposition from execution termination—enabling adaptive, verifiable hierarchical planning. Evaluated on the VirtualHome WAH-NL benchmark and a physical robot platform, our approach achieves 34% and 25% success rates, respectively, substantially outperforming prior methods.
Existing task decomposition methods for large language models predominantly focus on tool invocation and feedback mechanisms, overlooking the critical trade-off between performance and computational cost. Method: This paper proposes a “select-then-decompose” paradigm, establishing a closed-loop framework comprising three stages: selection, execution, and verification. First, we systematically categorize six decomposition patterns and identify key task features governing performance and cost. Second, we design a dynamic selection mechanism that adaptively identifies the optimal decomposition strategy per task. Third, we integrate a lightweight verification module to ensure result reliability. Contribution/Results: Our approach consistently achieves Pareto-optimal performance across multiple benchmarks—significantly improving inference efficiency while reducing computational overhead compared to fixed decomposition strategies. The implementation is publicly available.
This work investigates the role of task decomposition in program synthesis, examining its impact on generalization and subgoal validity by comparing the explicit decomposition framework ExeDec against the decomposition-free, execution-driven framework REGISM. Method: We propose a novel synthesis paradigm that jointly optimizes iterative execution feedback, subgoal modeling, and code generation, and introduce a cross-task generalization evaluation framework. Contribution/Results: (1) Execution-driven learning alone serves as a critical performance driver; (2) explicit decomposition substantially improves length generalization and compositional concept learning; (3) despite lacking explicit decomposition, REGISM matches or surpasses ExeDec across multiple metrics, and its implicit decomposition aligns more closely with human-annotated subgoal structures. Our study is the first to reveal the fundamental trade-off—decomposition is not strictly necessary but can yield measurable gains—thereby offering a new perspective on task structure modeling in program synthesis.
This study challenges the prevailing assumption that step-by-step monitoring is essential for adaptive performance in data-intensive tasks by systematically investigating planning horizon as an independent variable. Through controlled experiments, the authors compare full-horizon (FH) planning against single-horizon (SH) planning in knowledge-base question answering and multi-hop reasoning tasks. The results demonstrate that FH planning, augmented with on-demand replanning, achieves accuracy comparable to SH planning across varying task depths, breadths, and tool robustness conditions—while reducing token consumption by a factor of 2–3. These findings question the necessity of continuous, fine-grained monitoring in structured data tasks and suggest that broader planning horizons can yield substantial efficiency gains without compromising performance.
To address the challenges of zero-shot planning difficulty and weak dynamic adaptability in Embodied Instruction Following (EIF), this paper proposes the first training-free, vision-driven zero-shot embodied planning framework. Methodologically, it decomposes natural language instructions into executable high-level sub-goal sequences via a self-questioning-and-answering mechanism; introduces a visually grounded real-time re-planning mechanism that dynamically refines plans based on visual feedback during interaction; and designs RelaxedHLP—a novel evaluation metric that quantifies high-level planning quality for the first time. Experiments on the ALFRED benchmark demonstrate state-of-the-art zero-shot and few-shot performance, particularly on complex tasks requiring multi-step reasoning and environment responsiveness. Visual feedback significantly improves re-planning accuracy, with consistent gains over existing methods.
Automatically constructing high-quality, reusable skills from heterogeneous, fragmented interaction traces—often missing critical security behaviors—is highly challenging. This work proposes the W2S framework, which introduces a novel intermediate representation called RWSA to decouple skills into workflow structure, execution semantics, and runtime attachments, thereby enabling task decomposition, control-flow modeling, verification, rollback, and state management. W2S achieves efficient skill construction through trajectory segmentation, local skill draft generation, structural alignment, branch fusion, redundancy compression, and confidence-aware retention. Experimental evaluation across 70 skills demonstrates that W2S improves behavioral replay consistency by 10.5% compared to baseline approaches based on summarization and prompting.
This work addresses the systematic spatial reasoning errors exhibited by large language models when generating 3D structures from natural language instructions, which often manifest as coordinate inaccuracies that undermine structural reliability. To mitigate this, the authors propose a neuro-symbolic 2.5-D decomposition approach that disentangles deterministic physical constraints—such as gravity—from the language model’s output. The model is restricted to planning layouts in a 2D plane, while a symbolic executor determines vertical stacking based on column occupancy. This strategy significantly improves construction accuracy, achieving a 94.6% average structural correctness on the Build What I Mean benchmark—surpassing GPT-4o (90.3%) and the previous state-of-the-art system (76.3%). Notably, it retains 94.5% performance on Jetson Thor AGX edge hardware. Ablation studies attribute a 50.7-percentage-point accuracy gain to the proposed method, highlighting its potential for generalization to other physically constrained assembly tasks.
This work addresses the challenge of enabling effective human intervention in the planning processes of large language models (LLMs) within complex multi-agent systems, where existing approaches offer only outcome-level supervision and lack visibility into or control over intermediate reasoning. The paper presents the first systematic formulation of an interaction design space for human–LLM collaborative planning, structured along three dimensions—semantic vs. structural, global vs. local, and high-level vs. low-level—to support process-level supervision. A prototype system, AMBIPOM, is developed to instantiate diverse interaction modalities. User studies reveal a strong preference for mixed interaction strategies, while benchmark experiments quantitatively evaluate LLM plan revision efficacy under different editing strategies, uncovering a trade-off among effort, control, and risk. These findings provide both theoretical grounding and empirical evidence for transparent and controllable human–AI co-planning.
Existing language model agents struggle to efficiently execute complex instructions in long-horizon tasks due to insufficient planning capabilities. This work proposes a planner-centric multi-agent framework comprising a planner, an executor, and a memory manager. Through computational resource allocation analysis, we demonstrate that the planning component predominantly governs overall performance. Leveraging this insight, we apply reinforcement learning exclusively to the planner, incorporating trajectory-level rewards and a vision-language model-based evaluation mechanism to enable asymmetric computation allocation. The resulting approach achieves significant performance gains across diverse benchmarks—including web navigation, operating system control, and tool usage—thereby validating the efficacy and strong generalization of prioritizing high-level planning.