🤖 AI Summary
This study addresses the disconnect between planning declarations and execution in large language model (LLM) agents, where distinguishing planning errors from execution failures remains challenging. To this end, it proposes a "Planning as Routing" framework. The research reveals that the standard ReAct paradigm fails to maintain planning structures, prompting the introduction of a novel routing mechanism that decouples plan selection from execution. Specifically, an LLM declares one of four predefined planning modes—sequential, hierarchical, or search-based—which a deterministic router then dispatches to dedicated executors to ensure faithful implementation. Experimental results demonstrate that this approach increases success rates to 92% on ALFWorld and 44% on SWE-bench, substantially bridging the gap between planning and execution.
📝 Abstract
Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully. Existing planner--executor systems can fail at either stage, while final task success alone cannot distinguish selection from execution failures. We therefore study the Plan Declaration--Execution Gap and introduce Planning-as-Routing, where an LLM declares one of four planning modes: Predefined, Sequential, Hierarchical, or Search, and a deterministic router dispatches the task to the corresponding pattern-specific executor. Across four benchmarks and three LLMs, we find three consistent patterns. First, generic Plan+ReAct often fails to preserve declared planning structure, especially for longer plans: across three benchmarks, only (22)--(45%) of trajectories preserve it, whereas pattern-specific executors enforce the intended structure. Second, planning-mode effectiveness varies across environments and models: Search performs best on ALFWorld, Hierarchical on SWE-bench, and the strongest pattern can vary across models within the same benchmark. Third, the largest gains come from execution: pattern-specific executors improve task success from (0.48) to (0.92) on ALFWorld and from (0.36) to (0.44) on SWE-bench Verified over Plan+ReAct. Current LLMs, however, do not reliably select the strongest mode for each task, although few-shot examples improve selection in some benchmark--model combinations. Overall, reliable agent planning requires both effective mode selection and faithful execution: routing substantially closes the execution gap, while task-specific mode selection remains open.