Institution profile

ADIA

Research institution
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution

Sep 29, 2026

This study addresses the disconnect between planning declarations and execution in large language model (LLM) agents, where distinguishing planning errors from execution failures remains challenging. To this end, it proposes a "Planning as Routing" framework. The research reveals that the standard ReAct paradigm fails to maintain planning structures, prompting the introduction of a novel routing mechanism that decouples plan selection from execution. Specifically, an LLM declares one of four predefined planning modes—sequential, hierarchical, or search-based—which a deterministic router then dispatches to dedicated executors to ensure faithful implementation. Experimental results demonstrate that this approach increases success rates to 92% on ALFWorld and 44% on SWE-bench, substantially bridging the gap between planning and execution.

0 citationsRead paper
Recent publications

Latest Papers

Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution

Sep 29, 2026

This study addresses the disconnect between planning declarations and execution in large language model (LLM) agents, where distinguishing planning errors from execution failures remains challenging. To this end, it proposes a "Planning as Routing" framework. The research reveals that the standard ReAct paradigm fails to maintain planning structures, prompting the introduction of a novel routing mechanism that decouples plan selection from execution. Specifically, an LLM declares one of four predefined planning modes—sequential, hierarchical, or search-based—which a deterministic router then dispatches to dedicated executors to ensure faithful implementation. Experimental results demonstrate that this approach increases success rates to 92% on ALFWorld and 44% on SWE-bench, substantially bridging the gap between planning and execution.

0 citationsRead paper