🤖 AI Summary
This study addresses the degradation in reliability of large language models during complex task planning as task complexity increases. To this end, we propose GRASP, a strategy-aware multi-stage planning framework. This method introduces a pioneering context-isolated pipeline that decouples plan generation, revision, and evaluation modules, while incorporating a macro-regularization mechanism to mitigate performance degradation across multiple tasks. Through the synergistic optimization of GenPlan precompilation, RevPlan local policy exploration, and a VerPlan multi-criteria discriminator, the framework generates high-quality natural language execution plans. Experimental results demonstrate that GRASP establishes new state-of-the-art performance across multiple benchmark datasets, surpassing frontier reasoning models by 14.5% in accuracy.
📝 Abstract
Large Language Models (LLMs) typically exhibit a performance profile where reliability degrades as task complexity increases. We address the challenge of generating high-quality natural language executable plans for complex tasks by introducing $\textbf{GRASP}$, a strategy-aware, multi-stage planning framework. GRASP decouples the planning pipeline across specialized, context-isolated modules: it pre-compiles global macro-guidelines (GenPlan), explores alternative localized strategies within isolated context windows (RevPlan), and independently evaluates trajectories using a multi-criteria discriminator (VerPlan). Empirical evaluations show that GRASP consistently establishes a new state-of-the-art frontier across diverse datasets, yielding substantial accuracy gains over direct LLM planners on Natural Plan Calendar Scheduling ($\sim$12.4$\%$$\uparrow$), ZebraLogic ($\sim$30.8$\%$$\uparrow$), and SciBench Math. Crucially, under multi-task scaling-where standard planners suffer immediate performance collapse-GRASP completely flattens the multi-task degradation penalty. In interleaved dual-task environments, GRASP achieves an absolute accuracy gain of up to 16.7$\%$ over direct LLM planners. Furthermore, by isolating context and enforcing strict macro-regularization, GRASP outperforms frontier reasoning models (such as GPT-5-mini) by a margin of 14.5$\%$.