🤖 AI Summary
This study reevaluates the effectiveness of the large language model PlanGPT in automated planning tasks, critically examining the reliability of its original claim regarding plan coverage. Through systematic comparisons on standard planning benchmarks, the work assesses PlanGPT against classical planners and greedy search algorithms, using plan cost and generation time as primary evaluation metrics—criteria newly introduced into the PlanGPT evaluation framework. The results reveal that PlanGPT’s performance is comparable only to that of a greedy search strategy, demonstrating no significant advantage in either plan quality or computational efficiency. These findings challenge the purported practical utility of PlanGPT in automated planning and raise questions about its added value over simpler, well-established methods.
📝 Abstract
Automated Planning is a subfield of Artificial Intelligence (AI) where the main objective is generating a sequence of actions, known as a plan, that helps us reach a goal state from an initial state. A planning problem is defined by a set of objects, an initial state and a desired goal state. The objective is to compute a plan that'll lead us from the inital state to the goal state. Programs that generate plans are called planners.
In this paper, we did a complementary study to the state-of-the-art LLM called PlanGPT which was released last year. We redid some experiments to verify whether planning with LLMs is \textbf{pertinent} and \textbf{worthwhile}. We also check whether the results obtained in the official PlanGPT paper for plan coverage were correct, and we also performed a more comprehensive study on PlanGPT's performance: in our paper PlanGPT's performance was evaluated using two metrics: Plan Cost and Plan Generation Time. The results of planGPT were compared to those produced by a traditional planner for the same plans and same metrics. We discovered that PlanGPT is no better than a Greedy search strategy.