OntoPlan: An Ontology-Grounded Scene Representation and Agentic Framework for Scalable Robot Task Planning

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of large language model (LLM) agents in executing long-horizon tasks within large-scale environments, including inadequate spatial understanding, excessive token overhead, and difficulty satisfying action constraints. To this end, we propose a framework that integrates an ontological knowledge graph with an agentic workflow. The core innovation lies in constructing an ontology-aligned symbolic spatial reasoning mechanism, enabling an end-to-end closed loop encompassing instruction parsing, information retrieval, goal formalization, and executable plan generation. Experimental results demonstrate that the proposed method achieves an average success rate of 0.89 while reducing token consumption by 5.6 times. It significantly enhances planning robustness and efficiency in large-scale scenarios and effectively handles ambiguous or infeasible instructions.
📝 Abstract
Large language model (LLM)-based robot task planning is promising for open-ended instruction following, but degrades on long-horizon tasks in large environments. When spatial information is conveyed to the LLM through text, the model can fail to capture spatial context, and token cost grows with environment size. Generating action sequences directly with an LLM also makes it difficult to satisfy the current world state and action preconditions. We address this with an ontology-grounded scene representation that aligns objects, spaces, relations, and states in a shared symbolic vocabulary for spatial reasoning and task planning, and with OntoPlan, an agentic framework that interprets instructions, selectively retrieves task-relevant information, formalizes goals and constraints, and produces executable plans. Across 150 general tasks spanning five indoor environments and three scene scales, OntoPlan achieves 0.89 average task success, compared with 0.27 for the strongest baseline, while using 18.1k total tokens per task on average, about 5.6$\times$ fewer than the most efficient baseline. These advantages persist as scene scale increases, whereas prior methods degrade more sharply in success and remain far more costly in tokens. OntoPlan also responds appropriately to ambiguous or infeasible instructions by asking follow-up questions or reporting insufficient information rather than committing to invalid plans. Code available at https://github.com/namhyeongwoo/OntoPlan.
Problem

Research questions and friction points this paper is trying to address.

robot task planning
large language models
spatial reasoning
long-horizon tasks
scalability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Ontology-grounded Scene Representation
Agentic Framework
Robot Task Planning
Spatial Reasoning
Large Language Models
🔎 Similar Papers
No similar papers found.
H
Hyeongwoo Nam
School of Mechanical Engineering, Yonsei University, Seoul, Republic of Korea
W
Woongje Cho
School of Mechanical Engineering, Yonsei University, Seoul, Republic of Korea
J
Juwon Kim
School of Mechanical Engineering, Yonsei University, Seoul, Republic of Korea
Jongeun Choi
Jongeun Choi
Professor of Mechanical Engineering, Yonsei University
Machine LearningRobot LearningSystems and ControlAI in Healthcare