GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

๐Ÿ“… 2026-09-16
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
GAVEL้€š่ฟ‡ๆž„ๅปบๅ›พไธ–็•Œๆจกๅž‹ๆฅ้ชŒ่ฏๅ’Œไฟฎๅคๅคง่ฏญ่จ€ๆจกๅž‹็”Ÿๆˆ็š„้•ฟๆœŸไปปๅŠก่ฎกๅˆ’๏ผŒๆ้ซ˜ไบ†ๆœบๅ™จไบบ่ง„ๅˆ’็š„ๆˆๅŠŸ็އๅ’Œๆ•ˆ็އใ€‚
๐Ÿ“ Abstract
Large language models (LLMs) provide a flexible interface for long-horizon robot planning, but generated plans often fail to respect embodiment constraints, recover from planning errors, or reason effectively under partial observability. We present GAVEL, a framework for verifying and repairing long-horizon LLM planning built around an explicit graph world model. The graph represents relevant object-relations, action pre-conditions and effects, and probabilistic beliefs over unobserved object locations. This model can predict the consequences of LLM-generated actions before execution, detect violations, and repair those whose corrections follow directly from the world model. This method also reserves LLM replanning solely for errors requiring semantic reasoning. For multi-task instructions, GAVEL reasons over distributions of possible object locations to reorder remaining subtasks and minimize expected search cost. We evaluate GAVEL on BEHAVIOR-1K across 100 single long-horizon tasks and 500 multi-task instructions. With Qwen3-8B, GAVEL improves single-task success from 41.2% to 91.8% and multi-task success from 19.9% to 92.6%. Distributional belief reasoning also reduces travel distance by approximately 5.4% compared with a static variant. These improvements show that an explicit graph world model harness can substantially improve the reliability and efficiency of long-horizon embodied planning across compact and frontier hosted LLM capabilities.
Problem

Research questions and friction points this paper is trying to address.

Large language models
long-horizon planning
embodiment constraints
partial observability
planning errors
Innovation

Methods, ideas, or system contributions that make the work stand out.

graph world model
long-horizon planning
LLM verification and repair
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
R
Ruiyang Wang
H
Hao-Lun Hsu
S
Swarajh Mehta
Jiwoo Kim
Jiwoo Kim
์„ฑ๊ท ๊ด€๋Œ€ํ•™๊ต ์ธ๊ณต์ง€๋Šฅํ•™๊ณผ
Z
Zhihao Dou
M
Miroslav Pajic