Spatial Strategies, Not Actions: Vector-Quantized Geodesics as Tools for LLM-Driven Agents

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limited spatial understanding capabilities of large language model (LLM) agents by proposing a hierarchical decision-making framework that integrates geodesic geometry with LLMs. Methodologically, the learning process is decoupled into two levels: unsupervised tool discovery and online LLM reasoning. The lower level constructs a geometric behavioral repertoire by vector-quantizing geodesic trajectories, while the upper level employs Qwen3.6 to dynamically select tools and execute policies, thereby decoupling spatial planning from action execution. In grid-world simulations, the proposed approach achieves goal completion rates comparable to chain-of-thought baselines with second-level response times, substantially reducing computational overhead. This work offers an efficient new paradigm for spatial reasoning in embodied agents.
πŸ“ Abstract
Large language model (LLM) based agents are often criticized for lacking spatial understanding and mainly exploiting statistical text patterns. We investigate their spatial comprehension through an architecture combining geometrical tools with a LLM serving as a high-level orchestrator in grid-world environments. The agent first collects geodesic trajectories, which are then vector-quantized to extract a representative subset. Offline, the LLM associates a natural language description of the underlying behavioral patterns to each selected trajectory, making it a tool. Online, the LLM chooses the appropriate tool conditioned on the current state and goal. Low-level control is handled by primitive actions that execute the trajectory associated with the tool. From an agentic AI perspective, this approach separates learning into two levels: tool discovery is handled through unsupervised quantization of trajectories, while reasoning and decision-making are handled by the LLM. We test the approach in a partially observable dynamic 2D grid environment with an open vision-language model (Qwen3.6-35B-A3B). Pairing the geometry-derived tool library with an agent-centered zoom tool and a collision detection tool lets a fast, non-reasoning configuration match the goal-reaching rate of a much more costly chain-of-thought version, while cutting the cost of a decision from minutes to seconds.
Problem

Research questions and friction points this paper is trying to address.

Spatial understanding
LLM-driven agents
Grid-world navigation
Computational efficiency
Agentic AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vector-Quantized Geodesics
Tool Discovery
Spatial Strategies
LLM-Driven Agents
Two-Level Learning
πŸ”Ž Similar Papers
No similar papers found.