From Perception to Action: Spatial AI Agents and World Models

📅 2026-02-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the fragmented treatment of agent architectures and spatial intelligence in existing research, which lacks a unified framework integrating perception, reasoning, and physical action—thereby limiting the effectiveness of embodied agents in real-world 3D environments. Through a systematic review of over 2,000 papers, this work proposes the first triaxial taxonomy that explicitly distinguishes spatial embodiment (geometric and physical) from symbolic embodiment, and constructs an analytical framework combining graph neural networks (GNNs), large language models (LLMs), and world models. The research highlights the critical roles of hierarchical memory, GNN–LLM synergy, and world models in cross-scale spatial tasks, yielding three core insights and identifying six key challenges. These contributions establish a standardized evaluation benchmark and chart a roadmap for future advancements in robotics, autonomous driving, and geospatial intelligence.

Technology Category

Cognitive Modeling & Cognitive Systems: Agent ArchitecturesIntelligent Robots: Embodied AIKnowledge Representation and Reasoning: Geometric, Spatial, and Temporal Reasoning

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSearch and Retrieval-Augmented AI: Agentic search
📝 Abstract
While large language models have become the prevailing approach for agentic reasoning and planning, their success in symbolic domains does not readily translate to the physical world. Spatial intelligence, the ability to perceive 3D structure, reason about object relationships, and act under physical constraints, is an orthogonal capability that proves important for embodied agents. Existing surveys address either agentic architectures or spatial domains in isolation. None provide a unified framework connecting these complementary capabilities. This paper bridges that gap. Through a thorough review of over 2,000 papers, citing 742 works from top-tier venues, we introduce a unified three-axis taxonomy connecting agentic capabilities with spatial tasks across scales. Crucially, we distinguish spatial grounding (metric understanding of geometry and physics) from symbolic grounding (associating images with text), arguing that perception alone does not confer agency. Our analysis reveals three key findings mapped to these axes: (1) hierarchical memory systems (Capability axis) are important for long-horizon spatial tasks. (2) GNN-LLM integration (Task axis) is a promising approach for structured spatial reasoning. (3) World models (Scale axis) are essential for safe deployment across micro-to-macro spatial scales. We conclude by identifying six grand challenges and outlining directions for future research, including the need for unified evaluation frameworks to standardize cross-domain assessment. This taxonomy provides a foundation for unifying fragmented research efforts and enabling the next generation of spatially-aware autonomous systems in robotics, autonomous vehicles, and geospatial intelligence.
Problem

Research questions and friction points this paper is trying to address.

spatial intelligence
embodied agents
world models
agentic reasoning
spatial grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatial AI Agents
World Models
Spatial Grounding
GNN-LLM Integration
Three-axis Taxonomy
💼 Related Jobs
No related jobs found.
G
Gloria Felicia
AtlasPro AI
N
Nolan Bryant
AtlasPro AI
H
Handi Putra
AtlasPro AI
A
Ayaan Gazali
AtlasPro AI
E
Eliel Lobo
AtlasPro AI
E
Esteban Rojas
AtlasPro AI