spatial tool use

Designs, builds, or analyzes systems that enable agents or models (particularly vision and vision-language models) to select, invoke, and coordinate spatially-aware tools and interfaces operating on 2D and 3D data. This includes creating model-agnostic spatial interfaces, semantic planners and tool-calling frameworks, and producing tool-use or manipulation trajectories and evidence-requesting behaviors to ground and coordinate spatial operations.

spatialtooluse

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Foundation Model Driven Robotics: A Comprehensive Review

Jul 14, 2025
MT
Muhammad Tayyab Khan
🏛️ Nanyang Technological University | Texas A&M University

This paper systematically reviews bottlenecks in deploying foundation models (LLMs/VLMs) for robotics, identifying five core challenges: insufficient real-time responsiveness, weak perception–action coupling, poor cross-domain generalization, limited robustness, and deficient human–robot trust. To address these, we propose a system-level embodied intelligence framework integrating procedural scene generation, multimodal reasoning, policy generalization, and sim-to-real co-training—thereby enforcing closed-loop alignment between semantic understanding and physical execution. We introduce the first end-to-end evaluation taxonomy spanning perception, planning, control, and interaction, and identify three critical gaps: embodied representation modeling, scarcity of high-fidelity multimodal robotic data, and rigorous safety verification. Based on this analysis, we chart a pragmatic research roadmap. Our work provides both theoretical foundations and engineering blueprints to advance foundation models from “language intelligence” toward “physical intelligence.”

Bridging semantic reasoning with physical robot intelligenceChallenges in real-world application of multimodal roboticsHow foundation models enhance robotics perception and planning

Must-Read Papers

Most classic and influential ideas
View more

Current foundation models generate 3D content that, while visually plausible, lacks operability and thus struggles to support downstream tasks such as CAD or robotics. This work proposes Hylos, a novel system architecture that introduces contract-based design into spatial intelligence for the first time. Hylos employs a “spatial transaction” mechanism to maintain scene-level operational states, ensuring object recognizability, constraint satisfaction, and action feasibility. By integrating scene graph dependency tracking, constraint solving, and capability gap detection, the system guarantees verifiability and rollback capability for spatial modifications at transaction boundaries, while enabling upstream repairs grounded in causal dependencies. Experiments demonstrate that Hylos effectively identifies and corrects structural errors, transforming generated 3D content into a reliable foundation suitable for engineering applications and interactive world construction.

3D generationfoundation modelsoperability

This study addresses the fragmented treatment of agent architectures and spatial intelligence in existing research, which lacks a unified framework integrating perception, reasoning, and physical action—thereby limiting the effectiveness of embodied agents in real-world 3D environments. Through a systematic review of over 2,000 papers, this work proposes the first triaxial taxonomy that explicitly distinguishes spatial embodiment (geometric and physical) from symbolic embodiment, and constructs an analytical framework combining graph neural networks (GNNs), large language models (LLMs), and world models. The research highlights the critical roles of hierarchical memory, GNN–LLM synergy, and world models in cross-scale spatial tasks, yielding three core insights and identifying six key challenges. These contributions establish a standardized evaluation benchmark and chart a roadmap for future advancements in robotics, autonomous driving, and geospatial intelligence.

agentic reasoningembodied agentsspatial grounding

Current vision-language models (VLMs) are constrained by rigid action interfaces that hinder their ability to balance flexibility and accuracy in open-ended, complex 3D/4D spatial reasoning tasks. This work proposes SpatialClaw, a training-free framework that pioneers code as a stateful, interactive action interface. By leveraging a preloaded Python kernel containing input frames, VLM agents dynamically generate and execute composable perception-geometry operation units over multiple steps, adapting responsively to task demands. This approach overcomes the limitations of single-step execution or fixed tool invocation, achieving an average accuracy of 59.9% across 20 spatial reasoning benchmarks—surpassing prior methods by 11.2 percentage points. Consistent performance gains are observed across six diverse VLM backbones, all without task-specific fine-tuning or model adaptation.

3D/4D reasoningaction interfacespatial reasoning

Current LLM-driven autonomous agents lack scalable, secure, and maintainable architectures for tool orchestration, hindering large-scale deployment. To address this, we propose the novel “Control Plane as a Tool” paradigm, which— for the first time—abstracts tool scheduling, security policies, and extensibility mechanisms into a unified, pluggable tool interface, thereby decoupling control logic from the LLM agent core. Leveraging modular routing protocols and production-grade encapsulation, our approach significantly reduces integration complexity while enabling dynamic tool registration/removal and runtime policy updates. Empirical evaluation across diverse application scenarios demonstrates over 40% improvement in both horizontal scalability efficiency and security robustness. The architecture provides a reusable, evolution-aware foundation for controllable, embodied intelligent agents.

Addressing infrastructural and architectural challenges in AI agentsImproving scalability, safety, and extensibility in agent designManaging tool orchestration at scale in agentic AI systems

This work addresses the limited capacity of existing vision-language models (VLMs) and tool-augmented agents to perform effective spatial reasoning in continuous, dynamic 3D environments, as they are largely confined to static, single-frame understanding. We propose S-Agent, a novel paradigm that reframes the VLM as a semantic planner driven by hierarchical tool invocation. By orchestrating a spatial toolchain that integrates 2D perception, 3D geometric reconstruction, and spatiotemporal memory, S-Agent enables cross-frame evidence accumulation and high-level spatial knowledge construction without requiring any additional training. The framework also generates high-quality spatial reasoning trajectories suitable for supervised fine-tuning. Experiments demonstrate that S-Agent substantially improves both open- and closed-source VLMs on multiview and video-based spatial reasoning benchmarks. Fine-tuning on the S-300K dataset generated by our method yields the S-Agent-8B model, which outperforms same-scale baselines and rivals advanced closed-source systems such as GPT-5.4 and Gemini 3.

continuous 3D reasoningmulti-view perceptionspatial intelligence

Latest Papers

What's happening recently
View more

This study systematically investigates how tool architecture influences the behavior and performance of coding agents, holding underlying capabilities constant. Through controlled experiments on repository-scale program repair tasks, six distinct tool interfaces—ranging from bash and structured low-level APIs to natural language search, Python CodeAct, and cognitive scaffolding—are evaluated. Analysis of 11,700 agent trajectories reveals, for the first time, that the architectural design of tools—not merely their functional capacity—plays a critical role: structured low-level interfaces improve consistency across repeated attempts by 4.7×, natural language search increases access to relevant files by over 11%, and CodeAct substantially reduces both action steps (by 41.6%) and token consumption (by 56.3%), whereas cognitive scaffolding yields limited benefits.

agent behaviorcoding agentslarge language models

This work addresses the unreliability of large vision-language models in fine-grained spatial reasoning, stemming from their limited capacity for precise spatial perception and geometric computation. To overcome this limitation, we propose NeSy-Spatial, a novel framework that introduces, for the first time, a self-evolving neuro-symbolic skill mechanism. This mechanism abstracts tool invocation and geometric operations into executable atomic instructions, enabling dynamic composition, optimization, and reuse of skills through closed-loop reasoning and trajectory replay. By integrating neuro-symbolic systems, tool-augmented learning, and a skill retrieval-execution pipeline, NeSy-Spatial achieves substantial accuracy gains across three spatial reasoning benchmarks, demonstrating more precise tool usage and stronger cross-task generalization.

fine-grained geometric computationgeneralizationmultimodal reasoning

Hot Scholars

MD

Mustafa Doga Dogan

Adobe Research
human-computer interactionaugmented realityhuman-centered AIdigital fabrication
ES

Ehud Sharlin

Professor of Computer Science, University of Calgary
Interaction with Autonomous VehiclesHuman-Drone InteractionHuman-Robot InteractionTangible
UA

Usman Alim

Associate Professor, Department of Computer Science, University of Calgary
VisualizationComputer GraphicsImage ProcessingScientific Computing
KA

Karan Ahuja

Northwestern University
Human-Computer InteractionMachine Learning & SensingUbiquitous Computing
CZ

Chenliang Zhou

University of Cambridge
machine learninggenerative artificial intelligencecomputer visioncomputer graphics