Uruqi: Learning Spatial Cognition from Visual Experience

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing vision-language models in tracking self-motion and maintaining spatial relationship mappings during continuous movement. To overcome this, we propose a dense multi-turn supervision paradigm based on synthetic camera trajectories, which simulates continuous visual experiences to train atomic spatial reasoning, thereby transforming dynamic observations into scalable spatial cognition signals. Leveraging 11,738 theme-driven trajectories, we construct Uruqi, a benchmark comprising 52,000 questions, and fine-tune InternVL3-8B accordingly. Experimental results demonstrate that our approach increases accuracy on this benchmark from 15.84% to 50.41%, achieving performance comparable to GPT-6 Astra, while yielding an average relative improvement of 17.13% across external benchmarks.
📝 Abstract
Spatial intelligence requires maintaining a coherent understanding of the world as the embodied agent moves. Like humans, the agent must use its own motion to interpret changes across observations and update object locations and spatial relations accordingly. Despite spatial post-training having substantially broadened the spatial intelligence of vision-language models (VLMs), they still struggle with two atomic spatial capabilities: tracking self-motion and mapping the surrounding world during motion. To address this gap, we provide dense multi-turn supervision over interleaved atomic capabilities within each training episode, mimicking the visual experience of a continuously moving agent that reasons as it observes. To scale this up, we synthesize 11,738 motif-driven camera trajectories over a broad range of 3D scenes, supporting self-motion tracking, persistent object mapping, and rich spatial operations within each visual experience. By training models to reason over these atomic questions, our URUQI$_{\mathrm{Syn}}$-8B improves accuracy from 15.84% to 47.73% on our Uruqi benchmark comprising 52k questions across 2.7k episodes. URUQI-SI-Mix-8B further reaches 50.41%, comparable to the 50.08% achieved by GPT-6 Astra. Trained solely on our synthesized data, URUQI$_{\mathrm{Syn}}$-8B achieves an average relative accuracy improvement of 17.13% over its InternVL3-8B backbone across three external spatial benchmarks. These results highlight continuous visual experience as a scalable source of supervision for developing spatial cognition in VLMs.
Problem

Research questions and friction points this paper is trying to address.

spatial cognition
vision-language models
self-motion tracking
object mapping
embodied agent
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatial Cognition
Continuous Visual Experience
Self-Motion Tracking
Persistent Object Mapping
Synthetic Trajectories
🔎 Similar Papers
No similar papers found.