Adaptive World Memory 3D Foundation Model for Scalable 3D Mapping, Localization, and Rendering

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种具有自适应世界记忆机制的3D基础模型,用于解决大规模3D建图、定位和渲染问题,通过结合门控更新与时空调节来提高模型的持久性和可扩展性。
📝 Abstract
Recent 3D foundation models enable generalizable geometric reasoning from RGB images but remain limited in persistent memory, scalability, and renderable scene modeling. We present a memory-centric 3D foundation model for scalable robotic localization, reconstruction, and Gaussian rendering. Its core is an adaptive world memory mechanism that combines transformer-based gated updates with test-time temporal-spatial regulation. Learned gates control recurrent memory propagation, while temporal state evolution and spatial observation-state consistency regulate token-wise updates and forgetting over long image sequences. To support large-scale mapping, we organize memory into local submaps and integrate progressive mapping and tracking, loop closure, and SL(4)-based global refinement to maintain local accuracy and global consistency. A Gaussian reconstruction head decodes memory-enhanced features into renderable primitives, unifying camera pose estimation, dense point-cloud reconstruction, and photorealistic rendering within a single model. Experiments on public benchmarks and self-collected datasets from diverse robotic platforms demonstrate improved trajectory accuracy, reconstruction completeness, and rendering quality over existing 3D foundation reconstruction and SLAM baselines. These results support adaptive memory as a foundation for persistent robotic world modeling. The dataset and code will be made publicly available at \href{https://github.com/dtc111111/AWM-3DFM}{https://github.com/dtc111111/AWM-3DFM}.
Problem

Research questions and friction points this paper is trying to address.

3D foundation model
persistent memory
scalability
renderable scene modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive World Memory
Transformer-based Gated Updates
Gaussian Reconstruction
Scalable 3D Mapping
SLAM
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Tianchen Deng
Tianchen Deng
Shanghai Jiao Tong University
RoboticsComputer Vision
G
Guole Shen
School of Automation and Intelligent Sensing, Shanghai Jiao Tong University and State Key Laboratory of Avionics Integration and Aviation System-of-Systems Synthesis, Shanghai Key Laboratory of Navigation and Location Based Services, Shanghai 200240, China
Yilin Shen
Yilin Shen
AI Research Scientist
LLMMultimodal AIAgentOn-device AI
Wenhua Wu
Wenhua Wu
Shanghai Jiao Tong University
computer vision
Y
Yilin Fang
School of Automation and Intelligent Sensing, Shanghai Jiao Tong University and State Key Laboratory of Avionics Integration and Aviation System-of-Systems Synthesis, Shanghai Key Laboratory of Navigation and Location Based Services, Shanghai 200240, China
Z
Ziqi Ma
School of Automation and Intelligent Sensing, Shanghai Jiao Tong University and State Key Laboratory of Avionics Integration and Aviation System-of-Systems Synthesis, Shanghai Key Laboratory of Navigation and Location Based Services, Shanghai 200240, China
Tianjun Zhang
Tianjun Zhang
University of California, Berkeley
Reinforcement LearningMachine LearningArtificial Intelligence
S
Shenghai Yuan
Nanyang Technological University, Singapore
Wolfram Burgard
Wolfram Burgard
Professor of Computer Science, University of Technology Nuremberg
RoboticsArtificial IntelligenceAIMachine LearningComputer Vision
H
Hesheng Wang
School of Automation and Intelligent Sensing, Shanghai Jiao Tong University and State Key Laboratory of Avionics Integration and Aviation System-of-Systems Synthesis, Shanghai Key Laboratory of Navigation and Location Based Services, Shanghai 200240, China