OREN-X: Octree Residual Network for Real-Time Multi-Modal Mapping

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high computational overhead and lack of cross-modal synergy in existing methods that independently represent geometric, radiometric, and vision-language information. We propose OREN-X, an online mapping framework that unifies multi-modal environment storage and retrieval via a shared octree data structure. A cross-modal synergy mechanism is designed to enhance signed distance function (SDF) estimation by leveraging occupancy and radiance fields. Furthermore, GPU-accelerated ray-tracing traversal combined with online dictionary learning for feature compression enables efficient storage and precise querying. The proposed method supports real-time mapping, achieving SDF computation exceeding 80 fps while improving near-surface SDF accuracy by 33%. Additionally, open-vocabulary 3D mIoU and average accuracy increase by 71% and 61%, respectively, significantly advancing autonomous navigation capabilities.
📝 Abstract
To achieve general-purpose autonomy over long horizons, a robot needs to maintain spatial environment information that supports a variety of tasks: geometry for planning and control, radiance for rendering and relocalization, and vision-language features for open-vocabulary grounding. Existing methods represent and estimate each modality separately, multiplying memory and compute cost while forgoing potential synergy among the representations. We develop OREN-X, an online mapping method that uses an octree in 3D space as a shared data structure for indexing and storing a multi-modal field, capturing geometric, radiance, and vision-language information. OREN-X provides efficient unified storage and retrieval of these data in explicit/implicit and full/compressed form. Our unified representation yields cross-modality synergy: SDF estimates are sharpened by occupancy and radiance, while GPU-based ray-octree traversal and octree query enable real-time rendering. We also use online dictionary learning to compress the vision-language features, shrinking them 3.7x below full per-vertex storage while raising the query accuracy. On Replica, OREN-X maps in real time (80+ fps for SDF and 30+ fps for all four modalities), improves near-surface SDF accuracy by 33% over single-modality baselines, and improves mean open-vocabulary 3D mIoU by 71% and mean accuracy by 61% over the best prior method.
Problem

Research questions and friction points this paper is trying to address.

multi-modal mapping
spatial representation
robot autonomy
cross-modality synergy
real-time mapping
Innovation

Methods, ideas, or system contributions that make the work stand out.

Octree
Multi-modal Mapping
Cross-modality Synergy
Online Dictionary Learning
Real-time Rendering
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Zhirui Dai
Zhirui Dai
UC San Diego
Robotics
Q
Qihao Qian
Department of Electrical and Computer Engineering, University of California San Diego, La Jolla, CA 92093, USA
D
Dinh Minh Nguyen
VinMotion
Quan-Dung Pham
Quan-Dung Pham
VinMotion
K
Kiana Bronder
Parsons
Carlos Nieto-Granda
Carlos Nieto-Granda
U.S. Army Research Laboratory (ARL)
Multi-robot and multi-agent systemsAutonomous Navigation & ExplorationSLAMHuman-Robot Teams
Y
Yiyu Chen
VinMotion
Quan Nguyen
Quan Nguyen
University of Southern California
ControlRoboticsOptimization
N
Nikolay Atanasov
Department of Electrical and Computer Engineering, University of California San Diego, La Jolla, CA 92093, USA