AdaOcc: Adaptive 3D Occupancy Prediction for Embodied Tasks

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the deployment challenges of 3D occupancy prediction in embodied tasks, where adapting to heterogeneous sensors and dynamic computational budgets remains difficult. To this end, we propose a point-based adaptive semantic occupancy prediction framework. Methodologically, we design a geometry-guided dual-branch encoder to uniformly process multi-view RGB, depth, and LiDAR inputs. Furthermore, we integrate a progressive query decoding strategy with sparse point representations to enable flexible adjustment of computational costs, and introduce a novel inclusion loss function to constrain effective region modeling. Extensive experiments demonstrate that the proposed framework achieves state-of-the-art performance on the Occ-ScanNet dataset. Moreover, real-world evaluations validate its strong adaptability within practical embodied systems, highlighting its potential for efficient and robust deployment under varying hardware constraints.
📝 Abstract
Embodied tasks demand accurate, flexible, and semantically rich 3D scene representations. 3D semantic occupancy is well suited to this requirement, as it can model holistic 3D spaces by encoding geometric occupancy along with semantic categories. However, existing occupancy prediction methods struggle to meet practical deployment requirements, such as adapting to varying computing budgets, sensor setups, and observation views. In this paper, we propose a point-based Adaptive 3D Occupancy Prediction method, called AdaOcc, tailored for embodied scenarios. To accommodate heterogeneous sensor inputs, AdaOcc uses an adaptive geometry-guided dual-branch encoder that can support RGB images in various numbers of views with (estimated) depth maps or LiDAR scans. AdaOcc represents occupied regions via sparse semantic points trained with a progressive query learning strategy, allowing the prediction computational budget to be flexibly adjusted through query point numbers and decoder layers. To facilitate high-fidelity geometric modeling for lightweight point-based occupancy learning, we further propose a novel containment loss that regularizes predicted points to reside within valid occupied regions. Extensive experiments show that our method achieves a new state-of-the-art on Occ-ScanNet with considerable performance improvements over previous methods. Moreover, our framework demonstrates strong practical applicability as an adaptive 3D perception module in real-world embodied systems.
Problem

Research questions and friction points this paper is trying to address.

3D Occupancy Prediction
Embodied Tasks
Adaptive Perception
Heterogeneous Sensors
Computational Budget
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive 3D Occupancy Prediction
Embodied Tasks
Dual-branch Encoder
Progressive Query Learning
Containment Loss
🔎 Similar Papers
No similar papers found.
J
Jinglong Wang
Beihang University
Y
Yunjie Wang
Beijing Academy of Artificial Intelligence
Zhiyang Zhang
Zhiyang Zhang
Nanjing University
NLPLLMAgentAIOps
Jiawei He
Jiawei He
XYZ Embodied AI
Computer VisionEmbodied AI
Y
Ye Yuan
ShanghaiTech University
B
Bo Qiu
University of Science and Technology Beijing
Jing Zhang
Jing Zhang
School of Software, Beihang University
Computer VisionTransfer LearningDeep Learning