Score
Design and implement representations and algorithms that estimate which parts of space are occupied, free, or unknown, using discrete occupancy grids or continuous implicit occupancy networks. Build the sensor-fusion and probabilistic update pipelines that integrate range/depth/visibility measurements into an updatable map, and provide occupancy queries, uncertainty estimates, and evaluation of map accuracy for downstream tasks such as planning and perception.
This study addresses the challenge of unbounded spatial memory growth in visual robotic navigation within large-scale environments, which risks exhausting embedded platform resources. The authors conduct a systematic survey of 52 spatial memory representations from 1989 to 2025 and introduce, for the first time, a memory efficiency metric α—defined as the ratio of runtime memory to map storage—alongside a standardized evaluation protocol. Covering occupancy grids, neural implicit representations, 3D Gaussian Splatting (3DGS), and scene graphs, the analysis integrates GPU performance profiling and memory-completeness curves, revealing that neural methods exhibit α values spanning two orders of magnitude (2.3–215). This indicates that memory architecture, rather than representation paradigm, predominantly governs deployment feasibility. The work releases the first α benchmark dataset and provides α-aware budgeting algorithms with Pareto frontier analysis.
Autonomous systems face a fundamental trade-off among computational overhead, energy consumption, and model interpretability in occupancy grid map (OGM) modeling. To address this, we propose VSA-OGM—the first OGM framework integrating hyperdimensional computing (VSA) with Fourier-domain vector binding and Shannon entropy-driven probabilistic updating. Unlike conventional dense statistical inference or neural-network-based approaches requiring extensive domain-specific training, VSA-OGM achieves real-time inference, ultra-low power consumption, and strong interpretability without any training. Experiments show that VSA-OGM matches the accuracy of covariance propagation while reducing inference latency by 200× and memory footprint by 1000×. It further cuts latency by 3.7× compared to non-deformable traditional methods and outperforms state-of-the-art neural OGMs by 1.5× in speed—entirely training-free.
Existing LiDAR occupancy grid prediction methods predominantly rely on deterministic, grid-level optimization, often yielding physically implausible and scene-inconsistent artifacts that compromise safety-critical autonomous navigation. To address this, we propose the first latent-space disentangled generative occupancy prediction framework: it decouples representation learning from stochastic prediction, explicitly modeling uncertainty in the latent space while supporting multimodal conditional inputs—including RGB images and high-definition maps. Our approach employs a hybrid VAE-GAN architecture, enabling end-to-end training and zero-shot cross-platform transfer. Evaluated on NuScenes, Waymo Open Dataset, and a proprietary real-world vehicle dataset, our method achieves state-of-the-art performance, significantly improving prediction fidelity, physical plausibility, and scene consistency—key requirements for robust autonomous driving systems.
To address the challenge of predicting future robot states in complex dynamic environments, this paper proposes an uncertainty-aware stochastic occupancy prediction engine that jointly models robot ego-motion, dynamic object motion, and static scene geometry to generate multimodal distributions over future environmental states. Our method is the first lightweight, end-to-end framework for joint stochastic modeling of motion and geometry. Key innovations include a software-optimized stochastic occupancy map representation, a probabilistic propagation acceleration algorithm, and a unified training framework for heterogeneous multi-source data. Experimental results demonstrate significant efficiency improvements: 10× faster inference speed and 3× reduced memory footprint compared to prior approaches. On three real-world and simulated benchmarks, our method achieves superior prediction accuracy and robustness over state-of-the-art baselines, thereby enhancing the safety and reliability of downstream navigation policies.
Existing LiDAR occupancy prediction methods suffer from two key limitations: deterministic modeling fails to capture environmental stochasticity, and they inadequately fuse multimodal inputs—such as RGB images, high-definition maps, and planning trajectories. This paper proposes a self-supervised multi-future occupancy prediction framework that models uncertainty in latent space. Our core contributions are: (1) the first latent-space stochastic occupancy prediction paradigm; (2) a unified feature fusion mechanism supporting diverse multimodal conditional inputs; and (3) a dual-path decoder architecture—comprising a single-step decoder for real-time inference and a diffusion-enhanced batch decoder for temporal consistency. Evaluated on nuScenes and Waymo Open Dataset, our method achieves state-of-the-art performance, significantly mitigating compression artifacts and motion discontinuities. Both qualitative and quantitative results demonstrate consistent improvements across all major metrics.
To address inaccurate depth estimation, unmodeled uncertainty, and resulting global geometric inconsistency in visual-inertial SLAM—hindering real-time robot planning—this paper proposes an uncertainty-aware tightly coupled VIO-SLAM framework. Methodologically, it introduces the first deep integration of motion stereo vision and depth neural networks; pixel-wise depth and its uncertainty—output by the network—are jointly propagated via reprojection and IMU preintegration to voxel occupancy probabilities and submap alignment factors, enabling globally consistent and scalable dense submap representation. The framework synergistically integrates deep depth estimation, probabilistic graph optimization, voxel-hashed occupancy mapping, and nonlinear least-squares optimization. Evaluated on EuRoC and TUM-VI benchmarks, it outperforms state-of-the-art methods in both localization and mapping accuracy, while enabling real-time generation of high-fidelity, confidence-aware voxel occupancy maps directly usable for downstream robotic planning and control.
研究通过动态过滤策略解决机器人主动映射中占用模型错误导致的探索和移动问题,提高覆盖效率。
本文提出了一种基于采样的方法,通过统计估计而非穷举递归构建信息驱动的多分辨率概率占据网格层次表示,解决了大规模网格计算复杂性问题。
研究通过引入观察门控过滤器来解决自主3D主动映射中因预测占用图不准确导致的规划问题,该方法在无需重新训练或真实值的情况下改进了机器人导航和覆盖范围。
This work addresses the challenges of trajectory drift accumulation and computationally expensive global consistency optimization in large-scale SLAM, as well as limitations of existing discrete grid-based submap stitching methods—such as discontinuous gradients and neglect of occupancy uncertainty—by introducing the first continuous probabilistic submap stitching framework. The method jointly optimizes submap poses and a global occupancy field in an implicit log-odds space, compressing raw observations into informative sufficient statistics via sparse Bayesian inference and incorporating a variance-weighting mechanism to preserve posterior uncertainty. It enables analytical Jacobian computation and directly yields an optimal global map with closed-form mean and variance upon pose convergence. Experiments demonstrate significant improvements over state-of-the-art approaches in both simulated and real large-scale environments, achieving higher pose accuracy, enhanced global consistency, greater map compactness, and better-calibrated uncertainty.
Traditional 3D occupancy grid mapping faces significant challenges in unknown environments, including high memory consumption and substantial update latency, which hinder its applicability for autonomous robots requiring efficient and scalable mapping. This work proposes a boundary-based occupancy mapping framework that innovatively integrates truncated ray casting with a direct boundary update mechanism. By eliminating the need for auxiliary local voxel grids, the method avoids storing voxels across the entire space and bypasses exhaustive ray traversal. Experimental results on public datasets demonstrate that the proposed approach substantially outperforms existing baseline and boundary-aware methods, achieving comparable mapping accuracy while significantly reducing both memory usage and update time.