Score
Design, build, and analyze probabilistic 3D voxel/occupancy-grid mapping systems that fuse heterogeneous or distributed sensor observations into per-voxel occupancy probabilities and uncertainties using Bayesian/probabilistic fusion and grid-fusion techniques. Produce maps and associated per-cell likelihoods that preserve geometric detail during operations such as downsampling and ray integration, support occupancy tuning and efficient updates, and serve downstream needs like collision checking, planning, and extending perceptual reach.
Existing world modeling research predominantly focuses on 2D image/video generation, neglecting large-scale scene modeling using native 3D/4D representations—such as RGB-D, occupancy grids, and LiDAR point clouds—and lacks a unified definition and systematic taxonomy. Method: This paper introduces, for the first time, a standardized definition and a structured classification framework for 3D/4D world models, systematically categorizing generative paradigms into VideoGen, OccGen, and LiDARGen. It integrates generative modeling, 3D perception, and spatiotemporal modeling techniques, and synthesizes evaluation metrics and benchmark datasets. Contribution/Results: As the first comprehensive survey in this emerging field, it establishes WorldBench—an open-source literature platform—thereby filling a critical theoretical gap and providing foundational guidance and standardization pathways for 3D/4D world modeling research.
Autonomous systems face a fundamental trade-off among computational overhead, energy consumption, and model interpretability in occupancy grid map (OGM) modeling. To address this, we propose VSA-OGM—the first OGM framework integrating hyperdimensional computing (VSA) with Fourier-domain vector binding and Shannon entropy-driven probabilistic updating. Unlike conventional dense statistical inference or neural-network-based approaches requiring extensive domain-specific training, VSA-OGM achieves real-time inference, ultra-low power consumption, and strong interpretability without any training. Experiments show that VSA-OGM matches the accuracy of covariance propagation while reducing inference latency by 200× and memory footprint by 1000×. It further cuts latency by 3.7× compared to non-deformable traditional methods and outperforms state-of-the-art neural OGMs by 1.5× in speed—entirely training-free.
This work addresses the excessive memory consumption of conventional occupancy grid maps when deployed at high resolutions or large scales. To mitigate this, the authors propose a boundary-based map representation that explicitly stores only boundary voxels—such as occupied and frontier voxels—while implicitly encoding interior free and exterior unknown regions via two-dimensional closed surfaces. An efficient mechanism for occupancy state querying and updating is devised, integrated within a global–local mapping framework to enable real-time construction. Additionally, specialized data structures are introduced to enhance operational efficiency. The proposed method substantially reduces memory footprint while supporting efficient map construction, updates, and queries under real-time sensor input. The implementation has been made publicly available.
This work addresses the challenge of fairly comparing Bayesian log-odds and Dempster’s rule of combination in two-dimensional occupancy grid mapping by proposing a unified matching framework based on the pignistic transformation. By mapping the outputs of diverse fusion methods into a consistent decision probability space, the framework effectively isolates the influence of sensor-specific parameters, thereby enabling a direct comparison of the fusion rules themselves. For the first time, this approach facilitates systematic evaluation across paradigmatically distinct methods, with validation demonstrated through simulations, real LiDAR data, and path planning tasks. Experimental results reveal that under pignistic probability (BetP) matching, Bayesian methods significantly outperform Dempster-based approaches (15/15 consistent outcomes, p = 3.1e-5, effect size 0.001–0.022), whereas the conclusion reverses when using normalized belief matching—highlighting the strong dependence of fusion performance on the matching criterion and underscoring the necessity and novelty of the proposed framework.
To address inaccurate depth estimation, unmodeled uncertainty, and resulting global geometric inconsistency in visual-inertial SLAM—hindering real-time robot planning—this paper proposes an uncertainty-aware tightly coupled VIO-SLAM framework. Methodologically, it introduces the first deep integration of motion stereo vision and depth neural networks; pixel-wise depth and its uncertainty—output by the network—are jointly propagated via reprojection and IMU preintegration to voxel occupancy probabilities and submap alignment factors, enabling globally consistent and scalable dense submap representation. The framework synergistically integrates deep depth estimation, probabilistic graph optimization, voxel-hashed occupancy mapping, and nonlinear least-squares optimization. Evaluated on EuRoC and TUM-VI benchmarks, it outperforms state-of-the-art methods in both localization and mapping accuracy, while enabling real-time generation of high-fidelity, confidence-aware voxel occupancy maps directly usable for downstream robotic planning and control.
Existing LiDAR occupancy grid prediction methods predominantly rely on deterministic, grid-level optimization, often yielding physically implausible and scene-inconsistent artifacts that compromise safety-critical autonomous navigation. To address this, we propose the first latent-space disentangled generative occupancy prediction framework: it decouples representation learning from stochastic prediction, explicitly modeling uncertainty in the latent space while supporting multimodal conditional inputs—including RGB images and high-definition maps. Our approach employs a hybrid VAE-GAN architecture, enabling end-to-end training and zero-shot cross-platform transfer. Evaluated on NuScenes, Waymo Open Dataset, and a proprietary real-world vehicle dataset, our method achieves state-of-the-art performance, significantly improving prediction fidelity, physical plausibility, and scene consistency—key requirements for robust autonomous driving systems.
Traditional 3D occupancy grid mapping faces significant challenges in unknown environments, including high memory consumption and substantial update latency, which hinder its applicability for autonomous robots requiring efficient and scalable mapping. This work proposes a boundary-based occupancy mapping framework that innovatively integrates truncated ray casting with a direct boundary update mechanism. By eliminating the need for auxiliary local voxel grids, the method avoids storing voxels across the entire space and bypasses exhaustive ray traversal. Experimental results on public datasets demonstrate that the proposed approach substantially outperforms existing baseline and boundary-aware methods, achieving comparable mapping accuracy while significantly reducing both memory usage and update time.
This work addresses the challenges of trajectory drift accumulation and computationally expensive global consistency optimization in large-scale SLAM, as well as limitations of existing discrete grid-based submap stitching methods—such as discontinuous gradients and neglect of occupancy uncertainty—by introducing the first continuous probabilistic submap stitching framework. The method jointly optimizes submap poses and a global occupancy field in an implicit log-odds space, compressing raw observations into informative sufficient statistics via sparse Bayesian inference and incorporating a variance-weighting mechanism to preserve posterior uncertainty. It enables analytical Jacobian computation and directly yields an optimal global map with closed-form mean and variance upon pose convergence. Experiments demonstrate significant improvements over state-of-the-art approaches in both simulated and real large-scale environments, achieving higher pose accuracy, enhanced global consistency, greater map compactness, and better-calibrated uncertainty.
本文提出了一种基于采样的方法,通过统计估计而非穷举递归构建信息驱动的多分辨率概率占据网格层次表示,解决了大规模网格计算复杂性问题。
该研究解决了3D场景图中不确定性表示和传播问题,通过引入概率场景图(PSG)及高斯层次图(HGG),实现了实时感知与精确定位。
This work addresses the limitations of single-vehicle perception in vehicle-to-everything (V2X) systems, where restricted sensor coverage hinders reliable cooperative awareness. To overcome this challenge, the authors propose a Bayesian fusion–based collaborative perception framework that integrates heterogeneous sensor data from multiple agents to construct interpretable probabilistic occupancy grids. The system’s performance is evaluated through a hybrid validation approach combining CARLA-based virtual simulation with real-vehicle-in-the-loop testing, ensuring reproducible and certifiable assessment. In a roundabout scenario, collaboration among six vehicles expands perceptual coverage by 260% and significantly improves occupancy grid recall from 0.82 (single-vehicle baseline) to 0.94, thereby substantially enhancing environmental awareness in occluded regions beyond line-of-sight.