probabilistic voxel mapping

Designs, implements, and analyzes pipelines that voxelize 3D sensor data (point clouds, depth frames) and construct voxel backbones that project multi-view observations into a shared voxel space using depth and camera pose. These systems perform probabilistic fusion and source-aware semantic labeling to produce persistent volumetric semantic maps that support incremental queries, redundancy detection, voxel-based alignment, multi-resolution/local refinement, and other voxel-grid processing operations.

probabilisticvoxelmapping

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the lack of environment representation methods for mobile manipulators under edge computing constraints that simultaneously ensure timeliness, semantic richness, and geometric reliability. The authors propose KRVF, a task-oriented voxelized world representation that integrates occupancy, color, semantic evidence, temporal recency, and data provenance to enable closed-loop feedback between mapping and perception. Key innovations include decoupling observed occupancy from semantic priors to enhance object reasoning robustness under depth failure, introducing a map-prior-driven depth inpainting mechanism, and designing a task-aware semantic query interface. Built upon semantic voxel representations, source-aware fusion, and real-time ROS 2 integration, the system efficiently supports semantic object retrieval and grasp candidate generation on edge devices.

depth failureedge computingmobile manipulation

To address insufficient geometric-semantic joint understanding of robots in unstructured environments, this paper proposes an end-to-end modular RGB-D understanding pipeline. Our method introduces a novel hybrid mask generation mechanism integrating SAM2 with a lightweight semantic classifier for pixel-level semantic segmentation and instance awareness; incorporates ReID-enhanced cross-frame human tracking and semantic-weighted TSDF point cloud fusion to ensure geometric fidelity and semantic consistency; and outputs structured scene representations in USD format. Evaluated on ADE20K, our approach achieves 47.0% mIoU—surpassing SegFormer and OneFormer in boundary accuracy—and attains a reconstruction error of only 25.3 mm. It runs 1.81× faster than prior methods and has been validated for deployment feasibility on real-world Kinect data.

Enables efficient human tracking and point cloud fusionImproves semantic segmentation accuracy using hybrid SAM2 modelIntegrates semantic segmentation with geometric reconstruction for robotics

This work addresses the challenge of achieving open-vocabulary 3D semantic mapping in large-scale environments, where existing methods often fall short. The authors propose an online two-layer semantic mapping framework that unifies voxel-level dense semantics and instance-level open-vocabulary representations within a shared voxel map for the first time. By introducing a cross-layer semantic fusion mechanism within a sliding window, the method jointly optimizes the quality of both semantic layers. Coupled with multi-view semantic embeddings and voxelized 3D mapping, this approach significantly enhances accuracy and scalability. The framework demonstrates high-fidelity, generalizable open-vocabulary semantic mapping on standard 3D semantic segmentation benchmarks as well as large, multi-floor real-world scenes, effectively extending to unseen semantic concepts.

3D semantic mappingdense semantic mapinstance-level mapping

This work addresses the limitation of existing open-vocabulary 3D scene graph methods, which rely on a sequential “reconstruct-then-enrich” pipeline that precludes real-time querying during exploration. To overcome this, the authors propose an asynchronous architecture that decouples lightweight online mapping from heavyweight semantic refinement, executing them in parallel. A probabilistic voxel backbone maintains object identity consistency, while a background visual-language model (VLM) agent incrementally enriches semantic annotations. The approach further incorporates semantic loop closure to eliminate redundant trajectories and introduces a multi-objective frame scheduler to reduce VLM computational overhead. This is the first method to enable queryable scene graphs during active exploration, achieving state-of-the-art semantic segmentation performance on ScanNet and Replica, and significantly outperforming prior art by 15.3–18.8 A@0.25 on the Sr3D+, Nr3D, and ScanRefer benchmarks.

asynchronous processingincremental mappingopen-vocabulary 3D scene graph

VoxelNextFusion: A Simple, Unified, and Effective Voxel Fusion Framework for Multimodal 3-D Object Detection

Jan 05, 2024
ZS
Ziying Song
🏛️ Beijing Jiaotong University | Hebei University of Science and Technology | Lenovo Research | University of Macau

In LiDAR-camera fusion 3D detection, the sparsity of voxel features and the density of image features hinder cross-modal alignment and lead to semantic and continuity information loss—particularly for distant objects. To address this, we propose a voxel-space image feature processing framework: (1) image features are projected onto the voxel grid via voxelization; (2) pixel-level and patch-level multi-granularity features are extracted; (3) a cross-modal self-attention fusion mechanism enables fine-grained alignment of image features within the voxel space; and (4) a foreground-aware importance weighting module suppresses background interference. Evaluated on KITTI, our method achieves a +3.20% improvement in AP@0.7 for hard-category car detection over the Voxel R-CNN baseline. Furthermore, it demonstrates strong generalizability on nuScenes, validating its robustness across diverse driving scenarios and sensor configurations.

Address challenges in fusing sparse voxel and dense image featuresEnhance 3D object detection using LiDAR-camera fusionImprove detection performance, especially at long distances

Latest Papers

What's happening recently
View more

The field of 3D vision suffers from fragmented data representations, learning paradigms, and benchmarking protocols, leading to a lack of unified understanding regarding efficiency, fidelity, and scalability. This work proposes the first cohesive conceptual framework that integrates geometric representations—such as point clouds, meshes, voxels, and 3D Gaussians—with diverse learning paradigms—including 2D-supervised learning, implicit neural representations, and 4D modeling—and connects them to real-world application scenarios. By constructing a structured knowledge graph of 3D vision, the study systematically relates dataset design, supervision mechanisms, and task requirements, clarifying the trade-offs between efficiency and fidelity and charting pathways for multimodal geometric grounding. This framework offers systematic guidance for reconstruction, generation, and dynamic scene modeling, advancing the field toward a unified and efficient paradigm.

3D visionbenchmark fragmentationdata representation

该研究提出一种结合外部校准相机、对象检测、持续跟踪及本体驱动语义更新的混合管道,以构建动态语义世界模型,解决机器人在复杂环境中的交互问题。

contextual understandingdynamic ontologysemantic mapping

This work investigates how prior information from a fixed external RGB camera can enhance the initial completeness and exploration efficiency of 3D scene graphs in robotic active exploration. The method models external camera observations as a Common Prior Map (CPM), constructing semantic and geometric priors before exploration begins. A hardware-agnostic, RGB-only multi-view fusion framework seamlessly integrates both onboard and external viewpoints to enable incremental 3D scene graph generation. The key innovation lies in leveraging a fixed external camera as a universal prior source without requiring hardware modifications, coupled with a semantic uncertainty–guided active exploration strategy driven by partial scene graphs. Experiments demonstrate that a single external camera can improve initial object recall by up to 79%, substantially boosting both exploration efficiency and scene graph completeness.

3D Scene Graph GenerationActive ExplorationCommon Prior Maps

Hot Scholars

AV

Abhinav Valada

Professor & Director of Robot Learning Lab, University of Freiburg
RoboticsMachine LearningComputer VisionArtificial Intelligence
YC

Yu-Chiang Frank Wang

National Taiwan University & NVIDIA
Computer VisionDeep LearningMachine LearningArtificial Intelligence
KY

Kailun Yang

Professor. School of Artificial Intelligence and Robotics, Hunan University (HNU); KIT; UAH; ZJU
Computer VisionComputational OpticsIntelligent VehiclesAutonomous Driving
SL

Stefan Leutenegger

Associate Professor at ETH Zurich
Mobile RoboticsSpatial AIUnmanned Aerial Systems
LP

Lap-Pui Chau

The Hong Kong Polytechnic University
Visual Signal Processing