IndoorBEV: A Lightweight Real-Time LiDAR BEV Perception System for Indoor Mobile Robots

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of efficient perception in cluttered 3D environments for indoor mobile robots operating under strict latency and memory constraints. We propose a lightweight LiDAR perception framework that introduces statistical height features and multi-frequency encoding to explicitly preserve vertical geometric cues within a compact bird's-eye-view representation, enabling efficient 3D information processing via 2D convolutions. The architecture further integrates geometry-conditioned feature aggregation, a lightweight encoder, and decoupled dense prediction heads. With only 0.6M parameters, the model achieves an inference latency as low as 169.6 ms on the NVIDIA Orin platform, effectively balancing accuracy, low latency, and low power consumption while significantly outperforming existing methods.
📝 Abstract
Efficient indoor LiDAR perception is challenging because mobile robots must understand cluttered three-dimensional environments under strict latency and memory constraints. Existing point-based and voxel-based methods often incur substantial computational overhead, whereas conventional bird's-eye-view (BEV) representations improve efficiency at the cost of discarding vertical geometric information. We present IndoorBEV, a lightweight LiDAR perception framework that mitigates this tradeoff through a height-aware BEV representation and geometry-conditioned feature fusion. IndoorBEV summarizes the vertical point distribution in each BEV cell using statistical height features and multi-frequency height encoding, allowing informative three-dimensional cues to be processed efficiently by two-dimensional convolutions. A lightweight encoder then integrates complementary geometric features with multi-scale local representations and compact global scene context. Decoupled dense prediction heads jointly produce semantic BEV maps and oriented object bounding boxes. IndoorBEV contains only 0.6M parameters and requires 2.3 MB of model storage. On an NVIDIA AGX Orin, it uses 21.52 MB of GPU memory per inference and achieves a mean latency of 169.6 ms under a 200 ms perception deadline, with a deadline miss ratio of 1.8\%. Evaluations on simulated scenes, real-world robot scans, and an open-source indoor point-cloud dataset demonstrate a favorable tradeoff among perception accuracy, latency, and memory consumption. These results indicate that explicitly encoding vertical geometry within a compact BEV representation provides an effective approach to resource-efficient indoor LiDAR perception.
Problem

Research questions and friction points this paper is trying to address.

Indoor LiDAR perception
Bird's-eye-view representation
Mobile robots
Resource constraints
3D environment understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Height-aware BEV representation
Multi-frequency height encoding
Geometry-conditioned feature fusion
Lightweight LiDAR perception
Decoupled dense prediction heads
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.