🤖 AI Summary
This study addresses the deployment challenges and hardware lock-in issues of Bird's Eye View (BEV) perception models caused by operator mismatches, particularly with sparse convolutions. To this end, we propose BEVPIPE, a framework that introduces a novel architecture decoupling dense subgraphs from sparse operators to eliminate CUDA dependencies. By integrating model partitioning, portable GPU computation APIs, custom operator extensions, and shared memory mechanisms, BEVPIPE enables efficient and portable deployment of multimodal BEV perception within standard inference runtimes. Experimental results demonstrate that the proposed framework achieves a 19.5× end-to-end speedup while retaining 98.5% of the reference mean Average Precision (mAP). Furthermore, its compatibility across diverse GPU backends is empirically validated, highlighting its practicality for flexible, high-performance deployment in autonomous driving applications.
📝 Abstract
Modern autonomous driving systems rely on bird's-eye-view (BEV) perception models that fuse camera and LiDAR inputs to detect objects in 3D space. These models are accurate, but they cannot be deployed through standard inference runtimes. The reason is an operator mismatch between dense convolutions (which runtimes handle well), sparse 3D convolutions (which runtimes cannot represent), and geometric scatter operations (which runtimes have no vocabulary for). Today, every sparse convolution library is CUDA-only and PyTorch-coupled, locking BEV deployment to a single vendor's hardware and a single execution framework.
We present BEVPIPE, a framework for deploying multimodal BEV perception pipelines using portable GPU compute APIs and integrating them with production inference runtimes. BEVPIPE partitions the model into runtime-managed dense subgraphs and three external operator extensions (voxelizer, sparse encoder, BEV projector), connected through a shared GPU memory space. BEVPIPE achieves a 19.5x end-to-end speedup over conventional deployments while retaining 98.5% of reference mAP. We also showcase that BEVPIPE is portable across different GPU backends.