🤖 AI Summary
This study addresses the limitation of existing latent LiDAR generation models, where convolutional VAE decoders tend to blur depth discontinuities, resulting in "flying pixel" artifacts. To overcome this, we propose CRISP, a framework that replaces conventional convolutional modules with a pixel-space diffusion decoder. By introducing a backbone-agnostic latent adapter and a DiT-based denoiser, our method enables zero-shot, plug-and-play decoder replacement, effectively correcting edge depth drift while supporting masked prediction tasks. Extensive experiments across multiple benchmark datasets demonstrate that CRISP reduces flying pixel errors by an average of 50.5%, substantially improving point cloud reconstruction quality and effectively narrowing the simulation-to-reality domain gap.
📝 Abstract
Latent LiDAR pipelines suffer from flying pixels: convolutional VAEs blur sharp radial depth discontinuities, yielding edge depths that back-project to points floating between surfaces. We identify this as a major, directly correctable decoder bottleneck and introduce CRISP: a pixel-space diffusion decoder with a backbone-agnostic latent adapter, DiT-based denoiser, and support mask predictor. CRISP replaces video-VAE and LiDAR-native decoders alike while keeping the encoder and latent generator fixed. Across KITTI-360, SemanticKITTI, and nuScenes, replacing only the decoder reduces FSVD/FPVD by 50.5% on average across frozen backbones; for generic video VAEs, the reductions reach 71%/74%. On the LiDAR-native LiDM backbone, FRID drops by 71%, with the largest gains at depth discontinuities. In a pretrained LiDM world model, the same zero-shot replacement improves FSVD by 15.5%, narrowing the sim-to-real gap.