🤖 AI Summary
This study addresses the reliance on manual annotations and the limited perception of long-tail objects in autonomous driving by proposing a tri-modal unsupervised world model. The method integrates camera, LiDAR, and radar data, pioneering the use of these three sensors as mutual self-supervisory signals to achieve 4D occupancy prediction, scene flow estimation, and zero-shot segmentation. By requiring no additional annotations, it enables the perception of arbitrary obstacles, overcoming the long-tail bottleneck inherent in existing open-set approaches. Evaluated on benchmarks such as Argoverse 2, the proposed model achieves state-of-the-art performance across multiple 3D and 4D tasks while facilitating efficient zero-shot road obstacle segmentation.
📝 Abstract
We present TriO, a multi-modal unsupervised world model that predicts 4D occupancy, obstacle segmentation, flow and LiDAR. In contrast to prior work, TriO utilizes three distinct sensor modalities (camera, LiDAR, and RADAR) as both inputs and sources of self-supervision, eliminating the need for additional human annotations. Thanks to its novel supervision, the model is able to segment any occupancy from the drivable surface, overcoming the limitations of existing open-set methods in handling long-tail objects. TriO achieves state-of-the-art results in multiple 3D and 4D tasks, including occupancy, flow, and LiDAR prediction, as well as zero-shot road obstacle segmentation across multiple datasets such as Argoverse 2, and Spotting the Unexpected.