๐ค AI Summary
This work addresses the reliance on dense candidate proposals and complex post-processing in radar-only 3D object detection by proposing an end-to-end detection architecture based solely on a Transformer decoder. The method formulates detection as a set prediction task, leveraging learnable object queries and positional encoding. It introduces, for the first time, a Pyramid Token Fusion (PTF) module to effectively aggregate multi-scale radar features. As the first fully Transformer decoderโbased framework for radar-only 3D detection, the approach eliminates the need for dense proposal generation and non-maximum suppression (NMS). Evaluated on the RADDet dataset, the model significantly outperforms existing baselines, demonstrating both its effectiveness and novelty.
๐ Abstract
In this paper, we present a Transformer-based architecture for 3D radar object detection that uses a novel Transformer Decoder as the prediction head to directly regress 3D bounding boxes and class scores from radar feature representations. To bridge multi-scale radar features and the decoder, we propose Pyramid Token Fusion (PTF), a lightweight module that converts a feature pyramid into a unified, scale-aware token sequence. By formulating detection as a set prediction problem with learnable object queries and positional encodings, our design models long-range spatial-temporal correlations and cross-feature interactions. This approach eliminates dense proposal generation and heuristic post-processing such as extensive non-maximum suppression (NMS) tuning. We evaluate the proposed framework on the RADDet, where it achieves significant improvements over state-of-the-art radar-only baselines.