🤖 AI Summary
This study addresses the bottleneck of camera-only multi-view 3D detection, which typically relies on sequential training and persistent query memory. To overcome this limitation, the authors propose a progressive query refinement framework that supports random frame sampling and independent inference. Methodologically, a masked self-attention mechanism is designed to isolate denoising supervision, while stage-decoupled anchor embeddings are introduced to suppress interference from temporal attributes. Furthermore, built upon a ViT-L backbone, the framework incorporates motion-aligned high-confidence query propagation alongside a staged position-attribute injection strategy. By eliminating the dependence on stateful memory, this work achieves new state-of-the-art performance on the nuScenes test set, yielding 71.6 NDS and 64.9 mAP.
📝 Abstract
Temporal context is essential for camera-only multi-view 3D object detection. Existing streaming detectors maintain and propagate query states from one frame to the next, requiring sequence-aware training and chronological inference. We propose PQR3D, which performs progressive query refinement within referenceconditioned temporal windows. This design enables random frame sampling and independent inference without persistent query memory. Within each window, PQR3D progressively transfers motion-aligned high-confidence queries from earlier timestamps toward the target frame. We further introduce masked selfattention to regulate interactions among regular, propagated, and denoising queries while keeping denoising supervision isolated from detection queries. In addition, a stage-decoupled anchor embedding injects position before self-attention and size, orientation, and velocity afterward, reducing interference from temporally inconsistent attributes. With a ViT-L backbone, PQR3D sets a new state of the art on the nuScenes test set, achieving 71.6 NDS and 64.9 mAP. Source code is available at https://github.com/huiyegit/PQR3D