SGDet3D++: Geometry-Grounded Semantics for 4D Radar and Camera 3D Object Detection

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
该研究通过提出SGDet3D++,采用假设条件证据接地方法,解决了4D雷达与相机在3D物体检测中的模态对齐问题,提高了检测精度。
📝 Abstract
4D radar complements dense image semantics with long-range geometry and radial motion, but existing radar--camera detectors largely solve \emph{where} to align the modalities while leaving \emph{whether} a piece of evidence supports an evolving object hypothesis implicit. An image token may describe an occluder, a nearby radar return may belong to another object, and a pose-aligned memory slot may carry incompatible motion. We formulate \emph{hypothesis-conditioned evidence grounding}, which separates candidate access from evidence use: semantic, geometric, or temporal evidence is filtered or conditioned by the evolving 3D state before updating the corresponding query. \sgdetpp{} instantiates this principle through Anchor-Grounded Semantic Retrieval (AGR), which conditions deformable image retrieval on pooled anchor-consistent radar support; Geometry-Consistent Anchor Refinement (GCR), which attentively aggregates individual associated returns; and Doppler-Verified Correspondence (DVC), which replaces history only when current radial motion contradicts it. \sgdetpp{} improves the strongest compared method by 3.82 mAP and 6.82 ODS on OmniHD-Scenes and by 6.82 mAP and 9.22 NDS on ManTruckScenes, while also leading the listed methods in the TJ4DRadSet test comparison. Mechanism-targeted evaluations show that AGR improves strict AP in every projected-occlusion bin, the yaw-aligned box gate raises target-return purity from 29.95\% to 58.87\%, and DVC preserves 96.11\% of motion-consistent history while retaining 75.90\% contradiction recall. Code will be released.
Problem

Research questions and friction points this paper is trying to address.

4D radar
3D object detection
evidence grounding
multi-modal alignment
object hypothesis
Innovation

Methods, ideas, or system contributions that make the work stand out.

hypothesis-conditioned evidence grounding
Anchor-Grounded Semantic Retrieval (AGR)
Geometry-Consistent Anchor Refinement (GCR)
Doppler-Verified Correspondence (DVC)
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Xiaokai Bai
Xiaokai Bai
Zhejiang University Ph.D student
Multimodal Fusion3D object detection4D Radar Perceptionautonomous driving
Z
Zhenyu Fan
College of Information Science and Electronic Engineering, Zhejiang University
Lianqing Zheng
Lianqing Zheng
Tongji University Ph.D student
BEV/OCCVLA4D Radar PerceptionMultimodal FusionData Closed-Loop
S
Songkai Wang
College of Information Science and Electronic Engineering, Zhejiang University
Si-Yuan Cao
Si-Yuan Cao
Zhejiang University
image alignmenthomography estimationimage fusionplace recognition
H
Hui-liang Shen
College of Information Science and Electronic Engineering, Zhejiang University