RCVLA: 4D Radar-Grounded Semantic Reasoning and Trajectory Arbitration for Autonomous Driving

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the integration challenges between 4D radar and vision-language-action (VLA) models regarding semantic alignment and trajectory refinement by proposing the RCVLA framework alongside two newly constructed datasets. Methodologically, it introduces a pioneering bidirectional radar-language interaction mechanism that leverages a gated bidirectional Transformer with occupancy-aware queries to achieve radar-grounded semantic reasoning. Furthermore, a truncated diffusion model combined with radar risk calibration is incorporated to perform physical trajectory arbitration. Experimental results demonstrate that the proposed approach improves the CIDEr score by 9.92 points, reduces the velocity error of critical objects by 21.9%, and decreases the average collision rate to 0.175%.
📝 Abstract
4D radar provides geometric and motion cues that complement visual semantics, but integrating it into vision-language-action (VLA) models requires both radar--language alignment for semantic reasoning and explicit use of radar measurements for trajectory refinement and selection. To support these capabilities, we construct Cap4DR with 86,016 radar-image-text samples for alignment pretraining and OmniHD-QA with 520,161 question-answer pairs for instruction tuning across scene description, key-object reasoning, occupancy understanding, and trajectory planning. Building on these datasets, we propose RCVLA, a radar-camera VLA framework consisting of a radar-grounded semantic reasoning stage (RCVLA-Sem) and a trajectory arbitration stage (RCVLA-Phys). RCVLA-Sem performs gated bidirectional interaction between camera and radar tokens for driving question answering and reference trajectory generation, while auxiliary heads provide object and occupancy queries. RCVLA-Phys refines reference-guided trajectory candidates through truncated diffusion conditioned on these queries and cluster-level radar measurements, then calibrates candidate scores using radar-derived time-to-collision risk. On OmniHD-QA, RCVLA-Sem improves CIDEr by 9.92 points and reduces key-object velocity error by $21.9\%$ relative to OmniDrive. RCVLA-Phys further reduces average L2 error from $0.348$ to $0.259\,\mathrm{m}$ and average open-loop collision rate from $0.576\%$ to $0.175\%$ relative to RCVLA-Sem. Ablation studies further show that language-aligned radar tokens improve semantic reasoning, while cluster-level radar measurements and risk calibration improve trajectory arbitration. Code will be released.
Problem

Research questions and friction points this paper is trying to address.

4D radar
vision-language-action model
autonomous driving
semantic reasoning
trajectory planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

4D Radar
Vision-Language-Action
Semantic Reasoning
Trajectory Arbitration
Diffusion Model
🔎 Similar Papers
No similar papers found.