SafeGen: Goal-Conditioned Video Diffusion of Safety-Critical Scenarios for VLM-Based Autonomous Driving

πŸ“… 2026-07-21
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge that existing simulation-based methods for generating safety-critical driving scenarios struggle to realistically reproduce high-risk interactions involving vulnerable road users due to the sim-to-real gap. To overcome this limitation, the authors propose a goal-conditioned diffusion framework that leverages catastrophic end states as strong supervisory signals. By integrating vision-language models to analyze normal driving contexts and infer interaction vulnerabilities, the method enables context-anchored end-state reasoning and end-state-conditioned video evolution, yielding temporally coherent and physically plausible high-risk scenarios. Evaluated across three VLMAD systems, the approach improves the Judge Overall Score by an average of 24.25% and enhances the fine-tuned models’ performance on real-world scenarios by 15.9% on average.
πŸ“ Abstract
VLMs are increasingly deployed in AD systems, creating an urgent need for rigorous safety evaluation under rare yet safety-critical scenarios. Among these, interactions with vulnerable road users represent a major source of real-world failures. However, existing safety-critical scenario generation methods predominantly rely on simulator-based pipelines, which suffer from a substantial sim-to-real gap and often fail to capture realistic, diverse, and unforeseen human-vehicle interaction dynamics. We present SafeGen, a goal-conditioned diffusion framework for safety-critical scenario generation in VLMADs. Our key insight is to formulate scenario generation as a goal-conditioned diffusion process, where a predefined catastrophic end-state serves as a strong supervisory signal, guiding the generation of temporally coherent video trajectories that naturally evolve toward safety-critical outcomes. Building on this formulation, we introduce Context Grounded End State Reasoning, which leverages VLMs to analyze benign driving contexts and infer latent vulnerabilities in human-vehicle interactions, producing structured end-state specifications that induce high-risk scenarios. Conditioned on these targets, we further propose End State Conditioned Video Evolution, which grounds semantic threats into physically plausible visual dynamics. Specifically, we instantiate high-risk agents within the scene via depth-aware geometric projection, followed by boundary-conditioned diffusion to generate intermediate frames with consistent motion patterns and temporal coherence. Extensive experiments across 3 VLMADs demonstrate that SafeGen increases the Judge Overall Score, a metric using a VLM judge to evaluate VLMADs' understanding and decision-making, by 24.25% on average compared to SoTA baselines. Furthermore, fine-tuning a VLMAD improves performance in real-world driving scenes by an average of 15.9%.
Problem

Research questions and friction points this paper is trying to address.

safety-critical scenarios
VLM-based autonomous driving
sim-to-real gap
vulnerable road users
human-vehicle interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

goal-conditioned diffusion
safety-critical scenario generation
vision-language models
end-state reasoning
video diffusion
πŸ”Ž Similar Papers
No similar papers found.