Chat2Scenic: An Iterative RAG-Based Framework for Scenario Generation in Autonomous Driving

📅 2026-07-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of low compilation success rates and poor scalability in existing methods for automatically generating compilable autonomous driving test scenario scripts from regulatory texts. The authors propose an interactive framework based on iterative Retrieval-Augmented Generation (RAG), which integrates regulatory knowledge with domain-specific language (DSL) syntax to progressively refine compliant and executable scenario scripts through a conversational interface. The key innovation lies in the first-time adoption of an iterative RAG architecture to enable interactive fine-tuning, alongside the construction of an open-source benchmark dataset comprising 123 regulatory scenarios. Experimental results demonstrate that the proposed approach significantly outperforms current methods, achieving a compilation success rate of 76.42% and a framework accuracy of 58.17%. The code has been made publicly available to foster further research.
📝 Abstract
Validating autonomous driving systems requires diverse, regulation-compliant test scenarios. In simulation-based testing, scenarios are defined as executable scripts. Yet automatically generating such scripts from regulatory descriptions remains an open challenge, and existing approaches face fundamental trade-offs. Retrieval-assemble methods achieve reasonable compilation rates but lack scalability, whereas retrieval-based full-script generation suffers from low compilation success rates. We present Chat2Scenic, the first iterative retrieval-augmented framework to generate scenario scripts in Domain Specific Language (DSL). Specifically, Chat2Scenic provides a chatbot interface that supports interactive scenario refinement and integrates Retrieval-augmented Generation (RAG) to ground scenario generation in regulatory knowledge and DSL syntax. Furthermore, we propose an open benchmark for scenario generation comprising 123 scenarios from various regulations, including NHTSA and United Nations Vehicle Regulations, as well as other sources. Extensive evaluation with State-of-the-Art (SOTA) Large Language Models (LLMs) demonstrates that Chat2Scenic achieves 76.42% Compilation Success Rate (CSR) and 58.17% Framework Accuracy (FA), outperforming existing methods (Retrieval Assemble with 30.08% CSR, 11.03% FA and Retrieval full script generation with 16.26% CSR, 10.86% FA). To facilitate future research, we release our code as open source at https://github.com/TUM-AVS/chat2scenic.
Problem

Research questions and friction points this paper is trying to address.

autonomous driving
scenario generation
regulatory compliance
Domain Specific Language
simulation-based testing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Retrieval-Augmented Generation
Domain Specific Language
Scenario Generation
Autonomous Driving
Iterative Framework
Yuan Gao
Yuan Gao
Research Associate, Technical University of Munich
Autonomous DrivingFoundation ModelsScenario GenerationMotion planningControl
W
Wenting Miao
Professorship of Autonomous Vehicle Systems, TUM School of Engineering and Design, Technical University of Munich, 85748 Garching, Germany; Munich Institute of Robotics and Machine Intelligence (MIRMI)
Mattia Piccinini
Mattia Piccinini
TUM Global Post-doc Researcher, Technical University of Munich
Autonomous VehiclesArtificial IntelligenceRoboticsTrajectory PlanningMotion Control
H
Haoyu Wang
Professorship of Autonomous Vehicle Systems, TUM School of Engineering and Design, Technical University of Munich, 85748 Garching, Germany; Munich Institute of Robotics and Machine Intelligence (MIRMI)
Q
Qunying Song
University College London, London, United Kingdom
Johannes Betz
Johannes Betz
Professor, Autonomous Vehicle Systems, Technical University of Munich (TUM)
Autonomous SystemsMotion PlaningControlRobots