planning-ready dataset synthesis

Designs and builds simulation-derived datasets and augmentation pipelines that produce planning-ready training and evaluation data such as simulated trajectories, sensor/label pairs, physics-based field outputs, and annotated event sequences. Ensures datasets are tractable for planners by covering diverse operating and boundary conditions, injecting reachability constraints, balancing event sampling and transition ambiguity, and producing the labels and annotations needed for planner training and validation.

planning-readydatasetsynthesis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$215K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing text-to-simulation methods for generating high-fidelity, executable rare traffic scenarios often suffer from semantic inaccuracies, insufficient iterative refinement, and limited robustness. This work proposes a multi-stage closed-loop framework that leverages large language models (LLMs) to progressively translate natural language descriptions into executable Scenic scripts. The generated scenarios are automatically validated in a simulation environment, which provides structured diagnostic feedback to guide iterative LLM refinement. This approach significantly enhances the semantic accuracy, executability, and diversity of the generated scenarios. Experimental results demonstrate superior performance over existing baselines in rare traffic scenario generation, highlighting improved reliability and automation capabilities.

executable scenario generationnatural language to simulationrare scenario synthesis

Generating Traffic Scenarios via In-Context Learning to Learn Better Motion Planner

Dec 24, 2024
AA
Aizierjiang Aiersilan
🏛️ University of Macau

To address the scarcity of rare safety-critical scenarios in autonomous driving motion planning, the high cost and limited coverage of manual annotation for long-tail risks, this paper proposes a traffic scenario generation method leveraging in-context learning (ICL) with large language models (LLMs). The approach requires no model fine-tuning or handcrafted programming; instead, it synthesizes executable CARLA simulation scripts directly from natural-language scenario descriptions, enabling low-cost, highly diverse, and customizable critical scenario construction. To our knowledge, this is the first work to apply ICL to script-based traffic scenario modeling, substantially improving the realism and generalizability of synthetic data. Experimental results demonstrate that motion planners trained on the synthesized data achieve a 23.6% improvement in success rate on real-world critical-risk scenario evaluation, with marked gains in safety and robustness.

Generating diverse critical traffic scenarios for robust motion plannersImproving motion planner performance with synthetic and real-world dataReducing human costs in manual scenario composition for autonomous driving

Existing autonomous driving datasets lack sufficient diversity, coordination, and cross-domain support, limiting their utility for training multi-agent, multi-sensor systems. To address this gap, this work proposes a modular data generation pipeline built upon the AVstack framework and the CARLA simulator, capable of efficiently producing terabyte-scale, ground-truth-annotated multimodal data. The pipeline encompasses perspectives from ground vehicles, aerial platforms, and infrastructure sensors, and supports flexible single- or multi-agent configurations under controllable, complex scenarios. This approach represents the first scalable, cross-domain collaborative data generation methodology for autonomous driving, substantially enhancing the customization, training efficacy, and practical applicability of perception and sensor fusion models in cooperative autonomous systems.

autonomous systemslarge-scale learningmulti-agent

Existing autonomous driving datasets lack sufficient real-time perception capability and robustness against edge cases in high-definition map–free scenarios. Method: This paper proposes a scenario- and capability-driven dataset construction and evaluation methodology, systematically deriving perception requirements from ISO 21448 (SOTIF) and ISO/TR 4804, and pioneering the deep integration of SOTIF principles throughout the dataset development lifecycle. It establishes a reusable “scenario–capability” mapping framework to support both novel dataset creation and cross-dataset benchmarking. Contribution/Results: Empirical analysis reveals systemic deficiencies in mainstream lane detection datasets—particularly in real-world scenario coverage, annotation of ambiguous drivable boundaries, and representation of complex driving behaviors. The proposed methodology significantly enhances dataset safety alignment with real-world operational conditions and improves generalization across diverse, unstructured environments.

Addresses scarcity in dataset development for automated driving perceptionIdentifies limitations in current lane detection datasets' real-world relevanceProposes scenario-based method to enhance dataset applicability and safety

Latest Papers

What's happening recently
View more

The scarcity of real-world data severely hinders the widespread adoption of subsymbolic AI. To address this challenge, this work proposes a unified reference framework based on digital twins to systematically design and analyze simulation-based synthetic data generation methods for AI training. By integrating digital twin technology, high-fidelity simulation, and synthetic data generation, the framework delineates core components, advantages, and key challenges, offering a methodological foundation for producing high-quality, reproducible training data. This study not only fills the critical gap in the lack of systematic guidance for synthetic data generation but also provides a scalable and reusable technical pathway to mitigate reliance on real-world data.

AI trainingdata qualitydata volume

Existing planning benchmark datasets suffer from limitations in scalability, controllability, and automated verifiability, hindering effective evaluation or training of large language models (LLMs) on complex planning tasks. This work proposes the first constraint-driven synthetic framework that leverages a structured taxonomy of task types and constraint families to generate planning problems on demand—offering controllable difficulty levels, diverse scenarios, and built-in automatic verifiability. By integrating a constraint-driven data synthesis pipeline, quality filtering mechanisms, and instance-level verification checklists, the framework shifts planning data construction from static collection to dynamic generation. Experiments reveal that current LLMs exhibit limited planning capabilities under coupled constraints, whereas reinforcement learning trained on our dataset substantially improves their performance on unseen tasks and general instruction following.

constraint-driven synthesisLLM evaluationplanning benchmarks

This work addresses the challenges of real-world data scarcity, high acquisition costs, and privacy sensitivity in multimodal AI training by introducing Simula, a novel framework that pioneers inference-driven synthetic data generation without requiring any seed data. By integrating an agent-based architecture with a controllable generation pipeline, Simula enables fine-grained control over data characteristics and computational resource allocation, substantially enhancing the interpretability and scalability of synthetic data. Through a comprehensive multidimensional evaluation protocol, the framework simultaneously validates both the intrinsic quality of the generated data and its effectiveness in downstream tasks across multiple benchmarks, offering a practical pathway and design paradigm for AI development under data-constrained conditions.

data generationdata scarcitymulti-modal models

This work addresses the lack of a general, auditable dynamic control mechanism in existing training systems, which typically rely on framework-specific code. The authors propose the first cross-framework, open-source control plane that exposes training interfaces through a unified protocol, integrating declarative configuration, request validation, and secure control-point scheduling within the Aim workspace to enable metric monitoring, real-time intervention, and operational traceability. The system supports safe human and automated controller interventions during training while fully logging all operational trajectories. Experiments across five NLP and reinforcement learning tasks demonstrate its effectiveness, and the open-source implementation provides a foundation for reproducible human-in-the-loop training.

auditable trainingcontrol planehuman-in-the-loop

Hot Scholars

TW

Tai Wang

Shanghai AI Laboratory
Computer Vision3D VisionEmbodied AIDeep Learning
ZW

Ziqin Wang

Beihang University
Embodied AIRoboticsLarge Lanugage ModelComputer Vision
HZ

Hengshuang Zhao

The University of Hong Kong
Computer VisionMachine LearningArtificial Intelligence
JI

Jeffrey Ichnowski

Carnegie Mellon University
RoboticsManipulationMotion Planning
JN

J. Nathan Kutz

Professor of Applied Mathematics & Electrical and Computer Engineering
Dynamical SystemsData ScienceMachine LearningOptics