TaCarla: A comprehensive benchmarking dataset for end-to-end autonomous driving

📅 2026-02-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing autonomous driving datasets lack behavioral diversity and do not support closed-loop evaluation, limiting their utility for end-to-end model development. To address this gap, this work leverages the CARLA simulation platform to collect over 2.85 million frames of multi-sensor data across the diverse scenarios of Leaderboard 2.0, establishing the first comprehensive benchmark that unifies perception, prediction, planning, and vision-language action modeling. The dataset introduces a state rarity score to quantify scene distribution and supports both open-loop and closed-loop evaluation. Experimental results demonstrate its versatility and effectiveness across a range of perception and planning tasks, offering a high-quality, multifunctional benchmark resource to advance end-to-end autonomous driving research.

Technology Category

Humans and AI: Human-Aware Planning and Behavior PredictionIntelligent Robots: Multimodal Perception & Sensor FusionPlanning, Routing, and Scheduling: Planning with Language Models

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSecurity and Privacy: Data transparency and provenance
📝 Abstract
Collecting a high-quality dataset is a critical task that demands meticulous attention to detail, as overlooking certain aspects can render the entire dataset unusable. Autonomous driving challenges remain a prominent area of research, requiring further exploration to enhance the perception and planning performance of vehicles. However, existing datasets are often incomplete. For instance, datasets that include perception information generally lack planning data, while planning datasets typically consist of extensive driving sequences where the ego vehicle predominantly drives forward, offering limited behavioral diversity. In addition, many real datasets struggle to evaluate their models, especially for planning tasks, since they lack a proper closed-loop evaluation setup. The CARLA Leaderboard 2.0 challenge, which provides a diverse set of scenarios to address the long-tail problem in autonomous driving, has emerged as a valuable alternative platform for developing perception and planning models in both open-loop and closed-loop evaluation setups. Nevertheless, existing datasets collected on this platform present certain limitations. Some datasets appear to be tailored primarily for limited sensor configuration, with particular sensor configurations. To support end-to-end autonomous driving research, we have collected a new dataset comprising over 2.85 million frames using the CARLA simulation environment for the diverse Leaderboard 2.0 challenge scenarios. Our dataset is designed not only for planning tasks but also supports dynamic object detection, lane divider detection, centerline detection, traffic light recognition, prediction tasks and visual language action models . Furthermore, we demonstrate its versatility by training various models using our dataset. Moreover, we also provide numerical rarity scores to understand how rarely the current state occurs in the dataset.
Problem

Research questions and friction points this paper is trying to address.

autonomous driving
dataset
perception
planning
benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

end-to-end autonomous driving
CARLA simulation
comprehensive benchmarking dataset
closed-loop evaluation
rarity scoring
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.