From Imitation to Optimization: A Comparative Study of Offline Learning for Autonomous Driving

📅 2025-08-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In offline autonomous driving policy learning, behavioral cloning (BC) suffers from error accumulation, poor generalization, and insufficient robustness in closed-loop evaluation. To address these limitations, this paper proposes a systematic framework integrating imitation learning with conservative offline reinforcement learning (CQL). We introduce a Transformer-based entity-centric state representation and design a fine-grained reward function tailored for long-horizon driving tasks. Evaluated on 1,000 unseen scenarios from the Waymo Open Motion Dataset, our method achieves a 3.2× improvement in task success rate and a 7.4× reduction in collision rate over the best BC baseline. Moreover, it demonstrates significantly enhanced out-of-distribution state avoidance and error recovery capabilities. These results validate the effectiveness and practicality of conservative Q-learning for large-scale, real-world offline driving policy optimization.

Technology Category

Machine Learning: Imitation Learning & Inverse Reinforcement LearningHumans and AI: Human-Aware Planning and Behavior PredictionSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingResponsible Web: Machine-in-the-loop, human agency and autonomyEconomics, Online Markets and Human Computation: LLM based quality controls for crowd work
📝 Abstract
Learning robust driving policies from large-scale, real-world datasets is a central challenge in autonomous driving, as online data collection is often unsafe and impractical. While Behavioral Cloning (BC) offers a straightforward approach to imitation learning, policies trained with BC are notoriously brittle and suffer from compounding errors in closed-loop execution. This work presents a comprehensive pipeline and a comparative study to address this limitation. We first develop a series of increasingly sophisticated BC baselines, culminating in a Transformer-based model that operates on a structured, entity-centric state representation. While this model achieves low imitation loss, we show that it still fails in long-horizon simulations. We then demonstrate that by applying a state-of-the-art Offline Reinforcement Learning algorithm, Conservative Q-Learning (CQL), to the same data and architecture, we can learn a significantly more robust policy. Using a carefully engineered reward function, the CQL agent learns a conservative value function that enables it to recover from minor errors and avoid out-of-distribution states. In a large-scale evaluation on 1,000 unseen scenarios from the Waymo Open Motion Dataset, our final CQL agent achieves a 3.2x higher success rate and a 7.4x lower collision rate than the strongest BC baseline, proving that an offline RL approach is critical for learning robust, long-horizon driving policies from static expert data.
Problem

Research questions and friction points this paper is trying to address.

Learning robust driving policies from offline datasets
Overcoming brittleness of Behavioral Cloning in autonomous driving
Improving long-horizon performance with offline Reinforcement Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transformer-based model for structured state representation
Offline Reinforcement Learning with Conservative Q-Learning
Engineered reward function for error recovery
💼 Related Jobs
No related jobs found.
Independent Researcher