Mining the Long Tail: A Comparative Study of Data-Centric Criticality Metrics for Robust Offline Reinforcement Learning in Autonomous Motion Planning

📅 2025-08-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Offline reinforcement learning for autonomous driving motion planning suffers from poor policy robustness and safety due to scarcity of long-tail scenario data. To address this, we propose a data-driven criticality-weighted sampling framework that systematically evaluates six heuristic-, uncertainty-, and behavior pattern-based sampling strategies at both the timestep and full-scenario levels. We further integrate an attention mechanism with goal-conditioned Conservative Q-Learning (CQL) to enhance generalization and safety-aware decision-making. Experiments are conducted in the high-fidelity Waymax simulator using Waymo data. Results demonstrate that all proposed sampling methods significantly outperform uniform sampling: notably, model-uncertainty-based sampling reduces collision rate from 16.0% to 5.5%, yielding nearly threefold improvement in safety performance. Moreover, our analysis uncovers a fundamental trade-off between reactive safety and long-horizon planning—highlighting inherent limitations in current offline RL paradigms for autonomous driving.

Technology Category

Search and Optimization: Sampling/Simulation-based SearchPlanning, Routing, and Scheduling: Learning for Planning and SchedulingHumans and AI: Human-Aware Planning and Behavior Prediction

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingResponsible Web: Machine-in-the-loop, human agency and autonomyEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasets
📝 Abstract
Offline Reinforcement Learning (RL) presents a promising paradigm for training autonomous vehicle (AV) planning policies from large-scale, real-world driving logs. However, the extreme data imbalance in these logs, where mundane scenarios vastly outnumber rare "long-tail" events, leads to brittle and unsafe policies when using standard uniform data sampling. In this work, we address this challenge through a systematic, large-scale comparative study of data curation strategies designed to focus the learning process on information-rich samples. We investigate six distinct criticality weighting schemes which are categorized into three families: heuristic-based, uncertainty-based, and behavior-based. These are evaluated at two temporal scales, the individual timestep and the complete scenario. We train seven goal-conditioned Conservative Q-Learning (CQL) agents with a state-of-the-art, attention-based architecture and evaluate them in the high-fidelity Waymax simulator. Our results demonstrate that all data curation methods significantly outperform the baseline. Notably, data-driven curation using model uncertainty as a signal achieves the most significant safety improvements, reducing the collision rate by nearly three-fold (from 16.0% to 5.5%). Furthermore, we identify a clear trade-off where timestep-level weighting excels at reactive safety while scenario-level weighting improves long-horizon planning. Our work provides a comprehensive framework for data curation in Offline RL and underscores that intelligent, non-uniform sampling is a critical component for building safe and reliable autonomous agents.
Problem

Research questions and friction points this paper is trying to address.

Addressing data imbalance in autonomous driving logs for offline reinforcement learning
Evaluating data curation strategies to focus learning on critical scenarios
Improving safety and reliability in autonomous vehicle planning policies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Data curation strategies focusing on information-rich samples
Criticality weighting schemes including heuristic and uncertainty-based
Non-uniform sampling significantly reduces collision rates
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Independent Researcher