Odyssey: A Closed-Loop Benchmark for Long-Horizon Real-World Driving with Explicit Navigation Routes

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing closed-loop evaluations, which are constrained by short clips and ambiguous instructions that fail to capture the consequences of long-horizon driving decisions. Building upon nuPlan, this work constructs a hundred-scenario long-horizon closed-loop benchmark that unifies navigation objectives by replacing vague instructions with explicit SD map routes. It introduces route compliance metrics and a RouteDS score, while integrating 3D Gaussian Splatting with diffusion models to enhance simulation rendering and eliminate trajectory artifacts. The research reveals open challenges in route representation and system integration for end-to-end autonomous driving. Furthermore, the benchmark code and a Vision-Language-Action (VLA) adaptation baseline are released as open source.
📝 Abstract
Closed-loop evaluation of end-to-end driving requires continuous rollouts that reveal how earlier decisions affect subsequent driving. However, existing benchmarks evaluate only short segments and fail to capture later consequences. Ambiguous directional commands also obscure the intended navigation objective. We introduce Odyssey, a closed-loop benchmark for long-horizon driving comprising 100 scenarios, each reconstructed from a 100-second nuPlan driving log to preserve the context of navigation maneuvers and traffic interactions. To provide a consistent navigation objective, Odyssey replaces directional commands with explicit standard-definition (SD) map routes that specify which roads to follow, while sensor-based planning determines local driving actions. Throughout these rollouts, diffusion-based refinement of 3DGS-rendered images reduces rendering artifacts along the ego trajectory. To assess how effectively planners follow these routes and prepare for upcoming maneuvers, we introduce SD Route Compliance and Pre-Lane Change Score. These assessments are complemented by RouteDS, which extends the Driving Score with penalties for SD-route deviations and failed lane preparation. We adapt state-of-the-art planners, including vision-language-action (VLA) models, and evaluate their navigation performance using these metrics. Odyssey highlights open questions in route representation and integration for E2E driving. Benchmark code and adapted baselines will be released publicly.
Problem

Research questions and friction points this paper is trying to address.

closed-loop evaluation
end-to-end driving
long-horizon driving
navigation routes
autonomous driving benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Closed-loop benchmark
Long-horizon driving
SD map routes
3D Gaussian Splatting
Vision-language-action models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jungho Kim
Seoul National University, Republic of Korea
H
Hongjae Shin
Seoul National University, Republic of Korea
S
Seunghoon Yu
Seoul National University, Republic of Korea
Heecheol Yoo
Heecheol Yoo
MORAI Inc.
Computer VisionAutonomous Driving
M
Myeongjun Kim
Seoul National University, Republic of Korea
J
Jiyong Oh
Seoul National University, Republic of Korea
D
Donghyuk Kwak
Seoul National University, Republic of Korea
S
Seunghyeop Nam
Seoul National University, Republic of Korea
H
Haesung Oh
Seoul National University, Republic of Korea
Hyunju Kim
Hyunju Kim
Information Science, Cornell University
HCI
H
Hyungchan Cho
Seoul National University, Republic of Korea
Jaehyun Park
Jaehyun Park
Seoul National University, Republic of Korea
S
Soo Won Seo
Seoul National University, Republic of Korea
Jun Won Choi
Jun Won Choi
Department of Electrical and Computer Engineering, Seoul National University
Artificial intelligenceAutonomous driving robots/vehiclesRobot perceptionSensor fusion