🤖 AI Summary
This work addresses the high computational redundancy and inference latency caused by multi-inference trajectory generation in end-to-end autonomous driving, which hampers real-time system performance. The authors propose a single-inference architecture that restructures the action generation mechanism of diffusion models to significantly reduce latency while preserving trajectory diversity and prediction quality. They demonstrate that the single-inference approach incurs no substantial loss in diversity and introduce runtime optimization strategies to eliminate inter-block copy overhead and inefficient kernel execution. Experimental results show that the method achieves a 69.23% reduction in inference latency in both closed-loop and open-loop scenarios, effectively balancing efficiency and performance.
📝 Abstract
Reasoning-based end-to-end (E2E) autonomous driving has recently emerged as a promising approach to improving the interpretability of driving decisions as it can generate human-readable reasoning together with predicted trajectories. Such approaches commonly generate multiple trajectories to capture diverse future behaviors, and they fall into two categories: (1) multi-reasoning, where one reasoning sequence is generated per trajectory, and (2) single-reasoning, where a single reasoning is shared across all trajectories. The former offers richer diversity at the cost of redundant computation, while the latter is more efficient but is often assumed to sacrifice diversity. Alpamayo 1, a representative system, adopts the multi-reasoning approach and achieves competitive trajectory prediction performance. However, the efficiency of this design remains largely unexplored, making it a well-motivated subject for investigation. In this paper, we systematically analyze and improve Alpamayo 1 in two ways. First, we reduce inference latency while preserving trajectory diversity by redesigning Alpamayo 1 into a single-reasoning system. Through extensive experiments, we find that replacing multi-reasoning with single-reasoning does not meaningfully degrade trajectory diversity. Second, we accelerate diffusion-based action generation by eliminating inter-block overhead arising from unnecessary copy operations and inefficient kernel execution. Through closed-loop and open-loop experiments, we validate both optimizations, demonstrating a 69.23% reduction in inference latency while maintaining trajectory diversity and prediction quality. These results highlight the importance of jointly analyzing system architecture and runtime execution to improve the efficiency of reasoning-based E2E AD systems.