SlackDrive: Reclaiming Runtime Slack for Adaptive Driving Inference

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
针对驾驶推理中实时延迟与计算成本的冲突,提出SlackDrive方法,通过动态调整计算预算来优化模型执行效率。
📝 Abstract
Driving world-action models improve planning by coupling multimodal reasoning with future prediction, but their growing inference cost increasingly conflicts with the real-time latency requirements of vehicle control. Existing acceleration methods reduce tokens, layers, or sampling steps with policies selected prior to deployment, yet leave residual runtime variation largely unexploited after offline profiling and static scheduling on shared onboard compute. We observe that the largest admissible compute budget varies systematically with the residual runtime state, while recent realized latency provides a direct signal of the available compute slack. Motivated by this observation, we propose \textbf{SlackDrive}, a pre-inference compute allocator that reuses realized latency to select the compute budget of each control step before model execution. SlackDrive profiles the latency and planning utility of a small discrete budget set once, estimates online compute state from completed forwards, and selects the highest-utility budget predicted to remain within the admissible latency envelope, complementing existing profiling and resource scheduling while preserving the driving backbone and its compute actuator. On NAVSIM v2 with DriveDreamer-Policy, SlackDrive improves latency-constrained EPDMS by $21.7\%$ over the strongest baseline under a stringent latency regime, while the full-budget model and preconfigured token-pruning baselines exceed the admissible latency envelope under runtime contention.
Problem

Research questions and friction points this paper is trying to address.

inference cost
real-time latency
vehicle control
acceleration methods
runtime variation
Innovation

Methods, ideas, or system contributions that make the work stand out.

pre-inference compute allocator
realized latency
compute budget
adaptive driving inference
runtime slack
🔎 Similar Papers
No similar papers found.