🤖 AI Summary
This study addresses the challenge of transferring vision-language-action models from passenger vehicles to the complex scenarios encountered by Class 8 trucks. To this end, we propose an "adapt-then-guide" framework. Methodologically, we perform targeted fine-tuning on Alpamayo 1.5 to adapt the action space and design a compact flow-time-conditioned residual module to achieve flow-velocity-guided trajectory correction. This framework substantially reduces data requirements, halving the Average Displacement Error (ADE) and significantly lowering the Final Displacement Error (FDE). Under equivalent data budgets, it achieves 19–26% lower error than generic supervised baselines, yielding performance comparable to models trained on considerably larger datasets. Overall, this work enables efficient domain transfer for autonomous trucking applications.
📝 Abstract
Class 8 trucks differ from passenger cars in geometry, dynamics, and maneuvering requirements. As a result, vision-language-action (VLA) models trained for passenger vehicles do not readily transfer to Class 8 trucks, particularly in unstructured scenarios such as accident scenes and construction zones. Rather than training a truck-driving VLA from scratch, we propose an adapt-then-steer strategy that adapts an off-the-shelf VLA to generate trajectories for Class-8 trucks in these challenging scenarios. In the adapt stage, we use NVIDIA's Alpamayo 1.5 as the base model, fine-tuning only its action-generation stack on a few hundred real-world construction and accident-related highway scenarios. In the steer stage, we introduce Flow Velocity Steering (FVS) to further refine the model's predictions while holding the adapted VLA fixed. FVS is a compact, flow-time-conditioned residual module that adds learned corrections to the action-space flow velocity used to update the action sequence at each generation step. In open-loop evaluation on a scenario-disjoint held-out set, targeted fine-tuning more than halves single-candidate average displacement error (ADE) and final displacement error (FDE) over the entire 6.4 s horizon compared to the base model. Using the same targeted demonstrations, FVS further reduces the fine-tuned model's full-horizon ADE and FDE by 13.9% and 16.5%, respectively. At matched data budgets, targeted supervision yields 19-26% lower full-horizon ADE than general truck-driving supervision, while the targeted model remains competitive with a model fine-tuned on approximately 65 times as many general truck-driving scenarios. These results support adapt-then-steer for data-efficient vehicle-domain transfer to Class 8 trucks. Our project website is available at https://truckvla.github.io.