🤖 AI Summary
This study addresses the loss of visual details and high inference latency in traffic signal control, which hinder real-time responses to emergency vehicles. To this end, it proposes a lightweight, end-to-end Vision-Language-Action (VLA) framework. By eliminating intermediate textual descriptions and handcrafted features, the method integrates intersection observations and phase information through unified multi-view inputs, leveraging a 0.5B-parameter large language model to directly map visual data to discrete signal actions. Experimental results demonstrate that the proposed framework achieves real-time inference on local hardware, reduces emergency vehicle waiting time by 21.1% compared to cascaded approaches, and exhibits strong generalization capabilities to unseen intersection topologies and traffic flow patterns.
📝 Abstract
Traffic signal control (TSC) is essential for mitigating urban congestion. Recent advances in vision-language models (VLMs) enable richer interpretation of intersection scenes, opening new opportunities for visual-context-aware TSC. However, the loose coupling and repeated information conversion between modules can lead to the loss of fine-grained visual details, while sequential inference introduces substantial latency. To address these limitations, we propose VLALight, a lightweight end-to-end vision-language-action framework that directly maps intersection observations and signal-phase information to discrete signal actions. To handle the multi-view nature of TSC, VLALight combines multiple directional camera views into a unified visual input and uses textual instructions to establish their correspondence with traffic movements and signal phases. This design enables direct action prediction with a compact 0.5 B-parameter model, without intermediate image-to-text descriptions or handcrafted traffic-state representations. Experiments show that VLALight delivers the best emergency-vehicle service of all compared methods, reducing pooled emergency waiting time by 21.1% over the cascaded VLMLight while running in real time on local hardware and generalizing to unseen intersection topologies and traffic-flow patterns.