VLALight: Lightweight Vision-Language-Action Models for Emergency-Aware Traffic Signal Control

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the loss of visual details and high inference latency in traffic signal control, which hinder real-time responses to emergency vehicles. To this end, it proposes a lightweight, end-to-end Vision-Language-Action (VLA) framework. By eliminating intermediate textual descriptions and handcrafted features, the method integrates intersection observations and phase information through unified multi-view inputs, leveraging a 0.5B-parameter large language model to directly map visual data to discrete signal actions. Experimental results demonstrate that the proposed framework achieves real-time inference on local hardware, reduces emergency vehicle waiting time by 21.1% compared to cascaded approaches, and exhibits strong generalization capabilities to unseen intersection topologies and traffic flow patterns.
📝 Abstract
Traffic signal control (TSC) is essential for mitigating urban congestion. Recent advances in vision-language models (VLMs) enable richer interpretation of intersection scenes, opening new opportunities for visual-context-aware TSC. However, the loose coupling and repeated information conversion between modules can lead to the loss of fine-grained visual details, while sequential inference introduces substantial latency. To address these limitations, we propose VLALight, a lightweight end-to-end vision-language-action framework that directly maps intersection observations and signal-phase information to discrete signal actions. To handle the multi-view nature of TSC, VLALight combines multiple directional camera views into a unified visual input and uses textual instructions to establish their correspondence with traffic movements and signal phases. This design enables direct action prediction with a compact 0.5 B-parameter model, without intermediate image-to-text descriptions or handcrafted traffic-state representations. Experiments show that VLALight delivers the best emergency-vehicle service of all compared methods, reducing pooled emergency waiting time by 21.1% over the cascaded VLMLight while running in real time on local hardware and generalizing to unseen intersection topologies and traffic-flow patterns.
Problem

Research questions and friction points this paper is trying to address.

Traffic Signal Control
Vision-Language Models
Emergency Vehicle
Inference Latency
Multi-view
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Models
Traffic Signal Control
End-to-End Framework
Lightweight Model
Emergency Vehicle Priority
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kemou Jiang
School of Transportation Science and Engineering, Beihang University, Beijing, China
Maonan Wang
Maonan Wang
Unknown affiliation
Xingchen Zou
Xingchen Zou
The University of Hong Kong
LLM/VLM AgentAI4ScienceUrban Computing
J
Jiayue Zhu
School of Transportation Science and Engineering, Beihang University, Beijing, China
Y
Yuhang Fu
School of Transportation Science and Engineering, Beihang University, Beijing, China
S
Sicheng Wang
School of Transportation Science and Engineering, Beihang University, Beijing, China
Xi Chen
Xi Chen
Lingnan University, Hong Kong
Mechanics
Yirong Chen
Yirong Chen
Stanford University
Zhiyong Cui
Zhiyong Cui
Professor, Beihang University
Foundation ModelsAutonomous DrivingUrban ComputingTraffic PredictionTraffic Control