VLALight: A Vision-Language-Action Model for Traffic Signal Control

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the disconnection between visual observation and decision-making in existing traffic signal control methods, which typically rely on handcrafted states or isolated perception modules. We propose the first end-to-end Vision-Language-Action (VLA) model for traffic signal control that directly maps multi-view video inputs to coordinated signal actions. The model introduces a novel VLA architecture incorporating an adaptive fast-slow reasoning mechanism to balance decision quality with computational cost. Training integrates two-stage supervised cold-start initialization, cooperative multi-agent reinforcement learning, and topology-aware collaboration techniques. Extensive experiments across seven real-world datasets demonstrate that our approach significantly outperforms existing baselines, validating the potential of the VLA paradigm for physical traffic control applications.
📝 Abstract
Traffic signal control (TSC) is essential for improving urban mobility and reducing congestion. Although roadside cameras are widely deployed at signalized intersections and provide rich visual observations of evolving traffic, existing TSC methods typically rely on manually engineered traffic states or separate perception modules, creating a gap between physical observations and control decisions. We present VLALight, the first vision-language-action (VLA) model for end-to-end traffic signal control from multi-view roadside videos. VLALight directly maps visual observations to coordinated signal actions through multi-target spatiotemporal traffic reasoning and topology-aware cooperative perception across intersections. To establish this capability, we develop a two-stage supervised cold-start training strategy for visual traffic understanding and signal decision-making, followed by cooperative agentic reinforcement learning that jointly optimizes local control and network-wide traffic efficiency. Furthermore, VLALight introduces adaptive fast and slow reasoning modes, enabling the policy to allocate deeper reasoning only when additional deliberation provides sufficient control benefits. Through balanced mode-aware rollouts and relative advantage optimization, VLALight learns to trade off decision quality and inference cost. Extensive experiments on seven real-world traffic-flow datasets across three urban networks demonstrate that VLALight consistently outperforms transportation-based, RL-based, and LLM/VLM-based baselines. Ablation studies validate the effectiveness of cooperative perception, network-level optimization, and adaptive reasoning. These results demonstrate the potential of VLA models for real-world physical traffic control. Our project is available at https://github.com/usail-hkust/VLALight.git.
Problem

Research questions and friction points this paper is trying to address.

Traffic Signal Control
Vision-Language-Action Model
End-to-End Control
Cooperative Perception
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Model
Traffic Signal Control
Cooperative Perception
Agentic Reinforcement Learning
Adaptive Reasoning
💼 Related Jobs
No related jobs found.
P
Pan Zhang
The Hong Kong University of Science and Technology (Guangzhou)
Siqi Lai
Siqi Lai
Ph.D. student, The Hong Kong University of Science and Technology (Guangzhou)
Data MiningLLM AgentUrban Intelligence
K
Kemu Dong
Dalian University of Technology
H
Hao Liu
The Hong Kong University of Science and Technology (Guangzhou)