Score
Designs, trains, and evaluates adaptive traffic signal control systems that compute and update signal phase timings (for example green-phase durations) using reinforcement learning or other adaptive policies. Builds controllers that operate on local traffic observations to reduce performance metrics such as per-vehicle delay, queue lengths, and emissions.
To address the challenges of dynamics, non-stationarity, and scalability in multi-agent collaborative decision-making within intelligent transportation systems (ITS), this paper proposes a unified classification framework for multi-agent reinforcement learning (MARL) tailored to ITS. The framework systematically categorizes MARL approaches into four paradigms: value-based, policy-gradient-based, actor-critic-based, and communication-enhanced methods. It further maps these to key ITS applications—including traffic signal control, cooperative autonomous driving, logistics dispatching, and on-demand mobility. Empirical evaluation is conducted across mainstream simulation platforms (SUMO, CARLA, CityFlow), identifying critical bottlenecks such as sim-to-real transfer, credit assignment, and environmental non-stationarity. This work establishes the first structured taxonomy that jointly considers algorithmic principles and traffic-domain semantics, providing a comprehensive survey, standardized benchmarks, and actionable research directions for advancing both the theoretical foundations and real-world deployment of MARL in ITS.
Adaptive traffic signal control via multi-agent reinforcement learning (MARL) faces deployment challenges in real-world corridors due to oversimplified fixed-time assumptions and predominant reliance on value-based methods, which struggle with eight-phase intersection constraints and partially observable environments. Method: This paper proposes a cooperative control framework based on Multi-Agent Proximal Policy Optimization (MA-PPO), the first to apply MA-PPO to multi-intersection coordination under full phase constraints. It employs a centralized critic to model complex phase sequences and directly outputs deployable dynamic signal plans within the Centralized Training with Decentralized Execution (CTDE) paradigm. Results: Validated via Vissim-MaxTime hardware-in-the-loop simulation and field tests on a seven-intersection corridor, the approach reduces bidirectional main-stream travel time by 14% and 29%, respectively, compared to actuated coordinated control (ASC), demonstrating significantly enhanced robustness and adaptability to dynamic traffic fluctuations.
To address the limitations of conventional fixed-time and actuated traffic signal control in adapting to dynamic traffic flows, this paper proposes a reinforcement learning (RL)-based adaptive traffic signal control framework that minimizes the total queue length across all signal phases. Methodologically, we design a lightweight yet expressive multi-dimensional state representation—integrating an extended state space, an autoencoder, and K-Planes-inspired feature encoding—and optimize the control policy using the Proximal Policy Optimization (PPO) algorithm. Training and evaluation are conducted in the SUMO simulation environment with a queue-length-driven sparse reward function. Experimental results show that, under optimal configuration, our approach reduces average queue length by 29% compared to the Webster method and significantly outperforms existing RL-based baselines. The key contributions are: (1) an efficient, low-overhead state representation mechanism; and (2) a robust, generalizable PPO training paradigm tailored for real-world deployment.
This work addresses the limitations of existing reinforcement learning (RL) approaches for traffic signal control, which often lack robustness and training efficiency under complex phase structures and dynamic traffic demands, hindering real-world deployment. The authors propose an RL-based adaptive signal control algorithm tailored to the standard eight-phase ring-and-barrier structure, enhanced by a distributed asynchronous training architecture to improve learning efficiency. The method is systematically evaluated across diverse traffic volumes and origin-destination patterns to assess its generalization capability. Experimental results demonstrate that, in realistic intersection scenarios, the proposed approach reduces vehicle delay by 11%–32% compared to optimized actuated control. Notably, it maintains significant performance advantages even under unseen, highly heterogeneous traffic demands, thereby providing the first empirical validation of RL controller robustness and practicality within a realistic eight-phase signal framework.
To address the challenges in deep reinforcement learning (DRL)-based signal control at complex intersections—including heavy reliance on handcrafted reward functions, opaque decision-making, and extensive domain expertise—this paper proposes an interpretable phase-based traffic signal control method grounded in *phase urgency*. We introduce a novel tree-structured urgency model that dynamically computes phase priorities from real-time traffic flow features. Instead of gradient-based optimization, we employ genetic programming to evolve human-readable, fully traceable control logic while preserving high performance. Evaluated across diverse SUMO simulation scenarios and multiple public benchmarks, our approach achieves an average 12.7% reduction in vehicle delay and a 9.4% improvement in throughput, significantly outperforming state-of-the-art traffic controllers and mainstream DRL algorithms. This work bridges the gap between high control performance and full algorithmic transparency.
To address the physical and behavioral heterogeneity between traffic signal controllers and vehicle platoons, as well as their coordination challenges in urban traffic, this paper proposes the first region-level cooperative decision-making framework for real-time joint optimization. Methodologically, we design a heterogeneous graph neural network–driven multi-agent reinforcement learning (HeteroGNN-MARL) architecture and introduce a novel alternating optimization training mechanism between signal controllers and platoon agents, enabling adaptive policy co-evolution under dynamic traffic conditions. Theoretical modeling integrates fundamental traffic flow principles, and extensive SUMO-based microscopic simulations demonstrate that our approach reduces average travel time by 18.7% and fuel consumption by 15.2% compared to state-of-the-art adaptive signal control methods. Our core contributions include: (i) the first unified modeling of heterogeneity across signal and platoon agents; (ii) the realization of real-time, closed-loop signal-platoon coordination; and (iii) an open-source, extensible paradigm for heterogeneous cooperative decision-making.
This work addresses the limited generalization of existing methods in complex scenarios by proposing a novel architecture based on adaptive feature fusion and dynamic inference. The approach introduces a learnable context-aware weighting module to effectively integrate multi-scale semantic information and incorporates a lightweight dynamic network to allocate computational resources on demand. Experimental results demonstrate that the model significantly outperforms state-of-the-art methods across multiple benchmark datasets, achieving higher accuracy and robustness while maintaining inference efficiency. This study offers a new perspective on efficient and adaptive visual understanding, exhibiting strong theoretical value and practical potential.
This work addresses the limited interpretability of deep reinforcement learning in adaptive traffic signal control, which hinders its deployment in safety-critical scenarios. The authors propose an entity-centric, interpretable reinforcement learning framework that models intersection states as lane entities structured by phase timing. By integrating a two-stage attention mechanism—comprising multi-head cross-attention and self-attention—the approach captures inter-lane dependencies while incorporating action masking to enforce safe and compliant signal control. This is the first method to combine structured entity representations with attention mechanisms to enable visualizable decision-making. Experimental results in microscopic traffic simulation demonstrate significant performance gains over existing approaches, substantially reducing vehicle delay. Moreover, the learned attention weights align closely with established traffic engineering principles, thereby enhancing the system’s auditability and trustworthiness.
Conventional state-feedback controllers exhibit insufficient adaptability to time-varying traffic congestion scenarios. Method: This paper proposes a parameterized traffic controller adaptive tuning framework based on multi-agent reinforcement learning (MARL). It innovatively decouples high-frequency control execution from low-frequency parameter optimization to jointly ensure real-time responsiveness and environmental adaptability, and adopts a distributed multi-agent architecture to enhance robustness against local failures and system scalability. Results: Evaluated across diverse traffic network simulations, the framework enables dynamic online parameter adjustment. It significantly outperforms both uncontrolled and fixed-parameter baselines, achieves performance comparable to single-agent RL approaches, and demonstrates superior resilience and faster recovery under partial agent failures.
This study addresses the deployment of reinforcement learning controllers in urban arterial signal networks to enhance traffic throughput. It systematically evaluates centralized, fully decentralized, and parameter-sharing decentralized multi-agent reinforcement learning strategies against the classical MaxPressure method in terms of capacity region and average travel time. The work proposes a novel, generalizable parameter-sharing decentralized architecture that enables agents to spontaneously generate coordinated "green wave" effects without explicit communication or coordination. Experimental results demonstrate that the proposed approach not only outperforms baseline methods significantly but also maintains superior performance when transferred to larger, previously unseen road networks, exhibiting both high efficiency and strong generalization capability.
This work addresses the challenge of generalizing reinforcement learning methods to diverse intersection topologies and dynamic traffic demands in large-scale traffic signal control. To this end, the authors propose CROSS, a novel framework that integrates predictive contrastive clustering (PCC) with a scenario-adaptive mixture-of-experts (MoE) mechanism. PCC identifies latent traffic patterns, while MoE dynamically generates tailored control policies based on these patterns. Leveraging decentralized reinforcement learning, CROSS is evaluated on both synthetic and real-world datasets within the SUMO simulation platform. The results demonstrate significant improvements over existing approaches, achieving state-of-the-art performance in both control efficacy and cross-scenario generalization.