Reinforcement learning for traffic signal control: advancing efficiency through hybrid exploration strategies.
Saved in:
| Title: | Reinforcement learning for traffic signal control: advancing efficiency through hybrid exploration strategies. |
|---|---|
| Authors: | Thadikamalla, Saidulu1 (AUTHOR) saidulu.t@iiits.in, Joshi, Piyush1 (AUTHOR) piyush.j@iiits.in, Gangadharan, Deepak2 (AUTHOR) deepak.g@iiit.ac.in |
| Source: | Journal of Supercomputing. Oct2025, Vol. 81 Issue 15, p1-36. 36p. |
| Abstract: | Efficient traffic signal control is critical for reducing urban congestion, yet traditional rule-based and machine learning approaches fail to adapt to dynamic conditions. Reinforcement Learning (RL) provides a data-driven alternative, and this study evaluates classical exploration strategies Epsilon-Greedy, Thompson Sampling, Upper Confidence Bound (UCB), and Softmax alongside hybrid variants including Softmax with Temperature Annealing (SMX-TA), Softmax with Entropy Regularization (SMX-ER), and their combinations with UCB and Thompson Sampling. Our framework is based on value-based RL (Double Deep Q-Networks), chosen for scalability and stability compared to on-policy methods such as PPO and A2C, and implemented with parallel simulation in SUMO and CityFlow across synthetic grids ( 1 × 1 , 4 × 4 , 6 × 6 ) and real-world datasets (Hyderabad, New York). Benchmarking against classical baselines (FixedTime, MaxPressure) and recent RL models (FRAP, CoLight, GCN) shows that the proposed SMX-ER+UCB consistently yields the lowest average travel times and robust adaptability across traffic conditions. An ablation study confirms the complementary benefits of entropy regularization and UCB, while integration into scalable controllers such as CoLight demonstrates network-level generalization. These results highlight hybrid exploration as an effective and practical strategy for real-time, city-wide traffic optimization. Such large-scale optimization requires high-performance computing (HPC) to support parallel simulations, GPU-accelerated training, and multi-agent coordination, underscoring the necessity of HPC frameworks for intelligent transportation systems. [ABSTRACT FROM AUTHOR] |
| Copyright of Journal of Supercomputing is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
| Abstract: | Efficient traffic signal control is critical for reducing urban congestion, yet traditional rule-based and machine learning approaches fail to adapt to dynamic conditions. Reinforcement Learning (RL) provides a data-driven alternative, and this study evaluates classical exploration strategies Epsilon-Greedy, Thompson Sampling, Upper Confidence Bound (UCB), and Softmax alongside hybrid variants including Softmax with Temperature Annealing (SMX-TA), Softmax with Entropy Regularization (SMX-ER), and their combinations with UCB and Thompson Sampling. Our framework is based on value-based RL (Double Deep Q-Networks), chosen for scalability and stability compared to on-policy methods such as PPO and A2C, and implemented with parallel simulation in SUMO and CityFlow across synthetic grids ( 1 × 1 , 4 × 4 , 6 × 6 ) and real-world datasets (Hyderabad, New York). Benchmarking against classical baselines (FixedTime, MaxPressure) and recent RL models (FRAP, CoLight, GCN) shows that the proposed SMX-ER+UCB consistently yields the lowest average travel times and robust adaptability across traffic conditions. An ablation study confirms the complementary benefits of entropy regularization and UCB, while integration into scalable controllers such as CoLight demonstrates network-level generalization. These results highlight hybrid exploration as an effective and practical strategy for real-time, city-wide traffic optimization. Such large-scale optimization requires high-performance computing (HPC) to support parallel simulations, GPU-accelerated training, and multi-agent coordination, underscoring the necessity of HPC frameworks for intelligent transportation systems. [ABSTRACT FROM AUTHOR] |
|---|---|
| ISSN: | 09208542 |
| DOI: | 10.1007/s11227-025-07892-6 |