Reinforcement learning for traffic signal control: advancing efficiency through hybrid exploration strategies.

Saved in:
Bibliographic Details
Title: Reinforcement learning for traffic signal control: advancing efficiency through hybrid exploration strategies.
Authors: Thadikamalla, Saidulu1 (AUTHOR) saidulu.t@iiits.in, Joshi, Piyush1 (AUTHOR) piyush.j@iiits.in, Gangadharan, Deepak2 (AUTHOR) deepak.g@iiit.ac.in
Source: Journal of Supercomputing. Oct2025, Vol. 81 Issue 15, p1-36. 36p.
Abstract: Efficient traffic signal control is critical for reducing urban congestion, yet traditional rule-based and machine learning approaches fail to adapt to dynamic conditions. Reinforcement Learning (RL) provides a data-driven alternative, and this study evaluates classical exploration strategies Epsilon-Greedy, Thompson Sampling, Upper Confidence Bound (UCB), and Softmax alongside hybrid variants including Softmax with Temperature Annealing (SMX-TA), Softmax with Entropy Regularization (SMX-ER), and their combinations with UCB and Thompson Sampling. Our framework is based on value-based RL (Double Deep Q-Networks), chosen for scalability and stability compared to on-policy methods such as PPO and A2C, and implemented with parallel simulation in SUMO and CityFlow across synthetic grids ( 1 × 1 , 4 × 4 , 6 × 6 ) and real-world datasets (Hyderabad, New York). Benchmarking against classical baselines (FixedTime, MaxPressure) and recent RL models (FRAP, CoLight, GCN) shows that the proposed SMX-ER+UCB consistently yields the lowest average travel times and robust adaptability across traffic conditions. An ablation study confirms the complementary benefits of entropy regularization and UCB, while integration into scalable controllers such as CoLight demonstrates network-level generalization. These results highlight hybrid exploration as an effective and practical strategy for real-time, city-wide traffic optimization. Such large-scale optimization requires high-performance computing (HPC) to support parallel simulations, GPU-accelerated training, and multi-agent coordination, underscoring the necessity of HPC frameworks for intelligent transportation systems. [ABSTRACT FROM AUTHOR]
Copyright of Journal of Supercomputing is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: egs
DbLabel: Engineering Source
An: 188463077
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Reinforcement learning for traffic signal control: advancing efficiency through hybrid exploration strategies.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Thadikamalla%2C+Saidulu%22">Thadikamalla, Saidulu</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> saidulu.t@iiits.in</i><br /><searchLink fieldCode="AR" term="%22Joshi%2C+Piyush%22">Joshi, Piyush</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> piyush.j@iiits.in</i><br /><searchLink fieldCode="AR" term="%22Gangadharan%2C+Deepak%22">Gangadharan, Deepak</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> deepak.g@iiit.ac.in</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Journal+of+Supercomputing%22">Journal of Supercomputing</searchLink>. Oct2025, Vol. 81 Issue 15, p1-36. 36p.
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Efficient traffic signal control is critical for reducing urban congestion, yet traditional rule-based and machine learning approaches fail to adapt to dynamic conditions. Reinforcement Learning (RL) provides a data-driven alternative, and this study evaluates classical exploration strategies Epsilon-Greedy, Thompson Sampling, Upper Confidence Bound (UCB), and Softmax alongside hybrid variants including Softmax with Temperature Annealing (SMX-TA), Softmax with Entropy Regularization (SMX-ER), and their combinations with UCB and Thompson Sampling. Our framework is based on value-based RL (Double Deep Q-Networks), chosen for scalability and stability compared to on-policy methods such as PPO and A2C, and implemented with parallel simulation in SUMO and CityFlow across synthetic grids ( 1 × 1 , 4 × 4 , 6 × 6 ) and real-world datasets (Hyderabad, New York). Benchmarking against classical baselines (FixedTime, MaxPressure) and recent RL models (FRAP, CoLight, GCN) shows that the proposed SMX-ER+UCB consistently yields the lowest average travel times and robust adaptability across traffic conditions. An ablation study confirms the complementary benefits of entropy regularization and UCB, while integration into scalable controllers such as CoLight demonstrates network-level generalization. These results highlight hybrid exploration as an effective and practical strategy for real-time, city-wide traffic optimization. Such large-scale optimization requires high-performance computing (HPC) to support parallel simulations, GPU-accelerated training, and multi-agent coordination, underscoring the necessity of HPC frameworks for intelligent transportation systems. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Journal of Supercomputing is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=188463077
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s11227-025-07892-6
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 36
        StartPage: 1
    Titles:
      – TitleFull: Reinforcement learning for traffic signal control: advancing efficiency through hybrid exploration strategies.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Thadikamalla, Saidulu
      – PersonEntity:
          Name:
            NameFull: Joshi, Piyush
      – PersonEntity:
          Name:
            NameFull: Gangadharan, Deepak
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 10
              Text: Oct2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 09208542
          Numbering:
            – Type: volume
              Value: 81
            – Type: issue
              Value: 15
          Titles:
            – TitleFull: Journal of Supercomputing
              Type: main
ResultId 1