Submitted:
24 June 2026
Posted:
24 June 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- A hybrid RL approach that integrates Q-learning, SARSA, and Monte Carlo updates into a single framework.
- A simulation-based traffic signal control model using a dynamic and stochastic environment.
- A comprehensive performance evaluation using multiple traffic efficiency metrics.
- A comparative analysis demonstrating the effectiveness of the proposed method over classical RL approaches.
2. Background and Preliminaries
2.1. Reinforcement Learning
2.2. MDP: Markov Decision Process
- is the space of environmental states.
- is the space of actions the agent can take.
- P is a transition function (the probability of ending in state , given that action a is taken in state s.can be represented by a conditional probability mass function denoted by ).
- : The reward function: provides the expected reward conditioned on being in a specific state s and taking a particular action a.
- is a discount factor that specifies how much immediate rewards are favored over future rewards., When is equal to 1, it indicates that future rewards are considered as important as present rewards. Conversely, when is equal to 0, it means that only present rewards are valued.
2.3. Policies and value function
3. Traffic Signal Control Problem
3.1. TSC Types
3.2. Reinforcement Learning for TSC
- Environment: The traffic network constitutes the environment in which the RL agent operates. It may include intersections, vehicle and pedestrian flows, and external disturbances like weather or sudden congestion. The RL agent does not directly control this environment but learns to react to its dynamics by optimizing signal timings (Figure 3).
- RL Agent: The RL agent is the brain of the system, it receives the incoming traffic data and decides the best signal phases. It learns to minimize congestion, delays and emissions through continuous interaction with the environment depending on the method used.
- Action: Traffic light control systems offer options such as timing adjustment and synchronized green light lane design choices. Once an action is selected, the system performs signal adjustments (e.g. green time extension, phase skipping, cycle length adjustment, coordination of multiple intersections) as per the agent output.
- Environment Update: After applying the control decisions, the environment evolves. Vehicles move, congestion shifts, and new traffic enters, resulting in a new state that becomes the input for the agent’s next decision.
- Reward Calculation: The system evaluates the agent’s decision using predefined metrics. Common reward functions are reduction in waiting time, queue lengths, fuel consumption, emissions, and enhancement of pedestrian safety and traffic balance.
- Reward Signal Feedback: The computed reward is fed back into the RL algorithm, refining its policy. Through this iterative process, the agent improves its control strategy, enabling real-time, adaptive signal management.
3.3. RL Algorithms for Traffic Signal Control
3.3.1. Value-Based Algorithms
-
Q-Learning: An off-policy algorithm that iteratively updates an action-value function independently of the current policy. Suitable for single-intersection control, it learns state-action values via:Q is efficient in discrete environments but struggles with high-dimensional networks.
- Deep Q-Networks: DQN enhances Q-learning by approximating the Q-function with deep neural networks (NN). Techniques such as experience replay and target networks improve training stability, enabling coordination across multiple intersections. The update rule is given by:
3.3.2. Policy-Based Algorithms
- Proximal Policy Optimization (PPO) is a RL algorithm that constrains policy updates to maintain training stability. It is commonly employed in multi-intersection control systems and effectively balances exploration and exploitation within large-scale networks. The clipped objective function may be expressed as:where : and denotes the advantage function [9].
3.3.3. Actor-Critic Algorithms
-
Deep Deterministic Policy Gradient: The DDPG was specifically designed for continuous action spaces, making it well-suited to tasks such as the dynamic adjustment of signal phase durations at intersections. Its operation relies on two complementary, iterative updates. Firstly, the critic refines its estimate of the state-action value:Secondly, the actor updates its policy by maximising the gradient of this value:While powerful, DDPG is sensitive to tuning.
- Soft Actor-Critic (SAC): To address this sensitivity, SAC enriches the objective function by incorporating an entropy term. This modification naturally encourages broader exploration of control strategies, thereby improving robustness in potentially non-stationary multi-agent environments. The resulting optimisation objective takes the form:where the coefficient governs the trade-off between exploration and exploitation.
-
Advantage Actor-Critic (A2C): ntroduces yet another perspective through the notion of advantage, which quantifies the extent to which a given action outperforms the average expected return in a given state:By substituting this centred signal for the raw reward in the policy update, A2C reduces gradient estimate variance, yielding more stable learning dynamics:The A2C achieves strong results in real-time adaptive control, particularly in collaborative multi-agent architectures where variance reduction plays a decisive role in ensuring convergence.
4. Results and Discussion
4.1. Experimental Setup
4.2. Proposed Method
4.3. Traffic Performance Evaluation
- 1.
- The Average Delay measures the mean time vehicles spend waiting before passing through the intersection, reflecting overall congestion levels;
- 2.
- The Throughput represents the number of vehicles successfully processed per unit time, indicating the efficiency of traffic flow;
- 3.
- The Queue Length corresponds to the number of vehicles waiting in each lane or approach, providing insight into spatial congestion dynamics;
- 4.
- The Waiting Time captures the total time each vehicle spends in the system before being served, offering a microscopic view of individual vehicle experience.
4.4. Results
4.4.1. Learning Curves
4.4.2. Convergence Analysis
4.5. Comparative Analysis of RL Algorithms
5. Discussion
- Q-learning: fast convergence,
- SARSA: stability,
- Monte Carlo: long-term return estimation.
- Average Delay: The hybrid method achieves the lowest average delay among all evaluated algorithms. This indicates its superior capability in minimizing congestion and optimizing traffic signal timing. Q-learning performs competitively but may occasionally lead to suboptimal decisions due to value overestimation.
- Throughput: The throughput results indicate that the hybrid approach allows a higher number of vehicles to pass through the intersection. This reflects a more efficient allocation of green signal phases. SARSA also provides consistent throughput but remains slightly less efficient than the hybrid method.
- Queue Length: Queue length is a key indicator of congestion levels. The hybrid method maintains shorter and more stable queue lengths compared to other approaches. Monte Carlo exhibits higher variability, resulting in fluctuating queue sizes.
- Waiting Time: The total waiting time is significantly reduced using the hybrid approach. This improvement highlights the ability of the proposed method to dynamically adapt signal control decisions based on real-time traffic conditions.
6. Conclusion
Acknowledgments
Abbreviations
| RL | Reinforcement learning |
| DRL | Deep Reinforcement learning |
| TSC | Traffic Signal Control |
| MDP | Markov Decision Process |
| PPO | Proximal Policy Optimization |
| MC | Monte carlo |
| 1 | For a more comprehensive introduction to MDPs, please refer to [1]. |
References
- Richard S. Sutton and Andrew G. Barto, Reinforcement Learning: An Introduction; Second Edition. MIT Press, Cambridge, MA, 2018.
- A. Ben Moussa and A. Khazari, Reinforcement Learning for Autonomous Driving: Optimization Strategiesand Methodologies. Artificial Intelligence and Mathematics; Springer Nature Switzerland,2026, 369–385.
- P. Michailidis, I. Michailidis, and E. Kosmatopoulos, Traffic Signal Control via Reinforcement Learning: A Review on Applications and Innovations. Infrastructures 2025, 10, 114,2025. [CrossRef]
- Y. Elbaum, A. Novoselsky and E. Kagan. A Queueing Model for Traffic Flow Control in the Road Intersection, Mathematics 2022, 10(21), 3997. [CrossRef]
- Hashmi, H.T., Ud-Din, S., Khan, M.A., Khan, J.A., Arshad, M., and Hassan, M.U.,“Traffic Flow Optimization at Toll Plaza Using Proactive Deep Learning Strategies,” Infrastructures, Vol. 9, Article 87, pp. 1–18, 2024.
- S. El-Tantawy, B. Abdulhai, and H. Abdelgawad, Multiagent Reinforcement Learning for Integrated Network of Adaptive Traffic Signal Controllers (MARLIN-ATSC): Methodology and Large-Scale Application on Downtown Toronto;IEEE Transactions on Intelligent Transportation Systems, Vol. 14, No. 3, pp. 1140–1150, 2013.
- H. Wei, G. Zheng, H. Yao, and Z. Li, IntelliLight: A Reinforcement Learning Approach for Intelligent Traffic Light Control; In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD 2018), London, UK, pp. 2496–2505, 2018.
- N. Rouphail, A. Tarko, and J. Li, Traffic Flow at Signalized Intersections; Federal Highway Administration, Washington, DC, USA, 1992.
- H. Wei, X. Liu, L. Mashayekhy, and K. Decker, Mixed-Autonomy Traffic Control with Proximal Policy Optimization; In Proceedings of the 2019 IEEE Vehicular Networking Conference (VNC), Los Angeles, CA, USA, pp. 1–8, 2019.
- Transportation Research Board, Highway Capacity Manual; National Academy of Sciences, Washington, DC, USA, 2000.
- B. Abdulhai, R. Pringle, and G.J. Karakoulas, Reinforcement Learning for True Adaptive Traffic Signal Control; Journal of Transportation Engineering, Vol. 129, No. 3, pp. 278–285, 2003.
- M. Guo, P. Wang, C.-Y. Chan, and S. Askary, A Reinforcement Learning Approach for Intelligent Traffic Signal Control at Urban Intersections; In Proceedings of the IEEE Intelligent Transportation Systems Conference (ITSC), 2019.
- M. Guo, P. Wang, C.-Y. Chan, and S. Askary, A Reinforcement Learning Approach for Intelligent Traffic Signal Control at Urban Intersections; In Proceedings of the IEEE Intelligent Transportation Systems Conference (ITSC), 2019.
- S.M. Rayhanul Swapno et al., Traffic Light Control Using Reinforcement Learning; In Proceedings of the 2024 International Conference on Integrated Circuits and Communication Systems (ICICACS), pp. 1–7, IEEE, 2024.
- T.A. Haddad, D. Hedjazi, and S. Aouag, A Deep Reinforcement Learning-Based Cooperative Approach for Multi-Intersection Traffic Signal Control; Engineering Applications of Artificial Intelligence, Vol. 114, Article 105019, 2022.






| Method | Final Reward | Final Delay (s/veh) | Throughput (veh/ep) | Avg. Queue | Conv. Episode |
|---|---|---|---|---|---|
| Q-Learning | 2764 | ||||
| SARSA | 2718 | ||||
| Monte Carlo | 2875 | ||||
| Hybrid (ours) | 2938 |
| Algorithm | Avg. Reward | Std. Dev. Reward | Avg. Waiting Time (s) | Min. Waiting Time (s) | Max. Waiting Time (s) |
|---|---|---|---|---|---|
| Q-Learning | 245 | 407 | |||
| SARSA | 233 | 436 | |||
| Monte Carlo | 244 | 412 | |||
| Hybrid | 239 | 382 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).