Submitted:
14 July 2026
Posted:
16 July 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- CSAC-based Network Architecture for Heterogeneous Queuing: We model the high-priority queue’s delay violation as an independent episodic constraint, explicitly decoupled from the throughput objective, such that the binary “violation-or-not” semantics are embedded in the constraint term rather than in hand-tuned reward weights. Building upon this formulation, we develop a CSAC-based solver that integrates this constrained objective into the actor-critic learning process, ensuring stable and reliable policy optimization under stringent queuing delay constraints.
- Continuous-to-Discrete Action Mapping under Per-Time-Slot Link Capacity Constraints: We design a two-stage mapping that transforms the continuous policy logits into integer transmission quotas that are non-negative, do not exceed each queue’s backlog, and sum to the effective capacity available to backlogged queues, so they can be executed directly without post-processing.
- Comprehensive Experimental Validation under Heterogeneous Traffic: We systematically compare the proposed approach against an unconstrained soft actor-critic (SAC) baseline and simplified heuristic baselines under heterogeneous traffic conditions. The experimental scenarios include one queue following a Markov-modulated Poisson process (MMPP) with the highest priority and a stringent queuing delay requirement, as well as two queues following independent Poisson arrivals with lower priorities and best-effort queuing delay requirements. Results verify its advantages in throughput, queuing delay, and overflow packet count behavior.
2. Related Works
2.1. Non-Learning Methods
2.2. Learning-Based Methods
3. System Model
- 1.
- The agent observes system state , selects a continuous action based on and past actions, and maps to integer transmission quotas for queue i, , .
- 2.
- The HoL packets in queue i are transmitted in first-in, first-out (FIFO) order.
- 3.
- new packets arrive at queue i, and the length of queue i evolves aswhere the min operator accounts for the dropping of overflow packets.
- 4.
- The reward is obtained, along with the queuing delay violation cost of the highest-priority queue 0.
- 5.
- The system state is updated for time slot , denoted .
4. The Proposed Constrained Soft Actor-Critic Approach
4.1. Basic Architecture
4.1.1. Experience Collection
4.1.2. Parameter Update
4.2. State Space
4.2.1. Queue Occupancy History
4.2.2. Past Violation Feedback
4.2.3. Link Capacity History
4.3. Action Space
- 1.
- Compute the ideal allocation for .
- 2.
- Floor to for .
- 3.
- The residual capacity is distributed to queues with the largest remainders , each receiving one extra unit. For queues with the same remainder, allocation is performed in order of priority from highest to lowest.
4.4. Reward and Queuing Delay Violation Cost
4.4.1. Throughput-Oriented Reward
4.4.2. Queuing Delay Violation Cost for Highest-Priority Queue 0
4.5. Parameter Update
4.6. Approach Execution
| Algorithm 1 CSAC-HQS |
|
5. Evaluation
5.1. Queuing Delay Violation Rate for Queue 0
5.2. Average Queuing Delay per Best-Effort Queue
5.3. Throughput and Overflow Packet Count per Queue
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| CoDel | Controlled Delay |
| CRL | Constrained Reinforcement Learning |
| CSAC | Constrained Soft Actor-Critic |
| CSAC-HQS | Constrained Soft Actor-Critic for Heterogeneous Queuing Systems |
| CVaR | Conditional Value at Risk |
| DRL | Deep Reinforcement Learning |
| DRR | Deficit Round Robin |
| EDF | Earliest-Deadline-First |
| FIFO | First-In, First-Out |
| HoL | Head-of-Line |
| IoT | Internet of Things |
| MDP | Markov Decision Process |
| MMPP | Markov-Modulated Poisson Process |
| QoS | Quality of Service |
| QADRA | QoS-Aware Deep Reinforcement learning Agent |
| RSD4 | Recurrent Softmax Delayed Deep Double Deterministic policy gradient |
| SAC | Soft Actor-Critic |
| SAC-FSO | Soft Actor-Critic-based Flow Scheduling Optimization |
| TSN | Time-Sensitive Networking |
| WCSAC | Worst-Case Soft Actor-Critic |
| WFQ | Weighted Fair Queuing |
References
- Foukas, X.; Patounas, G.; Elmokashfi, A.; Marina, M.K. Network slicing in 5G: Survey and challenges. IEEE Commun. Mag. 2017, 55, 94–100. [Google Scholar] [CrossRef]
- Lu, Y.; Yang, L.; Yang, S.X.; Hua, Q.; Sangaiah, A.K.; Guo, T.; Yu, K. An intelligent deterministic scheduling method for ultra-low latency communication in edge enabled industrial internet of things. IEEE Trans. Ind. Inform. 2023, 19, 1756–1767. [Google Scholar] [CrossRef]
- Mohammadpour, E.; Stai, E.; Le Boudec, J.Y. Improved network-calculus nodal delay-bounds in time-sensitive networks. IEEE/ACM Trans. Netw. 2023, 31, 2902–2917. [Google Scholar] [CrossRef]
- Giambene, G. Queuing Theory and Telecommunications: Networks and Applications, 3rd ed.; Springer: Cham, Switzerland, 2021. [Google Scholar] [CrossRef]
- Zhang, X.; Lu, Y. Asynchronous channel-aware and queue-aware deficit round robin scheduling strategy for different TSN traffics in the TSN-5G network. Electron. Lett. 2025, 61, e70418. [Google Scholar] [CrossRef]
- Mou, S.; Maguluri, S.T. Heavy-traffic queue length behavior in a switch under Markovian arrivals. Adv. Appl. Probab. 2024, 56, 1106–1152. [Google Scholar] [CrossRef]
- Li, Z.; Gurushankar, K.; Harchol-Balter, M.; Scheller-Wolf, A. Improving upon the generalized cμ rule: A Whittle approach. ACM SIGMETRICS Perform. Eval. Rev. 2025, 53, 122–124. [Google Scholar] [CrossRef]
- Deng, L.; Zeng, G.; Kurachi, R.; Takada, H.; Xiao, X.; Li, R.; Xie, G. Enhanced real-time scheduling of AVB flows in time-sensitive networking. ACM Trans. Des. Autom. Electron. Syst. 2024, 29, 33:1–33:26. [Google Scholar] [CrossRef]
- Chen, W.; Tian, Y.; Yu, X.; Zheng, B.; Zhang, X. Enhancing fairness for approximate weighted fair queueing with a single queue. IEEE/ACM Trans. Netw. 2024, 32, 3901–3915. [Google Scholar] [CrossRef]
- Bouillard, A.; Boyer, M.; Le Corronc, E. Deterministic Network Calculus: From Theory to Practical Implementation; Wiley-ISTE: Hoboken, NJ, USA, 2018. [Google Scholar] [CrossRef]
- Tabatabaee, S.M.; Le Boudec, J.Y. Deficit round-robin: A second network calculus analysis. In Proceedings of the 2021 IEEE 27th Real-Time and Embedded Technology and Applications Symposium (RTAS); IEEE, 2021; pp. 171–183. [Google Scholar] [CrossRef]
- Hu, W.; Sun, L.; Wang, J.; Chen, W.; Li, W. Upper-bound latency analysis for TSN with improved network calculus-based model. Electron. Lett. 2025, 61, e70302. [Google Scholar] [CrossRef]
- Yang, Q.; Jiang, X.; Liu, R.; Li, T.; Yang, H.; Quan, W.; Sun, Z. A performance-balanced scheduling algorithm for diverse real-world TSN scenarios. IEEE Trans. Parallel Distrib. Syst. 2025, 36, 2469–2481. [Google Scholar] [CrossRef]
- Xu, Z.; Tang, J.; Meng, J.; Zhang, W.; Wang, Y.; Liu, C.H.; Yang, D. Experience-driven networking: A deep reinforcement learning based approach. In Proceedings of the IEEE INFOCOM 2018 - IEEE Conference on Computer Communications. IEEE, 2018; pp. 1871–1879. [Google Scholar] [CrossRef]
- Mao, H.; Alizadeh, M.; Menache, I.; Kandula, S. Resource management with deep reinforcement learning. In Proceedings of the Proceedings of the 15th ACM Workshop on Hot Topics in Networks (HotNets), 2016; ACM; pp. 50–56. [Google Scholar] [CrossRef]
- Stigenberg, J.; Saxena, V.; Tayamon, S.; Ghadimi, E. QoS-aware scheduling in new radio using deep reinforcement learning. In Proceedings of the 2021 IEEE 32nd Annual International Symposium on Personal, Indoor and Mobile Radio Communications (PIMRC); IEEE, 2021; pp. 991–997. [Google Scholar] [CrossRef]
- Forero, P.A.; Zhang, P.; Radosevic, D. Active queue-management policies for undersea networking via deep reinforcement learning. In Proceedings of the OCEANS 2021: San Diego–Porto; IEEE, 2021; pp. 1–8. [Google Scholar] [CrossRef]
- Wang, X.; Zhang, J.; Lu, X.; Li, F.; Chen, C.; Guan, X. Towards wireless time-sensitive networking: Multi-link deterministic scheduling via deep reinforcement learning. Comput. Netw. 2025, 261, 111119. [Google Scholar] [CrossRef]
- Mawlood, M.A.; Mahmood, D.A. Optimizing weighted fair queuing with deep reinforcement learning for dynamic bandwidth allocation. Telecom 2025, 6, 46. [Google Scholar] [CrossRef]
- Yang, Q.; Simão, T.D.; Tindemans, S.H.; Spaan, M.T.J. WCSAC: Worst-case soft actor-critic for safety-constrained reinforcement learning. Proc. Proc. AAAI Conf. Artif. Intell. 2021, Vol. 35, 10639–10646. [Google Scholar] [CrossRef]
- Zhang, Q.; Leng, S.; Ma, X.; Liu, Q.; Wang, X.; Liang, B.; Liu, Y.; Yang, J. CVaR-constrained policy optimization for safe reinforcement learning. IEEE Trans. Neural Netw. Learn. Syst. 2025, 36, 830–841. [Google Scholar] [CrossRef] [PubMed]
- Liu, Y.; Ding, J.; Zhang, Z.L.; Liu, X. CLARA: A constrained reinforcement learning based resource allocation framework for network slicing. In Proceedings of the 2021 IEEE International Conference on Big Data (Big Data). IEEE, 2021; pp. 1427–1437. [Google Scholar] [CrossRef]
- Hu, P.; Chen, Y.; Pan, L.; Fang, Z.; Xiao, F.; Huang, L. Multi-user delay-constrained scheduling with deep recurrent reinforcement learning. IEEE/ACM Trans. Netw. 2024, 32, 2344–2359. [Google Scholar] [CrossRef]
- Alvi, N.M.; Alvi, W.M.; Zhou, X.; Li, J.; Wei, Y. Constrained soft actor-critic for joint computation offloading and resource allocation in UAV-assisted edge computing. Sensors 2026, 26, 1149. [Google Scholar] [CrossRef] [PubMed]
- Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. Proc. Proc. Int. Conf. Mach. Learn. (ICML) 2018, Vol. 80, 1861–1870. [Google Scholar] [CrossRef]
- Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. PyTorch: An imperative style, high-performance deep learning library. Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) 2019, Vol. 32, 8024–8035. [Google Scholar]





| Category | Parameter | Value |
|---|---|---|
| Queuing System | Number of queues (D) | 3 |
| Buffer capacity (B) | 1000 packets | |
| Time slot duration () | 5 ms | |
| Episode length (H) | 200 time slots | |
| Link capacity lower bound () | 338 packets | |
| Link capacity upper bound () | 438 packets | |
| Maximum allowable queuing delays for queue 0, 1, 2 () | 10, 20, 30 ms | |
| MMPP transition probabilities of queue 0 (, ) 1 | , | |
| MMPP normal and burst arrival rates | 5, 20 packets/ms | |
| Arrival rates of queue 1, 2 | 30, 40 packets/ms | |
| CSAC-HQS | Optimizer | Adam |
| Optimizer learning rate | ||
| Discount factor () | 0.99 | |
| Replay buffer size | ||
| Mini-batch size (N) | 256 | |
| Soft-update coefficient () | 0.005 | |
| Entropy temperature () | auto-tuned, initialized at 0.2 | |
| Total training steps | ||
| Actor hidden layers | [64, 64, 64, 64] | |
| Critic hidden layers | [128, 128, 128, 128] | |
| Sharpening factor () | 1, 5 | |
| History window of the link capacity () | 10 time slots | |
| Warm-up steps () | 1000 time slots | |
| Update interval (G) | 10 time slots | |
| Gradient steps per update | 1 | |
| Per-episode queuing delay violation budget () | 2 | |
| Multiplier learning rate () | ||
| Entropy temperature learning rate () | ||
| Maximum multiplier () | 5 | |
| Reward-scaling coefficient () | 0.01 | |
| Priority weights () | ||
| Throughput utility coefficient () | 1.0 | |
| Queuing delay violation penalty coefficient () | 1.0 | |
| Drop penalty coefficient () | 0.5 | |
| Queuing delay violation penalty clip multiplier () | 10 |
| SAC | CSAC-HQS | |
|---|---|---|
| Sharpening factor | 1 | 1, 5 |
| Masked-softmax | yes | yes |
| Queuing delay violation penalty indicator of queue 0 () | 1 | 0 |
| Queuing delay violation penalty indicator of queue 1 () | 1 | 1 |
| Queuing delay violation penalty indicator of queue 2 () | 1 | 1 |
| Constraint critic | — | yes |
| Lagrangian update on | — | yes |
| Warm-up steps | 1000 | 1000 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).