Submitted:
19 May 2026
Posted:
20 May 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- We consider a CPN task scheduling system that integrates computing and networking resources. Based on this system, we describe the specific processes of task communication and computation, and formulate an optimization problem with the objectives of maximizing the task success rate and minimizing the average processing delay.
- To solve the joint optimization problem in the CPN, we propose the TS-DQNF method. Specifically, the TS-DQNF first uses the DQN algorithm to determine the target computation node for each task. Then, it introduces a congestion-aware mechanism to improve the Floyd algorithm and compute the shortest routing path for each task. Finally, it gradually finds an appropriate task scheduling scheme through alternating iterative optimization.
- We verify the effectiveness of the proposed TS-DQNF method through simulation experiments. The results show that the TS-DQNF achieves a higher task success rate and a lower average processing delay under different network scenarios and different scales of UDs, while also exhibiting good convergence performance.
2. Related Work
2.1. Computing Power Network
2.2. Task Scheduling
3. System Model and Problem Formulation
3.1. CPN Model
3.2. Communication Model
3.2.1. From UDs to CPN Routers
3.2.2. From CPN routers to CPN nodes
3.3. Computing Model
3.4. Problem formulation
3.4.1. Success rate
3.4.2. Average Processing Delay
3.4.3. Optimization Problem
4. Proposed Method
4.1. DRL-Based Markov Decision Process for Task Scheduling
-
State space: To ensure that the CPN controller has comprehensive awareness of the system, it is necessary to fully capture the state information of the system environment, enabling the controller to make informed decisions. Therefore, the system state space in time slot t can be defined as:where indicates the state information of tasks, indicates the state information of CPN routers,indicates the state information of CPN nodes.
- Action space: In the CPN system environment, an action refers to the scheduling decision formulated by the CPN controller, which determines where a task will finally be scheduled for execution. Therefore, the action space for task scheduling is defined as:where indicates the action space of task , including , and indicates a task scheduling action executed by the system in time slot t.
- State-transition function: The state-transition function models the transition from the current state to the next state after actions are executed at time step t. We define the state transition function as , which returns the next system state after executing action .
- Reward function: In DRL, the reward is the feedback from the environment after the agent executes an action, which is used to evaluate the impact of its scheduling decision on the system. As described above, this paper aims to improve the task success rate while minimizing the average processing delay of successfully executed tasks. Therefore, we define the immediate reward of the system environment in time slot t as:
4.2. Decision Model Training of Task Scheduling Based on DQN
| Algorithm 1 Training Process of Task Scheduling Decision Model Based on DQN |
|
4.3. Dynamic Congestion-Aware Shortest Routing Path Planning
| Algorithm 2 Dynamic Congestion-Aware Shortest Path Planning via Floyd |
|
5. Performance Evaluation
- RQ 1: How is the convergence performance of the TS-DQNF? (Section 5.2)
- RQ 2: How is the performance of the TS-DQNF and other methods in different network scenarios? (Section 5.3)
- RQ 3: How is the effect of different UD numbers on the performance of the TS-DQNF and other methods? (Section 5.4)
5.1. Simulation Settings
- PSO[27]: This method is a commonly used heuristic method for solving task scheduling problems. It is adopting the particle swarm optimization algorithm to determine the target CPN node for each task and, at the same time, compute the routing path for task scheduling.
- KDRL[25]: The core of this method is a joint optimization algorithm for routing and scheduling. It is first adopting the K-shortest path (KSP) algorithm to determine a set of candidate routing paths. Then, a DRL agent is used to jointly select the target node and routing path of each task from the generated candidate set.
- RS: This method is adopting a random scheduling strategy, namely randomly selecting a CPN node for each task and randomly generating a routing path.
- GS: This method is adopting a greedy scheduling strategy. Specifically, it selects the CPN node with the lowest processing delay for each task each time, and selects the link with the lowest transmission delay each time as part of the path to generate the final routing path.
5.2. RQ 1: Convergence Performance of the TS-DQNF
5.3. RQ2: Performance Analysis of Each Method in Different Network Scenarios
5.4. RQ3: Performance Analysis of Different UD Numbers on Each Method
6. Conclusion
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Wang, L.; Xu, W.; Liu, Y.; Wang, M. Artificial intelligence for virtual reality: A review. Sci. China Inf. Sci. 2025, 69, 111101. [Google Scholar] [CrossRef]
- Badue, C.; Guidolini, R.; Carneiro, R.V.; Azevedo, P.; Cardoso, V.B.; Forechi, A.; Jesus, L.; Berriel, R.; Paixão, T.M.; Mutz, F.; De Paula Veronese, L.; Oliveira-Santos, T.; De Souza, A.F. Self-driving cars: A survey. Expert Syst. With Appl. 2021, 165, 113816. [Google Scholar] [CrossRef]
- Liu, F.; Zhao, Q.; Liu, X.; Zeng, D. Joint face alignment and 3D face reconstruction with application to face recognition. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 42, 664–678. [Google Scholar] [CrossRef] [PubMed]
- Ye, J. Artificial intelligence-generated content (AIGC) in biomedical research, healthcare delivery, and clinical practices: Technologies, applications, and regulatory considerations. Artif. Intell. Rev. 2026, 59, 86. [Google Scholar] [CrossRef]
- Wang, W.-X.; Wu, J. Survey on hypergraph algorithms: Foundations, advances, and applications in cloud computing. J. Comput. Sci. Technol. 2026. [Google Scholar] [CrossRef]
- Tesfay, C.H.; Xiang, Z.; Yang, L.; Mahmood, J.; Ren, M.; Zhang, S.; Das, A.K.; Chaudhry, S.A. Task offloading and optimization methods in UAV-enabled mobile edge computing: A comprehensive survey. Comput. Commun. 2026, 254, 108537. [Google Scholar] [CrossRef]
- Xiao, H.; Xu, C.; Ma, Y.; Yang, S.; Zhong, L.; Muntean, G.-M. Edge intelligence: A computational task offloading scheme for dependent IoT application. IEEE Trans. Wirel. Commun. 2022, 21, 7222–7237. [Google Scholar] [CrossRef]
- Tang, X.; Cao, C.; Wang, Y.; Zhang, S.; Liu, Y.; Li, M.; He, T. Computing power network: The architecture of convergence of computing and networking towards 6G requirement. China Commun. 2021, 18, 175–185. [Google Scholar] [CrossRef]
- Sun, Y.; Lei, B.; Liu, J.; Huang, H.; Zhang, X.; Peng, J.; Wang, W. Computing power network: A survey. China Commun. 2024, 21, 109–145. [Google Scholar] [CrossRef]
- Sun, W.; Li, Z.; Wang, Q.; Zhang, Y. FedTAR: Task and resource-aware federated learning for wireless computing power networks. IEEE Internet Things J. 2023, 10, 4257–4270. [Google Scholar] [CrossRef]
- Emara, F.A.; Gad-Elrab, A.A.A.; Sobhi, A.; Alsharkawy, A.S.; Embabi, M.E.; El-Baky, M.A.A. Multi-objective task scheduling algorithm for load balancing in cloud computing based on improved Harris hawks optimization. J. Supercomput. 2025, 81, 790. [Google Scholar] [CrossRef]
- Talaat, F.M.; Hamza, A.A. CloudSched-GA: An adaptive genetic optimizer for efficient and balanced task scheduling in cloud ecosystems. Neural Comput. Appl. 2025, 37, 28269–28293. [Google Scholar] [CrossRef]
- Sahraei, S.H.; Kashani, M.M.R.; Rezazadeh, J.; Farahbakhsh, R. Efficient job scheduling in cloud computing based on genetic algorithm. Int. J. Commun. Netw. Distrib. Syst. 2019, 22, 447–467. [Google Scholar] [CrossRef]
- Zhang, W.; Ou, H. Reinforcement learning based multi objective task scheduling for energy efficient and cost effective cloud edge computing. Sci. Rep. 2025, 15, 41716. [Google Scholar] [CrossRef]
- Mangalampalli, S.; Karri, G.R.; Ratnamani, M.V.; Mohanty, S.N.; Jabr, B.A.; Ali, Y.A.; Ali, S.; Abdullaeva, B.S. Efficient deep reinforcement learning based task scheduler in multi cloud environment. Sci. Rep. 2024, 14, 21850. [Google Scholar] [CrossRef] [PubMed]
- Qi, Q.; Zhang, L.; Wang, J.; Sun, H.; Zhuang, Z.; Liao, J.; Yu, F.R. Scalable parallel task scheduling for autonomous driving using multi-task deep reinforcement learning. IEEE Trans. Veh. Technol. 2020, 69, 13861–13874. [Google Scholar] [CrossRef]
- ITU-T. Comput. Power Netw.-Framew. Archit. Y 2021.
- Xia, L.; Guo, D.; Wang, Y.; Sun, D.; Zhen, W.; Jing, C. Optimal load scheduling based on mobile edge computing technology in 5G dense networking. In Proceedings of the 2022 3rd Asia Conference on Computers and Communications (ACCC), 2022; pp. 137–142. [Google Scholar]
- Tang, Q.; Xie, R.; Feng, L.; Yu, F.R.; Chen, T.; Zhang, R.; Huang, T. SIaTS: A service intent-aware task scheduling framework for computing power networks. IEEE Netw. 2024, 38, 233–240. [Google Scholar] [CrossRef]
- Liu, J.; Sun, Y.; Su, J.; Li, Z.; Zhang, X.; Lei, B.; Wang, W. Computing power network: A testbed and applications with edge intelligence. In Proceedings of the IEEE INFOCOM 2022–IEEE Conference on Computer Communications Workshops (INFOCOM WKSHPS), 2022; pp. 1–2. [Google Scholar]
- Li, J.; Zeng, Q.; Lu, L.; Zhang, E.; Xu, A.; Li, W. Task offloading and computing resource allocation of the joint UAV with computing power network. J. Supercomput. 2025, 81, 884. [Google Scholar] [CrossRef]
- Zhao, W.; Huang, X.; Li, D.; Zhang, P.; Xie, K. Intelligent task scheduling towards distributed computing power network. In Advanced Intelligent Computing Technology and Applications; Huang, D.-S., Zhang, C., Zhang, Q., Pan, Y., Eds.; Springer Nature Singapore: Singapore, 2025; pp. 495–506. [Google Scholar]
- Liu, H.; Zhang, S.; Li, L.; Sun, T.; Xue, W.; Yao, X.; Xu, Y. Computing power network dynamic resource scheduling integrating time series mixing dynamic state estimation and hierarchical reinforcement learning. Sci. Rep. 2026, 16, 2905. [Google Scholar] [CrossRef]
- Yin, M.; Gao, X.; Wu, S.; Wang, H.; Zhou, T.; Lin, W. Key technologies for terminal computing power network. In Proceedings of the 15th International Conference on Computer Engineering and Networks; Cheng, Z., Liu, X., Su, J., Eds.; Springer Nature Singapore: Singapore, 2026; pp. 627–636. [Google Scholar]
- Feng, L.; Xie, R.; Tang, Q.; Huang, T.; Xiong, Z.; Chen, T.; Zhang, R.; Tan, S.; Fang, Z. CaRCS: Joint optimization of computing-aware routing and collaborative scheduling in computing power networks. IEEE Netw. 2025, 39, 270–278. [Google Scholar] [CrossRef]
- Zhang, Y.; Zhang, H.; Song, C. A hybrid algorithm for multi-objective task scheduling in heterogeneous cloud computing. J. Supercomput. 2025, 81, 1143. [Google Scholar] [CrossRef]
- Malik, M.; Nandan, D.; Prabha, C.; Uddin, M.; Acharya, B.; Hu, Y.-C. A bio-inspired metaheuristic approach for cloud task scheduling using lateral hyena based particle swarm optimization. Multimed. Tools Appl. 2025, 84, 20023–20046. [Google Scholar] [CrossRef]
- Yu, X.; Mi, J.; Tang, L.; Long, L.; Qin, X. Dynamic multi objective task scheduling in cloud computing using reinforcement learning for energy and cost optimization. Sci. Rep. 2025, 15, 45387. [Google Scholar] [CrossRef]
- Tian, B.; Xiao, G.; Shen, Y. A deep reinforcement learning approach for dynamic task scheduling of flight tests. J. Supercomput. 2024, 80, 18761–18796. [Google Scholar] [CrossRef]
- Xiao, H.; Xu, C.; Ma, Y.; Yang, S.; Zhong, L.; Muntean, G.-M. Edge intelligence: A computational task offloading scheme for dependent IoT application. IEEE Trans. Wirel. Commun. 2022, 21, 7222–7237. [Google Scholar] [CrossRef]
- Sun, Y.; Xu, J.; Cui, S. User association and resource allocation for MEC-enabled IoT networks. IEEE Trans. Wirel. Commun. 2022, 21, 8051–8062. [Google Scholar] [CrossRef]
- Mas, L.; Vilaplana, J.; Mateo, J.; Solsona, F. A queuing theory model for fog computing. J. Supercomput. 2022, 78, 11138–11155. [Google Scholar] [CrossRef]
- Cao, X.; Tang, G.; Guo, D.; Li, Y.; Zhang, W. Edge federation: Towards an integrated service provisioning model. IEEE/ACM Trans. Netw. 2020, 28, 1116–1129. [Google Scholar] [CrossRef]
- Sun, Z.; Mo, Y.; Yu, C. Graph-reinforcement-learning-based task offloading for multiaccess edge computing. IEEE Internet Things J. 2023, 10, 3138–3150. [Google Scholar] [CrossRef]
- Liu, G.; Deng, W.; Xie, X.; Huang, L.; Tang, H. Human-level control through directly trained deep spiking Q-networks. IEEE Trans. Cybern. 2023, 53, 7187–7198. [Google Scholar] [CrossRef]
- Floyd, R.W. Algorithm 97, shortest path algorithms. Commun. ACM 1962, 5, 345. [Google Scholar] [CrossRef]
- Orlowski, S.; Pióro, M.; Tomaszewski, A.; Wessäly, R. SNDlib 1.0–survivable network design library. Networks 2010, 55, 276–286. [Google Scholar] [CrossRef]
- Nguyen, D.C.; Ding, M.; Pathirana, P.N.; Seneviratne, A.; Li, J.; Poor, H.V. Cooperative task offloading and block mining in blockchain-based edge computing with multi-agent deep reinforcement learning. IEEE Trans. Mob. Comput. 2023, 22, 2021–2037. [Google Scholar] [CrossRef]
- Laili, Y.; Wang, X.; Zhang, L.; Ren, L. DSAC-configured differential evolution for cloud–edge–device collaborative task scheduling. IEEE Trans. Ind. Inform. 2024, 20, 1753–1763. [Google Scholar] [CrossRef]








| Symbol | Connotation |
|---|---|
| Set of UDs | |
| Set of CPN routers | |
| Set of CPN nodes | |
| Task generated by UD n in time slot t | |
| Data volume of task | |
| Computing volume of task | |
| Tolerable delay of task | |
| The i-th CPN router | |
| Forwarding capability of CPN router | |
| Maximum load capacity of CPN router | |
| Real-time load of CPN router | |
| Congestion rate of CPN router | |
| The i-th CPN node | |
| Computing capacity of CPN node | |
| Real-time load of CPN node | |
| Load rate of CPN node | |
| Scheduling scheme of task | |
| W | Network link transmission rate matrix |
| Task success rate | |
| Task average processing delay | |
| & | Weighting coefficients for and |
| Parameter | Value |
|---|---|
| Number of UDs | [10,30] |
| Data volume of task | [1.0,8.0] MB |
| Computing volume of task | [1.0,5.0] GCPU |
| Tolerable delay of task | 2 s |
| Forwarding capability of CPN router | [75,85] MB/s |
| Network link transmission rate | [400,500] MB/s |
| Computing capacity of CPN node | [10,15] GCPU/s |
| Weighting coefficients | 0.5,0.5 |
| Discount factor | 0.95 |
| Exploration rate | 0.1 |
| Learning rate | 0.0001 |
| Replay batch size | 64 |
| Training episodes | 1000 |
| Scenarios | Name | No. of CPN Routers | No. of CPN Nodes | No. of CPN UDs |
|---|---|---|---|---|
| 1 | France | 21 | 4 | 20 |
| 2 | India35 | 30 | 5 | 20 |
| 3 | Cost266 | 30 | 7 | 20 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).