Submitted:
03 July 2023
Posted:
04 July 2023
You are already at the latest version
Abstract
Keywords:
1. Introduction
- We developed a global dictionary during training to discover the TSPWP’s best strategy for solving different realizations. The dictionary comprises letters representing the available hotspots, tokens representing local paths, and words depicting the complete trajectories and order of hotspots. By studying the dictionary, we can comprehend the decision-maker’s grammar (i.e., the TSPWP strategy) and how it uses the available letters to form tokens and words.
- We have designed a novel hierarchical representation structuring the acquired knowledge (the global dictionary) to accurately depict the properties of the TSPWP graphs at various levels of abstraction and time scales.
- We tested the proposed method on different scenarios with varying hotspots. Our method outperformed traditional Q-learning by providing fast, stable, and reliable solutions with good generalization ability.
2. Literature Review

3. System Model and Problem Formulation
4. Proposed Goal-Directed Trajectory Design Method
4.1. TSP with profits instances
4.2. World Model
4.2.1. Dictionary Learning
4.2.2. The proposed graphical representation

4.3. Active Inference
4.3.1. Action selection
4.3.2. Prediction and Perception
4.3.3. Abnormality measures and action update
5. Numerical Results and Discussion
5.1. Comparison with modified Q-learning
6. Conclusions and Future Directions
Author Contributions
Funding
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| UAV | Unmanned aerial vehicle |
| LoS | Line of sight |
| NOMA | Non-orthogonal multiple access |
| GPS | Global positioning system |
| IoT | Internet of things |
| AI | Artificial intelligence |
| ML | Machine learning |
| RL | Reinforcement learning |
| TSPWP | Travel salesman problem with profits |
| GDBN | Generalized dynamic Bayesian network |
| C-MGDBN | Coupled multi-scale generalized dynamic Bayesian network |
| DP | Dynamic programming |
| WSN | Wireless sensor node |
| MILP | Mixed integer linear programming |
| TSP | Travel salesman problem |
| GA | Genetic algorithm |
| PSO | Particle swarm optimization |
| ACO | Ant colony optimization |
| QoE | Quality of experience |
| QL | Q-learning |
| DQL | Deep Q-learning |
| FBS | Flying base station |
| GU | Ground users |
| RB | Resource block |
| OFDMA | Orthogonal frequency division multiple access |
| NLoS | Non line of sight |
| AWGN | Additive White Gaussian Noise |
| C-GDBN | Coupled Generalized dynamic Bayesian network |
| M-GDBN | Multi-scale generalized dynamic Bayesian network |
| GNG | Growing neural gas |
| POMDP | Partially observable Markov decision process |
| KF | Kalman filter |
| PF | Particle filter |
References
- Li, B.; Fei, Z.; Zhang, Y. UAV Communications for 5G and Beyond: Recent Advances and Future Trends. IEEE Internet of Things Journal 2019, 6, 2241–2263. [Google Scholar] [CrossRef]
- Krayani, A.; Baydoun, M.; Marcenaro, L.; Gao, Y.; Regazzoni, C.S. Smart Jammer Detection for Self-Aware Cognitive UAV Radios. In Proceedings of the 2020 IEEE 31st Annual International Symposium on Personal, Indoor and Mobile Radio Communications; 2020; pp. 1–7. [Google Scholar] [CrossRef]
- Zhou, Y.; Yeoh, P.L.; Chen, H.; Li, Y.; Schober, R.; Zhuo, L.; Vucetic, B. Improving Physical Layer Security via a UAV Friendly Jammer for Unknown Eavesdropper Location. IEEE Transactions on Vehicular Technology 2018, 67, 11280–11284. [Google Scholar] [CrossRef]
- Khawaja, W.; Ozdemir, O.; Guvenc, I. UAV Air-to-Ground Channel Characterization for mmWave Systems. In Proceedings of the 2017 IEEE 86th Vehicular Technology Conference (VTC-Fall); 2017; pp. 1–5. [Google Scholar] [CrossRef]
- Cheng, F.; Zhang, S.; Li, Z.; Chen, Y.; Zhao, N.; Yu, F.R.; Leung, V.C.M. UAV Trajectory Optimization for Data Offloading at the Edge of Multiple Cells. IEEE Transactions on Vehicular Technology 2018, 67, 6732–6736. [Google Scholar] [CrossRef]
- Osseiran, A.; Boccardi, F.; Braun, V.; Kusume, K.; Marsch, P.; Maternia, M.; Queseth, O.; Schellmann, M.; Schotten, H.; Taoka, H.; et al. Scenarios for 5G mobile and wireless communications: The vision of the METIS project. IEEE Communications Magazine 2014, 52, 26–35. [Google Scholar] [CrossRef]
- Zeng, Y.; Zhang, R.; Lim, T.J. Wireless communications with unmanned aerial vehicles: Opportunities and challenges. IEEE Communications Magazine 2016, 54, 36–42. [Google Scholar] [CrossRef]
- Yang, D.; Wu, Q.; Zeng, Y.; Zhang, R. Energy Tradeoff in Ground-to-UAV Communication via Trajectory Design. IEEE Transactions on Vehicular Technology 2018, 67, 6721–6726. [Google Scholar] [CrossRef]
- Wang, Q.; Chen, Z.; Li, H.; Li, S. Joint Power and Trajectory Design for Physical-Layer Secrecy in the UAV-Aided Mobile Relaying System. IEEE Access 2018, 6, 62849–62855. [Google Scholar] [CrossRef]
- Yi, W.; Liu, Y.; Bodanese, E.; Nallanathan, A.; Karagiannidis, G.K. A Unified Spatial Framework for UAV-Aided MmWave Networks. IEEE Transactions on Communications 2019, 67, 8801–8817. [Google Scholar] [CrossRef]
- Kandeepan, S.; Gomez, K.; Reynaud, L.; Rasheed, T. Aerial-terrestrial communications: Terrestrial cooperation and energy-efficient transmissions to aerial base stations. IEEE Transactions on Aerospace and Electronic Systems 2014, 50, 2715–2735. [Google Scholar] [CrossRef]
- Zhang, S.; Zeng, Y.; Zhang, R. Cellular-Enabled UAV Communication: A Connectivity-Constrained Trajectory Optimization Perspective. IEEE Transactions on Communications 2019, 67, 2580–2604. [Google Scholar] [CrossRef]
- Yuan, X.; Yang, T.; Hu, Y.; Xu, J.; Schmeink, A. Trajectory Design for UAV-Enabled Multiuser Wireless Power Transfer With Nonlinear Energy Harvesting. IEEE Transactions on Wireless Communications 2021, 20, 1105–1121. [Google Scholar] [CrossRef]
- Li, L.; Li, W.; Wang, J.; Chen, X.; Peng, Q.; Huang, W. UAV Trajectory Optimization for Spectrum Cartography: A PPO Approach. IEEE Communications Letters 2023, 1. [Google Scholar] [CrossRef]
- Wang, J.; Wang, X.; Liu, X.; Cheng, C.T.; Xiao, F.; Liang, D. Trajectory Planning of UAV-enabled Data Uploading for Large-scale Dynamic Networks: A Trend Prediction Based Learning Approach. IEEE Transactions on Vehicular Technology 2023, 1–6. [Google Scholar] [CrossRef]
- Yin, D.; Yang, X.; Yu, H.; Chen, S.; Wang, C. An Air-to-Ground Relay Communication Planning Method for UAVs Swarm Applications. IEEE Transactions on Intelligent Vehicles 2023, 8, 2983–2997. [Google Scholar] [CrossRef]
- Chen, G.; Zhai, X.B.; Li, C. Joint Optimization of Trajectory and User Association via Reinforcement Learning for UAV-Aided Data Collection in Wireless Networks. IEEE Transactions on Wireless Communications 2023, 22, 3128–3143. [Google Scholar] [CrossRef]
- Zhang, Z.; Xu, C.; Li, Z.; Zhao, X.; Wu, R. Deep Reinforcement Learning for Aerial Data Collection in Hybrid-Powered NOMA-IoT Networks. IEEE Internet of Things Journal 2023, 10, 1761–1774. [Google Scholar] [CrossRef]
- Zhu, B.; Bedeer, E.; Nguyen, H.H.; Barton, R.; Gao, Z. UAV Trajectory Planning for AoI-Minimal Data Collection in UAV-Aided IoT Networks by Transformer. IEEE Transactions on Wireless Communications 2023, 22, 1343–1358. [Google Scholar] [CrossRef]
- Afifi, G.; Gadallah, Y. Cellular Network-Supported Machine Learning Techniques for Autonomous UAV Trajectory Planning. IEEE Access 2022, 10, 131996–132011. [Google Scholar] [CrossRef]
- Hu, S.; Yuan, X.; Ni, W.; Wang, X. Trajectory Planning of Cellular-Connected UAV for Communication-Assisted Radar Sensing. IEEE Transactions on Communications 2022, 70, 6385–6396. [Google Scholar] [CrossRef]
- Qin, Y.; Zhang, Z.; Li, X.; Huangfu, W.; Zhang, H. Deep Reinforcement Learning Based Resource Allocation and Trajectory Planning in Integrated Sensing and Communications UAV Network. IEEE Transactions on Wireless Communications 2023, 1. [Google Scholar] [CrossRef]
- Krayani, A.; Alam, A.S.; Marcenaro, L.; Nallanathan, A.; Regazzoni, C. An Emergent Self-Awareness Module for Physical Layer Security in Cognitive UAV Radios. IEEE Transactions on Cognitive Communications and Networking 2022, 8, 888–906. [Google Scholar] [CrossRef]
- Krayani, A.; William, N.J.; Alam, A.S.; Marcenaro, L.; Qin, Z.; Nallanathan, A.; Regazzoni, C. Generalized Filtering with Transport Planning for Joint Modulation Conversion and Classification in AI-enabled Radios. In Proceedings of the ICC 2022 - IEEE International Conference on Communications; 2022; pp. 3759–3765. [Google Scholar] [CrossRef]
- Li, X.; Wang, Q.; Liu, J.; Zhang, W. Trajectory Design and Generalization for UAV Enabled Networks:A Deep Reinforcement Learning Approach. In Proceedings of the 2020 IEEE Wireless Communications and Networking Conference (WCNC); 2020; pp. 1–6. [Google Scholar] [CrossRef]
- Ji, S.; Pan, S.; Cambria, E.; Marttinen, P.; Yu, P.S. A Survey on Knowledge Graphs: Representation, Acquisition, and Applications. IEEE Transactions on Neural Networks and Learning Systems 2022, 33, 494–514. [Google Scholar] [CrossRef] [PubMed]
- Griffiths, T.L.; Chater, N.; Kemp, C.; Perfors, A.; Tenenbaum, J.B. Probabilistic models of cognition: Exploring representations and inductive biases. Trends in Cognitive Sciences 2010, 14, 357–364. [Google Scholar] [CrossRef] [PubMed]
- Parr, T.; Friston, K.J. Uncertainty, epistemics and active inference. Journal of The Royal Society Interface 2017, 14, 20170376. [Google Scholar] [CrossRef] [PubMed]
- Friston, K.; FitzGerald, T.; Rigoli, F.; Schwartenbeck, P.; Pezzulo, G. Active Inference: A Process Theory. Neural Computation 2017, 29, 1–49. [Google Scholar] [CrossRef]
- Friston, K. Active inference and free energy. Behavioral and Brain Sciences 2013, 36, 212–213. [Google Scholar] [CrossRef]
- Parr, T.; Friston, K.; Pezzulo, G. Generative models for sequential dynamics in active inference. Cognitive Neurodynamics 2023, 1–14. [Google Scholar] [CrossRef]
- Friston, K.J.; Parr, T.; de Vries, B. The graphical brain: Belief propagation and active inference. Network Neuroscience 2017, 1, 381–414. [Google Scholar] [CrossRef]
- Feillet, D.; Dejax, P.; Gendreau, M. Traveling Salesman Problems with Profits. Transportation Science 2005, 39, 188–205. [Google Scholar] [CrossRef]
- Krayani, A.; Alam, A.S.; Calipari, M.; Marcenaro, L.; Nallanathan, A.; Regazzoni, C. Automatic Modulation Classification in Cognitive-IoT Radios using Generalized Dynamic Bayesian Networks. In Proceedings of the 2021 IEEE 7th World Forum on Internet of Things (WF-IoT); 2021; pp. 235–240. [Google Scholar] [CrossRef]
- Krayani, A.; Baydoun, M.; Marcenaro, L.; Alam, A.S.; Regazzoni, C. Self-Learning Bayesian Generative Models for Jammer Detection in Cognitive-UAV-Radios. In Proceedings of the GLOBECOM 2020 - 2020 IEEE Global Communications Conference; 2020; pp. 1–7. [Google Scholar] [CrossRef]
- Baydoun, M.; Campo, D.; Sanguineti, V.; Marcenaro, L.; Cavallaro, A.; Regazzoni, C. Learning Switching Models for Abnormality Detection for Autonomous Driving. In Proceedings of the 2018 21st International Conference on Information Fusion (FUSION); 2018; pp. 2606–2613. [Google Scholar] [CrossRef]
- Tran, D.H.; Vu, T.X.; Chatzinotas, S.; ShahbazPanahi, S.; Ottersten, B. Coarse Trajectory Design for Energy Minimization in UAV-Enabled. IEEE Transactions on Vehicular Technology 2020, 69, 9483–9496. [Google Scholar] [CrossRef]
- Zixuan, Z.; Qinhao, W.; Bo, Z.; Xiaodong, Y.; Yuhua, T. UAV flight strategy algorithm based on dynamic programming. Journal of Systems Engineering and Electronics 2018, 29, 1293–1299. [Google Scholar] [CrossRef]
- De Waen, J.; Dinh, H.T.; Cruz Torres, M.H.; Holvoet, T. Scalable multirotor UAV trajectory planning using mixed integer linear programming. In Proceedings of the 2017 European Conference on Mobile Robots (ECMR); 2017; pp. 1–6. [Google Scholar] [CrossRef]
- Dhulkefl, E.; Durdu, A.; Terzioğlu, H. Dijkstra algorithm using UAV path planning. Konya Mühendislik Bilimleri Dergisi 2020, 8, 92–105. [Google Scholar] [CrossRef]
- Karur, K.; Sharma, N.; Dharmatti, C.; Siegel, J.E. A survey of path planning algorithms for mobile robots. Vehicles 2021, 3, 448–468. [Google Scholar] [CrossRef]
- Ibrahim, N.S.A.; Saparudin, F.A. Review on path planning algorithm for unmanned aerial vehicles. Indonesian Journal of Electrical Engineering and Computer Science 2021, 24. [Google Scholar] [CrossRef]
- Xie, J.; Garcia Carrillo, L.R.; Jin, L. Path Planning for UAV to Cover Multiple Separated Convex Polygonal Regions. IEEE Access 2020, 8, 51770–51785. [Google Scholar] [CrossRef]
- Johnson, D.S.; McGeoch, L.A. The traveling salesman problem: A case study. In Local Search in Combinatorial Optimization; Aarts, E., Lenstra, J.K., Eds.; Princeton University Press: Princeton, 2003; pp. 215–310. [Google Scholar] [CrossRef]
- Chen, J.; Ye, F.; Li, Y. Travelling salesman problem for UAV path planning with two parallel optimization algorithms. In Proceedings of the 2017 Progress in Electromagnetics Research Symposium - Fall (PIERS - FALL); 2017; pp. 832–837. [Google Scholar] [CrossRef]
- Pehlivanoglu, Y.V.; Pehlivanoglu, P. An enhanced genetic algorithm for path planning of autonomous UAV in target coverage problems. Applied Soft Computing 2021, 112, 107796. [Google Scholar] [CrossRef]
- Phung, M.D.; Ha, Q.P. Safety-enhanced UAV path planning with spherical vector-based particle swarm optimization. Applied Soft Computing 2021, 107, 107376. [Google Scholar] [CrossRef]
- Yue, L.; Chen, H. Unmanned vehicle path planning using a novel ant colony algorithm. EURASIP Journal on Wireless Communications and Networking 2019, 2019, 1–9. [Google Scholar] [CrossRef]
- Bayerlein, H.; De Kerret, P.; Gesbert, D. Trajectory Optimization for Autonomous Flying Base Station via Reinforcement Learning. In Proceedings of the 2018 IEEE 19th International Workshop on Signal Processing Advances in Wireless Communications (SPAWC); 2018; pp. 1–5. [Google Scholar] [CrossRef]
- Colonnese, S.; Cuomo, F.; Pagliari, G.; Chiaraviglio, L. Q-SQUARE: A Q-learning approach to provide a QoE aware UAV flight path in cellular networks. Ad Hoc Networks 2019, 91, 101872. [Google Scholar] [CrossRef]
- Abeywickrama, H.V.; He, Y.; Dutkiewicz, E.; Jayawickrama, B.A.; Mueck, M. A Reinforcement Learning Approach for Fair User Coverage Using UAV Mounted Base Stations Under Energy Constraints. IEEE Open Journal of Vehicular Technology 2020, 1, 67–81. [Google Scholar] [CrossRef]
- Zhang, Q.; Saad, W.; Bennis, M.; Lu, X.; Debbah, M.; Zuo, W. Predictive Deployment of UAV Base Stations in Wireless Networks: Machine Learning Meets Contract Theory. IEEE Transactions on Wireless Communications 2021, 20, 637–652. [Google Scholar] [CrossRef]
- Liu, X.; Liu, Y.; Chen, Y.; Hanzo, L. Trajectory Design and Power Control for Multi-UAV Assisted Wireless Networks: A Machine Learning Approach. IEEE Transactions on Vehicular Technology 2019, 68, 7957–7969. [Google Scholar] [CrossRef]
- Hu, Y.; Chen, M.; Saad, W.; Poor, H.V.; Cui, S. Meta-Reinforcement Learning for Trajectory Design in Wireless UAV Networks. In Proceedings of the GLOBECOM 2020 - 2020 IEEE Global Communications Conference; 2020; pp. 1–6. [Google Scholar] [CrossRef]
- Yin, S.; Zhao, S.; Zhao, Y.; Yu, F.R. Intelligent Trajectory Design in UAV-Aided Communications With Reinforcement Learning. IEEE Transactions on Vehicular Technology 2019, 68, 8227–8231. [Google Scholar] [CrossRef]
- Mozaffari, M.; Saad, W.; Bennis, M.; Debbah, M. Wireless Communication Using Unmanned Aerial Vehicles (UAVs): Optimal Transport Theory for Hover Time Optimization. IEEE Transactions on Wireless Communications 2017, 16, 8052–8066. [Google Scholar] [CrossRef]
- Applegate, D.L.; Bixby, R.E.; Chvatál, V.; Cook, W.J. The Traveling Salesman Problem: A Computational Study; Princeton University Press, 2006. [Google Scholar]
- Krayani, A.; Alam, A.S.; Marcenaro, L.; Nallanathan, A.; Regazzoni, C. Automatic Jamming Signal Classification in Cognitive UAV Radios. IEEE Transactions on Vehicular Technology 2022, 71, 12972–12988. [Google Scholar] [CrossRef]
- Zeng, Y.; Zhang, R. Energy-Efficient UAV Communication With Trajectory Optimization. IEEE Transactions on Wireless Communications 2017, 16, 3747–3760. [Google Scholar] [CrossRef]
- Watkins, C.; Dayan, P. Technical Note: Q-Learning. Machine Learning 1992, 8, 279–292. [Google Scholar] [CrossRef]























| Symbol | Meaning |
|---|---|
| Ground Users (GUs) | |
| N | Number of hotspots |
| T | battery life time |
| UAV’s initial location | |
| UAV’s final location | |
| UAV’s trajectory | |
| Sequence of hotspots served by the UAV | |
| nth hotspot serverd by the UAV | |
| total number of hotspots served along the trajectory | |
| set of possible trajectories to follow by the UAV | |
| Probability to move toward hotspot after visiting at time | |
| Remaining time to go back to the original location after serving | |
| The set of available hotspot areas | |
| The set of GUs distributed across the total geographical area | |
| The set of GUs belonging to the nth hotspot | |
| The coordinate of GU belonging to the | |
| Center of nth hotspot | |
| Radius of the nth hotspot | |
| The set of the average data rate of all the available hotspots | |
| Data rate of the nth hotspot | |
| t | Time slot |
| u | UAV |
| Channel gain between GU () and UAV (u) | |
| Channel factor | |
| Carrier frequency | |
| c | Speed of light |
| Path loss exponent | |
| Probability of Line of Sight | |
| Probability of Non Line of Sight | |
| Additional attenuation for line of sight links | |
| Additional attenuation for non line of sight links | |
| Distance between GU and UAV u at time t | |
| Achievable data rate in hotspot n | |
| The bandwidth of the resource block (RB) allocated to user | |
| Transmit power of user | |
| Power spectral density of the additive white Gaussian noise | |
| Training set of realizations representing M examples | |
| Set of the sequences of hotspots selected by TSPWP to solve M examples | |
| Set of trajectory instances generated by TSPWP | |
| Set of clusters generated by GNG | |
| Generalized letter | |
| Adjacency matrix | |
| Global adjacency matrix | |
| Global transition matrix | |
| Degree matrix | |
| Tokens | |
| Tokens transition matrix | |
| Words on order | |
| Words on motion | |
| Coupling word |
| Parameter | Value | Parameter | Value |
|---|---|---|---|
| 1 W | 2 | ||
| 180 KHz | dBm | ||
| 3 | 23 | ||
| N | 80 | M | 1000 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).