Submitted:
20 July 2026
Posted:
21 July 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Related Work
3. Framework for Reviewing and Evaluating Existing Methods
3.1. Levels of Autonomy
3.2. Methodology for Literature Selection
3.3. Fundamentals of Ship Path Planning
4. Ship Path Planning with Emphasis on DRL-Based Decision-Making
4.1. Model-Based Methods as Planning Priors and Safety Constraints
4.2. Value-Based DRL for Discrete Path-Planning Decisions
4.3. Policy-Gradient and Actor–Critic DRL for Continuous Ship Control
4.4. Multi-Agent and Hybrid DRL for Complex Maritime Interactions
5. Discussion and Outlook
5.1. Cross-Cutting Design Lessons
5.2. Challenges for Real-World Deployment
5.3. Priority Directions for Future Research
6. Conclusion
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Campbell, S.; Naeem, W.; Irwin, G.W. A review on improving the autonomy of unmanned surface vehicles through intelligent collision avoidance manoeuvres. Annu. Rev. Control 2012, 36, 267–283. [Google Scholar] [CrossRef]
- Liu, Z.; Zhang, Y.; Yu, X.; Yuan, C. Unmanned surface vehicles: An overview of developments and challenges. Annu. Rev. Control 2016, 41, 71–93. [Google Scholar] [CrossRef]
- Ramos, M.A.; Utne, I.B.; Mosleh, A. Collision avoidance on maritime autonomous surface ships: Operators’ tasks and human failure events. Saf. Sci. 2019, 116, 33–44. [Google Scholar] [CrossRef]
- Negenborn, R.R.; Goerlandt, F.; Johansen, T.A.; Slaets, P.; Valdez Banda, O.A.; et al. Autonomous ships are on the horizon: Here’s what we need to know. Nature 2023, 615, 30–33. [Google Scholar] [CrossRef] [PubMed]
- Alamoush, A.S.; Ölçer, A.I. Maritime autonomous surface ships: Architecture for autonomous navigation systems. J. Mar. Sci. Eng. 2025, 13(no. 1). [Google Scholar] [CrossRef]
- Huang, Y.; Chen, L.; Chen, P.; Negenborn, R.R.; Van Gelder, P. Ship collision avoidance methods: State-of-the-art. Saf. Sci. 2020, 121, 451–473. [Google Scholar] [CrossRef]
- Zhou, C.; Gu, S.; Wen, Y.; Du, Z.; Xiao, C.; et al. The review unmanned surface vehicle path planning: Based on multi-modality constraint. Ocean Eng. 2020, 200, 107043. [Google Scholar] [CrossRef]
- Vagale, A.; Bye, R.T.; Oucheikh, R.; Osen, O.L.; Fossen, T.I. Path planning and collision avoidance for autonomous surface vehicles II: A comparative study of algorithms. J. Mar. Sci. Technol. 2021, 26, 1307–1323. [Google Scholar] [CrossRef]
- Vagale, A.; Oucheikh, R.; Bye, R.T.; Osen, O.L.; Fossen, T.I. Path planning and collision avoidance for autonomous surface vehicles I: A review. J. Mar. Sci. Technol. 2021, 26, 1292–1306. [Google Scholar] [CrossRef]
- Wróbel, K.; Gil, M.; Huang, Y.; Wawruch, R. The vagueness of COLREG versus collision avoidance techniques—A discussion on the current state and future challenges concerning the operation of autonomous ships. Sustainability 2022, 14, 16516. [Google Scholar] [CrossRef]
- Hashali, S.D.; Yang, S.; Xiang, X. Route planning algorithms for unmanned surface vehicles: A comprehensive analysis. J. Mar. Sci. Eng. 2024, 12, 382. [Google Scholar] [CrossRef]
- Li, L.; Wu, D.; Huang, Y.; Yuan, Z.-M. A path planning strategy unified with a COLREGS collision avoidance function based on deep reinforcement learning and artificial potential field. Appl. Ocean Res. 2021, 113, 102759. [Google Scholar] [CrossRef]
- Cui, Y.; Osaki, S.; Matsubara, T. Autonomous boat driving system using sample-efficient model predictive control-based reinforcement learning approach. J. Field Robot. 2021, 38, 331–354. [Google Scholar]
- Xu, X.; Lu, Y.; Liu, G.; Cai, P.; Zhang, W. COLREGs-abiding hybrid collision avoidance algorithm based on deep reinforcement learning for USVs. Ocean Eng. 2022, 247, 110749. [Google Scholar] [CrossRef]
- Xue, D.; Wu, D.; Yamashita, A.S.; Li, Z. Proximal policy optimization with reciprocal velocity obstacle based collision avoidance path planning for multi-unmanned surface vehicles. Ocean Eng. 2023, 273, 114005. [Google Scholar] [CrossRef]
- Wu, C.; Yu, W.; Li, G.; Liao, W. Deep reinforcement learning with dynamic window approach based collision avoidance path planning for maritime autonomous surface ships. Ocean Eng. 2023, 284, 115208. [Google Scholar] [CrossRef]
- Vaaler, A.; Husa, S.J.; Menges, D.; Larsen, T.N.; Rasheed, A. Modular control architecture for safe marine navigation: Reinforcement learning with predictive safety filters. Artif. Intell. 2024, 336, 104201. [Google Scholar] [CrossRef]
- Wu, Y.; Wang, T.; Liu, S. A review of path planning methods for marine autonomous surface vehicles. J. Mar. Sci. Eng. 2024, 12, 833. [Google Scholar] [CrossRef]
- Burmeister, H.-C.; Constapel, M. Autonomous collision avoidance at sea: A survey. Front. Robot. AI 2021, 8, 739013. [Google Scholar] [CrossRef] [PubMed]
- Akdağ, M.; Solnør, P.; Johansen, T.A. Collaborative collision avoidance for maritime autonomous surface ships: A review. Ocean Eng. 2022, 250, 110920. [Google Scholar] [CrossRef]
- Sarhadi, P.; Naeem, W.; Athanasopoulos, N. A survey of recent machine learning solutions for ship collision avoidance and mission planning. IFAC-PapersOnLine 2022, 55, 257–268. [Google Scholar] [CrossRef]
- Lyu, H.; Hao, Z.; Li, J.; Li, G.; Sun, X.; et al. Ship autonomous collision-avoidance strategies—A comprehensive review. J. Mar. Sci. Eng. 2023, 11, 830. [Google Scholar] [CrossRef]
- Zhu, Q.; Xi, Y.; Weng, J.; Han, B.; Hu, S.; et al. Intelligent ship collision avoidance in maritime field: A bibliometric and systematic review. Expert Syst. With Appl. 2024, 252, 124148. [Google Scholar] [CrossRef]
- Li, Y.; Wu, D.; You, Z.; Chen, G.; Wu, D. Deep reinforcement learning for collision avoidance in unmanned surface vehicles: State-of-the-art. Appl. Ocean Res. 2025, 164, 104778. [Google Scholar] [CrossRef]
- Chaal, M.; Ren, X.; BahooToroody, A.; Basnet, S.; Bolbot, V.; et al. Research on risk, safety, and reliability of autonomous ships: A bibliometric review. Saf. Sci. 2023, 167, 106256. [Google Scholar] [CrossRef]
- Li, Z.; Zhang, D.; Han, B.; Wan, C. Risk and reliability analysis for maritime autonomous surface ship: A bibliometric review of literature from 2015 to 2022. Accid. Anal. Prev. 2023, 187, 107090. [Google Scholar] [CrossRef] [PubMed]
- Xue, J.; Yang, P.; Li, Q.; Song, Y.; Van Gelder, P.; et al. Machine learning in maritime safety for autonomous shipping: A bibliometric review and future trends. J. Mar. Sci. Eng. 2025, 13, 746. [Google Scholar] [CrossRef]
- Enevoldsen, T.T.; Blanke, M.; Galeazzi, R. Autonomy for ferries and harbour buses: A collision avoidance perspective. arXiv 2023, arXiv:2301.02711. [Google Scholar]
- Ding, G.; Li, R.; Li, C.; Yang, B.; Li, Y.; et al. Review of ship navigation safety in fog. J. Navig. 2025, First View. 1–21. [Google Scholar]
- Jovanović, I.; Perčić, M.; BahooToroody, A.; Fan, A.; Vladimir, N. Review of research progress of autonomous and unmanned shipping and identification of future research directions. J. Mar. Eng. Technol. 2024, 23, 82–97. [Google Scholar] [CrossRef]
- Hagen, I.B.; Vassbotn, O.; Skogvold, M.; Johansen, T.A.; Brekke, E.F. Safety and COLREG evaluation for marine collision avoidance algorithms. Ocean Eng. 2023, 288, 115991. [Google Scholar] [CrossRef]
- Clement, B.; Chaffre, T.; Sarhadi, P.; Dubromel, M. ColSim, a simulator for hybrid navigation acceptability and safety. IFAC-PapersOnLine 2024, 58, 147–152. [Google Scholar] [CrossRef]
- Vekinis, A.A.; Perantonis, S. Aeolus Ocean—A simulation environment for the autonomous COLREG-compliant navigation of unmanned surface vehicles using deep reinforcement learning and maritime object detection. arXiv 2023, arXiv:2307.06688. [Google Scholar]
- Raza, M.; Prokopova, H.; Huseynzade, S.; Azimi, S.; Lafond, S. Towards integrated digital-twins: An application framework for autonomous maritime surface vessel development. J. Mar. Sci. Eng. 2022, 10, 1469. [Google Scholar] [CrossRef]
- Wei, G.; Kuo, W. COLREGs-compliant multi-ship collision avoidance based on multi-agent reinforcement learning technique. J. Mar. Sci. Eng. 2022, 10, 1431. [Google Scholar] [CrossRef]
- Niu, Y.; Zhu, F.; Wei, M.; Du, Y.; Zhai, P. A multi-ship collision avoidance algorithm using data-driven multi-agent deep reinforcement learning. J. Mar. Sci. Eng. 2023, 11, 2101. [Google Scholar] [CrossRef]
- Verma, S.; Samvedi, A. Cooperative collision avoidance for autonomous vessels in a mixed traffic environment. In Proceedings of the 2024 Winter Simulation Conference; 2024; pp. 549–559. [Google Scholar]
- De La Fuente, N.; Noguer i Alonso, M.; Casadellà, G. Game theory and multi-agent reinforcement learning: From Nash equilibria to evolutionary dynamics. arXiv 2024, arXiv:2412.20523. [Google Scholar]
- Malviya, A.; Rajendran, S. Multi-Agent reinforcement learning for collision avoidance and path following of autonomous surface vehicles. In Proceedings of the 2025 IEEE Underwater Technology; 2025; pp. 1–6. [Google Scholar]
- Pei, D.; He, J.; Liu, K.; Chen, M.; Zhang, S. Application of large language models and assessment of their ship-handling theory knowledge and skills for connected maritime autonomous surface ships. Mathematics 2024, 12, 2381. [Google Scholar] [CrossRef]
- Sanchez-Heres, L.; Weber, R.; Ahlgren, F.; Olsson, F.; Lundström, O. COLREG3: Exploring the potential of large language models in marine navigation systems. In Lighthouse Reports; Lighthouse: Gothenburg, Sweden, 2024. [Google Scholar]
- Christensen, K.A.; Gusev, A.; Tufte, A.G.; Alsos, O.A.; Steinert, M. AI Captain: Conversational mission planning and execution system for autonomous surface vehicles. Ocean Eng. 2025, 338, 121988. [Google Scholar] [CrossRef]
- Agyei, K.; Sarhadi, P.; Naeem, W. Large language model-based decision-making for COLREGs and the control of autonomous surface vehicles. arXiv 2024, arXiv:2411.16587. [Google Scholar]
- Agyei, K.; Sarhadi, P.; Naeem, W. CORALL: A COLREGs-guided risk-aware LLM for decision-making in maritime autonomous surface ships. Authorea Prepr. 2025. [Google Scholar] [CrossRef]
- Din, M.U.; Akram, W.; Bakht, A.B.; Dong, Y.; Hussain, I. Maritime mission planning for unmanned surface vessel using large language model. In Proceedings of the 2025 IEEE International Conference on Simulation, Modeling, and Programming for Autonomous Robots; 2025; pp. 1–6. [Google Scholar]
- Higaki, T.; Hashimoto, H. Human-like route planning for automatic collision avoidance using generative adversarial imitation learning. Appl. Ocean Res. 2023, 138, 103620. [Google Scholar] [CrossRef]
- Zheng, K.; Zhang, X.; Wang, C.; Li, Y.; Cui, J.; et al. Adaptive collision avoidance decisions in autonomous ship encounter scenarios through rule-guided vision supervised learning. Ocean Eng. 2024, 297, 117096. [Google Scholar] [CrossRef]
- Rong, W.; Zheng, J.; Chen, Y.; Liu, Y.; Zhang, Z. Autonomous collision avoidance decision-making method with human-like attention distribution for MASSs based on GMA-TD3 algorithm. Ocean Eng. 2025, 330, 121118. [Google Scholar] [CrossRef]
- Kim, H.; Lee, K.; Park, J.; Li, J.; Park, J. Human implicit preference-based policy fine-tuning for multi-agent reinforcement learning in USV swarm. arXiv 2025, arXiv:2503.03796. [Google Scholar]
- Yoshioka, H.; Hashimoto, H. Explainable AI for ship collision avoidance: Decoding decision-making processes and behavioral intentions. Appl. Ocean Res. 2025, 156, 104471. [Google Scholar] [CrossRef]
- Wang, C.; Wang, N.; Gao, H.; Wang, L.; Zhao, Y.; et al. Knowledge transfer enabled reinforcement learning for efficient and safe autonomous ship collision avoidance. Int. J. Mach. Learn. Cybern. 2024, 15, 3715–3731. [Google Scholar] [CrossRef]
- Jiang, Y.; Zhang, K.; Zhao, M.; Qin, H. Adaptive meta-reinforcement learning for AUVs 3D guidance and control under unknown ocean currents. Ocean Eng. 2024, 309, 118498. [Google Scholar] [CrossRef]
- Menges, D.; Sætre, S.M.; Rasheed, A. Digital twin for autonomous surface vessels to generate situational awareness. Proc. Int. Conf. Offshore Mech. Arct. Eng. 2023, 86878, V005T06A025. [Google Scholar]
- Menges, D.; Von Brandis, A.; Rasheed, A. Digital twin of autonomous surface vessels for safe maritime navigation enabled through predictive modeling and reinforcement learning. Proc. Int. Conf. Offshore Mech. Arct. Eng. 2024, 87837, V05BT06A064. [Google Scholar]
- Vasanthan, C.; Nguyen, D.T. Combining supervised learning and digital twin for autonomous path-planning. IFAC-PapersOnLine 2021, 54, 7–15. [Google Scholar] [CrossRef]
- Larsen, T.N.; Teigen, H.Ø.; Laache, T.; Varagnolo, D.; Rasheed, A. Risk-based convolutional perception models for collision avoidance in autonomous marine surface vessels using deep reinforcement learning. IFAC-PapersOnLine 2023, 56, 10033–10038. [Google Scholar] [CrossRef]
- Lee, P.; Theotokatos, G.; Boulougouris, E. Robust decision-making for the reactive collision avoidance of autonomous ships against various perception sensor noise levels. J. Mar. Sci. Eng. 2024, 12, 557. [Google Scholar] [CrossRef]
- Fan, Z.; Wu, D.; Li, Y.; You, Z.; Zhong, S. Memory-based deep reinforcement learning for COLREGs-compliant obstacle avoidance in USV with limited environmental knowledge. Ocean Eng. 2025, 338, 121978. [Google Scholar] [CrossRef]
- Wu, X.; Wei, C.; Guan, D.; Ji, Z. Risk-aware deep reinforcement learning for mapless navigation of unmanned surface vehicles in uncertain and congested environments. Ocean Eng. 2025, 322, 120446. [Google Scholar] [CrossRef]
- Jin, K.; Liu, Z.; Wang, J. Predictive obstacle avoidance algorithm for underactuated unmanned surface vehicle under disturbances via reinforcement learning. J. Field Robot. 2025. early view. [Google Scholar]
- Jin, K.; Wang, J.; Wang, H.; Liang, X.; Guo, Y.; et al. Soft formation control for unmanned surface vehicles under environmental disturbance using multi-task reinforcement learning. Ocean Eng. 2022, 260, 112035. [Google Scholar] [CrossRef]
- Wang, Q.; Liu, C.; Meng, Y.; Ren, X.; Wang, X. Reinforcement learning-based moving-target enclosing control for an unmanned surface vehicle in multi-obstacle environments. Ocean Eng. 2024, 304, 117920. [Google Scholar] [CrossRef]
- Zhang, J.; Ren, J.; Cui, Y.; Fu, D.; Cong, J. Multi-USV task planning method based on improved deep reinforcement learning. IEEE Internet Things J. 2024, 11, 18549–18567. [Google Scholar] [CrossRef]
- Yin, S.; Xiang, Z. Real-time distributed decision-making for simultaneous target assignment and path planning in multiple unmanned surface vehicles. Expert Syst. With Appl. 2025, 279, 127457. [Google Scholar] [CrossRef]
- Tao, Y.; Du, J.; Lewis, F.L. Integrated intelligent guidance and motion control of USVs with anticipatory collision avoidance decision-making. IEEE Trans. Intell. Transp. Syst. 2024. early access. [Google Scholar]
- Liu, J.; Shi, G.; Zhu, K.; Shi, J. Research on MASS collision avoidance in complex waters based on deep reinforcement learning. J. Mar. Sci. Eng. 2023, 11, 779. [Google Scholar] [CrossRef]
- Zhao, Y.; Han, F.; Han, D.; Peng, X.; Zhao, W. Decision-making for the autonomous navigation of USVs based on deep reinforcement learning under IALA maritime buoyage system. Ocean Eng. 2022, 266, 112557. [Google Scholar] [CrossRef]
- Hao, S.; Guan, W.; Cui, Z.; Lu, J. USV collision avoidance decision-making based on the improved PPO algorithm in restricted waters. J. Mar. Sci. Eng. 2024, 12, 1428. [Google Scholar] [CrossRef]
- Xu, X.; Cao, Y.; Cai, P.; Zhang, W.; Chen, H. Research on real-time collision avoidance and path planning of USVs in multi-obstacle ships environment. Ocean Eng. 2024, 295, 116890. [Google Scholar] [CrossRef]
- Chen, X.; Yin, S.; Li, Y.; Xiang, Z. Dynamic path planning for multi-USV in complex ocean environments with limited perception via proximal policy optimization. Ocean Eng. 2025, 326, 120907. [Google Scholar] [CrossRef]
- Felski, A.; Zwolak, K. The ocean-going autonomous ship—Challenges and threats. J. Mar. Sci. Eng. 2020, 8, 41. [Google Scholar] [CrossRef]
- Rødseth, Ø.J.; Nordahl, H. Definitions for Autonomous Merchant Ships; Norwegian Forum for Autonomous Ships, 2017. [Google Scholar]
- Singh, Y.; Sharma, S.; Sutton, R.; Hatton, D.; Khan, A. A constrained A* approach towards optimal path planning for an unmanned surface vehicle in a maritime environment containing dynamic obstacles and ocean currents. Ocean Eng. 2018, 169, 187–201. [Google Scholar] [CrossRef]
- Zhao, L.; Bai, Y.; Paik, J.K. Global-local hierarchical path planning scheme for unmanned surface vehicles under dynamically unforeseen environments. Ocean Eng. 2023, 280, 114750. [Google Scholar] [CrossRef]
- Kuwata, Y.; Wolf, M.T.; Zarzhitsky, D.; Huntsberger, T.L. Safe maritime autonomous navigation with COLREGS, using velocity obstacles. IEEE J. Ocean. Eng. 2013, 39, 110–119. [Google Scholar] [CrossRef]
- Kufoalor, D.K.M.; Brekke, E.F.; Johansen, T.A. Proactive collision avoidance for ASVs using a dynamic reciprocal velocity obstacles method. In Proceedings of the 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); 2018; pp. 2402–2409. [Google Scholar]
- Tengesdal, T.; Johansen, T.A.; Brekke, E.F. Ship collision avoidance utilizing the cross-entropy method for collision risk assessment. IEEE Trans. Intell. Transp. Syst. 2021, 23, 11148–11161. [Google Scholar] [CrossRef]
- Tam, C.; Bucknall, R. Collision risk assessment for ships. J. Mar. Sci. Technol. 2010, 15, 257–270. [Google Scholar] [CrossRef]
- Du, L.; Banda, O.A.V.; Goerlandt, F.; Huang, Y.; Kujala, P. A COLREG-compliant ship collision alert system for stand-on vessels. Ocean Eng. 2020, 218, 107866. [Google Scholar] [CrossRef]
- Fossen, T.I. Handbook of Marine Craft Hydrodynamics and Motion Control; Wiley, 2011. [Google Scholar]
- Tsolakis, A.; Negenborn, R.R.; Reppa, V.; Ferranti, L. Model predictive trajectory optimization and control for autonomous surface vessels considering traffic rules. IEEE Trans. Intell. Transp. Syst. 2024. early access. [Google Scholar]
- Meyer, E.; Heiberg, A.; Rasheed, A.; San, O. COLREG-compliant collision avoidance for unmanned surface vehicle using deep reinforcement learning. IEEE Access 2020, 8, 165344–165364. [Google Scholar] [CrossRef]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [PubMed]
- Van Hasselt, H.; Guez, A.; Silver, D. Deep reinforcement learning with double Q-learning. Proc. AAAI Conf. Artif. Intell. 2016, 30, 2094–2100. [Google Scholar] [CrossRef]
- Wang, Z.; Schaul, T.; Hessel, M.; Van Hasselt, H.; Lanctot, M.; et al. Dueling network architectures for deep reinforcement learning. Proc. Mach. Learn. Res. 2016, 48, 1995–2003. [Google Scholar]
- Woo, J.; Kim, N. Collision avoidance for an unmanned surface vehicle using deep reinforcement learning. Ocean Eng. 2020, 199, 107001. [Google Scholar] [CrossRef]
- Xu, X.; Lu, Y.; Liu, X.; Zhang, W. Intelligent collision avoidance algorithms for USVs via deep reinforcement learning under COLREGs. Ocean Eng. 2020, 217, 107704. [Google Scholar] [CrossRef]
- Gao, M.; Kang, Z.; Zhang, A.; Liu, J.; Zhao, F. MASS autonomous navigation system based on AIS big data with dueling deep Q networks prioritized replay reinforcement learning. Ocean Eng. 2022, 249, 110834. [Google Scholar] [CrossRef]
- Yang, X.; Lou, M.; Hu, J.; Ye, H.; Zhu, Z.; et al. A human-like collision avoidance method for USVs based on deep reinforcement learning and velocity obstacle. Expert Syst. With Appl. 2024, 124388. [Google Scholar] [CrossRef]
- Fortunato, M.; Azar, M.G.; Piot, B.; Menick, J.; Osband, I.; et al. Noisy networks for exploration. In Proceedings of the International Conference on Learning Representations; 2018. [Google Scholar]
- Sutton, R.S. Dyna, an integrated architecture for learning, planning, and reacting. ACM SIGART Bull. 1991, 2, 160–163. [Google Scholar] [CrossRef]
- Waltz, M.; Okhrin, O. Spatial–temporal recurrent reinforcement learning for autonomous ships. Neural Netw. 2023, 165, 634–653. [Google Scholar] [CrossRef] [PubMed]
- Li, Y.; Wu, D.; Wang, H.; Lou, J. Dynamic collision avoidance for maritime autonomous surface ships based on deep Q-network with velocity obstacle method. Ocean Eng. 2025, 320, 120335. [Google Scholar] [CrossRef]
- Guo, S.; Zhang, X.; Du, Y.; Zheng, Y.; Cao, Z. Path planning of coastal ships based on optimized DQN reward function. J. Mar. Sci. Eng. 2021, 9, 210. [Google Scholar] [CrossRef]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar]
- Mnih, V.; Badia, A.P.; Mirza, M.; Graves, A.; Lillicrap, T.; et al. Asynchronous methods for deep reinforcement learning. Proc. Mach. Learn. Res. 2016, 48, 1928–1937. [Google Scholar]
- Lillicrap, T.P.; Hunt, J.J.; Pritzel, A.; Heess, N.; Erez, T.; et al. Continuous control with deep reinforcement learning. In Proceedings of the International Conference on Learning Representations; 2016. [Google Scholar]
- Fujimoto, S.; van Hoof, H.; Meger, D. Addressing function approximation error in actor-critic methods. Proc. Mach. Learn. Res. 2018, 80, 1587–1596. [Google Scholar]
- Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. Proc. Mach. Learn. Res. 2018, 80, 1861–1870. [Google Scholar]
- Meyer, E.; Robinson, H.; Rasheed, A.; San, O. Taming an autonomous surface vehicle for path following and collision avoidance using deep reinforcement learning. IEEE Access 2020, 8, 41466–41481. [Google Scholar] [CrossRef]
- Heiberg, A.; Larsen, T.N.; Meyer, E.; Rasheed, A.; San, O.; et al. Risk-based implementation of COLREGs for autonomous surface vehicles using deep reinforcement learning. Neural Netw. 2022, 152, 17–33. [Google Scholar] [CrossRef] [PubMed]
- Teitgen, R.; Monsuez, B.; Kukla, R.; Pasquier, R.; Foinet, G. Dynamic trajectory planning for ships in dense environment using collision grid with deep reinforcement learning. Ocean Eng. 2023, 281, 114807. [Google Scholar] [CrossRef]
- Wang, W.; Li, M.; Chen, G.; Yang, S.; Suo, Y.; et al. Cognitive entropy proximal policy optimization for autonomous ship collision avoidance based on deep reinforcement learning. Eng. Appl. Artif. Intell. 2026, 172, 114416. [Google Scholar] [CrossRef]
- Xie, S.; Chu, X.; Zheng, M.; Liu, C. A composite learning method for multi-ship collision avoidance based on reinforcement learning and inverse control. Neurocomputing 2020, 411, 375–392. [Google Scholar] [CrossRef]
- Xu, X.; Cai, P.; Cao, Y.; Chu, Z.; Zhu, W.; et al. Real-time planning and collision avoidance control method based on deep reinforcement learning. Ocean Eng. 2023, 281, 115018. [Google Scholar] [CrossRef]
- Lou, M.; Yang, X.; Hu, J.; Shen, H.; Xu, B.; et al. Design and field test of collision avoidance method with prediction for USVs: A deep deterministic policy gradient approach. IEEE Internet Things J. 2024. early access. [Google Scholar]
- Sun, X.; Li, G.; Liu, Z.; Zhang, L.; Yu, H.; et al. Path planning algorithm for unmanned surface vessels based on COLREGs and Meta-TD3. Ocean Eng. 2025, 334, 121580. [Google Scholar] [CrossRef]
- Waltz, M.; Paulig, N.; Okhrin, O. 2-level reinforcement learning for ships on inland waterways: Path planning and following. Expert Syst. With Appl. 2025, 274, 126933. [Google Scholar] [CrossRef]
- Zhao, Y.; Han, F.; Han, D.; Peng, X.; Zhao, W.; et al. A port water navigation solution based on priority sampling SAC: Taking Yantai port environment as an example. Robot. Auton. Syst. 2025, 188, 104956. [Google Scholar] [CrossRef]
- Jin, K.; Liu, Z.; Wang, J.; Wang, H. Unmanned surface vehicle navigation under disturbances: World model enhanced reinforcement learning. IEEE/ASME Transactions on Mechatronics 2025. early access. [Google Scholar] [CrossRef]
- Rashid, T.; Samvelyan, M.; Schroeder de Witt, C.; Farquhar, G.; Foerster, J.; et al. QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning. Proc. Mach. Learn. Res. 2018, 80, 4295–4304. [Google Scholar]
- Wang, Y.; Zhao, Y. Multiple ships cooperative navigation and collision avoidance using multi-agent reinforcement learning with communication. Ocean Eng. 2025, 320, 120244. [Google Scholar] [CrossRef]
- Lowe, R.; Wu, Y.; Tamar, A.; Harb, J.; Abbeel, P.; et al. Multi-agent actor-critic for mixed cooperative-competitive environments. Adv. Neural Inf. Process. Syst. 2017, 30, 6379–6390. [Google Scholar]
- Yu, C.; Velu, A.; Vinitsky, E.; Gao, J.; Wang, Y.; et al. The surprising effectiveness of PPO in cooperative, multi-agent games. Adv. Neural Inf. Process. Syst. 2022, 35. [Google Scholar]
- Wang, W.; Wu, H.; Yang, S.; Mei, X.; Han, D.; et al. LNPP: Logical neural path planning of mobile beacon for ocean sensor networks in uncertain environments using hierarchical reinforcement learning. IEEE Trans. Netw. Sci. Eng. 2025, 12, 2606–2621. [Google Scholar] [CrossRef]
- MahmoudZadeh, S.; Abbasi, A.; Yazdani, A.; Wang, H.; Liu, Y. Uninterrupted path planning system for multi-USV sampling mission in a cluttered ocean environment. Ocean Eng. 2022, 254, 111328. [Google Scholar] [CrossRef]
- Sawada, R.; Sato, K.; Majima, T. Automatic ship collision avoidance using deep reinforcement learning with LSTM in continuous action spaces. J. Mar. Sci. Technol. 2021, 26, 509–524. [Google Scholar]
- Kordabad, A.B.; Esfahani, H.N.; Lekkas, A.M.; Gros, S. Reinforcement learning based on scenario-tree MPC for ASVs. In Proceedings of the 2021 American Control Conference; 2021; pp. 1985–1990. [Google Scholar]
- Hart, F.; Waltz, M.; Okhrin, O. Two-step dynamic obstacle avoidance. arXiv 2024, arXiv:2311.16841v2. [Google Scholar]
- Zhang, P.; Chen, Q.; Macdonald, T.; Lau, Y.-Y.; Tang, Y.-M. Game change: A critical review of applicable collision avoidance rules between traditional and autonomous ships. J. Mar. Sci. Eng. 2022, 10, 1655. [Google Scholar] [CrossRef]
- Ames, A.D.; Coogan, S.; Egerstedt, M.; Notomista, G.; Sreenath, K.; et al. Control barrier functions: Theory and applications. In Proceedings of the 18th European Control Conference; 2019; pp. 3420–3431. [Google Scholar]








| Method Paradigm | Representative Methods | Primary Role | Main Limitations |
|---|---|---|---|
| Search- and sampling-based | A*, D*, PRM, and RRT* | Global route generation | High computational cost for dynamic replanning |
| Geometry- and rule-based | APF, VO, DWA, and COLREGs | Local collision avoidance and rule enforcement | Limited adaptability to multi-vessel scenarios |
| Optimization-based | MPC, GA, PSO, and DE | Trajectory planning under dynamic constraints | High online computational burden |
| Learning-driven | DQN, PPO, TD3, and SAC | Complex interaction modeling and continuous decision-making | Limited generalization capability and insufficient safety guarantees |
| Hybrid and hierarchical | A*–DRL, MPC–RL, and safety filters | Coordinated global–local decision-making | Complex system interfaces and verification procedures |
| Autonomy Level | Decision Characteristics | Role of Path Planning | Key Requirements |
|---|---|---|---|
| Level 1 | Crew-led operation with system assistance | Route recommendations and risk warnings | Interpretability and ease of verification |
| Level 2 | Human supervision with constrained autonomous operation | Local route adjustment and assisted collision avoidance | Stability and ease of human takeover |
| Level 3 | Remote supervision with autonomous operation | Dynamic replanning and rule-compliant collision avoidance | Real-time responsiveness and safe degradation |
| Level 4 | Fully autonomous operation | Closed-loop global–local path planning | Robustness, verifiability, and generalizability |
| Screening Category | Inclusion Criteria | Exclusion Criteria |
|---|---|---|
| Research platform | USVs, ASVs, MASSs, and related surface vessels | Air, ground, or underwater platforms without direct relevance to ship planning |
| Research task | Path, route, or trajectory planning; collision avoidance; integrated planning and control | Path following, sensing, communication, or task management without planning decisions |
| Method category | Search, sampling, geometric, rule-based, optimization, DRL, MARL, hybrid, and safety methods | Control or perception methods unrelated to path planning |
| Evaluation | Clearly defined scenarios, constraints, metrics, or comparative experiments | Insufficient methodological detail, unclear scenarios, or unsupported results |
| Publication type | Journal papers, conference papers, reviews, and selected relevant preprints | Duplicates, overlapping studies, or publications with limited relevance |
| Component | Representative Studies | Integration with DRL | Primary Contribution |
|---|---|---|---|
| Global search and sampling | Singh et al. [73]; Zhao et al. [74] | Supply reference routes, navigable-space priors, and mission-level guidance | Preserve global route consistency and prevent purely local, short-sighted decisions |
| Artificial potential fields | Li et al. [12] | Encode goal attraction and obstacle repulsion in states or rewards | Improve goal-directed exploration and accelerate convergence |
| Velocity obstacles and dynamic windows | Xue et al. [15]; Wu et al. [16] | Provide collision-risk features or restrict the admissible action space | Anticipate unsafe velocities and reduce ineffective exploration |
| COLREGs modules | Meyer et al. [82]; Xu et al. [14] | Encode encounter responsibilities through rewards, state features, or action constraints | Improve rule compliance and the interpretability of avoidance manoeuvres |
| MPC and safety filters | Cui et al. [13]; Vaaler et al. [17] | Correct or reject policy actions that violate model-based constraints | Enforce dynamic feasibility and safety at execution time |
| Study | Method | Application Setting | Main Contribution |
|---|---|---|---|
| Woo and Kim [86] | Double and Dueling DQN | Representative encounters and multi-vessel collision avoidance | Established the feasibility of DQN variants for discrete USV collision-avoidance decisions |
| Xu et al. [87] | DQN with COLREGs constraints | Rule-compliant USV collision avoidance | Embedded navigational responsibilities into discrete policy learning |
| Guo et al. [94] | Reward-optimized DQN | Coastal ship path planning | Balanced navigational safety and route efficiency through reward design |
| Li et al. [12] | DQN + APF + COLREGs | Static and dynamic obstacle environments | Combined goal guidance, geometric risk information, and rule-compliant decision-making |
| Gao et al. [88] | Dueling DQN + prioritized replay | AIS-driven MASS navigation | Improved value estimation and sample efficiency by prioritizing informative transitions |
| Liu et al. [66] | Dyna-DQN | MASS collision avoidance in complex waters | Augmented value learning with model-generated experience |
| Yang et al. [89] | Dueling DQN + velocity obstacle | Multi-vessel encounters | Integrated geometric risk prediction with value-based decision-making |
| Li et al. [93] | DQN + velocity obstacle | Dynamic collision avoidance for MASSs | Improved prospective identification of unsafe vessel motions |
| Study | Method | Action Output | Application Setting | Main Contribution |
|---|---|---|---|---|
| Meyer et al. [82] | PPO | Continuous avoidance actions | COLREGs-compliant USV collision avoidance | Learned rule-compliant policies through COLREGs-informed reward design |
| Meyer et al. [100] | PPO | Thrust and control moments | Integrated path following and collision avoidance | Unified trajectory tracking and collision avoidance within one policy |
| Heiberg et al. [101] | PPO | Continuous control actions | Risk-based COLREGs-compliant navigation | Embedded navigational rules into risk indicators for policy learning |
| Xue et al. [15] | PPO + RVO | Continuous local actions | Multi-USV path planning | Introduced reciprocal-velocity-obstacle information to improve local safety |
| Teitgen et al. [102] | DRL / PPO | Local planning actions | Dense maritime traffic | Represented dynamic collision risk using collision-grid observations |
| Wu et al. [16] | PPO + DWA | Velocity and steering actions | Local collision avoidance for MASSs | Constrained local decisions using dynamic-window information |
| Lee et al. [57] | PPO | Continuous avoidance actions | Navigation under perception noise | Evaluated policy robustness across different sensor-noise levels |
| Vaaler et al. [17] | PPO + safety filter | Safety-corrected actions | Safety-oriented marine navigation | Corrected unsafe policy outputs using predictive safety filtering |
| Wang et al. [103] | CEPPO | Continuous avoidance actions | Autonomous-ship collision avoidance | Adapted exploration intensity through cognitive entropy |
| Xie et al. [104] | A3C | Continuous avoidance actions | Multi-vessel encounters | Increased experience-collection throughput through asynchronous training |
| Xu et al. [105] | DDPG | Rudder angle and thrust | COLREGs-compliant dynamic avoidance | Generated collision-avoidance commands directly in a continuous action space |
| Lou et al. [106] | DDPG | Continuous control actions | Predictive USV collision avoidance | Combined motion prediction with continuous manoeuvring and field validation |
| Sun et al. [107] | Meta-TD3 | Continuous avoidance actions | COLREGs-compliant USV planning | Used meta-learning to improve adaptation across encounter scenarios |
| Waltz et al. [108] | LSTM-TD3 | Planning and tracking actions | AIS-replay inland-waterway navigation | Combined temporal memory with hierarchical planning and path following |
| Rong et al. [48] | GMA-TD3 | Continuous avoidance actions | Complex multi-vessel encounters | Used recurrent attention to prioritize high-risk vessels and regions |
| Zhao et al. [109] | SAC | Thrust and rudder commands | Port-water navigation | Integrated IALA buoyage rules and environmental disturbances |
| Jin et al. [110] | World model + SAC | Continuous control actions | Navigation under environmental disturbances | Improved disturbance adaptation by learning latent environmental dynamics |
| Study | Method | Interaction Setting | Coordination or Safety Mechanism | Main Contribution |
|---|---|---|---|---|
| Multi-agent methods | ||||
| Wei and Kuo [35] | MADRL | COLREGs-compliant multi-vessel encounters | Cooperative multi-agent decision-making | Applied MARL to rule-constrained multi-vessel collision avoidance |
| Niu et al. [36] | Data-driven MADRL | Multi-vessel collision avoidance | Cooperative policies learned from interaction data | Learned coordinated avoidance behaviour without prescribing fixed manoeuvring rules |
| Verma and Samvedi [37] | Cooperative MARL | Mixed maritime traffic | Coordination among heterogeneous traffic participants | Extended cooperative avoidance beyond homogeneous autonomous fleets |
| Wang and Zhao [112] | Communication-enhanced MADRL | Cooperative multi-vessel navigation | Inter-vessel exchange of intent and risk information | Improved coordination under limited local observations |
| Malviya and Rajendran [39] | MAPPO / MADRL | ASV collision avoidance and path following | Centralized training with decentralized execution | Extended multi-agent policy optimization from conflict resolution to cooperative navigation |
| Yin and Xiang [64] | Distributed MARL | Target assignment and path planning for multiple USVs | Distributed real-time decision-making | Unified mission allocation and cooperative path planning |
| Hybrid methods | ||||
| Li et al. [12] | DQN + APF + COLREGs | Navigation among static and dynamic obstacles | Potential-field guidance and rule-based constraints | Combined goal attraction, geometric avoidance, and regulatory guidance |
| Xu et al. [14] | Hybrid DRL | COLREGs-compliant collision avoidance | Rule-based module coupled with a learned policy | Improved regulatory consistency of learned manoeuvres |
| Xue et al. [15] | PPO + RVO | Multi-USV local path planning | Reciprocal-velocity-obstacle constraints | Restricted policy decisions to geometrically safer velocity regions |
| Wu et al. [16] | PPO + DWA | Dynamic collision avoidance for MASSs | Dynamic-window-based action constraints | Improved local feasibility and manoeuvring smoothness |
| Vaaler et al. [17] | RL + predictive safety filter | Safety-critical marine navigation | Online verification and correction of policy actions | Enforced safety and feasibility constraints at execution time |
| Design Dimension | Typical Formulation | Representative Studies | System Role | Persistent Limitation |
|---|---|---|---|---|
| Objectives and constraints | ||||
| Global efficiency | Path length, travel time, and reference routes | Singh et al. [73]; Zhao et al. [74] | Provide mission-level guidance | Limited adaptation to rapidly changing traffic |
| Collision safety | DCPA, TCPA, ship domains, and collision probability | Tam and Bucknall [78]; Huang et al. [6]; Tengesdal et al. [77] | Estimate and constrain encounter risk | Sensitivity to thresholds and prediction errors |
| Regulatory compliance | Rule-based rewards, action masking, and encounter classification | Kuwata et al. [75]; Meyer et al. [82]; Xu et al. [14] | Generate rule-consistent manoeuvres | Ambiguous and context-dependent COLREGs semantics |
| Dynamic feasibility | Motion models, curvature limits, and control penalties | Fossen [80]; Tsolakis et al. [81]; Meyer et al. [100] | Produce executable and smooth trajectories | Model mismatch and efficiency–smoothness trade-offs |
| Energy efficiency | Speed, thrust, and energy costs | MahmoudZadeh et al. [116]; Waltz et al. [108] | Reduce propulsion or mission cost | Strong sensitivity to objective weights |
| State and environment representation | ||||
| Compact kinematic states | Relative distance, bearing, heading, and speed | Woo and Kim [86]; Xu et al. [87] | Support stable and sample-efficient learning | Limited representation of complex traffic |
| Spatial risk representations | Occupancy maps, risk maps, and collision grids | Teitgen et al. [102]; Larsen et al. [56] | Represent dense and spatially distributed hazards | High sensing and computational cost |
| Temporal and multisource states | LSTM, GRU, history windows, AIS, radar, LiDAR, and ENC | Sawada et al. [117]; Waltz and Okhrin [92]; Lee et al. [57]; Fan et al. [58] | Address partial observability and temporal interactions | Noise sensitivity and long-horizon credit assignment |
| Prior knowledge and learning mechanisms | ||||
| Geometric priors | APF, VO, RVO, and DWA | Li et al. [12]; Xue et al. [15]; Wu et al. [16] | Reduce unsafe or ineffective exploration | Dependence on local geometric assumptions |
| Predictive constraints | MPC and predictive safety filters | Kordabad et al. [118]; Vaaler et al. [17] | Verify and correct policy actions online | Computational burden and model dependence |
| Exploration | Entropy regularization, parameter noise, and stochastic policies | Fortunato et al. [90]; Zhao et al. [109] | Avoid premature convergence and local optima | Sensitivity to entropy and noise parameters |
| Sample efficiency | Dyna, digital twins, and world models | Sutton [91]; Menges et al. [54]; Jin et al. [110] | Reduce dependence on costly real interactions | Bias introduced by learned or simulated models |
| Interaction and validation | ||||
| Multi-vessel coordination | CTDE, communication, and centralized critics | Wei and Kuo [35]; Wang and Zhao [112] | Coordinate coupled vessel decisions | Non-stationarity and limited scalability |
| Human-compatible decisions | Imitation learning, attention, and preference learning | Higaki and Hashimoto [46]; Rong et al. [48]; Kim et al. [49] | Improve interpretability and behavioural acceptability | Unclear relationship between human likeness and safety |
| Scenario-based validation | Synthetic encounters, AIS replay, and disturbance injection | Heiberg et al. [101]; Hart et al. [119]; Wu et al. [59] | Support controlled training and evaluation | Absence of standardized benchmarks |
| Real-system validation | Hardware-in-the-loop, port tests, and vessel trials | Lou et al. [106]; Zhao et al. [109] | Assess operational feasibility | Limited scale and diversity of trials |
| Challenge | Deployment Risk | Representative Studies | Research Priority | Required Evidence |
|---|---|---|---|---|
| Rules and interactive decision-making | ||||
| Rule semantics and responsibility transfer | Ambiguous COLREGs terms and changing roles may produce inconsistent manoeuvres | Wróbel et al. [10]; Zhang et al. [120]; Akdağ et al. [20]; Wei and Kuo [35] | Context-aware rule models and dynamic responsibility reasoning | Compliance, role consistency, manoeuvre timing, and expert agreement |
| Intent uncertainty and mixed traffic | Unexpected or heterogeneous behaviour can invalidate motion predictions | Huang et al. [6]; Waltz and Okhrin [92]; Verma and Samvedi [37] | Probabilistic intent prediction and heterogeneous-agent modelling | Prediction accuracy and safety under behavioural uncertainty |
| Reward specification and rare events | Policies may exploit rewards and underperform in hazardous encounters | Meyer et al. [82]; Xu et al. [14]; Xue et al. [15]; Teitgen et al. [102] | Preference learning, inverse RL, and risk-focused scenario generation | Reward transferability, rare-event coverage, and worst-case safety |
| Robustness and generalization | ||||
| Partial observability and perception error | Occlusion, missed detections, and sensor errors distort risk estimates | Sawada et al. [117]; Fan et al. [58]; Lee et al. [57]; Wu et al. [59] | Uncertainty-aware perception, memory, fusion, and belief-state planning | Robustness to noise, delay, occlusion, and detection errors |
| Environmental and vessel generalization | Policies may fail under unseen dynamics, waterways, or disturbances | Fossen [80]; Jin et al. [61,110]; Tsolakis et al. [81] | Domain randomization, meta-learning, adaptive models, and world models | Performance degradation across vessels, waterways, and disturbances |
| Benchmarking and reproducibility | Inconsistent models, scenarios, thresholds, and metrics hinder comparison | Hagen et al. [31]; Waltz et al. [108] | Standardized scenarios, dynamics, metrics, and reporting protocols | Repeated-run statistics, common metrics, code, and computational cost |
| Safety assurance and deployment | ||||
| Hard safety constraints | High expected return cannot exclude rare constraint violations | Vaaler et al. [17]; Ames et al. [121] | Safety filters, barrier functions, runtime monitors, and fallback policies | Violation rate, minimum separation, intervention, and recovery success |
| Real data and Sim-to-Real transfer | Simulator bias may cause substantial deployment performance loss | Heiberg et al. [101]; Menges et al. [54]; Vasanthan and Nguyen [55]; Zhao et al. [109] | Digital-twin loops, hardware-in-the-loop testing, and staged trials | Transfer loss, data fidelity, latency, and disturbance robustness |
| Multi-agent scalability and credit assignment | Joint-state growth and unclear contributions destabilize coordination | Niu et al. [36]; Wang and Zhao [112]; Rashid et al. [111]; Yu et al. [114] | Scalable CTDE, sparse communication, graphs, and value decomposition | Fleet-size scaling, communication cost, and agent-level safety |
| Interpretability and human acceptance | Opaque policies are difficult to audit, predict, and trust | Yoshioka and Hashimoto [50]; Rong et al. [48] | Faithful explanations and navigator-centred interfaces | Explanation fidelity, expert agreement, and takeover performance |
| Semantic mission reasoning | LLM outputs may be inconsistent, unverifiable, or dynamically infeasible | Pei et al. [40]; Agyei et al. [44] | Constrained LLM supervision linked to verified planners | Semantic consistency, reasoning accuracy, and intervention rate |
| Certification and assurance | No unified assurance process exists for adaptive planners | Negenborn et al. [4]; Alamoush and Ölçer [5] | Staged assurance combining simulation, formal analysis, HIL, and trials | Scenario coverage, traceability, failure handling, and maturity |
| System integration and error propagation | Errors may accumulate across perception, planning, and control modules | Tao et al. [65]; Yin and Xiang [64] | Verified interfaces, uncertainty propagation, monitoring, and degradation | End-to-end latency, interface failures, stability, and degraded-mode safety |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).