Submitted:
14 April 2023
Posted:
17 April 2023
You are already at the latest version
Abstract
The capabilities of autonomous mobile robotic systems have been steadily improving due to recent advancements in computer science, engineering, and related disciplines such as cognitive science. In controlled environments, autonomous robots have been able to achieve relatively high levels of autonomy. In more unstructured environments, however, the realisation of autonomous mobile robots remains challenging due to limitations in the robots’ external environment understanding. Many autonomous mobile robots use classical, learning-based or hybrid approaches for navigation. The classical navigation approach typically includes robot perception, localisation, environmental mapping, path planning and motion control stages. More recent learning-based methods may replace the complete navigation pipeline or selected stages of the classical approach. For effective deployment, autonomous robots need to be able to understand their external environments at a sophisticated level according to their intended applications. Therefore, in addition to robot perception, scene analysis and higher-level scene understanding (e.g., traversable/non-traversable, and rough or smooth terrain) are required for autonomous robot navigation in unstructured outdoor environments. A wide number of alternative approaches have been proposed in recent years to attempt to address these scene understanding requirements. This paper provides a comprehensive review and critical analysis of these methods in the context of their applications to the problems of robot perception and scene understanding in unstructured environments, and the related problems of localisation, environment mapping and path planning. State-of-the-art sensor fusion methods and multimodal scene understanding approaches are also discussed and evaluated within this context. The paper concludes with an in-depth discussion regarding the current state of the autonomous ground robot navigation challenge in unstructured outdoor environments and the most promising future research directions to overcome these challenges.
Keywords:
unstructured environments
; mobile robots
; robot navigation
; perception
; scene understanding
; path planning
; autonomous robots
; ground robots
1. Introduction
The use of autonomous mobile robotic systems is rapidly expanding in many scientific and commercial fields, and the capabilities of these robots have been growing due to continuous research and industrial efforts. The range of different research application areas in autonomous mobile robotics is wide, including areas such as surveillance, space exploration, defence, petrochemical, industrial automation, disaster management, construction, marine, personal assistance, extreme environments, sports entertainment, agriculture, transportation logistics, and many other industrial and non-industrial applications [1]. Many different types of robot platforms have been and continue to be developed for these applications and autonomous mobile robotics is a global and continuously evolving scientific field. Robots that move in contact with ground surfaces are commonly referred to as mobile ground robots, and these robots may be deployed in different working environments, including indoor or outdoor, structured or unstructured, and in proximity to static or dynamic actors. Each of these environments creates various challenges for robot applications. The overall system configuration of a robot mostly depends on the relevant operating environment, for example, a robot designed for a static environment may not effectively adapt to dynamic situations that arise within the environment [2].
The development of ground robotic systems is particularly challenging when the intended application area of autonomous vehicles/robots falls into the unstructured off-road category, making this a very active area of research. The difficulties mainly arise due to the weak scene understanding of robots in these unstructured environment scenarios. The agriculture industry is one example of a relevant application area to deploy scene understanding based off-road autonomous robotic systems. In recent times, the agriculture industry has shown a growing adoption of robotics technologies, but in general, the involvement of novel technologies depends on economic sustainability. Ground mobile robots and manipulators are already used in precision farming to pick fruits, harvest vegetables and for weeding, but their application areas are generally fairly narrow due to limited scene understanding capability. In the past few years, self-driving vehicles have been gradually growing in technological capability and market size within the automobile industry [3,4]. Despite these developments, autonomous driving is still a very challenging problem, and in complex scenarios the performance level remains below that of an average human operator. This is particularly true when the intended application area of autonomous vehicles/robots includes off-road areas. This low performance mainly arises due to the weak external environment understanding of robots. In general, application areas such as disaster management, environment exploration, defence, mining and transportation are associated with complex and unstructured environments. Therefore, these fields are also driving further research on mobile robot scene understanding, as it relates to the broader topic of autonomous robot navigation.
Several challenges of the classical autonomous robot navigation pipeline remain for current robotic systems under the topics of
- perception,
- localisation and mapping, and
- scene understanding.
In robot perception, robots sense environments using different sensors and extract actionable information via the sensor data. Perception plays an important role in the realisation of autonomous robots. For robots to perform effectively in unstructured outdoor environments in real-time, it is essential that they possess accurate, reliable, and robust perception capabilities. To achieve these characteristics, in general, autonomous robots in complex scenarios are equipped with several sensor modalities (that can be exteroceptive or proprioceptive) [5,6]. Different modalities such as sound, pressure, temperature, light, and contact have been used in robot environmental perception applications [7]. Sensor fusion (combining different sensor modalities) has been applied in many recent autonomous mobile robot/self-driving applications [8]. Multimodal sensor fusion brings the complementary properties of different sensors together to achieve better environment perception across a range of conditions [9]. Many recent deep learning-based sensor fusion methods have shown higher robustness in perception than conventional mono-sensor methodologies [6,10]. Camera, Light Detection and Ranging (LiDAR), radar, ultrasonic, Global Navigation Satellite System (GNSS), inertial measurement unit (IMU) and odometry sensors are used in many mobile ground robot perception applications. The most frequently used robot vision-based sensor fusion methods combine camera images with LiDAR point clouds. Research into sensor fusion techniques remains an important component of achieving better sensing capabilities for unstructured outdoor environments.
In the second stage of the autonomous robot navigation pipeline, the information retrieved from perception sensors is used for robot localisation and to map their external environments. Autonomous mobile robots require accurate and reliable environmental mappings and localisation at a sophisticated level based on the application context. The Simultaneous Localisation and Mapping (SLAM) approach is a commonly used technique in autonomous mobile robot systems to represent the robot positions and the map of their external environments when both the robot pose and environmental map are previously unknown. Many SLAM systems use LiDAR sensors and vision-based sensors such as Red-Green-Blue (RGB)/RGB-Depth (RGB-D) cameras that can retrieve visual or proximity information, and the fusion of these sensors has achieved better robustness than using either camera or LiDAR sensors alone [11].
A fundamental aim of robotic vision is to interpret the semantic information present in a scene to provide scene understanding. Scene understanding goes beyond object detection and requires analysis and elaboration of the data retrieved by the sensors [12]. This concept is used in many practical applications such as self-driving vehicles, surveillance, transportation, mobile robot navigation, and activity recognition. Understanding scenes using images or videos is a complex problem, however, and requires more steps than just recording and extracting features [13]. Scene understanding can be aided by taking advantage of multiple sensor modalities, and this is usually termed multimodal scene understanding [14]. One of the requirements for autonomous robots operating in unstructured environments is a capacity to understand the surrounding environment. Scene understanding comprises subtasks such as depth estimation, scene categorisation, object detection, object tracking, and event categorisation [15]. These scene understanding sub-tasks can describe different aspects of a scene acquired by perception sensors. In scene understanding, a representation is given to a scene by carrying out some of the above tasks jointly to get a holistic understanding of the retrieved scene. To generate this overall representation, the information observed from the above-mentioned scene understanding subtasks must be combined meaningfully. Some of the early methods used to obtain scene understanding include using a block world assumption [16], or bottom-up top-down inference [17], and many of these early works have depended on heuristics rather than learning-based methods and, thus, were not suitable for generalisation to unstructured real-world scenarios [15]. The most used recent methods try to acquire information from deep learning-based (i.e., convolutional neural networks (CNNs), graph neural networks, vision transformers) approaches. Scene understanding in unstructured outdoor environments using multimodal scene understanding concepts is a challenging task. Many modern scene understanding methods use feature-based high-level representations of environments. In unstructured environments, however, detection of useful object features is challenging. Therefore, it is difficult to reliably interpret visual information from unstructured or dynamic environments [18]. However, to attain real-world effectiveness, robots should understand their operating environments up to a level that is accurate enough to execute real-time and goal-oriented decisions.
One of the key capabilities required for autonomous robot operation is autonomous path planning for robot navigation. Path planning is generally separated into global and local planning. For global planning, previous knowledge of the operating environment is necessary, and this planning method is also identified as an off-line mode for robot path planning. Robot local path planning, also known as online robot path planning, allows for real-time decisions to be made by the robot in response to perception of the local environment. Autonomous robot local path planning for optimal terrain traversal in unstructured outdoor environments is an important challenge to solve for robots operating in off-road conditions. This is due to the limitations of standard path planning algorithms, which are incapable of performing the desired tasks in dynamic or unstructured environments where the system lacks prior knowledge and/or already existing maps [19]. Artificial potential fields, simulated annealing, fuzzy logic, artificial neural networks, and dynamic window [20] approaches are some of the algorithms that have been used in robot local path planning [21]. An optimal local path planning approach should enable robots to adaptably deal with their environments, such as assisting in avoiding dynamic obstacles or identifying traversable routes through varying terrain conditions. In unstructured environments, classical path planning concepts such as Rapidly exploring Random Tree (RRT) variants, Batch Informed Tree (BIT), D* algorithm variants, artificial potential field methods, A* algorithm variants and learning-based path planning methods have been used [22]. Over the past years, deep learning has been used to enhance the performance of sensor fusion [23], multimodal scene understanding [24] and robot local path planning [25] techniques.
For autonomous navigation, a robot must have the ability to understand both its pose and enough about the external environment to determine an optimal and traversable path to safely reach a goal position without human assistance in the robot control loop. However, fully autonomous navigation in unstructured environments has not yet been achieved despite significant advancements in computing and engineering technology. Autonomous vehicle navigation on urban roads has been of great interest due to its emerging commercial applications around the world. As a result, autonomous perception technologies are developing continuously due to the competitiveness of this industry [26]. Compared to the level of development of methods for autonomous navigation in urban settings [27], however, autonomous navigation in off-road scenarios has not been studied to the same extent, meaning there are significant opportunities for new research in this area. A large number of autonomous navigation techniques have been explored and these can be broadly divided into two subsections according to the approach used to execute control commands following the processing of the input sensory data.
As shown in Figure 1, autonomous navigation methods can be categorised as the classical modular pipeline or end-to-end learning-based approaches. The modular pipeline architecture includes intermediate representations that humans can interpret and that can provide information related to the failure modes of the approach. Furthermore, the system development can be parallelised among several expert teams because of its modular nature. Perception, localisation and mapping, path planning and robot control modules may consist of classical, learning-based or hybrid methods [8,28]. Due to algorithmic and modelling limitations, however, the modular approach may not be optimal for general autonomous navigation applications. Also, machine learning-based modules must be separately trained and validated using auxiliary loss functions, which can create suboptimality across the system as a whole. In contrast, end-to-end learning-based systems learn policies from observations of the outcomes that result from the actions of autonomous systems [16,29]. In these systems, Deep Reinforcement Learning (DRL) is used to refine the control algorithms used to determine autonomous actions. However, commonly employed methods like imitation learning can suffer from overfitting and cause problems with the poor generalisation of the system behaviour when deploying in different environment scenarios [30]. Also, when errors occur in an end-to-end system, it can be complex to investigate because the origin of these errors is hidden in the holistic neural network-based architectures [8,31]. It is clear, however, that future development of autonomous robot navigation will rely on further advancements in robot environmental perception and machine learning.
The remainder of this paper is structured in the following manner. A discussion of robot vision and the types of active ranging sensors that are used for robot environment perception is presented in Section 2. This section also discusses deep learning-based camera and LiDAR sensor fusion methods for depth estimation, object detection, and semantic and instance segmentation. Section 3 describes modern mobile ground robot scene understanding techniques and scene representations that are utilised in robot navigation. Robot path planning algorithms at the global and local levels are discussed in Section 4. In Section 5, the details of robot vision and ranging sensors, fusion methods, scene understanding concepts, and local navigation approaches are summarised. Section 6 provides an overview of the research challenges related to the topic. Additionally, it offers some potential future research directions. Finally, conclusion of the review can be found in Section 7.
4. Mobile Robot Local Path Planning
Approaches for path planning of robots in local environments can be broadly classified into either classical methods or learning-based methods that attempt to modify the robots’ path-planning based on environmental conditions [129,130]. The classical methods follow a modular architecture with environmental perception, planning of paths relative to generated global maps, and trajectory following. Many classical techniques are applicable for robot navigation in static environments however, their suitability can substantially diminish in unstructured or dynamic environments [131]. In general, these methods can be used effectively in indoor mobile robot applications but may not be suitable for outdoor off-road navigation conditions (e.g., terrain with grass might be traversable, despite appearing blocked, and ground with mud or sand may not be suitable for pathing, despite appearing clear), hence the requirement to investigate machine learning-based navigation approaches. Several global and local motion planning methods for ground robot navigation are compared in Table 4 and Table 5.
The A* and Dijkstra algorithms have been well-researched in the past decades, and these methods have shown their potential by extensively being applied with the Robot Operating System (ROS) in many real-world robotics applications. Combining these two path-planning methods with heuristic searching is effective in relatively low complexity 2D environments. However, these methods require heavy computational capabilities in high-dimensional environments or may struggle in unstructured and dynamic working environments.
Random sampling-based path planning algorithms generally consist of BITs and RRTs. Regionally Accelerated Batch Informed Trees (RABIT) are more widely used and perform better in high dimensional and dynamic environments [132] in comparison to graph search-based path planning algorithms. Bionic-based intelligent robot path planning methods simulate the behaviours of insects to generate evolving paths. These evolutionary methods include the ant colony, particle swarm optimisation, genetic and artificial bee colony algorithms. Many other optimised versions of these algorithms have been proposed to improve calculation efficiency and avoid local minimum problems. A welding robot system developed in [133] had a combination of genetic and particle swarm optimisation algorithms to solve for the shortest route, while avoiding obstacles. The artificial potential field method, ant colony optimisation, and geometry optimisation path methods were combined in another research study [134] to search the globally optimal path in 2D scenarios using computer simulations. This method showed fast solutions and reduced the risk of trapping robots in local minimum points. A multi-objective hierarchical particle swarm optimisation method has been proposed by [135] to plan global optimal path trajectories in cluttered environments. This method utilises three layers to generate the robot navigation trajectories. The triangular decomposition method [136] is applied in the first layer, the Dijkstra algorithm is used in the second, and a modified particle swarm optimisation algorithm is applied in the last layer.
In general, the local path planning strategies that have been discussed use the available sensor data of robots regarding their surroundings to map, understand, and generate local paths while avoiding obstacles. These local path planning methods are effective in mobile robotic applications because the data captured by sensors varies in real time in response to the dynamics of environments. In comparison to global path planning strategies, local path planning is more critical for practical robot operation and usually serves as a bridge between global planning and the direct control of robots. However, local path planners have one notable drawback in that they often lead robots to local minimum points. Many classical local path planning algorithms can generate optimal paths in the local environments while avoiding local minimum problems. These methods include techniques such as the Fuzzy Logic algorithm, artificial potential field method, and simulated annealing algorithm. In general, these methods do not evaluate the relative velocities between robots and dynamic environment objects, however, which can lead to difficulties. In many worst-case scenarios, even the velocity profiles of these obstacles can be hard to acquire for the robots. Visual-inertial odometry methods show success in outdoor environment navigation scenarios, but are not suitable for navigating in off-road conditions without pre-built maps or GPS assistance.
Machine learning methods have been applied for mobile robot navigation to learn semantic information [137,138,139] and statistical patterns [140,141] of environments. Several other research works [142,143,144,145] have used machine learning to achieve robustness in path following. In recent years, many Reinforcement Learning (RL) and imitation learning methods and approaches based on self-supervised learning [146,147,148,149] have been applied in mobile robot navigation policy design and for training supervision. Many classical modular and deep learning-based approaches have been utilised in the navigation modules of outdoor mobile robots [150]. RL uses strategies to learn optimal robot decisions from experiences. The interconnection between environments and robots is modelled as a Markov Decision Process (MDP). Robots receive rewards as feedback signals for training while traversing different environments. The basic RL process is illustrated in Figure 2. RL methods can be separated into model-free and model-based learning scenarios. In model-free RL, the robots are not required to evaluate the MDP model rewards or policies directly and can obtain these directly through what the robot experiences. Model-free RL approaches have several subcategories such as value-based, policy-based, and Actor-Critic (improved versions of the policy-based algorithms) [151,152].
In value-based methods, optimal policies are obtained by iteratively updating the value functions. Policy gradient-based methods directly approximate a policy network and update the policy parameters to get an optimal policy that maximises the reward value. Deep Q network (DQN) and Double DQN are the two main value-based DRL methods. A DQN-based end-to-end navigation method has been introduced in [153]. In this work, a feature-extracting network used an edge segmentation method to improve the efficiency of the network training process. The simulated models were transferred to the real world without significant performance loss. Discrete robot actions were implemented based on a grid map. The robot exploration framework is divided into decision, planning, and mapping modules. The learning-based decision module has shown good performance, efficiency, and adaptability in novel environments. A graph-based technique has been implemented for mapping module. The applied path planning module includes the A* algorithm as the global planner along with a timed elastic band [154] local planner.
Value-based DRL methods produce discrete actions meaning they are not appropriate for the continuous robot action space. The policy-based DRL methods provide continuous motion commands for robots. Combining the policy gradient with the value function creates the Actor-Critic RL method. In general, Actor-Critic algorithms are well-suited for the continuous motion space of autonomous mobile robots. The Actor approximates the policy used to generate actions in the environment. The Critic architecture is accountable for using a reward function to evaluate robot actions iteratively and guides the Actor in the successive iteration. The Critic uses deep learning value functions such as DQN, and Double DQN to evaluate each iteration step. The Actor-Critic approach includes learning methods like Depth Deterministic Policy Gradient algorithm (DDPG), Trust Region Policy Optimisation (TRPO), Proximal Policy Optimisation (PPO), Asynchronous Advantage Actor-Critic (A3C), and Soft Actor-Critic (SAC) [15]. An extended Actor-Critic algorithm was proposed in [155]. The visual navigation module of the deep neural network consisted of depth map prediction and semantic segmentation auxiliary tasks. The proposed learning network requires an image of the target and an observed image as inputs. This network architecture was proposed to obtain a visual navigation policy for indoor environments. A summary of DRL motion planning methods is included in Table 6.
Model-based robot RL methods generate models of external environments using supervised training and implement value functions to learn actions that maximise the return. Model-based RL has faster convergence and high sample efficiency. [156] presents a biped robot that learns on a rotating platform by combining model-based and model-free machine learning methods. This paper has addressed the overfitting problem of model-based robots. The research has simulated the robot in a 2D scenario and has shown a reduction in learning time than model-free RL algorithms. Imitation learning is another related approach which uses demonstrations by experts operating robots for specific tasks to obtain policy functions. Mobile robots learn mappings between observations of these demonstrations and appropriate robot actions. Path trajectories can be rapidly generated by manipulating learned policies from expert demonstrations. However, the robot’s capacity to practically record these demonstration results from real-world experiments can be challenging. [157] has introduced a method to acquire a mapping from actions to states using ego-centric videos collected by a human demonstrator using a mobile phone camera. The policy learning step was executed within a simulation platform, then the developed navigation policy was tested on a Clearpath Jackal wheeled robot in an indoor environment. The robot and the trained videos had view-point mismatches, but the model was robust to those changes. The system was successfully able to map camera sensor inputs to actuator commands using the developed imitation learning policy.
Inverse RL is another approach capable of learning reward functions from expert demonstrations. However, inverse RL is more appropriate for exploring novel environments because imitation learning tries to directly follow a demonstrator rather than improve beyond that knowledge level. The creation of an inverse RL network by employing a vision-based imitation learning method was presented by [157]. This method approximates a value function from one middle layer of a policy trained by imitation learning. The related experiments were carried out on a real-world setup and using the ROS Gazebo simulation platform. This proposed method has shown the capability of generalising the system for unseen environments by producing usable cost maps. Traditional geometric-based SLAM procedures lack the capacity to capture dynamic objects in external environments and are thus only suitable for planning robot actions in static environments. RL-based techniques, however, can learn policies by considering dynamic obstacles in outdoor environments, although the sample efficiency is lower compared to imitation learning and model-based learning. The implementation of learning-based maps and traditional path planners is one way to improve sample efficiency. A mobile robot affordance map generation using RL policy-based method was introduced by [158]. This research used metric cost maps (attaching semantics and geometry) for robot navigation. The A* classical path planner was implemented for path generation. This learning-based SLAM approach was implemented using a simulation software. Semantic, dynamic, and behavioural attributes of unseen environments have been learnt by the model using the simulation scheme. Local path planning remains a highly challenging task, particularly in environments with dynamic actors or challenging terrain features such as unpredictably varying topography or varying ground conditions. Modern machine learning approaches, coupled with advanced SLAM techniques show promise for enabled effective navigation planning, but methods that can effectively handle all of these challenges do not yet exist and considerable further work will be required to enable fully autonomous navigation in all conditions.
5. Summary of the Current State-of-the-Art Techniques
Autonomous navigation can be broadly categorised into the classical modular pipeline methods and end-to-end learning approaches. The modular pipeline approach has limitations due to the requirement for high levels of human intervention in designing the modules, the loss of information through each module, and the overall lack of robustness when conditions vary beyond anticipated limits. The end-to-end learning approach suffers from problems such as over-fitting and poor generalisation. Both modular and end-to-end robot navigation approaches rely on sensors to capture information about the environment or internal robot attributes. Researchers have investigated different sensor modalities to improve robot perception, including raw input types like sound, pressure, light, and magnetic fields, as well as common robot perception sensor modalities like cameras, LiDAR, radar, sonar, GNSS, IMU, and odometry sensors. It is crucial to have a reliable real-time understanding of external 3D environments to ensure safe robot navigation. While cameras are commonly used in mobile robotics, relying on a single sensor is not always robust. Different sensor modalities are often combined in robotic applications to improve perception reliability and robustness. Camera and LiDAR fusion has been a popular approach in the robotics community. This fusion method has shown better performance than other vision-based dual sensor fusion approaches, and recent advances in deep learning algorithms have further improved its performance. Monocular and stereo cameras are commonly used with LiDAR sensors to fuse images and point cloud data. However, the technical challenges and cost of sensors and processing power requirements still limit the widespread application of these methods for regular use. Computer vision has experienced rapid growth in the past decade, with deep learning methods accelerating its development. Object detection, depth estimation, semantic segmentation, instance segmentation, scene reconstruction, motion estimation, object tracking, scene understanding, and end-to-end learning are among the subtopics of computer vision. These methods have been widely applied in autonomous navigation, but their accuracy and reliability can be limited, prompting ongoing efforts to improve their quality. Benchmarking datasets, such as KITTI, Waymo, A2D2, nuScenes, Cityscapes, RELLIS-3D, RUGD, Freiburg [24], and WildDash [159], have been used to compare the performance of different vision methods in autonomous urban driving and off-road driving applications.
Semantic SLAM techniques provide geometric and semantic details of external environments. However, existing semantic SLAM techniques may not be adequate for the safe navigation of robots in outdoor unstructured environments due to challenges in feature detection, uneven terrain conditions (which can lead to robot localisation errors, detection errors, etc.) and effects of wind on vegetation, which can lead to sensor noise and ambiguous results. Therefore, terrain traversability analysis, and improved scene understanding using deep learning methods have been developed that are showing promise with assisting robot navigation in outdoor unstructured environments.
Many classical robot path-planning approaches are used for navigation in local environments. These classical approaches rely on a modular architecture to perceive the environment, plan paths relative to generated maps, and follow trajectories. However, these classical methods are often challenged in dynamic or deformable environments and are not well-suited for off-road navigation conditions where terrains may be unpredictable. Learning-based methods have been developed to improve robot path-planning abilities under different environmental conditions and have shown promise in addressing these challenges. In particular, learning-based navigation methods have been shown to be better suited for off-road navigation conditions, where terrains can be highly variable and unpredictable.
7. Conclusion
This paper has provided a comprehensive review of the current state-of-the-art in autonomous mobile ground robot navigation, identified research gaps and challenges, and suggested promising future research directions for improved autonomous navigation in outdoor unstructured terrains. A broad review of robot sensing, camera-LiDAR sensor fusion, robot scene understanding and local path planning techniques has been provided to deliver a comprehensive discussion of their essential methodologies and current capabilities and limitations. The use of deep learning, multimodal sensor fusion, incremental scene understanding concepts, scene representations that preserve input data topology and spatial geometry, and learning-based hierarchical path planning concepts are identified as promising research domains to investigate in order to realise fully autonomous navigation in unstructured outdoor environments. Our review has indicated that applying these techniques to outdoor unstructured terrain robot navigation research can likely improve robot domain adaptability, scene understanding and conscious decision-making abilities.
Author Contributions
Conceptualisation, A.R., D.C. and L.W.; writing-original, L.W.; writing-review and editing, A.R., D.C.; supervision, A.R., D.C. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Data Availability Statement
Data sharing not applicable.
Conflicts of Interest
The authors declare no conflict of interest.
References
- Rubio, F.; Valero, F.; Llopis-Albert, C. A review of mobile robots: Concepts, methods, theoretical framework, and applications. International Journal of Advanced Robotic Systems 2019, 16, 1–22. [Google Scholar] [CrossRef]
- Cai, K.; Wang, C.; Cheng, J.; De Silva, C.W.; Meng, M.Q.H. Mobile robot path planning in dynamic environments: A survey. arXiv arXiv:2006.14195 2020.
- Quarles, N.; Kockelman, K.M.; Lee, J. America’s fleet evolution in an automated future. Research in Transportation Economics 2021, 90, 1–12. [Google Scholar] [CrossRef]
- Pavel, M.I.; Tan, S.Y.; Abdullah, A. Vision-based autonomous vehicle systems based on deep learning: A systematic literature review. Applied Sciences 2022, 12. [Google Scholar] [CrossRef]
- Zhang, S.; Yao, J.; Wang, R.; Liu, Z.; Ma, C.; Wang, Y.; Zhao, Y. Design of intelligent fire-fighting robot based on multi-sensor fusion and experimental study on fire scene patrol. Robotics and Autonomous Systems 2022, 154, 1–18. [Google Scholar] [CrossRef]
- Li, Q.; Kroemer, O.; Su, Z.; Veiga, F.F.; Kaboli, M.; Ritter, H.J. A review of tactile information: Perception and action through touch. IEEE Transactions on Robotics 2020, 36, 1619–1634. [Google Scholar] [CrossRef]
- Alatise, M.B.; Hancke, G.P. A review on challenges of autonomous mobile robot and sensor fusion methods. IEEE Access 2020, 8, 39830–39846. [Google Scholar] [CrossRef]
- Feng, D.; Haase-Schütz, C.; Rosenbaum, L.; Hertlein, H.; Glaeser, C.; Timm, F.; Wiesbeck, W.; Dietmayer, K. Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges. IEEE Transactions on Intelligent Transportation Systems 2020, 22, 1341–1360. [Google Scholar] [CrossRef]
- Hu, C.; Yang, C.; Li, K.; Zhang, J. A forest point cloud real-time reconstruction method with single-line LiDAR based on visual-IMU fusion. Applied Sciences 2022, 12. [Google Scholar] [CrossRef]
- Jin, X.B.; Su, T.L.; Kong, J.L.; Bai, Y.T.; Miao, B.B.; Dou, C. State-of-the-art mobile intelligence: Enabling robots to move like humans by estimating mobility with artificial intelligence. Applied Sciences 2018, 8. [Google Scholar] [CrossRef]
- Yang, T.; Li, Y.; Zhao, C.; Yao, D.; Chen, G.; Sun, L.; Krajnik, T.; Yan, Z. 3D ToF LiDAR in mobile robotics: A review. arXiv arXiv:2202.11025 2022.
- Moon, J.; Lee, B.H. PDDL planning with natural language-based scene understanding for UAV-UGV cooperation. Applied Sciences 2019, 9. [Google Scholar] [CrossRef]
- Yang, M.; Rosenhahn, B.; Murino, V. Multimodal scene understanding: Algorithms, applications and deep learning; Academic Press: United Kingdom, 2019; pp. 1–7. [Google Scholar]
- Zhang, Y.; Sidibé, D.; Morel, O.; Mériaudeau, F. Deep multimodal fusion for semantic image segmentation: A survey. Image and Vision Computing 2021, 105, 1–17. [Google Scholar] [CrossRef]
- Sun, H.; Zhang, W.; Yu, R.; Zhang, Y. Motion planning for mobile robots—Focusing on deep reinforcement learning: A systematic review. IEEE Access 2021, 9, 69061–69081. [Google Scholar] [CrossRef]
- Janai, J.; Güney, F.; Behl, A.; Geiger, A.; others. Computer vision for autonomous vehicles: Problems, datasets and state of the art. Foundations and Trends® in Computer Graphics and Vision 2020, 12, 1–308. [Google Scholar] [CrossRef]
- Gupta, A.; Efros, A.A.; Hebert, M. Blocks world revisited: Image understanding using qualitative geometry and mechanics. In Proceedings of the 11th European Conference on Computer Vision, Heraklion, Crete, Greece, 5-11 September 2010; pp. 482–496. [Google Scholar]
- Kocić, J.; Jovičić, N.; Drndarević, V. Sensors and sensor fusion in autonomous vehicles. 26th Telecommunications Forum (TELFOR),Serbia, Belgrade, 20-, pp. 420–425. In Proceedings of the 26th Telecommunications Forum (TELFOR), Serbia, Belgrade, 20-21 November 2018; pp. 420–425. [Google Scholar]
- Muñoz-Bañón, M.Á.; Candelas, F.A.; Torres, F. Targetless camera-LiDAR calibration in unstructured environments. IEEE Access 2020, 8, 143692–143705. [Google Scholar] [CrossRef]
- Li, A.; Cao, J.; Li, S.; Huang, Z.; Wang, J.; Liu, G. Map construction and path planning method for a mobile robot based on multi-sensor information fusion. Applied Sciences 2022, 12. [Google Scholar] [CrossRef]
- Wang, W.; Shen, J.; Cheng, M.M.; Shao, L. An iterative and cooperative top-down and bottom-up inference network for salient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); pp. 5968–5977.
- Santos, L.C.; Santos, F.N.; Pires, E.S.; Valente, A.; Costa, P.; Magalhães, S. Path planning for ground robots in agriculture: A short review. IEEE International Conference on Autonomous Robot Systems and Competitions (ICARSC), Ponta Delgada, Portugal, 15-16 April 2020; pp. 61–66. [Google Scholar]
- Fayyad, J.; Jaradat, M.A.; Gruyer, D.; Najjaran, H. Deep learning sensor fusion for autonomous vehicle perception and localization: A review. Sensors 2020, 20, 1–35. [Google Scholar] [CrossRef]
- Valada, A.; Oliveira, G.L.; Brox, T.; Burgard, W. Deep multispectral semantic scene understanding of forested environments using multimodal fusion. International Symposium on Experimental Robotics, Tokyo, Japan, 3-6 October 2016; pp. 465–477. [Google Scholar]
- Lei, X.; Zhang, Z.; Dong, P. Dynamic path planning of unknown environment based on deep reinforcement learning. Journal of Robotics 2018, 2018, 1–10. [Google Scholar] [CrossRef]
- Crespo, J.; Castillo, J.C.; Mozos, O.M.; Barber, R. Semantic information for robot navigation: A survey. Applied Sciences 2020, 10, 1–28. [Google Scholar] [CrossRef]
- Galvao, L.G.; Abbod, M.; Kalganova, T.; Palade, V.; Huda, M.N. Pedestrian and vehicle detection in autonomous vehicle perception systems—A review. Sensors 2021, 21, 1–47. [Google Scholar] [CrossRef] [PubMed]
- Hewawasam, H.; Ibrahim, M.Y.; Appuhamillage, G.K. Past, present and future of path-planning algorithms for mobile robot navigation in dynamic environments. IEEE Open Journal of the Industrial Electronics Society 2022, 3, 353–365. [Google Scholar] [CrossRef]
- Martini, M.; Cerrato, S.; Salvetti, F.; Angarano, S.; Chiaberge, M. Position-Agnostic Autonomous Navigation in Vineyards with Deep Reinforcement Learning. IEEE 18th International Conference on Automation Science and Engineering (CASE), 20-24 August 2022, pp. 477–484.
- Huang, Z.; Lv, C.; Xing, Y.; Wu, J. Multi-modal sensor fusion-based deep neural network for end-to-end autonomous driving with scene understanding. IEEE Sensors Journal 2020, 21, 11781–11790. [Google Scholar] [CrossRef]
- Hamza, A. Deep reinforcement learning for mapless mobile robot navigation. Master’s thesis, Luleå University of Technology, Sweden, 2022.
- Carrasco, P.; Cuesta, F.; Caballero, R.; Perez-Grau, F.J.; Viguria, A. Multi-sensor fusion for aerial robots in industrial GNSS-denied environments. Applied Sciences 2021, 11. [Google Scholar] [CrossRef]
- Li, R.; Wang, S.; Gu, D. DeepSLAM: A robust monocular SLAM system with unsupervised deep learning. IEEE Transactions on Industrial Electronics 2020, 68, 3577–3587. [Google Scholar] [CrossRef]
- Aguiar, A.; Santos, F.; Sousa, A.J.; Santos, L. FAST-FUSION: An improved accuracy omnidirectional visual odometry system with sensor fusion and GPU optimization for embedded low cost hardware. Applied Sciences 2019, 9. [Google Scholar] [CrossRef]
- Li, Y.; Brasch, N.; Wang, Y.; Navab, N.; Tombari, F. Structure-slam: Low-drift monocular slam in indoor environments. IEEE Robotics and Automation Letters 2020, 5, 6583–6590. [Google Scholar] [CrossRef]
- Zaffar, M.; Ehsan, S.; Stolkin, R.; Maier, K.M. Sensors, SLAM and long-term autonomy: A review. NASA/ESA Conference on Adaptive Hardware and Systems (AHS), United Kingdom, 6-9 August 2018, pp. 285–290.
- Sabattini, L.; Levratti, A.; Venturi, F.; Amplo, E.; Fantuzzi, C.; Secchi, C. Experimental comparison of 3D vision sensors for mobile robot localization for industrial application: Stereo-camera and RGB-D sensor. 12th International Conference on Control Automation Robotics & Vision (ICARCV), Guangzhou, China, 5-7 December 2012, pp. 823–828.
- Tölgyessy, M.; Dekan, M.; Chovanec, L.; Hubinskỳ, P. Evaluation of the azure kinect and its comparison to kinect v1 and kinect v2. Sensors 2021, 21, 1–23. [Google Scholar] [CrossRef]
- Evangelidis, G.D.; Hansard, M.; Horaud, R. Fusion of range and stereo data for high-resolution scene-modeling. IEEE Transactions on Pattern Analysis and Machine Intelligence 2015, 37, 2178–2192. [Google Scholar] [CrossRef]
- Glover, A.; Bartolozzi, C. Robust visual tracking with a freely-moving event camera. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada, 24–28 September 2017, pp. 3769–3776.
- Gallego, G.; Delbrück, T.; Orchard, G.; Bartolozzi, C.; Taba, B.; Censi, A.; Leutenegger, S.; Davison, A.J.; Conradt, J.; Daniilidis, K.; others. Event-based vision: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 2020, 44, 154–180. [Google Scholar] [CrossRef]
- Yuan, W.; Li, J.; Bhatta, M.; Shi, Y.; Baenziger, P.S.; Ge, Y. Wheat height estimation using LiDAR in comparison to ultrasonic sensor and UAS. Sensors 2018, 18, 1–20. [Google Scholar] [CrossRef] [PubMed]
- Moosmann, F.; Stiller, C. Velodyne slam. IEEE Intelligent Vehicles Symposium,Baden-Baden, Germany, 5-9 June 2011, pp. 393–398.
- Li, K.; Li, M.; Hanebeck, U.D. Towards high-performance solid-state-lidar-inertial odometry and mapping. IEEE Robotics and Automation Letters 2021, 6, 5167–5174. [Google Scholar] [CrossRef]
- Poulton, C.V.; Yaacobi, A.; Cole, D.B.; Byrd, M.J.; Raval, M.; Vermeulen, D.; Watts, M.R. Coherent solid-state LIDAR with silicon photonic optical phased arrays. Optics letters 2017, 42, 4091–4094. [Google Scholar] [CrossRef] [PubMed]
- Behroozpour, B.; Sandborn, P.A.; Wu, M.C.; Boser, B.E. Lidar system architectures and circuits. IEEE Communications Magazine 2017, 55, 135–142. [Google Scholar] [CrossRef]
- Xu, X.; Zhang, L.; Yang, J.; Cao, C.; Wang, W.; Ran, Y.; Tan, Z.; Luo, M. A review of multi-sensor fusion slam systems based on 3D LIDAR. Remote Sensing 2022, 14, 1–27. [Google Scholar] [CrossRef]
- Li, Y.; Yu, A.W.; Meng, T.; Caine, B.; Ngiam, J.; Peng, D.; Shen, J.; Lu, Y.; Zhou, D.; Le, Q.V. others. Deepfusion: LiDAR-camera deep fusion for multi-modal 3D object detection. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, Louisiana, USA, 18-24 June 2022, pp. 17182–17191.
- Zheng, W.; Xie, H.; Chen, Y.; Roh, J.; Shin, H. PIFNet: 3D object detection using joint image and point cloud features for autonomous driving. Applied Sciences 2022, 12. [Google Scholar] [CrossRef]
- Cui, Y.; Chen, R.; Chu, W.; Chen, L.; Tian, D.; Li, Y.; Cao, D. Deep learning for image and point cloud fusion in autonomous driving: A review. IEEE Transactions on Intelligent Transportation Systems 2021, 23, 722–739. [Google Scholar] [CrossRef]
- Du, X.; Ang, M.H.; Karaman, S.; Rus, D. A general pipeline for 3D detection of vehicles. 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, Australia, 21-25 May 2018, pp. 3194–3200.
- Yang, Z.; Sun, Y.; Liu, S.; Shen, X.; Jia, J. Ipod: Intensive point-based object detector for point cloud. arXiv arXiv:1812.05276 2018.
- Qi, C.R.; Liu, W.; Wu, C.; Su, H.; Guibas, L.J. Frustum pointnets for 3D object detection from RGB-D data. IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18-23 June 2018, pp. 918–927.
- Shin, K.; Kwon, Y.P.; Tomizuka, M. Roarnet: A robust 3D object detection based on region approximation refinement. IEEE intelligent vehicles symposium (IV), Paris, France, 9-12 June 2019, pp. 2510–2515.
- Qi, C.R.; Su, H.; Mo, K.; Guibas, L.J. Pointnet: Deep learning on point sets for 3D classification and segmentation. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 21-26 July 2017, pp. 652–660.
- Tran, D.; Bourdev, L.; Fergus, R.; Torresani, L.; Paluri, M. Learning spatiotemporal features with 3D convolutional networks. IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 2015, pp. 4489–4497.
- Maturana, D.; Scherer, S. Voxnet: A 3D convolutional neural network for real-time object recognition. IEEE/RSJ international conference on intelligent robots and systems (IROS), Hamburg, Germany, 28 Sept-2 Oct 2015, pp. 922–928.
- Xu, D.; Anguelov, D.; Jain, A. Pointfusion: Deep sensor fusion for 3D bounding box estimation. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 18-23 June 2018, pp. 244–253.
- Ku, J.; Mozifian, M.; Lee, J.; Harakeh, A.; Waslander, S.L. Joint 3D proposal generation and object detection from view aggregation. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS),Madrid, Spain, 1-5 October 2018, pp. 1–8.
- Liang, M.; Yang, B.; Wang, S.; Urtasun, R. Deep continuous fusion for multi-sensor 3D object detection. 15th European conference on computer vision (ECCV), Munich, Germany, 8-14 September 2018, pp. 641–656.
- Sindagi, V.A.; Zhou, Y.; Tuzel, O. Mvx-net: Multimodal voxelnet for 3D object detection. International Conference on Robotics and Automation (ICRA), Montreal, Canada, 20-24 May 2019, pp. 7276–7282.
- Chen, X.; Ma, H.; Wan, J.; Li, B.; Xia, T. Multi-view 3D object detection network for autonomous driving. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 21-26 July 2017, pp. 1907–1915.
- Geiger, A.; Lenz, P.; Stiller, C.; Urtasun, R. Vision meets robotics: The kitti dataset. The International Journal of Robotics Research 2013, 32, 1231–1237. [Google Scholar] [CrossRef]
- Sun, P.; Kretzschmar, H.; Dotiwalla, X.; Chouard, A.; Patnaik, V.; Tsui, P.; Guo, J.; Zhou, Y.; Chai, Y.; Caine, B. others. Scalability in perception for autonomous driving: Waymo open dataset. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13-19 June 2020, pp. 2446–2454.
- Geyer, J.; Kassahun, Y.; Mahmudi, M.; Ricou, X.; Durgesh, R.; Chung, A.S.; Hauswald, L.; Pham, V.H.; Mühlegg, M.; Dorn, S. ; others. A2d2: Audi autonomous driving dataset. arXiv arXiv:2004.06320 2020.
- Caesar, H.; Bankiti, V.; Lang, A.H.; Vora, S.; Liong, V.E.; Xu, Q.; Krishnan, A.; Pan, Y.; Baldan, G.; Beijbom, O. nuscenes: A multimodal dataset for autonomous driving. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13-19 June 2020, pp. 11621–11631.
- Cordts, M.; Omran, M.; Ramos, S.; Rehfeld, T.; Enzweiler, M.; Benenson, R.; Franke, U.; Roth, S.; Schiele, B. The cityscapes dataset for semantic urban scene understanding. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27-30 June 2016, pp. 3213–3223.
- Ma, F.; Cavalheiro, G.V.; Karaman, S. Self-supervised sparse-to-dense: Self-supervised depth completion from lidar and monocular camera. International Conference on Robotics and Automation (ICRA), Montreal, Canada, 20-24 May 2019, pp. 3288–3295.
- Ma, F.; Karaman, S. Sparse-to-dense: Depth prediction from sparse depth samples and a single image. IEEE international conference on robotics and automation (ICRA), Brisbane, Australia, 21-25 May 2018, pp. 4796–4803.
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27-30 June 2016, pp. 770–778.
- Cheng, X.; Wang, P.; Yang, R. Depth estimation via affinity learned with convolutional spatial propagation network. European Conference on Computer Vision (ECCV), Munich, Germany, 8-14 September 2018, pp. 103–119.
- Cheng, X.; Wang, P.; Guan, C.; Yang, R. CSPN++: Learning context and resource aware convolutional spatial propagation networks for depth completion. 34th AAAI Conference on Artificial Intelligence, New York, USA, 7-12 February 2020, pp. 10615–10622.
- Cheng, X.; Zhong, Y.; Dai, Y.; Ji, P.; Li, H. Noise-aware unsupervised deep LiDAR-stereo fusion. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15-20 June 2019, pp. 6339–6348.
- Jalal, A.S.; Singh, V. The state-of-the-art in visual object tracking. Informatica 2012, 36, 1–22. [Google Scholar]
- Tang, P.; Wang, X.; Wang, A.; Yan, Y.; Liu, W.; Huang, J.; Yuille, A. Weakly supervised region proposal network and object detection. 15th European conference on computer vision (ECCV), Munich, Germany, 8-14 September 2018, pp. 352–368.
- Uijlings, J.R.; Van De Sande, K.E.; Gevers, T.; Smeulders, A.W. Selective search for object recognition. International Journal of Computer Vision 2013, 104, 154–171. [Google Scholar] [CrossRef]
- Hong, M.; Li, S.; Yang, Y.; Zhu, F.; Zhao, Q.; Lu, L. SSPNet: Scale selection pyramid network for tiny person detection from UAV images. IEEE Geoscience and Remote Sensing Letters 2021, 19, 1–5. [Google Scholar] [CrossRef]
- Girshick, R. Fast R-CNN. IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 7-13 December 2015, pp. 1440–1448.
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis and Machine Intelligence 2015, 39, 1–9. [Google Scholar] [CrossRef] [PubMed]
- Kim, J.; Cho, J. Exploring a multimodal mixture-of-YOLOs framework for advanced real-time object detection. Applied Sciences 2020, 10. [Google Scholar] [CrossRef]
- Gupta, S.; Girshick, R.; Arbeláez, P.; Malik, J. Learning rich features from RGB-D images for object detection and segmentation. 13th European conference on computer vision (ECCV), Zurich, Switzerland, 6-12 September 2014, pp. 345–360.
- Simonyan, K.; Zisserman, A. Very deep convolutional networks for large-scale image recognition. arXiv arXiv:1409.1556 2014.
- Zhou, Y.; Tuzel, O. Voxelnet: End-to-end learning for point cloud based 3D object detection. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA,, 18-23 June 2018, pp. 4490–4499.
- Meyer, G.P.; Charland, J.; Hegde, D.; Laddha, A.; Vallespi-Gonzalez, C. Sensor fusion for joint 3D object detection and semantic segmentation. IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 16-17 June 2019, pp. 1–8.
- Meyer, G.P.; Laddha, A.; Kee, E.; Vallespi-Gonzalez, C.; Wellington, C.K. Lasernet: An efficient probabilistic 3D object detector for autonomous driving. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, California, 15-20 June 2019, pp. 12677–12686.
- Guo, Z.; Huang, Y.; Hu, X.; Wei, H.; Zhao, B. A survey on deep learning based approaches for scene understanding in autonomous driving. Electronics 2021, 10, 1–29. [Google Scholar] [CrossRef]
- Zoph, B.; Vasudevan, V.; Shlens, J.; Le, Q.V. Learning transferable architectures for scalable image recognition. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18-23 June 2018, pp. 8697–8710.
- Valada, A.; Mohan, R.; Burgard, W. Self-supervised model adaptation for multimodal semantic segmentation. International Journal of Computer Vision 2020, 128, 1239–1285. [Google Scholar] [CrossRef]
- Caltagirone, L.; Bellone, M.; Svensson, L.; Wahde, M. LIDAR–camera fusion for road detection using fully convolutional neural networks. Robotics and Autonomous Systems 2019, 111, 125–131. [Google Scholar] [CrossRef]
- Dai, A.; Nießner, M. 3DMV: Joint 3D-multi-view prediction for 3D semantic scene segmentation. European Conference on Computer Vision (ECCV), Munich, Germany, 8-14 September 2018, pp. 452–468.
- Chiang, H.Y.; Lin, Y.L.; Liu, Y.C.; Hsu, W.H. A unified point-based framework for 3D segmentation. International Conference on 3D Vision (3DV), Québec, Canada,, 16-19 September 2019, pp. 155–163.
- Qi, C.R.; Yi, L.; Su, H.; Guibas, L.J. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in Neural Information Processing Systems 2017, 30, 1–10. [Google Scholar]
- Jaritz, M.; Gu, J.; Su, H. Multi-view pointnet for 3D scene understanding. IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), Seoul, Korea (South), 27-28 October 2019, pp. 1–9.
- Su, H.; Jampani, V.; Sun, D.; Maji, S.; Kalogerakis, E.; Yang, M.H.; Kautz, J. Splatnet: Sparse lattice networks for point cloud processing. IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18-23 June 2018, pp. 2530–2539.
- Hou, J.; Dai, A.; Nießner, M. 3D-SIS: 3D semantic instance segmentation of RGB-D scans. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15-20 June 2019, pp. 4421–4430.
- Narita, G.; Seno, T.; Ishikawa, T.; Kaji, Y. Panopticfusion: Online volumetric semantic mapping at the level of stuff and things. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Macau, China, 4-8 November 2019, pp. 4205–4212.
- Elich, C.; Engelmann, F.; Kontogianni, T.; Leibe, B. 3D bird’s-eye-view instance segmentation. 41st DAGM German Conference on Pattern Recognition, Dortmund, Germany, 10-13 September 2019, pp. 48–61.
- Comaniciu, D.; Meer, P. Mean shift: A robust approach toward feature space analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence 2002, 24, 603–619. [Google Scholar] [CrossRef]
- Kochanov, D.; Ošep, A.; Stückler, J.; Leibe, B. Scene flow propagation for semantic mapping and object discovery in dynamic street scenes. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Daejeon, Korea, 9-14 October 2016, pp. 1785–1792.
- Yue, Y.; Zhao, C.; Li, R.; Yang, C.; Zhang, J.; Wen, M.; Wang, Y.; Wang, D. A hierarchical framework for collaborative probabilistic semantic mapping. IEEE international conference on robotics and automation (ICRA), Paris, France, 31 May - 31 August 2020, pp. 9659–9665.
- Rosinol, A.; Violette, A.; Abate, M.; Hughes, N.; Chang, Y.; Shi, J.; Gupta, A.; Carlone, L. Kimera: From SLAM to spatial perception with 3D dynamic scene graphs. The International Journal of Robotics Research 2021, 40, 1510–1546. [Google Scholar] [CrossRef]
- Tian, Y.; Chang, Y.; Arias, F.H.; Nieto-Granda, C.; How, J.P.; Carlone, L. Kimera-multi: Robust, distributed, dense metric-semantic slam for multi-robot systems. IEEE Transactions on Robotics 2022, 38, 2022–2038. [Google Scholar] [CrossRef]
- Kim, U.H.; Park, J.M.; Song, T.J.; Kim, J.H. 3-D scene graph: A sparse and semantic representation of physical environments for intelligent agents. IEEE Transactions on Cybernetics 2019, 50, 4921–4933. [Google Scholar] [CrossRef] [PubMed]
- Rosinol, A.; Gupta, A.; Abate, M.; Shi, J.; Carlone, L. 3D dynamic scene graphs: Actionable spatial perception with places, objects, and humans. arXiv arXiv:2002.06289 2020.
- Liu, H.; Yao, M.; Xiao, X.; Cui, H. A hybrid attention semantic segmentation network for unstructured terrain on Mars. Acta Astronautica 2023, 204, 492–499. [Google Scholar] [CrossRef]
- Humblot-Renaux, G.; Marchegiani, L.; Moeslund, T.B.; Gade, R. Navigation-oriented scene understanding for robotic autonomy: learning to segment driveability in egocentric images. IEEE Robotics and Automation Letters 2022, 7, 2913–2920. [Google Scholar] [CrossRef]
- Badrinarayanan, V.; Kendall, A.; Cipolla, R. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence 2017, 39, 2481–2495. [Google Scholar] [CrossRef]
- Guan, T.; Kothandaraman, D.; Chandra, R.; Sathyamoorthy, A.J.; Weerakoon, K.; Manocha, D. GA-Nav: Efficient terrain segmentation for robot navigation in unstructured outdoor environments. IEEE Robotics and Automation Letters 2022, 7, 8138–8145. [Google Scholar] [CrossRef]
- Wigness, M.; Eum, S.; Rogers, J.G.; Han, D.; Kwon, H. A RUGD dataset for autonomous navigation and visual perception in unstructured outdoor environments. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Macau, China, 4-8 November 2019, pp. 5000–5007.
- Jiang, P.; Osteen, P.; Wigness, M.; Saripalli, S. RELLIS-3D dataset: Data, benchmarks and analysis. IEEE international conference on robotics and automation (ICRA), Xi’an, China, May 31 - June 4 2021, pp. 1110–1116.
- Ma, L.; Stückler, J.; Kerl, C.; Cremers, D. Multi-view deep learning for consistent semantic mapping with RGB-D cameras. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada, 24-28 September 2017, pp. 598–605.
- Hazirbas, C.; Ma, L.; Domokos, C.; Cremers, D. Fusenet: Incorporating depth into semantic segmentation via fusion-based cnn architecture. 13th Asian Conference on Computer Vision, Taipei, Taiwan, 20-24 November 2016, pp. 213–228.
- Zhang, J.; Henein, M.; Mahony, R.; Ila, V. VDO-SLAM: A visual dynamic object-aware SLAM system. arXiv arXiv:2005.11052 2020.
- Maturana, D.; Scherer, S. Voxnet: A 3D convolutional neural network for real-time object recognition. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Hamburg, Germany, Sept 28 - Oct 2 2015, pp. 922–928.
- Huang, J.; You, S. Point cloud labeling using 3D convolutional neural network. 23rd International Conference on Pattern Recognition (ICPR), Cancun, Mexico, 4-8 December 2016, pp. 2670–2675.
- LeCun, Y.; Bottou, L.; Bengio, Y.; Haffner, P. Gradient-based learning applied to document recognition. Proceedings of the IEEE 1998, 86, 2278–2324. [Google Scholar] [CrossRef]
- Song, S.; Yu, F.; Zeng, A.; Chang, A.X.; Savva, M.; Funkhouser, T. Semantic scene completion from a single depth image. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, Hawaii, 21-26 July 2017, pp. 1746–1754.
- Riegler, G.; Osman Ulusoy, A.; Geiger, A. OctNet: Learning deep 3D representations at high resolutions. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, Hawaii, 21-26 July 2017, pp. 3577–3586.
- Tatarchenko, M.; Park, J.; Koltun, V.; Zhou, Q.Y. Tangent convolutions for dense prediction in 3D. IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, Utah, USA, 18-23 June 2018, pp. 3887–3896.
- Wang, F.; Yang, Y.; Wu, Z.; Zhou, J.; Zhang, W. Real-time semantic segmentation of point clouds based on an attention mechanism and a sparse tensor. Applied Sciences 2023, 13. [Google Scholar] [CrossRef]
- Wu, W.; Qi, Z.; Fuxin, L. Pointconv: Deep convolutional networks on 3D point clouds. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15-20 June 2019, pp. 9621–9630.
- Hua, B.S.; Tran, M.K.; Yeung, S.K. Pointwise convolutional neural networks. IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18-23 June 2018, pp. 984–993.
- Zamorski, M.; Zięba, M.; Klukowski, P.; Nowak, R.; Kurach, K.; Stokowiec, W.; Trzciński, T. Adversarial autoencoders for compact representations of 3D point clouds. Computer Vision and Image Understanding 2020, 193, 1–8. [Google Scholar] [CrossRef]
- Ye, X.; Li, J.; Huang, H.; Du, L.; Zhang, X. 3D recurrent neural networks with context fusion for point cloud semantic segmentation. 15th European Conference on Computer Vision (ECCV), Munich, Germany, 8-14 September 2018, pp. 403–417.
- Wang, Y.; Sun, Y.; Liu, Z.; Sarma, S.E.; Bronstein, M.M.; Solomon, J.M. Dynamic graph cnn for learning on point clouds. ACM Transactions on Graphics 2019, 38, 1–12. [Google Scholar] [CrossRef]
- Qi, X.; Liao, R.; Jia, J.; Fidler, S.; Urtasun, R. 3D graph neural networks for RGB-D semantic segmentation. IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22-29 October 2017, pp. 5199–5208.
- Landrieu, L.; Simonovsky, M. Large-scale point cloud semantic segmentation with superpoint graphs. IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18-23 June 2018, pp. 4558–4567.
- Li, J.; Chen, B.M.; Lee, G.H. SO-Net: Self-organizing network for point cloud analysis. IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18-23 June 2018, pp. 9397–9406.
- Thrun, S. Probabilistic robotics. Communications of the ACM 2002, 45, 52–57. [Google Scholar] [CrossRef]
- Siegwart, R.; Nourbakhsh, I.R.; Scaramuzza, D. Introduction to autonomous mobile robots, 2nd ed.; MIT press, 2011.
- Sivakumar, A.N.; Modi, S.; Gasparino, M.V.; Ellis, C.; Velasquez, A.E.B.; Chowdhary, G.; Gupta, S. Learned visual navigation for under-canopy agricultural robots. 17th Robotics: Science and Systems, 12-16 July 2021.
- Atas, F.; Grimstad, L.; Cielniak, G. Evaluation of sampling-based optimizing planners for outdoor robot navigation. arXiv arXiv:2103.13666 2021.
- Wang, X.; Shi, Y.; Ding, D.; Gu, X. Double global optimum genetic algorithm–particle swarm optimization-based welding robot path planning. Engineering Optimization 2016, 48, 299–316. [Google Scholar] [CrossRef]
- Zhu, S.; Zhu, W.; Zhang, X.; Cao, T. Path planning of lunar robot based on dynamic adaptive ant colony algorithm and obstacle avoidance. International Journal of Advanced Robotic Systems 2020, 17, 1–14. [Google Scholar] [CrossRef]
- Mac, T.T.; Copot, C.; Tran, D.T.; De Keyser, R. A hierarchical global path planning approach for mobile robots based on multi-objective particle swarm optimization. Applied Soft Computing 2017, 59, 68–76. [Google Scholar] [CrossRef]
- Ghita, N.; Kloetzer, M. Trajectory planning for a car-like robot by environment abstraction. Robotics and Autonomous Systems 2012, 60, 609–619. [Google Scholar] [CrossRef]
- Zhu, Y.; Mottaghi, R.; Kolve, E.; Lim, J.J.; Gupta, A.; Fei-Fei, L.; Farhadi, A. Target-driven visual navigation in indoor scenes using deep reinforcement learning. IEEE international conference on robotics and automation (ICRA), Singapore, 29 May - 3 June 2017, pp. 3357–3364.
- Wijmans, E.; Kadian, A.; Morcos, A.; Lee, S.; Essa, I.; Parikh, D.; Savva, M.; Batra, D. Dd-ppo: Learning near-perfect pointgoal navigators from 2.5 billion frames. arXiv arXiv:1911.00357 2019.
- Gupta, S.; Davidson, J.; Levine, S.; Sukthankar, R.; Malik, J. Cognitive mapping and planning for visual navigation. International Journal of Computer Vision 2020, 128, 1311–1330. [Google Scholar] [CrossRef]
- Datta, S.; Maksymets, O.; Hoffman, J.; Lee, S.; Batra, D.; Parikh, D. Integrating egocentric localization for more realistic point-goal navigation agents. 4th Conference on Robot Learning (CoRL), 16 - 18 November 2020, pp. 313–328.
- Kumar, A.; Gupta, S.; Fouhey, D.; Levine, S.; Malik, J. Visual memory for robust path following. Advances in Neural Information Processing Systems 2018, 31, 1–10. [Google Scholar]
- Pan, Y.; Cheng, C.A.; Saigol, K.; Lee, K.; Yan, X.; Theodorou, E.; Boots, B. Agile autonomous driving using end-to-end deep imitation learning. arXiv arXiv:1709.07174 2017.
- Sadeghi, F.; Levine, S. CAD2RL: Real single-image flight without a single real image. Robotics: Science and Systems, Cambridge, Massachusetts, USA, 12-16 July 2017.
- Ross, S.; Melik-Barkhudarov, N.; Shankar, K.S.; Wendel, A.; Dey, D.; Bagnell, J.A.; Hebert, M. Learning monocular reactive uav control in cluttered natural environments. IEEE International Conference on Robotics and Automation (ICRA), Karlsruhe, Germany, 6-10 May 2013, pp. 1765–1772.
- Gandhi, D.; Pinto, L.; Gupta, A. Learning to fly by crashing. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada, 24-28 September 2017, pp. 3948–3955.
- Gasparino, M.V.; Sivakumar, A.N.; Liu, Y.; Velasquez, A.E.; Higuti, V.A.; Rogers, J.; Tran, H.; Chowdhary, G. Wayfast: Navigation with predictive traversability in the field. IEEE Robotics and Automation Letters 2022, 7, 10651–10658. [Google Scholar] [CrossRef]
- Sathyamoorthy, A.J.; Weerakoon, K.; Guan, T.; Liang, J.; Manocha, D. TerraPN: Unstructured terrain navigation using online self-supervised learning. IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Kyoto, Japan, 23-27 October 2022, pp. 7197–7204.
- Hirose, N.; Shah, D.; Sridhar, A.; Levine, S. ExAug: Robot-conditioned navigation policies via geometric experience augmentation. arXiv arXiv:2210.07450 2022.
- Heimann, D.; Hohenfeld, H.; Wiebe, F.; Kirchner, F. Quantum deep reinforcement learning for robot navigation tasks. arXiv arXiv:2202.12180 2022.
- Gyagenda, N.; Hatilima, J.V.; Roth, H.; Zhmud, V. A review of GNSS-independent UAV navigation techniques. Robotics and Autonomous Systems 2022, 152, 1–17. [Google Scholar] [CrossRef]
- Zhu, K.; Zhang, T. Deep reinforcement learning based mobile robot navigation: A review. Tsinghua Science and Technology 2021, 26, 674–691. [Google Scholar] [CrossRef]
- Li, Y. Deep reinforcement learning: An overview. arXiv arXiv:1701.07274 2017.
- Li, H.; Zhang, Q.; Zhao, D. Deep reinforcement learning-based automatic exploration for navigation in unknown environment. IEEE transactions on Neural Networks and Learning Systems 2019, 31, 2064–2076. [Google Scholar] [CrossRef] [PubMed]
- Wu, J.; Ma, X.; Peng, T.; Wang, H. An improved timed elastic band (TEB) algorithm of autonomous ground vehicle (AGV) in complex environment. Sensors 2021, 21, 1–12. [Google Scholar] [CrossRef]
- Kulhánek, J.; Derner, E.; De Bruin, T.; Babuška, R. Vision-based navigation using deep reinforcement learning. European Conference on Mobile Robots (ECMR), Prague, Czech Republic, 4-6 September 2019, pp. 1–8.
- Xi, A.; Mudiyanselage, T.W.; Tao, D.; Chen, C. Balance control of a biped robot on a rotating platform based on efficient reinforcement learning. IEEE/CAA Journal of Automatica Sinica 2019, 6, 938–951. [Google Scholar] [CrossRef]
- Lee, K.; Vlahov, B.; Gibson, J.; Rehg, J.M.; Theodorou, E.A. Approximate inverse reinforcement learning from vision-based imitation learning. IEEE International Conference on Robotics and Automation (ICRA), Xi’an, China, 31 May - 4 June 2021, pp. 10793–10799.
- Qi, W.; Mullapudi, R.T.; Gupta, S.; Ramanan, D. Learning to move with affordance maps. arXiv arXiv:2001.02364 2020.
- Zendel, O.; Honauer, K.; Murschitz, M.; Steininger, D.; Dominguez, G.F. Wilddash-creating hazard-aware benchmarks. 15th European Conference on Computer Vision (ECCV), Munich, Germany, 8-14 September 2018, pp. 402–416.
- Tang, J.; Chen, Y.; Kukko, A.; Kaartinen, H.; Jaakkola, A.; Khoramshahi, E.; Hakala, T.; Hyyppä, J.; Holopainen, M.; Hyyppä, H. SLAM-aided stem mapping for forest inventory with small-footprint mobile LiDAR. Forests 2015, 6, 4588–4606. [Google Scholar] [CrossRef]
- Chen, W.; Shang, G.; Ji, A.; Zhou, C.; Wang, X.; Xu, C.; Li, Z.; Hu, K. An overview on visual SLAM: From tradition to semantic. Remote Sensing 2022, 14, 1–47. [Google Scholar] [CrossRef]
- Chghaf, M.; Rodriguez, S.; Ouardi, A.E. Camera, LiDAR and multi-modal SLAM systems for autonomous ground vehicles: A survey. Journal of Intelligent & Robotic Systems 2022, 105, 1–35. [Google Scholar]
- Xue, H.; Hein, B.; Bakr, M.; Schildbach, G.; Abel, B.; Rueckert, E. Using deep reinforcement learning with automatic curriculum learning for mapless navigation in intralogistics. Applied Sciences 2022, 12. [Google Scholar] [CrossRef]
Figure 1.
Classical and end-to-end autonomous navigation approaches.

Figure 2.
Robot RL approach.

Table 1.
Depth sensor modalities.
| Technique | Typical Sensors | Advantages | Disadvantages |
|---|---|---|---|
| Structured light | Kinect v1, Xtion PROLive, RealSense SR300 and F200 | High accuracy and precision in indoor environments | Limited range, not suitable for outdoor environment due to noise from ambient light, interference from reflections and other light sources |
| ToF | Kinect v2 | Good for indoor outdoor applications, long range, robust to illumination changes | Lower image resolution than structured light cameras, high power consumption, cost varies with resolution, rain fog can affect sensor performance |
| Active infrared stereo | RealSense R200, RealSense D435, D435i | Compact, lightweight, dense depth images | Stereo matching requires high processing power, struggle at high occlusions and featureless environments, relatively low range especially outdoors |
Table 2.
Camera configurations.
| Configuration | Advantages | Disadvantages |
|---|---|---|
| Monocular | Compactness, low hardware requirements | No direct depth measurements |
| Stereo | Depth measurements, low occlusions | Fails in featureless environments, CPU intensive, accuracy/range depends on camera quality |
| RGB-D | Color and depth information per pixel | Limited range, reflection problems on transparent, shiny, or very matte and absorbing objects |
| Event | High temporal resolution, suitable for changing light intensities, low latency [41] | No direct depth information, costly, not suitable for static scenes, requires non-traditional algorithms |
| Omni-directional | Wide angle view (alternative to rotating cameras) | Lower resolution, need special methods to compensate for image distortions |
Table 3.
LiDAR sensor types.
| Configuration | Advantages | Disadvantages |
|---|---|---|
| Pulsed | High frame rate | Low depth resolution, higher inference from other LiDAR sensors |
| AMCW | Not limited by low SNRs, however, not effective at very low SNRs | Low accuracy than FMCW, lower depth resolution than FMCW |
| FMCW | Velocity and range detection in a single shot, higher accuracy than AMCW, higher depth resolution | Currently at the research and development stage |
Table 4.
Global path planning algorithms.
| Algorithms | Advantages | Disadvantages |
|---|---|---|
| Dijkstra | The calculation strategy is not complex and give the shortest path | The increment of traversal nodes complicates the calculations |
| A* | In static environments the algorithm search efficiency is high | Not appropriate for dynamic environments |
| D* | Good for dynamic environment path planning and more efficient than A* | Planning longer paths via D* creates challenges |
| RRT | Fast convergence, high search capability | Algorithm efficiency is low in unstructured environments |
| Genetic | Appropriate for complex environments, good for finding optimal paths | Low algorithm convergence speed, low search ability in local paths |
| Ant colony | Appropriate for complex environments, can be combined with other heuristic-based path planners | Slow convergence rate, easily trapped in local minima |
| Particle swarm optimisation | High convergence rate, good robustness | Frequently, solutions converge into local optimal solutions |
Table 5.
Local path planning algorithms.
| Algorithms | Advantages | Disadvantages |
|---|---|---|
| Artificial potential field | Can be implemented for 3D path planning, and can solve the local minimum problem | Cannot guarantee the optimal solution |
| Simulated annealing | Flexible and easy implementation, can deal with noisy data and non-linear models | Can produce unstable results, the trade-off between accuracy and speed |
| Fuzzy logic | Strong robustness, decrease the dependencies between environmental data | Needs accurate prior knowledge, poor learning capabilities |
| Neural network | Strong robustness, and learning ability from experiences | Low path planning efficiency |
| Dynamic window | Good self-adaptation to environments | Not appropriate for unstructured complex environments |
Table 6.
DRL motion planning methods.
| Algorithms | Advantages | Disadvantages |
|---|---|---|
| DQN | Updates are done offline, are not complex, and are reliable | Only discrete motions |
| DDPG | High sample efficiency, less data correlation and faster convergence compared to DQN | The poor generalisation of novel environments |
| TRPO | Ensure stable convergence | Too many assumptions, may create large errors |
| PPO | Simplified solution process, good performance and easier to implement compared to TRPO | Low sampling efficiency |
| A3C | Asynchronous parallel network training, fast convergence, suitable for multi-robot systems | Require large training data, Difficult to migrate model to real world |
| SAC | Better robustness and sample efficiency compared to the above methods | Bulky model size |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.