OER·harvester

← Back to the library
arXiv HTML resource

A Survey on Deep Learning for Localization and Mapping: Towards the Age of Spatial Machine Intelligence

Deep learning based localization and mapping has recently attracted significant attention. Instead of creating hand-designed algorithms through exploitation of physical models or geometric theories, deep learning based solutions provide an alternative to solve the problem in a data-driven way. Benefiting from ever-increasing volumes of data and computational power, these methods are fast evolving into a new area tha…

Licence
OPEN CC-BY-4.0
Authors
Changhao Chen, Bing Wang, Chris Xiaoxuan Lu, Niki Trigoni, Andrew Mar…
Published
2020-06-22 · arXiv
Language
en
Length
23410 words
Type
narrative text

Cites 19 works

inferred
Open original ↗

II Taxonomy of Existing Approaches

We provide a new taxonomy of existing deep learning approaches, relevant to localization and mapping, to connect the fields of robotics, computer vision and machine learning. Broadly, they can be categorized into odometry estimation, mapping, global localization and SLAM, as illustrated by the taxonomy shown in Figure 2:

  1. Odometry estimation concerns the calculation of the relative change in pose, in terms of translation and rotation, between two or more frames of sensor data. It continuously tracks self-motion, and is followed by a process to integrate these pose changes with respect to an initial state to derive global pose, in terms of position and orientation. This is widely known as the so-called dead reckoning solution. Odometry estimation can be used in providing pose information and as odometry motion model to assist the feedback loop of robot control. The key problem is to accurately estimate motion transformations from various sensor measurements. To this end, deep learning is applied to model the motion dynamics in an end-to-end fashion or extract useful features to support a pre-built system in a hybrid way.

  2. Mapping builds and reconstructs a consistent model to describe the surrounding environment. Mapping can be used to provide environment information for human operators and high-level robot tasks, constrain the error drifts of odometry estimation, and retrieve the inquiry observation for global localization [39]. Deep learning is leveraged as an useful tool to discover scene geometry and semantics from high-dimensional raw data for mapping. Deep learning based mapping methods are sub-divided into geometric, semantic, and general mapping, depending on whether the neural network learns the explicit geometry or semantics of a scene, or encodes the scene into an implicit neural representation respectively.

  3. Global localization retrieves the global pose of mobile agents in a known scene with prior knowledge. This is achieved by matching the inquiry input data with a pre-built 2D or 3D map, other spatial references, or a scene that has been visited before. It can be leveraged to reduce the pose drift of a dead reckoning system or solve the ’kidnapped robot’ problem[40]. Deep learning is used to tackle the tricky data association problem that is complicated by the changes in views, illumination, weather and scene dynamics, between the inquiry data and map.

  4. Simultaneous Localisation and Mapping (SLAM) integrates the aforementioned odometry estimation, global localization and mapping processes as front-ends, and jointly optimizes these modules to boost performance in both localization and mapping. Except these abovementioned modules, several other SLAM modules perform to ensure the consistency of the entire system as follows: local optimization ensures the local consistency of camera motion and scene geometry; global optimization aims to constrain the drift of global trajectories, and in a global scale; keyframe detection is used in keyframe-based SLAM to enable more efficient inference, while system error drifts can be mitigated by global optimization, once a loop closure is detected by loop-closure detection; uncertainty estimation provides a metric of belief in the learned poses and mapping, critical to probabilistic sensor fusion and back-end optimization in SLAM systems.

Fig. 4: The typical structure of supervised learning of visual odometry, i.e. DeepVO [24] and unsupervised learning of visual odometry, i.e. SfmLearner [29].

Despite the different design goals of individual components, the above components can be integrated into a spatial machine intelligence system (SMIS) to solve real-world challenges, allowing for robust operation, and long-term autonomy in the wild. A conceptual figure of such an integrated deep-learning based localization and mapping system is indicated in Figure 3, showing the relationship of these components. In the following sections, we will discuss these components in details.