OER·harvester

← Back to the library
arXiv HTML resource

A Survey on Deep Learning for Localization and Mapping: Towards the Age of Spatial Machine Intelligence

Deep learning based localization and mapping has recently attracted significant attention. Instead of creating hand-designed algorithms through exploitation of physical models or geometric theories, deep learning based solutions provide an alternative to solve the problem in a data-driven way. Benefiting from ever-increasing volumes of data and computational power, these methods are fast evolving into a new area tha…

Licence
OPEN CC-BY-4.0
Authors
Changhao Chen, Bing Wang, Chris Xiaoxuan Lu, Niki Trigoni, Andrew Mar…
Published
2020-06-22 · arXiv
Language
en
Length
23410 words
Type
narrative text

Cites 19 works

inferred
Open original ↗

III Odometry Estimation

We begin with odometry estimation, which continuously tracks camera egomotion and produces relative poses. Global trajectories are reconstructed by integrating these relative poses, given an initial state, and thus it is critical to keep motion transformation estimates accurate enough to ensure high-prevision localization in a global scale. This section discusses deep learning approaches to achieve odometry estimation from various sensor data, that are fundamentally different in their data properties and application scenarios. The discussion mainly focuses on odometry estimation from visual, inertial and point-cloud data, as they are the common choices of sensing modalities on mobile agents.

III-A Visual Odometry

Visual odometry (VO) estimates the ego-motion of a camera, and integrates the relative motion between images into global poses. Deep learning methods are capable of extracting high-level feature representations from images, and thereby provide an alternative to solve the VO problem, without requiring hand-crafted feature extractors. Existing deep learning based VO models can be categorized into end-to-end VO and hybrid VO, depending on whether they are purely neural-network based or whether they are a combination of classical VO algorithms and deep neural networks. Depending on the availability of ground-truth labels in the training phase, end-to-end VO systems can be further classified into supervised VO and unsupervised VO.

III-A1 Supervised Learning of VO

We start with the introduction of supervised VO, one of the most predominant approaches to learning-based odometry, by training a deep neural network model on labelled datasets to construct a mapping function from consecutive images to motion transformations directly, instead of exploiting the geometric structures of images as in conventional VO systems[41]. At its most basic, the input of deep neural network is a pair of consecutive images, and the output is the estimated translation and rotation between two frames of images.

One of the first works in this area was Konda et al. [44]. This approach formulates visual odometry as a classification problem, and predicts the discrete changes of direction and velocity from input images using a convolutional neural network (ConvNet). Costante et al. [45] used a ConvNet to extract visual features from dense optical flow, and based on these visual features to output frame-to-frame motion estimation. Nonetheless, these two works have not achieved end-to-end learning from images to motion estimates, and their performance is still limited.

DeepVO [24] utilizes a combination of convolutional neural network (ConvNet) and recurrent neural network (RNN) to enable end-to-end learning of visual odometry. The framework of DeepVO becomes a typical choice in realizing supervised learning of VO, due to its specialization in end-to-end learning. Figure 4 (a) shows the architecture of this RNN+ConvNet based VO system, which extracts visual features from pairs of images via a ConvNet, and passes features through RNNs to model the temporal correlation of features. Its ConvNet encoder is based on a FlowNet structure to extract visual features suitable for optical flow and self-motion estimation. Using a FlowNet based encoder can be regarded as introducing the prior knowledge of optical flow into the learning process, and potentially prevents DeepVO from being overfitted to the training datasets. The reccurent model summarizes the history information into its hidden states, so that the output is inferred from both past experience and current ConvNet features from sensor observation. It is trained on large-scale datasets with groundtruthed camera poses as labels. To recover the optimal parameters $\boldsymbol{\theta}^{*}$ of framework, the optimization target is to minimize the Mean Square Error (MSE) of the estimated translations $\mathbf{\hat{p}}\in\mathbb{R}^{3}$ and euler angle based rotations $\hat{\boldsymbol{\varphi}}\in\mathbb{R}^{3}$:

$$ \boldsymbol{\theta}^{*}=\operatorname*{arg\,min}_{\boldsymbol{\theta}}\frac{1}{N}\displaystyle\sum_{i=1}^{N}\displaystyle\sum_{t=1}^{T}\|\hat{\mathbf{p}}_{t}-\mathbf{p}_{t}\|_{2}^{2}+\|\hat{\boldsymbol{\varphi}}_{t}-\boldsymbol{\varphi}_{t}\|_{2}^{2}, $$

where $(\hat{\mathbf{p}}_{t},\hat{\boldsymbol{\varphi}}_{t})$ are the estimates of relative pose at the timestep $t$, $(\mathbf{p},\boldsymbol{\varphi})$ are the corresponding groundtruth values, $\boldsymbol{\theta}$ are the parameters of the DNN framework, $N$ is the number of samples.

DeepVO reports impressive results on estimating the pose of driving vehicles, even in previously unseen scenarios. In the experiment on the KITTI odometry dataset[46], this data-driven solution outperforms conventional representative monocular VO, e.g. VISO2[47] and ORB-SLAM (without loop closure) [21]. Another advantage is that supervised VO naturally produces trajectory with the absolute scale from monocular camera, while classical VO algorithm is scale-ambiguous using only monocular information. This is because deep neural network can implicitly learn and maintain the global scale from large collection of images, which can be viewed as learning from past experience to predict current scale metric.

Based on this typical model of supervised VO, a number of works further extended this approach to improve the model performance. To improve the generalization ability of supervised VO, [48] incorporates curriculum learning (i.e. the model is trained by increasing the data complexity) and geometric loss constraints. Knowledge distillation (i.e. a large model is compressed by teaching a smaller one) is applied into the supervised VO framework to greatly reduce the number of network parameters, making it more amenable for real-time operation on mobile devices [49]. Furthermore, Xue et al. [50] introduced a memory module that stores global information, and a refining module that improves pose estimates with the preserved contextual information.

In summary, these end-to-end learning methods benefit from recent advances in machine learning techniques and computational power, to automatically learn pose transformations directly from raw images that can tackle challenging real-world odometry estimation.

III-A2 Unsupervised Learning of VO

There is growing interest in exploring unsupervised learning of VO. Unsupervised solutions are capable of exploiting unlabelled sensor data, and thus it saves human effort on labelling data, and has better adaptation and generalization ability in new scenarios, where no labelled data are available. This has been achieved in a self-supervised framework that jointly learns depth and camera ego-motion from video sequences, by utilizing view synthesis as a supervisory signal[29].

As shown in Figure 4 (b), a typical unsupervised VO solution consists of a depth network to predict depth maps, and a pose network to produce motion transformations between images. The entire framework takes consecutive images as input, and the supervision signal is based on novel view synthesis - given a source image $\mathbf{I}_{s}$, the view synthesis task is to generate a synthetic target image $\mathbf{I}_{t}$. A pixel of source image $\mathbf{I}_{s}(p_{s})$ is projected onto a target view $\mathbf{I}_{t}(p_{t})$ via:

$$ p_{s}\sim\mathbf{K}\mathbf{T}_{t\to s}\mathbf{D}_{t}(p_{t})\mathbf{K}^{-1}p_{t} $$

where $\mathbf{K}$ is the camera’s intrinsic matrix, $\mathbf{T}_{t\to s}$ denotes the camera motion matrix from target frame to source frame, and $\mathbf{D}_{t}(p_{t})$ denotes the per-pixel depth maps in the target frame. The training objective is to ensure the consistency of the scene geometry by optimizing the photometric reconstruction loss between the real target image and the synthetic one:

$$ \mathcal{L}_{\text{photo}}=\sum_{<\mathbf{I}_{1},...,\mathbf{I}_{N}>\in S}\sum_{p}|\mathbf{I}_{t}(p)-\hat{\mathbf{I}}_{s}(p)|, $$

where p denotes pixel coordinates, $\mathbf{I}_{t}$ is the target image, and $\hat{\mathbf{I}}_{s}$ is the synthetic target image generated from the source image $\mathbf{I}_{s}$.

Model Sensor Supervision Scale Performance Contributions
Seq09 Seq10
VO Konda et al.[44] MC Supervised Yes - - formulate VO as a classification problem
Costante et al.[45] MC Supervised Yes 6.75 21.23 extract features from optical flow for VO estimates
Backprop KF[51] MC Hybrid Yes - - a differentiable Kalman filter based VO
DeepVO[24] MC Supervised Yes - 8.11 combine RNN and ConvNet for end-to-end learning
SfmLearner[29] MC Unsupervised No 17.84 37.91 novel view synthesis for self-supervised learning
Yin et al.[52] MC Hybrid Yes 4.14 1.70 introduce learned depth to recover scale metric
UnDeepVO[53] SC Unsupervised Yes 7.01 10.63 use fixed stereo line to recover scale metric
Barnes et al.[54] MC Hybrid Yes - - integrate learned depth and ephemeral masks
GeoNet[55] MC Unsupervised No 43.76 35.6 geometric consistency loss and 2D flow generator
Zhan et al.[56] SC Unsupervised No 11.92 12.45 use fixed stereo line for scale recovery
DPF[57] MC Hybrid Yes - - a differentiable particle filter based VO
Yang et al.[58] MC Hybrid Yes 0.83 0.74 use learned depth into classical VO
Zhao et al.[59] MC Supervised Yes - 4.38 generate dense 3D flow for VO and mapping
Struct2Depth[60] MC Unsupervised No 10.2 28.9 introduce 3D geometry structure during learning
Saputra et al.[48] MC Supervised Yes - 8.29 curriculum learning and geometric loss constraints
GANVO[61] MC Unsupervised No - - adversarial learning to generate depth
CNN-SVO[62] MC Hybrid Yes 10.69 4.84 use learned depth to initialize SVO
Xue et al.[50] MC Supervised Yes - 3.47 memory and refinement module
Wang et al.[63] MC Unsupervised Yes 9.30 7.21 integrate RNN and flow consistency constraint
Li et al.[64] MC Unsupervised No - - global optimization for pose graph
Saputra et al.[49] MC Supervised Yes - - knowledge distilling to compress deep VO model
Gordon[65] MC Unsupervised No 2.7 6.8 camera matrix learning
Koumis et al.[66] MC Supervised Yes - - 3D convolutional networks
Bian et al.[30] MC Unsupervised Yes 11.2 10.1 scale recovery from only monocular images
Zhan et al.[67] MC Hybrid Yes 2.61 2.29 integrate learned optical flow and depth
D3VO[25] MC Hybrid Yes 0.78 0.62 integrate learned depth, uncertainty and pose
VIO VINet[68] MC+I Supervised Yes - - formulate VIO as a sequential learning problem
VIOLearner[69] MC+I Unsupervised Yes 1.51 2.04 online correction module
Chen et al.[70] MC+I Supervised Yes - - feature selection for deep sensor fusion
DeepVIO[71] SC+I Unsupervised Yes 0.85 1.03 learn VIO from stereo images and IMU
LO Velas et al.[72] L Supervised Yes 4.94 3.27 ConvNet to estimate odometry from point clouds
LO-Net[73] L Supervised Yes 1.37 1.80 geometric constraint loss
DeepPCO[74] L Supervised Yes - - parallel neural network
Valente et al.[75] MC+L Supervised Yes - 7.60 sensor fusion for LIDAR and camara
  • Model: VO, VIO and LO represent visual odometry, visual-inertial odometry and LIDAR odometry respectively.
  • Sensor: MC, SC, I and L represent monocular camera, stereo camera, inertial measurement unit, and LIDAR respectively.
  • Supervision represents whether this work is a purely neural network based model trained with groundtruth labels (Supervised) or without labels (Unsupervised), or it is a combination of classical and deep neural network (Hybrid)
  • Scale indicates whether a trajectory with a global scale can be produced.
  • Performance reports the localization error (a small number is better), i.e. the averaged translational RMSE drift (%) on lengths of 100m-800m on the KITTI odometry dataset[46]. Most works were evaluated on the Sequence 09 and 10, and thus we took the results on these two sequences from their original papers for a performance comparison. Note that the training sets may be different in each work.
  • Contributions summarize the main contributions of each work compared with previous research.

TABLE I: A summary of existing methods on deep learning for odometry estimation.

However, there are basically two main problems that remained unsolved in the original work[29]: 1) this monocular image based approach is not able to provide pose estimates in a consistent global scale. Due to the scale ambiguity, no physically meaningful global trajectory can be reconstructed, limiting its real use. 2) The photometric loss assumes that the scene is static and without camera occlusions. Although the authors proposed the use of an explainability mask to remove scene dynamics, the influence of these environmental factors is still not addressed completely, which violates the assumption. To address these concerns, an increasing number of works [53, 55, 56, 58, 59, 61, 64, 76, 77] extended this unsupervised framework to achieve better performance.

To solve the global scale problem, [53, 56] proposed to utilize stereo image pairs to recover the absolute scale of pose estimation. They introduced an additional spatial photometric loss between the left and right pairs of images, as the stereo baseline (i.e. motion transformation between the left and right images) is fixed and known throughout the dataset. Once the training is complete, the network produces pose predictions using only monocular images. Thus, although it is unsupervised in the context of not having access to ground-truth, the training dataset (stereo) is different to the test set (mono). [30] tackles the scale issue by introducing a geometric consistency loss, that enforces the consistency between predicted depth maps and reconstructed depth maps. The framework transforms the predicted depth maps into a 3D space, and projects them back to produce reconstructed depth maps. In doing so, the depth predictions are able to remain scale-consistent over consecutive frames, enabling pose estimates to be scale-consistent meanwhile.

The photometric consistency constraint assumes that the entire scenario consists only of rigid static structures, e.g. buildings and lanes. However, in real-world applications, environmental dynamics (e.g. pedestrians and vehicles), will distort the photometric projection and degrade the accuracy of pose estimation. To address this concern, GeoNet [55] divides its learning process into two sub-tasks by estimating static scene structures and motion dynamics separately through a rigid structure reconstructor and a non-rigid motion localizer. In addition, GeoNet enforces a geometric consistency loss to mitigate the issues caused by camera occlusions and non-Lambertian surfaces. [59] adds a 2D flow generator along with a depth network to generate 3D flow. Benefiting from better 3D understanding of environment, their framework is able to produce more accurate camera pose, along with a point cloud map. GANVO [61] employs a generative adversarial learning paradigm for depth generation, and introduces a temporal recurrent module for pose regression. Li et al. [76] also utilized a generative adversarial network (GAN) to generate more realistic depth maps and poses, and further encourage more accurate synthetic images in the target frame. Instead of a hand-crafted metric, a discriminator is adopted to evaluate the quality of synthetic images generation. In doing so, the generative adversarial setup facilitates the generated depth maps to be more texture-rich and crisper. In this way, high-level scene perception and representation are accurately captured and environmental dynamics are implicitly tolerated.

Although unsupervised VO still cannot compete with supervised VO in performance, as illustrated in Figure 5, its concerns of scale metric and scene dynamics problem have been largely resolved. With the benefits of self-supervised learning, and ever-increasing improvement on performance, unsupervised VO would be a promising solution in providing pose information, and tightly coupled with other modules in spatial machine intelligence system.

III-A3 Hybrid VO

Unlike end-to-end VO that only relies on a deep neural network to interpret pose from data, hybrid VO integrates classical geometric models with deep learning framework. Based on mature geometric theory, they use a deep neural network to expressively replace parts of a geometry model.

A straightforward way is to incorporate the learned depth estimates into a conventional visual odometry algorithm to recover the absolute scale metric of poses [52]. Learning depth estimation is a well-researched area in the computer vision community. For example, [78, 79, 80, 81] provide per-pixel depths in a global scale by employing a trained deep neural model. Thus the so-called scale problem of conventional VO is mitigated. Barnes et al. [54] utilize both the predicted depth maps and ephemeral masks (i.e. the area of moving objects) into a VO system to improve its robustness to moving objects. Zhan et al. [67] integrate the learned depth and optical flow predictions into a conventional visual odometry model, achieving competitive performance over other baselines. Other works combine physical motion models with deep neural network e.g. via a differentiable Kalman filter [82], and a particle filter [83]. The physical model serves as an algorithmic prior in the learning process. Furthermore, D3VO [25] incorporates the deep predictions of depth, pose, and uncertainty into a direct visual odometry.

Combining the benefits from both geometric theory and deep learning, hybrid models are normally more accurate than end-to-end VO at this stage, as shown in Table 1. It is notable that hybrid models even outperform the state-of-the-art conventional monocular VO or visual-inertial odometry (VIO) systems on common benchmarks, for example, D3VO[25] defeats several popular conventional VO/VIO systems, such as DSO[84], ORB-SLAM[21], VINS-Mono[15]. This demonstrates the rapid rate of progress in this area.

III-B Visual-Inertial Odometry

Integrating visual and inertial data as visual-inertial odometry (VIO) is a well-defined problem in mobile robotics. Both cameras and inertial sensors are relatively low-cost, power-efficient and widely deployed. These two sensors are complementary: monocular cameras capture the appearance and structure of a 3D scene, while they are scale-ambiguous, and not robust to challenging scenarios, e.g. strong lighting changes, lack of texture and high-speed motion; In contrast, IMUs are completely ego-centric, scene-independent, and can also provide absolute metric scale. Nevertheless, the downside is that inertial measurements, especially from low-cost devices, are plagued by process noise and biases. An effective fusion of the measurements from these two complementary sensors is of key importance to accurate pose estimation. Thus, according to their information fusion methods, conventional model based visual-inertial approaches are roughly segmented into three different classes : filtering approaches [12], fixed-lag smoothers [13] and full smoothing methods [14].

Data-driven approaches have emerged to consider learning 6-DoF poses directly from visual and inertial measurements without human intervention or calibration. VINet [68] is the first work that formulated visual-inertial odometry as a sequential learning problem, and proposed a deep neural network framework to achieve VIO in an end-to-end manner. VINet uses a ConvNet based visual encoder to extract visual features from two consecutive RGB images, and an inertial encoder to extract inertial features from a sequence of IMU data with a long short-term memory (LSTM) network. Here, the LSTM aims to model the temporal state evolution of inertial data. The visual and inertial features are concatenated together, and taken as the input into a further LSTM module to predict relative poses, conditioned on the history of system states. This learning approach has the advantage of being more robust to calibration and relative timing offset errors. However, VINet has not fully addressed the problem of learning a meaningful sensor fusion strategy.

To tackle the deep sensor fusion problem, Chen et al. [70] proposed selective sensor fusion, a framework that selectively learns context-dependent representations for visual inertial pose estimation. Their intuition is that the importance of features from different modalities should be considered according to the exterior (i.e., environmental) and interior (i.e., device/sensor) dynamics, by fully exploiting the complementary behaviors of two sensors. Their approach outperforms those without a fusion strategy, e.g. VINet, avoiding catastrophic failures.

Similar to unsupervised VO, Visual-inertial odometry can also be solved in a self-supervised fashion using novel view synthesis. VIOLearner [69] constructs motion transformations from raw inertial data, and converts source images into target images with the camera matrix and depth maps via the Equation 2 mentioned in Section III-A2. In addition, an online error correction module corrects the intermediate errors of the framework. The network parameters are recovered by optimizing a photometric loss. Similarly, DeepVIO [71] incorporates inertial data and stereo images into this unsupervised learning framework, and is trained with a dedicated loss to reconstruct trajectories in a global scale.

Learning-based VIO cannot defeat the state-of-the-art classical model based VIOs, but they are generally more robust to real issues[68, 70, 71] such as measurement noises, bad time synchronization, thanks to the impressive ability of DNNs in feature extraction and motion modelling.

III-C Inertial Odometry

Beyond visual odometry and visual-inertial odometry, an inertial-only solution, i.e. inertial odometry provides an ubiquitous alternative to solve the odometry estimation problem. Compared with visual methods, an inertial sensor is relatively low-cost, small, energy efficient and privacy preserving. It is relatively immune to environmental factors, such as lighting conditions or moving objects. However, low-cost MEMS inertial measurement units (IMU) widely found on robots and mobile devices are corrupted with high sensor bias and noise, leading to unbounded error drifts in the strapdown inertial navigation system (SINS), if inertial data are doubly integrated.

Chen et al. [85] formulated inertial odometry as a sequential learning problem with a key observation that 2D motion displacements in the polar coordinate (i.e. polar vector) can be learned from independent windows of segmented inertial data. The key observation is that when tracking human and wheeled configurations, the frequency of their vibrations is relevant to the moving speed, which is reflected by inertial measurements. Based on this, they proposed IONet, a LSTM based framework for end-to-end learning of relative poses from sequences of inertial measurements. Trajectories are generated by integrating motion displacements. [86] leveraged deep generative models and domain adaptation technique to improve the generalization ability of deep inertial odometry in new domains. [87] extends this framework by an improved triple-channel LSTM network to predict polar vectors for drone localization from inertial data and sampling time. RIDI [88] trains a deep neural network to regress linear velocities from inertial data, calibrates the collected accelerations to satisfy the constraints of the learned velocities, and doubly integrates the accelerations into locations with a conventional physical model. Similarly, [89] compensates the error drifts of the classical SINS model with the aid of learned velocities. Other works have also explored the usage of deep learning to detect zero-velocity phase for navigating pedestrians [90] and vehicles [91]. This zero-velocity phase provides context information to correct system error drifts via Kalman filtering.

Inertial only solution can be a backup plan to offer pose information in extreme environments, where visual information is not available or is highly distorted. Deep learning has proven its capability to learn useful features from noisy IMU data, and compensate the error drifts of inertial dead reckoning, which is difficult to solve by classical algorithms.

Fig. 5: A comparison of the performance of deep learning based visual odometry with an evaluation on the Trajectory 10 of the KITTI dataset.

III-D LIDAR Odometry

LIDAR sensors provide high-frequency range measurements, with the benefits of working consistently in complex lighting conditions and optically featureless scenarios. Mobile robots and self-driving vehicles are normally equipped with LIDAR sensors to obtain relative self-motion (i.e. LIDAR odometry) and global pose with respect to a 3D map (LIDAR relocalization). The performance of LIDAR odometry is sensitive to point cloud registration errors due to non-smooth motion. In addition, the data quality of LIDAR measurements is also affected by extreme weather conditions, for example, heavy rain or fog/mist.

Traditionally, LIDAR odometry relies on point cloud registration to detect feature points, e.g. line and surface segments, and uses a matching algorithm to obtain the pose transformation by minimizing the distance between two consecutive point-cloud scans. Data-driven methods consider solving LIDAR odometry in an end-to-end fashion, by leveraging deep neural networks to construct a mapping function from point cloud scan sequences to pose estimates [72, 73, 74]. As point cloud data are challenging to be directly ingested by neural networks due to their sparse and irregularly sampled format, these methods typically convert point clouds into a regular matrix through cylindrical projection, and adopt ConvNets to extract features from consecutive point cloud scans. These networks regress relative poses and are trained via ground-truth labels. LO-Net [73] reports competitive performance over the conventional state-of-the-art algorithm, i.e. the LIDAR Odometry and Mapping (LOAM) algorithm [16].

III-E Comparison of Odometry Estimation

Table I compares existing work on odometry estimation, in terms of their sensor type, model, whether a trajectory with an absolute scale is produced, and their performance evaluation on the KITTI dataset, where available. As deep inertial odometry has not been evaluated on the KITTI dataset, we do not include inertial odometry in this table. The KITTI dataset [46] is a common benchmark for odometry estimation, consisting of a collection of sensor data from car-driving scenarios. As most data-driven approaches adopt the trajectory 09 and 10 of the KITTI dataset to evaluate model performance, we compared them according to the averaged Root Mean Square Error (RMSE) of the translation for all the subsequences of lengths (100, 200, .., 800) meters, which is provided by the official KITTI VO/SLAM evaluation metrics.

We take visual odometry as an example. Figure 5 reports the translational drifts of deep visual odometry models over time on the 10th trajectory of the KITTI dataset. Clearly, hybrid VO shows the best performance over supervised VO and unsupervised VO, as the hybrid model benefits from both the mature geometry models of traditional VO algorithms and the strong capacity for feature extraction of deep learning. Although supervised VO still outperforms unsupervised VO, the performance gap between them is diminishing as the limitations of unsupervised VO are gradually addressed. For example, it has been found that unsupervised VO now can recover global scale from monocular images [30]. Overall, data-driven visual odometry shows a remarkable increase in model performance, indicating the potentials of deep learning approaches in achieving more accurate odometry estimation in the future.