OER·harvester

← Back to the library
arXiv HTML resource

A Survey on Deep Learning for Localization and Mapping: Towards the Age of Spatial Machine Intelligence

Deep learning based localization and mapping has recently attracted significant attention. Instead of creating hand-designed algorithms through exploitation of physical models or geometric theories, deep learning based solutions provide an alternative to solve the problem in a data-driven way. Benefiting from ever-increasing volumes of data and computational power, these methods are fast evolving into a new area tha…

Licence
OPEN CC-BY-4.0
Authors
Changhao Chen, Bing Wang, Chris Xiaoxuan Lu, Niki Trigoni, Andrew Mar…
Published
2020-06-22 · arXiv
Language
en
Length
23410 words
Type
narrative text

Cites 19 works

inferred
Open original ↗

VI SLAM

Simultaneously tracking self-motion and estimate the structure of surroundings constructs a simultaneous localization and mapping (SLAM) system. The individual modules of localization and mapping discussed in the above sections can be viewed as modules of a complete SLAM systems. This section overviews the SLAM systems using deep learning, with the main focus on the modules that contribute to the integration of a SLAM system, including local/global optimization, keyframe/loop closure detection and uncertainty estimation. Table VI summarizes the existing approaches that employ the deep learning based SLAM modules discussed in this section.

VI-A Local Optimization

When jointly optimizing estimated camera motion and scene geometry, SLAM systems enforce them to satisfy a certain constraint. This is done by minimizing a geometric or photometric loss to ensure their consistency in the local area - the surroundings of camera poses, which can be viewed as a bundle adjustment (BA) problem[233]. Learning based approaches predict depth maps and ego-motion through two individual networks [29] trained above large datasets. During the testing procedure when deployed online, there is a requirement that enforces the predictions to satisfy the local constraints. To enable local optimization, traditionally, the second-order solvers, e.g. Gauss-Newton (GN) method or Levenberg-Marquadt (LM) algorithm [234], are applied to optimize motion transformations and per-pixel depth maps

To this end, LS-Net [235] tackled this problem via a learning based optimizer by integrating the analytical solvers into its learning process. It learns a data-driven prior, followed by refining the DNN predictions with an analytical optimizer to ensure photometric consistency. BA-Net [236] integrates a differentiable second-order optimizer (LM algorithm) into a deep neural network to achieve end-to-end learning. Instead of minimizing geometric or photometric error, BA-Net is performed on feature space to optimize the consistency loss of features from multiview images extracted by ConvNets. This feature-level optimizer can mitigate the fundamental problems of geometric or photometric solution, i.e. some information may be lost in the geometric optimization, while environmental dynamics and lighting changes may impact the photometric optimization). These learning based optimizers provide an alternative to solve bundle adjustment problem.

Modules Employed by
Local optimization [235, 236]
Global optimization [123, 237, 64, 238]
Keyframe detection [77]
Loop-closure detection [239, 240, 241, 242]
Uncertainty Estimation [243, 135, 137]

TABLE VI: A summary of existing approaches on deep learning for SLAM

VI-B Global Optimization

Odometry estimation suffers from the accumulative error drifts during long-term operations, due to the fundamental problems of path integration, i.e. the system error accumulate without effective restrictions. To address this, graph-SLAM [42] constructs a topological graph to represent camera poses or scene features as graph nodes, which are connected by edges (measured by sensors) to constrain the poses. This graph-based formulation can be optimized to ensure the global consistency of graph nodes and edges, mitigating the possible errors on pose estimates and the inherent sensor measurement noise. A popular solver for global optimization is through Levenberg-Marquardt (LM) algorithm.

In the era of deep learning, deep neural networks excel at extracting features, and constructing functions from observations to poses and scene representations. A global optimization upon the DNN predictions is necessary to reducing the drifts of global trajectories and support large-scale mapping. Compared with a variety of well-researched solutions in classical SLAM, optimizing deep predictions globally is underexplored.

Existing works explored to combine learning modules into a classical SLAM system at different levels - in the front-end, DNNs produce predictions as priors, followed by incorporating these deep predictions into the back-end for next step optimization and refinement. One good example is CNN-SLAM [123], which utilizes the learned per-pixel depths into LSD-SLAM [124], a full SLAM system to support loop closing and graph optimization. Camera poses and scene representations are jointly optimized with depth maps to produce consistent scale metrics. In DeepTAM [237], both the depth and pose predictions from deep neural networks are introduced into a classical DTAM system [121], that is optimized globally by the back-end to achieve more accurate scene reconstruction and camera motion tracking. A similar work can be found on integrating unsupervised VO with a graph optimization back-end [64]. DeepFactors [238] vice versa integrates the learned optimizable scene representation (their so-called code representation) into a different style of back-end - probabilistic factor graph for global optimization. The advantage of the factor-graph based formulation is its flexibility to include sensor measurements, state estimates, and constraints. In a factor graph bach-end, it is quite easy and convenient to add new sensor modalities, pairwise constraints and system states into the graph for optimization. However, these back-end optimizers are not yet differentiable.

VI-C Keyframe and Loop-closure Detection

Detecting keyframe and loop-closing is of key importance to the back-end optimization of SLAM systems.

Keyframe selection facilitates SLAM systems to be more efficient. In the key-frame based SLAM systems, pose and scene estimates are only refined when a keyframe is detected. [77] provides a learning solution to detect key-frames together with unsupervised learning of ego-motion tracking and depths estimation [29]. Whether an image is the keyframe is determined by comparing its feature similarity with existing keyframes (i.e. if the similarity is below a threshold, this image will be treated as a new keyframe).

Loop-closure detection or place recognition is also an important module in SLAM back-end to reduce open-loop errors. Conventional works are based on bag-of-words (BoW) to store and use the visual features from the hand-crafted detectors. However, this problem is complicated by the changes of illumination, weather, viewpoints and moving objects in real-world scenarios. To solve this, previous researchers such as [239] proposed to use the ConvNet features instead, that are from a pre-trained model on a generic large-scale image processing dataset. These methods are more robust against the variance of viewpoints and conditions due to the high-level representations extracted by deep neural networks. Other representative works [240, 241, 242] are built on deep auto-encoder structure to extract a compact representation, that compresses scene in an unsupervised manner. Deep learning based loop closing contributes more robust and effective visual features, and achieves state-of-the-art performance in place recognition, which is suitable to be integrated in SLAM systems.

VI-D Uncertainty Estimation

Safety and interpretability are an critical step towards the practical deployment of mobile agents in everyday life: the former enables agents to live and act with human reliably, while the latter allows users to have better understanding over the model behaviours. Although deep learning models achieve state-of-the-art performance in a wide range of regression and classification tasks, some corner cases should be given enough attention as well. In these failure cases, errors from one component will propagate to the other downstream modules, causing catastrophic consequences. To this end, there is an emerging need to estimate uncertainty for deep neural networks to ensure safety and provide interpretability.

Deep learning models usually only produce the mean values of predictions, for example, the output of a DNN-based visual odometry model is a 6-dimensional relative pose vector, i.e. the translation and rotation. In order to capture the uncertainty of deep models, learning models can be augmented into a Bayesian model [244, 245]. The uncertainty from Bayesian models is broadly categorized into Aleatoric uncertainty and epistemic uncertainty: Aleatoric uncertainty reflects observation noises, e.g. sensor measurement or motion noises; epistemic uncertainty captures the model uncertainty [245]. In the context of this survey, we focus on the work of estimating uncertainty on the specific task of localization and mapping, with regard to their usages, i.e. whether they capture the uncertainty with the purpose of motion tracking or scene understanding.

The uncertainty of DNN-based odometry estimation has been explored by [243, 246]. They adopted a common strategy to convert the target predictions into a Gaussian distribution, conditioned on the mean value of pose estimates and its covariance. The parameters inside the framework are optimized via the loss function with a combination of mean and covariance. By minimizing the error function to find the best combination, the uncertainty is automatically learned in an unsupervised fashion. In this way, the uncertainty of motion transformation is recovered. The motion uncertainty plays a vital role in probabilistic sensor fusion or the back-end optimization of SLAM systems. To validate the effectiveness of uncertainty estimation in SLAM systems, [243] integrated the learned uncertainty into a graph-SLAM as the covariances of odometry edges. Based on these covariances a global optimization is then performed to reduce system drifts. It also confirms that uncertainty estimation improves the performance of SLAM systems over the baseline with a fixed predefined value of covariance. Similar Bayesian models are applied to the global relocalization problem. As illustrated in [135, 137], the uncertainty from deep models are able to reflect the global location errors, in which the unreliable pose estimates are avoided with this belief metric.

In addition to the uncertainty for motion/relocalization, estimating the uncertainty for scene understanding also contributes to SLAM systems. This uncertainty offers a belief metric in to what extent the environmental perception and scene structure should be trusted. For example, in the semantic segmentation and depth estimation tasks, uncertainty estimation provides per-pixel uncertainties for the DNN predictions [247, 245, 248, 249]. Further more, scene uncertainty is applicable to building hybrid SLAM systems. For example, photometric uncertainty can be learned to capture the variance of intensity on each image pixel, and hence enhances the robustness of SLAM system to observation noise [25].