OER·harvester

← Back to the library
arXiv HTML resource

A Survey on Deep Learning for Localization and Mapping: Towards the Age of Spatial Machine Intelligence

Deep learning based localization and mapping has recently attracted significant attention. Instead of creating hand-designed algorithms through exploitation of physical models or geometric theories, deep learning based solutions provide an alternative to solve the problem in a data-driven way. Benefiting from ever-increasing volumes of data and computational power, these methods are fast evolving into a new area tha…

Licence
OPEN CC-BY-4.0
Authors
Changhao Chen, Bing Wang, Chris Xiaoxuan Lu, Niki Trigoni, Andrew Mar…
Published
2020-06-22 · arXiv
Language
en
Length
23410 words
Type
narrative text

Cites 19 works

inferred
Open original ↗

I Introduction

Localization and mapping is a fundamental need for human and mobile agents. As a motivating example, humans are able to perceive their self-motion and environment via multimodal sensory perception, and rely on this awareness to locate and navigate themselves in a complex three-dimensional space [1]. This ability is part of human spatial ability. Furthermore, the ability to perceive self-motion and their surroundings plays a vital role in developing cognition, and motor control [2]. In a similar vein, artificial agents or robots should also be able to perceive the environment and estimate their system states using on-board sensors. These agents could be any form of robot, e.g. self-driving vehicles, delivery drones or home service robots, sensing their surroundings and autonomously making decisions [3]. Equivalently, as emerging Augmented Reality (AR) and Virtual Reality (VR) technologies interweave cyber space and physical environments, the ability of machines to be perceptually aware underpins seamless human-machine interaction. Further applications also include mobile and wearable devices, such as smartphones, wristbands or Internet-of-Things (IoT) devices, providing users with a wide range of location-based-services, ranging from pedestrian navigation [4], to sports/activity monitoring [5], to animal tracking [6], or emergency response [7] for first-responders.

Fig. 1: A spatial machine intelligence system exploits on-board sensors to perceive self-motion, global pose, scene geometry and semantics. (a) Conventional model based solutions build hand-designed algorithms to convert input sensor data to target values. (c) Data-driven solutions exploit learning models to construct this mapping function. (b) Hybrid approaches combine both hand-crafted algorithms and learning models. This survey discusses (b) and (c).

Fig. 2: A taxonomy of existing works on deep learning for localization and mapping.

Enabling a high level of autonomy for these and other digital agents requires precise and robust localization, and incrementally building and maintaining a world model, with the capability to continuously process new information and adapt to various scenarios. Such a quest is termed as ‘Spatial Machine Intelligence System (SMIS)’ in our work or recently as Spatial AI in [8]. In this work, broadly, localization refers to the ability to obtain the internal system states of robot motion, including locations, orientations and velocities, whilst mapping indicates the capacity to perceive external environmental states and capture the surroundings, including the geometry, appearance and semantics of a 2D or 3D scene. These components can act individually to sense the internal or external states respectively, or jointly as in simultaneous localization and mapping (SLAM) to track pose and build a consistent environmental model in a global frame.

I-A Why to Study Deep Learning for Localization and Mapping

The problems of localization and mapping have been studied for decades, with a variety of intricate hand-designed models and algorithms being developed, for example, odometry estimation (including visual odometry [9, 10, 11], visual-inertial odometry [12, 13, 14, 15] and LIDAR odometry [16]), image-based localization[17, 18], place recognition[19], SLAM[20, 10, 21], and structure from motion (SfM)[22, 23]. Under ideal conditions, these sensors and models are capable of accurately estimating system states without time bound and across different environments. However, in reality, imperfect sensor measurements, inaccurate system modelling, complex environmental dynamics and unrealistic constraints impact both the accuracy and reliability of hand-crated systems.

The limitations of model based solutions, together with recent advances in machine learning, especially deep learning, have motivated researchers to consider data-driven (learning) methods as an alternative to solve problem. Figure 1 summarizes the relation between input sensor data (e.g. visual, inertial, LIDAR data or other sensors) and output target values (e.g. location, orientation, scene geometry or semantics) as a mapping function. Conventional model-based solutions are achieved by hand-designing algorithms and calibrating to a particular application domain, while the learning based approaches construct this mapping function by learned knowledge. The advantages of learning based methods are three-fold:

First of all, learning methods can leverage highly expressive deep neural network as an universal approximator, and automatically discover features relevant to task. This property enables learned models to be resilience to circumstances, such as featureless areas, dynamic lightning conditions, motion blur, accurate camera calibration, which are challenging to model by hand [3]. As a representative example, visual odometry has achieved notable improvements in terms of robustness by incorporating data-driven methods in its design [24, 25], outperforming the state-of-the-art conventional algorithms. Moreover, learning approaches are able to connect abstract elements with human understandable terms[26, 27], such as semantics labelling in SLAM, which is hard to describe in a formal mathematical way.

Fig. 3: High-level conceptual illustration of a spatial machine intelligence system (i.e. deep learning based localization and mapping). Rounded rectangles represent a function module, while arrow lines connect these modules for data input and output. It is not necessary to include all modules to perform this system.

Secondly, learning methods allow spatial machine intelligence systems to learn from past experience, and actively exploit new information. By building a generic data-driven model, it avoids human effort on specifying the full knowledge about mathematical and physical rules[28], to solve domain specific problem, before being deployed. This ability potentially enables learning machines to automatically discover new computational solutions, further develop themselves and improve their models, within new scenarios or confronting new circumstances. A good example is that by using novel view synthesis as a self-supervision signal, self-motion and depth can be recovered from unlabelled videos[29, 30]. In addition, the learned representations can further support high-level tasks, such as path planning[31], and decision making[32], by constructing task-driven maps.

The third benefit is its capability of fully exploiting the increasing amount of sensor data and computational power. Deep learning or deep neural network has the capacity to scale to large-scale problems. The huge amount of parameters inside a DNN framework are automatically optimized by minimizing a loss function, by training on large datasets through backpropagation and gradient-descent algorithms. For example, the recent released GPT-3[33], the largest pretrained language model, with incredibly over 175 Billion parameters, achieves the state-of-the-art results on a variety of natural language processing (NLP) tasks, even without fine-tuning. In addition, a variety of large-scale datasets relevant to localization and mapping have been released, for example, in the autonomous vehicles scenarios, [34, 35, 36] are with a collection of rich combinations of sensor data, and motion and semantic labels. This gives us an imagination that it would be possible to exploit the power of data and computation in solving localization and mapping.

However, it must also be pointed out that these learning techniques are reliant on massive datasets to extract statistically meaningful patterns and can struggle to generalize to out-of-set environments. There is lack of model interpretability. Additionally, although highly parallelizable, they are also typically more computationally costly than simpler models. Details of limitations are discussed in Section 7.

I-B Comparison with Other Surveys

There are several survey papers that have extensively discussed model-based localization and mapping approaches. The development of SLAM problem in early decades has been well summarized in [37, 38]. The seminal survey [39] provides a thorough discussion on existing SLAM work, reviews the history of development and charts several future directions. Although this paper contains a section which briefly discusses deep learning models, it does not overview this field comprehensively, especially due to the explosion of research in this area of the past five years. Other SLAM survey papers only focus on individual flavours of SLAM systems, including the probabilistic formulation of SLAM [40], visual odometry [41], pose-graph SLAM [42], and SLAM in dynamic environments [43]. We refer readers to these surveys for a better understanding of the conventional model based solutions. On the other hand, [3] has a discussion on the applications of deep learning to robotics research; however, its main focus is not on localization and mapping specifically, but a more general perspective towards the potentials and limits of deep learning in a broad context of robotics, including policy learning, reasoning and planning.

Notably, although the problem of localization and mapping falls into the key notion of robotics, the incorporation of learning methods progresses in tandem with other research areas such as machine learning, computer vision and even natural language processing. This cross-disciplinary area thus imposes non-trivial difficulty when comprehensively summarizing related works into a survey paper. To the best of our knowledge, this is the first survey article that thoroughly and extensively covers existing work on deep learning for localization and mapping.

I-C Survey Organization

The remainder of the paper is organized as follows: Section 2 offers an overview and presents a taxonomy of existing deep learning based localization and mapping; Sections 3, 4, 5, 6 discuss the existing deep learning works on relative motion (odometry) estimation, mapping methods for geometric, semantic and general, global localization, and simultaneous localization and mapping with a focus on SLAM back-ends respectively; Open questions are summarized in Section 7 to discuss the limitations and future prospects of existing work; and finally Section 8 concludes the paper.