IV Mapping
Mapping refers to the ability of a mobile agent to build a consistent environmental model to describe the surrounding scene. Deep learning has fostered a set of tools for scene perception and understanding, with applications ranging from depth prediction, to semantic labelling, to 3D geometry reconstruction. This section provides an overview of existing works relevant to deep learning based mapping methods. We categorize them into geometric mapping, semantic mapping, and general mapping. Table II summarizes the existing methods on deep learning based mapping.
| Output Representation | Employed by | |
|---|---|---|
| Geometric Map | Depth Representation | [78, 92, 79], [80, 81], [29], [53, 55, 56, 58, 59, 93, 61, 64], [76, 77] |
| Voxel Representation | [94, 95], 96, [97], 98, 99, 100 | |
| Point Representation | 101 | |
| Mesh Representation | 102, 103, 104, 105, [106, 107] | |
| Semantic Map | Semantic Segmentation | [26, 27, 108] |
| Instance Segmentation | [109, 110, 111] | |
| Panoptic Segmentation | [112] | |
| General Map | Neural Representation | [113],[114, 115, 116], [117, 118, 31, 32] |
- Object indicates that this method is only validated on reconstructing single objects rather than a scene.
TABLE II: A summary of existing methods on deep learning for mapping.
IV-A Geometric Mapping
Broadly, geometric mapping captures the shape and structural description of a scene. Typical choices of the scene representations used in geometric mapping include depth, voxel, point and mesh. We follow this representational taxonomy and categorize deep learning for geometric mapping into the above four classes. Figure 6 demonstrates these geometric representations on the Stanford Bunny benchmark.
Fig. 6: An illustrations of scene representations on the Stanford Bunny benchmark: (a) original model, (b) depth representation, (c) voxel representation (d) point representation (e) mesh representation.
IV-A1 Depth Representation
Depth maps play a pivotal role in understanding the scene geometry and structure. Dense scene reconstruction has been achieved by fusing depth and RGB images [119, 120]. Traditional SLAM systems represent scene geometry with dense depth maps (i.e. 2.5D), such as DTAM [121]. In addition, accurate depth estimation can contribute to the absolute scale recovery for visual SLAM.
Learning depth from raw images is a fast evolving area in computer vision community. The earliest work formulates depth estimation as a mapping function of input single images, constructed by a multi-scale deep neural network [78] to output the per-pixel depth maps from single images. More accurate depth prediction is achieved by jointly optimizing the depth and self-motion estimation [79]. These supervised learning methods [78, 92, 79] can predict per-pixel depth by training deep neural networks on large data collections of images with corresponding depth labels. Although they are found outperforming the traditional structure based methods, such as [122], their effectiveness are largely reliant on model training and can be difficult to generalize to new scenarios in absence of labeled data.
On the other side, recent advances in this field focus on unsupervised solutions, by reformulating depth prediction as a novel view synthesis problem. [80, 81] utilized photometric consistency loss as a self-supervision signal for training neural models. With stereo images and a known camera baseline, [80, 81] synthesize the left view from the right image, and the predicted depth maps of the left view. By minimizing the distance between synthesized images and real images, i.e. the spatial consistency, the parameters of the networks can be recovered via this self-supervision in an end-to-end manner. Besides the spatial consistency, [29] proposed to apply temporal consistency as a self-supervised signal, by synthesizing the image in the target time frame from the source time frame. At the same time, egomotion is recovered along with the depth estimation. This framework only requires monocular images to learn both the depth maps and egomotion. A number of following works [53, 55, 56, 58, 59, 93, 61, 64, 76, 77] extended this framework and achieved better performance in depth and egomotion estimation. We refer the readers to Section III-A2, in which a variety of additional constraints haven been discussed.
With the depth maps predicted by ConvNets, learning based SLAM systems can integrate depth information to address some limitations of classical monocular solution. For example, CNN-SLAM [123] utilizes the learned depths from single images into a monocular SLAM framework (i.e. LSD-SLAM [124]). Their experiment shows how the learned depth maps contribute to mitigate the absolute scale recovery problem in pose estimates and scene reconstruction. CNN-SLAM achieves dense scene predictions even in texture-less areas, which is normally hard for a conventional SLAM system.
IV-A2 Voxel Representation
Voxel-based formulation is a natural way to represent 3D geometry. Similar to the usage of pixel (i.e. 2D element) in images, voxel is a volume element in a three-dimensional space. Previous works have explored to use multiple input views, to reconstruct the volumetric representation of scene [94, 95] and objects [96]. For example, SurfaceNet [94] learns to predict the confidence of a voxel to determine whether it is on surface or not, and reconstruct the 2D surface of a scene. RayNet [95] reconstructs the scene geometry by extracting view-invariant features while imposing geometric constraints. Recent works focus on generating high-resolution 3D volumetric models [98, 97]. For example, Tatarchenko et al. [97] designed a convolutional decoder based on octree-based formulation to enable scene reconstruction in much higher resolution. Other work can be found on scene completion from RGB-D data [99, 100]. One limitation of voxel representation is the high computational requirement, especially when attempting to reconstruct a scene in high resolution.
IV-A3 Point Representation
Point-based formulation consists of the 3-dimensional coordinates (x, y, z) of points in 3D space. Point representation is easy to understand and manipulate, but suffers from the ambiguity problem, which means that different forms of point clouds can represent a same geometry. The pioneer work, PointNet [125], processes unordered point data with a single symmetric function - max pooling, to aggregate point features for classification and segmentation. Fan et al. [101] developed a deep generative model that generates 3D geometry in point-based formulation from single images. In their work, a loss function based on Earth Mover’s distance is introduced to tackle the problem of data ambiguity. However, their method is only validated on the reconstruction task of single objects. No work on point generation for scene reconstruction has been found yet.
IV-A4 Mesh Representation
Mesh-based formulation encodes the underlying structure of 3D models, such as edges, vertices and faces. It is a powerful representation that naturally captures the surface of 3D shape. Several works considered the problem of learning mesh generation from images [102, 103] or point clouds data [104, 105]. However, these approaches are only able to reconstruct single objects, and limited to generating models with simple structures or from familiar classes. To tackle the problem of scene reconstruction in mesh representation, [106] integrates the sparse features from monocular SLAM with dense depth maps from ConvNet for the update of 3D mesh representation. The depth predictions are fused into the monocular SLAM system to recover the absolute scale of pose and scene features estimation. To allow efficient computation and flexible information fusion, [107] utilizes 2.5D mesh to represent scene geometry. In their approach, the image plane coordinates of mesh vertices are learned by deep neural networks, while the depth maps are optimized as free variables.
IV-B Semantic Map
Fig. 7: (b) semantic segmentation, (c) instance segmentation and (d) panoptic segmentation for semantic mapping [126].
Semantic mapping connects semantic concepts (i.e. object classification, material composition etc) with the geometry of environments. This is treated as a data association problem. The advances in deep learning greatly fosters the developments of object recognition and semantic segmentation. Maps with semantic meanings enable mobile agents to have high-level understandings of their environments beyond pure geometry, and allow for a greater range of functionality and autonomy.
SemanticFusion [26] is one of the early works that combined the semantic segmentation labels from deep ConvNet with the dense scene geometry from a SLAM system. It incrementally integrates per-frame semantic segmentation predictions into a dense 3D map by probabilistically associating the 2D frames with the 3D map. This combination not only generates a map with useful semantic information, but also shows the integration with a SLAM system helps to enhance the single frame segmentation. The two modules are loosely coupled in SemanticFusion. [27] proposed a self-supervised network that predicts consistent semantic labels for a map, by imposing constraints on the consistency of semantic predictions in multiple views. DA-RNN [108] introduces recurrent models into semantic segmentation framework to learn the temporal connections over multiple view frames, producing more accurate and consistent semantic labelling for volumetric maps from KinectFusion [127]. Yet these methods provide no information on object instances, which means that they are not able to distinguish among different ojects from the same category.
With the advances in instance segmentation, semantic mapping evolves into the instance level. A good example is [109] that offers object-level semantic mapping by identifying individual objects via a bounding box detection module and an unsupervised geometric segmentation module. Unlike other dense semantic mapping approaches, Fusion++ [110] builds a semantic graph-based map, which predicts only object instances and maintains a consistent map via loop closure detection, pose-graph optimization and further refinement. [111] presented a framework that achieves instance-aware semantic mapping, and enables novel object discovery. Recently, panoptic segmentation [126] attracts a lot of attentions. PanopticFusion [112] advanced semantic mapping to the level of stuff and things level that classifies static objects, e.g. walls, doors, lanes as stuff classes, and other accountable objects as things classes, e.g. moving vehicles, human and tables. Figure 7 compares semantic segmentation, instance segmentation and panoptic segmentation.
IV-C General Map
Beyond the explicit geometric and semantic map representation, deep learning models are able to encode the whole scene into an implicit representation, i.e. a general map representation to capture the underlying scene geometry and appearance.
Utilizing deep autoencoders can automatically discover the high-level compact representation of high-dimensional data. A notable example is CodeSLAM [113] that encodes observed images into a compact and optimizable representation to contain the essential information of a dense scene. This general representation is further used into a keyframe-based SLAM system to infer both pose estimates and keyframe depth maps. Due to the reduced size of learned representations, CodeSLAM allows efficient optimization of tracking camera motion and scene geometry for a global consistency.
Neural rendering models are another family of works that learn to model 3D scene structure implicitly by exploiting view synthesis as a self-supervision signal. The target of neural rendering task is to reconstruct a new scene from an unknown viewpoint. The seminar work, Generative Query Network (GQN) [128] learns to capture representation and render a new scene. GQN consists of a representation network and a generation network: the representation network encodes the observations from reference views into a scene representation; the generation network which is based on recurrent model, reconstructs the scene from a new view conditioned on the scene representation and a stochastic latent variable. Taking inputs as the observed images from several viewpoints, and the camera pose of a new view, GQN predicts the physical scene of this new view. Intuitively, through end-to-end training, the representation network can capture the necessary and important factors of 3D environment for the scene reconstruction task via the generation network. GQN is extended by incorporating a geometric-aware attention mechanism to allow more complex environment modelling [114], as well as including multimodal data for scene inference [115]. Scene representation network (SRN) [116] tackles the scene rendering problem via a learned continuous scene representation that connects a camera pose and its corresponding observation. A differentiable Ray Marching algorithm is integrated into SRN to enforce the network to model 3D structure consistently. However, these frameworks can only be applied to synthetic datasets due to the complexity of real-world environments.
Last but not least, in the quest of ‘map-less’ navigation, task-driven maps emerge as a novel map representation. This representation is jointly modelled by deep neural networks with respect to the task at hand. Generally those tasks leverage location information, such as navigation or path planning, requiring mobile agents to understand the geometry and semantics of environment. Navigation in unstructured environments (even in a city scale) is formulated as an policy learning problem in these works [117, 118, 31, 32], and solved by deep reinforcement learning. Different from traditional solutions that follow a procedure of building an explicit map, planning path and making decisions, these learning based techniques predict control signals directly from sensor observations in an end-to-end manner, without explicitly modelling the environment. The model parameters are optimized via sparse reward signals, for example, whenever agents reach a destination, a positive reward will be given to tune the neural network. Once a model is trained, the actions of agents can be determined conditioned on the current observations of environment, i.e. images. In this case, all of environmental factors, such as the geometry, appearance and semantics of a scene, are embedded inside the neurons of a deep neural network and suitable for solving the task at hand. Interestingly, the visualization of the neurons inside a neural model that is trained on the navigation task via reinforcement learning, has similar patterns as the grid and place cells inside human brain. This provides cognitive cues to support the effectiveness of neural map representation.