V Global Localization
Global localization concerns the retrieval of absolute pose of a mobile agent within a known scene. Different from odometry estimation that relies on estimating the internal dynamical model and can perform in an unseen scenario, in global localization, prior knowledge about the scene is provided and exploited, through a 2D or 3D scene model. Broadly, it describes the relation between the sensor observations and map, by matching a query image or view against a pre-built model, and returning an estimate of global pose.
We categorize deep learning based global localization into three categories, according to the types of inquiry data and map: 2D-to-2D localization queries 2D images against an explicit database of geo-referenced images or implicit neural map; 2D-to-3D localization establishes correspondences between 2D pixels of images and 3D points of a scene model; and 3D-to-3D localization matches 3D scans to a pre-built 3D map. Table III, IV and V summarize the existing approaches on deep learning based 2D-to-2D localization, 2D-to-3D localization and 3D-to-3D localization respectively.
V-A 2D-to-2D Localization
2D-to-2D localization regresses the camera pose of an image against a 2D map. Such 2D map is explicitly built by a geo-referenced database or implicitly encoded in a neural network.
Fig. 8: The typical architectures of 2D-to-2D based localization through (a) explict map, i.e. RelocNet [129] and (b) implicit map, i.e. e.g. PoseNet [130]
| Model | Agnostic | Performance (m/degree) | Contributions | |||
|---|---|---|---|---|---|---|
| 7Scenes | Cambridge | |||||
| 2D-to-2D Localization | Explicit Map | NN-Net [131] | Yes | 0.21/9.30 | - | combine retrieval and relative pose estimation |
| DeLS-3D [132] | No | - | - | jointly learn with semantics | ||
| AnchorNet [133] | Yes | 0.09/6.74 | 0.84/2.10 | anchor point allocation | ||
| RelocNet [129] | Yes | 0.21/6.73 | - | camera frustum overlap loss | ||
| CamNet [134] | Yes | 0.04/1.69 | - | multi-stage image retrieval | ||
| Implicit Map | PoseNet [130] | No | 0.44/10.44 | 2.09/6.84 | first neural network in global pose regression | |
| Bayesian PoseNet [135] | No | 0.47/9.81 | 1.92/6.28 | estimate Bayesian uncertainty for global pose | ||
| BranchNet [136] | No | 0.29/8.30 | - | multi-task learning for orientation and translation | ||
| VidLoc [137] | No | 0.25/- | - | efficient localization from image sequences | ||
| Geometric PoseNet [138] | No | 0.23/8.12 | 1.63/2.86 | geometry-aware loss | ||
| SVS-Pose [139] | No | - | 1.33/5.17 | data augmentation in 3D space | ||
| LSTM PoseNet [140] | No | 0.31/9.85 | 1.30/5.52 | spatial correlation | ||
| Hourglass PoseNet [141] | No | 0.23/9.53 | - | hourglass-shaped architecture | ||
| VLocNet [142] | No | 0.05/3.80 | 0.78/2.82 | jointly learn global localization and odometry | ||
| MapNet [143] | No | 0.21/7.77 | 1.63/3.64 | impose spatial and temporal constraints | ||
| SPP-Net [144] | No | 0.18/6.20 | 1.24/2.68 | synthetic data augmentation | ||
| GPoseNet [145] | No | 0.30/9.90 | 2.00/4.60 | hybrid model with Gaussian Process Regressor | ||
| VLocNet++ [146] | No | 0.02/1.39 | - | jointly learn with odometry and semantics | ||
| LSG [147] | No | 0.19/7.47 | - | odometry-aided localization | ||
| PVL [148] | No | - | 1.60/4.21 | prior-guided dropout mask to improve robustness | ||
| AdPR [149] | No | 0.22/8.8 | - | adversarial architecture | ||
| AtLoc [150] | No | 0.20/7.56 | - | attention-guided spatial correlation |
- Agnostic indicates whether it can generalize to new scenarios.
- Performance reports the position (m) and orientation (degree) error (a small number is better) on the 7-Scenes (Indoor)[151] and Cambridge (Outdoor) dataset[130]. Both datasets are split into training and testing set. We report the averaged error on the testing set.
- Contributions summarize the main contributions of each work compared with previous research.
TABLE III: A summary on existing methods on deep learning for 2D-to-2D global localization
V-A1 Explicit Map Based Localization
Explicit map based 2D-to-2D localization typically represents the scene by a database of geo-tagged images (references) [152, 153, 154]. Figure 8 (a) illustrates the two stages of this localization with 2D references: image retrieval determines the most relevant part of a scene represented by reference images to the visual queries; pose regression obtains the relative pose of query image with respect to the reference images.
One problem here is how to find suitable image descriptors for image retrieval. Deep learning based approaches [155, 156] are based on a pre-trained ConvNet model to extract image-level features, and then use these features to evaluate the similarities against other images. In challenging situations, local descriptors are first extracted, followed by being aggregated to obtain robust global descriptors. A good example is NetVLAD [157] that designs a trainable generalized VLAD (the Vector of Locally Aggregated Descriptors) layer. This VLAD layer can be plugged into the off-the-shelf ConvNet architecture to encourage better descriptors learning for image retrieval.
In order to obtain more precise poses of the queries, additional relative pose estimation with respect to the retrieved images is required. Traditionally, relative pose estimation is tackled by epipolar geometry, relying on the 2D-2D correspondences determined by local descriptors [158, 159]. In contrast, deep learning approaches regress the relative poses straightforwardly from pairwise images. For example, NN-Net [131] utilized neural network to estimate the pairwise relative poses between the query and the top N ranked references. A triangulation-based fusion algorithm coalesces the predicted N relative poses and the ground truth of 3D geometry poses, and the absolute query pose can be naturally calculated. Furthermore, Relocnet [129] introduces a frustum overlap loss to assist global descriptors learning that are suitable for camera localization. Motivated by these, CamNet [134] applies two stages retrieval, image-based coarse retrieval and pose-based fine retrieval, to select the most similar reference frames for finally precise pose estimation. Without the need of training on specific scenarios, reference-based approaches are naturally scalable and flexible to be utilized in new scenarios. Since reference-based methods need to maintain a database of geo-tagged images, they are more trivial to scale to large-scale scenarios, compared with the structure-based counterparts. Overall, these image retrieval based methods achieve a trade-off between accuracy and scalability.
V-A2 Implicit Map Based Localization
Implicit map based localization directly regresses camera pose from single images, by implicitly representing the structure of entire scene inside a deep neural network. The common pipeline is illustrated in Figure 8 (b) - the input to a neural network is single images, while the output is the global position and orientation of query images.
PoseNet [130] is the first work to tackle camera relocalization problem by training a ConvNet to predict camera pose from single RGB images in an end-to-end manner. PoseNet is based on the main structure of GoogleNet [160] to extract visual features,, but removes the last softmax layers. Instead, a fully connected layer was introduced to output a 7 dimensional global pose, consisting of position and orientation vector in 3 and 4 dimensions respectively. However, PoseNet was designed with a naive regression loss function without any consideration for geometry, in which the hyper-parameters inside requires expensive hand-engineering to be tuned. Furthermore, it also suffers from the overfitting problem due to the high dimensionality of the feature embedding and limited training data. Thus various extensions enhance the original pipeline by exploiting LSTM units to reduce the dimensionality [140], applying synthetic generation to augment training data [139, 136, 144], replacing backbone with ResNet34 [141], modelling pose uncertainty [135, 145] and introducing geometry-aware loss function [138]. Alternatively, Atloc [150] associates the features in spatial domain with attention mechanism, which encourages the network to focus on parts of the image that are temporally consistent and robust. Similarly, a prior guided dropout mask is additionally adopted in RVL [148] to further eliminate the uncertainty caused by dynamic objects. Different such methods only considering spatial connections, VidLoc [137] incorporates temporal constraints of image sequences to model the temporal connections of input images for visual localization. Moreover, additional motion constraints, including spatial constraints and other sensor constraints from GPS or SLAM systems are exploited in MapNet [143], to enforce the motion consistency between predicted poses. Similar motion constraints are also added by jointly optimizing a relocalization network and visual odometry network [142, 147]. However, being application-specific, scene representations learned from localization tasks may ignore some useful features they are not designed for. Out of this, VLocNet++ [146] and FGSN [161] additionally exploits the inter-task relationship among learning semantics and regressing poses, achieving impressive results.
Implicit map based localization approaches take the advantages of deep learning in automatically extracting features, that play a vital role in global localization in featureless environments, where conventional methods are prone to fail. However, the requirement of scene-specific training prohibits it from generalizing to unseen scenes without being retrained. Also, current implicit map based approaches have not shown comparable performance over other explicit map based methods [162].
V-B 2D-to-3D Localization
2D-to-3D localization refers to methods that recover the camera pose of a 2D image with respect to a 3D scene model. This 3D map is pre-built before performing global localization, via approaches such as structure from motion (SfM)[43]. As shown in Figure 9, 2D-to-3D approaches establish 2D-3D correspondences between the 2D pixels of query image and the 3D points of scene model through local descriptor matching [163, 164, 165] or by regressing 3D coordinates from pixel patches [166, 167, 151, 168]. Such 2D-3D matches are then used to calculate camera pose by applying a Perspective-n-Point (PnP) solver [169, 170] inside a RANSAC loop [171].
![Fig. 9: The typical architectures of 2D-to-3D based localization through (a) descriptor matching, i.e. HF-Net [172] and (b) scene coordinate regression, i.e. Confidence SCR [173].](https://arxiv.org/html/2006.12567v2/structure.png)
Fig. 9: The typical architectures of 2D-to-3D based localization through (a) descriptor matching, i.e. HF-Net [172] and (b) scene coordinate regression, i.e. Confidence SCR [173].
| Model | Agnostic | Performance (m/degree) | Contributions | |||
|---|---|---|---|---|---|---|
| 7Scenes | Cambridge | |||||
| 2D-3D Localization | Descriptor Based | NetVLAD [157] | Yes | - | - | differentiable VLAD layer |
| DELF [174] | Yes | - | - | attentive local feature descriptor | ||
| InLoc [175] | Yes | 0.04/1.38 | 0.31/0.73 | dense data association | ||
| SVL [176] | No | - | - | leverage a generative model for descriptor learning | ||
| SuperPoint [177] | Yes | - | - | jointly extract interest points and descriptors | ||
| Sarlin et al. [178] | Yes | - | - | hierarchical localization | ||
| NC-Net [179] | Yes | - | - | neighbourhood consensus constraints | ||
| 2D3D-MatchNet [180] | Yes | - | - | jointly learn the descriptors for 2D and 3D keypoints | ||
| Unsuperpoint [181] | Yes | - | - | unsupervised detector and descriptor learning | ||
| HF-Net [172] | Yes | - | - | coarse-to-fine localization | ||
| D2-Net [182] | Yes | - | - | jointly learn keypoints and descriptors | ||
| Speciale et al [183] | No | - | - | privacy preserving localization | ||
| OOI-Net [184] | No | - | - | objects-of-interest annotations | ||
| Camposeco et al. [185] | Yes | - | 0.56/0.66 | hybrid scene compression for localization | ||
| Cheng et al. [186] | Yes | - | - | cascaded parallel filtering | ||
| Taira et al. [187] | Yes | - | - | comprehensive analysis of pose verification | ||
| R2D2 [188] | Yes | - | - | learn a predictor of the descriptor discriminativeness | ||
| ASLFeat [189] | Yes | - | - | leverage deformable convolutional networks | ||
| Scene Coordinate Regression | DSAC [190] | No | 0.20/6.3 | 0.32/0.78 | differentiable RANSAC | |
| DSAC++ [191] | No | 0.08/2.40 | 0.19/0.50 | without using a 3D model of the scene | ||
| Angle DSAC++ [192] | No | 0.06/1.47 | 0.17/0.50 | angle-based reprojection loss | ||
| Dense SCR [193] | No | 0.04/1.4 | - | full frame scene coordinate regression | ||
| Confidence SCR [173] | No | 0.06/3.1 | - | model uncertainty of correspondences | ||
| ESAC [194] | No | 0.034/1.50 | - | integrates DSAC in a Mixture of Experts | ||
| NG-RANSAC [195] | No | - | 0.24/0.30 | prior-guided model hypothesis search | ||
| SANet [196] | Yes | 0.05/1.68 | 0.23/0.53 | scene agnostic architecture for camera localization | ||
| MV-SCR [197] | No | 0.05/1.63 | 0.17/0.40 | multi-view constraints | ||
| HSC-Net [198] | No | 0.03/0.90 | 0.13/0.30 | hierarchical scene coordinate network | ||
| KFNet [199] | No | 0.03/0.88 | 0.13/0.30 | extends the problem to the time domain |
- Agnostic indicates whether it can generalize to new scenarios.
- Performance reports the position (m) and orientation (degree) error (a small number is better) on the 7-Scenes (Indoor)[151] and Cambridge (Outdoor) dataset[130]. Both datasets are split into training and testing set. We report the averaged error on the testing set.
- Contributions summarize the main contributions of each work compared with previous research.
TABLE IV: A summary on existing methods on deep learning for 2D-to-3D global localization
V-B1 Descriptor Matching Based Localization
Descriptor matching methods mainly rely on feature detector and descriptor, and establish the correspondences between the features from 2D input and 3D model. They can be further divided into three types: detect-then-describe, detect-and-describe, and describe-to-detect, according to the role of detector and descriptor in the learning process.
Detect-then-describe approach first performs feature detection and then extracts a feature descriptor from a patch centered around each keypoint [200, 201]. The keypoint detector is typically responsible for providing robustness or invariance against possible real issues such as scale transformation, rotation, or viewpoint changes by normalizing the patch accordingly. However, some of these responsibilities might also be delegated to the descriptor. The common pipeline varies from using hand-crafted detectors [202, 203] and descriptors [204, 205], replacing either the descriptor [206, 207, 208, 179, 209, 210] or detector [211, 212, 213] with a learned alternative, or learning both the detector and descriptor [214, 215]. For efficiency, the feature detector often considers only small image regions and typically focuses on low-level structures such as corners or blobs [216]. The descriptor then captures higher level information in a larger patch around the keypoint.
In contrast, detect-and-describe approaches advance description stage. By sharing a representation from deep neural network, SuperPoint [177], UnSuperPoint [181] and R2D2 [188] attempt to learn a dense feature descriptor and a feature detector. However, they rely on different decoder branches which are trained independently with specific losses. On the contrary, D2-net [182] and ASLFeat [189] shares all parameters between detection and description and uses a joint formulation that simultaneously optimizes for both tasks.
Similarly, the describe-to-detect approach, e.g. D2D [217], also postpones the detection to a later stage but applies such detector on pre-learned dense descriptors to extract a sparse set of keypoints and corresponding descriptors. Dense feature extraction foregoes the detection stage and performs the description stage densely across the whole image [218, 219, 220, 176]. In practice, this approach has shown to lead to better matching results than sparse feature matching, particularly under strong variations in illumination [221, 222]. Different from these works, which purely rely on image features, 2D3D-MatchNet [180] proposed to learn local descriptors that allow direct matching of key points across a 2D image and 3D point cloud. Similarly, LCD [223] introduced a dual auto-encoder architecture to extract cross-domain local descriptors. However, they still require pre-defined 2D and 3D keypoints separately, which will result in poor matching results caused by inconsistent keypoint selection rules.
![Fig. 10: The typical architecture of 3D-to-3D localization, e.g. L3-Net [224].](https://arxiv.org/html/2006.12567v2/Lidar_loc.png)
Fig. 10: The typical architecture of 3D-to-3D localization, e.g. L3-Net [224].
| Models | Agnostic | Contributions |
|---|---|---|
| LocNet[225] | No | convert 3D points into 2D matrix, search in global prior map |
| PointNetVLAD[226] | Yes | learn global descriptor from point clouds |
| Barsan et al.[227] | No | learn from LIDAR intensity maps and online point clouds |
| L3-Net[224] | No | extract feature by PointNet |
| PCAN[228] | Yes | predict the significance of each local point based on context |
| DeepICP[229] | Yes | generate matching correspondence from learned matching probabilities |
| DCP[230] | Yes | a learning based iterative closest point |
| D3Feat[231] | Yes | jointly learn detector and descriptors for 3D points |
- Agnostic indicates whether it can generalize to new scenarios.
- Contributions summarize the main contributions of each work compared with previous research.
TABLE V: A summary of existing approaches on deep learning for 3D-to-3D localization
V-B2 Scene Coordinate Regression Based Localization
Different from match-based methods that establish 2D-3D correspondences before calculating pose, scene coordinate regression approaches estimate the 3D coordinates of each pixel from the query image within the world coordinate system, i.e. the scene coordinates. It can be viewed as learning a transformation from the query image to the global coordinates of the scene. DSAC [190] utilizes a ConvNet model to regress scene coordinates, followed by a novel differentiable RANSAC to allow end-to-end training of the whole pipeline. Such common pipeline was then improved by introducing the reprojection loss [191, 232, 192] or multi-view geometric constraints [197] to enable unsupervised learning, jointly learning the observation confidences [173, 195] to enhance the sampling efficiency and accuracy, exploiting Mixture of Experts (MoE) strategy [194] or hierarchical coarse-to-fine [198] to eliminate environment ambiguities. Different from these, KFNet [199] extends the scene coordinate regression problem to the time domain and thus bridges the existing performance gap between temporal and one-shot relocalization approaches. However, they still trained for a specific scene and cannot be generalized to unseen scenes without retraining. To build a scene agnostic method, SANet [196] regress the scene coordinate map of the query by interpolating the 3D points associated with retrieved scene images. Unlike aforementioned methods trained in a patch-based manner, Dense SCR [193] propose to perform the scene coordinate regression in a full-frame manner to make the computation efficient at test time and, more importantly, to add more global context to the regression process to improve the robustness.
Scene coordinate regression methods often perform better robustness and higher accuracy under small indoor scenarios, outperforming traditional algorithms such as [18]. But they have not yet proven their capacity in large-scale scenes.
V-C 3D-to-3D Localization
3D-to-3D localization (or LIDAR localization) refers to methods that recover the global pose of 3D points (i.e. LIDAR point cloud scans) against a pre-built 3D map by establishing a 3D-to-3D correspondence matching. Figure 10 shows the pipeline of 3D-to-3D localization: online scans or predicted coarse poses are applied to query the most similar 3D map data, which are further used for precise localization by calculating the offset between predicted poses and ground truths or estimating the relative poses between online scans and queried scene.
By formulating LIDAR localization as a recursive Bayesian inference problem, [227] embeds both LIDAR intensity maps and online point cloud sweeps in a sharing space for fully differentiable pose estimation. Instead of operating on 3D data directly, LocNet [225] converts point cloud scans to 2D rotational invariant representation for searching similar frames in the global prior map, and then performs the iterative closest point (ICP) methods to calculate global pose. Towards proposing a learning based LIDAR localization framework that directly processes point clouds, L3-Net [224] processes point cloud data with PointNet [125] to extract feature descriptors that encode certain useful properties, and models the temporal connections of motion dynamics via a recurrent neural network. It optimizes the loss between the predicted poses and ground truth values by minimizing the matching distance between the point cloud input and the 3D map. Some techniques, such as PointNetVLAD [226], PCAN [228] and D3Feat [231] explored to retrieve the reference scene at the beginning, while other techniques such as DeepICP [229] and DCP [230] allow to estimate relative motion transformations from 3D scans. Compared with image-based relocalization including 2D-to-3D and 2D-to-2D localization, 3D-to-3D localization is relatively underexplored.