VII Open Questions
Although deep learning has brought great success to the research in localization and mapping, as aforementioned, existing models are not sophisticated enough to completely solve the problem at hand. The current form of deep solutions is still in its infancy state. Towards great autonomy in the wild, there are numerous challenges for future researchers to investigate. Practical applications of these techniques should be considered a systematic research problem. We discuss several open questions that likely lead the further development in this area.
-
End-to-end model vs. hybrid model. End-to-end learning models are able to predict self-motion and scene directly from raw data, without any hand-engineering. Benefited from the advances of deep learning, end-to-end models are evolving fast to achieve increasing performance in accuracy, efficiency and robustness. Meanwhile, these models have been shown easier to be integrated with other high-level learning tasks, e.g. path planning and navigation[31]. Fundamentally, there exist underlying physical or geometric models to govern localization and mapping systems. Whether we should develop end-to-end models relying only on the power of data-driven approaches or integrate deep learning modules into the pre-built physical/geometric models as a hybrid model is a critical question for future research. As we can see, hybrid models already achieved the state-of-the art results in many tasks, e.g. visual odometry[25], and global localization[191]. Thus, it is reasonable to investigate in the way to take better advantage of the prior empirical knowledge from deep learning for hybrid models. On the other side, pure end-to-end models are data hunger. The performance of current models can be limited by the size of training dataset, and it is essential to create large and diverse dataset to enlarge the capacity of data-driven models.
-
Unifying evaluation benchmark and metric. Finding suitable evaluating benchmark and metric is always a concern for SLAM systems. This is especially the case for DNN based systems. The predictions from DNNs are affected by the characteristics of both training and test data, including the dataset size, hyperparameters (batch size and learning rate etc.), and the difference in testing scenarios. Therefore, it is hard to fairly compare them when considering dataset differences, training/testing configuration, or evaluation metric adopted in each work. For example, the KITTI dataset is a common choice to evaluate visual odometry, but previous works split training and testing data in different ways (e.g. [24, 48, 50] used Sequence 00, 02, 08, 09 as training set, and Sequence 03, 04, 05, 06, 07, 10 as testing set, while [30, 25] used Sequence 00 - 08 as training, and left 09, and 10 as testing set). Some of them are even based on different evaluation metrics (e.g. [24, 48, 50] applied the KITTI official evaluation metric, while [29, 56] applied absolute trajectory error (ATE) as their evaluation metric). All these factors bring difficulties to a direct and fair comparison across them. Moreover, the KITTI dataset is relatively simple (the vehicle only moves in 2D translation) and in small size. It is not convincing if only results on KITTI benchmark are provided without a comprehensive evaluation in the long-term real-world experiment. In fact, there is a growing need in creating a benchmark for a though system evaluation covering various environments, self-motions and dynamics.
-
Real-world deployment. Deploying deep learning models in real-world environments is a systematic research problem. In the existing research, the prediction accuracy is always their ‘golden rule’ to follow, while other crucial issues are overlooked, such as whether the model structure and the parameter number of framework is optimal. The computational and energy consumption have to be considered on resource-constrained systems, such as low-cost robots or VR wearable devices. The prallerization opportunities, such as convolutional filters or other parallel neural network modules should be exploited in order to take better use of GPUs. Examples for consideration include in which situations the feed-back should be returned to fine-tune the systems, how to incorporate the self-supervised models into the systems and whether the systems allow the real-time online learning.
-
Lifelong learning. Most previous works we discussed so far have only been validated on simple closed-form dataset, such as visual odometry and depth predictions are performed on the KITTI dataset. However, in an open world, the mobile agents will confront everchanging environmental factors, and moving dynamics. This will require the DNN models to continuously and coherently learn and adapt to the changes of the world. Moreover, new concepts and objects will appear unexpectedly, requiring an object discovery and new knowledge extension phase for robots.
-
New sensors: Beyond the common choice of on-board sensors, such as cameras, IMU and LIDAR, the emerging new sensors provide an alternative to construct a more accurate and robust multimodal system. New sensors including event camera[250], thermo camera[251], mm-wave device[252], radio signals[253], magnetic sensor[254], have distinct properties and data format compared to predominant SLAM sensors such as cameras, IMU and LIDAR. Nevertheless, the effective learning approaches to processing these unusual sensors are still underexplored.
-
Scalability. Both the learning based localization and mapping models have now achieved promising results on the evaluation benchmark. However, they are restricted to some scenarios. For example, odometry estimation is always evaluated in the city area or on the roads. Whether these techniques could be applied to other environments, e.g. rural area or forest area is still an open question. Moreover, existing works on scene reconstruction are restricted on single-objects, synthetic data or room level. It is worthy exploring whether these learning methods are capable of scaling to much more complex and large-scale reconstruction problems.
-
Safety, reliability and interpretability. Safety and reliability are critical to practical applications, e.g. self-driving vehicles. In these scenarios, even a small error of pose or scene estimates will cause disasters to the entire system. Deep neural networks have been long-critisized as ’black-box’, exacerbating the safety concerns for critical tasks. Some initial efforts explored the interpretability on deep models [255]. For example, uncertainty estimation[244, 245] can offer a belief metric, representing to what extent we trust our models. In this way, the unreliable predictions (with low uncertainty) are avoided in order to ensure the systems to stay safe and reliable.