A Survey on Deep Learning for Localization and Mapping: Towards the Age of Spatial Machine Intelligence
Deep learning based localization and mapping has recently attracted significant attention. Instead of creating hand-designed algorithms through exploitation of physical models or geometric theories, deep learning based solutions provide an alternative to solve the problem in a data-driven way. Benefiting from ever-increasing volumes of data and computational power, these methods are fast evolving into a new area tha…
Licence
OPEN
CC-BY-4.0
Authors
Changhao Chen, Bing Wang, Chris Xiaoxuan Lu, Niki Trigoni, Andrew Mar…
This work comprehensively overviews the area of deep learning for localization and mapping, and provides a new taxonomy to cover the relevant existing approaches from robotics, computer vision and machine learning communities. Learning models are incorporated into localization and mapping systems to connect input sensor data and target values, by automatically extracting useful features from raw data without any human effort. Deep learning based techniques have so far achieved the state-of-the-art performance in a variety of tasks, from visual odometry, global localization to dense scene reconstruction. Due to the highly expressive capacity of deep neural networks, these models are capable of implicitly modelling the factors such as environmental dynamics or sensor noises, that are hard to be modelled by hand, and thus are relatively more robust in real-world applications. In addition, high-level understanding and interaction are easy to perform for mobile agents with the learning based framework. The fast development of deep learning provides an alternative to solve classical localization and mapping problem in a data-driven way, and meanwhile paves the road towards a next-generation AI based spatial perception solution.
Metadata record
One description, two standard projections
Built from what the sources declared and what the gates observed.
Nothing absent has been filled in here.
1 value read out of the text by the enrichment rules and 19 links to or from other resources — a lab's files, the pages it links, the works it cites stand under their elements, marked inferred, and are kept apart in the exports.
DCMI Metadata Terms. Dublin Core has no element that separates the original file from the text extracted out of it, and none for LOM's educational characterisation. Both survive here as provenance statements and in the record itself, not in the projection.
the standard ↗
The groups below are this library's, for reading. DCMI Terms itself has no categories; each term keeps its standard name.
Deep learning based localization and mapping has recently attracted significant attention. Instead of creating hand-designed algorithms through exploitation of physical models or geometric theories, deep learning based solutions provide an alternative to solve the problem in a data-driven way. Benefiting from ever-increasing volumes of data and computational power, these methods are fast evolving into a new area that offers accurate and robust systems to track motion and estimate scenes and their structure for real-world applications. In this work, we provide a comprehensive survey, and propose a new taxonomy for localization and mapping using deep learning. We also discuss the limitations of current models, and indicate possible future directions. A wide range of topics are covered, from learning odometry estimation, mapping, to global localization and simultaneous localization and mapping (SLAM). We revisit the problem of perceiving self-motion and scene understanding with on-board sensors, and show how to solve it by integrating these modules into a prospective spatial machine intelligence system (SMIS). It is our hope that this work can connect emerging works from robotics, computer vision and machine learning communities, and serve as a guide for future researchers to apply deep learning to tackle localization and mapping problems.
The abstract, and the depositor's additional notes after it when the source has a field for them.
Where the work was published, in the source's own words: a journal with its volume and pages, a conference, an imprint.
Relations and custody
1/9
Is part ofdcterms:isPartOf
none found — the source did not declare it
The repository, book or record this resource was found inside.
Has partdcterms:hasPart
none found — the source did not declare it
What this resource is made of, when the source lists its parts.
Is version ofdcterms:isVersionOf
none found — the source did not declare it
Has versiondcterms:hasVersion
none found — the source did not declare it
Referencesdcterms:references
inferred
read its reference list, line 532citation
arXiv:2005.14165
arXiv:1906.11435
arXiv:1811.04370
arXiv:1909.03557
arXiv:1411.1509
arXiv:1908.01293
arXiv:1805.08443
arXiv:1809.01019
arXiv:1907.04011
arXiv:1906.06195
arXiv:2003.10071
arXiv:1802.03237
and 7 more, every one of them in the export
Works this one cites, when the source declares them as relations. What its text links to and its reference list cites is inferred, and stands under it apart.
Is referenced bydcterms:isReferencedBy
none found — the source did not declare it
What links to or cites this one, when a source declares it; inferred from the texts held otherwise, and set apart.
Requiresdcterms:requires
none found — the source did not declare it
What this resource needs in order to be used, when a source declares it. The files a lab works on are inferred, and stand under it apart.
Is required bydcterms:isRequiredBy
none found — the source did not declare it
What needs this resource, when a source declares it. The lab a component belongs to is inferred, and stands under it apart.
Provenancedcterms:provenance
this engine
source
conversion
latexml-html
Retrieved from arXiv on 2026-10-09 in response to the search string “(all:"artificial intelligence" OR all:"machine learning" OR all:"generative AI" OR all:"deep learning" OR all:"reinforcement learning" OR all:"large language model") AND (all:"AI concepts" OR all:"types of AI" OR all:"AI fundamentals" OR all:"recognizing AI" OR all:"recognising AI" OR all:"general versus narrow AI" OR all:"narrow AI" OR all:"general AI" OR all:"machine intelligence" OR all:"AI strengths and weaknesses" OR all:"traditional software" OR all:"rule-based systems" OR all:"introduction to AI" OR all:"introduction to artificial intelligence" OR all:"artificial intelligence introduction" OR all:"AI primer" OR all:"foundations of artificial intelligence" OR all:"overview of AI" OR all:"understanding AI" OR all:"history of AI" OR all:"AI essentials" OR all:"AI terminology" OR all:"metaphors for AI" OR all:"AI fundamental concepts" OR all:"AI key concepts" OR all:"philosophy of AI" OR all:"critical AI literacy")”. arXiv served the resource and is not asserted to be its publisher or author.
Text extracted from latexml-html to Markdown by arxiv-html; the original is retained unchanged beside it.
Where it was collected from, what was converted, and what container it came out of — the custody statements that would otherwise be mistaken for authorship.
IEEE 1484.12.1 Learning Object Metadata. LOM has no element for an SPDX identifier or a licence URI, so both are written into 6.3 Rights.Description. Flattening this record into simple Dublin Core would lose more again, which is why the two projections exist side by side rather than one being generated from the other.
the standard ↗
Deep learning based localization and mapping has recently attracted significant attention. Instead of creating hand-designed algorithms through exploitation of physical models or geometric theories, deep learning based solutions provide an alternative to solve the problem in a data-driven way. Benefiting from ever-increasing volumes of data and computational power, these methods are fast evolving into a new area that offers accurate and robust systems to track motion and estimate scenes and their structure for real-world applications. In this work, we provide a comprehensive survey, and propose a new taxonomy for localization and mapping using deep learning. We also discuss the limitations of current models, and indicate possible future directions. A wide range of topics are covered, from learning odometry estimation, mapping, to global localization and simultaneous localization and mapping (SLAM). We revisit the problem of perceiving self-motion and scene understanding with on-board sensors, and show how to solve it by integrating these modules into a prospective spatial machine intelligence system (SMIS). It is our hope that this work can connect emerging works from robotics, computer vision and machine learning communities, and serve as a guide for future researchers to apply deep learning to tackle localization and mapping problems.
The abstract, and the depositor's additional notes after it as a second LangString when the source has a field for them.
Role, entity and date per declared contribution. A role outside LOM's vocabulary is reported in the entry's description instead.
3 Meta-metadata
4/4
Identifier3.1
this engine
resource_id
URI: tag:aim-pro.eu,2026:oer/43dfc5065704/record
The identifier of this metadata record — the resource's own, with /record after it, because the record is a description of the resource and not the resource.
Contribute3.2
this engine
source
AIM-PRO WP3 OER harvester (arXiv) — creator
Who generated this record and when — a statement about the record, not about the resource.
Metadata schema3.3
LOMv1.0
aimpro-oer-profile/1
LOMv1.0, and the profile this was built against.
Language3.4
language gate
read the declared field
en
4 Technical
3/7
Format4.1
conversion
latexml-html
text/markdown
One value per form held: the original as the source published it, and the Markdown this engine extracted.
Size4.2
conversion
latexml-html
155694
Bytes. The original's, because the resource is the file and not our conversion of it.
Location4.3
this engine
resource_id
https://arxiv.org/abs/2006.12567
Where the source serves it.
Requirement4.4
not collected — this library does not fill it
Software or hardware needed to use it. No source declares it.
Installation remarks4.5
not collected — this library does not fill it
No source declares it.
Other platform requirements4.6
not collected — this library does not fill it
No source declares it.
Duration4.7
not collected — this library does not fill it
Playing time, for audio and video. The corpus holds neither.
5 Educational
1/11
Interactivity type5.1
not collected — this library does not fill it
Active, expositive or mixed. A judgement about how the resource is used; no source declares it.
Yes unless the licence reserves nothing — attribution is a restriction. The conditions after the dash are the licence gate's reading; the export carries LOM's bare term.
The SPDX id, the licence URI and the conditions. LOM has no element for any of the three, so this is where they survive.
7 Relation
0/2
Kind7.1
none found — the source did not declare it
Resource7.2
inferred
read its reference list, line 532citation
references: arXiv:2005.14165
references: arXiv:1906.11435
references: arXiv:1811.04370
references: arXiv:1909.03557
references: arXiv:1411.1509
references: arXiv:1908.01293
references: arXiv:1805.08443
references: arXiv:1809.01019
references: arXiv:1907.04011
references: arXiv:1906.06195
references: arXiv:2003.10071
references: arXiv:1802.03237
and 7 more, every one of them in the export
The container a file was found inside, and the Markdown extracted from the original. What a lab requires, and the lab a component belongs to, are inferred and stand apart.
8 Annotation
0/3
Entity8.1
not collected — this library does not fill it
Comments on the resource's educational use, by whoever made them. The platform's review grades competencies, which are classification (9), and writes no comment here.
Date8.2
not collected — this library does not fill it
Description8.3
not collected — this library does not fill it
9 Classification
0/4
Purpose9.1
not collected — this library does not fill it
Empty in the record: no source declares a competency. The alignment reads the resource for them and stands beside the record, never in it, and a taxon path derived from the search string that found it would be a claim about the query.
Taxon path9.2
not collected — this library does not fill it
Where the competency framework goes. Empty in the record for the reason above.
Description9.3
not collected — this library does not fill it
Keyword9.4
not collected — this library does not fill it
What could not be established3
Where the source's metadata could not be carried
over as it was — missing, contradictory, with no matching term in the
standard, restructured, or taken from the repository — and what was done
instead. Without these notes, an empty element would look like something
the harvester missed.
Status
Field
Why
Not available
rights_holder
no rights holder is named at source; the licence is recorded without one rather than attributed to the platform that served it
Not available
publisher
the source named no publisher of the work; where it was collected from is recorded as collection provenance instead, which is a different claim
Not available
educational
the source declared no educational metadata — no resource type, audience, context, difficulty or learning time. Nothing here estimates them