Submitted:
23 August 2026
Posted:
25 August 2026
You are already at the latest version
Abstract
Worldwide image geolocalization aims to predict the geographic location of an image taken anywhere on Earth, expressed as GPS coordinates, a geographic cell, or an administrative region. The task is open-world by nature: no reference collection provides complete imagery coverage of the planet, so a model must generalize to locations it has never observed. Classical methods pursue this generalization statistically, learning visual-to-geographic mappings with classification heads, retrieval embeddings, or continuous probabilistic models. More recently, foundation models have opened a second route based on knowledge-driven reasoning, in which the location answer is generated from world knowledge internalized during pretraining, ranging from retrieval-augmented generation (RAG) to tool-using agents. This paradigm changes the mechanism by which location answers are produced. This survey provides a systematic review of worldwide image geolocalization with a focus on the foundation-model era. We introduce a two-level taxonomy that organizes classical paradigms by their output mechanism and foundation-model-era methods by the role the foundation model plays, covering RAG, reasoning, agentic, and hybrid designs. We further present a unified review of datasets and benchmarks, a cross-method comparison of reported results, and a dedicated discussion of privacy, fairness, and ethics. Finally, we outline open challenges and future directions. A continuously updated paper list is also available at https://github.com/Jia-py/Awesome-Worldwide-Image-Geolocalization.
Keywords:
image geolocalization
; worldwide geolocalization
; foundation models
; large vision-language models
; LVLMs
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.