Submitted:
09 September 2026
Posted:
10 September 2026
You are already at the latest version
Abstract
Visual geolocalization aims to estimate the geographic location where an image was captured solely from its visual content. By enabling location inference without GPS signals or geographic metadata, it supports a wide range of applications, including autonomous navigation, urban computing, and location-based services. Over the past decade, visual geolocalization has witnessed remarkable progress driven by advances in visual representation learning, large-scale geo-tagged datasets, and, more recently, large vision-language models (LVLMs), giving rise to increasingly diverse methodological paradigms and problem settings. However, existing surveys primarily focus on specific application scenarios or traditional geolocalization approaches, leaving recent developments in reasoning-based geolocalization, multimodal foundation models, and emerging benchmarks insufficiently covered. In this survey, we provide a comprehensive review of visual geolocalization from the perspectives of both datasets and methodologies. We first present a unified taxonomy that organizes existing methods into three representative paradigms: retrieval-based, prediction-based, and reasoning-based localization. Building upon this taxonomy, we systematically review representative algorithms, benchmark datasets, evaluation protocols, and practical applications, while highlighting the evolution of the field and the relationships among different methodological paradigms. Furthermore, we discuss current challenges and identify promising future research directions, with particular emphasis on multimodal reasoning, agentic geolocalization, and open-world localization. We hope this survey provides a unified understanding of the rapidly evolving field and serves as a valuable resource for both newcomers and experienced researchers.
Keywords:
visual geolocalization
; visual place recognition
; cross-view geolocalization
; image geolocalization
; survey
1. Introduction
Location information serves as a fundamental source of spatial context for intelligent systems, supporting decision-making across a wide range of real-world applications, including autonomous navigation, urban computing, and disaster response [1,2]. Although modern positioning technologies such as GPS provide reliable localization under favorable conditions, they are often unavailable or unreliable for web-sourced imagery, historical photographs, or images shared through social media, where geographic metadata is missing or inaccessible [3,4]. Consequently, enabling machines to infer locations directly from visual observations has become an increasingly important research problem.
Motivated by this challenge, Visual Geolocalization has emerged as a fundamental task that aims to estimate the geographic location where an image was captured solely from its visual content [5,6,7,8]. Unlike conventional localization systems that depend on external sensors or metadata, Visual Geolocalization exploits diverse geographic cues embedded in images, including landmarks, architectural styles, vegetation, road layouts, terrain, weather conditions, and cultural characteristics. As illustrated in Figure 1, existing approaches can be broadly categorized into three representative methodological paradigms: retrieval-based, prediction-based, and reasoning-based localization. Benefiting from its ability to localize images without auxiliary geographic information, Visual Geolocalization has found broad applications in tourism recommendation [9,10,11], location-based services, digital forensics, social media analysis, and crisis response [12,13].
As shown in Figure 2, Visual Geolocalization has witnessed sustained growth in recent years, reflecting the increasing research interest in this field. Alongside this rapid development, the methodological landscape has evolved substantially. Two major methodological paradigms have evolved in parallel. Retrieval-based methods estimate image locations by matching a query image against large geo-referenced databases, while prediction-based methods directly infer geographic locations through classification, regression, or generative modeling. With the rapid development of convolutional neural networks (CNNs), Vision Transformers (ViTs), and vision-language pretraining such as CLIP, both paradigms have achieved substantial improvements in representation learning, localization accuracy, and scalability. Nevertheless, these methods predominantly rely on learned correlations between visual appearance and geographic locations, making them vulnerable to visually ambiguous scenes, limited reference coverage, and significant distribution shifts.
More recently, the emergence of Large Vision-Language Models (LVLMs) [14,15,16,17] has introduced a new reasoning-based paradigm for Visual Geolocalization. Rather than relying solely on visual similarity or learned geographic priors, these methods progressively infer locations by integrating scene understanding, world knowledge, commonsense reasoning, and external tools such as web search, and map services [18,19,20,21,22,23,24]. This emerging paradigm significantly enhances model interpretability, adaptability, and open-world generalization, transforming Visual Geolocalization from a purely appearance-driven recognition task into a geographically grounded reasoning problem.
Despite the rapid progress of the field, existing surveys remain fragmented in their coverage. Most primarily focus on traditional retrieval-based or prediction-based methods, specific application scenarios, or cross-view localization, while recent advances in LVLM-driven reasoning, agentic geolocalization, and newly introduced datasets and benchmarks have not yet been systematically reviewed. Moreover, the rapid expansion of methodological paradigms has blurred the overall landscape of Visual Geolocalization, making it increasingly difficult for researchers to understand the relationships, strengths, and limitations of different approaches. These observations highlight the need for a comprehensive and up-to-date survey that unifies recent developments under a coherent methodological framework.
To address this gap, this survey presents a comprehensive review of Visual Geolocalization from the dual perspectives of methodologies and datasets. Unlike previous surveys, we organize existing methods into three representative paradigms, retrieval-based, prediction-based, and reasoning-based localization, providing a unified perspective on the evolution of the field while systematically reviewing representative methods, benchmark datasets, evaluation protocols, practical applications, and emerging research directions. We hope this survey serves as a valuable resource for researchers and practitioners, facilitating the development of Visual Geolocalization systems that are more accurate, robust, interpretable, and trustworthy.
The main contributions of this survey are summarized as follows:
- A unified taxonomy of visual geolocalization. We present a comprehensive taxonomy that categorizes existing methods into retrieval-based, prediction-based, and reasoning-based paradigms, providing a coherent framework for understanding the rapidly evolving methodological landscape.
- A systematic review of datasets and methodologies. We comprehensively review representative datasets, benchmark protocols, and methodological developments, covering both traditional approaches and the latest LVLM-driven reasoning frameworks.
- Insights into future research directions. We discuss current limitations and identify promising research opportunities across datasets, methodologies, and foundation-model-based geolocalization, with particular emphasis on multimodal reasoning, agentic systems, and open-world localization.
2. Background
This section introduces the fundamental concepts of Visual Geolocalization that will be used throughout this survey, including the unified problem formulation, common data organizations, representative methodological paradigms, and widely adopted evaluation protocols.
2.1. Unified Problem Formulation
Visual Geolocalization aims to estimate the geographic location associated with a visual observation. Given a visual input I and optional contextual information C, the task can be generally formulated as
where denotes a Visual Geolocalization model and y represents the predicted geographic target. The above formulation intentionally covers diverse task settings. Depending on the application, the visual input I may consist of a ground-level image, street-view panorama, video sequence, satellite image, UAV image, or a pair of images from different viewpoints. The contextual information C is optional and may include a geo-referenced image database, metadata, textual descriptions, maps, retrieved candidates, or reasoning history. Likewise, the prediction target y may take different forms, including geographic coordinates, geocells, administrative regions, landmark identities, addresses, reference images, ranked candidate lists, or natural-language location descriptions.
2.2. Data Organization
The organization of data largely determines the learning paradigm and evaluation protocol adopted by a Visual Geolocalization system. Existing datasets can generally be grouped into two complementary organizations, as illustrated in Figure 3.
Structured retrieval datasets explicitly establish query-reference relationships by providing query images together with geo-referenced galleries, positive pairs, negative pairs, or cross-view correspondences. Such datasets naturally support retrieval-based localization.
Geospatially annotated datasets directly associate visual observations with geographic labels, including coordinates, administrative regions, landmarks, textual descriptions, question-answer pairs, or reasoning traces. These datasets primarily support prediction-based and reasoning-based localization.
These two data organizations naturally correspond to different learning paradigms and evaluation protocols, which will be discussed in the following sections.
2.3. Methodological Paradigms
Visual Geolocalization methods differ primarily in how geographic locations are estimated. Accordingly, this survey groups existing approaches into three representative methodological paradigms.
2.3.1. Retrieval-Based Localization
Retrieval-based methods estimate image locations by matching a query image against a geo-referenced database. Instead of directly predicting geographic locations, these methods infer the target location from retrieved reference samples with similar visual or semantic characteristics.
2.3.2. Prediction-Based Localization
Prediction-based methods directly infer geographic locations from visual observations without explicit database retrieval during inference. Depending on the prediction target, the output may correspond to geographic coordinates, discrete geographic cells, administrative regions, landmarks, or other location labels.
2.3.3. Reasoning-Based Localization
Reasoning-based methods extend Visual Geolocalization beyond appearance matching by integrating scene understanding, semantic reasoning, world knowledge, and external information sources. Recent large vision-language models further enhance this paradigm through multi-step reasoning and tool interaction, enabling more interpretable and adaptive location inference.
These paradigms emphasize different strategies for geographic inference, although modern Visual Geolocalization systems often integrate multiple paradigms within a unified framework. Building upon this taxonomy, the following sections first review representative methods under each paradigm, followed by benchmark datasets, evaluation protocols, applications, and future research directions.
2.4. Evaluation Protocols
Evaluation protocols in Visual Geolocalization are closely related to the organization of benchmark datasets and the form of model outputs. Existing evaluation metrics can be broadly grouped into three categories: ranking-based retrieval metrics, distance-based localization metrics, and category-based localization metrics.
2.4.1. Ranking-Based Retrieval Metrics
Structured retrieval benchmarks explicitly provide a query set together with a geo-referenced reference set, also referred to as a database or gallery. The objective is to retrieve the correct reference images or place candidates for each query. Consequently, performance is typically evaluated using ranking-based metrics such as Recall@K and mean Average Precision (mAP), which measure whether the correct reference appears among the top-ranked retrieval results. In visual place recognition and cross-view geolocalization benchmarks, variants such as top-K recall, hit rate, and rank-based localization accuracy are also commonly adopted. Although these metrics may share similar names with classification metrics, they evaluate the quality of a ranked reference list rather than the correctness of a discrete category prediction.
2.4.2. Distance-Based Localization Metrics
When the prediction target is geographic coordinates, localization performance is commonly measured by the great-circle distance between the predicted coordinate and the ground-truth coordinate . The Haversine distance is computed as
where , , and R denotes the Earth’s radius. Localization performance is commonly reported as accuracy under predefined distance thresholds, such as street-, city-, region-, country-, and continent-level accuracy, indicating the percentage of predictions whose localization errors fall below each threshold.
2.4.3. Category-Based Localization Metrics
For methods formulated as classification problems, the prediction target is usually a discrete geographic category, such as a geocell, administrative region, city, country, landmark, or point of interest. These methods are typically evaluated using category-level metrics, including top-1 accuracy, top-K accuracy, precision, recall, and F1 score, depending on whether the task is defined as single-label classification, multi-label prediction, or hierarchical geographic recognition. In some settings, discrete predictions can also be converted into representative coordinates, such as cell centroids, region centers, or landmark locations, and further evaluated using distance-based localization metrics.
Recently, reasoning-based Visual Geolocalization has begun to introduce additional evaluation dimensions, such as reasoning faithfulness, answer validity, uncertainty estimation, and tool-use correctness. However, these protocols remain under active development and have not yet been standardized.
3. Methods: A Survey
This section presents a comprehensive review of Visual Geolocalization methods organized under the three-paradigm taxonomy introduced in Section 2 and illustrated in Figure 4. We begin with a methodological overview that traces the evolution of the field (Section 3.1), followed by detailed discussions of retrieval-based (Section 3.2), prediction-based (Section 3.3), and reasoning-based (Section 3.4) localization. We conclude with a discussion that compares the strengths and limitations of the three paradigms (Section 3.5).
3.1. Method Overview
Visual Geolocalization has undergone a notable methodological shift over the past two decades. Since the earliest days of the field, retrieval-based localization has been a foundational paradigm, in which a query image is matched against a geo-referenced reference collection to infer its location. Early systems relied on handcrafted visual descriptors such as GIST and VLAD for this matching process [25,26,27], while the rise of deep learning transformed the quality of the learned embedding space, replacing engineered features with data-driven representations [6,28]. In parallel, deep learning also enabled a second, prediction-based paradigm that directly maps visual content to geographic labels or coordinates through classification, regression, or generative modeling, bypassing the need for an explicit reference database during inference. Both paradigms have been substantially advanced by convolutional neural networks (CNNs) [77,78], Vision Transformers (ViTs), and vision-language pretraining such as CLIP [79].
More recently, the emergence of Large Vision-Language Models (LVLMs) [14,15,16,17] has introduced a third paradigm: reasoning-based localization. Rather than relying solely on visual similarity or learned geographic priors, these methods progressively infer locations by integrating scene understanding, world knowledge, commonsense reasoning, and external tools such as web search and map services [18,20,21,23,24]. This evolution transforms Visual Geolocalization from a purely appearance-driven matching task into a geographically grounded reasoning problem.
As illustrated in Figure 4, the three paradigms differ primarily in how geographic locations are estimated: retrieval-based methods infer locations from matched references, prediction-based methods produce direct geographic outputs, and reasoning-based methods derive locations through multi-step inference. While modern Visual Geolocalization systems increasingly integrate multiple paradigms within a unified framework, we discuss each paradigm independently in the following subsections to highlight their respective design principles, representative methods, technical characteristics, and limitations.
3.2. Retrieval-Based Localization
Retrieval-based methods estimate image locations by matching a query image against a geo-referenced database. Instead of directly predicting geographic coordinates, these methods infer the target location from retrieved reference samples with similar visual or semantic characteristics, as illustrated in Figure 5. Depending on the modality of the query and reference data, existing retrieval-based methods can be further divided into image-to-image retrieval and cross-modal retrieval, as shown in Figure 4.
3.2.1. Image-to-Image Retrieval
Image-to-image retrieval methods learn a visual embedding space in which images captured at nearby locations are mapped to similar representations. Location is then inferred by aggregating the geographic information of the top-ranked reference images. Based on the viewpoint relationship between query and reference images, this line of work can be further subdivided into same-view and cross-view matching.
Same-view retrieval. Same-view retrieval matches query and reference images captured from the same viewpoint modality, most commonly ground-level photos. Before the deep learning era, IM2GPS [25] established the foundational framework by comparing a query image against over six million GPS-tagged Flickr images using handcrafted global scene descriptors. Yagcioglu et al. [26] combined GIST [80] and Tiny Image [81] features with deformable spatial pyramid matching for city-scale localization, while Torii et al. [27] improved robustness to appearance changes by synthesizing virtual views and encoding local descriptors with VLAD [82,83]. Vo et al. [6] revisited this framework in the deep learning era, demonstrating that features trained with classification objectives outperform contrastive embeddings for nearest-neighbor matching (as shown in Figure 6). To improve robustness against viewpoint variations, GeoWarp [28] introduced a trainable warping module for dense local feature matching, serving as an effective re-ranking mechanism. More recently, Waheed et al. [31] leveraged vision-language model guidance for planet-scale visual place recognition, and Pan et al. [30] proposed a three-stage framework leveraging maximal clique theory for large-scale remote sensing image geo-localization. GeoRanker [29] further advanced retrieval precision by employing a large multimodal model to perform distance-aware ranking over retrieved candidates, bridging retrieval and reasoning.
Cross-view retrieval. Cross-view retrieval matches images captured from different viewpoint modalities, most commonly ground-to-aerial or street-view-to-satellite matching. This setting is motivated by the dense and wide geographic coverage provided by geo-referenced aerial and satellite imagery. Early studies primarily focused on learning robust representations and reducing the large geometric discrepancy across viewpoints. Lin et al. [32] introduced an early deep learning approach for ground-to-aerial geo-localization, while Vo and Hays [34] explored deep CNN architectures and orientation-aware matching strategies for jointly localizing and orienting street views using overhead imagery. CVM-Net [33] employed NetVLAD for cross-view feature aggregation, and Cai et al. [35] introduced a hard exemplar reweighting triplet loss to improve metric learning. To further reduce geometric discrepancies, Polar SAFA [36] combined polar transformation with spatial-aware feature aggregation.
Another line of research seeks to bridge the viewpoint gap through generative modeling or extend cross-view localization to richer visual settings. SelectionGAN [37] employed multi-channel attention selection with cascaded semantic guidance for cross-view image translation, while GPG2A [40] incorporated geometry and text guidance into diffusion-based ground-to-aerial image synthesis. Beyond conventional ground-to-satellite matching, PCL [38] integrated UAV-to-satellite view synthesis with cross-view geo-localization, while GAMa [39] extended cross-view geo-localization from individual images to ground video sequences.
Transformer-based representation learning has subsequently become increasingly prominent. L2LTR [41] introduced a layer-to-layer Transformer with positional encoding and self-cross attention to model cross-view geometric configurations and inter-layer dependencies. TransGeo [42] proposed a pure Transformer-based architecture for cross-view matching, while Sample4Geo [43] introduced hard-negative sampling strategies within a contrastive learning framework. SIGN [44] further introduced a saliency-aware global-local framework with Saliency-Aware Partition for fine-grained cross-view representation learning.
More recently, cross-view geo-localization has begun to move beyond purely visual correspondence by incorporating language, semantic knowledge, and foundation models. Ye et al. [45] introduced text-based cross-view localization, using natural-language scene descriptions to retrieve corresponding satellite images or OSM data. GLEAM [46] unified multi-view and multimodal matching while further combining cross-view correspondence prediction with explainable reasoning. GeoBridge [47] developed a semantic-anchored multi-view foundation model that bridges multi-view representations through textual descriptions, supporting bidirectional cross-view matching and language-to-image retrieval.
3.2.2. Cross-Modal Retrieval
Cross-modal retrieval moves beyond visual-to-visual matching by aligning images with non-visual geographic representations such as text, coordinates, or addresses. This direction has been primarily driven by vision-language models. StreetCLIP [48] adapted CLIP [79] for the geolocation domain by fine-tuning on image-caption pairs derived from geographic data, enabling zero-shot retrieval where the query image is matched against textual location descriptions. GeoCLIP [8] extended this paradigm by treating GPS coordinates as a learnable modality, constructing a shared embedding space between images and sinusoidally encoded coordinates. AddressCLIP [50] generalized this direction to city-scale localization by aligning image features with structured address strings. Jia et al. [49] proposed a concept-aware global image-GPS alignment framework for interpretable geo-localization, GT-Loc [51] unified temporal and spatial grounding in a joint embedding space, and GEOMR [52] integrated image geographic features with human reasoning knowledge. Unlike traditional systems that require storing and searching millions of reference images, these CLIP-based approaches perform retrieval in a highly compressed, parametric embedding space, redefining the retrieval target from visual instances to semantic or coordinate-level representations.
Technical characteristics. The performance of retrieval-based approaches is shaped by three factors: the structure of the learned embedding space, the modality and granularity of the reference data, and the design of the retrieval strategy. Embeddings are typically learned using contrastive, triplet, or cross-modal alignment losses, often augmented with hard negative mining [43] to improve discrimination. For cross-view matching, polar transformations and generative image synthesis are commonly employed to bridge the viewpoint gap. With the introduction of vision-language models, retrieval is increasingly guided by soft prompts or region-level semantics, enabling more abstract and flexible representations.
Strengths and limitations. Retrieval-based methods offer a conceptually straightforward approach by grounding predictions in reference data. They support fine-grained localization and benefit from the generalization ability of pre-trained encoders. However, their performance is sensitive to the density and quality of the reference database, with reduced effectiveness in visually ambiguous or sparsely covered areas. Interpretability also remains limited because predictions are guided by similarity rather than step-by-step reasoning.
3.3. Prediction-Based Localization
Prediction-based methods directly infer geographic locations from visual observations without explicit database retrieval during inference. Depending on the form of the prediction target, these methods can be divided into discrete location prediction, which formulates geolocalization as a classification problem over predefined geographic regions, and continuous location prediction, which models the geographic output as continuous coordinates through regression or generative processes.
3.3.1. Discrete Location Prediction
Discrete location prediction divides the Earth’s surface into a finite set of geographic regions (e.g., S2 cells, administrative boundaries) and trains a classifier to predict the region label for a given input image, as illustrated in Figure 7.
Main approaches. PlaNet [5] introduced a foundational approach by formulating geolocation as a classification task over discretized geographic cells (as shown in Figure 8). It employed a CNN trained on millions of geo-tagged images to predict a probability distribution over geographic regions, demonstrating that visual content alone can achieve competitive localization accuracy. To improve resolution, CPlaNet [7] proposed a combinatorial partitioning scheme that generates fine-grained geoclasses by intersecting multiple coarse-grained divisions, allowing more precise localization while preserving training sample density. Johns et al. [53] explored multi-modal geolocation estimation using deep neural networks, and ISNs [54] incorporated hierarchical spatial structures and scene classification, leveraging coarse-to-fine spatial cues and environmental context to guide region prediction. MvMF [55] introduced a probabilistic model using a mixture of von Mises-Fisher distributions that better respects the Earth’s spherical geometry and enables smoother location predictions. SemP [56] focused on interpretability by deriving region partitions from real-world geographic entities (e.g., roads, rivers, cities) using OpenStreetMap data.
Several works have enriched the classification framework with auxiliary modalities. [57] aligned image features with human-written guidebook texts via attention mechanisms for geolocation via guidebook grounding. TransLocator [58] and GeoDecoder [59] employed dual-branch architectures and hierarchical attention to process semantic segmentation maps and learn region-specific features across geographic levels. CityGuessr [61] extended the classification framework to the video domain, integrating soft text-label supervision for global-scale city-level prediction. GeoToken [60] reformulated hierarchical geolocalization as a next-token prediction task, unifying multi-level geographic reasoning within an autoregressive framework. Bianco et al. [62] enhanced worldwide geolocation by ensembling satellite-based ground-level attribute predictors.
Technical characteristics. The defining characteristic of discrete prediction methods is the representation of the geographic target space as a set of discrete regions. Design choices center on the partitioning strategy and the granularity of the grid. Finer partitions provide higher spatial resolution but often lead to class imbalance and increased ambiguity near boundaries. To mitigate this, some methods adopt hierarchical [54] or semantically informed [56] region definitions. Many approaches also incorporate multimodal inputs such as paired texts [57] or segmentation maps [58] to enhance spatial discrimination.
3.3.2. Continuous Location Prediction
Continuous location prediction bypasses the quantization limits of fixed geographic grids by modeling the geographic output as continuous coordinates. These methods either learn coordinate-aware representations that preserve spherical geometry or formulate geolocation as a conditional generation task.
Main approaches. Sphere2Vec [63] proposed a general-purpose location representation learning method over the spherical surface, designing encoder functions that preserve the distance and topology of geographic coordinates for large-scale geospatial predictions. This representation enables downstream models to better distinguish locations that are geographically close but fall on opposite sides of a cell boundary. More recently, generative approaches have emerged to model the continuous posterior distribution of coordinates directly. Dufour et al. [64] formulated geolocation as a denoising process on the spherical manifold, training a diffusion model to iteratively refine a noisy distribution into precise latitude and longitude coordinates, as illustrated in Figure 9. LocDiff [65] extended this direction by diffusing in the Hilbert space, which maps the spherical surface to a one-dimensional representation that preserves spatial locality while enabling efficient generative modeling. Instead of selecting from a finite set of predefined classes, these models learn to sample from a continuous coordinate distribution, effectively bypassing the resolution constraints of grid-based classification.
Technical characteristics. Continuous prediction methods focus on modeling the complex, multi-modal posterior distribution of geographic coordinates. For representation-based methods such as Sphere2Vec [63], the technical design centers on encoder functions that respect spherical geometry and preserve spatial relationships. For generative methods [64,65], the design involves formulating noise schedules and iterative sampling processes on appropriate geometric spaces. Model outputs range from continuous coordinate embeddings to refined latitude-longitude samples, providing finer resolution than discrete classification.
Strengths and limitations. Prediction-based methods are efficient and scalable, requiring no external reference database during inference. Their end-to-end nature simplifies training and deployment, and probabilistic outputs enable uncertainty estimation. However, discrete classification methods are limited by the resolution of their discretized label space, and the fixed label space can be suboptimal in regions with uneven data distributions. While continuous generative approaches effectively resolve this quantization issue by enabling continuous prediction, they may introduce trade-offs in inference latency due to the iterative sampling process required. Furthermore, the outputs of prediction-based methods generally lack interpretability in terms of specific geographic cues.
3.4. Reasoning-Based Localization
Reasoning-based methods extend Visual Geolocalization beyond appearance matching and direct prediction by integrating scene understanding, semantic reasoning, world knowledge, and external information sources. Enabled primarily by LVLMs, these methods reformulate geolocation as a step-by-step inference task in which the model identifies visual cues, generates location hypotheses, and justifies predictions in a human-interpretable manner. As shown in Figure 4, reasoning-based methods can be further divided into chain-of-thought reasoning, retrieval-augmented reasoning, and agentic reasoning.
3.4.1. Chain-of-Thought Reasoning
Chain-of-thought (CoT) reasoning methods prompt or fine-tune LVLMs to produce explicit, step-by-step reasoning traces that connect visual observations to geographic conclusions, as illustrated in Figure 10.
Main approaches. GeoReasoner [18] introduced a multi-stage framework that fine-tunes a vision-language model using human reasoning patterns extracted from geolocalization games, enabling a comprehensive reasoning pipeline that generates structured justifications and verifies them against candidate regions (as shown in Figure 11). GeoCoT [20] advanced the paradigm by introducing explicit CoT prompting combined with GeoComp, a large-scale dataset collected from human gameplay, enabling direct supervision at the reasoning trace level. GeoChain [66] proposed a multimodal chain-of-thought framework that decomposes geographic reasoning into a sequence of grounded visual-linguistic steps. GeoLocSFT [67] explored efficient visual geolocation via supervised fine-tuning of multimodal foundation models, demonstrating that lightweight adaptation can elicit strong reasoning behavior. GaGA [22] extended the reasoning process into the interactive domain by incorporating iterative correction and user feedback, supported by a large image-text dataset that promotes explainability and response adaptability. GLOBE [24] introduced a reinforcement learning strategy based on GRPO [84] to enhance both recognition and reasoning capabilities in LVLMs, applying geolocation-aware reward signals to align model behavior with spatial accuracy. Wang et al. [68] proposed the GRE suite, which combines fine-tuned vision-language models with enhanced reasoning chains for geo-localization inference. GAEA [21] reframed the task as a visual question-answering problem, enriching the reasoning process with language-based explanation supervision derived from OpenStreetMap attributes.
Technical characteristics. CoT reasoning methods are typically grounded in LVLMs that have been pre-trained on vast multimodal corpora, equipping them with rich world knowledge. These models can associate detected visual elements, such as language scripts, architectural styles, vegetation, or cultural symbols, with geographic concepts, grounding their predictions in structured reasoning trajectories. Key technical elements include entity recognition, visual-textual alignment, spatial filtering, and supervision based on annotated reasoning trajectories. Training strategies include supervised fine-tuning on human-generated reasoning traces [18,23] and reinforcement learning with geolocation-aware rewards [24]. The final outputs go beyond location coordinates to include human-readable explanations, ranked candidates, or structured reasoning paths.
3.4.2. Retrieval-Augmented Reasoning
Retrieval-augmented reasoning methods combine a retrieval system that provides contextual evidence with a reasoning module (typically an LVLM) that synthesizes this evidence to infer the final location. This design bridges the gap between recognition (visual matching) and cognition (semantic reasoning), offering a flexible solution for handling the complexity of global-scale geolocalization.
Main approaches. Img2Loc [69] pioneered this direction by reformulating geolocalization as a text generation task. It uses an image retrieval system to fetch relevant geo-coordinates and visual references, which are then constructed into a prompt for a large multimodal model that synthesizes the retrieved context to generate the final coordinates. G3 [70] introduced an adaptive framework designed to tackle the visual aliasing problem, in which visually similar but geographically distant locations are confused. G3 first retrieves a set of candidate locations and then employs a large multimodal model to perform detailed visual reasoning and verification, adaptively engaging the reasoning module only for ambiguous samples. Cheng et al. [71] proposed a test-time scaling approach that guides multimodal LLMs for efficient visual place recognition without fine-tuning, leveraging retrieved context to guide inference. These methods implement a cascaded inference mechanism: by progressively reducing the search space from the entire globe to a small set of candidates, they enable the application of computationally intensive reasoning models on a manageable subset of data.
Technical characteristics. The defining characteristic of retrieval-augmented reasoning is its multi-stage hierarchical architecture, which decomposes the geolocation problem into retrieval and reasoning sub-tasks. The retrieval module provides a set of visual context candidates, which are then analyzed by the reasoning module to resolve ambiguities that pure visual similarity cannot handle. Technically, the system transitions from a visual similarity space to a semantic verification space, combining the coverage of retrieval with the precision and interpretability of reasoning.
3.4.3. Agentic Reasoning
Agentic reasoning methods represent the most recent frontier, in which Visual Geolocalization is formulated as a multi-agent or tool-augmented inference process. These systems decompose localization into specialized roles, such as perception, knowledge retrieval, hypothesis generation, and verification, and orchestrate multiple agents or tool calls to collaboratively solve the geolocation task.
Main approaches. NAVIG [23] synthesized earlier trends by modularizing the reasoning pipeline into three components: REASONER (generating hypotheses), SEARCHER (retrieving knowledge from external resources), and GUESSER (producing the final prediction), trained on NAVIcues, a dataset with expert reasoning traces. smileGeo [73] introduced a multi-agent collaborative framework based on swarm intelligence, in which multiple large vision-language models cooperate to refine geolocation predictions through debate and consensus. LocationAgent [74] proposed a hierarchical agent that decouples reasoning strategy from evidence retrieval, separating parametric knowledge from external evidence to reduce hallucination. GraphGeo [75] employed heterogeneous graph neural networks to model the interactions among multiple agents in a debate framework for visual geo-localization. GeoVista [72] introduced web-augmented agentic visual reasoning, enabling the system to actively search the web for geographic evidence during inference. Ji et al. [76] proposed a reinforced parallel map-augmented agent that leverages map-based tools and parallel reasoning paths to improve geolocalization accuracy.
Technical characteristics. Agentic reasoning systems are characterized by their modular, multi-step, and interactive nature. They typically incorporate tool-use capabilities (such as web search, map services, and image retrieval), multi-agent communication protocols, and feedback loops that allow iterative refinement. The architectural complexity is higher than that of CoT or retrieval-augmented methods, but the modularity enables more flexible and transparent inference, as each agent’s contribution can be inspected and verified.
Strengths and limitations. Reasoning-based approaches offer strong interpretability, modularity, and improved generalization to ambiguous or previously unseen scenes. By decomposing the localization process into human-readable steps, they facilitate error analysis and enable interactive geolocation scenarios. Furthermore, by integrating external knowledge sources or linguistic priors, these methods can perform effectively even in sparsely covered regions where retrieval-based models typically struggle. However, reasoning-based models often require substantial computational resources and careful prompt engineering or pipeline design. Their effectiveness depends heavily on the quality of training data and supervision strategy; inadequate or noisy reasoning traces can cause even powerful LVLMs to default to shallow heuristics rather than performing true spatial inference. Additionally, the lack of standardized evaluation protocols for multi-step reasoning remains an open challenge.
3.5. Discussion
The three paradigms reviewed above emphasize different strategies for geographic inference, and each carries distinct strengths and limitations.
3.5.1. Retrieval vs. Prediction
Retrieval-based methods ground predictions in real-world reference data, supporting fine-grained localization and benefiting from the generalization ability of pre-trained encoders. However, their performance is bounded by the density and coverage of the reference database, and they struggle in sparsely covered or visually ambiguous regions. Prediction-based methods, by contrast, encode geographic knowledge directly into model parameters and require no external database during inference, making them efficient and scalable. Their main limitations are the resolution constraints of discrete label spaces and the lack of interpretability in model outputs. While continuous prediction methods [64,65] mitigate quantization issues, they introduce additional inference complexity.
3.5.2. From Recognition to Reasoning
The emergence of reasoning-based methods marks a paradigm shift from appearance-driven recognition to geographically grounded reasoning. By leveraging LVLMs, these methods integrate visual perception with world knowledge, commonsense reasoning, and external tools, enabling more interpretable and adaptive location inference. Retrieval-augmented and agentic reasoning further combine the coverage of retrieval with the precision and interpretability of reasoning, representing a promising direction for unified Visual Geolocalization systems.
3.5.3. Convergence of Paradigms
Modern Visual Geolocalization systems increasingly integrate multiple paradigms within a unified framework. For example, GeoRanker [29] combines retrieval with reasoning-based re-ranking, G3 [70] couples retrieval with adaptive reasoning verification, and NAVIG [23] integrates reasoning with external knowledge retrieval. This convergence suggests that the boundaries between paradigms are becoming less rigid, and future systems are likely to leverage the complementary strengths of retrieval, prediction, and reasoning in a principled manner.
4. Datasets and Benchmarks
Datasets serve as the empirical foundation of Visual Geolocalization research. Unlike tasks with abstract labels, Visual Geolocalization requires linking visual inputs to real-world coordinates, which demands spatially meaningful and semantically rich supervision. Following the two data organizations introduced in Section 2.2, this section reviews representative datasets under two complementary categories: Structured Retrieval Datasets (Section 4.1) that establish explicit query-reference relationships, and Geospatially Annotated Datasets (Section 4.2) that directly associate visual observations with geographic labels. Table 1 provides a summary of representative datasets across the four subcategories. We conclude with a comparison and discussion of their characteristics, coverage, and limitations (Section 4.3).
4.1. Structured Retrieval Datasets
Structured retrieval datasets explicitly establish query-reference relationships by providing query images together with geo-referenced galleries, positive pairs, negative pairs, or cross-view correspondences. These datasets naturally support retrieval-based localization and are widely used in visual place recognition (VPR) and cross-view geo-localization (CVGL). Depending on whether the query and reference images share the same viewpoint modality, existing structured retrieval datasets can be further divided into same-view and cross-view datasets.
4.1.1. Same-View Retrieval Datasets
Same-view retrieval datasets provide query and reference images captured from the same viewpoint modality, most commonly ground-level photos or street-view panoramas. The reference gallery is typically geo-tagged, and location is inferred by matching the query against the gallery.
Early same-view datasets were built on geo-tagged photo collections from social media platforms. Oxford5k [85] and Paris6k [86], sourced from Flickr, established the standard evaluation protocol for instance-level landmark retrieval, providing 5K and 6K reference images with 55 queries each. Holidays [87] introduced 1,491 personal holiday photos with 500 queries, covering diverse scene types to test robustness against various image transformations. European Cities 50k [88] extended this line with 51K Flickr images covering 14 European cities, while Paris500k [91] scaled to 501K geotagged Flickr images within the Paris geographic boundary.
City-scale street-view datasets further expanded the reference gallery size. San Francisco Landmark [89] provided 1.7M reference images from Google Street View combined with mobile photos. Pitts250k [90] extracted 250K street-view images from 10,586 panoramic locations in Pittsburgh, and Pitts30k [92] provided a 30K subset that became a standard training and evaluation split. The 24/7 Tokyo dataset [27] specifically addressed day-night robustness by providing 315 query images spanning day, dusk, and night against a 76K daytime-only database.
With the growth of large-scale visual place recognition, landmark-centric datasets became prominent. Landmarks-full [93] collected 169K images covering 586 landmarks from an image search engine. Google Landmarks [94] scaled to 1.1M images spanning 4,872 cities across 187 countries, and Google Landmarks Dataset v2 [95] further expanded to 4.2M images covering approximately 200K landmark instances from Wikimedia Commons, establishing the largest instance-level retrieval benchmark to date.
4.1.2. Cross-View Retrieval Datasets
Cross-view retrieval datasets establish correspondences between images captured from different viewpoint modalities, most commonly ground-level (or street-view) and aerial (or satellite) imagery. These datasets enable the evaluation of cross-view matching algorithms that bridge the viewpoint gap.
The field originated with the Charleston & SF Dataset [96], which provided street-view, satellite, and land-cover attribute map triplets. CVUSA [97] became a widely adopted benchmark by pairing Flickr street-view panoramas with Bing Maps satellite imagery across the United States, comprising 2.48M image pairs. UrbanGeo [98] collected 18K ground-aerial pairs from Google Maps with building boundary annotations, and CVACT [99] refined the cross-view setting with higher resolution and strict north-aligned ground-aerial pairs from Australian cities. VIGOR [101] moved beyond one-to-one retrieval by introducing one-to-many matching with meter-level offset annotations.
A significant branch of cross-view datasets targets UAV-to-satellite matching. University-1652 [100] provided 146K images of 1,652 university buildings from drone, satellite, and street-view perspectives, becoming the standard multi-view benchmark. SUES-200 [105] offered drone imagery at multiple altitudes (150–300m), and CVOGL [104] introduced ground, UAV, and satellite triplets with object-level bounding box annotations for cross-view object geo-localization.
Sequence-based cross-view datasets have also emerged to leverage temporal information. SeqGeo [102] adapted CVUSA and VIGOR for image sequence geo-localization with 157K images, and KITTI-CVL [103] paired KITTI driving sequences with satellite imagery for video-based camera localization. More recently, text-augmented cross-view datasets have integrated natural language into cross-view matching: GeoText-1652 [106] added fine-grained text descriptions to University-1652, CVG-Text [45] paired street-view imagery with human-written descriptions across New York, Brisbane, and Tokyo, and DReSS [107] introduced panoramic street-view and very-high-resolution satellite imagery in decentralized settings.
4.2. Geospatially Annotated Datasets
Geospatially annotated datasets directly associate visual observations with geographic labels, including coordinates, administrative regions, landmarks, textual descriptions, question-answer pairs, or reasoning traces. These datasets primarily support prediction-based and reasoning-based localization and have evolved from simple image-coordinate pairs toward richer, semantically grounded annotations. Based on the annotation richness, we divide them into image-GPS datasets (with coordinates and optionally administrative labels) and image-GPS-text datasets (with additional textual descriptions, reasoning traces, or QA pairs).
4.2.1. Image–GPS Datasets
Image–GPS datasets associate each image with geographic coordinates and optionally with hierarchical administrative labels (e.g., country, state, city).
The earliest and most influential image–GPS datasets are built on Flickr photo collections. IM2GPS [25] provided 237 test images with GPS coordinates, establishing the first widely used evaluation set for ground-level geolocalization. MP16 [108] sampled 4.72M images from YFCC100M as the standard training set for Visual Geolocalization models. IM2GPS3K [6] (3K test images) and YFCC4K [6] (4K test images) provided larger and more diverse evaluation sets, while YFCC26K [56] further extended the evaluation scale to 26K images.
Street-view-sourced datasets addressed the geographic bias of Flickr collections. GWS15K [59] collected 15K images from Google Street View across 193 countries, providing a more balanced global distribution by sampling cities proportional to country surface area. OSV-5M [109] scaled this approach to 5.1M images from Mapillary crowdsourced street-view imagery, supplemented with rich geographic metadata including administrative divisions, land cover, climate, and soil type.
Hierarchical and address-level annotations have further enriched image–GPS datasets. MP16-Pro [70] extended MP16 with neighborhood, city, county, state, region, country, and continent labels, facilitating multimodal training for vision-language models. Pitts-IAL and SF-IAL-Base [50] supplemented street-view imagery with structured address information (building name, house number, street, neighborhood, city, county, country), enabling city-wide address localization. GeoRanking [29] constructed query-candidate pairs from MP16-Pro for distance-aware ranking evaluation. CityGuessr68k [61] broadened the modality to video, providing 68K YouTube clips with hierarchical continent/country/state/city labels. GeoGlobe [73] provided 293K images covering nearly 150 countries with diverse man-made landmarks and natural landscapes. MAPBench [76] provided 5K POI-centered street-view images with easy/hard splits, designed specifically for evaluating map-augmented agents.
4.2.2. Image–GPS–Text Datasets
Image–GPS–Text datasets go beyond coordinates by associating each image with textual descriptions, reasoning traces, question-answer pairs, or dialogue data. These datasets have emerged to support reasoning-based Visual Geolocalization and represent a significant shift toward cognitively meaningful supervision.
MP16-Reason [24] extended MP16-Pro by adding reasoning trajectory annotations, providing 46K samples with step-by-step geographic inference traces. NAVICLUES [23] contributed 1.1K images sourced from GeoGuessr expert gameplay videos, each annotated with visual cues, retrieval prompts, and final decisions aligned with the modular NAVIG pipeline. Geolocation Hub [110] contributed 50K Google Street View images, of which 30K have detailed text descriptions. GAEA-Bench [21] provided a 3.5K QA evaluation set derived from MP-16, Google Landmarks v2, and CityGuessr, with QA pairs grounded in OpenStreetMap attributes. GREval-Bench [68] offered a 3K evaluation set with GPS coordinates and chain-of-thought reasoning steps. GeoChain [66] contributed 1.46M images from Mapillary with 150 class labels and 30M QA pairs across 21 question types, providing the largest multimodal reasoning resource to date. MG-Geo [22] constructed 5M image-text pairs from OSV-5M metadata, GeoGuessr clues, and GPT-4V-generated conversational data.
Several datasets have introduced novel evaluation dimensions beyond standard localization. TimeSpot [112] introduced geo-temporal understanding with 1.5K images covering 80 countries, annotated with temporal attributes (season, month, time of day) alongside geographic labels. EarthWhere [111] probed geolocation skills across scales with 500 country-level multiple-choice questions and 310 street-level multi-step reasoning tasks. GTPred [113] provided 370 images for benchmarking interpretable geo-localization and time-of-capture prediction, jointly evaluating spatial and temporal reasoning.
4.3. Comparison and Discussion
4.3.1. Complementary Roles
Structured retrieval datasets and geospatially annotated datasets play complementary roles in Visual Geolocalization research. Structured retrieval datasets provide the reference infrastructure needed for retrieval-based methods, enabling fine-grained localization through visual or cross-modal matching. Geospatially annotated datasets, by contrast, provide the supervision signals needed for prediction-based and reasoning-based methods, supporting direct geographic inference and stepwise reasoning. Modern Visual Geolocalization systems often rely on both types of datasets: retrieval-augmented reasoning methods [69,70] use structured retrieval datasets to provide candidate references and geospatially annotated datasets to train the reasoning module.
4.3.2. Scale and Coverage
The scale of Visual Geolocalization datasets has grown steadily over the past two decades across both structured retrieval and geospatially annotated categories. In structured retrieval, reference galleries have expanded from Oxford5k’s 5K images [85] to Google Landmarks Dataset v2’s 4.2M images [95], enabling instance-level retrieval at unprecedented scale. In geospatially annotated data, training sets have grown from MP16’s 4.72M images [108] to OSV-5M’s 5.1M images [109], while evaluation benchmarks have progressed from IM2GPS’s 237 test images [25] to IM2GPS3K’s 3K [6] and YFCC26K’s 26K [56], though they remain intentionally compact to ensure manageable and reproducible evaluation. However, scale does not guarantee coverage. Flickr-sourced datasets such as Oxford5k [85] and IM2GPS3K [6] over-represent popular tourist destinations and developed regions, while street-view-sourced datasets such as GWS15K [59] and OSV-5M [109] are concentrated in urban areas of developed countries. Datasets sourced from gameplay and OpenStreetMap, such as NAVICLUES [23] and GAEA-Bench [21], offer broader geographic diversity but may introduce biases related to player demographics and map annotation completeness.
4.3.3. Annotation Evolution
The annotation richness of Visual Geolocalization datasets has evolved significantly. Early datasets provided only raw GPS coordinates [25,108], limiting supervision to location-level signals. Subsequent work enriched annotations with hierarchical administrative labels [50,70] and structured address strings. The most recent trend is the inclusion of reasoning trajectory supervision, multi-step traces that reflect how humans infer location from visual cues and prior knowledge, as exemplified by MP16-Reason [24], NAVICLUES [23], and GeoChain [66]. This progression reflects a fundamental shift from what the location is to why and how it can be inferred.
4.3.4. Modality Expansion
The input modality of Visual Geolocalization datasets has expanded from single ground-level images to include panoramas, satellite imagery, drone footage, video sequences, and text. Cross-view datasets [97,99,100,101] established multi-view correspondence as a core setting, while text-augmented datasets [45,106] integrated natural language into cross-view matching. Video-based datasets such as CityGuessr68k [61] and KITTI-CVL [103] introduced temporal dynamics, challenging models to leverage sequential visual cues.
4.3.5. Open Challenges
Despite the rapid growth in dataset quantity and diversity, several challenges persist. First, many images lack identifiable geographic cues, reducing supervision effectiveness; principled methods for filtering non-localizable samples or reducing their impact during training remain an open problem. Second, the gap between large-scale but weakly supervised data and small-scale but richly annotated data remains a fundamental tension. Third, standardized evaluation protocols for multi-step reasoning are still underdeveloped, making it difficult to assess reasoning faithfulness and tool-use correctness. Finally, privacy and ethical considerations related to inferring sensitive location information from visual content deserve greater attention as Visual Geolocalization systems become more capable.
5. Challenges and Outlook
Despite the rapid progress reviewed in preceding sections, Visual Geolocalization remains a complex and evolving research problem. This section discusses open challenges facing current methods (Section 5.1), the ongoing transition from recognition-driven to reasoning-driven geolocalization (Section 5.2), and the broader vision toward spatial intelligence (Section 5.3).
5.1. Open Challenges
5.1.1. Geographic Coverage and Data Bias
A fundamental challenge in Visual Geolocalization is the uneven geographic distribution of available data. Existing datasets, particularly those sourced from social media platforms such as Flickr [6,25,108], are heavily biased toward popular tourist destinations and developed regions. This bias leads to inflated performance on iconic scenes while masking significant difficulties in rural, remote, or underrepresented areas. Street-view-sourced datasets [59,109] provide more balanced coverage but remain limited by the availability of street-view imagery, which is concentrated in urban areas of developed countries. Addressing this coverage gap requires both broader data collection and debiasing strategies that can compensate for uneven geographic representation during training and evaluation.
5.1.2. Visual Ambiguity and Localizability
Not all images contain geographically discriminative cues. Many scenes, such as indoor environments, close-up objects, or generic natural landscapes, lack identifiable visual signals that enable reliable localization. These non-localizable samples act as noise during training and weaken the learning signal [20]. While recent datasets have begun to filter out such samples based on human localizability ratings, a more principled understanding of what makes an image localizable, and how to reduce the impact of non-localizable samples without explicit removal, remains an open problem. Closely related is the challenge of visual aliasing, in which visually similar but geographically distant locations are confused by appearance-based methods [70]. Resolving visual aliasing requires going beyond visual similarity to incorporate semantic reasoning and contextual verification.
5.1.3. Reasoning Supervision and Evaluation
The shift toward reasoning-based Visual Geolocalization introduces new challenges in supervision and evaluation. Current reasoning-oriented datasets such as NAVICLUES [23], GAEA-Bench [21], and GREval-Bench [68] provide reasoning traces, but these annotations are costly to collect and may reflect the biases of individual annotators. More critically, there is no standardized way to evaluate the correctness, completeness, or faithfulness of multi-step reasoning processes. While distance-based and category-based metrics are well established for measuring localization accuracy, metrics for reasoning quality, such as reasoning faithfulness, uncertainty estimation, and tool-use correctness, remain under active development and have not yet been standardized. Without reliable evaluation of reasoning, it is difficult to distinguish genuine spatial inference from shallow pattern matching.
5.1.4. Scalability and Efficiency
As Visual Geolocalization systems incorporate larger foundation models, retrieval databases, and multi-agent pipelines, scalability and efficiency become increasingly pressing concerns. Reasoning-based and agentic methods [23,72,73] typically incur high computational costs due to multi-step inference, tool interactions, and inter-agent communication. Generative prediction methods [64,65] require iterative sampling that increases inference latency. Balancing the accuracy and interpretability gains of reasoning-based approaches against their computational overhead is a key challenge for deploying Visual Geolocalization systems in real-world, latency-sensitive applications.
5.1.5. Privacy and Ethical Considerations
Visual Geolocalization inherently involves inferring sensitive location information from visual content, raising privacy concerns [3,4]. Images shared on social media may be geolocated without the creator’s consent, potentially exposing personal information such as home addresses or travel patterns. Developing privacy-preserving geolocalization techniques, and establishing ethical guidelines for data collection, model deployment, and result dissemination, is an important but underexplored direction.
5.2. From Recognition to Reasoning
5.2.1. The Paradigm Shift
The evolution of Visual Geolocalization reflects a broader shift in computer vision from recognition to reasoning. Early methods treated geolocalization as a visual recognition problem, relying on handcrafted features [25,26,27] or learned embeddings [6,8] to match query images against reference data. Prediction-based methods [5,7] encoded geographic knowledge directly into model parameters through classification. While effective in many settings, these approaches fundamentally operate by correlating visual appearance with geographic locations, making them vulnerable to visually ambiguous scenes, limited reference coverage, and distribution shifts.
The emergence of LVLMs [14,15,16,17] has enabled a fundamentally different approach. Reasoning-based methods [18,20,21,23,24] decompose the localization process into interpretable steps: identifying visual cues, generating region-level hypotheses, and refining predictions through contextual analysis and external knowledge. This shift from recognition to reasoning improves transparency, supports error analysis, and enables new forms of supervision through reasoning annotations.
5.2.2. Inference-Stage Strategies
A key trend is the use of inference-stage strategies to elicit geographic reasoning from LVLMs without additional training. Prompt engineering, chain-of-thought (CoT) prompting [114], and retrieval-augmented generation (RAG) [115,116] have been successfully applied to Visual Geolocalization. Methods such as GeoCoT [20] use CoT prompting to produce structured inferences aligned with human spatial thinking, while Img2Loc [69] and G3 [70] use RAG to enrich prompts with retrieved geographic context. Cheng et al. [71] further showed that test-time scaling can guide multimodal LLMs for efficient visual place recognition without fine-tuning. However, these strategies face challenges: prompt design remains heuristic, CoT outputs can be unreliable, and RAG methods depend on retrieval quality.
5.2.3. Training-Stage Strategies
Training-based approaches adapt LVLMs to Visual Geolocalization through supervised fine-tuning (SFT) or reinforcement learning (RL). SFT methods [18,23,67] fine-tune models on curated reasoning traces, improving interpretability and reasoning structure, but depend on costly annotations and typically optimize for accuracy rather than reasoning quality. RL methods [24,68] offer an alternative by directly optimizing reasoning behaviors through reward signals. GLOBE [24] applies GRPO-based RL [84] with geolocation-aware rewards, while the GRE suite [68] combines fine-tuning with enhanced reasoning chains. Techniques such as DPO [117] and PPO [118] also offer potential for aligning model behavior with cognitive reasoning objectives. Despite their promise, RL approaches face challenges in designing stable, interpretable reward functions for multi-step reasoning, and training can be unstable.
5.2.4. Toward Unified Systems
The boundaries between paradigms are becoming increasingly blurred. Modern Visual Geolocalization systems often integrate retrieval, prediction, and reasoning within a unified framework: GeoRanker [29] combines retrieval with reasoning-based re-ranking, G3 [70] couples retrieval with adaptive reasoning, and NAVIG [23] integrates reasoning with external knowledge retrieval. Agentic systems [72,73,74,75,76] further extend this integration by orchestrating multiple specialized agents and tools. This convergence suggests that the future of Visual Geolocalization lies not in any single paradigm but in the principled integration of recognition, prediction, and reasoning capabilities.
5.3. Toward Spatial Intelligence
5.3.1. Geographic World Knowledge
A defining limitation of current Visual Geolocalization models is the gap between visual perception and geographic world knowledge. While LVLMs have demonstrated impressive scene understanding capabilities, their geographic knowledge is often shallow, fragmented, or biased toward well-represented regions. Closing this gap requires integrating structured geographic knowledge bases, such as OpenStreetMap [21], administrative boundaries, climate zones, and cultural-linguistic atlases, into the reasoning process. Recent work such as GAEA-Bench [21] and LocationAgent [74] has begun to explore this direction by grounding reasoning in map attributes and decoupling parametric knowledge from external evidence, but a comprehensive framework for geographic knowledge integration remains an open goal.
5.3.2. Multimodal and Multi-Source Fusion
Real-world geolocalization often benefits from multiple information sources beyond a single image, including temporal cues, weather conditions, textual metadata, and multi-view observations. Current methods predominantly operate on single ground-level images, leaving the potential of multimodal and multi-source fusion largely untapped. Future systems should be able to jointly reason over images, videos, text, maps, and metadata, leveraging the complementary information each modality provides. Cross-view methods [42,43,47] and cross-modal methods [8,50] represent early steps in this direction, but full multimodal fusion, particularly the integration of temporal and environmental context, remains underexplored.
5.3.3. Open-World and Continual Localization
Existing Visual Geolocalization systems are typically evaluated on fixed benchmark datasets with predefined geographic coverage. In open-world settings, models must handle continuously evolving geographic environments, previously unseen locations, and distribution shifts caused by seasonal changes, urban development, or natural disasters. Developing Visual Geolocalization systems that can generalize to open-world conditions and continually update their geographic knowledge without catastrophic forgetting is an important frontier. This requires advances in continual learning, domain adaptation, and uncertainty-aware inference, as well as evaluation protocols that explicitly measure open-world generalization.
5.3.4. Agentic and Tool-Augmented Geolocalization
The emergence of agentic Visual Geolocalization systems [23,72,73,74,75,76] points toward a future in which geolocalization is performed by autonomous agents that actively gather evidence, consult external tools, and collaborate with other agents. These systems offer several advantages: modularity, transparency, and the ability to leverage specialized tools such as web search, map services, and image retrieval. However, they also introduce challenges in agent coordination, tool-use reliability, and computational efficiency. A key direction is the development of more robust and efficient agentic architectures, including hierarchical agent design [74], graph-based multi-agent debate [75], and parallel map-augmented reasoning [76].
5.3.5. Trustworthy and Interpretable Geolocalization
As Visual Geolocalization systems are increasingly deployed in real-world applications, including autonomous navigation [119], urban computing [2,120], and crisis response [12,13], trustworthiness and interpretability become critical requirements. Users need to understand not only where a model predicts an image was taken, but also why. Reasoning-based methods [18,20,21] offer a natural path toward interpretable geolocalization by producing human-readable explanations. However, ensuring that these explanations are faithful to the model’s actual decision process, and that models do not produce plausible-sounding but incorrect justifications, remains an open challenge. Developing evaluation protocols that jointly assess localization accuracy and reasoning faithfulness, along with techniques for uncertainty estimation and calibration, will be essential for building trustworthy Visual Geolocalization systems.
5.3.6. Toward Spatial Intelligence
Looking forward, we envision Visual Geolocalization evolving into a core capability of spatial intelligence, the ability to perceive, understand, reason about, and act within geographic space. Spatial intelligence goes beyond predicting coordinates: it encompasses the ability to recognize geographic context, understand spatial relationships, infer cultural and environmental characteristics, and integrate diverse information sources to make geographically grounded decisions. Achieving this vision will require advances across the full stack of Visual Geolocalization research: richer datasets with multimodal and reasoning-oriented supervision, unified methodologies that integrate recognition, prediction, and reasoning, and evaluation frameworks that measure not only localization accuracy but also reasoning quality, robustness, and trustworthiness. As foundation models, multimodal reasoning, and agentic systems continue to mature, we expect Visual Geolocalization to play an increasingly central role in bridging visual perception and geographic understanding, ultimately enabling more intelligent, trustworthy, and spatially aware AI systems.
6. Conclusions
Visual Geolocalization has become an important research topic at the intersection of computer vision, geospatial intelligence, and multimodal artificial intelligence. Rapid advances in deep representation learning, large-scale geo-tagged datasets, and large vision-language models have significantly expanded the capability and scope of Visual Geolocalization, while also introducing new opportunities and challenges for geographic understanding. In this survey, we presented a comprehensive review of Visual Geolocalization from the perspectives of methodologies and datasets. We established a unified formulation of the task and organized existing methods into three representative methodological paradigms: retrieval-based, prediction-based, and reasoning-based localization. Based on this taxonomy, we systematically reviewed representative methods, benchmark datasets, evaluation protocols, and practical applications, and further discussed the strengths, limitations, and open challenges of existing approaches. Looking forward, we expect Visual Geolocalization to continue benefiting from advances in foundation models, multimodal reasoning, and agentic systems. Future research is likely to place increasing emphasis on integrating visual perception, semantic understanding, world knowledge, and external tools to improve localization accuracy, robustness, interpretability, and open-world generalization. We hope this survey provides a unified perspective on the rapidly evolving landscape of Visual Geolocalization and serves as a valuable reference for future research toward more intelligent, trustworthy, and spatially aware geolocalization systems.
References
- Hao, X.; Jiang, Y.; Zou, X.; Liu, J.; Yin, Y.; Liang, Y. Unlocking Location Intelligence: A Survey from Deep Learning to The LLM Era. arXiv 2025, arXiv:2505.09651. [Google Scholar]
- Zou, X.; Yan, Y.; Hao, X.; Hu, Y.; Wen, H.; Liu, E.; Zhang, J.; Li, Y.; Li, T.; Zheng, Y.; et al. Deep learning for cross-domain data fusion in urban computing: Taxonomy, advances, and outlook. Inf. Fusion 2025, 113, 102606. [Google Scholar] [CrossRef]
- Lorestani, M.A.; Ranbaduge, T.; Rakotoarivelo, T. Privacy risk in GeoData: A survey. arXiv 2024, arXiv:2402.03612. [Google Scholar]
- Jin, F.; Hua, W.; Francia, M.; Chao, P.; Orlowska, M.E.; Zhou, X. A survey and experimental study on privacy-preserving trajectory data publishing. IEEE Trans. Knowl. Data Eng. 2022, 35, 5577–5596. [Google Scholar] [CrossRef]
- Weyand, T.; Kostrikov, I.; Philbin, J. Planet-photo geolocation with convolutional neural networks. In Proceedings of the Proceedings of the European Conference on Computer Vision (ECCV), 2016; Springer; pp. 37–55. [Google Scholar]
- Vo, N.; Jacobs, N.; Hays, J. Revisiting IM2GPS in the Deep Learning Era. In Proceedings of the Proceedings of the IEEE International Conference on Computer Vision, 2017; pp. 2621–2630. [Google Scholar]
- Seo, P.H.; Weyand, T.; Sim, J.; Han, B. Cplanet: Enhancing image geolocalization by combinatorial partitioning of maps. In Proceedings of the Proceedings of the European Conference on Computer Vision (ECCV), 2018; pp. 536–551. [Google Scholar]
- Vivanco Cepeda, V.; Nayak, G.K.; Shah, M. Geoclip: Clip-inspired alignment between locations and images for effective worldwide geo-localization. Adv. Neural Inf. Process. Syst. 2023, 36, 8690–8701. [Google Scholar] [CrossRef]
- Yim, H.S.; Ahn, H.J.; Kim, J.W.; Park, S.J. Agent-based adaptive travel planning system in peak seasons. Expert Syst. With Appl. 2004, 27, 211–222. [Google Scholar] [CrossRef]
- Wang, K.; Shen, Y.; Lv, C.; Zheng, X.; Huang, X.J. Triptailor: A real-world benchmark for personalized travel planning. Proc. Find. Assoc. Comput. Linguist. ACL 2025, 2025, 9705–9723. [Google Scholar] [CrossRef]
- Yuan, L.; Han, D.J.; Brinton, C.G.; Brunswicker, S. LLMAP: LLM-Assisted Multi-Objective Route Planning with User Preferences. arXiv 2025, arXiv:2509.12273. [Google Scholar]
- Firmansyah, H.B.; Fernandez-Marquez, J.L.; Mulayim, M.O.; Gomes, J.; Ribeiro, J.; Lorini, V. Empowering Crisis Response Efforts: A Novel Approach to Geolocating Social Media Images for Enhanced Situational Awareness. In Proceedings of the Proceedings of the International ISCRAM Conference, 2024. [Google Scholar]
- Sapru, D. GeoSense-AI: Fast Location Inference from Crisis Microblogs. arXiv 2025, arXiv:2512.18225. [Google Scholar]
- Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F.L.; et al. Gpt-4 technical report. arXiv 2023, arXiv:2303.08774. [Google Scholar]
- Liu, H.; Li, C.; Wu, Q.; Lee, Y.J. Visual instruction tuning. Adv. Neural Inf. Process. Syst. 2023, 36, 34892–34916. [Google Scholar] [CrossRef]
- Bai, J.; Bai, S.; Chu, Y.; Cui, Z.; et al. Qwen technical report. arXiv 2023, arXiv:2309.16609. [Google Scholar]
- Chen, Z.; Wu, J.; Wang, W.; et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024; pp. 24185–24198. [Google Scholar]
- Li, L.; Ye, Y.; Jiang, B.; Zeng, W. Georeasoner: Geo-localization with reasoning in street views using a large vision-language model. In Proceedings of the Forty-first International Conference on Machine Learning, 2024. [Google Scholar]
- Liu, Y.; Ding, J.; Deng, G.; Li, Y.; Zhang, T.; Sun, W.; Zheng, Y.; Ge, J.; Liu, Y. Image-Based Geolocation Using Large Vision-Language Models. arXiv 2024, arXiv:2408.09474. [Google Scholar]
- Song, Z.; Yang, J.; Huang, Y.; Tonglet, J.; Zhang, Z.; Cheng, T.; Fang, M.; Gurevych, I.; Chen, X. Geolocation with real human gameplay data: A large-scale dataset and human-like reasoning framework. arXiv 2025, arXiv:2502.13759. [Google Scholar]
- Campos, R.; Vayani, A.; Kulkarni, P.P.; Gupta, R.; Zafar, A.; Dutta, A.; Shah, M. Gaea: A geolocation aware conversational assistant. In Proceedings of the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2026; pp. 5236–5246. [Google Scholar]
- Dou, Z.; Wang, Z.; Han, X.; Qiang, C.; Wang, K.; Li, G.; Huang, Z.; Han, Z. GaGA: Towards Interactive Global Geolocation Assistant. In Proceedings of the Proceedings of the European conference on computer vision (ECCV), 2026. [Google Scholar]
- Zhang, Z.; Li, R.; Kabir, T.; Boyd-Graber, J. NAVIG: Natural Language-guided Analysis with Vision Language Models for Image Geo-localization. arXiv 2025, arXiv:2502.14638. [Google Scholar]
- Li, L.; Zhou, Y.; Liang, Y.; Tsung, F.; Wei, J. Recognition through reasoning: Reinforcing image geo-localization with large vision-language models. Adv. Neural Inf. Process. Syst. 2026, 38, 62132–62159. [Google Scholar]
- Hays, J.; Efros, A.A. Im2gps: estimating geographic information from a single image. In Proceedings of the 2008 ieee conference on computer vision and pattern recognition; IEEE, 2008; pp. 1–8. [Google Scholar]
- Yagcioglu, S.; Erdem, E.; Erdem, A. City scale image geolocalization via dense scene alignment. In Proceedings of the 2015 IEEE Winter Conference on Applications of Computer Vision; IEEE, 2015; pp. 726–732. [Google Scholar]
- Torii, A.; Arandjelovic, R.; Sivic, J.; Okutomi, M.; Pajdla, T. 24/7 place recognition by view synthesis. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2015; pp. 1808–1817. [Google Scholar]
- Berton, G.; Masone, C.; Paolicelli, V.; Caputo, B. Viewpoint invariant dense matching for visual geolocalization. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021; pp. 12169–12178. [Google Scholar]
- Jia, P.; Park, S.; Gao, S.; Zhao, X.; Li, S. Georanker: Distance-aware ranking for worldwide image geolocalization. Adv. Neural Inf. Process. Syst. 2026, 38, 17673–17699. [Google Scholar]
- Pan, K.; Guo, W.; Liu, Y.; Zhang, X.; Cheng, X.; Wu, R. Large-Scale Geo-Localization of Remote Sensing Images: A Three-Stage Framework Leveraging Maximal Clique Theory. IEEE Transactions on Geoscience and Remote Sensing, 2025. [Google Scholar]
- Waheed, S.; An, N.M.; Milford, M.; Ramchurn, S.D.; Ehsan, S. VLM-Guided Visual Place Recognition for Planet-Scale Geo-Localization. arXiv 2025, arXiv:2507.17455. [Google Scholar]
- Lin, T.Y.; Cui, Y.; Belongie, S.; Hays, J. Learning deep representations for ground-to-aerial geolocalization. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2015; pp. 5007–5015. [Google Scholar]
- Hu, S.; Feng, M.; Nguyen, R.M.; Lee, G.H. Cvm-net: Cross-view matching network for image-based ground-to-aerial geo-localization. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2018; pp. 7258–7267. [Google Scholar]
- Vo, N.N.; Hays, J. Localizing and orienting street views using overhead imagery. In Proceedings of the European conference on computer vision, 2016; Springer; pp. 494–509. [Google Scholar]
- Cai, S.; Guo, Y.; Khan, S.; Hu, J.; Wen, G. Ground-to-aerial image geo-localization with a hard exemplar reweighting triplet loss. In Proceedings of the Proceedings of the IEEE/CVF international conference on computer vision, 2019; pp. 8391–8400. [Google Scholar]
- Shi, Y.; Liu, L.; Yu, X.; Li, H. Spatial-aware feature aggregation for image based cross-view geo-localization. Adv. Neural Inf. Process. Syst. 2019, 32. [Google Scholar]
- Tang, H.; Xu, D.; Sebe, N.; Wang, Y.; Corso, J.J.; Yan, Y. Multi-channel attention selection gan with cascaded semantic guidance for cross-view image translation. In Proceedings of the Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019; pp. 2417–2426. [Google Scholar]
- Tian, X.; Shao, J.; Ouyang, D.; Shen, H.T. UAV-satellite view synthesis for cross-view geo-localization. IEEE Trans. Circuits Syst. Video Technol. 2021, 32, 4804–4815. [Google Scholar] [CrossRef]
- Vyas, S.; Chen, C.; Shah, M. Gama: Cross-view video geo-localization. In Proceedings of the European Conference on Computer Vision, 2022; Springer; pp. 440–456. [Google Scholar]
- Arrabi, A.; Zhang, X.; Sultani, W.; Chen, C.; Wshah, S. Cross-view meets diffusion: Aerial image synthesis with geometry and text guidance. In Proceedings of the 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV); IEEE, 2025; pp. 5356–5366. [Google Scholar]
- Yang, H.; Lu, X.; Zhu, Y. Cross-view geo-localization with evolving transformer. arXiv 2021, arXiv:2107.00842. [Google Scholar]
- Zhu, S.; Shah, M.; Chen, C. Transgeo: Transformer is all you need for cross-view image geo-localization. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022; pp. 1162–1171. [Google Scholar]
- Deuser, F.; Habel, K.; Oswald, N. Sample4geo: Hard negative sampling for cross-view geo-localisation. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023; pp. 16847–16856. [Google Scholar]
- Mo, Z.; Sun, Y.; Xu, M.; Jia, S. SIGN: Saliency-Aware Integrated Global-Local Network for Cross-View Geo-Localization. In Proceedings of the IGARSS 2025-2025 IEEE International Geoscience and Remote Sensing Symposium; IEEE, 2025; pp. 6296–6300. [Google Scholar]
- Ye, J.; Lin, H.; Ou, L.; Chen, D.; Wang, Z.; Zhu, Q.; He, C.; Li, W. Where am i? cross-view geo-localization with natural language descriptions. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025; pp. 5890–5900. [Google Scholar]
- Lu, X.; Zheng, Z.; Wan, Y.; Yao, Y.; Wang, A.; Zhang, R.; Xia, P.; Wu, Q.; Li, Q.; Lin, W.; et al. GLEAM: Learning to Match and Explain in Cross-View Geo-Localization. arXiv 2025, arXiv:2509.07450. [Google Scholar]
- Song, Z.; Zhang, J.; Wang, D.; Zhou, Z.; Liu, W.; Guo, H.; Wang, E.; Du, B. Geobridge: A semantic-anchored multi-view foundation model bridging images and text for geo-localization. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026; pp. 27793–27803. [Google Scholar]
- Haas, L.; Alberti, S.; Skreta, M. Learning generalized zero-shot learners for open-domain image geolocalization. arXiv 2023, arXiv:2302.00275. [Google Scholar]
- Jia, F.; Liu, L.; Hou, C.; Zhang, F.; Liu, X.; Liu, Y. Towards Interpretable Geo-localization: a Concept-Aware Global Image-GPS Alignment Framework. arXiv 2025, arXiv:2509.01910. [Google Scholar]
- Xu, S.; Zhang, C.; Fan, L.; Meng, G.; Xiang, S.; Ye, J. Addressclip: Empowering vision-language models for city-wide image address localization. In Proceedings of the European Conference on Computer Vision, 2024; Springer; pp. 76–92. [Google Scholar]
- Shatwell, D.G.; Dave, I.R.; Swetha, S.; Shah, M. Gt-loc: Unifying when and where in images through a joint embedding space. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025; pp. 1–11. [Google Scholar]
- Fang, J.; Qian, S.; Liu, S. GEOMR: Integrating Image Geographic Features and Human Reasoning Knowledge for Image Geolocalization. Knowl.-Based Syst. 2026, 115391. [Google Scholar] [CrossRef]
- Johns, J.M.; Rounds, J.; Henry, M.J. Multi-modal Geolocation Estimation Using Deep Neural Networks. arXiv 2017, arXiv:1712.09458. [Google Scholar]
- Muller-Budack, E.; Pustu-Iren, K.; Ewerth, R. Geolocation estimation of photos using a hierarchical model and scene classification. In Proceedings of the Proceedings of the European conference on computer vision (ECCV), 2018; pp. 563–579. [Google Scholar]
- Izbicki, M.; Papalexakis, E.E.; Tsotras, V.J. Exploiting the earth’s spherical geometry to geolocate images. In Proceedings of the Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2019, Würzburg, Germany, September 16–20, 2019, Proceedings, Part II; Springer; pp. 3–19.
- Theiner, J.; Müller-Budack, E.; Ewerth, R. Interpretable semantic photo geolocation. In Proceedings of the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2022; pp. 750–760. [Google Scholar]
- Luo, G.; Biamby, G.; Darrell, T.; Fried, D.; Rohrbach, A. G3: Geolocation via Guidebook Grounding. Proc. Find. Assoc. Comput. Linguist. EMNLP 2022, 2022, 5841–5853. [Google Scholar] [CrossRef]
- Pramanick, S.; Nowara, E.M.; Gleason, J.; Castillo, C.D.; Chellappa, R. Where in the world is this image? transformer-based geo-localization in the wild. In Proceedings of the European Conference on Computer Vision, 2022; Springer; pp. 196–215. [Google Scholar]
- Clark, B.; Kerrigan, A.; Kulkarni, P.P.; Cepeda, V.V.; Shah, M. Where we are and what we’re looking at: Query based worldwide image geo-localization using hierarchies and scenes. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023; pp. 23182–23190. [Google Scholar]
- Ghasemi, N.; Ziashahabi, A.; Avestimehr, S.; Shahabi, C. GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction. arXiv 2025, arXiv:2511.01082. [Google Scholar]
- Kulkarni, P.P.; Nayak, G.K.; Shah, M. CityGuessr: City-Level Video Geo-Localization on a Global Scale. In Proceedings of the European Conference on Computer Vision, 2024; Springer; pp. 293–311. [Google Scholar]
- Bianco, M.J.; Eigen, D.; Gormish, M. Enhancing worldwide image geolocation by ensembling satellite-based ground-level attribute predictors. In Proceedings of the 2025 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW); IEEE, 2025; pp. 535–543. [Google Scholar]
- Mai, G.; Xuan, Y.; Zuo, W.; He, Y.; Song, J.; Ermon, S.; Janowicz, K.; Lao, N. Sphere2Vec: A general-purpose location representation learning over a spherical surface for large-scale geospatial predictions. ISPRS J. Photogramm. Remote Sens. 2023, 202, 439–462. [Google Scholar] [CrossRef]
- Dufour, N.; Kalogeiton, V.; Picard, D.; Landrieu, L. Around the world in 80 timesteps: A generative approach to global visual geolocation. In Proceedings of the Proceedings of the Computer Vision and Pattern Recognition Conference, 2025; pp. 23016–23026. [Google Scholar]
- Wang, Z.; Liu, Z.; Zhang, J.; Zhou, Z.; Cao, Q.; Wu, N.; Mu, L.; Song, Y.; Xie, Y.; Lao, N.; et al. LocDiff: Identifying Locations on Earth by Diffusing in the Hilbert Space. Adv. Neural Inf. Process. Syst. 2026, 38, 620–647. [Google Scholar]
- Yerramilli, S.; Pande, N.; Grover, R.; Tamarapalli, J.S. GeoChain: Multimodal Chain-of-Thought for Geographic Reasoning. In Proceedings of the EMNLP (Findings), 2025; pp. 23624–23639. [Google Scholar]
- Yi, Q.; Shan, L. Geolocsft: Efficient visual geolocation via supervised fine-tuning of multimodal foundation models. arXiv 2025, arXiv:2506.01277. [Google Scholar]
- Wang, C.; Ye, X.; Pan, X.; Pan, Z.; Wang, H.; Song, Y. Gre suite: Geo-localization inference via fine-tuned vision-language models and enhanced reasoning chains. Adv. Neural Inf. Process. Syst. 2026, 38, 92379–92409. [Google Scholar]
- Zhou, Z.; Zhang, J.; Guan, Z.; Hu, M.; Lao, N.; Mu, L.; Li, S.; Mai, G. Img2Loc: Revisiting image geolocalization using multi-modality foundation models and image-based retrieval-augmented generation. In Proceedings of the Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2024; pp. 2749–2754. [Google Scholar]
- Jia, P.; Liu, Y.; Li, X.; Zhao, X.; Wang, Y.; Du, Y.; Han, X.; Wei, X.; Wang, S.; Yin, D. G3: an effective and adaptive framework for worldwide geolocalization using large multi-modality models. Adv. Neural Inf. Process. Syst. 2024, 37, 53198–53221. [Google Scholar] [CrossRef]
- Cheng, J.; Li, W.; Luo, J.; Tang, X.; He, Z.; Wu, J.; Zou, Y.; Zhang, W. Scale, don’t fine-tune: Guiding multimodal llms for efficient visual place recognition at test-time. IFAC-PapersOnLine 2025, 59, 97–102. [Google Scholar] [CrossRef]
- Wang, Y.; Liu, Z.; Wang, Z.; Hu, H.; Liu, P.; Rao, Y. GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization. arXiv 2025, arXiv:2511.15705. [Google Scholar]
- Han, X.; Zhu, C.; Zhu, H.; Zhao, X. Swarm intelligence in geo-localization: A multi-agent large vision-language model collaborative framework. In Proceedings of the Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2; 2025; pp. 814–825. [Google Scholar]
- Li, Q.; Xiao, Z.; Wang, X.; Ma, Z.; Yang, C.; Li, H. LocationAgent: A Hierarchical Agent for Image Geolocation via Decoupling Strategy and Evidence from Parametric Knowledge. arXiv 2026, arXiv:2601.19155. [Google Scholar]
- Zheng, H.; Shi, Y.; Gu, X.; You, H.; Zhang, Z.; Gan, L.; Zhang, H.; Huang, W.; Huang, J. GraphGeo: Multi-Agent Debate Framework for Visual Geo-localization with Heterogeneous Graph Neural Networks. arXiv 2025, arXiv:2511.00908. [Google Scholar]
- Ji, Y.; Wang, Y.; Ma, Z.; Hu, Y.; Huang, H.; Hu, X.; Chen, G.; Wu, L.; Chu, X. Thinking with map: Reinforced parallel map-augmented agent for geolocalization. Proc. Find. Assoc. Comput. Linguist. ACL 2026, 2026, 17991–18004. [Google Scholar] [CrossRef]
- Krizhevsky, A.; Sutskever, I.; Hinton, G.E. Imagenet classification with deep convolutional neural networks. Adv. Neural Inf. Process. Syst. 2012, 25. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2016; pp. 770–778. [Google Scholar]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the International conference on machine learning; PmLR, 2021; pp. 8748–8763. [Google Scholar]
- Oliva, A.; Torralba, A. Building the gist of a scene: The role of global image features in recognition. Prog. Brain Res. 2006, 155, 23–36. [Google Scholar] [CrossRef] [PubMed]
- Torralba, A.; Fergus, R.; Freeman, W.T. Tiny images. 2007.
- Arandjelovic, R.; Zisserman, A. All about VLAD. In Proceedings of the Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2013; pp. 1578–1585. [Google Scholar]
- Jégou, H.; Douze, M.; Schmid, C.; Pérez, P. Aggregating local descriptors into a compact image representation. In Proceedings of the 2010 IEEE computer society conference on computer vision and pattern recognition; IEEE, 2010; pp. 3304–3311. [Google Scholar]
- Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Bi, X.; Zhang, H.; et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv 2024, arXiv:2402.03300. [Google Scholar]
- Philbin, J.; Chum, O.; Isard, M.; Sivic, J.; Zisserman, A. Object retrieval with large vocabularies and fast spatial matching. In Proceedings of the 2007 IEEE conference on computer vision and pattern recognition; IEEE, 2007; pp. 1–8. [Google Scholar]
- Philbin, J.; Chum, O.; Isard, M.; Sivic, J.; Zisserman, A. Lost in quantization: Improving particular object retrieval in large scale image databases. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2008; pp. 1–8. [Google Scholar]
- Jegou, H.; Douze, M.; Schmid, C. Hamming embedding and weak geometric consistency for large scale image search. In Proceedings of the European Conference on Computer Vision, 2008; Springer; pp. 304–317. [Google Scholar]
- Avrithis, Y.; Tolias, G.; Kalantidis, Y. Feature map hashing: Sub-linear indexing of appearance and global geometry. In Proceedings of the Proceedings of the 18th ACM international conference on Multimedia, 2010; pp. 231–240. [Google Scholar]
- Chen, D.M.; Baatz, G.; Köser, K.; Tsai, S.S.; Vedantham, R.; Pylvänäinen, T.; Roimela, K.; Chen, X.; Bach, J.; Pollefeys, M.; et al. City-scale landmark identification on mobile devices. In Proceedings of the CVPR 2011; IEEE, 2011; pp. 737–744. [Google Scholar]
- Torii, A.; Sivic, J.; Pajdla, T.; Okutomi, M. Visual place recognition with repetitive structures. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2013; pp. 883–890. [Google Scholar]
- Weyand, T.; Leibe, B. Visual landmark recognition from internet photo collections: A large-scale evaluation. Comput. Vis. Image Underst. 2015, 135, 1–15. [Google Scholar] [CrossRef]
- Arandjelovic, R.; Gronat, P.; Torii, A.; Pajdla, T.; Sivic, J. NetVLAD: CNN architecture for weakly supervised place recognition. In Proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, 2016; pp. 5297–5307. [Google Scholar]
- Gordo, A.; Almazán, J.; Revaud, J.; Larlus, D. Deep image retrieval: Learning global representations for image search. In Proceedings of the European conference on computer vision, 2016; Springer; pp. 241–257. [Google Scholar]
- Noh, H.; Araujo, A.; Sim, J.; Weyand, T.; Han, B. Large-scale image retrieval with attentive deep local features. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV); IEEE, 2017; pp. 3476–3485. [Google Scholar]
- Weyand, T.; Araujo, A.; Cao, B.; Sim, J. Google Landmarks Dataset v2 – A Large-Scale Benchmark for Instance-Level Recognition and Retrieval. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020; pp. 2575–2584. [Google Scholar]
- Lin, T.Y.; Belongie, S.; Hays, J. Cross-view image geolocalization. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2013; pp. 891–898. [Google Scholar]
- Workman, S.; Souvenir, R.; Jacobs, N. Wide-area image geolocalization with aerial reference imagery. In Proceedings of the Proceedings of the IEEE International Conference on Computer Vision, 2015; pp. 3961–3969. [Google Scholar]
- Tian, Y.; Chen, C.; Shah, M. Cross-view image matching for geo-localization in urban environments. In Proceedings of the Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017; pp. 3608–3616. [Google Scholar]
- Liu, L.; Li, H.; Dai, Y. Lending orientation to neural networks for cross-view geo-localization. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019; pp. 7112–7121. [Google Scholar]
- Zheng, Z.; Wei, Y.; Sui, W.; Chen, R.; Shen, Y. University-1652: A Multi-view Multi-source Benchmark for Drone-based Geo-localization. In Proceedings of the Proceedings of the 28th ACM International Conference on Multimedia, 2020; pp. 1399–1408. [Google Scholar]
- Zhu, S.; Shen, W.; Diao, L.; Li, W.; Chen, X.; Hu, J.; Lu, X.; Liu, X. VIGOR: Cross-View Image Geo-Localization Beyond One-to-One Retrieval. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021; pp. 5312–5321. [Google Scholar]
- Zhang, X.; Sultani, W.; Wshah, S. Cross-view image sequence geo-localization. In Proceedings of the 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV); IEEE, 2023; pp. 2913–2922. [Google Scholar]
- Shi, Y.; Yu, X.; Wang, S.; Li, H. Cvlnet: Cross-view semantic correspondence learning for video-based camera localization. In Proceedings of the Asian Conference on Computer Vision, 2022; Springer; pp. 123–141. [Google Scholar]
- Sun, Y.; Ye, Y.; Kang, J.; Fernandez-Beltran, R.; Feng, S.; Li, X.; Luo, C.; Zhang, P.; Plaza, A. Cross-view object geo-localization in a local region with satellite imagery. IEEE Trans. Geosci. Remote Sens. 2023, 61, 1–16. [Google Scholar] [CrossRef]
- Zhu, R.; Yin, L.; Yang, M.; Wu, F.; Yang, Y.; Hu, W. SUES-200: A multi-height multi-scene cross-view image benchmark across drone and satellite. IEEE Trans. Circuits Syst. Video Technol. 2023, 33, 4825–4839. [Google Scholar] [CrossRef]
- Chu, M.; Zheng, Z.; Ji, W.; Wang, T.; Chua, T.S. Towards natural language-guided drones: Geotext-1652 benchmark with spatial relation matching. In Proceedings of the European Conference on Computer Vision, 2024; Springer; pp. 213–231. [Google Scholar]
- Xia, P.; Yu, L.; Wan, Y.; Wu, Q.; Chen, P.; Zhong, L.; Yao, Y.; Wei, D.; Liu, X.; Ru, L.; et al. Cross-view geo-localization with panoramic street-view and VHR satellite imagery in decentrality settings. ISPRS J. Photogramm. Remote Sens. 2025, 227, 1–11. [Google Scholar] [CrossRef]
- Larson, M.; Soleymani, M.; Gravier, G.; Ionescu, B.; Jones, G.J. The benchmarking initiative for multimedia evaluation: MediaEval 2016. IEEE Multimed. 2017, 24, 93–96. [Google Scholar] [CrossRef]
- Sévelov, Y.; Paret, A.; Klinger, V.; Jeni, L.A.; Lepetit, V.; Hamza, M.; Kostinger, A.; LeCun, Y. OpenStreetView-5M: The Many Roads to Global Visual Geolocation. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024. [Google Scholar]
- Liu, Y.; Deng, G.; Ding, J.; Li, Y.; Zhang, T.; Sun, W.; Zheng, Y.; Ge, J. Mission: Impossible–image-based geolocation with large vision language models. Proceedings on Privacy Enhancing Technologies, 2025. [Google Scholar]
- Qian, T.; et al. Where on Earth? A Vision-Language Benchmark for Probing Model Geolocation Skills Across Scales. arXiv 2025, arXiv:2506.05068. [Google Scholar]
- Wasi, A.T.; Ridoy, S.Z.; Tonmoy, K.A.; Tshering, K.; Hasan, S.; Faisal, W.; Mohiuddin, T.; Parvez, M.R. TimeSpot: Benchmarking Geo-Temporal Understanding in Vision–Language Models in Real-World Settings. In Proceedings of the Forty-third International Conference on Machine Learning, 2026. [Google Scholar]
- Zhang, o. GTPred: Benchmarking MLLMs for Interpretable Geo-localization and Time-of-capture Prediction. arXiv 2026, arXiv:2601.00000. [Google Scholar]
- Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q.V.; Zhou, D.; et al. Chain-of-thought prompting elicits reasoning in large language models. Adv. Neural Inf. Process. Syst. 2022, 35, 24824–24837. [Google Scholar] [CrossRef]
- Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.t.; Rocktäschel, T.; et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Adv. Neural Inf. Process. Syst. 2020, 33, 9459–9474. [Google Scholar]
- Gao, Y.; Xiong, Y.; Gao, X.; Jia, K.; Pan, J.; Bi, Y.; Dai, Y.; Sun, J.; Wang, H.; Wang, H. Retrieval-augmented generation for large language models: A survey. arXiv 2023, arXiv:2312.109972. [Google Scholar]
- Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C.D.; Ermon, S.; Finn, C. Direct preference optimization: Your language model is secretly a reward model. Adv. Neural Inf. Process. Syst. 2023, 36, 53728–53741. [Google Scholar] [CrossRef]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar]
- Miao, J.; Jiang, K.; Wen, T.; Wang, Y.; Jia, P.; Wijaya, B.; Zhao, X.; Cheng, Q.; Xiao, Z.; Huang, J.; et al. A survey on monocular re-localization: From the perspective of scene map representation. IEEE Transactions on Intelligent Vehicles, 2024. [Google Scholar]
- Zhang, F.; Salazar-Miranda, A.; Duarte, F.; Vale, L.; Hack, G.; Chen, M.; Liu, Y.; Batty, M.; Ratti, C. Urban visual intelligence: Studying cities with artificial intelligence and street-level imagery. Ann. Am. Assoc. Geogr. 2024, 114, 876–897. [Google Scholar] [CrossRef]
Figure 1.
An illustration of Visual Geolocalization. Given a query image, the goal is to estimate its geographic location. Existing approaches can be broadly categorized into retrieval-based, prediction-based, and reasoning-based paradigms.
Figure 1.
An illustration of Visual Geolocalization. Given a query image, the goal is to estimate its geographic location. Existing approaches can be broadly categorized into retrieval-based, prediction-based, and reasoning-based paradigms.

Figure 2.
Publication trend of Visual Geolocalization in the reviewed corpus from 2007 to 2026. The rapid increase in publications highlights the growing attention devoted to this research area.
Figure 2.
Publication trend of Visual Geolocalization in the reviewed corpus from 2007 to 2026. The rapid increase in publications highlights the growing attention devoted to this research area.

Figure 3.
Two types of dataset organization in Visual Geolocalization: structured retrieval datasets and geospatially annotated datasets.
Figure 3.
Two types of dataset organization in Visual Geolocalization: structured retrieval datasets and geospatially annotated datasets.

Figure 4.
A taxonomy of representative approaches for Visual Geolocalization.

Figure 5.
A typical pipeline of retrieval-based image geolocalization. A reference set of geo-tagged images is first encoded into feature embeddings. Given a query image, its embedding is matched against the reference database to retrieve visually similar examples. The associated geographic information of retrieved references is aggregated to infer the location of the query image.
Figure 5.
A typical pipeline of retrieval-based image geolocalization. A reference set of geo-tagged images is first encoded into feature embeddings. Given a query image, its embedding is matched against the reference database to retrieve visually similar examples. The associated geographic information of retrieved references is aggregated to infer the location of the query image.

Figure 6.
An example result of retrieval-based image geolocalization, adapted from Vo et al. [6]. The input image (left) is matched to its nearest neighbors (right). The predicted location is shown as a red star (*), and the ground truth as a green circle (o).
Figure 6.
An example result of retrieval-based image geolocalization, adapted from Vo et al. [6]. The input image (left) is matched to its nearest neighbors (right). The predicted location is shown as a red star (*), and the ground truth as a green circle (o).

Figure 7.
A typical pipeline of discrete location prediction for image geolocalization. The Earth’s surface is divided into predefined geographic regions, each treated as a distinct class. A neural network is trained to assign the input image to one of these regions, and the predicted region is mapped to a representative coordinate.
Figure 7.
A typical pipeline of discrete location prediction for image geolocalization. The Earth’s surface is divided into predefined geographic regions, each treated as a distinct class. A neural network is trained to assign the input image to one of these regions, and the predicted region is mapped to a representative coordinate.

Figure 8.
An illustration of classification-based image geolocalization, adapted from PlaNet [5]. Given a query image (left), the model outputs a probability distribution over geographic regions on the Earth’s surface (right).
Figure 8.
An illustration of classification-based image geolocalization, adapted from PlaNet [5]. Given a query image (left), the model outputs a probability distribution over geographic regions on the Earth’s surface (right).

Figure 9.
The generative framework introduced by Dufour et al. [64], which formulates geolocation as a denoising process on the spherical manifold to achieve continuous coordinate prediction.
Figure 9.
The generative framework introduced by Dufour et al. [64], which formulates geolocation as a denoising process on the spherical manifold to achieve continuous coordinate prediction.

Figure 10.
A typical pipeline of reasoning-based approaches for image geolocalization. Given an input image and a geolocation query, a pretrained LVLM performs multi-step reasoning to infer the most likely geographic location, producing both a location prediction and a natural language explanation.
Figure 10.
A typical pipeline of reasoning-based approaches for image geolocalization. Given an input image and a geolocation query, a pretrained LVLM performs multi-step reasoning to infer the most likely geographic location, producing both a location prediction and a natural language explanation.

Figure 11.
Examples of reasoning-based approaches for image geolocalization (adapted from GeoReasoner [18]). Green highlights indicate predictions that match the ground truth, while blue highlights denote reasoning components containing valid geographic clues.
Figure 11.
Examples of reasoning-based approaches for image geolocalization (adapted from GeoReasoner [18]). Green highlights indicate predictions that match the ground truth, while blue highlights denote reasoning components containing valid geographic clues.

Table 1.
Summary of representative Visual Geolocalization datasets across four categories.
| Type | Dataset | Scale | Data Characteristics | Venue | Year |
|---|---|---|---|---|---|
| Structured Retrieval Datasets | Same-view | ||||
| Oxford5k [85] | 5K | 11 Oxford landmarks from Flickr | CVPR | 2007 | |
| Paris6k [86] | 6K | 11 Paris landmarks from Flickr | CVPR | 2008 | |
| Holidays [87] | 1.5K | Personal holiday photos, diverse scenes | ECCV | 2008 | |
| European Cities 50k [88] | 51K | Flickr images from 14 European cities | MM | 2010 | |
| San Francisco Landmark [89] | 1.7M | Google Street View + mobile photos | CVPR | 2011 | |
| Pitts250k [90] | 250K | Google Street View, 10,586 viewpoints | CVPR | 2013 | |
| 24/7 Tokyo [27] | 76K | Daytime database, day/dusk/night queries | CVPR | 2015 | |
| Paris500k [91] | 501K | Geotagged Flickr images | CVIU | 2015 | |
| Pitts30k [92] | 30K | Subset of Pitts250k | CVPR | 2016 | |
| Landmarks-full [93] | 169K | 586 landmarks from image search | ECCV | 2016 | |
| Google Landmarks [94] | 1.1M | 4,872 cities across 187 countries | ICCV | 2017 | |
| Google Landmarks Dataset v2 [95] | 4.2M | 200K landmark instances, Wikimedia Commons images | CVPR | 2020 | |
| Cross-view | |||||
| Charleston & SF Dataset [96] | 24.8K | Street, satellite, and land-cover maps | CVPR | 2013 | |
| CVUSA [97] | 2.48M | Flickr ground + Bing satellite imagery | ICCV | 2015 | |
| UrbanGeo [98] | 18K | Google Maps imagery with building semantics | CVPR | 2017 | |
| CVACT [99] | 128K | Google Street View + aligned aerial imagery | CVPR | 2019 | |
| University-1652 [100] | 146K | Drone + satellite + street-view imagery | MM | 2020 | |
| VIGOR [101] | 48K | Ground + satellite imagery, one-to-many matching | CVPR | 2021 | |
| SeqGeo [102] | 157K | Ground-image sequences adapted from CVUSA/VIGOR | WACV | 2023 | |
| KITTI-CVL [103] | 31K | KITTI ground sequences + satellite imagery | ICCV | 2023 | |
| CVOGL [104] | 16K | Ground + UAV + satellite imagery with object annotations | TGRS | 2023 | |
| SUES-200 [105] | 13K | Multi-height UAV + satellite imagery | MM | 2023 | |
| GeoText-1652 [106] | 150K | Drone + satellite imagery with fine-grained text descriptions | ECCV | 2024 | |
| DReSS [107] | 96K | Panoramic street-view + VHR satellite imagery | ISPRS | 2025 | |
| CVG-Text [45] | 30K | Street-view imagery with human-written descriptions | ICCV | 2025 | |
| Geospatially Annotated Datasets | Image–GPS | ||||
| IM2GPS [25] | 237 | Geotagged Flickr images | CVPR | 2008 | |
| MP16 [108] | 4.72M | Geotagged Flickr images from YFCC100M | MM | 2016 | |
| IM2GPS3K [6] | 3K | Geotagged Flickr images, no overlap with IM2GPS | ICCV | 2017 | |
| YFCC4K [6] | 4K | Random subset of YFCC100M | ICCV | 2017 | |
| YFCC26K [56] | 26K | Random subset of YFCC100M | WACV | 2022 | |
| GWS15K [59] | 15K | Google Street View images | CVPR | 2023 | |
| MP16-Pro [70] | 4.12M | MP16 with hierarchical geographic labels | NeurIPS | 2024 | |
| OSV-5M [109] | 5.1M | Mapillary images with rich geographic metadata | CVPR | 2024 | |
| Pitts-IAL [50] | 253K | Pittsburgh street-view images with address labels | ECCV | 2024 | |
| SF-IAL-Base [50] | 205K | San Francisco street-view images with address labels | ECCV | 2024 | |
| CityGuessr68k [61] | 68K | YouTube videos with hierarchical geographic labels | ECCV | 2024 | |
| GeoGlobe [73] | 293K | Internet images of landmarks and natural scenes | KDD | 2025 | |
| GeoRanking [29] | 100K | MP16-Pro with image-text candidate locations | NeurIPS | 2025 | |
| MAPBench [76] | 5K | POI-centered street-view images | ACL Findings | 2026 | |
| Image–GPS–Text | |||||
| MP16-Reason [24] | 46K | MP16-Pro with reasoning trajectories | NeurIPS | 2025 | |
| NAVICLUES [23] | 1.1K | GeoGuessr expert gameplay with reasoning traces | arXiv | 2025 | |
| Geolocation Hub [110] | 50K | Google Street View with text descriptions | PoPETs | 2025 | |
| GREval-Bench [68] | 3K | Images with GPS coordinates and CoT reasoning | NeurIPS | 2025 | |
| EarthWhere [111] | 810 | Country- and street-level geolocation tasks | arXiv | 2025 | |
| GeoChain [66] | 1.46M | Mapillary street-level images with multimodal QA | EMNLP Findings | 2025 | |
| GAEA-Bench [21] | 3.5K | MP-16, GLDv2, and CityGuessr with QA pairs | WACV | 2026 | |
| TimeSpot [112] | 1.5K | Geo-temporal attributes across 80 countries | ICML | 2026 | |
| MG-Geo [22] | 5M | OSV-5M metadata + GeoGuessr clues + conversational data | ECCV | 2026 | |
| GTPred [113] | 370 | Geo-localization and time-of-capture labels | arXiv | 2026 | |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.