Preprint
Review

This version is not peer-reviewed.

Foundation Models and AI Agents in Geographic Science: A Review

Submitted:

25 August 2026

Posted:

26 August 2026

You are already at the latest version

Abstract
Large language models, multimodal foundation models, and agent systems are increasingly being integrated into geographic research by linking natural-language interaction with remote sensing, geospatial data, and specialized analytical tools. This review combines bibliometric analysis with qualitative synthesis to examine this convergence through a Perception--Reasoning--Action--Decision framework. A Web of Science search covering 2022--2026 year-to-date yielded a bibliometric corpus of 1,147 records. From this corpus, 151 representative studies were purposively selected for detailed narrative synthesis across four analytical stages: multimodal perception, geospatial reasoning, agentic action, and operational decision support. The review compares developments in cross-modal alignment, geographic cognition, spatial code generation, tool use, multi-agent collaboration, and applications in urban governance, transportation, disaster response, environmental monitoring, agriculture, energy, and satellite scheduling. Across these areas, recurring limitations concern precise spatial grounding, cross-sensor and cross-region generalization, hallucination, workflow verification, computational efficiency, real-time deployment, and responsible decision-making. The literature therefore points toward three priorities for future geospatial intelligence: explicit spatiotemporal grounding, verifiable tool-augmented workflows, and reliable integration of multimodal observations with professional geographic models and human expertise.
Keywords: 
;  ;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.