Submitted:
24 August 2026
Posted:
26 August 2026
You are already at the latest version
Abstract
With the rapid growth of remote sensing data and multimodal information, foundation models for Earth observation and vision–language geospatial reasoning models have become key drivers of intelligent remote sensing analysis. However, the specific roles and applications of reasoning models in remote sensing remain insufficiently explored. This survey defines RS-Reasoning operationally as task-dependent, multi-step inference whose conclusions are supported by traceable visual, temporal, spatial, or tool-derived evidence and are evaluated beyond final-answer accuracy. We organize the literature into three non-exclusive paradigms according to the dominant mechanism by which reasoning is acquired and executed: supervised reasoning, reinforcement learning–driven reasoning, and agentic or tool-augmented reasoning. We further distinguish reasoning-specific methods from vision–language models that provide semantic alignment, generation, or grounding foundations but do not themselves demonstrate multi-step inference. Across urban analysis, disaster assessment, environmental monitoring, and spatiotemporal question answering, we identify the conditions under which reasoning may add value while separating benchmark capability from validated operational deployment. Our synthesis exposes four persistent limitations: frequent use of synthetic supervision with unevenly reported provenance, outcome-dominant evaluation, weak verification of cross-modal evidence, and fragile long-horizon tool use. We therefore outline a research agenda centered on geography-aware process supervision, multimodal constraint checking, calibrated uncertainty, expert feedback, and robustness tests across sensors, regions, and time. For convenient reference and further research, we maintain a curated collection of related resources at https://github.com/ML4Sustain/Awesome-RS-Reasoning-Models.
Keywords:
remote sensing
; earth observation
; vision–language models
; multimodal reasoning
; geospatial intelligence
; reinforcement learning
; tool-augmented agents
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.